Appearance
Databases in AWS — Theory (Bản gốc slide / Original slide)
1. Chọn đúng loại Database (Choosing the Right Database)
- AWS có rất nhiều managed database để lựa chọn
- Các câu hỏi cần trả lời để chọn đúng database cho kiến trúc của bạn:
- Workload read-heavy, write-heavy hay balanced? Nhu cầu throughput? Có thay đổi theo thời gian trong ngày, cần scale hay fluctuate không?
- Lưu bao nhiêu dữ liệu, trong bao lâu? Dữ liệu có tăng trưởng không? Kích thước object trung bình? Được truy cập như thế nào?
- Data durability? Đâu là source of truth cho dữ liệu?
- Yêu cầu về latency? Số concurrent user?
- Data model? Query dữ liệu như thế nào? Có join không? Structured hay Semi-Structured?
- Cần schema chặt chẽ hay linh hoạt hơn? Cần reporting? Search? RDBMS hay NoSQL?
- Chi phí license? Có nên chuyển sang Cloud Native DB như Aurora?
- AWS has a lot of managed databases to choose from
- Questions to answer to choose the right database for your architecture:
- Read-heavy, write-heavy, or balanced workload? Throughput needs? Will it change, does it need to scale or fluctuate during the day?
- How much data to store and for how long? Will it grow? Average object size? How are they accessed?
- Data durability? Source of truth for the data?
- Latency requirements? Concurrent users?
- Data model? How will you query the data? Joins? Structured? Semi-Structured?
- Strong schema? More flexibility? Reporting? Search? RDBMS / NoSQL?
- License costs? Switch to a Cloud Native DB such as Aurora?
2. Các loại Database trên AWS (Database Types)
- RDBMS (= SQL / OLTP): RDS, Aurora — rất tốt cho join
- NoSQL database — không join, không SQL: DynamoDB (~JSON), ElastiCache (key/value pairs), Neptune (graph), DocumentDB (cho MongoDB), Keyspaces (cho Apache Cassandra)
- Object Store: S3 (cho object lớn) / Glacier (cho backup/archive)
- Data Warehouse (= SQL Analytics / BI): Redshift (OLAP), Athena, EMR
- Search: OpenSearch (JSON) — full-text, unstructured search
- Graph: Amazon Neptune — hiển thị mối quan hệ giữa các dữ liệu
- Ledger: Amazon Quantum Ledger Database
- Time series: Amazon Timestream
Lưu ý: một số database trong danh sách trên sẽ được trình bày kỹ hơn ở phần Data & Analytics.
- RDBMS (= SQL / OLTP): RDS, Aurora — great for joins
- NoSQL database — no joins, no SQL: DynamoDB (~JSON), ElastiCache (key/value pairs), Neptune (graphs), DocumentDB (for MongoDB), Keyspaces (for Apache Cassandra)
- Object Store: S3 (for big objects) / Glacier (for backups/archives)
- Data Warehouse (= SQL Analytics / BI): Redshift (OLAP), Athena, EMR
- Search: OpenSearch (JSON) — free text, unstructured searches
- Graphs: Amazon Neptune — displays relationships between data
- Ledger: Amazon Quantum Ledger Database
- Time series: Amazon Timestream
Note: some databases in this list are discussed in more detail in the Data & Analytics section.
3. Amazon RDS — Tổng kết
- Managed PostgreSQL / MySQL / Oracle / SQL Server / DB2 / MariaDB / Custom
- Cấu hình sẵn RDS Instance Size và EBS Volume Type & Size
- Có khả năng auto-scaling cho storage
- Hỗ trợ Read Replicas và Multi-AZ
- Bảo mật qua IAM, Security Groups, KMS, SSL in transit
- Automated Backup với tính năng point-in-time restore (tối đa 35 ngày)
- Manual DB Snapshot cho phục hồi dài hạn
- Bảo trì được quản lý và lên lịch (có downtime)
- Hỗ trợ IAM Authentication, tích hợp với Secrets Manager
- RDS Custom cho phép truy cập và tuỳ biến sâu vào instance bên dưới (Oracle & SQL Server)
- Use case: lưu dữ liệu quan hệ (RDBMS/OLTP), thực hiện SQL query, transaction
- Managed PostgreSQL / MySQL / Oracle / SQL Server / DB2 / MariaDB / Custom
- Provisioned RDS Instance Size and EBS Volume Type & Size
- Auto-scaling capability for storage
- Support for Read Replicas and Multi-AZ
- Security through IAM, Security Groups, KMS, SSL in transit
- Automated Backup with point-in-time restore feature (up to 35 days)
- Manual DB Snapshot for longer-term recovery
- Managed and scheduled maintenance (with downtime)
- Support for IAM Authentication, integration with Secrets Manager
- RDS Custom for access to and customization of the underlying instance (Oracle & SQL Server)
- Use case: store relational datasets (RDBMS/OLTP), perform SQL queries, transactions
4. Amazon Aurora — Tổng kết
- API tương thích PostgreSQL / MySQL, tách biệt storage và compute
- Storage: dữ liệu lưu thành 6 bản sao trên 3 AZ — highly available, self-healing, auto-scaling
- Compute: Cluster gồm nhiều DB Instance trên nhiều AZ, auto-scaling Read Replicas
- Cluster: có custom endpoint riêng cho writer và reader DB instance
- Cùng các tính năng security / monitoring / maintenance như RDS
- Cần nắm các tuỳ chọn backup & restore của Aurora
- Aurora Serverless — cho workload không dự đoán được / gián đoạn, không cần capacity planning
- Aurora Global: tối đa 16 DB Read Instance ở mỗi region, storage replication dưới 1 giây
- Aurora Machine Learning: chạy ML bằng SageMaker & Comprehend ngay trên Aurora
- Aurora Database Cloning: tạo cluster mới từ cluster hiện có, nhanh hơn restore từ snapshot
- Use case: giống RDS, nhưng ít bảo trì hơn / linh hoạt hơn / hiệu năng cao hơn / nhiều tính năng hơn
- Compatible API for PostgreSQL / MySQL, separation of storage and compute
- Storage: data is stored in 6 replicas, across 3 AZs — highly available, self-healing, auto-scaling
- Compute: cluster of DB instances across multiple AZs, auto-scaling of Read Replicas
- Cluster: custom endpoints for writer and reader DB instances
- Same security / monitoring / maintenance features as RDS
- Know the backup & restore options for Aurora
- Aurora Serverless — for unpredictable/intermittent workloads, no capacity planning
- Aurora Global: up to 16 DB Read Instances in each region, < 1 second storage replication
- Aurora Machine Learning: perform ML using SageMaker & Comprehend on Aurora
- Aurora Database Cloning: new cluster from an existing one, faster than restoring a snapshot
- Use case: same as RDS, but with less maintenance / more flexibility / more performance / more features
5. Amazon ElastiCache — Tổng kết
- Managed Redis / Memcached (tương tự RDS, nhưng dành cho cache)
- In-memory data store, latency dưới mili-giây
- Chọn một ElastiCache instance type (ví dụ
cache.m6g.large) - Hỗ trợ Clustering (Redis), Multi-AZ, Read Replicas (sharding)
- Bảo mật qua IAM, Security Groups, KMS, Redis Auth
- Backup / Snapshot / Point-in-time restore
- Bảo trì được quản lý và lên lịch
- Cần sửa code ứng dụng mới tận dụng được
- Use case: key/value store, đọc nhiều-ghi ít, cache kết quả DB query, lưu session data cho website, không dùng được SQL
- Managed Redis / Memcached (similar offering as RDS, but for caches)
- In-memory data store, sub-millisecond latency
- Select an ElastiCache instance type (e.g.,
cache.m6g.large) - Support for Clustering (Redis) and Multi-AZ, Read Replicas (sharding)
- Security through IAM, Security Groups, KMS, Redis Auth
- Backup / Snapshot / Point-in-time restore feature
- Managed and scheduled maintenance
- Requires some application code changes to be leveraged
- Use Case: key/value store, frequent reads, less writes, cache results for DB queries, store session data for websites, cannot use SQL
6. Amazon DynamoDB — Tổng kết
- Công nghệ độc quyền của AWS, managed serverless NoSQL database, latency mili-giây
- Capacity modes: provisioned capacity (kèm auto-scaling tuỳ chọn) hoặc on-demand capacity
- Có thể thay thế ElastiCache làm key/value store (ví dụ lưu session data, dùng tính năng TTL)
- Highly Available, mặc định Multi-AZ, read/write tách biệt (decoupled), hỗ trợ transaction
- DAX cluster cho read cache, latency đọc ở mức micro-giây
- Bảo mật, authentication và authorization đều thông qua IAM
- Event Processing: DynamoDB Streams để tích hợp với AWS Lambda, hoặc Kinesis Data Streams
- Global Table: mô hình active-active
- Automated backup tối đa 35 ngày với PITR (restore ra table mới), hoặc on-demand backup
- Export ra S3 mà không tốn RCU trong khung PITR, import từ S3 mà không tốn WCU
- Rất phù hợp để tiến hoá schema nhanh chóng
- Use case: phát triển ứng dụng serverless (document nhỏ, cỡ vài trăm KB), distributed serverless cache
- AWS proprietary technology, managed serverless NoSQL database, millisecond latency
- Capacity modes: provisioned capacity with optional auto-scaling, or on-demand capacity
- Can replace ElastiCache as a key/value store (storing session data for example, using the TTL feature)
- Highly Available, Multi-AZ by default, reads and writes are decoupled, transaction capability
- DAX cluster for read cache, microsecond read latency
- Security, authentication and authorization is done through IAM
- Event Processing: DynamoDB Streams to integrate with AWS Lambda, or Kinesis Data Streams
- Global Table feature: active-active setup
- Automated backups up to 35 days with PITR (restore to a new table), or on-demand backups
- Export to S3 without using RCU within the PITR window, import from S3 without using WCU
- Great to rapidly evolve schemas
- Use Case: serverless application development (small documents, 100s of KB), distributed serverless cache
7. Amazon S3 — Tổng kết
- S3 về bản chất là... key/value store cho object
- Rất tốt cho object lớn, không tối ưu cho nhiều object nhỏ
- Serverless, scale vô hạn, kích thước object tối đa 50 TB, hỗ trợ versioning
- Tiers: S3 Standard, S3 Infrequent Access, S3 Intelligent, S3 Glacier + lifecycle policy
- Tính năng: Versioning, Encryption, Replication, MFA-Delete, Access Logs…
- Bảo mật: IAM, Bucket Policies, ACL, Access Points, Object Lambda, CORS, Object/Vault Lock
- Mã hoá: SSE-S3, SSE-KMS, SSE-C, client-side, TLS in transit, default encryption
- Batch operation trên object bằng S3 Batch, liệt kê file bằng S3 Inventory
- Hiệu năng: Multi-part upload, S3 Transfer Acceleration, S3 Select
- Automation: S3 Event Notifications (SNS, SQS, Lambda, EventBridge)
- Use case: file tĩnh, key/value store cho file lớn, website hosting
- S3 is a... key/value store for objects
- Great for bigger objects, not so great for many small objects
- Serverless, scales infinitely, max object size is 50 TB, versioning capability
- Tiers: S3 Standard, S3 Infrequent Access, S3 Intelligent, S3 Glacier + lifecycle policy
- Features: Versioning, Encryption, Replication, MFA-Delete, Access Logs…
- Security: IAM, Bucket Policies, ACL, Access Points, Object Lambda, CORS, Object/Vault Lock
- Encryption: SSE-S3, SSE-KMS, SSE-C, client-side, TLS in transit, default encryption
- Batch operations on objects using S3 Batch, listing files using S3 Inventory
- Performance: Multi-part upload, S3 Transfer Acceleration, S3 Select
- Automation: S3 Event Notifications (SNS, SQS, Lambda, EventBridge)
- Use Cases: static files, key/value store for big files, website hosting
8. Amazon DocumentDB
- Aurora là bản triển khai "kiểu AWS" của PostgreSQL / MySQL…
- DocumentDB cũng tương tự vậy nhưng dành cho MongoDB (một NoSQL database)
- MongoDB dùng để lưu trữ, truy vấn và index dữ liệu JSON
- Có khái niệm triển khai (deployment concepts) tương tự Aurora
- Fully managed, highly available với replication trên 3 AZ
- Storage của DocumentDB tự động tăng theo từng bước 10GB
- Tự động scale để đáp ứng workload lên tới hàng triệu request/giây
- Aurora is an "AWS-implementation" of PostgreSQL / MySQL...
- DocumentDB is the same for MongoDB (which is a NoSQL database)
- MongoDB is used to store, query, and index JSON data
- Similar "deployment concepts" as Aurora
- Fully managed, highly available with replication across 3 AZs
- DocumentDB storage automatically grows in increments of 10GB
- Automatically scales to workloads with millions of requests per second
9. Amazon Neptune
- Fully managed graph database
- Một dataset dạng graph phổ biến là mạng xã hội:
- User có bạn bè (friends)
- Post có comment
- Comment có like từ user
- User chia sẻ và like post…
- Highly available trên 3 AZ, tối đa 15 read replica
- Xây dựng và chạy ứng dụng làm việc với dữ liệu kết nối chặt chẽ (highly connected) — tối ưu cho các query phức tạp, khó
- Có thể lưu tới hàng tỷ relationship và query graph với latency mili-giây
- Highly available với replication trên nhiều AZ
- Rất phù hợp cho knowledge graph (như Wikipedia), fraud detection, recommendation engine, social networking
- Fully managed graph database
- A popular graph dataset would be a social network:
- Users have friends
- Posts have comments
- Comments have likes from users
- Users share and like posts…
- Highly available across 3 AZs, with up to 15 read replicas
- Build and run applications working with highly connected datasets — optimized for these complex and hard queries
- Can store up to billions of relations and query the graph with millisecond latency
- Highly available with replications across multiple AZs
- Great for knowledge graphs (Wikipedia), fraud detection, recommendation engines, social networking
10. Amazon Neptune — Streams
- Chuỗi thay đổi có thứ tự, real-time cho toàn bộ graph data
- Thay đổi có sẵn ngay sau khi ghi
- Không trùng lặp, thứ tự nghiêm ngặt (strict order)
- Dữ liệu stream truy cập được qua HTTP REST API
- Use case:
- Gửi notification khi có thay đổi nhất định
- Giữ đồng bộ dữ liệu graph sang một data store khác (ví dụ S3, OpenSearch, ElastiCache)
- Replicate dữ liệu giữa các region trong Neptune
- Real-time ordered sequence of every change to your graph data
- Changes are available immediately after writing
- No duplicates, strict order
- Streams data is accessible in an HTTP REST API
- Use cases:
- Send notifications when certain changes are made
- Maintain your graph data synchronized in another data store (e.g., S3, OpenSearch, ElastiCache)
- Replicate data across regions in Neptune
11. Amazon Keyspaces (cho Apache Cassandra)
- Apache Cassandra là một NoSQL distributed database mã nguồn mở
- Amazon Keyspaces là managed service tương thích Apache Cassandra
- Serverless, Scalable, highly available, được AWS quản lý toàn phần
- Tự động scale up/down table theo traffic của ứng dụng
- Table được replicate 3 lần trên nhiều AZ
- Sử dụng Cassandra Query Language (CQL)
- Latency single-digit mili-giây ở mọi quy mô, hàng nghìn request/giây
- Capacity: On-demand mode hoặc provisioned mode kèm auto-scaling
- Encryption, backup, PITR (Point-In-Time Recovery) tối đa 35 ngày
- Use case: lưu thông tin thiết bị IoT, dữ liệu time-series…
- Apache Cassandra is an open-source NoSQL distributed database
- A managed Apache Cassandra-compatible database service
- Serverless, Scalable, highly available, fully managed by AWS
- Automatically scale tables up/down based on the application's traffic
- Tables are replicated 3 times across multiple AZs
- Using the Cassandra Query Language (CQL)
- Single-digit millisecond latency at any scale, 1000s of requests per second
- Capacity: On-demand mode or provisioned mode with auto-scaling
- Encryption, backup, Point-In-Time Recovery (PITR) up to 35 days
- Use cases: store IoT devices info, time-series data, ...
12. Amazon Timestream
- Fully managed, nhanh, scalable, serverless — database chuyên cho time series
- Tự động scale up/down để điều chỉnh capacity
- Lưu trữ và phân tích tới hàng nghìn tỷ (trillions) event mỗi ngày
- Nhanh hơn hàng nghìn lần & rẻ bằng 1/10 so với relational database
- Scheduled queries, multi-measure record, tương thích SQL
- Phân tầng lưu trữ dữ liệu: dữ liệu gần đây giữ trong memory, dữ liệu lịch sử lưu ở storage tối ưu chi phí
- Có sẵn hàm phân tích time series (giúp phát hiện pattern trong dữ liệu gần thời gian thực)
- Mã hoá in transit và at rest
- Use case: ứng dụng IoT, ứng dụng vận hành (operational), phân tích real-time…
- Fully managed, fast, scalable, serverless time series database
- Automatically scales up/down to adjust capacity
- Store and analyze trillions of events per day
- 1000s times faster & 1/10th the cost of relational databases
- Scheduled queries, multi-measure records, SQL compatibility
- Data storage tiering: recent data kept in memory and historical data kept in a cost-optimized storage
- Built-in time series analytics functions (helps you identify patterns in your data in near real-time)
- Encryption in transit and at rest
- Use cases: IoT apps, operational applications, real-time analytics, ...
13. Amazon Timestream — Kiến trúc (Architecture)
Dữ liệu time series đổ về Timestream từ nhiều nguồn: AWS IoT, Kinesis Data Streams (trực tiếp hoặc qua Lambda), Prometheus, Kinesis Data Streams + Kinesis Data Analytics for Apache Flink, Amazon MSK. Sau khi lưu trong Timestream, dữ liệu được truy vấn/trực quan hoá bởi Amazon QuickSight, Amazon SageMaker, hoặc bất kỳ kết nối JDBC nào.
Time series data flows into Timestream from multiple sources: AWS IoT, Kinesis Data Streams (directly or via Lambda), Prometheus, Kinesis Data Streams + Kinesis Data Analytics for Apache Flink, Amazon MSK. Once stored in Timestream, data is queried/visualized by Amazon QuickSight, Amazon SageMaker, or any JDBC connection.