Appearance
Serverless — Amazon DynamoDB — Theory (Bản gốc slide / Original slide)
1. Amazon DynamoDB — Tổng quan (Overview)
- Fully managed, tính sẵn sàng cao với replication qua nhiều AZ
- Là NoSQL database — không phải relational database — nhưng có hỗ trợ transaction
- Scale tới workload cực lớn, là database phân tán (distributed)
- Hàng triệu request/giây, hàng nghìn tỷ (trillion) row, hàng trăm TB storage
- Hiệu năng nhanh và ổn định (độ trễ single-digit millisecond)
- Tích hợp IAM cho bảo mật, phân quyền và quản trị
- Chi phí thấp và có khả năng auto-scaling
- Không cần bảo trì hay vá lỗi, luôn sẵn sàng
- Có Standard & Infrequent Access (IA) Table Class
- Fully managed, highly available with replication across multiple AZs
- A NoSQL database — not a relational database — but with transaction support
- Scales to massive workloads, a distributed database
- Millions of requests per second, trillions of rows, 100s of TB of storage
- Fast and consistent performance (single-digit millisecond latency)
- Integrated with IAM for security, authorization and administration
- Low cost and auto-scaling capabilities
- No maintenance or patching, always available
- Standard & Infrequent Access (IA) Table Class
2. DynamoDB — Basics
- DynamoDB được cấu thành từ các Table
- Mỗi table có một Primary Key (phải quyết định ngay lúc tạo)
- Mỗi table có thể chứa số lượng item (= row) không giới hạn
- Mỗi item có các attribute (có thể thêm theo thời gian — có thể null)
- Kích thước tối đa của một item: 400KB
- Data type hỗ trợ:
- Scalar Types — String, Number, Binary, Boolean, Null
- Document Types — List, Map
- Set Types — String Set, Number Set, Binary Set
- Nhờ vậy, DynamoDB cho phép thay đổi schema nhanh chóng
- DynamoDB is made of Tables
- Each table has a Primary Key (must be decided at creation time)
- Each table can have an infinite number of items (= rows)
- Each item has attributes (can be added over time — can be null)
- Maximum size of an item is 400KB
- Data types supported:
- Scalar Types — String, Number, Binary, Boolean, Null
- Document Types — List, Map
- Set Types — String Set, Number Set, Binary Set
- Therefore, in DynamoDB you can rapidly evolve schemas
3. DynamoDB — Ví dụ Table
Ví dụ table lưu kết quả game, với Partition Key = User_ID, Sort Key = Game_ID (hợp thành Primary Key), và các attribute Score, Result:
| User_ID (Partition Key) | Game_ID (Sort Key) | Score | Result |
|---|---|---|---|
7791a3d6-… | 4421 | 92 | Win |
873e0634-… | 1894 | 14 | Lose |
873e0634-… | 4521 | 77 | Win |
Example table storing game results, with Partition Key = User_ID, Sort Key = Game_ID (together forming the Primary Key), and attributes Score, Result:
| User_ID (Partition Key) | Game_ID (Sort Key) | Score | Result |
|---|---|---|---|
7791a3d6-… | 4421 | 92 | Win |
873e0634-… | 1894 | 14 | Lose |
873e0634-… | 4521 | 77 | Win |
4. DynamoDB — Read/Write Capacity Modes
Kiểm soát cách quản lý capacity (throughput đọc/ghi) của table:
- Provisioned Mode (mặc định)
- Bạn tự chỉ định số read/write mỗi giây
- Cần lên kế hoạch capacity từ trước
- Trả tiền theo Read Capacity Unit (RCU) & Write Capacity Unit (WCU) đã cấp phát
- Có thể bật auto-scaling cho RCU & WCU
- On-Demand Mode
- Read/write tự động scale lên/xuống theo workload
- Không cần lên kế hoạch capacity
- Trả tiền theo mức dùng thực tế, đắt hơn ($$$)
- Rất tốt cho workload khó dự đoán, spike đột ngột
Control how you manage your table's capacity (read/write throughput):
- Provisioned Mode (default)
- You specify the number of reads/writes per second
- You need to plan capacity beforehand
- Pay for provisioned Read Capacity Units (RCU) & Write Capacity Units (WCU)
- Possibility to add auto-scaling mode for RCU & WCU
- On-Demand Mode
- Reads/writes automatically scale up/down with your workloads
- No capacity planning needed
- Pay for what you use, more expensive ($$$)
- Great for unpredictable workloads, steep sudden spikes
5. DynamoDB Accelerator (DAX)
- In-memory cache cho DynamoDB, fully-managed, highly available, seamless
- Giúp giải quyết read congestion bằng cách cache dữ liệu
- Độ trễ mili-giây (microseconds) cho dữ liệu đã cache
- Không cần sửa đổi logic ứng dụng (tương thích với API DynamoDB hiện có)
- TTL cache mặc định 5 phút
DAX vs. ElastiCache:
- DAX: cache item riêng lẻ và kết quả Query & Scan — tích hợp sẵn, chuyên biệt cho DynamoDB
- ElastiCache: lưu kết quả tổng hợp (aggregation) do ứng dụng tự tính toán — linh hoạt hơn nhưng cần code tự quản lý cache logic
- In-memory cache for DynamoDB, fully-managed, highly available, seamless
- Helps solve read congestion by caching data
- Microseconds latency for cached data
- Doesn't require application logic modification (compatible with existing DynamoDB APIs)
- 5 minutes TTL for cache (default)
DAX vs. ElastiCache:
- DAX: caches individual objects and Query & Scan results — built-in, specialized for DynamoDB
- ElastiCache: stores aggregation results computed by the application — more flexible but requires your own cache logic
6. DynamoDB — Stream Processing
- Là stream có thứ tự của các thay đổi ở cấp item (create/update/delete) trong table
- Use cases: phản ứng real-time với thay đổi (email chào mừng user), phân tích usage real-time, insert vào table dẫn xuất, cross-region replication, invoke Lambda khi table thay đổi
DynamoDB Streams vs. Kinesis Data Streams (mới hơn):
| DynamoDB Streams | Kinesis Data Streams |
|---|---|
| Retention 24 giờ | Retention 1 năm |
| Giới hạn số lượng consumer | Nhiều consumer |
| Xử lý bằng Lambda Trigger, hoặc DynamoDB Stream Kinesis adapter | Xử lý bằng Lambda, Kinesis Data Analytics, Kinesis Data Firehose, AWS Glue Streaming ETL… |
Kiến trúc điển hình: Table → DynamoDB Streams → xử lý bằng Lambda (messaging/notification qua SNS, filter/transform ghi vào table khác) hoặc qua Kinesis Data Streams → Kinesis Data Firehose (phân tích đổ vào Redshift, archive vào S3, index vào OpenSearch)
- An ordered stream of item-level modifications (create/update/delete) in a table
- Use cases: react to changes in real-time (welcome email to users), real-time usage analytics, insert into derivative tables, implement cross-region replication, invoke AWS Lambda on table changes
DynamoDB Streams vs. Kinesis Data Streams (newer):
| DynamoDB Streams | Kinesis Data Streams |
|---|---|
| 24 hours retention | 1 year retention |
| Limited # of consumers | High # of consumers |
| Processed via AWS Lambda Triggers, or DynamoDB Stream Kinesis adapter | Processed via AWS Lambda, Kinesis Data Analytics, Kinesis Data Firehose, AWS Glue Streaming ETL… |
Typical architecture: Table → DynamoDB Streams → processed via Lambda (messaging/notification via SNS, filter/transform into another table) or via Kinesis Data Streams → Kinesis Data Firehose (analytics into Redshift, archiving into S3, indexing into OpenSearch)
7. DynamoDB Global Tables
- Giúp một DynamoDB table truy cập được với độ trễ thấp ở nhiều region
- Replication kiểu Active-Active
- Ứng dụng có thể ĐỌC và GHI vào table ở bất kỳ region nào
- Phải bật DynamoDB Streams như một điều kiện tiên quyết
- Make a DynamoDB table accessible with low latency in multiple regions
- Active-Active replication
- Applications can READ and WRITE to the table in any region
- Must enable DynamoDB Streams as a pre-requisite
8. DynamoDB — Time To Live (TTL)
- Tự động xoá item sau một mốc thời gian hết hạn (expiry timestamp)
- Use cases: giảm dữ liệu lưu trữ bằng cách chỉ giữ item hiện hành, tuân thủ quy định, xử lý web session…
- Cơ chế: một tiến trình quét & đánh dấu hết hạn (expire) các item quá hạn, sau đó một tiến trình khác quét & xoá chúng
- Automatically delete items after an expiry timestamp
- Use cases: reduce stored data by keeping only current items, adhere to regulatory obligations, web session handling…
- Mechanism: an expiration process scans & marks expired items, then a deletion process scans & deletes them
9. DynamoDB — Backups & Tích hợp với S3
Backups cho Disaster Recovery:
- Continuous backups dùng Point-in-Time Recovery (PITR)
- Tuỳ chọn bật, giữ trong 35 ngày gần nhất
- Khôi phục về bất kỳ thời điểm nào trong khoảng backup
- Quá trình khôi phục tạo ra một table mới
- On-demand backups
- Full backup cho lưu trữ dài hạn, giữ tới khi bị xoá tường minh
- Không ảnh hưởng hiệu năng/latency
- Có thể cấu hình & quản lý trong AWS Backup (hỗ trợ copy cross-region)
- Quá trình khôi phục tạo ra một table mới
Tích hợp với Amazon S3:
- Export sang S3 (phải bật PITR)
- Hoạt động với bất kỳ thời điểm nào trong 35 ngày gần nhất
- Không ảnh hưởng read capacity của table
- Dùng để phân tích dữ liệu, giữ snapshot cho audit, ETL trên S3 trước khi import ngược lại DynamoDB
- Export dạng DynamoDB JSON hoặc ION
- Import từ S3
- Import file CSV, DynamoDB JSON, hoặc ION
- Không tốn write capacity
- Tạo ra một table mới
- Lỗi import được log trong CloudWatch Logs
Backups for disaster recovery:
- Continuous backups using Point-in-Time Recovery (PITR)
- Optionally enabled, for the last 35 days
- Recovery to any time within the backup window
- The recovery process creates a new table
- On-demand backups
- Full backups for long-term retention, until explicitly deleted
- Doesn't affect performance or latency
- Can be configured/managed in AWS Backup (enables cross-region copy)
- The recovery process creates a new table
Integration with Amazon S3:
- Export to S3 (must enable PITR)
- Works for any point in time in the last 35 days
- Doesn't affect the read capacity of your table
- Perform data analysis on top of DynamoDB, retain snapshots for auditing, ETL on top of S3 data before importing back into DynamoDB
- Export in DynamoDB JSON or ION format
- Import from S3
- Import CSV, DynamoDB JSON, or ION format
- Doesn't consume any write capacity
- Creates a new table
- Import errors are logged in CloudWatch Logs
10. Ví dụ: Xây dựng một Serverless API
Kiến trúc serverless API chuẩn: Client → API Gateway (REST API) → Lambda (proxy request) → DynamoDB (CRUD). Đây chính là bước đệm dẫn vào phần AWS API Gateway tiếp theo.
Standard serverless API architecture: Client → API Gateway (REST API) → Lambda (proxy request) → DynamoDB (CRUD). This is the lead-in to the next AWS API Gateway part.