Skip to content

Serverless — Amazon DynamoDB — Theory (Bản gốc slide / Original slide)

1. Amazon DynamoDB — Tổng quan (Overview)

  • Fully managed, tính sẵn sàng cao với replication qua nhiều AZ
  • NoSQL databasekhông phải relational database — nhưng có hỗ trợ transaction
  • Scale tới workload cực lớn, là database phân tán (distributed)
  • Hàng triệu request/giây, hàng nghìn tỷ (trillion) row, hàng trăm TB storage
  • Hiệu năng nhanh và ổn định (độ trễ single-digit millisecond)
  • Tích hợp IAM cho bảo mật, phân quyền và quản trị
  • Chi phí thấp và có khả năng auto-scaling
  • Không cần bảo trì hay vá lỗi, luôn sẵn sàng
  • Standard & Infrequent Access (IA) Table Class
  • Fully managed, highly available with replication across multiple AZs
  • A NoSQL databasenot a relational database — but with transaction support
  • Scales to massive workloads, a distributed database
  • Millions of requests per second, trillions of rows, 100s of TB of storage
  • Fast and consistent performance (single-digit millisecond latency)
  • Integrated with IAM for security, authorization and administration
  • Low cost and auto-scaling capabilities
  • No maintenance or patching, always available
  • Standard & Infrequent Access (IA) Table Class

2. DynamoDB — Basics

  • DynamoDB được cấu thành từ các Table
  • Mỗi table có một Primary Key (phải quyết định ngay lúc tạo)
  • Mỗi table có thể chứa số lượng item (= row) không giới hạn
  • Mỗi item có các attribute (có thể thêm theo thời gian — có thể null)
  • Kích thước tối đa của một item: 400KB
  • Data type hỗ trợ:
    • Scalar Types — String, Number, Binary, Boolean, Null
    • Document Types — List, Map
    • Set Types — String Set, Number Set, Binary Set
  • Nhờ vậy, DynamoDB cho phép thay đổi schema nhanh chóng
  • DynamoDB is made of Tables
  • Each table has a Primary Key (must be decided at creation time)
  • Each table can have an infinite number of items (= rows)
  • Each item has attributes (can be added over time — can be null)
  • Maximum size of an item is 400KB
  • Data types supported:
    • Scalar Types — String, Number, Binary, Boolean, Null
    • Document Types — List, Map
    • Set Types — String Set, Number Set, Binary Set
  • Therefore, in DynamoDB you can rapidly evolve schemas

3. DynamoDB — Ví dụ Table

Ví dụ table lưu kết quả game, với Partition Key = User_ID, Sort Key = Game_ID (hợp thành Primary Key), và các attribute Score, Result:

User_ID (Partition Key)Game_ID (Sort Key)ScoreResult
7791a3d6-…442192Win
873e0634-…189414Lose
873e0634-…452177Win

Example table storing game results, with Partition Key = User_ID, Sort Key = Game_ID (together forming the Primary Key), and attributes Score, Result:

User_ID (Partition Key)Game_ID (Sort Key)ScoreResult
7791a3d6-…442192Win
873e0634-…189414Lose
873e0634-…452177Win

4. DynamoDB — Read/Write Capacity Modes

Kiểm soát cách quản lý capacity (throughput đọc/ghi) của table:

  • Provisioned Mode (mặc định)
    • Bạn tự chỉ định số read/write mỗi giây
    • Cần lên kế hoạch capacity từ trước
    • Trả tiền theo Read Capacity Unit (RCU) & Write Capacity Unit (WCU) đã cấp phát
    • Có thể bật auto-scaling cho RCU & WCU
  • On-Demand Mode
    • Read/write tự động scale lên/xuống theo workload
    • Không cần lên kế hoạch capacity
    • Trả tiền theo mức dùng thực tế, đắt hơn ($$$)
    • Rất tốt cho workload khó dự đoán, spike đột ngột

Control how you manage your table's capacity (read/write throughput):

  • Provisioned Mode (default)
    • You specify the number of reads/writes per second
    • You need to plan capacity beforehand
    • Pay for provisioned Read Capacity Units (RCU) & Write Capacity Units (WCU)
    • Possibility to add auto-scaling mode for RCU & WCU
  • On-Demand Mode
    • Reads/writes automatically scale up/down with your workloads
    • No capacity planning needed
    • Pay for what you use, more expensive ($$$)
    • Great for unpredictable workloads, steep sudden spikes

5. DynamoDB Accelerator (DAX)

ApplicationDAX ClusterNodes (in-memorycache)AmazonDynamoDB
  • In-memory cache cho DynamoDB, fully-managed, highly available, seamless
  • Giúp giải quyết read congestion bằng cách cache dữ liệu
  • Độ trễ mili-giây (microseconds) cho dữ liệu đã cache
  • Không cần sửa đổi logic ứng dụng (tương thích với API DynamoDB hiện có)
  • TTL cache mặc định 5 phút

DAX vs. ElastiCache:

  • DAX: cache item riêng lẻkết quả Query & Scan — tích hợp sẵn, chuyên biệt cho DynamoDB
  • ElastiCache: lưu kết quả tổng hợp (aggregation) do ứng dụng tự tính toán — linh hoạt hơn nhưng cần code tự quản lý cache logic
  • In-memory cache for DynamoDB, fully-managed, highly available, seamless
  • Helps solve read congestion by caching data
  • Microseconds latency for cached data
  • Doesn't require application logic modification (compatible with existing DynamoDB APIs)
  • 5 minutes TTL for cache (default)

DAX vs. ElastiCache:

  • DAX: caches individual objects and Query & Scan results — built-in, specialized for DynamoDB
  • ElastiCache: stores aggregation results computed by the application — more flexible but requires your own cache logic

6. DynamoDB — Stream Processing

ApplicationDDB TableDynamoDB StreamsLambdaKinesis Data StreamsKinesis FirehoseSNS (notify)DDB TableS3Redshift
  • stream có thứ tự của các thay đổi ở cấp item (create/update/delete) trong table
  • Use cases: phản ứng real-time với thay đổi (email chào mừng user), phân tích usage real-time, insert vào table dẫn xuất, cross-region replication, invoke Lambda khi table thay đổi

DynamoDB Streams vs. Kinesis Data Streams (mới hơn):

DynamoDB StreamsKinesis Data Streams
Retention 24 giờRetention 1 năm
Giới hạn số lượng consumerNhiều consumer
Xử lý bằng Lambda Trigger, hoặc DynamoDB Stream Kinesis adapterXử lý bằng Lambda, Kinesis Data Analytics, Kinesis Data Firehose, AWS Glue Streaming ETL

Kiến trúc điển hình: Table → DynamoDB Streams → xử lý bằng Lambda (messaging/notification qua SNS, filter/transform ghi vào table khác) hoặc qua Kinesis Data Streams → Kinesis Data Firehose (phân tích đổ vào Redshift, archive vào S3, index vào OpenSearch)

  • An ordered stream of item-level modifications (create/update/delete) in a table
  • Use cases: react to changes in real-time (welcome email to users), real-time usage analytics, insert into derivative tables, implement cross-region replication, invoke AWS Lambda on table changes

DynamoDB Streams vs. Kinesis Data Streams (newer):

DynamoDB StreamsKinesis Data Streams
24 hours retention1 year retention
Limited # of consumersHigh # of consumers
Processed via AWS Lambda Triggers, or DynamoDB Stream Kinesis adapterProcessed via AWS Lambda, Kinesis Data Analytics, Kinesis Data Firehose, AWS Glue Streaming ETL

Typical architecture: Table → DynamoDB Streams → processed via Lambda (messaging/notification via SNS, filter/transform into another table) or via Kinesis Data Streams → Kinesis Data Firehose (analytics into Redshift, archiving into S3, indexing into OpenSearch)

7. DynamoDB Global Tables

GLOBAL TABLETableUS-EAST-1TableAP-SOUTHEAST-2two-way replication
  • Giúp một DynamoDB table truy cập được với độ trễ thấp ở nhiều region
  • Replication kiểu Active-Active
  • Ứng dụng có thể ĐỌC và GHI vào table ở bất kỳ region nào
  • Phải bật DynamoDB Streams như một điều kiện tiên quyết
  • Make a DynamoDB table accessible with low latency in multiple regions
  • Active-Active replication
  • Applications can READ and WRITE to the table in any region
  • Must enable DynamoDB Streams as a pre-requisite

8. DynamoDB — Time To Live (TTL)

  • Tự động xoá item sau một mốc thời gian hết hạn (expiry timestamp)
  • Use cases: giảm dữ liệu lưu trữ bằng cách chỉ giữ item hiện hành, tuân thủ quy định, xử lý web session
  • Cơ chế: một tiến trình quét & đánh dấu hết hạn (expire) các item quá hạn, sau đó một tiến trình khác quét & xoá chúng
  • Automatically delete items after an expiry timestamp
  • Use cases: reduce stored data by keeping only current items, adhere to regulatory obligations, web session handling…
  • Mechanism: an expiration process scans & marks expired items, then a deletion process scans & deletes them

9. DynamoDB — Backups & Tích hợp với S3

Backups cho Disaster Recovery:

  • Continuous backups dùng Point-in-Time Recovery (PITR)
    • Tuỳ chọn bật, giữ trong 35 ngày gần nhất
    • Khôi phục về bất kỳ thời điểm nào trong khoảng backup
    • Quá trình khôi phục tạo ra một table mới
  • On-demand backups
    • Full backup cho lưu trữ dài hạn, giữ tới khi bị xoá tường minh
    • Không ảnh hưởng hiệu năng/latency
    • Có thể cấu hình & quản lý trong AWS Backup (hỗ trợ copy cross-region)
    • Quá trình khôi phục tạo ra một table mới

Tích hợp với Amazon S3:

  • Export sang S3 (phải bật PITR)
    • Hoạt động với bất kỳ thời điểm nào trong 35 ngày gần nhất
    • Không ảnh hưởng read capacity của table
    • Dùng để phân tích dữ liệu, giữ snapshot cho audit, ETL trên S3 trước khi import ngược lại DynamoDB
    • Export dạng DynamoDB JSON hoặc ION
  • Import từ S3
    • Import file CSV, DynamoDB JSON, hoặc ION
    • Không tốn write capacity
    • Tạo ra một table mới
    • Lỗi import được log trong CloudWatch Logs

Backups for disaster recovery:

  • Continuous backups using Point-in-Time Recovery (PITR)
    • Optionally enabled, for the last 35 days
    • Recovery to any time within the backup window
    • The recovery process creates a new table
  • On-demand backups
    • Full backups for long-term retention, until explicitly deleted
    • Doesn't affect performance or latency
    • Can be configured/managed in AWS Backup (enables cross-region copy)
    • The recovery process creates a new table

Integration with Amazon S3:

  • Export to S3 (must enable PITR)
    • Works for any point in time in the last 35 days
    • Doesn't affect the read capacity of your table
    • Perform data analysis on top of DynamoDB, retain snapshots for auditing, ETL on top of S3 data before importing back into DynamoDB
    • Export in DynamoDB JSON or ION format
  • Import from S3
    • Import CSV, DynamoDB JSON, or ION format
    • Doesn't consume any write capacity
    • Creates a new table
    • Import errors are logged in CloudWatch Logs

10. Ví dụ: Xây dựng một Serverless API

ClientAPI GatewayLambdaDynamoDBREST APIPROXY REQUESTSCRUD

Kiến trúc serverless API chuẩn: Client → API Gateway (REST API) → Lambda (proxy request) → DynamoDB (CRUD). Đây chính là bước đệm dẫn vào phần AWS API Gateway tiếp theo.

Standard serverless API architecture: Client → API Gateway (REST API) → Lambda (proxy request) → DynamoDB (CRUD). This is the lead-in to the next AWS API Gateway part.

Personal notes by thanhlt