SDSystem Design Studio
Search guide and handbook titles, headings, and text
Complete book · 13. Master System Design Review Checklist

Your handbook progress

0 of 31 sections complete.

Loading saved progress…

Guided learning paths

Interview preparation

Practice a repeatable design flow and the trade-offs most often explored in interviews.

12 sections · 3–5 hours · 0/12 complete

Continue path
  1. 1. Practical system-design workflow· Not complete
  2. 2. The 12-question system design loop· Not complete
  3. 3. 1. Requirements: FRs, NFRs, Constraints, and Assumptions· Not complete
  4. 4. 2B. Data Modeling, Indexing, and Partitioning· Not complete
  5. 5. 3. Concurrency· Not complete
  6. 6. 4. Transactions and Consistency· Not complete
  7. 7. 5. APIs, Contracts, and Idempotency· Not complete
  8. 8. 6. Messaging and Asynchronous Work· Not complete
  9. 9. 7. Failure Handling and Resilience· Not complete
  10. 10. 8. Scale, Capacity, Performance, and Caching· Not complete
  11. 11. 13. Master System Design Review Checklist· Not complete
  12. 12. Design review outcome template· Not complete

Architecture review

Review an architecture systematically from boundaries through operability and evolution.

17 sections · 5–7 hours · 0/17 complete

Continue path
  1. 1. 1. Requirements: FRs, NFRs, Constraints, and Assumptions· Not complete
  2. 2. 2. Boundaries, State, and Data· Not complete
  3. 3. 2A. Networking and Communication· Not complete
  4. 4. 2B. Data Modeling, Indexing, and Partitioning· Not complete
  5. 5. 2C. Time, Clocks, and Ordering· Not complete
  6. 6. 4. Transactions and Consistency· Not complete
  7. 7. 5. APIs, Contracts, and Idempotency· Not complete
  8. 8. 6. Messaging and Asynchronous Work· Not complete
  9. 9. 7. Failure Handling and Resilience· Not complete
  10. 10. 8. Scale, Capacity, Performance, and Caching· Not complete
  11. 11. 9. Security· Not complete
  12. 12. 10. Observability and Reliability· Not complete
  13. 13. 11. Deployment, Migration, and Evolution· Not complete
  14. 14. 12. Cost, Simplicity, and Operability· Not complete
  15. 15. 13. Master System Design Review Checklist· Not complete
  16. 16. Architecture Decision Record — short template· Not complete
  17. 17. Design review outcome template· Not complete

Agentic systems

Design agent and LLM systems with explicit contracts, failure boundaries, and review gates.

9 sections · 3–4 hours · 0/9 complete

Continue path
  1. 1. 1. Requirements: FRs, NFRs, Constraints, and Assumptions· Not complete
  2. 2. 5. APIs, Contracts, and Idempotency· Not complete
  3. 3. 6. Messaging and Asynchronous Work· Not complete
  4. 4. 7. Failure Handling and Resilience· Not complete
  5. 5. 9. Security· Not complete
  6. 6. 10. Observability and Reliability· Not complete
  7. 7. 15. LLM and Agentic Systems· Not complete
  8. 8. 16. Spec-Driven Development for Agentic Systems· Not complete
  9. 9. 17. Agent-System Design Review Checklist· Not complete

Complete handbook · Section 22 of 31

13. Master System Design Review Checklist

Use this section before Technical Design approval, Architecture Review, or implementation sign-off. Items marked “if relevant” are not mandatory for every system.

13.1 Scope and requirements

  • Problem, users, scope, and out-of-scope are clear.

  • Fixed constraints and important assumptions are explicit.

  • Critical flows are named.

  • Traffic, peak, data volume, and growth assumptions are recorded.

  • Availability, latency, throughput, RPO, and RTO are measurable where relevant.

  • Security/compliance requirements are explicit.

  • Cost and delivery constraints are explicit.

13.2 Architecture and boundaries

  • Context/container diagrams are current.

  • Service responsibilities and ownership are clear.

  • Source of truth is known for important data.

  • Trust boundaries are shown.

  • Critical dependency chains are known.

  • Failure domains are intentional.

13.3 Data and concurrency

  • Business invariants are identified.

  • Concurrent update scenarios are handled.

  • Lost updates are prevented.

  • Time-zone, expiry, and business-ordering semantics are explicit where relevant.

  • Transaction boundaries are clear.

  • Isolation level and relevant anomalies are understood.

  • Locks/version checks/constraints are justified and scoped.

  • Core access patterns are known and the selected data model fits them.

  • Indexes are designed for critical queries and their write/storage cost is understood.

  • Sharding/partitioning is used only for a concrete scale/isolation reason.

  • Partition key avoids hotspots and supports dominant access patterns.

13.4 Distributed consistency

  • Cross-service consistency requirements are explicit.

  • DB + message dual writes are protected.

  • Eventual consistency is acceptable to the business where used.

  • Saga/compensation exists only where a multi-step distributed workflow needs it.

  • Retry of aborted transactions is safe.

  • Partition/multi-region behavior states what becomes unavailable, stale, or rejected when coordination is lost.

13.5 APIs

  • API contract is clear and consistent.

  • Idempotency/retry behavior is explicit.

  • Pagination exists for potentially large collections.

  • Timeout and error semantics are defined.

  • Versioning/backward compatibility strategy exists.

  • Authentication, authorization, rate limits, and telemetry are defined.

  • Protocol/connection style is justified: request-response, polling, SSE, WebSocket, or other.

  • Persistent connection reconnect/failover behavior is defined where relevant.

13.6 Messaging (if relevant)

  • Reason for async communication is explicit.

  • Delivery semantics are understood.

  • Consumers are idempotent.

  • Ordering requirement is explicit.

  • DLQ/poison-message handling exists.

  • Retry/backlog/replay behavior is bounded and observable.

  • Message schema evolution is defined.

  • Long-running jobs have status, cancellation, timeout, and worker-capacity rules.

  • Async processing is used because it solves a concrete duration/load/decoupling problem.

13.7 Failure and overload

  • Every remote call has a finite timeout.

  • Retries are selective, bounded, and use backoff.

  • Retry storms are prevented.

  • Circuit breaker/fail-fast is used where persistent failures would cause damage.

  • Pools/concurrency are bounded.

  • Overload strategy exists: throttle, queue, shed, or reject.

  • Optional features degrade without breaking critical paths where possible.

13.8 Scale and performance

  • Load model exists.

  • Bottlenecks and quotas are known.

  • Autoscaling uses meaningful signals.

  • Downstream capacity remains safe when upstream scales.

  • Cache strategy includes invalidation/TTL and staleness.

  • Load/stress testing is planned or completed.

  • Capacity estimates are tied to architectural decisions and checked against current service limits.

  • Hot partitions/keys and cross-partition operations are considered.

13.9 Security

  • Identity and access model is explicit.

  • Least privilege is applied.

  • Secrets management and rotation are defined.

  • Data in transit/at rest protection meets requirements.

  • Input validation exists at trust boundaries.

  • Security events are observable.

  • Threat/security review matches system risk.

13.10 Reliability and observability

  • SLIs/SLOs exist for critical flows.

  • Logs, metrics, and traces support diagnosis.

  • Alerting is actionable.

  • Health/readiness semantics are correct.

  • Backup/restore are defined and tested where needed.

  • Recovery design can meet RPO/RTO.

13.11 Deployment and change

  • CI/CD path is repeatable.

  • Rollout and rollback/roll-forward strategy exist.

  • DB/API/message changes work during mixed-version deployment.

  • Destructive schema changes are sequenced only after old application versions no longer depend on the old schema.

  • Canary/progressive deployment has measurable gates where appropriate.

  • Configuration and infrastructure are controlled and observable.

13.12 Cost and operations

  • Architecture complexity is justified.

  • Team can operate the selected technologies.

  • Ownership and escalation are clear.

  • Operational runbooks exist for critical failure modes.

  • Cost model includes resilience and observability overhead.

  • Trade-offs and revisit triggers are recorded.

13.13 Requirements, ADRs, and implementation plan

  • FRs are traceable to acceptance criteria, design, implementation, and tests.

  • Critical NFRs have measurable targets, conditions, verification methods, and production signals.

  • Architecturally significant decisions are captured in ADRs with alternatives and trade-offs.

  • Accepted ADRs are preserved; changed decisions are superseded rather than silently rewritten.

  • A TIP or equivalent implementation plan translates the approved design into ordered implementation, migration, testing, observability, and rollout work.

  • Production evidence feeds back into requirements, ADR assumptions, and future design work.

Evidence: S46, S47, S48, S49, S50, S52, S53

Evidence: S2, S24, S29, S30, S31, S32, S45, S53, S55, S56, S57

Practice after learning

Section learning lab

Build the idea, test your recall, and keep page-specific notes.

The canvas is horizontally scrollable on narrow screens. Select a node and use arrow keys or the move controls; dragging also works.

Diagram ready.

Loading saved work…