SDSystem Design Studio
Search guide and handbook titles, headings, and text
Complete book · 17. Agent-System Design Review Checklist

Your handbook progress

0 of 31 sections complete.

Loading saved progress…

Guided learning paths

Interview preparation

Practice a repeatable design flow and the trade-offs most often explored in interviews.

12 sections · 3–5 hours · 0/12 complete

Continue path
  1. 1. Practical system-design workflow· Not complete
  2. 2. The 12-question system design loop· Not complete
  3. 3. 1. Requirements: FRs, NFRs, Constraints, and Assumptions· Not complete
  4. 4. 2B. Data Modeling, Indexing, and Partitioning· Not complete
  5. 5. 3. Concurrency· Not complete
  6. 6. 4. Transactions and Consistency· Not complete
  7. 7. 5. APIs, Contracts, and Idempotency· Not complete
  8. 8. 6. Messaging and Asynchronous Work· Not complete
  9. 9. 7. Failure Handling and Resilience· Not complete
  10. 10. 8. Scale, Capacity, Performance, and Caching· Not complete
  11. 11. 13. Master System Design Review Checklist· Not complete
  12. 12. Design review outcome template· Not complete

Architecture review

Review an architecture systematically from boundaries through operability and evolution.

17 sections · 5–7 hours · 0/17 complete

Continue path
  1. 1. 1. Requirements: FRs, NFRs, Constraints, and Assumptions· Not complete
  2. 2. 2. Boundaries, State, and Data· Not complete
  3. 3. 2A. Networking and Communication· Not complete
  4. 4. 2B. Data Modeling, Indexing, and Partitioning· Not complete
  5. 5. 2C. Time, Clocks, and Ordering· Not complete
  6. 6. 4. Transactions and Consistency· Not complete
  7. 7. 5. APIs, Contracts, and Idempotency· Not complete
  8. 8. 6. Messaging and Asynchronous Work· Not complete
  9. 9. 7. Failure Handling and Resilience· Not complete
  10. 10. 8. Scale, Capacity, Performance, and Caching· Not complete
  11. 11. 9. Security· Not complete
  12. 12. 10. Observability and Reliability· Not complete
  13. 13. 11. Deployment, Migration, and Evolution· Not complete
  14. 14. 12. Cost, Simplicity, and Operability· Not complete
  15. 15. 13. Master System Design Review Checklist· Not complete
  16. 16. Architecture Decision Record — short template· Not complete
  17. 17. Design review outcome template· Not complete

Agentic systems

Design agent and LLM systems with explicit contracts, failure boundaries, and review gates.

9 sections · 3–4 hours · 0/9 complete

Continue path
  1. 1. 1. Requirements: FRs, NFRs, Constraints, and Assumptions· Not complete
  2. 2. 5. APIs, Contracts, and Idempotency· Not complete
  3. 3. 6. Messaging and Asynchronous Work· Not complete
  4. 4. 7. Failure Handling and Resilience· Not complete
  5. 5. 9. Security· Not complete
  6. 6. 10. Observability and Reliability· Not complete
  7. 7. 15. LLM and Agentic Systems· Not complete
  8. 8. 16. Spec-Driven Development for Agentic Systems· Not complete
  9. 9. 17. Agent-System Design Review Checklist· Not complete

Complete handbook · Section 26 of 31

17. Agent-System Design Review Checklist

Use this together with Chapter 13. Mark each item PASS, RISK, N/A, or DECISION REQUIRED.

17.1 Problem and autonomy

  • The system solves a named user/business problem with observable success.
  • The need for an LLM is supported by task characteristics or evaluation evidence.
  • The need for an agent rather than a fixed workflow is justified.
  • The need for multiple agents is justified by separability, specialization, or parallelism.
  • Non-goals and prohibited outcomes are explicit.
  • Read, propose, stage, write, communicate, spend, and delete authority are separately defined.
  • Completion, abstention, escalation, and cancellation conditions are testable.
  • Turn, token, time, cost, tool, delegation, and concurrency budgets are enforced outside the model.

17.2 Model, instructions, and outputs

  • Each approved model/provider/version is evaluated on the real task distribution.
  • Routing and fallback paths meet their own quality and safety thresholds.
  • Instruction priority and scope are explicit.
  • Untrusted content is never treated as authoritative instruction.
  • Output schemas are validated before downstream use.
  • Free text is escaped or constrained before entering interpreters or renderers.
  • Uncertainty, missing evidence, and source conflicts have defined behavior.
  • Model/prompt/spec versions are recorded in every trace.

17.3 Context, retrieval, and memory

  • Every context source has an owner, trust level, access rule, and freshness policy.
  • Retrieval enforces source authorization and tenant isolation.
  • Ingestion, parsing, chunking, embedding, reranking, and deletion are versioned.
  • Citations resolve to the exact source/version used for generation.
  • Retrieval quality and grounding are evaluated separately from answer fluency.
  • Context limits, truncation, compaction, and conflict handling are tested.
  • Run state, conversation state, working notes, user memory, knowledge, and audit history are distinct.
  • Durable memory has consent, purpose, retention, correction, deletion, confidence, and provenance.
  • Sensitive data is minimized in prompts, memory, caches, traces, and evaluation corpora.

17.4 Tools and actions

  • Every tool has a narrow purpose and machine-validatable input/output schema.
  • The tool validates identity, tenant, authorization, and business invariants at execution time.
  • Read and write capabilities are separated where practical.
  • Side effects are classified by impact and reversibility.
  • High-impact actions require contextual, argument-specific approval.
  • Idempotency and ambiguous-timeout behavior are specified.
  • Timeouts, retries, limits, error categories, and compensation are explicit.
  • Tool output is bounded, normalized, and treated as untrusted.
  • Generic shell, browser, filesystem, network, or admin access is sandboxed and allowlisted.
  • Secrets do not enter model context and credentials are short-lived and audience-restricted.
  • Every side effect has a durable audit record and verified final outcome.

17.5 Orchestration and interoperability

  • Deterministic code owns policy, permissions, budgets, and critical state transitions.
  • The chosen orchestration pattern matches the dependency graph.
  • Delegation includes scope, context, authority, budget, artifact, and return condition.
  • Shared mutable state has a single writer or conflict-resolution rule.
  • Parent/manager accountability survives delegation and handoff.
  • Loops and recursive delegation have hard termination limits.
  • Long-running work checkpoints safely and resumes without replaying effects.
  • Cancellation propagates to queued, running, and delegated work.
  • MCP, A2A, OpenAPI, JSON Schema, OAuth, and telemetry versions are pinned where used.
  • Remote tools and agents are treated as separate trust, identity, data, and reliability boundaries.
  • Emerging protocol features are not assumed merely because they appear on a roadmap.

17.6 Security, privacy, and human control

  • The threat model covers model, data, context, tools, remote agents, sandbox, telemetry, and humans.
  • Prompt injection, data exfiltration, privilege escalation, confused deputy, and excessive agency are tested.
  • Supply-chain risk covers models, SDKs, tools, skills, plugins, prompts, and retrieved sources.
  • Resource and spend exhaustion are rate-limited and observable.
  • Approval UI shows the exact action, target, consequence, and validated arguments.
  • Changed arguments invalidate prior approval.
  • Rejection, timeout, interruption, malformed pending state, and resume fail safely.
  • A kill switch, credential revocation, incident owner, and recovery procedure exist.
  • Privacy notices and user controls match actual storage, provider, training, and deletion behavior.
  • Audit/telemetry access is restricted and sensitive fields are redacted.

17.7 Evaluation and reliability

  • Evaluations reproduce realistic initial state, tools, permissions, and failures.
  • Final external state is graded rather than trusting the agent’s self-report.
  • Critical invariants use deterministic graders where possible.
  • Nuanced human/model graders have documented rubrics and calibration evidence.
  • Multiple trials expose variance and rare critical failures.
  • Development, held-out, adversarial, and production-regression sets are separated.
  • Unauthorized, unnecessary, duplicate, and unrequested actions are scored.
  • Recovery from model/tool timeout, partial failure, interruption, and resume is tested.
  • Quality, safety, latency, cost, and reliability thresholds are release gates.
  • Evaluation datasets and traces follow privacy, retention, and access rules.

17.8 Operations and evolution

  • Runs correlate model, tool, retrieval, approval, handoff, and business-outcome telemetry.
  • SLOs measure user outcomes and critical safety properties.
  • Alerts detect policy violations, runaway loops, provider degradation, tool failures, and cost anomalies.
  • Backpressure, quotas, fair scheduling, and dependency protection are implemented.
  • Production cost is measured per successful outcome.
  • Model, prompt, tool, retrieval, policy, and orchestration changes can be canaried or disabled.
  • Rollback behavior is tested, including state/schema compatibility.
  • Production failures feed the specification, threat model, and regression set.
  • Owners and review dates exist for provider assumptions and evolving specifications.

Evidence: S63, S64, S66, S67, S68, S70, S71, S72, S73, S74, S75, S76, S77, S81, S82, S83, S84, S89, S90, S92

Practice after learning

Section learning lab

Build the idea, test your recall, and keep page-specific notes.

The canvas is horizontally scrollable on narrow screens. Select a node and use arrow keys or the move controls; dragging also works.

Diagram ready.

Loading saved work…