This domain is 27% of the exam. All five weightings:
This is the biggest domain, so it repays the most build time. The exam tests whether you can design an agent system that keeps working when production gets messy — not whether you can name the pieces.
The agentic loop
An agent is a loop: send the conversation to Claude, check why it stopped, act on that, and repeat. If stop_reason is tool_use, run the requested tool(s), append the results and call Claude again. If it is end_turn, the agent has finished. Know the other stop reasons (max_tokens, stop_sequence) and what each means for your loop — a truncated response is not a finished one.
Coordinator and subagents
In a multi-agent design a coordinator breaks work into pieces and hands them to specialised subagents (research, analysis, synthesis, reporting). The coordinator is also where results are merged and where gaps are noticed. Questions often ask which design is right: one agent holding every tool is usually the wrong answer, because tool overload hurts selection accuracy and a single context mixes unrelated work.
Subagents start with a blank page
The most-tested idea in this domain: a subagent knows nothing you do not tell it. It has not seen the user's request, earlier findings or tool results. Everything it needs — the task, the documents, the output format, the quality bar — must be in the prompt you give it, every time. When an answer choice says the subagent will "inherit" or "remember" context, it is a distractor.
- Spawn independent subagents in parallel (several spawn calls in one coordinator turn) when their work does not depend on each other.
- With the Agent SDK, the coordinator's allowed tools must include the tool used to spawn subagents, and each subagent is defined with its own name, system prompt and restricted tool list.
Rules that cannot fail belong in code
Prompts guide; code enforces. "Always verify the customer's identity before issuing a refund" followed 95% of the time is a production incident waiting to happen. For money, compliance and hard business rules, use programmatic enforcement: hooks that intercept calls before they run (a pre-tool-use hook can block a refund until a verification step has succeeded) and hooks that inspect results before the model sees them. "Strengthen the system prompt" is almost never the best answer in these questions.
Escalation and error propagation
- Escalate on evidence, not on self-reported confidence. Claude's stated confidence is not calibrated. Escalate on explicit triggers: policy exceptions, customer requests for a human, repeated failure, or no progress.
- Propagate errors with context. When a subagent fails, return a structured error — what failed, what was attempted, any partial results — so the coordinator can retry, route around it or report the gap. Silently dropping the failure, or killing the whole run on one failed branch, are both usually wrong.
- Report coverage gaps. A research synthesis should say which sources could not be read instead of presenting partial coverage as complete.
Sessions
Know when to resume a session, when to fork one to explore an alternative, and when a fresh, independent session is better — for example, an independent reviewer should not share the generator's session, because it would be biased toward the original reasoning.
Build to learn it
Build a coordinator that spawns two subagents with deliberately different context and watch the outputs diverge; then add a hook that blocks an action until a precondition is met. See the study plan.