This domain is 20% of the exam. All five weightings:
This domain is about getting dependable, machine-readable results from Claude — and knowing what prompting can and cannot guarantee.
Be explicit about criteria
"Be thorough" or "be conservative" are not instructions Claude can verify. State concrete criteria: what counts as a defect, which categories to report, what to ignore, what format to return. Vague review prompts produce inconsistent findings; precise ones can be tested.
Few-shot examples
Two to four well-chosen examples — including a borderline case and what the right handling is — usually do more than another paragraph of rules. Examples should cover the variety of inputs, not just the happy path.
Structured output with tool use and JSON schema
For reliable structure, define a tool whose input schema is the output you want and have Claude call it. The schema guarantees the shape; it does not guarantee the content is correct. Know the tool_choice options:
"auto"— Claude decides whether to call a tool or answer in text."any"— Claude must call some tool (use when you must get structure back).{"type":"tool","name":"..."}— Claude must call that specific tool.
Design schemas that do not force fabrication
If a field is required but the document does not contain the value, Claude may invent one. Make fields that may be absent nullable or optional, and add an "other"/"unclear" category with a detail field so unusual cases are captured rather than forced into the wrong bucket.
Validation-retry loops
Validate the extraction (for example with Pydantic), and on failure retry with the original document, the failed output and the specific validation error. Know what retry fixes (format and structural mistakes) and what it cannot fix (information that is simply missing from the source — retrying only pushes Claude to guess).
Batch strategy and review
- The Message Batches API is 50% cheaper but can take up to 24 hours and has no latency guarantee. It suits non-blocking, bulk work; use
custom_idto match results to requests. - Anything a person is waiting on (pre-merge checks, live support) needs the real-time API instead.
- Route low-confidence fields to human review rather than treating a whole document as pass/fail.
- For review quality, use a separate, independent Claude instance rather than asking the generator to critique itself, and consider multiple focused passes over one giant pass.
Build to learn it
Build the validation-retry loop: extract with tool_use and a JSON schema, validate, and on failure retry with the document, the failed extraction and the exact error. Feel which errors it fixes and which it cannot. See the Structured Data Extraction scenario.