The failure mode of a medical assistant is not the answer that is obviously wrong. It is the answer that is fluent, specific, and unsourced — the one you cannot check without doing the work yourself, which is the work you were trying to avoid.
Receipts, not footnotes
When an answer consults an outside source, the result distils into a citation the reader can follow: a drug-safety statement links to the specific US FDA label revision on DailyMed, a pharmacogenomic claim to the CPIC guideline behind it, and a passage retrieved from your own files back to the file it came from.
Only the citation metadata is kept — the retrieved payloads are not persisted. The client-side parser accepts only http and https URLs, so a citation that arrived poisoned cannot become a clickable script.
Graded on where the claim came from
The evaluation harness scores tool use as its own dimension: a drug-safety claim is marked down when it comes from the model's memory rather than from a live label lookup, even when the claim happens to be correct. Being right for the wrong reason is not a passing answer, because it does not repeat reliably.
What this deliberately does not do
- Receipts cover tool-backed claims — general explanation carries no citation and should not be read as sourced
- A cited label is the revision current at lookup time, not a permanent record
- Citations show where a claim came from, not that it applies to you