
Articles · Research · 18 min
When should I not trust an AI research agent's brief?
Do not treat a research agent's brief as settled when the claim is legal, medical, or market, or when the trace shows one source talking to itself. Open the tool log. Map every claim to a primary document. Send those three classes of claim to a person who can own them. The brief is a draft until those checks pass.
By Eric · Rome · Aug 29, 2026
I still get briefs that look finished: four pages, nine citations, a calm voice. Then I open the tool log and find six of the nine URLs are the same post, rewritten by two other sites. The statute is a law firm's marketing page.
The agent did the job I gave it: it searched, wrote, and stopped. The finish line was a file in a folder, and the file was not evidence.
A research agent is allowed to draft. I am not allowed to ship the draft as a finding when the claim would bind a person, a patient, or a number someone will trade on. The rest of this page is the checks I run.
A research agent searches, reads, and writes a brief. Trust it when you can name the sources, see the tool calls, and still have a person who owns the claim. Legal, medical, and market statements fail that test until a primary document and a qualified owner sit on them.
When I refuse the brief
I refuse four classes of sentence as settled: legal, medical, market, and anything whose trace is one source talking to itself. The agent can still fetch, and a person still owns the claim. Valid for a reading list is not valid for advice, and intended use is the test.
I refuse the brief as a finding in those four cases, and I do not refuse the worker. The same loop can fetch a statute, a paper, and a 10-K. It cannot tell me those documents apply, that a dose is safe, or that a market will clear.
r/LocalLLaMA, 20 Dec 2024, is a thread about Anthropic's Building effective agents, posted the day after the essay shipped. It is not a legal-risk case study.
I already wrote what a research agent is: a job, a short list of tools, a brief that has to exist, a stop rule. This page is not that definition. This page is when the artifact is still a draft even though the file landed.
NIST's AI Risk Management Framework 1.0, published January 2023 as NIST AI 100-1, puts valid and reliable at the base of trustworthy AI. Validation, in that document, is confirmation through objective evidence that the requirements for a specific intended use have been fulfilled. A reading list and a legal memo are different intended uses.
Deployment of AI systems which are inaccurate, unreliable, or poorly generalized to data and settings beyond their training creates and increases negative AI risks and reduces trustworthiness.
| Claim class | The agent may | Do not treat as |
|---|---|---|
| Legal | Fetch the instrument, the date, the court or code | Advice that it applies to this matter |
| Medical | Fetch the paper, the label, the guideline | A diagnosis, a dose, a plan of care |
| Market | Fetch the filing line or the account | A forecast, a TAM, a they-will |
| Single-source loop | Show you the one page it found | Triangulation, confirmation, a finding |
Legal claims stay with counsel
A legal sentence in a brief is a retrieval, not representation. I want the instrument, the date, and the court or code, and if the citation never appears in a tool result it is not a source. Counsel owns the conclusion about what the text does to this matter.
A statute has a jurisdiction, an effective date, and a text, and a law-firm blog has a marketing calendar. If the tool log shows a search that landed on the blog, the brief has a recap, not the law. Recaps drop conditions, and conditions are the job.
I ask three questions of every legal sentence: which instrument, which date, which court or code. If the agent cannot point at the primary text, I do not let the sentence leave the draft. The claim that a practice is legal in Texas is a conclusion, and conclusions belong to a person with a license and a file.
The agent can fetch the code section and the last amendment. I still need counsel to say what it does to this contract, this employee, this filing.
Fetching is tool use, and applying is representation. I do not blur them because the prose is calm.
Case law is worse: headnotes, SEO pages, and case summaries compress holdings until they lie. I want the opinion, the court, the year, and the later history if we are going to cite it, and the brief should quote the opinion, not a blog about the opinion.
I also watch for invented citations, because models will emit a case name that looks right. The check is mechanical: does the tool result contain that citation, or did the model write it after the search, and OpenAI Agents SDK tracing records function tool calls as their own spans, with inputs and outputs. If the citation is only in the generation span and never in a function span, it is not a source.
Dates and forum are part of the instrument: an undated HTML page is not current law, and a federal opinion in one circuit is not the law of another. I want the as-of and the forum next to the citation, taken from the document, not from the model's prior.
Privilege is a runtime problem, not a prompt problem. If the research is on a live matter, the prompt and the documents may be confidential.
The Agents SDK tracing docs note that generation spans and function spans can store inputs and outputs, and that you can turn that capture off with RunConfig.trace_include_sensitive_data. I keep traces for audit. I do not dump a privileged memo into a dashboard I have not reviewed.
The brief is useful as a map of sections, opinions, and what I could not find. Not found is a valid outcome, invented found is not, and the map is not the advice. Counsel still signs.
Medical claims stay with a clinician
A medical sentence becomes a safety problem the moment someone might act on it. Papers, labels, and guidelines are different objects, so I want groundedness to the object, coverage of the contrary finding, and a clinician who owns the next step. The agent collects, and it does not treat.
NIST AI RMF 1.0 (2023) treats safety as a system that should not, under defined conditions, lead to a state in which human life, health, property, or the environment is endangered. Safety risks that can injure or kill get the most thorough process. A research brief about a drug, a device, or a diagnosis sits in that class the moment someone might act on it.
A preprint is not a label, a blog about a paper is not the paper, and a meta-analysis is not a treatment plan for this patient. I want the agent to fetch the object. I want a clinician to say what it means for this body, this history, this other drug.
Anthropic's Demystifying evals for AI agents, published 9 January 2026, is blunt about research quality: what counts as comprehensive, well-sourced, or even correct depends on context. Experts may disagree.
Ground truth shifts as reference content changes. Those are reasons to keep a human who can read the paper, not reasons to skip the agent.
Groundedness is the first check. Every clinical sentence in the brief has to point at a retrieved source, not at the model's prior.
If the sentence is not in the tool result, it does not ship. Coverage is the second: the contrary finding, the contraindication, the population the paper actually enrolled. A brief that reports a positive trial and skips the exclusion criteria is a coverage miss.
Anthropic's research evals define coverage as the key facts a good answer must include. I write those facts into the finish line before the run.
Source quality is the third: a journal article, a regulator's page, a specialty guideline. Not the first search result. Anthropic says source quality checks should confirm the consulted sources are authoritative, rather than simply the first retrieved.
I also watch retractions and errata, because a paper can sit in an index after it has been pulled. If the job is clinical, a second tool that checks retraction status is the intended use, not extra work. The brief should say when it last checked.
I do not let the agent dose or diagnose. I do let it collect the papers, extract the tables, and list what it could not find. If a user is asking the agent as if it were a doctor, I change the job to a packet for a clinician, with quotes and links, not an answer that sounds like care.
Market numbers stay with a filing
A number is not a market fact until a filing, an account, or an equivalent primary document holds the line. Newsletters and decks are loops with extra steps, and exact match works for reported figures. Forward-looking language stays labeled as a quote, dated, and off the finding list.
Market claims move money: revenue, share, TAM, they raised, the category will. Those sentences get treated as facts because they have digits in them, and digits are not facts until they have a document. I want a 10-K, a 10-Q, an 8-K, a prospectus, a Companies House account, or the equivalent in that jurisdiction.
A newsletter that cites a deck that cites a blog is a loop with extra steps: three URLs, one class of document.
I count classes, not footnotes. The finish line says so, or the agent will optimize for more links.
A reported quarterly figure is a claim with an exact-match test. Anthropic's evals post says exact match works for objectively correct answers of that shape.
The agent either retrieved the line or it did not. I need the line, the period, and the document type.
Forward-looking language is a different object. A sentence about what the market will do is not a number in a filing, so I will not let the brief state it as a finding.
The agent can quote an analyst. I still label it as a quote, with a date. Mixing guidance and actuals is how a brief becomes a pitch.
TAM is usually a vendor's story. I treat it as marketing unless I can see the method and the primary data.
Private-company figures are worse: a round announcement is not revenue, and a press page is not an audit. I would rather ship not found than a confident guess.
Dates matter here as much as they do in law, and last year's 10-K is not this quarter. Anthropic notes that ground truth for research shifts as reference content changes. I date every market sentence, and if the tool log cannot show the date of the source, the sentence stays in draft.
I also watch currency, fiscal year, and non-GAAP. Agents mix them: they will take a non-GAAP line and call it revenue, and they will convert currency on the wrong day. The check is the table in the filing versus the sentence in the brief.
A single source is a loop
One ranking page, cloned by two other sites, is still one source. I count document classes, not footnote counts, and I force a contrary search, because source quality means authoritative, not first retrieved. A run that cannot show a second class of document fails the finish line.
This is the failure I see most. The agent searches, a page ranks, and that page cites itself or two SEO clones cite each other. The brief has three footnotes and one fact: the prose is smooth, and the trace is a circle.
Anthropic's research-agent evals call this out as source quality: confirm the consulted sources are authoritative, rather than simply the first retrieved. First retrieved is what a search tool does unless you constrain it. I constrain it in the finish line, not in a paragraph of hopes at the top of the prompt.
I count distinct domains in the function spans, then I count distinct classes of document. A blog, a tweet, and a newsletter are three URLs and one class. A statute, a paper, and a 10-K are three classes.
I want classes of document. A pile of URLs is how a loop hides.
Same publisher, different paths, still one class. A press release republished forty times is still one release.
Wikipedia is an index. I will follow its citations, and I will not cite it as the finding. Mirror sites and syndicated copies do not add a class.
I force a contrary search. When the job is the rule, the next query is the exception, and when the job is they lead the category, the next query is who reports a different number. Coverage, in Anthropic's terms, is the key facts a good answer must include, and the missing contrary fact is a coverage miss.
Memory can become a loop: a previous brief sitting in the worker's memory is not a primary source. If the second run cites the first run, I have a hall of mirrors. I want the tool to go back to the object: the statute, the paper, the filing.
Shared memory is a pile, and the packet should hold evidence ids, not a vibe summary of last Tuesday.
How to build a research agent is where I put the tools and the stop rule. On this page I only care that the finish line names the classes.
At least two classes of primary document, or the run is a fail. That sentence does more than a page of style instructions.
Circular citation is easy to miss if you only read the prose: two tools, same domain, same paragraph, rewritten. That is a loop, and I throw the brief back.
I do not scold the model. I change the done condition and I log the miss into the eval set. A single-source loop is a missing check, not a moral failing of the system.
Read the trace, not the prose
The prose will always look calmer than the log. I read generations and function spans for the query, the result, and the quote, and I score whether the primary document exists, not whether the brief says confirmed. If tracing is off, the brief does not ship.
OpenAI Agents SDK tracing docs describe a record of the run: LLM generations, tool calls, handoffs, guardrails, custom events. Traces are the workflow, spans are the steps, and function spans carry the arguments and the results. That is the audit trail for a brief.
The Agents SDK includes built-in tracing, collecting a comprehensive record of events during an agent run: LLM generations, tool calls, handoffs, guardrails, and even custom events that occur.
Without the query I cannot trust the citation, and without the result I cannot tell whether the model quoted a page or invented a page. The dashboard is not a nicety. It is the only way the brief can be checked without redoing the job by hand.
Anthropic splits transcript from outcome. A flight agent can say the flight is booked while the reservation table is empty. A research agent can say confirmed while no filing exists.
I score the outcome: does the primary document exist, and does the sentence match it. The last paragraph of the brief is not the outcome.
I also run more than one trial, and Anthropic's pass^k is the probability that every trial succeeds. A brief I will send to counsel, a clinician, or a buyer has to survive more than one run. One pretty trace is a demo, and pass@k, the chance of at least one success in k tries, is the wrong headline when a person will read a single packet and act.
I read failures, because Anthropic's evals post says you will not know if graders work unless you read transcripts. When a brief fails, I want to know if the agent missed the filing or if my checker rejected a valid quote. That reading is the work, and Arena is the same job: same brief, competing traces, the artifact as the score.
If tracing is off, I do not ship. The SDK lets you disable tracing with an environment variable, a global flag, or a per-run config, which is useful for a local experiment and not useful for a brief someone else will read.
Organizations on a Zero Data Retention policy will not get this tracing. Then I need another log of tool calls that I own.
Sensitive research needs a decision: keep the spans, or redact them. The default in the SDK includes generation and function inputs and outputs.
Privileged legal files and clinical notes do not belong in a vendor dashboard by accident. I set the flag on purpose. I still keep a log I control: tool name, arguments, a hash of the result, timestamps.
How I score the brief
I score groundedness, coverage, source quality, and claim class. Legal, medical, and market sentences need a qualified owner on top of a citation, and I run more than one trial. I calibrate any model judge against a person who can read the primary document.
I score claims, not vibe. Anthropic's three research graders are the skeleton: groundedness, coverage, source quality. I add a fourth for this page: claim class.
A medical sentence cannot pass on groundedness alone. It needs a clinician on the edge, and the same rule holds for counsel and for a person who can read a filing.
- 01
Freeze a small set of briefs
Real jobs you already know, including ones the agent failed. Anthropic's evals post treats 20 to 50 tasks from real failures as a start. I write the finish line for each: the artifact, the classes of source, the owner for legal, medical, and market sentences.
- 02
Run more than one trial
Same job, several traces. I care about pass^k when a person will act on one packet. A single pretty run is a demo. I log every tool call: query, result, error, cost, steps.
- 03
Grade the claims against the objects
Grounded in a tool result. Coverage of the contrary fact. Source quality by class, not by footnote count. Exact match for a number that lives in a filing. Partial credit if the filing is in the packet and the sentence is wrong.
- 04
Keep a human on the three claim classes
Counsel, a clinician, or a filing reader sits on a sample. Anthropic says LLM rubrics for research should be calibrated against expert human judgment, often. If the judge and the expert diverge, the judge is wrong.
- 05
Promote only what holds
Capability tasks that the agent used to fail can graduate to a regression suite. I do not ship a new prompt because the prose improved. I ship when the artifact beats last week's baseline on the frozen set.
I do not grade the path, because Anthropic warns that checking a fixed sequence of tool calls is brittle, and agents find valid routes you did not draw. I grade whether the filing is in the packet and whether the sentence matches the line. A creative route that still cites the 10-K is a pass, and a neat search path that never fetches the 10-K is a zero.
I do use a model as a judge for partial credit: missing hedges, over-claim, tone that sounds like advice. I do not use it as the only score.
I wrote how to evaluate AI agents for that rule. A judge that grades writing will pass a brief that never fetched the object.
NIST's Measure function in AI RMF 1.0 (2023) asks that the system to be deployed is demonstrated valid and reliable, and that limitations of generalizability are documented. The limitation I document for a research agent is simple: it is not counsel, not a clinician, not a filing sign-off, and it is not to be trusted on a single-source loop.
What I still let the agent write
The agent can fetch, extract, list contradictions, and write a packet with hedges that match the sources. It can say not found, and I still read legal, medical, and market sentences before they leave the desk. Construction of the worker is a different page from this refusal list.
The agent can build the reading list: it can fetch the opinion, the paper, the 10-K, extract tables, list contradictions, and say it did not find a source. Not found is a clean outcome. Invented found is the failure mode this page is about.
I let it draft the brief with hedges that match the sources, and I do not let it drop the hedges to sound sure. Sure is a style, and style is not a source. A preprint stays labeled preprint, and guidance stays labeled guidance.
How to build a research agent is the construction side: tools, stop rules, the brief as an artifact, the grants. This page is the refusal side. Build the worker, then refuse to treat the output as settled in the four cases above.
Checks I run before anyone else sees it
Open the trace and map each claim to a tool result. Count document classes, date the sources, and run the contrary query. Hand legal, medical, and market sentences to a person who can own them, then freeze a small set of briefs and score the artifact, not the vibe.
- Tracing is on, and I can see generation spans and function spans for this run.
- Every citation appears in a tool result, not only in the prose.
- At least two classes of primary document, or the run is marked fail.
- Every legal, medical, and market sentence has a date taken from the object.
- A contrary query ran, and missing contrary facts are listed as not found or as open questions.
- Invented citations: none. If a name is not in a function span, it comes out.
- Claim class is tagged. Counsel, a clinician, or a filing reader has the sentences that belong to them.
- Sensitive inputs did not land in a dashboard I have not approved.
- The same job was run more than once if a person will act on a single packet.
- The brief is stored with the trace id. No orphan PDFs.
I keep failing briefs. They become tasks. Without that set you wait for complaints, you fix one miss, and you cannot tell noise from a real regression.
The construction path is how to build a research agent, and the scoring path is how to evaluate AI agents. The definition is what is a research agent. This page is only the refusal list, and the list is short on purpose.
Questions
Yes, as a fetcher. It can collect statutes, opinions, papers, labels, and guidelines, and it can write a packet with quotes and links. It cannot represent, diagnose, dose, or sign. Counsel or a clinician owns those sentences.
Count classes of document, not footnotes. A blog rewritten three times is one class. A statute plus an opinion, or a paper plus a label, or a 10-K plus an 8-K, is a start. One domain in the trace is a loop. Mark that run as fail.
No. A larger model still needs a tool result for every citation. OpenAI Agents SDK tracing is how you see generations and function calls. If the citation exists only in the prose, it is not a source, whatever the model size.
A filing, an account, a prospectus, or the equivalent in that jurisdiction. A newsletter, a deck, and a round-up blog are not the line. Exact match against the document. Date it. Keep forward-looking quotes labeled as quotes.
Next

