claim check
Paste an AI-drafted document; every checkable claim comes back with a verdict and quoted evidence that code has proven exists on a page it fetched. Plausible is its job. Correct is yours.
team one internal tool, built as the leave-behind for an agency AI training session and deployed on the agency's internal aws/eks platform. usage is not measured; the numbers below come from the seeded test documented in the acceptance criteria.
Session 06 of the agency's AI Summer Sessions was about judgment: “Judgment in the Loop.” The anchor lines were plausible is its job, correct is yours and you sign the work, not the model. The demo was a memo with five errors planted in it, the kind a fluent draft carries without blinking: an inflated statistic, an article title that doesn't exist, a paraphrase presented as a direct quote, a regulation read too broadly, and a statistic with its meaning quietly swapped. Three true claims sat beside them as anchors.
A talk about verification needed something people could use the next morning. Claim Check was built as the leave-behind, and v0.1 shipped the day after the session.
The easy version asks a model to fact-check the document. That version fails the exact way the memo does: a model asked whether something is true will answer fluently from memory, and a model asked for sources will produce citations that look right. Asking it harder changes nothing.
The reframe was to keep the model away from both jobs it's bad at. It never testifies from memory, and it never vouches for its own sources. Retrievers suggest pages. The app fetches those pages itself. Code checks that every quoted passage is actually on the page it claims to come from. The judge model sees only passages that survived that check, and if none survive, there is no judge call at all: the claim is unverifiable, which the report treats as a red flag, not a pass.
Two retrievers, both demoted to leads. Every claim is researched through OpenAI web search and through Gemini with Google Search grounding, and the candidate URLs are deduplicated. A third-party search API was in the first build and was swapped out for Gemini grounding two days later when its credit model didn't hold up. Either way, nothing a retriever returns counts as evidence until the app has fetched the page itself.
Prove the quote. A passage counts only if it matches the fetched page text exactly, or near-verbatim (0.88 or better) after normalizing quotes, whitespace, and case. Passages that fail are dropped, and the drop count is shown to the user on purpose: it is the clearest signal of how much the retrievers' summaries drift from what the pages say.
Rules the model can't argue with. Four verdicts: verified, partly true, unverifiable, false. Verified requires at least one supporting passage; false requires at least one contradicting passage. The judge proposes, then code enforces those rules and records any downgrade. Source tiers (primary, established, other, weak) are domain lists in code, and evidence strength is computed from them. The model never grades its own sources.
One pipeline, run from the web UI or a terminal: documents (pasted text, .txt, .md, .docx, .pdf) → extract → search → fetch → evidence → judge → report. Every extracted claim carries a verbatim span located in the original document, so the report can highlight the text by verdict and jump from a highlight to its claim card. Each verdict cites its passages by index and, when the claim is wrong, states what the sources actually say. Exports come out as JSON, CSV, and a Markdown verification log in the same table shape the session taught.
A thorough mode adds an escalation pass for claims still unverifiable after the standard run: a high-effort research call whose citations go through the same fetch, prove, and judge path. Word, claim, and escalation caps keep the cost bounded.
The privacy contract is the product. People paste drafts they haven't shared yet, so the tool has to be trustworthy before it's useful. No user document or report is ever written to disk. Jobs live in memory behind random 128-bit ids, there is no listing or “recent” view, responses are marked no-store, and reports purge 60 minutes after they finish, or immediately on delete now. The landing page says plainly that claim text goes to the two model providers. The only thing that touches disk is a cache of public web pages.
Day one in production, a user got a 504 and then 404s. Page parsing and passage matching are CPU-heavy, and they were running on the web server's event loop. The loop blocked, the platform's liveness probe failed, the pod restarted, and every in-memory report went with it. The memory-only design meant the data loss was total, which is the price of the privacy promise and the reason the fix couldn't be “persist it.” Parsing and matching moved to worker threads the same day, and the rule went into the agent notes so no later change puts blocking work back on the loop.
On the session's seeded memo, the extractor surfaced all five planted claims plus the three true anchors, and the verdicts landed 8 for 8 in their expected bands: the planted errors flagged as false, partly true, or unverifiable, the anchors verified. That run took about three minutes, checked 11 claims, and cost $0.98. The deterministic parts (passage matching, source tiers, verdict enforcement, the privacy guarantees) have their own tests; every acceptance criterion is marked passing with a date.
What it doesn't claim: adoption. Usage isn't measured, and one seeded document is a demonstration, not an accuracy rate. The point it proves is narrower and more useful: a verification tool can be built so that every link it shows was fetched and every quote it shows is really on the page, and it can say “I couldn't verify this” instead of guessing.
A model's citation is a lead, not evidence. Make code prove the quoted passage is on a page you fetched yourself, let the judge see only what survived, and show the count of what didn't. The dropped-passage number is the honesty metric.
the insight