Loading
Loading
Your feedback directly shapes Sporos.
Sign in to track your feedback history
Trust & verification
Answers are built only from retrieved primary government records, and an independent audit pass checks the most load-bearing claims — with a deterministic test that each supporting quote exists verbatim in the evidence. Not the model’s word — the code’s. This is the part of the product a general counsel can defend.
The problem this solves
of queries to leading AI legal-research tools (Lexis+ AI, Westlaw) returned hallucinated content in an independent Stanford study.
Stanford RegLab ↗An incumbent policy publisher's AI report tools were shut down after a landmark arbitration found they shipped reports with “glaring errors” and no editorial review — the public cautionary tale for ungrounded policy AI.
Nieman Lab ↗“The limiting factor for AI adoption is no longer awareness or access — it is trust.” The buyer named the spec; this is built to it.
FiscalNote, 2026 State of Government Affairs ↗How it works
Every answer is built only from records retrieved out of primary government corpora — Congress.gov, the Federal Register, SEC EDGAR, the LDA lobbying database, the courts — joined through one canonical entity graph. The model writes over evidence; it does not write from memory.
Citations resolve through a registry that only holds sources actually retrieved from the database — a citation cannot be fabricated into existence. A reference number that doesn't resolve is caught deterministically and flagged on the answer, in the open. Most sources deep-link to the publisher's own page; where a corpus only supports an index-level link, the citation says so. One exception is labeled, never hidden: a corpus total (“457 filings on record”) cites a Corpus aggregate source — a measurement our engine computed over the entity graph at retrieval time, badged as such and dated, with no government page to open because it is our measurement, not a government record. Every individual record behind the count remains separately retrievable and citable.
A second, independent audit model samples the most load-bearing factual claims (up to 12 per output) and must paste the exact verbatim text that supports each. Then code — not a model's opinion — confirms that quoted text actually exists in the evidence. A claim that can't be grounded verbatim is flagged, even if the model believes it. The support judgment itself is the auditor's; the verbatim-existence check is deterministic.
Where coverage is thin, the answer says so. A name-matched record is labeled differently from an entity-graph-linked one; an empty search is never dressed up as a clean record. We surface the limit instead of papering over it — because a confidently-wrong answer is the worst possible failure for this buyer.
Anatomy of a verified claim
For each load-bearing claim, the auditor must paste the exact text that supports it — then code, not a model, confirms that text is really in the evidence. Here is the mechanism on two claims: one that grounds, one that doesn’t. (Illustrative of the pipeline.)
The claim
“The company reported $2.7M in federal lobbying spend in Q2 2025.”[4]
Verbatim span in evidence
…LDA filing, 2025 Q2: total reported lobbying expenses $2,700,000…
✓ Grounded
The pasted text is found verbatim in the retrieved record. Certified by a deterministic string match — not the model’s judgment.
The claim
“The drug was approved by the FDA in March 2024.”
Verbatim span in evidence
No matching text in the retrieved evidence — the date appears to come from the model’s training, not the sources.
Flagged · verify
The supporting text can’t be located in the evidence, so the claim is surfaced for review — even though the model “believed” it. This is the catch a same-model rubber-stamp misses.
In a measured run, an Eli Lilly position memo grounded 10 of 12 load-bearing claims verbatim and the verifier caught the other two — a reformatted date list and a fact imported from training knowledge — exactly the two failure modes shown above.
What we measured
A separate judge model (claude-sonnet-4-6 — same vendor as the analyst, which we disclose rather than hide) reads the hydrated content of each cited source and rules supported, partial, or unsupported, across 15 questions spanning ten industries and a negative control. Stanford’s 17–34% hallucination finding for incumbent legal tools is directional context, not an apples-to-apples benchmark — the task and definitions differ. The live span-grounding verifier separately grounded 174 of 183 load-bearing claims (95.1%) verbatim across the same runs — 163 of them traced to the exact source [n] the claim cites, the stricter correspondence check.
Zero invented reference numbers across all 15 eval questions — every [n] resolves to a source that was actually retrieved. The negative control (an entity with no US footprint) produced a clean 'no records found' answer with zero fabricated sources.
Supported + 0.5·partial over 454 cited claims, judged by a separate model (claude-sonnet-4-6) against the hydrated content of each cited record; corpus totals must cite their badged aggregate source, and every cited aggregate was re-counted against the live database by code — 67 of 67 clean. Up from 75.7% a day earlier after the citation mechanics were fixed at the root.
17 of 454 claims where the cited source did not back the sentence — mostly transcription slips like a filing year misread from context. We publish the failure modes, not just the rate; each one is in the committed eval results.
Where we’re honest about the edge: citation validity proves the [n] points at a real retrieved source; citation faithfulness proves that source actually supports the sentence. We measure both — and we report the flags, the grounding count, and the partials rather than rounding them away. Known limits of the measurement itself: the judge reads bounded excerpts of each record, not entire documents; only sentences that carry a citation are faithfulness-scored (roughly a third of substantive sentences do); and faithfulness says nothing about corpus currency or completeness — it measures whether cited sources support their sentences against a fixed snapshot. The grounding check confirms a claim is supported in the retrieved evidence, and additionally traces each grounded claim to the specific source [n] it cites — reporting how many trace verbatim to their exact cited record, not merely somewhere in the evidence. Where a claim legitimately draws on more than one source it still grounds globally, so this stricter count is reported as additional assurance, never as a flag. The check is not a completeness guarantee; it is the reason you can verify in seconds instead of taking the model’s word.
The invariants
The audit trail
Every finished artifact is written to your account’s run history the moment it completes: the full artifact as generated, the evidence horizon it was built against, the verification result, the model, and a content digest. Opening a stored run replays that record — the model is not called again and the corpus is not re-queried, so a memo you circulated in March is still the memo you circulated in March.
The dossier, the brief, the comment letter, the memo — each one shows its verification state inline. Pick one and follow a citation to its publisher.