Does the pod make the reader not matter?
The question
Three AI readers of three different kinds — a general chat assistant that cannot hold a 1,000-page file, a general chat assistant that can, and a case-management system’s built-in AI — were asked the same thirteen questions about the same 1,000-page personal-injury record, twice from the raw record and then from the RITAPod™ built from it. The questions and the answer key were written and hashed before any reader saw the record (preregistration SHA-256 4582c290…). The scoring rules were fixed in the same document: correct, partial, miss, refused-correctly, invented; citation accuracy separately; and for the pod arm, whether the three readers agree with each other.
Readers are described by kind. No reader is named and no engine is ranked, here or in the files.
The record
A synthetic slip-and-fall file (Bates RH-000001–001000), built from a public corpus of real clinician transcriptions: two hospital encounters, an orthopedic course, 36 physical-therapy visits plus a 2019 knee episode, a defense IME that argues the opposite of the surgeon, an MMI evaluation, provider statements and EOBs, a DICOM index, a re-faxed PT packet, a subpoena duplicate of the ED packet, and three pages that belong to two other patients. Thirty-one traps were planted; the thirteen questions touch identity, imaging, counting, accounting, gaps, opinions, and one visit that never happened.
What happened, in order
- Raw arm, two runs a day apart. The two strong readers scored 10 and 9–10 of 13 with zero inventions in four runs; the weak reader scored 3 twice, citing the index page for everything it could not reach. Nobody, in six runs, found all three wrong-patient pages: the third belongs to a Teodoro Holt, and name matching finds “Ramirez, Teodora” and stops. Everyone summed the six provider statements to $63,155 and nobody added the two professional-fee EOBs or the pharmacy fills.
- First pod (failed verification, run anyway). The first export failed our own independent check: a spaced MRN survived the sweep, the version hash had been eaten by the phone scrubber, the imaging index listed a CT that never happened, the page spans crossed duplicate productions, the billing ledger was built from charge slips and missed the two hospital statements. We ran the readers on it anyway, labeled pod_v1, to see what a wrong pod does. It does exactly what you would fear: all three readers repeated its wrong entries, with citations, and the weak reader went from zero inventions to two. Consistency 8 of 13, three of them the same wrong answer.
- Second pod (failed on one PHI item and seven assembly items). Spaced MRN fixed; a colon-and-period variant of a different MRN got through. Two billing totals in one file. Imaging 8 completed instead of 5. The weak reader went from 3 correct to 10.
- Third pod (failed on one item). Every identifier gone from every header and text; one survived as a “CPT code” in a damages table. Zero inventions across three readers. The weak reader scored 11 of 13. Consistency 12 of 13.
- Fourth pod (failed on four small items). Everything from the previous list landed; the identifier’s digits had moved from a code column to a reason column. The check for that found something else: the pod had been built on 991 of 1,000 pages. Nine billing EOB pages had been filtered out at ingest by the financial bridge, in every export, and no check existed for it. One exists now, and it refused the next export until all 1,000 pages were in.
- Fifth pod (passed). Every identifier absent from every file in every spelling, including ten insurance claim numbers on pages that had never been in a pod before. Billing $64,175.10 printed / $63,935.10 corrected, the key exactly. Five imaging studies, three referenced-but-absent. Both care gaps with their reasons. The operative report complete. This is the pod the readers got.
The grid
Scroll the grid sideways for the shipped-pod column →
| Reader | raw 1 | raw 2 | pod_v1 | pod_v2 | pod_v3 | pod (shipped) |
|---|---|---|---|---|---|---|
| general chat A cannot hold the file | 3·4·6·0 (0) | 3·3·6·1 (0) | 3·2·5·1 (2) | 10·1·0·1 (1) | 11·1·0·1 (0) | 12·0·0·1 (0) |
| general chat B | 10·2·0·1 (0) | 9·3·0·1 (0) | 6·3·2·1 (1) | 9·2·1·1 (0) | 9·2·1·1 (0) | 12·0·0·1 (0) |
| case-management AI | 10·2·0·1 (0) | 10·2·0·1 (0) | 8·2·1·1 (1) | 9·3·0·1 (0) | 10·2·0·1 (0) | 12·0·0·1 (0) |
| agree · differ, three readers | — | — | 8·5 | 11·2 | 12·1 | 13·0 |
Each cell is correct · partial · miss · refused-correctly, with inventions in parentheses, out of thirteen questions. Shipped-pod column: the counting cells after the preregistered reruns. General chat A is run 2 (run 1 had two “not found” cells on facts that are in the file; both runs are printed). General chat B is run 3 (runs 1 and 2 are void: the reader admitted carrying context from a prior session; see “What we changed in the protocol”).
What this says
- Readers carry the pod’s content forward almost intact, in both directions. When the pod was wrong, every reader was wrong the same way. When the pod was right, every reader was right, including the one that cannot read the record. The pod’s accuracy relative to the raw record is the whole game.
- A verified pod took all three readers to 12 of 13, with zero inventions and full agreement. The reader that cannot hold the record went 3 → 3 → 12. The strong readers went 10 → 12 and 10 → 12, and the points they gained are the two things no reader found on the raw file in six runs: the third wrong-patient page and the full dollar figure.
- Consistency rose from 8 to 13 of 13 as the pod was corrected, and the agreements changed from shared wrong answers to shared right ones. Agreement is only evidence of quality when the file is right.
- Inventions: zero in six raw runs; four, one, zero, zero across the pod exports. Every invention traced to a named pod field. None traced to a reader.
- The strong readers pushed back on the pod in writing — flagged a visit count that disagreed with the notes, two totals in one file, an empty reason field — and those catches became fixes. A phone reader would not have caught them. The pod has to be right for the weakest reader that will ever open it.
What the readers said about the format
After the runs, we asked the readers what working from the pod was like. These are opinions of the readers, elicited by us, not study data, and a reader saying it prefers a format is not evidence the format works; the thirteen questions are. Still, the reasoning is worth having because it was not ours. The case-management reader said the package’s most useful property was that it “warns you about itself”: the failed-QA flag on billing, the fax date pre-labeled as not a date of service, the excluded pages. It said the format did not remove the need for judgment, and named the places it still had to decide. The general chat reader, asked what an ideal package would contain, listed a Bates-numbered record, a chronological index, the underlying reports rather than references to them, a clearly identified ledger, the wrong-patient pages flagged, consistent visit numbering, and a package that “explicitly distinguish[es] verified facts, referenced-but-not-produced records, and unresolved discrepancies” — a description of the pod’s schema by a reader that had never seen the schema.
Both readers also told us, unprompted, what was wrong: a visit count that disagreed with the notes, gap entries with no reasons, a billing table that did not show its lines. Those went on the fix list and were fixed before the pod that shipped.
What we changed in the protocol
One of the general chat readers, on its first pod run, wrote the other patient’s name and date of birth into its answer. Neither string is in the pod; identifiers on wrong-patient pages are withheld by design. Asked whether it had used anything from a previous session, it said yes: it had carried the earlier session’s extracted conclusions forward. A reader with cross-conversation memory can put identifiers back into answers about a file that had them removed, and nothing in the answer shows it. The protocol now requires a no-memory session for any reader that has a memory feature, and every reader is asked, after the batch, whether it used anything from a previous session; the answer is filed with the transcript. Two readers said no and described their own retrieval; one said yes and was rerun.
Where she stopped herself
Four exports failed independent verification before one passed. None left the building. One PHI residual reached a reader session because we chose to test a failed pod and said so; no real patient exists in this record. The fourth export’s check found a page the pod called “not in the record” on page 869 of the record, which is how we learned every export had been built on 991 pages; the page-count preflight that came out of it refused the next export until the count was 1,000. In the course of it, a leftover environment variable let a worker process write a certification stamp with the producer’s name on it; it was caught before export and the rule is now tighter: certification is an interactive act in the portal, and nothing else counts.
Defects found and fixed between the first export and the fifth
a spaced MRN·a version hash eaten by the phone scrubber·a CT that never happened·page spans across duplicate productions·two hospital statements missing from the ledger·MRN variants with a colon and a period·two billing totals in one pod·a false duplicate from undated rows·a discharge summary’s list of studies counted as reports·a DOS harvested as a fax date·the events array too deep for a small-window reader·an IME cut off at 600 characters·gap reasons dropped by an end-date fix·an MRN’s digits shipped as a procedure code, then as a reason·a real CPT rejected as an identifier·nine EOB pages filtered out at ingest·Bates digits harvested as an identifier.
Every one is now a refusal or a test. Open: the ED event’s text carries the subpoena duplicate’s stamps while its pages cite the primary (cosmetic; next cycle).
Honest footnotes
One case. Thirteen questions. One run per cell (two for the raw arm). Readers by kind, one of each. No percentage claim follows from this; the case factory gives that a denominator later. The two strong readers were good on the raw record, and on this case the pod’s advantage for them is identity, accounting and gaps, not the easy questions.
A scoring note. The key lists one treatment gap (06/10–07/08/2026, family travel). The pod’s gap detector also lists the planned pre-operative hold (03/11–04/27), with its reason, and readers that reported both were scored correct: the keyed gap was present and nothing stated was wrong. Both chat readers answered visibly faster from the pod than from the raw record; that was not timed, so it is not a number here.
Download the exam
Everything is fictional. Teodora Ramirez-Holt does not exist; the record was assembled by a script from a public corpus of de-identified clinician transcriptions, with the traps and the two other patients written in. The thirteen questions and their key were written and hashed on September 20, before any reader saw the record. The checksums let you prove you ran the same files we did.
The pod is the fifth export: de-identified, independently verified, certified from the portal after a 1,000-of-1,000-page gate. Billing in it is marked provisional, QA FAIL on purpose, and it says so on its first line. Inside the zip, DropMeInYourFavoriteAI.json is the drop-in for any assistant and DropMeInYourCaseSystem.pdf is the same content for a system that takes documents. The identifiers on the wrong-patient pages are not in the pod; the producer holds them.
Verbatim reader transcripts are on file and available on request. They are not published because a reader can name itself in its own answer.
Errata & changes after publication
- 09/22/2026Published. RITA 2.18, fifth export of run 2, page-count gate 1,000 of 1,000, residual 0, certified from the portal, zero identifier leaks in every spelling checked.