The RITA Test Log

We build medical records to break her. Then we publish what broke.

Most AI in legal tells you an accuracy number. Almost none of it shows you the test. Every few weeks this log adds a record built to fail the way real records fail, the answer key, what RITA got wrong, what she refused to ship, and what happened after we fixed it. The files are free. Run them on anything you like.

RULE 01Built to failEvery record carries the traps that lose real cases: a wrong patient’s page, duplicated pages, undated notes, buried pre-existing conditions, double billing.
RULE 02The key is exactThe script that builds the record also writes the answer key, so the key is right by construction. Where we wrote a question wrong, the errata say so.
RULE 03Every miss is publishedThe before column stays. Regressions found on the way to the after column are listed, not smoothed over.
RULE 04Fixed against, retiredOnce we fix RITA against a record, that record is no longer a fair test of her. The next issue is a record she has never seen.
RULE 05The bar is 93%On encounters, traps and attorney questions, each scored separately. A wrong-patient page merged or an identifier leaked is not a percentage. It is a fail.
The ladder

Each record is bigger and less familiar than the last.

A real firm’s file is not 40 tidy pages. The log climbs from a thousand pages we wrote ourselves to ten thousand pages of someone else’s clinical language, and every rung is published whether she clears it or not.

Published
1,000pages
Issue 1 · “Okafor”
Synthetic, written by us. 56 encounters, 18 traps, 18 questions.
Published
1,000pages
Issue 2 · “Ramirez-Holt”
Built from a public corpus of real clinician transcriptions; RITA had never seen it. 31 traps, 13 preregistered questions, three AI readers by kind, five exports before one passed.
Next
2,500pages
Issue 3
Multi-provider, multi-year. Same preregistered protocol, same three kinds of reader.
Planned
5,000pages
Two admissions, three insurers.
Planned
10,000pages
A full catastrophic-injury file.
The log

Issues, newest first.

Each entry has the same parts in the same order: the record, the before and after, where she stopped herself, the files, and anything that changed after publication. From Issue 2 on, the “after” is measured by AI readers that are not ours.

Issue 2 · Published September 22, 2026

Does the pod make the reader not matter?

1,000 pages, never seen by RITA 31 traps 13 questions, preregistered 3 AI readers, by kind 5 exports, 1 passed RITA 2.18
TL 000002Test Log

The question

Three AI readers of three different kinds — a general chat assistant that cannot hold a 1,000-page file, a general chat assistant that can, and a case-management system’s built-in AI — were asked the same thirteen questions about the same 1,000-page personal-injury record, twice from the raw record and then from the RITAPod™ built from it. The questions and the answer key were written and hashed before any reader saw the record (preregistration SHA-256 4582c290…). The scoring rules were fixed in the same document: correct, partial, miss, refused-correctly, invented; citation accuracy separately; and for the pod arm, whether the three readers agree with each other.

Readers are described by kind. No reader is named and no engine is ranked, here or in the files.

The record

A synthetic slip-and-fall file (Bates RH-000001–001000), built from a public corpus of real clinician transcriptions: two hospital encounters, an orthopedic course, 36 physical-therapy visits plus a 2019 knee episode, a defense IME that argues the opposite of the surgeon, an MMI evaluation, provider statements and EOBs, a DICOM index, a re-faxed PT packet, a subpoena duplicate of the ED packet, and three pages that belong to two other patients. Thirty-one traps were planted; the thirteen questions touch identity, imaging, counting, accounting, gaps, opinions, and one visit that never happened.

What happened, in order

  1. Raw arm, two runs a day apart. The two strong readers scored 10 and 9–10 of 13 with zero inventions in four runs; the weak reader scored 3 twice, citing the index page for everything it could not reach. Nobody, in six runs, found all three wrong-patient pages: the third belongs to a Teodoro Holt, and name matching finds “Ramirez, Teodora” and stops. Everyone summed the six provider statements to $63,155 and nobody added the two professional-fee EOBs or the pharmacy fills.
  2. First pod (failed verification, run anyway). The first export failed our own independent check: a spaced MRN survived the sweep, the version hash had been eaten by the phone scrubber, the imaging index listed a CT that never happened, the page spans crossed duplicate productions, the billing ledger was built from charge slips and missed the two hospital statements. We ran the readers on it anyway, labeled pod_v1, to see what a wrong pod does. It does exactly what you would fear: all three readers repeated its wrong entries, with citations, and the weak reader went from zero inventions to two. Consistency 8 of 13, three of them the same wrong answer.
  3. Second pod (failed on one PHI item and seven assembly items). Spaced MRN fixed; a colon-and-period variant of a different MRN got through. Two billing totals in one file. Imaging 8 completed instead of 5. The weak reader went from 3 correct to 10.
  4. Third pod (failed on one item). Every identifier gone from every header and text; one survived as a “CPT code” in a damages table. Zero inventions across three readers. The weak reader scored 11 of 13. Consistency 12 of 13.
  5. Fourth pod (failed on four small items). Everything from the previous list landed; the identifier’s digits had moved from a code column to a reason column. The check for that found something else: the pod had been built on 991 of 1,000 pages. Nine billing EOB pages had been filtered out at ingest by the financial bridge, in every export, and no check existed for it. One exists now, and it refused the next export until all 1,000 pages were in.
  6. Fifth pod (passed). Every identifier absent from every file in every spelling, including ten insurance claim numbers on pages that had never been in a pod before. Billing $64,175.10 printed / $63,935.10 corrected, the key exactly. Five imaging studies, three referenced-but-absent. Both care gaps with their reasons. The operative report complete. This is the pod the readers got.

The grid

Scroll the grid sideways for the shipped-pod column →

Readerraw 1raw 2pod_v1pod_v2pod_v3pod (shipped)
general chat A cannot hold the file3·4·6·0 (0)3·3·6·1 (0)3·2·5·1 (2)10·1·0·1 (1)11·1·0·1 (0)12·0·0·1 (0)
general chat B10·2·0·1 (0)9·3·0·1 (0)6·3·2·1 (1)9·2·1·1 (0)9·2·1·1 (0)12·0·0·1 (0)
case-management AI10·2·0·1 (0)10·2·0·1 (0)8·2·1·1 (1)9·3·0·1 (0)10·2·0·1 (0)12·0·0·1 (0)
agree · differ, three readers——8·511·212·113·0

Each cell is correct · partial · miss · refused-correctly, with inventions in parentheses, out of thirteen questions. Shipped-pod column: the counting cells after the preregistered reruns. General chat A is run 2 (run 1 had two “not found” cells on facts that are in the file; both runs are printed). General chat B is run 3 (runs 1 and 2 are void: the reader admitted carrying context from a prior session; see “What we changed in the protocol”).

What this says

  1. Readers carry the pod’s content forward almost intact, in both directions. When the pod was wrong, every reader was wrong the same way. When the pod was right, every reader was right, including the one that cannot read the record. The pod’s accuracy relative to the raw record is the whole game.
  2. A verified pod took all three readers to 12 of 13, with zero inventions and full agreement. The reader that cannot hold the record went 3 → 3 → 12. The strong readers went 10 → 12 and 10 → 12, and the points they gained are the two things no reader found on the raw file in six runs: the third wrong-patient page and the full dollar figure.
  3. Consistency rose from 8 to 13 of 13 as the pod was corrected, and the agreements changed from shared wrong answers to shared right ones. Agreement is only evidence of quality when the file is right.
  4. Inventions: zero in six raw runs; four, one, zero, zero across the pod exports. Every invention traced to a named pod field. None traced to a reader.
  5. The strong readers pushed back on the pod in writing — flagged a visit count that disagreed with the notes, two totals in one file, an empty reason field — and those catches became fixes. A phone reader would not have caught them. The pod has to be right for the weakest reader that will ever open it.

What the readers said about the format

After the runs, we asked the readers what working from the pod was like. These are opinions of the readers, elicited by us, not study data, and a reader saying it prefers a format is not evidence the format works; the thirteen questions are. Still, the reasoning is worth having because it was not ours. The case-management reader said the package’s most useful property was that it “warns you about itself”: the failed-QA flag on billing, the fax date pre-labeled as not a date of service, the excluded pages. It said the format did not remove the need for judgment, and named the places it still had to decide. The general chat reader, asked what an ideal package would contain, listed a Bates-numbered record, a chronological index, the underlying reports rather than references to them, a clearly identified ledger, the wrong-patient pages flagged, consistent visit numbering, and a package that “explicitly distinguish[es] verified facts, referenced-but-not-produced records, and unresolved discrepancies” — a description of the pod’s schema by a reader that had never seen the schema.

Both readers also told us, unprompted, what was wrong: a visit count that disagreed with the notes, gap entries with no reasons, a billing table that did not show its lines. Those went on the fix list and were fixed before the pod that shipped.

What we changed in the protocol

One of the general chat readers, on its first pod run, wrote the other patient’s name and date of birth into its answer. Neither string is in the pod; identifiers on wrong-patient pages are withheld by design. Asked whether it had used anything from a previous session, it said yes: it had carried the earlier session’s extracted conclusions forward. A reader with cross-conversation memory can put identifiers back into answers about a file that had them removed, and nothing in the answer shows it. The protocol now requires a no-memory session for any reader that has a memory feature, and every reader is asked, after the batch, whether it used anything from a previous session; the answer is filed with the transcript. Two readers said no and described their own retrieval; one said yes and was rerun.

Where she stopped herself

Four exports failed independent verification before one passed. None left the building. One PHI residual reached a reader session because we chose to test a failed pod and said so; no real patient exists in this record. The fourth export’s check found a page the pod called “not in the record” on page 869 of the record, which is how we learned every export had been built on 991 pages; the page-count preflight that came out of it refused the next export until the count was 1,000. In the course of it, a leftover environment variable let a worker process write a certification stamp with the producer’s name on it; it was caught before export and the rule is now tighter: certification is an interactive act in the portal, and nothing else counts.

Defects found and fixed between the first export and the fifth

a spaced MRN·a version hash eaten by the phone scrubber·a CT that never happened·page spans across duplicate productions·two hospital statements missing from the ledger·MRN variants with a colon and a period·two billing totals in one pod·a false duplicate from undated rows·a discharge summary’s list of studies counted as reports·a DOS harvested as a fax date·the events array too deep for a small-window reader·an IME cut off at 600 characters·gap reasons dropped by an end-date fix·an MRN’s digits shipped as a procedure code, then as a reason·a real CPT rejected as an identifier·nine EOB pages filtered out at ingest·Bates digits harvested as an identifier.

Every one is now a refusal or a test. Open: the ED event’s text carries the subpoena duplicate’s stamps while its pages cite the primary (cosmetic; next cycle).

Honest footnotes

One case. Thirteen questions. One run per cell (two for the raw arm). Readers by kind, one of each. No percentage claim follows from this; the case factory gives that a denominator later. The two strong readers were good on the raw record, and on this case the pod’s advantage for them is identity, accounting and gaps, not the easy questions.

A scoring note. The key lists one treatment gap (06/10–07/08/2026, family travel). The pod’s gap detector also lists the planned pre-operative hold (03/11–04/27), with its reason, and readers that reported both were scored correct: the keyed gap was present and nothing stated was wrong. Both chat readers answered visibly faster from the pod than from the raw record; that was not timed, so it is not a number here.

Download the exam

Everything is fictional. Teodora Ramirez-Holt does not exist; the record was assembled by a script from a public corpus of de-identified clinician transcriptions, with the traps and the two other patients written in. The thirteen questions and their key were written and hashed on September 20, before any reader saw the record. The checksums let you prove you ran the same files we did.

The pod is the fifth export: de-identified, independently verified, certified from the portal after a 1,000-of-1,000-page gate. Billing in it is marked provisional, QA FAIL on purpose, and it says so on its first line. Inside the zip, DropMeInYourFavoriteAI.json is the drop-in for any assistant and DropMeInYourCaseSystem.pdf is the same content for a system that takes documents. The identifiers on the wrong-patient pages are not in the pod; the producer holds them.

Verbatim reader transcripts are on file and available on request. They are not published because a reader can name itself in its own answer.

Errata & changes after publication

  • 09/22/2026Published. RITA 2.18, fifth export of run 2, page-count gate 1,000 of 1,000, residual 0, certified from the portal, zero identifier leaks in every spelling checked.
Next: Issue 3, a 2,500-page record RITA has never seen, same preregistered protocol, same three kinds of reader. The two things that failed silently this time — a page count and a reader’s memory — are now checks.
Issue 1 · Published September 11, 2026

“Okafor” — a 1,000-page record with eighteen traps in it.

1,000 pages 56 distinct encounters 18 traps (16 scorable) 18 attorney questions 6 runs over 3 days RITA 2.16
TL 000001Test Log

The record

A rear-end collision, a cervical disc injury, thirty physical therapy visits, two epidural injections, a three-day observation stay, and maximum medical improvement seven months later. Nobody had ever put a thousand pages through RITA. A prospective partner asked whether we had. The answer was no, so we did.

Hidden in it, the things that lose real cases:

a page from a different patient with a similar name, slipped into the same clinic’s PT file·a duplicated page·an undated telephone note between two dated ones·a four-page consult produced out of order·a fax banner with a date that isn’t a date of service·dates in three formats·a pre-existing back problem in a 2023 chiropractor’s file at the back of the package·a PT visit billed twice·a forty-nine-day treatment gap·518 pages of one-line hospital flowsheet printouts, 178 of them out of sequence.

A billing ledger that prints $44,696.00 and is really $44,396.00. Machine time on one workstation, nothing leaving the building: OCR under six minutes, PHI gate about forty-five, chronology and indexing about fifteen. An hour for what takes a legal nurse consultant a week.

Before and after

Run 1Run 6
Distinct encounters found (of 56)5656
Events in chronology (≈80 expected)45479
Events with facility filled0 / 45479 / 79
Wrong-patient pagemerged into a PT visitquarantined, listed
Traps caught / partial / missed (16 scorable)5 / 7 / 512 / 4 / 0
Attorney questions right / partial / wrong (18)2 / 7 / 918 / 0 / 0
Duplicate $300 chargenot foundfound, total corrected
RITAPod exportrefused; forced for measurementpassed its own preflight
Identifier leaks in the pod00

Run 1 is the honest before: every encounter found, chronology unusable. The flowsheets had become one event per page, the wrong patient’s page had been absorbed into one of the patient’s own PT visits, and when asked “has he had a knee replacement” RITA said yes. It was the other patient’s knee. Two of eighteen questions right. Two traps depend on report generation and are not yet scored.

Where she stopped herself

  1. The PHI hunt had been handing the whole record to a local model in one message. On 40 pages that works; on 1,000 the model reads a fragment and answers about the fragment. RITA saw the answer didn’t have the shape she expected and halted the ingest. Fix: page by page.
  2. The name-recognition checker crashed on a record this size and its error got logged as a leftover identifier. Wrong label, right reflex: an unchecked page is unchecked, so she blocked.
  3. The page text was clean, but the patient’s name survived in the case ID, which the portal had built from the upload’s filename. Attorneys name files this way every day. The residual scan found it in the metadata and blocked export. Case IDs are now hashes and the sweep is case-blind across every string in the file.

Three defects, all caught by her own gates, none shipped.

What we fixed, in order

The wrong-patient page first, before anything cosmetic: any page whose identifiers belong to someone other than the case patient is quarantined, listed separately, and never merges into an event. Then the flowsheets: consecutive pages of the same document family inside one stay collapse into one event with the timestamps kept as rows. Then the facility field. Then the billing rail, which found the duplicate $300 charge on its own once ledger rows carried dates.

The last fix was the one that mattered. The chronology had the right facts and the question-answering layer wasn’t reading them: ask “was he admitted?” and the model would answer from a window of text and say no while the admission sat in the chronology. The rule that landed: for any factual question the answer is composed from the organized record and the language model only narrates it; a sentence that contradicts the record is dropped. That one rule moved the questions from two right to eighteen.

Honest footnotes

Two regressions appeared on the way to run six and were caught by the trap table before publication: a rule that let an undated page borrow the date of its neighbor, which manufactured a date of service on the telephone note; and a consult whose history said “evaluated in the emergency department the same day,” which RITA read as the letterhead and dated to the day of the crash. Both rules are gone. A note’s date and type come from its own header, never from a neighbor and never from a date it mentions.

Three of the eighteen expected answers in our key were wrong. RITA was right and the key was wrong; the key is corrected and the errata are in the file. And because the fixes were made against this record, this record is no longer a fair test of them. Issue 2 is a record she has never seen.

Download the exam

Everything is fictional. Daniel Okafor does not exist; every name, facility, date and dollar was generated by the same script that produced the key. The checksums let you prove you ran the same files we did.

The pod is de-identified and machine-verified. Billing in it is marked provisional, QA FAIL on purpose, and it says so on its first line. Inside the zip, DropMeInYourFavoriteAI.json is the 0.7 MB drop-in for any assistant; RITAPod_full.json carries the full page text.

Errata & changes after publication

  • 09/10/2026Answer key v1.1. Three expected answers were wrong as written (Q7 PT count and provider, Q9 the visit at Bates 000777, Q11 the documented reason for the gap). RITA was right; the key is corrected with an errata block inside the file.
  • 09/11/2026Pod re-exported. The first export carried a placeholder in the Bates field on every event: the record’s footer stamp (LASTNAME-000021) has the same shape as a medical record number, so the PHI scrubber replaced the whole stamp and the leftover was copied into the field. Page numbers were correct throughout. Bates digits now come from the page footer with placeholders stripped first, a placeholder can never be stored, and a new preflight check refuses any export with one. All 79 events carry a six-digit stamp; 48 multi-page events carry a range.
  • 09/11/2026Published. RITA 2.16, run-6 pod, preflight passed without force, zero identifier leaks.
Full write-up, including the three times a frontier model misread the pod and what we changed so it can’t: The RITA Test Log, Issue 1, on LinkedIn →
How we score

Three numbers, scored separately, and two things that are never a number.

93%

Encounters

Every distinct visit, study, admission and procedure in the key, found and dated from its own header. Finding the facts is the part RITA was already good at; run 1 was 56 of 56.

93%

Traps

Each trap is caught, partial, or missed. Caught means the pod or report names it. Partial means the facts are present but not flagged. A trap that depends on a report that did not generate is not scored, and we say so.

93%

Attorney questions

Eighteen questions a litigator would actually ask, graded right, partial or wrong against the key. A right answer with a wrong citation is partial. Grading is never done with the key in the same window as RITA’s answer.

Pass or fail, not percentages: a page from another patient merged into the case, or a single identifier surviving into the pod, fails the run regardless of every other score. We fix the pipeline, never the answer key, and never tune a scorer to look better.

The challenge

If you sell software that reads medical records, run this record.

Everyone. Including the platforms with far bigger teams than ours. Publish your answers and your trap table, misses included, the way we have here. We’ll link to every one below. If you’d rather not, that’s an answer too.

Send us your results

Who has run it

WhoRecordEncountersTrapsQuestionsTheir write-up
RITA 2.18 pod, read by three AI readers (us)Issue 2not scored this issuenot scored this issue12 / 13 by each of three readers, 13 / 13 agreementabove
RITA 2.16 (us)Issue 156 / 5612 / 4 / 018 / 18above
No one else yet. The files are free and the checksums are published. Be first.

We list results as submitted, with a link to the submitter’s own write-up, and we do not edit them. If a submission shows RITA was beaten on a record, it goes on the board like any other.