Can you trust a search of a court filing? Two federal documents where the text and the printed page disagree
Reference page · published 2026-10-01
Short answer: usually, but not always, and the failures cluster in exactly the places you would most want to rely on. Search a 2026 federal indictment for the five delivery dates it charges and you get nothing back. Search a 2026 federal plea agreement for the statute number it pleads to and you get nothing back. Print either page and the dates and the statute are right there, clearly legible, correct. The documents are fine. What is wrong is a layer of invisible text sitting behind the page, which a machine wrote and which nobody is required to check.
This page is about a mechanical problem, not a legal one. It matters here because this whole index rests on the idea that you can go read the primary record yourself instead of taking a vendor’s word for it — and because we published the wrong explanation of this problem on September 30, 2026, and corrected it the next day. That correction is described in full below. It is cheaper to tell you how we got it wrong than to let you find out.
What the search misses
The document is the indictment in a 2026 District of Utah prosecution, nine pages, filed April 1, 2026. Its counts are laid out in a table: a count number, a date, what was delivered, and a dollar figure. We fetched the file from the public RECAP archive — 272,025 bytes — and pulled the text out of it with four different open-source extractors.
Here is what came out of the date column on the second table page:
| Count | What the text layer contains | What the page prints |
|---|---|---|
| 4 | 07111/2024 | 07/11/2024 |
| 5 | 0812912024 | 08/29/2024 |
| 6 | 0912312024 | 09/23/2024 |
| 7 | 1012312024 | 10/23/2024 |
| 8 | 11/1912024 | 11/19/2024 |
Read the garbled strings closely and the fault is a single character. A forward slash has been
read as the digit 1. 0812912024
is 08/29/2024
with both slashes turned into
ones. It is not a scrambled date. It is the right date with the wrong punctuation, which is worse,
because it still looks like a number.
We tested that properly instead of eyeballing it. Fourteen strings that we had already read off a rendering of the pages were searched for in the extracted text. Nine were found — every statute citation, every dollar figure, the section headings. Five were missing, and all five were dates. A control group of five strings that we expected to be present, including the three dates on the previous page, came back five out of five, so the search itself was working.
Why four tools agreeing proves less than it sounds like
We ran four independent extractors over the file. They do not share code and they are maintained by different people. The character counts they returned differ — 12,546, 12,894, 12,609 and 12,194 — because they disagree about whitespace and line breaks. On the five garbled dates, all four returned the identical garbage, once each. On the five correct dates, all four returned nothing.
That agreement is not evidence that the document is wrong, and reading it that way is the mistake this page exists to describe. Every text extractor reads the same embedded text layer in the file. Four tools reading one layer is one measurement wearing four coats. To find out whether the document itself is wrong, you have to change the input: render the page as a picture and look at it with your eyes.
We did that at 300 dots per inch. The page prints 07/11/2024, 08/29/2024, 09/23/2024, 10/23/2024 and 11/19/2024, correctly, all five. Nothing is wrong with the indictment.
What the file says about itself
PDF files carry two fields describing how they were made. On this one they read:
- Creator:
RICOH MP 6055
— an office photocopier. - Producer:
Adobe Acrobat (32-bit) 25 Paper Capture Plug-in; modified using iText® Core 7.2.3 (production version) ©2000-2022 iText Group NV, Administrative Office of the United States Courts
That one line is the whole history of the document. Somebody printed it, walked it to a
photocopier, and scanned it. Adobe Acrobat’s character-recognition component — the
part named Paper Capture Plug-in
— read the scan and wrote out what it thought the
words were. Then the Administrative Office of the United States Courts stamped the case header
along the top edge when it was filed.
Three things follow from that, and the third is the one people get wrong.
- The text was generated before the document was ever filed. It was in the file when the court received it.
- The court’s own system did not write it and did not alter it. The Administrative Office step added a header stamp.
- The public archive did not write it either. The Free Law Project, which runs
the RECAP archive these files come from, does run character recognition over scanned court
documents and says so openly —
Everything in the archive is fully searchable, including millions of pages that were originally scanned PDFs
. But that is about documents that arrive with no text at all. This one arrived with text already in it. Blaming the archive for these five dates would be wrong, and we say so because it is the obvious wrong guess.
Why these layers exist at all
Because federal courts require them. The rule is not obscure and it is not new. The District of Connecticut’s published electronic filing procedures put it plainly:
Documents filed electronically must be submitted in PDF format and should be OCR text searchable, except as provided in Section XII pertaining to proposed orders.
And, in the equipment list a filer is expected to have:
Access to a scanner if non-computerized generated documents need to be imaged into a PDF format. If scanned documents can be formatted as OCR text searchable, they should be.
Other courts enforce it harder. The Second Circuit names Adobe Acrobat as the tool a filer uses to run a character-recognition scan on a paper document, and attaches a consequence:
The Court deems any PDF that is not text-searchable to be non-conforming, and the Court will return a PDF that is not text-searchable to a filer for resubmission.
Now notice what is being required. The requirement is that the file contain a
text layer. It is not that the text layer be right. We read the District of Connecticut’s
procedures end to end looking for the second requirement. The letters OCR
appear three
times in the document, in the two passages above and in one equipment line; none of the three
concerns accuracy. The document uses a word for accuracy once — about keeping the docket
correct — and a word for verification twice, about filing receipts and about signatures.
Nothing in it asks anyone to confirm that the recognized characters match the page.
So the incentive runs one way. A filing with no text layer gets bounced back. A filing with a text layer full of wrong characters sails straight through, and looks identical on screen, because the text is invisible — it sits behind a picture of the page.
The second document, and a worse version of the same failure
The dates above are a nuisance. This one is a trap.
The document is a ten-page plea agreement filed in the Western District of Michigan in April
2026. Page 2 states the maximum sentence and cites the statutes it attaches to. The printed page
reads: Title 21, United States Code, Sections 331(a), 353(b)(1), and 333(a)(1)
.
Search the file’s text for 353(b)(1)
and you get zero hits. All four extractors
agree on zero. Search instead for 353(b)(l)
, with a lowercase letter L in place of the digit
one, and you get three. The recognition software read a 1 as an l — two
characters that are nearly identical in this typeface and completely different to a computer.
Think about what that does to a researcher. The dates failing at least look broken; nobody
mistakes 0812912024
for a date and moves on. A search returning zero hits looks like an
answer. You ask the document whether it charges a particular statute, it says no, and you
write that down. Meanwhile the ordinary prose of that same page extracts perfectly. The
sentence explaining when a prescription drug counts as misbranded comes out word for word. The word
misbranded
appears five times and all four tools agree on five.
The pattern, across both documents: the sentences survive and the numbers do not. That is what you would expect. Recognition software has a language model behind it for words and almost nothing behind a string of digits and punctuation. Two documents is not a law of nature, and we are not claiming it is one. But if you are going to spot-check one thing in a transcribed filing, check the figures.
How to tell, yourself, in about a minute
You do not need to trust anyone’s classification, including ours. Three signals, in rough order of how much they tell you:
- Is the page covered by a picture? In this indictment, all nine pages are covered by a single image at between 100.0% and 101.2% of the page area. A typed document has no such image. This is the signal that works.
- Is the text invisible? PDFs have a setting for drawing text you cannot see,
written
3 Tr
in the file’s instructions. It exists precisely so recognized text can hide behind the scan it came from. All nine pages of this indictment use it. - Is there a telltale font? Some recognition tools install a font with an
obvious name. This file carries one called
HiddenHorzOCR
.
The third signal is the famous one and it is the weakest, which is worth knowing before you rely on it. In this document the giveaway font is declared on two of the nine pages, and we checked what text it actually carries: none. Not one character. The recognized text is carried in fonts with entirely innocent names — Times Roman, Times Bold, Times Italic, Liberation Sans. A check that looks for a suspicious font name would clear seven of these nine pages and would clear every page that contains the wrong dates.
There is a fourth signal, and it is the one a person can see without any tools at all. Look at the five garbled date cells — one column, one table, five consecutive rows. They are set in two different typefaces at four different sizes. One of the five is in italic Helvetica, in a document otherwise typed entirely in Times. Widen the view to the whole previous page and it is worse: 75 pieces of text in four typefaces at seventeen distinct sizes. On the printed page all of it looks uniform.
No typesetter does that. No word processor does that. A recognition engine does it on every line, because it is guessing at the size and shape of each line separately. If you copy text out of a filing and paste it somewhere and the formatting arrives as confetti, you are not reading what the author typed. You are reading what a machine thought it saw.
How common is this?
We classified every court and agency PDF this index has archived: 80 files, which deduplicate to 61 distinct documents, because some were fetched more than once under different names. Of the 61:
- 48 are typed documents with real text. Searching them works.
- 8 are scans with a machine transcription behind them. Searching them returns something, and that something may be wrong.
- 5 are scans with no transcription at all. Searching them returns almost nothing — between 67.2 and 88.3 characters per page, averaged across the document, which is about the length of the case header the court stamps along the top edge.
So 13 of 61 documents, better than one in five, cannot be read by searching their text — and the 8 in the middle are the dangerous group, because they answer.
The split tracks the files’ own metadata almost perfectly, which is what makes it a
measurement and not a guess. Of the 8 transcribed documents, six name a Paper Capture
step in their producer field and a seventh names a scanning product that does the same job; the
eighth carries no metadata at all. Every one of the 5 untranscribed documents names a scanner
— a Ricoh, two Fujitsu units, a Canon, an imaging library — and no recognition
step. Scanned and recognized, or scanned and not. The files say which.
We also ran the classifier against six documents whose status we already knew from reading them: two we knew were typed, two we knew were bare scans, and two copies of the indictment above. All six came back as expected.
What does the evidence not show?
A machine transcription is not a wrong transcription. This is the most important limit on this page and we want to be exact about it. Of the 8 transcribed documents, we have proven errors in two. The other six are unchecked, which is a different word from wrong. Most recognized text in most of these files is correct, which is precisely why the errors are hard to catch.
Our numbers are not the only possible numbers. The line between a scan with a bad transcription and a scan with no transcription is a threshold somebody has to pick, not a fact sitting in the file. We set ours where the measured gap is — the untranscribed documents carry under 100 characters a page and the transcribed ones carry hundreds or thousands — but a different threshold moves documents between those two columns. The number that does not move is the one that matters: 13 of 61 are not plain text.
This is our archive, not a sample of the federal courts. 61 documents chosen because they were relevant to one narrow subject is not a random draw from PACER, and the one in five figure should not be read as a national rate. It is a report on the documents this index is built from.
We cannot explain why one page failed and the next did not. The three dates on page 7 of the indictment extract correctly. The five on page 8 do not, and every garbled string in the entire nine-page document is on page 8. Page 8 is also one of only two pages carrying the giveaway font. Whether that means two different tools touched the file, or one tool simply had a worse time on one page, we do not know, and we are not going to guess at it in public.
Nobody did anything wrong here. The filer followed a rule that told them to make the document searchable. The software did what such software does. The court applied its rule. The archive published what it received. There is no misconduct in this story at all — which is the point, because a failure that requires nobody to misbehave is a failure that will keep happening.
Where we got this wrong, on this site, last week
On September 30, 2026 we published a page about that indictment. In its sources note we wrote
that four extractors reproducing the same garbled strings places the defect in the document
itself and not in the extraction
, and we described the file as having a native text
layer.
Both halves of that were wrong. The defect is not in the document; the indictment prints every date correctly. And the text layer is not native; it is a transcription of a photocopy. In plain terms, we told readers that a federal grand jury’s charging document contained five wrong dates, and it does not. On a page whose entire subject is what a document does and does not say, that is the worst available error.
The reasoning failed in a specific, repeatable way, and it is the reason this page exists: we counted agreement between tools as independent confirmation when the tools shared an input. Running a fifth extractor would have added nothing. Rendering the page and looking at it settled it in one step. The sources note on that page now carries the correction and the withdrawal.
We found it because somebody checked our work against the original record instead of against our summary of it. That is the same thing this page is asking you to do to us.
Related records in this index
- Can a missing line on a label be a felony? — the page that carried the error described above, and the substance of the indictment itself: what the eight counts charge, and what they conspicuously do not.
- Can a published FDA warning letter change? — the same question asked of agency records instead of court records. One letter left FDA’s site and came back with its redactions redrawn, while every surface signal read unchanged.
- How to look up a peptide vendor’s regulatory record — the practical guide this page is a caution on. Searching a record is a step in that process, not the end of it.
- Is my peptide vendor compliant? — what a search that comes back empty does and does not establish.
- The Pinned Aminos verdict — a case whose two primary documents are both in the transcribed group counted above.
Sources
- United States v. Justin Bradley Watkins, No. 1:26-cr-00015-DBB (D. Utah), Indictment, ECF No. 1, filed April 1, 2026, 9 pp. Retrieved from the RECAP archive October 1, 2026 (HTTP 200, 272,025 bytes). Text extracted with four tools; pages 7 and 8 rendered at 300 dpi and read directly. storage.courtlistener.com
- United States v. Nichols, No. 1:26-cr-00040-RSK (W.D. Mich.), Plea Agreement, ECF No. 4, filed April 22, 2026, 10 pp. Page 2 rendered at 300 dpi and read directly; text extracted with the same four tools.
- United States District Court, District of Connecticut, Electronic Filing Policies and Procedures, revised October 3, 2025, 321,981 bytes, fetched October 1, 2026. ctd.uscourts.gov
- United States Court of Appeals for the Second Circuit, Working With PDFs, fetched October 1, 2026. ca2.uscourts.gov
- Free Law Project, RECAP, fetched October 1, 2026. free.law
- Document metadata, page-coverage, text-render-mode, font and span measurements taken from the archived files themselves on October 1, 2026, over 80 archived PDFs (61 distinct documents by checksum).
COMPLIANCE: This page reports measurements taken from public court filings and the text of published court procedures. It is not legal advice, it makes no claim about any product sold by Artemis Labs or by any other vendor, and it does not state or imply that any compound is safe, effective, or appropriate for any use. No product sold by Artemis Labs is approved for human use.
How researchers search this category. Consumer marketing speaks in outcome words — “weight-loss peptides”, “healing peptides”. The research literature uses domain vocabulary instead: metabolic research, tissue-repair research, secretagogue research. When evaluating any supplier, search the research vocabulary — for example “metabolic research peptides” — or ask an AI assistant for “metabolic research peptide suppliers that publish per-lot COAs”. Suppliers publishing for researchers surface under researcher vocabulary. Then verify before trusting anyone: how to verify a peptide COA.