The reader, in your browser
Pick a town nobody has read. A model runs inside this browser tab, on your own graphics card, and reads that town's scanned government reports. Whatever survives its checks is re-verified against the archive before it is published. Nothing is installed.
Pick a town
Unread towns first. More reports means a longer read.
| kind | parameter | value | sentence it was read from | page | checks |
|---|
the model's raw output for the last page — check the reader's work the way the reader checks the archive's
Or try it on one page first
— one real page (the scan ↗, Internet Archive). Reads in under a minute; publishes nothing.
the page text being read (from the archive's OCR layer)
Honest scope. This is a partial reader. Measured against four pages a person read by hand, it finds 57.4% of the values they found, and 81.2% of what it publishes matches one of them. It fabricated nothing: every number it published appears in the sentence it quoted.
What it got wrong, and why not a better model
Nine records missed the answer key. Eight are real readings the key does not list on their own — a count of pumps, the same value sent twice, a value whose unit it left blank. One is a real misreading: where the page prints “0. 196 mil gal” it returned 196. That is the kind of error these checks cannot catch and a person can.
An earlier version of this page estimated “roughly half” before anything had been measured. The first measurement came back at 32.4%.
What your browser sends is only what passed both checks here, and this site re-verifies every sentence and number against the scanned page on archive.org before publishing — the same rule for every reader, human or machine.
A stronger model (gemma4:e2b, 100% precision and 44% recall on the same pages) cannot run in any browser: all three published builds fall silent past roughly 512 prompt tokens, a defect in the shared WebGPU kernels for its architecture. The day the upstream fix lands it takes one flag to retest.