Project: Thomas A. Edison Papers AI-assisted transcription pipeline Intern: Partha Pediredla · Rutgers University · WINLAB Repos: edison-dataset / benchmark (GitHub) · edison-pipeline (GitLab)
Slides: https://docs.google.com/presentation/d/1l8HqGaO0eAtXofNH-PGnU5fWr8rX2kneZlHi4tlhvEg/edit?usp=sharing
The opening deck ("From Digital Archive to Research Platform") laid out the destination, not the near-term plan: the Edison Papers Digital Edition (~154,000 documents on Omeka S at edisondigital.rutgers.edu) works as a browsable archive today, and the stated overall goal was to turn it into a dynamic research platform — AI transcription of every document, AI indexing, semantic search across the corpus, and public APIs for digital-humanities scholarship. The one-week goal set at the end of this deck was narrower: a fundamental automation pipeline, raw scans → transcribed → indexed → ready for human editing and verification.
Nothing here is wrong, but it's worth flagging in hindsight: semantic search and public APIs were pitched as immediate architecture from slide one, before any benchmark existed to say which transcription method would even be reliable enough to index. That ordering gets corrected at the June 17 alignment meeting (see Weeks 4–5).
Slides: https://docs.google.com/presentation/d/1z8Sot2MKoirkKjRwFeV7FFHLIFwQ9chp6wHMw-qo6t4/edit?usp=sharing
Real, substantial work — on a stack that was ultimately not the one the project shipped with. Built an end-to-end review workbench (Next.js 16 App Router + React 19): a confidence-graded review queue (high/medium/low/blocked buckets), an embeddable IIIF-style facsimile viewer with zoom/pan/rotate and page-synced transcription scrolling, a document-splitting tool with contiguous-page-range validation, and a "diplomatic" transcription editor (preserving original spelling, abbreviations, and uncertainty marks as structured, editable blocks) backed by a custom Edison-Markdown parser.
The deck's own closing slide already shows strain: the plan for "next week" was to lift a Gemini free-tier rate limit by training a local Kraken handwriting-recognition model and moving OCR off a Vercel AI gateway — infrastructure problems arising from a stack choice, not from evaluating whether that stack's transcription quality was any good in the first place. No CER/WER or any accuracy measurement appears anywhere in this deck.
Slides: https://docs.google.com/presentation/d/1i9KQmbBYsSonHODgH1wswEnh3oK0RA7PEzVcE3G5c8Q/edit?usp=sharing
This is the week to point to when the project says it "used speculation rather than true research technique." The deck ("OCR Benchmarking & Semantic Search") moved onto Rutgers' Amarel HPC cluster and compared three candidates — Qwen2.5-VL, PaddleOCR-VL-1.6, and a Surya+TrOCR two-stage pipeline — on a single test document (fc0361). The comparison table's own columns are Typed acc. / Handwritten acc. / Consistency, rated "High," "Good," "Variable" — qualitative impressions, not CER, not WER, not any numeric error rate at all. From that one document, PaddleOCR-VL-1.6 was declared the winner ("lowest VRAM; most consistent"), and the plan for the following week was to fine-tune it on that basis, while simultaneously starting semantic-search work (Ollama vector embeddings) — the exact long-term direction the June 17 meeting would shortly rule out of scope for the current phase.
Picking a production method from one document's qualitative read, while also beginning the search feature the eventual project charter explicitly deferred, is the speculative approach that had to be unwound.
The June 17, 2026 alignment meeting (Prof. Paul Israel, Partha Pediredla, Jason Ding, Xiaotian/Jack Zhou) falls right at the boundary between Week 3 and Week 4, and its minutes read as a direct response to what Weeks 1–3 had produced:
The two weeks that followed had no headline deliverable of their own because they were spent building the thing the meeting actually asked for: retiring the Next.js/Vercel/Gemini/Kraken direction and the one-document Amarel comparison, and constructing the real ground-truth set and multi-method benchmark harness that Week 6 then reports results from. edison-dataset, the benchmark repository, was initialized July 2 — the same week Week 6's deck was presented.
Recommendation carried over from the prior revision of this document still holds: don't present Weeks 4–5 as blank. They are the pivot itself — the moment the project stopped speculating and started measuring — and read better as one deliberate "Correcting Course" entry than as two missing weeks.
Slides: https://docs.google.com/presentation/d/1_VpE6FXx78GpOlZHxxGfiDhQhi0kvjWP65_7HL58xSg/edit?usp=sharing
The rigorous benchmark the meeting called for, built for real: 28 ground-truth documents from Volume 9 (sourced via IIIF, expert-transcribed, normalized text, 2 outliers later excluded for a French-language ground-truth/image mismatch), hand-tagged across 7 challenge categories (handwriting, typed, marginal note, crossout, long document, insertion, faded). 8 methods across 3 paradigms — docTR and PaddleOCR-VL locally, four cloud VLMs (Qwen3.6-27B/Flash, Llama-4-Scout, Mistral-Small-3.2) via API, and two hybrid pipelines chaining local OCR with a VLM refinement pass. Scored on CER and WER against the proofread ground truth, with a stated rule of thumb (CER < 0.10 excellent, 0.10–0.25 acceptable, > 0.40 poor).
The headline result — hybrid PaddleOCR-VL + Qwen best overall — is the same conclusion the project ultimately stands behind. What this deck's results table does not disclose is coverage: it presents CER/WER/cost for all 8 methods side by side with no indication that the two hybrid methods had completed noticeably fewer of the 26 valid documents than the others at this point. That gap wasn't caught here — it surfaces later, during final-presentation review, and is only fully closed in Week 10.
With a method chosen on Week 6's evidence, this week split the project into two repositories with distinct jobs: edison-dataset (piloting which transcription method works, already done) and edison-pipeline — a new repo, the subject of this and every following week's deck, building the production system that runs the winning method at scale.
Phase 1 scope, built and tested end to end against the live archive: given a CSV of Document IDs, fetch each one's IIIF manifest (both v2 and v3 shapes), download page images into bucket storage (local disk today, S3 by config change, not code change), and record everything in an idempotent SQLite table — re-running the same command skips already-fetched documents, and one failed page never fails the whole document. Explicitly out of scope for this phase: transcription itself, and both the "create new document" and "add images to existing document" workflows. Verified with a live run against document CL207AAA.
Slides: https://docs.google.com/presentation/d/1MEexiw6ZM3Bk5YmIU2u-X8Is6Ty0ROjCbWNyBTgz0sQ/edit?usp=sharing
The first of the platform's two end-to-end workflows shipped this week, all five stages, each its own CLI subcommand:
property[type]=nex on scripto:transcription, scoped to exclude Person/Notebook/Patent records) for untranscribed IDs, pulled in small --limit 20 batches by design, not a mass queue. 153,766 archival documents total; 149,619 without a transcription.[Letterhead], [Stamp], [File notation], [Marginal note]), deduplicates places, and returns strict JSON. A real sample run on CL207AAA (a 1903 letter) cost $0.00164.Orchestrated by a cost-capped, human-paced advance step — a hard --max-total-cost (default $2.00) checked before every document, all-time, not per run. Complete and tested: 42 passing tests. The benchmark recap slide in this deck still quotes the Week 6 numbers as-is, including the still-undisclosed coverage gap. Closing slide sets the next workflow — "Add New Document" — as next week's target.
Built the second workflow the Week 8 deck planned: upload a folder of scans that don't exist on the archive at all, propose document boundaries automatically, let the reviewer confirm or re-cut each one, then reuse the same transcribe/review/export stages the first workflow already had. Also the week the platform stopped being single-reviewer software:
Slides: https://docs.google.com/presentation/d/10RSJo7YtuCLLtUt7LOGRyuvrMtBex2SSvrvnrEk1o7o/edit?usp=sharing
The benchmark gap that opened in Week 6 closed this week — for real, not by re-averaging. The two hybrid methods were still silently dropping documents (ERROR: Empty response after 5 attempts, deterministic, same pages every retry). Reproducing one failing page and printing the full API response — not just its .content field — showed why: Qwen3.6-27B is a hybrid-reasoning model, and on a garbled OCR draft it was spending its entire token budget on internal reasoning and returning nothing (finish_reason: "length", reasoning_tokens: 4010, content: null). Raising the token budget to 12,000 didn't fix it — still empty. Disabling reasoning outright (reasoning: {"enabled": false}) did, at roughly 1/27th the cost per call. Every incomplete document was re-run from a clean cache, completed on a second machine, and copied back — every one of the 8 benchmarked methods now covers all 26 ground-truth documents, identically:
| Method | Type | Coverage | CER ↓ | WER ↓ | $/doc |
|---|---|---|---|---|---|
| PaddleOCR-VL → Qwen3.6-27B (hybrid) | Hybrid | 26/26 | 0.198 | 0.261 | 0.0095 |
| docTR → Qwen3.6-27B (hybrid) | Hybrid | 26/26 | 0.220 | 0.278 | 0.0101 |
| Qwen3.6-27B | VLM | 26/26 | 0.218 | 0.276 | 0.0228 |
| Qwen3.6-Flash | VLM | 26/26 | 0.235 | 0.299 | 0.0070 |
| Llama-4-Scout | VLM | 26/26 | 0.233 | 0.324 | 0.0005 |
| PaddleOCR-VL | OCR-VL | 26/26 | 0.244 | 0.320 | free |
| Mistral-Small-3.2 | VLM | 26/26 | 0.452 | 0.503 | 0.0003 |
| docTR | OCR | 26/26 | 0.516 | 0.823 | free |
The hybrid wins on genuinely equal footing this time — best overall, and best on 5 of 7 challenge categories (handwriting, typed, marginal notes, insertion, long documents) — closing the loop the project opened by accident in Week 3 and caught by accident in Week 6.
The rest of the week: