J. L. BELL is a Massachusetts writer who specializes in (among other things) the start of the American Revolution in and around Boston. He is particularly interested in the experiences of children in 1765-75. He has published scholarly papers and popular articles for both children and adults. He was consultant for an episode of History Detectives, and contributed to a display at Minute Man National Historic Park.

Subscribe thru Follow.it





•••••••••••••••••



Saturday, October 10, 2026

Improving Transcriptions of Handwritten Documents

Even as I muse on the value of A.I. programming for transcribing historical handwritten documents, the Fold3 service announces:
We’ve unlocked 2.4 million pages of Revolutionary War Pension Files with a state-of-the-art, patent-pending Transcription Viewer.
For paid subscribers, the service offers multiple ways to see the scanned documents and its transcriptions at the same time—that seems to be the patent-pending part. But beneath that function are the transcriptions themselves.

Fold3 launched “Full-Text Search” of the Revolutionary War Pension Files in March. However, its staff says, “Handwriting recognition technology has really advanced since” that time, so the service now provides “Dramatically improved line-by-line transcription quality.”

The improved transcriptions were produced, says the staff, “primarily with AI with lots of human interaction.” It’s not clear what form that human interaction took—staff hours, volunteers, feedback from users, advanced programming?

Human oversight still seems to be key. When I asked Dr. Jeffrey M. Griffith of the John Hancock Papers about his impression of the latest handwriting-recognition software, he said:
I have not used AI tools to transcribe the manuscripts, as I have found that widely-accessible large scale models like Chat GPT, Claude, and Grok are not as accurate as I desire. What’s even more troubling for my project in particular, is that the same prompt can result in different outputs. Most problematically, in some instances where I have tested the systems, if the system is uncertain, it just creates contents that may not even be relevant to the particular manuscript.

Knowing how specialized AI systems can get, I am sure there are some extremely powerful (and equally expensive to use) models continue to improve reading handwriting from various decades and centuries, but I have not explored those. Admittedly, even if an AI system is to be used for an initial transcription, the content needs to be thoroughly reviewed and compared to the original text.
Griffith judges the basic transcription tools available now as “much better than those very early efforts by Clements Library at University of Michigan.” Transkribus or a similar service might today produce better results from the Gage Papers. Conversely, if “lots of human interaction” is key to continual improvement, that site might benefit from providing a clear channel for user feedback on texts.

Because the profession of documentary editing puts such value on accuracy, people working in the field would naturally be wary of A.I.-produced transcriptions until they’re much more reliable. A text with errors could mislead readers. More likely, in a time of tight budgets, an institution might decide that a text with errors is good enough and not put funds toward professional-quality transcriptions.

All that said, I do use imperfect transcriptions like those on Google Books or the Gage Papers as my starting point when I quote documents. I did that with texts produced through O.C.R. programs, and I’m doing that with texts produced through A.I.

I find those imperfect databases easier to search, read, and quote than online archives without any transcriptions, like the New York Public Library’s scans of the Samuel Adams Papers and Boston Committee of Correspondence Papers.

So for me the bottom line is the more information that’s available in an online archive, the better. Scans better than mere lists of documents. Scans and bad transcriptions better than scans alone. Good, authoritative transcriptions without scans very useful, and good transcriptions with scans the best of all.

TOMORROW: Other ways of using A.I.

No comments: