Historic Hood River
OCR and the Oregon Trail

Notes
Last Monday Morganne shared the first page from Polly Coon’s manuscript/ diary describing her journey across the Great Plains to the Pacific Coast, and many of you wanted to see more. Morganne found this old transcription of the journal. Some modern technology lets me share the entire document with you in a searchable form. Since “Morganne’s Mondays” are all about the behind the scenes activities at the museum, let’s dive into how this is done. If you’re not interested in that, just click here and enjoy the journey.
The first step is to use the Museum’s scanning photocopier, which despite having trouble copying a single page without jamming has no problem scanning an entire manuscript like this and emailing it anywhere in the world in well under a minute.
The next step is to use optical character recognition (OCR) software to convert the scanned image into text, and then embed that OCR’ed text in the original scanned image so you can search for text as well as select and copy text right out of the image. OCR software gets better every year. The paid version of Adobe Acrobat includes a good OCR engine, but I chose to use an open software product called Tesseract for this project.
Tesseract comes bundled in the open software product OCRmyPDF. While this software requires a bit of computer expertise (if “github” means nothing to you, don’t bother), OCRmyPDF offers the ability to steer the OCR engine to deal with things like multiple column manuscripts, skewed originals, and “dirty paper” documents. All this means it only took a minute or two to finish a roughly 99% accurate version of this transcription. Its only real problems were with punctuation. The typewriter which was used for this transcription had a misadjusted shift carriage, which is why all the capital letters are “flying” a little, along with the periods at the end of the sentences. The OCR engine has no problem with the capitals, but interprets many of the periods as hyphens.
By coincidence I rebuilt a 1913 Underwood Model 5 typewriter with a similar problem just last year. There is a simple but delicate adjustment to correct this problem– but unfortunately the OCR engine doesn’t (yet) have such an adjustment. Funny thing is when people want to make a document look old they insert flying capitals in them, since it was such a common misadjustment a typed document doesn’t look old without it.
Wondering about doing OCR directly on handwritten documents? This is another technology which gets better every year. iPhones have a built in engine to do this, and while it is not nearly as accurate as the OCR we used for this post, it can save some time. Today you’ll need to manually verify the results, but this will probably improve with every release.


Jeffrey Bryant
Thank you for the entire journal. It was interesting reading. My Liberty Vaughan family also came to Oregon in 1852. His son, Cyrus Vaughan, born 1844, moved to Hood River in 1902, and died there in 1917.
L. E.
Thank you for all the advice on scanning. I got really excited about the OCR software but I know nothing about github, so I guess I will forget that.
I have been reading the diary of Missionary Cyrus Shepard who traveled the Oregon Trail in 1834.
Thank goodness, someone transcribed the hand written diary. It makes it so much more enjoyable to read.
Arthur Babitz
LE, for most uses commercial software like Adobe Acrobat will provide good OCR results. It probably would have been fine for this document, but we’ll see some documents next week which require more control of the OCR engine.