How the transcriptions and the page scans relate to each other.
Page numbers are contiguous and zero-padded, 000 to
223, and every page exists in every form. For page N:
| Scan, original | /images/pages/original/medres/diaryNNN.jpg — see
metadata.json for the exact per-resolution paths |
|---|---|
| Scan, transcribed | /images/pages/transcribed/medres/diaryNNN.jpg |
| Transcript, HTML | /text/HTML/diaryNNN.html |
| Transcript, text | /text/txt/diaryNNN.txt |
| This site | /diary/page/NNN/ |
A page's transcription is a sequence of sections. Section anchors are
#section-1, #section-2 and so on, numbered in document
order. They are generated deterministically from the stored transcript, so a
given input always yields the same anchors.
Each stored page transcript is a single block of text with no line structure,
so a #line-12 style anchor would have to be invented. Invented line
numbers are unstable: any correction to a transcription would silently move every
anchor after it, breaking existing citations. Anchoring per section is the finest
granularity the source data actually supports.
19 pages are blank. They are given URLs, are listed in the sitemap, and render their scan with an explicit note, so the corpus does not look smaller than it is.