About the text
Every issue of Pacific Islands Monthly held by the National
Library of Australia, 840 issues from September 1930 to 2000, rebuilt into a browsable,
searchable archive. Nothing here is newly written. It is the magazine's own pages,
re-presented. This note explains where the text comes from and what we changed.
Where the text comes from
The National Library digitised the print run and made it available through Trove as
record nla.obj-310385031. During
digitisation each scanned page was run through Optical Character Recognition (OCR), which
turns the image of a page into machine-readable text and groups it into articles. That OCR
is the Library's, produced when the magazine was digitised. We did not re-scan or re-read
the pages.
The per-issue text came from Trove via the
GLAM Workbench trove-periodicals
tools, and the per-page layout data (word positions, article grouping, advertisement zones)
from Trove's page OCR, which lets us reassemble articles in reading order and keep
advertisements out of the contents lists.
What we changed
Our changes are structural. We reshaped how the text reads without altering the words
themselves.
- Reflow The raw download breaks text at page edges rather than
paragraphs. We rejoin lines, mend words hyphenated across a line break, and re-form
paragraphs.
- Articles Each issue appears as its articles, in reading order, using
Trove's zoning, with advertisements collapsed out of the tables of contents.
- Headlines OCR sometimes mis-reads a headline or cuts it off mid-phrase.
We drop unreadable ones from the contents list, and where a headline was truncated
(“Influenza Epidemic at”) we complete it from the article's own opening line
(“Influenza Epidemic at Rabaul”).
- Names The people, places, organisations and ships in the
Names index were found automatically by a language model reading
the article text. Names appearing fewer than three times are left out of that index, though
they stay findable in the full-text search.
What we did not fix
- The characters themselves are still the original OCR, not hand-corrected.
- Around 29,000 pages were scanned with the right-hand edge cut off, concentrated in the
1960s.
- A few pages returned no text from Trove at all, and one 1951 issue is truncated at the
source. Affected issues carry a note.