Every page traces back to a shelf.
Provenance is the part of a corpus that is hardest to add afterwards, so we build it in from the first scan. This page sets out what ships with the data, how we arrive at our token figures, and the source behind every number we publish.
What ships with every corpus
Four things travel with the images themselves, in every delivery.
Item-level record
Title, author, imprint, place, date, shelf mark, holding institution and capture date — one record per item, delivered alongside the images as JSON Lines or MARC.
Page images retained
Archival masters in JPEG 2000 or TIFF are kept permanently, so higher-resolution derivatives can be issued years after delivery.
Named holding agreement
The agreement under which the material was digitized, naming the institution and the scope of use granted. Documented, not asserted.
Capture record
Resolution, format, color reference and capture date, for every page. Documented rather than assumed.
How we count
A shelf holds items, not tokens. Getting from one to the other means two adjustments, and we publish both rather than quoting a headline figure.
Deduplication
A library holds multiple copies, editions and reprints of the same work. We apply a 0.55 distinct-title ratio, which is our largest open assumption — the honest band runs from 0.40 to 0.75. The catalog query that settles it will be published whichever way it goes.
Tokenizer fertility
Arabic fragments differently across tokenizers, so the same text can yield wildly different counts. We quote at a fertility of 1.65 and will recompute against your tokenizer on request, because a token price without a named tokenizer is not a price.
The record
Every figure we publish, with its source. Institutional holdings are self-reported and unaudited; we label them that way rather than rounding them into something cleaner.
- Verified
- Recomputed by us from the primary artifact.
- Reported
- Published by the institution. Self-reported and unaudited.
- Estimate
- Derived by us under the stated assumptions.
Arabic monographs, Egyptian National Library
ENLA official collections page. Present unchanged in archived captures since August 2019.
Scoped to the library's acquisitions department, and the unit does not distinguish title from physical item.
Manuscript volumes
ENLA official: «51193 رقم حفظ = 59321 مجلدًا = 88164 عنوانًا» — 51,193 accession numbers, equalling 59,321 volumes, equalling 88,164 titles.
We publish the volume count, and say which unit we are using.
Arabic periodicals, bound
ENLA official, across more than 6,000 titles.
Published as a floor rather than an exact count.
Legal deposit intake, 2025
ENLA annual report, published 1 January 2026.
The report states some months were not yet tallied, so this is a floor.
Tokens per distinct volume
Weighted across four material classes at tokenizer fertility 1.65. Words per page constrained by 82 measured Arabic volumes — 23,226 pages, 214.5 words per page in aggregate — and an independent derivation at 187.5.
Distinct-title ratio
Central assumption, band 0.40 to 0.75.
Our largest open assumption. It carries a 1.9x spread, and the catalog query that settles it will be published whichever way it goes.
Bibliotheca Alexandrina printed items
Bibliotheca Alexandrina official FAQ, English and Arabic. The figure covers books, manuscripts, maps, periodicals and theses together.
al-Azhar rare manuscripts
Press reporting, 2013. al-Azhar publishes no current holdings statistics of its own.
AUC print volumes
American University in Cairo, published library statistics.