Every page traces back to a shelf.

Provenance is the part of a corpus that is hardest to add afterwards, so we build it in from the first scan. This page sets out what ships with the data, how we arrive at our token figures, and the source behind every number we publish.

What ships with every corpus

Four things travel with the images themselves, in every delivery.

Item-level record

Title, author, imprint, place, date, shelf mark, holding institution and capture date — one record per item, delivered alongside the images as JSON Lines or MARC.

Page images retained

Archival masters in JPEG 2000 or TIFF are kept permanently, so higher-resolution derivatives can be issued years after delivery.

Named holding agreement

The agreement under which the material was digitized, naming the institution and the scope of use granted. Documented, not asserted.

Capture record

Resolution, format, color reference and capture date, for every page. Documented rather than assumed.

How we count

A shelf holds items, not tokens. Getting from one to the other means two adjustments, and we publish both rather than quoting a headline figure.

1,351,921
Arabic-language items identified
×
0.55
distinct-title ratio, for copies, editions and reprints
=
866,251
distinct volumes
×
92,000
tokens per volume, at tokenizer fertility 1.65
=
79.7 billion
tokens after transcription, deduplicated

Deduplication

A library holds multiple copies, editions and reprints of the same work. We apply a 0.55 distinct-title ratio, which is our largest open assumption — the honest band runs from 0.40 to 0.75. The catalog query that settles it will be published whichever way it goes.

Tokenizer fertility

Arabic fragments differently across tokenizers, so the same text can yield wildly different counts. We quote at a fertility of 1.65 and will recompute against your tokenizer on request, because a token price without a named tokenizer is not a price.

The record

Every figure we publish, with its source. Institutional holdings are self-reported and unaudited; we label them that way rather than rounding them into something cleaner.

Verified
Recomputed by us from the primary artifact.
Reported
Published by the institution. Self-reported and unaudited.
Estimate
Derived by us under the stated assumptions.

Arabic monographs, Egyptian National Library

ENLA official collections page. Present unchanged in archived captures since August 2019.

Scoped to the library's acquisitions department, and the unit does not distinguish title from physical item.

>1,052,600
Reported

Manuscript volumes

ENLA official: «51193 رقم حفظ = 59321 مجلدًا = 88164 عنوانًا» — 51,193 accession numbers, equalling 59,321 volumes, equalling 88,164 titles.

We publish the volume count, and say which unit we are using.

59,321
Reported

Arabic periodicals, bound

ENLA official, across more than 6,000 titles.

Published as a floor rather than an exact count.

240,000
Reported

Legal deposit intake, 2025

ENLA annual report, published 1 January 2026.

The report states some months were not yet tallied, so this is a floor.

7,218
Reported

Tokens per distinct volume

Weighted across four material classes at tokenizer fertility 1.65. Words per page constrained by 82 measured Arabic volumes — 23,226 pages, 214.5 words per page in aggregate — and an independent derivation at 187.5.

~92,000
Estimate

Distinct-title ratio

Central assumption, band 0.40 to 0.75.

Our largest open assumption. It carries a 1.9x spread, and the catalog query that settles it will be published whichever way it goes.

0.55
Estimate

Bibliotheca Alexandrina printed items

Bibliotheca Alexandrina official FAQ, English and Arabic. The figure covers books, manuscripts, maps, periodicals and theses together.

~2,000,000
Reported

al-Azhar rare manuscripts

Press reporting, 2013. al-Azhar publishes no current holdings statistics of its own.

~34,000
Reported

AUC print volumes

American University in Cairo, published library statistics.

566,949
Reported