Indus script
Thousands of inscriptions, none longer than a line, in a language no one knows — and a standing debate over whether it is writing at all.
Curatorial note
The Indus script is the great short-text problem. The Indus Valley Civilisation left thousands of inscribed seals, tablets and objects, but almost every text is tiny — on the order of five signs, the longest just over a dozen on one line. Statistical decipherment needs length and variety; here there is neither, and the underlying language is unknown with no bilingual to anchor it.
There is even a genuine scholarly dispute — argued most sharply by Farmer, Sproat and Witzel (2004) — over whether the signs encode language at all, or are a non-linguistic system of emblems. The institute records that dispute without resolving it.
Four thousand inscriptions, and not one sentence among them.
And because no code points exist, we cannot even show you the signs here: the box above is the tofu you meet in Floor 4's Tofu Room. The case is filed, unopened, in the Undeciphered Vault.
Why it resists
A decipherment needs three thingsWhat survives
The corpusRoughly 4,000–5,000 inscribed objects are known, catalogued and re-catalogued since the site of Harappa was first excavated in 1924. The sign inventory is unsettled — commonly cited counts run from about 419 (Mahadevan) through 386 (Parpola) to 694 (Wells), i.e. somewhere in the 400–700 range depending on where you draw the line between a distinct sign and a variant.
The texts themselves are minute: the average inscription runs to about five signs, the longest reaching only ~17 on a single continuous surface (seal M-314), or ~26 if you add up the three separate faces of one moulded terracotta prism. Most inscriptions sit on small stamp seals, set alongside an animal motif — a unicorn-like bovine, an elephant, a tiger — from the main sites of Mohenjo-daro and Harappa.
In the building
In the Reading Room
Is It Even Writing?
Thousands of seals, texts five signs long, a language no one knows — and a serious argument that the Indus signs may not encode speech at all.
Plaque sources: The Unicode Standard — the Indus script is not encoded (no code points; M. Everson's 1999 proposal remains unaccepted) · ISO 15924 assigns the code Inds · 610 (registered 2004-05-01; a code, not characters) · A. Parpola and I. Mahadevan on the corpus and sign counts (Mahadevan 419; Parpola 386; Wells 694; ~4,000–5,000 objects; average ~5 signs; longest ~17 on one line, M-314; ~26 across a prism's faces) · S. Farmer, R. Sproat & M. Witzel, "The Collapse of the Indus-Script Thesis" (EJVS, 2004) for the non-linguistic argument · ScriptSource (SIL). Sources: unicode.org/iso15924/iso15924-codes.html · safarmer.com/indus-longestinscription · scriptsource.org (Inds). Figures are approximate and the interpretation is contested. Compiled 2026-07.