Decode a script
Paste text in any writing system. ROSETTA identifies the script from its Unicode signature, transliterates it into the standard scholarly romanisation, and — where the language is one of the twelve ancient tongues in its lexicon — glosses it word by word.
Write in an ancient script
Type romanised text and render it in the sign inventory of a historical writing system. This is transliteration in reverse — it spells your sounds with their signs. It is not translation: writing "king" in cuneiform gives you the sounds k-i-n-g in wedges, not the Sumerian word LUGAL.
Translate
Two different machines. Living languages go to an online translation service (about 130 languages; needs internet). Ancient languages are glossed offline from ROSETTA's own lexicons — word-for-word, with no grammar, which is exactly what an interlinear gloss is.
Reverse lexicon
Search in English; see how twelve ancient languages said it. Useful for checking whether an idea even existed as a single word — Sumerian has no word for "privacy", Latin has no word for "yes".
Proto-Indo-European roots
The words below were never written down by anyone. They are reconstructions — inferred backwards from systematic sound correspondences across the daughter languages. When Latin has p where English has f in word after word (pater/father, piscis/fish, pes/foot), that regularity is the evidence. Grimm's Law, 1822.
Atlas of writing systems
Every script ROSETTA knows, with its sign chart where one exists. Writing has been independently invented perhaps four times in human history — Mesopotamia, Egypt (arguably), China, and Mesoamerica. Everything else on this page is a descendant or a deliberate imitation.
Language families of Earth
Roughly 7,150 languages are spoken today. They fall into about 140 families, plus ~130 isolates with no known relatives at all. Around 40% are endangered; a language dies roughly every three months. The families below cover the overwhelming majority of living speakers.
The unread
Scripts we cannot read, and why. Decipherment needs one of three things: a bilingual text, a known language written in an unknown script, or a known script writing an unknown language. Where none is available, the wall is usually permanent.
Numerals across civilisations
How twelve systems wrote the same quantity. Only two cultures independently invented a true positional zero: the Maya and India. Babylon had a placeholder but never a number.
What this is, and what it honestly is not
So this is built around what is actually tractable, and it does that part properly:
- Script decoding is solvable, and it is done here in full. Given text in a known writing system, converting signs to sounds is a deterministic mapping. ROSETTA carries real transliteration tables for — writing systems, from Egyptian hieroglyphs and Sumerian cuneiform to Linear B, Brahmi, Ogham and Cherokee, in the standard scholarly romanisations (IAST, Gardiner, Ventris–Chadwick, ISO 259, DMG).
- Structural engines, not just lookup tables. Ten Brahmic scripts share one abugida engine that handles inherent vowels and virama. Geʽez is generated from its consonant × vowel grid. Hangul is decomposed algorithmically into jamo. These are implementations of how the scripts actually work.
- Ancient-language glossing is honest word-for-word. Twelve curated lexicons (— headwords) cover Latin, Ancient Greek, Sanskrit, Old English, Old Norse, Gothic, Sumerian, Akkadian, Middle Egyptian, Hittite, Old Church Slavonic and Biblical Hebrew. It does light stem-stripping, then tells you exactly which words it could not find. It does not parse grammar and does not pretend to.
- Living languages go online. ~130 languages via public translation services. This is the one module that needs a network.
- The unread are documented as unread. With the corpus size, the sign count, and the specific reason each one resists.
Things that will be wrong
- Cuneiform signs are polyvalent. 𒀭 is AN "sky", DINGIR "god", or the syllable /an/, and it also silently marks the next word as divine. Any single reading shown here is one of several. This module identifies signs; it does not disambiguate them.
- Egyptian has no vowels. "nfr" is conventionally read "nefer", but the actual vowels are unrecoverable except through Coptic and Greek transcriptions.
- Abjads drop short vowels. Hebrew, Arabic, Phoenician and Syriac transliterate consonants; vowel points are stripped, as they are in ordinary writing.
- Right-to-left scripts are transliterated in logical (memory) order, not visual order.
- Chinese covers the ~750 commonest characters. Beyond that you get [字?].
- Thai, Khmer, Burmese, Tibetan and Sinhala use character-level approximations, not their full orthographic rules.
- Proto-Indo-European is reconstructed, not attested. Every asterisked form is an inference, not a record.
If a script shows as boxes
Your system lacks the font. Windows ships Segoe UI Historic, which covers most ancient blocks — it is first in the font stack here. For full coverage install Google's Noto historical families. The transliteration still works regardless of whether the glyphs render.
Sources for the tables
Unicode Standard code charts; Gardiner's Egyptian Grammar sign list; Ventris & Chadwick for Linear B; Borger / ePSD conventions for cuneiform; IAST for Indic; ISO 259 (Hebrew), DMG/ALA-LC (Arabic), Hübschmann–Meillet (Armenian), Wylie (Tibetan), Hepburn (Japanese), Revised Romanisation (Korean), ISO 9 (Cyrillic).