Skip to the content
[ ]

LSPHIL

LSPHIL › Writing systems & orthography › The Hidden Barrier Between a Script and Your Keyboard

The Hidden Barrier Between a Script and Your Keyboard

A community can use a writing system for centuries and still find it untypeable. Unicode, the standard that governs how text is stored and exchanged, does not automatically include every script that exists. The Unicode Consortium's Script Encoding Working Group sets three rigid conditions before a script can even be proposed: proven community usage, stability of the writing system, and a demonstrated need for interchange in plain text. These criteria explain why scholars, developers, and ordinary users sometimes encounter languages that remain outside the digital text ecosystem entirely.

English · 875 words

Why Some Scripts Cannot Be Exchanged as Text

When a script lacks Unicode encoding, it cannot function as ordinary text. Operating systems, search engines, and databases treat unencoded characters as binary data or images, not as searchable, sortable, comparable text. The practical consequences cascade through every layer of digital infrastructure.

ASCII Code Chart-Quick ref card
ASCII Code Chart-Quick ref card. an unknown officer or employee of the United States Government, Namazu-tron (scanned) · Public domain · Wikimedia Commons

Normalization compounds the problem. Unicode Standard Annex #15 specifies that equivalent text must produce identical binary representations under standard normalization forms, yet semantically equivalent strings can use different code point sequences. The W3C's Character Model for the World Wide Web: String Matching advises applications to normalize before comparison, converting strings to Unicode code points first, then performing normalization, then checking for identity. Without encoded characters, none of this processing applies. A script outside Unicode cannot reliably match against itself, let alone integrate with other text.

This creates a threshold effect. Below it, a script survives as photographs, PDFs, or proprietary font hacks. Above it, the script becomes data. The gap between these states is not technical limitation but bureaucratic process.

The Evidence Burden for New Scripts

Proposing a script requires documentation that would satisfy an academic peer review. The UTC Script Encoding Working Group demands a short summary, prior related documents, and an introduction citing modern sources. Proposal templates ask for comparison to visually similar existing characters, suggested character properties, preferred ordering, and concrete examples of each proposed character in printed text.

The examples matter. They establish that these shapes need digital interchange, not merely that they exist. Templates also require script type, directionality, combining diacritics, phonetic values, and a proposed name and glyph for each character. For entire scripts, proposals must cite modern, definitive sources. Even dead or obsolete scripts, if already partially encoded, require citation of the most important modern sources for any additions.

The proposal must also demonstrate that each shape is already a character by Unicode's definition and does not already exist in the Standard. This last requirement catches many hopeful submitters who discover their "new" script overlaps with existing encoded blocks or unified characters.

From Submission to Decision

The formal process begins only after paperwork. Every submitter must sign a Unicode Contributor License Agreement before creating a submission, and the same agreement extends to co-authors and any entity claiming intellectual property rights in the proposed material. This legal prerequisite filters out casual or conflicting claims before technical review even starts.

Once submitted, a working group member reviews the proposal. The full Script Encoding Working Group then discusses it. Only proposals deemed "mature" and ready for UTC review typically appear in the Unicode document register. The working group's recommendations carry weight but remain non-binding; the Unicode Technical Committee makes final decisions.

This two-tier structure means delays at multiple points: initial completeness, working group consensus, maturity designation, and UTC scheduling. A script can spend years in this pipeline while its community continues analog use.

Why Matching Requires More Than Code Points

Normalization resolves a problem that pure encoding does not. The same visual text can be constructed through different sequences: a base character plus combining mark versus a precomposed character, for instance. The W3C matching procedure explicitly converts strings to code points, normalizes, then compares. Without this step, databases return false negatives, passwords fail to match, and search engines miss relevant documents.

The requirement that equivalent text produce identical binary representations under normalization forms is what makes encoded text reliable. Unencoded scripts bypass this entire mechanism, remaining visually legible but computationally opaque.

Checking a Script or Character Yourself

Readers can verify encoding status through several public resources. The Unicode code charts, published as Standard Annexes, list every assigned character by block and script. The Unicode Character Database provides machine-readable data on properties and assignments. For scripts in progress, the document register contains mature proposals awaiting UTC review, and rejected or withdrawn submissions remain accessible as historical records.

If a character appears nowhere in these sources, it is unencoded. If it appears only in a proposal document, it is under consideration. If it appears in the code charts, it is assigned but may require attention to normalization form. The distinction between "in a font" and "in the Standard" matters: private-use slots and custom fonts can display shapes that remain unencoded, creating false confidence.

The Persistence of Visual Equality

Two strings can look identical, encode differently, and fail to match. This remains true even for encoded scripts when normalization is skipped. For unencoded scripts, the problem is absolute: no encoding means no normalization, no matching, no reliable text processing. The committee review that grants a script Unicode membership is what moves it from visual artifact to manipulable data. Until that review concludes, the script exists in the world but not in the systems that handle text.