arXiv:2609.14967cs.CL6 pagesAnnounced September 2026

A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator

Abstract

Quranic text is distributed in two orthographic forms that are byte-level distinct: the Uthmani script used in every printed mushaf, and the Standard (Imla’i) Arabic form that every mainstream Arabic NLP tool is built for. The gap is concentrated in one Unicode character, U+0670 (superscript alef), which appears in some of the most frequently recited words in the Quran and is silently mishandled by general-purpose Arabic normalizers. We release a 2,290-pair, corpus-aligned Uthmani-to-Standard word mapping constructed by aligning the complete 6,236-verse Quran across both orthographic forms, together with a seven-step text normalization pipeline built on it. Normalizing both forms of all 6,236 verses through that pipeline yields identical strings for 90.9% of verses, and we characterize the residual divergence rather than assert that it is closed. On top of the normalized text, we build a deterministic, LLM-free Quranic recitation validator using a four-layer verse-matching search (exact, morphological, relaxed, fuzzy) and word-error-rate-graded feedback across five severity tiers. The validator scores 98.4% (122/124) on a 124-case suite emitted by the released test harness, and both failures share one mechanism: a single substitution error can make a different verse an exact match. A full-corpus census additionally quantifies an inherent text-only ambiguity affecting 16.5% of verses, and on 34 recitation transcripts drawn from a deployed Arabic ASR system the validator identifies the correct verse in every case. We release the mapping, the script that builds it, the validator, and the evaluation harness under open licenses; every number in this paper except the deployment measurement, whose transcripts are not ours to publish, is reproduced by running them.

In numbers

Uthmani ↔ Standard word pairs
2,290
Verses aligned — the full Quran
6,236
Verses identical after normalization
90.9%
Validator accuracy (122 / 124)
98.4%
Verses with inherent text-only ambiguity
16.5%
Real ASR transcripts, correct verse
34 / 34

Contributions

  1. A 2,290-pair, corpus-aligned Uthmani-to-Standard word mapping, built by aligning the complete Quran across both orthographic forms.
  2. A deterministic recitation validator on top of it: a seven-step normalization pipeline and a four-layer verse search, with no language model anywhere in the pipeline.
  3. A released evaluation harness, so every number in the paper except the deployment measurement is reproduced by running the repository.

Released artifacts

  • Code & datamuslim-quran-validatorThe mapping, the script that rebuilds it, the validator, and the 124-case evaluation harness.

Data CC BY 4.0 · code MIT

Cite

Yahya Mohamed Elnawasany. A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator. arXiv:2609.14967 [cs.CL], 2026. doi.org/10.48550/arXiv.2609.14967

BibTeX
@misc{elnawasany2026uthmani,
  title         = {A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping
                   and a Deterministic Recitation Validator},
  author        = {Elnawasany, Yahya Mohamed},
  year          = {2026},
  eprint        = {2609.14967},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  doi           = {10.48550/arXiv.2609.14967},
  url           = {https://arxiv.org/abs/2609.14967}
}