The present invention provides a computer-implemented method and
system for generating a Unified Script Code (USC) representation of multilingual text to enable consistent, script-neutral, and phonemically accurate encoding across diverse languages and writing systems The
system converts
Unicode-encoded text from one or more Indic scripts into a Devanagari-based
intermediate form using predefined or bitwise mapping. It then normalizes the text by inserting inherent vowels, converting dependent
vowel signs (matras) into independent vowels, and removing halant characters. The resulting USC representation explicitly encodes
consonant-
vowel sequences, reduces script specific variation, and preserves phonetic integrity. This approach improves tokenization efficiency, enhances performance in
natural language processing and
machine learning tasks, and supports reversible conversion to original scripts.