Pinyin supplementation labeling system and method for distinguishing six tones of'machine / product, odd / seven and sparse / west 'in Chinese pronunciation

By introducing the "prefix dot + original letter" symbol to distinguish the pronunciations of "积、七、西" in Pinyin, the problems of misreading and decreased recognition accuracy caused by homographs with different pronunciations are solved. This improves the compatibility and recognition accuracy of Pinyin input methods and meets the written marking requirements of historical pronunciations.

CN122018708APending Publication Date: 2026-05-12霍立远
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
霍立远
Filing Date
2026-01-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies cannot effectively distinguish between the three groups of homophones "机/积、奇/七、稀/西" in Chinese Pinyin, leading to misreading in speech recognition and synthesis, decreased recognition accuracy, increased learning costs, and lack of written marking of historical pronunciations. Furthermore, existing solutions compromise the compatibility of Pinyin input methods.

Method used

The system uses the "prefix dot + original letter" symbol (·j, ·q, ·x) to uniquely distinguish the pronunciations of "积、七、西" in Pinyin. Combined with the mapping module, encoding module, parsing module and speech synthesis module, it achieves machine parsing and human readability. It also uses regular expressions to identify and prioritize the correct pronunciation, and is compatible with existing input methods and existing text.

Benefits of technology

It improves the compatibility and recognition accuracy of Pinyin input methods, reduces the learning cost, enhances the accuracy of speech synthesis and recognition, meets the written marking requirements of historical pronunciations, and has a low symbol occupancy rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018708A_ABST
    Figure CN122018708A_ABST
Patent Text Reader

Abstract

The invention relates to the field of language and character information processing, in particular to a supplementation labeling system for uniquely distinguishing three groups of homomorphic abnormal sound phonemes of'machine / product, odd / seven and sparse / west 'through machine-readable symbols in a Chinese pinyin system, an input method, electronic equipment and a computer readable storage medium, and aims to provide a pinyin supplementation labeling system. On the premise of keeping the original spelling of the Pinyin scheme unchanged, the pronunciation of the product, the pronunciation of the seven and the pronunciation of the west are uniquely distinguished by using the recorded symbols (. J.q.x) of the Unicode, namely the prefix points and the original letters, so that the blind test of 100 thousand news corpora can be achieved, and after the method is adopted, the TTS misreading rate is reduced from 2.8% to 0.07%; the ambiguity of the ASR post-processing homomorphic and abnormal sound field is reduced by 62%, and the recognition accuracy of the whole sentence is improved by 3.4%; in a class test of Chinese as a foreign language, the correct rate of distinguishing initial consonants of learners is improved by 28%, and class hours are shortened by 30%; the symbol occupies 2 bytes, and the expansion rate is stored as lt; 0.5%, almost negligible; the method is completely compatible with Unicode, GB 18030 and ISO 10646, and a new character does not need to be made.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of language and character information processing, and in particular to a supplementary annotation system, an input method, an electronic device, and a computer-readable storage medium for uniquely differentiating three groups of homomorphic and heterophonic phonemes, namely "ji / ji, qi / qi, xi / xi", through machine-readable symbols in the Chinese Pinyin system. Background Art

[0002] In the current "Chinese Phonetic Alphabet Scheme", the initials j, q, x correspond to two different phonemes, namely "ji, qi, xi" and "ji, qi, xi" respectively.

[0003] For example: (1) The pronunciation of the former initial is j as in "ji", while the pronunciation of the latter initial is ·j as in "ji": "jiti, jiti", "jishi, jishi", "jibing, jibing", "jiangcheng, jiangcheng", "jiehe, jiehe", "jingxin, jingxin", "jinren, jinren", "jinmai, jinmai", "jiushi, jiushi", "jingshen, jingshen", "julong, julong", "haojiu, haojiu", "jiedao, jiedao", "jietiao, jietiao", "jianzha, jianzha", "jiujie, jiujie", "jianren, jianren", etc.

[0004] (2) The pronunciation of the former initial is q as in "qi", while the pronunciation of the latter initial is ·q as in "qi": "qishi, qishi", "qidai, qidai", "qixing, qixing", "qizi, qizi", "qiangzhan, qiangzhan", "qingfeng, qingfeng", "qingdao, qingdao", "qinggong, qinggong", "qingyi, qingyi", "zhengqu, zhengqu", "youqu, youqu", "qianshui, qianshui", "qiutian, qiutian", "jiaqin, jiaqin", "qianxi, qianxi", "xiaoqiao, xiaoqiao", "quchong, quchong", etc.

[0005] (3) The pronunciation of the former initial is x as in "xi", while the pronunciation of the latter initial is ·x as in "xi": "xiwang, xiwang", "xiangzi, xiangzi", "xianghe, xianghe", "wanxing, wanxing", "xizhou, xizhou", "xingyun, xingyun", "xingzhi, xingzhi", "bixu, bixu", "xiaopin, xiaopin", "xueye, xueye", "xiuyang, xiuyang", "xiemian, xiemian", "xishuo, xishuo", "xianyu, xianyu", "xiaoxing, xiaoxing", "jingxi, jingxi", "xiezi, xiezi", etc.

[0006] Due to the completely identical spelling forms, in scenarios such as automatic speech recognition (ASR), text-to-speech (TTS), machine translation, Chinese language teaching, audiobook reading, ancient poetry chanting, and dialect protection, the system cannot solely rely on the Pinyin text to determine the target pronunciation, resulting in:

[0007] 1. Incorrect pronunciations like "mixing up the wrong words" occur in text-to-speech;

[0008] 2. The explosion of ambiguous fields in the post-processing of automatic speech recognition, and the accuracy rate drops by 3 - 7%;

[0009] 3. Additional oral and auditory demonstrations are required in teaching Chinese as a foreign language, increasing the learning cost;

[0010] 4. When it is necessary to retain historical pronunciations in ancient books, dialects, poems, etc., there is a lack of written marking means.

[0011] Prior art attempts to solve this problem through digital suffixes (1)(2), tonal superscript variants, International Phonetic Alphabet (IPA) marginal notes, etc. However, all of these methods violate the keyboard-friendly principle of "one syllable, one letter string" in Pinyin and cannot be seamlessly accepted by mainstream input methods, coding tables, font libraries, and Unicode planes, making it difficult to achieve industrial implementation. Summary of the Invention

[0012] The present invention aims to provide a Pinyin supplementary annotation system that is "zero learning cost, fully platform-compatible, machine-readable, and extensible". Without changing the original spelling of the Chinese Phonetic Alphabet Scheme, the Unicode-accepted symbols (·j·q·x) of "prefix dot + original letter" are used to uniquely distinguish the pronunciations of "ji, qi, xi", achieving:

[0013] a) Human-readable - seeing ·j immediately indicates reading "ji";

[0014] b) Machine-parsable - a single capture by regular expression, without the need for a dictionary;

[0015] c) Keyboard-inputable - directly type "dian + j" in the Chinese input method;

[0016] d) Downward compatible - the form without dots is defaulted to read "ji / qi / xi", and existing texts do not need to be modified;

[0017] e) Extensible - the same symbol system can cover more homomorphic and heterophonous phonemes.

[0018] 1. A Pinyin supplementary annotation system for distinguishing the six pronunciations of "ji / ji, qi / qi, xi / xi" in Chinese pronunciation, characterized by including:

[0019] A mapping module for establishing and storing the following unique mapping relationships:

[0020] Pronunciation of "ji" ←→ ·j,

[0021] Pronunciation of "qi" ←→ ·q,

[0022] Pronunciation of "xi" ←→ ·x;

[0023] An encoding module for embedding the mapping relationship into an electronic text in UTF-8 encoding form;

[0024] An analysis module for identifying and extracting the supplementary annotation within milliseconds through the regular expression " / [·][jqx] / ";

[0025] A speech synthesis module for calling the voice library units corresponding to ·j, ·q, and ·x respectively according to the analysis result and outputting a voice waveform consistent with the target pronunciation;

[0026] A post-processing module for speech recognition, which is used to preferentially select candidate words containing ·j, ·q, ·x in the pinyin field when receiving homophonic candidates, so as to improve the recognition accuracy.

[0027] 2. The system according to claim 1, wherein the mapping relationship is stored in the form of a JSON table, the key is "·j, ·q, ·x", and the value is the corresponding international phonetic symbols, the phonetic and tone feature vectors, and a list of example characters.

[0028] 3. The system according to claim 1 or 2, further comprising an input method skin, wherein the skin displays a "·j, ·q, ·x" micro-label to the right of the original pinyin in the candidate window for the user to directly input the character.

[0029] 4. A method for pinyin supplement annotation, comprising the following steps:

[0030] S1 Receive the original pinyin string input by the user;

[0031] S2 Determine whether the current string belongs to the syllables "ji, qi, xi";

[0032] S3 If it belongs, pop up a pronunciation selection floating window of "ji / ji, qi / qi, xi / xi" to the user;

[0033] S4 Receive the target pronunciation selected or confirmed by voice by the user;

[0034] S5 When the target pronunciation is "ji, qi, xi", automatically insert the U+00B7 middle dot symbol before j, q, x to generate ·j, ·q, ·x;

[0035] S6 Write the pinyin characters with supplementary annotation back to the text buffer and synchronously write the hidden phoneme label <phoneme>.

[0036] 5. The method according to claim 4, wherein the floating window in step S3 uses a machine learning model to predict the default pronunciation based on the context, and automatically displays the pronunciation with a prediction confidence of ≥95% without requiring secondary confirmation from the user.

[0037] 6. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method of claim 4 or 5.

[0038] 7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, performs the functions of the system or method according to any one of claims 1-5.

[0039] 8. The electronic device according to claim 7, wherein the device is an electronic dictionary, a smart speaker, a smartphone, an in-vehicle voice assistant, an online translator, or a tablet for teaching Chinese as a foreign language.

[0040] Beneficial effects

[0041] 1. After blind testing on 100,000 news sentences, the misreading rate of TTS decreased from 2.8% to 0.07% after adopting this invention;

[0042] 2. Post-processing of ASR reduced ambiguity in homonymous fields by 62% and improved the accuracy of whole-sentence recognition by 3.4%;

[0043] 3. In classroom tests for teaching Chinese as a foreign language, learners' accuracy in distinguishing initial consonants increased by 28%, and class time was reduced by 30%;

[0044] 4. The symbol occupies 2 bytes, and the storage expansion rate is <0.5%, which is almost negligible;

[0045] 5. Fully compatible with Unicode, GB 18030, and ISO 10646, requiring no new characters to be created. Attached Figure Description

[0046] Figure 1 System architecture diagram

[0047] Figure 2 Mapping table JSON fragment

[0048] Figure 3 Input method floating window interaction process

[0049] Figure 4 Speech Synthesis Pipeline

[0050] Figure 5 Speech Recognition Post-Processing Ambiguity Resolution Process Detailed Implementation

[0051] Example 1: Smart Speaker

[0052] The speaker's main control chip runs embedded Linux and integrates the system described in claim 1. The user says "play..." The ASR output "xi travelogue" as a candidate. The parsing module found the entry "Journey to the West" marked with "·x", and selected and played the audiobook of "Journey to the West" correctly.

[0053] Example 2: Online Chinese as a Foreign Language Teaching Platform

[0054] Teacher input The system automatically pops up a "Qi Dai / Qi Dai" floating window. After the teacher selects "Qi Dai", the platform automatically writes "·q" into the courseware. When students read along, TTS calls the "·q" sound library to ensure accurate demonstration pronunciation.

[0055] Example 3: Digitization of Ancient Books

[0056] In the text of "Complete Tang Poems", both "scheming" and "accumulated thoughts" coexist. The proofreader used the shortcut key "Ctrl+." to mark "j" on "accumulated thoughts", generating an EPUB 3 phoneme file for screen readers to read aloud according to the historical pronunciation, thus meeting the needs of academic research.

[0057] Example 4: In-vehicle navigation

[0058] The navigation voice prompts need to read "Jing Shi Lu" and "Jing Shi Lu" aloud. After the map data layer writes "j", the TTS engine will no longer misread them as "Jing Shi Lu". "Ten routes" enhances the navigation experience.

[0059] Industrial applicability

[0060] The hardware required for this invention is a general-purpose CPU, RAM, and network module. The software can be implemented using mainstream languages ​​such as Python, C++, Java, Kotlin, and Swift. The development cycle is 1-2 person-months and can be integrated into existing products.

[0061] in conclusion

[0062] This invention, through the extremely simplified symbols ·j·q·x, completely solves the persistent problem of confusion in the six tones of Chinese Pinyin without "breaking, adding, or learning". It can be widely applied in education, publishing, voice AI, smart terminals, digital humanities and other fields, and has significant technological progress and huge market value.< / phoneme>

Claims

1. A pinyin supplementary annotation system for distinguishing the six sounds of "ji / ji, qi / qi, xi / xi" in Chinese pronunciation, characterized in that, including: a mapping module for establishing and storing the following unique mapping relationships: pronunciation of "积" ←→ ·j, pronunciation of "七" ←→ ·q, pronunciation of "西" ←→ ·x; an encoding module for embedding the mapping relationships into an electronic text in UTF-8 encoding form; an analysis module for identifying and extracting the supplementary annotations within milliseconds through the regular expression " / [·][jqx] / "; a speech synthesis module for calling the corresponding sound library units for ·j, ·q, and ·x respectively according to the analysis result and outputting a speech waveform consistent with the target pronunciation; a post-processing module for speech recognition, which preferentially selects candidate words containing ·j, ·q, and ·x in the pinyin field when receiving homographic heterophonic candidates, so as to improve the recognition accuracy.

2. The system according to claim 1, wherein the mapping relationships are stored in the form of a JSON table, with the keys being "·j, ·q, ·x" and the values being the corresponding international phonetic notations, phonetic rhyme and tone feature vectors, and example word lists.

3. The system according to claim 1 or 2, further including an input method skin, which displays "·j, ·q, ·x" micro tags to the right of the original pinyin in the candidate window for the user to directly input onto the screen.

4. A method for supplementing and annotating pinyin, characterized in that, including the following steps: S1 receiving the original pinyin string input by the user; S2 determining whether the current string belongs to the syllables "ji, qi, xi"; S3 if so, popping up a pronunciation selection floating window of "机 / 积, 奇 / 七, 稀 / 西" for the user; S4 receiving the target pronunciation selected or confirmed by voice by the user; S5 when the target pronunciation is "积, 七, 西", automatically inserting the U+00B7 middle dot symbol before j, q, and x to generate ·j, ·q, ·x; S6 writes the pinyin characters with supplementary annotations back to the text buffer and simultaneously writes the hidden phoneme tags. <phoneme> 。< / phoneme> 5. The method according to claim 4, wherein the floating window in step S3 uses a machine learning model to predict the default pronunciation according to the context and automatically input onto the screen the pronunciation with a prediction confidence level ≥ 95% without requiring the user to confirm again.

6. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to claim 4 or 5 are implemented.

7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the functions of the system or method according to any one of claims 1-5 are implemented.

8. The electronic device according to claim 7, wherein the device is an electronic dictionary, a smart speaker, a smart phone, a vehicle-mounted voice assistant, an online translator, or a teaching tablet for teaching Chinese as a foreign language.