Computer Chinese character input method and processing method
By obtaining the roots of Chinese characters and combining pronunciation and letter combinations and ending letter combinations, it is solved by overusing Chinese character encoding resources in the prior art, and the efficiency of Chinese character input and natural language processing is improved.
Patent Information
- Application Number
- CN202510219032.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, Chinese character encoding occupies too much resources, resulting in low efficiency in Chinese character input and natural language processing.
By obtaining the roots of Chinese characters and combining pronunciation and letter combinations and ending letter combinations, it is converted into Latin letter encoding to realize Chinese character input and processing.
It improves the accuracy of Chinese character input and the efficiency of computer Chinese characters' natural language processing, reduces the occupation of coding resources, and achieves accurate coverage of all Chinese characters.
Smart Images

Figure CN120122835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a computer Chinese character input method and a processing method. Background Art
[0002] Information technology has become the most important material and technological foundation in the world today. The rapid development of natural language and text processing information technology, in particular, has promoted the maturity of large language models. A new era of artificial intelligence technology revolution has arrived, which will have a subversive, profound and long-term impact on all aspects of production and life in human society.
[0003] In order to process Chinese characters using computer information technology, they must also be encoded. The current international standard for Chinese character encoding is Unicode / ISO10646, and the number of encoded Chinese characters is as high as 70,000+. This is because Chinese characters have the particularity of graphic ideograms compared to many other languages in the world. All relevant standards and international Unicode encode each Chinese character as an independent character, so the encoding resources occupied are more than 2,000 times that of various phonetic alphabetic languages (2,700 times that of English). It can be seen that the current Chinese character encoding has the problem of occupying too many encoding resources and a large number, resulting in very low processing efficiency, which also affects the efficiency of Chinese character and natural language processing in artificial intelligence.
[0004] Therefore, it is necessary to propose a more efficient computer Chinese character input method and Chinese character processing method. Summary of the invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a computer Chinese character input method and processing method that improves the accuracy of Chinese character input and the efficiency of computer Chinese character natural language processing.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A computer Chinese character input method comprises the following steps:
[0008] Obtaining a Chinese character root, the root being determined based on a Chinese character creation method type;
[0009] Inputting the phonetic letter combination and the root letter combination of the Chinese character in sequence, thereby inputting the Chinese character, wherein the phonetic letter combination is the Chinese phonetic spelling of the pronunciation of the Chinese character, and the root letter combination is the Chinese phonetic spelling of the root pronunciation of the Chinese character;
[0010] If there are multiple Chinese characters obtained according to the phonetic letter combination and the radical letter combination, then the suffix letter combination of the Chinese character is input to input the Chinese character.
[0011] Further, if there are multiple pronunciations of a certain Chinese character, the Chinese Pinyin spelling is determined according to the pronunciation of the phonetic component of the phono-semantic compound character or the common pronunciation.
[0012] Further, if there are multiple pronunciations of a certain Chinese character, the Chinese Pinyin spelling is determined by any one of the existing pronunciations.
[0013] Further, the root of a certain Chinese character is determined based on the following method:
[0014] If the Chinese character is a phono-semantic compound character, its root is the semantic component of the phono-semantic compound character, a part of the semantic component, or a custom Chinese character component;
[0015] If the Chinese character is a semantic compound character, its root is the main semantic component, the visible semantic component, or a custom Chinese character component;
[0016] If the Chinese character is a pictograph, its root is a character formed by itself, a visible semantic component, or a custom Chinese character component;
[0017] If the Chinese character is an ideograph, its root is the main semantic component or a custom Chinese character component.
[0018] Further, the combination of ending letters is a single-letter or double-letter combination composed of v, r, and x.
[0019] Further, the combination of ending letters is a combination of any one or more letters of the 26 Latin letters.
[0020] The present invention also provides a computer Chinese character processing method, including the following steps:
[0021] Obtain the natural Chinese language to be processed, convert each Chinese character therein into a Latin letter code to obtain a letter Chinese language, and perform machine recognition and processing based on the letter Chinese language;
[0022] Among them, the conversion of the Latin letter code of the Chinese character is obtained based on a pre-set coding mapping relationship. In the coding mapping relationship, each Chinese character corresponds to each Latin letter code one by one. The Latin letter code is obtained through the following steps:
[0023] 1) Obtain the combination of phonetic letters and the combination of root letters of the Chinese character, splice the combination of phonetic letters and the combination of root letters to obtain a first combination, and determine whether the Chinese character corresponding to the first combination is unique. If so, use the first combination as the Latin letter code of the Chinese character. If not, perform step 2);
[0024] 2) Add the combination of ending letters of the Chinese character to each Chinese character corresponding to the first combination, splice the combination of phonetic letters, the combination of root letters, and the combination of ending letters to obtain a second combination, and use the second combination as the Latin letter code of the Chinese character;
[0025] The radical letter combination is the Chinese phonetic spelling of the radical pronunciation of the Chinese character, and the radical of a certain Chinese character is determined based on the type of Chinese character creation method.
[0026] Furthermore, the phonetic letter combination is the pinyin spelling of the Chinese character pronunciation.
[0027] Furthermore, if there are multiple pronunciations of a Chinese character, the Chinese phonetic spelling is determined according to the pronunciation of the phonetic symbol of the phono-semantic character or the common pronunciation, or the Chinese phonetic spelling is determined according to any existing pronunciation.
[0028] Further, the radical of a Chinese character is determined based on the following method:
[0029] If the Chinese character is a phono-semantic character, its root is the pictogram or part of the pictogram or a custom Chinese character component of the phono-semantic character;
[0030] If the Chinese character is a pictophonetic character, its root is a main semantic character or a visible semantic character or a custom Chinese character component;
[0031] If the Chinese character is a pictographic character, its root is a self-contained character or a visible semantic symbol or a custom Chinese character component;
[0032] If the Chinese character is a pictographic character, its root is the main semantic symbol or a custom Chinese character component.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1. The present invention realizes Chinese character input based on the combination of the phonetic letter combination and the radical letter combination of Chinese characters or the combination of the phonetic letter combination, the radical letter combination and the suffix letter combination, which can realize Chinese character input more accurately, avoid the problem of using homophones incorrectly in the existing pinyin input method, and improve the input accuracy.
[0035] 2. The present invention converts each Chinese character into Latin letter code to obtain alphabetic Chinese character language, realizes the Latin letter coding of square Chinese characters, and compresses the coding resource space of 70,000+ Chinese characters to 26 Latin letters equivalent to English. Compared with the original Chinese character coding scheme, it effectively improves the efficiency of computer recognition and processing of square Chinese characters, and the efficiency is improved by several thousand times. As the most basic technology for Chinese character information processing, it will have an immeasurable positive impact on Chinese character processing.
[0036] 3. The present invention realizes Chinese character Latin letter encoding by designing phonetic letter combination, radical letter combination and / or suffix letter combination for square Chinese characters, which can accurately cover the processing of all Chinese characters and has high reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0038] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0039] 1. Conceptual basis of the present invention
[0040] 1. It is a common fact that there are many homophones in Chinese characters.
[0041] Taking the Xinhua Dictionary (compiled by the Institute of Linguistics, Chinese Academy of Social Sciences, and published by the Commercial Press, 10th edition) as an example, although it only contains more than 10,000 characters, there are a large number of homophones such as: 90 characters with bi sound; 85 characters with fu sound; 109 characters with ji sound; 80 characters with li sound; 84 characters with xi sound; 115 characters with yi sound; 99 characters with yu sound; 82 characters with zhi sound, and so on.
[0042] Take the "GB13000.1 Character Set Chinese Character Table" (including 20,902 Chinese characters) jointly issued by the Ministry of Education of the People's Republic of China and the National Language and Writing Commission as an example:
[0043] There are 346 characters for the Chinese character "ji"; 292 characters for the Chinese character "yi"; 256 characters for the Chinese character "jian"; 243 characters for the Chinese character "li"; 221 characters for the Chinese character "yu"; 214 characters for the Chinese character "fu"; 205 characters for the Chinese character "xin"; 204 characters for the Chinese character "yan"; 195 characters for the Chinese character "qi"; 194 characters for the Chinese character "chi"; 193 characters for the Chinese character "shi"; 185 characters for the Chinese character "ju"; and 180 characters for the Chinese character "zhi"
[0044] 2. The reason why there are so many homophones in Chinese characters (homophones in a broad sense, not counting the four tones).
[0045] According to the "Hanyu Pinyin Scheme", there are 21 initials and 35 finals, theoretically the number of possible combinations is 735, but the actual number of valid combinations is only over 400. The total number of Chinese characters is about 100,000, so 100,000 / 400=250, and each syllable (excluding the four tones) corresponds to an average of 250 Chinese characters, which is the objective reason for the existence of a large number of homophones.
[0046] 3. Research results and inspirations of "Shuowen Jiezi".
[0047] Shuowen Jiezi was written around 100 AD, more than 1,900 years ago. It is the greatest ancient Chinese classic that has been handed down to this day and still has irreplaceable value. It is the soul and cornerstone of Chinese character research. Among the many achievements of Shuowen Jiezi, there are two great results: one is the first interpretation of the origin and original meaning of Chinese characters using the six-character method of Chinese character creation; the other is the establishment of 540 radicals, which systematically classified Chinese characters and created a model for the compilation of Chinese character dictionaries in later generations. To this day, modern Chinese character dictionaries still follow its radical retrieval paradigm. These two major achievements also laid the foundation for the present invention.
[0048] 4. The history of the evolution and development of Chinese characters.
[0049] The inventor of the present invention has conducted in-depth research on the history of the evolution and development of Chinese characters. The great practice of the Chinese ancestors in creating and using Chinese characters is summarized as the six-character theory of Chinese character creation and use, namely, pictographic, indicative, ideographic, phono-semantic, transliteration and loan. Among them, pictographic, indicative, ideographic and phono-semantic are the methods of Chinese character formation and creation, transliteration is the method of exegesis, and loan is the method of using characters. Initially, pictographic characters that depict the image of specific things were created, such as: sun, moon, mountain, river, person, wood, horse, cow, sheep, etc. Then, based on pictographic characters, indicative characters were created, such as: root, end, blade, etc. Then, ideographic characters were created and developed by combining pictographic characters and indicative characters, such as the combination of two pictographic characters "木" to form the ideographic character "林"; three "木" to form the ideographic character "森"; or a pictographic character "人" and a pictographic character "木" to form the ideographic character "休"; a pictographic character "人" and an indicative character "本" to form the ideographic character "体", and so on. As the things described by Chinese characters become more and more complex and numerous, the pictographic and indicative methods of creating characters are unsustainable, and eventually the phono-semantic characters, which are composed of the phonetic radical (indicating pronunciation) and the semantic radical (indicating the meaning of the character), become the dominant method of creating characters, such as: "按" is a left-shaped character and a right-shaped character, the left shape "扌" represents the meaning of the character or the association of the meaning of the character, and the right sound "安" represents the pronunciation of the character "按", thus opening up an infinite space for Chinese character creation. According to research statistics, the proportion of Chinese phono-semantic characters exceeds 80%. Based on the core law of the Chinese phono-semantic character creation method, the inventor has innovatively proposed the technical solution of the present invention after rigorous demonstration and empirical verification of a large number of Chinese characters.
[0050] 2. Specific technical solutions of the present invention
[0051] Example 1
[0052] like Figure 1 As shown, this embodiment provides a computer Chinese character input method, comprising the following steps:
[0053] S1, obtaining a radical of a Chinese character, the radical being determined based on a type of Chinese character creation method;
[0054] S2, using a keyboard with 26 English letters, typing the phonetic letter combination and the root letter combination of the Chinese character in sequence, thereby inputting the Chinese character, wherein the root letter combination is the Chinese phonetic spelling of the root pronunciation of the Chinese character;
[0055] S3. If there are multiple Chinese characters obtained according to the phonetic letter combination and the radical letter combination, then the suffix letter combination of the Chinese character is input to determine the final Chinese character.
[0056] In the above input method, the phonetic letter combination is the Chinese phonetic spelling of the Chinese character pronunciation as the head. Specifically, for general Chinese characters, the phonetic letter spelling of the Chinese character's unique pronunciation (excluding the four tones) can be selected. When the Chinese character is a polyphonic character, the phonetic letter spelling that conforms to the pronunciation of the phonetic symbol of the phono-semantic character or the phonetic letter spelling of the common pronunciation can be selected. The Latin letter code of a Chinese character must have one and only one pronunciation pinyin spelling, and the pronunciation pinyin does not consider the four tones.
[0057] In other implementations, if there are multiple pronunciations of a Chinese character, the Chinese Pinyin spelling is determined by any existing pronunciation to accommodate more situations.
[0058] A radical is a Chinese character component that describes the meaning or meaning category of a Chinese character. The radical letter combination is the Chinese phonetic spelling of the radical pronunciation of the Chinese character. The radical of a Chinese character is determined based on the type of Chinese character creation method. Specifically, the radical of a Chinese character is determined based on the following method:
[0059] (a) If the Chinese character is a phono-semantic character, its root is the pictographic component (semantic component) of the phono-semantic character, that is, the pictographic character or Chinese character component formed by its pictographic component (radical) is taken as its root;
[0060] (b) if the Chinese character is a pictophonetic character, its root is a main semantic character, a visible semantic character, or a custom Chinese character component;
[0061] (c) If the Chinese character is a pictographic character, its root is a self-contained character or a visible semantic symbol or a custom Chinese character component;
[0062] (d) If the Chinese character is a pictographic character, its root is the main semantic symbol or a custom Chinese character component.
[0063] In the above input method, the radical letter combination of a Chinese character is determined based on the following method:
[0064] (a) if the Chinese character is a phono-semantic character, the letter combination of its root is the phonetic spelling of the pronunciation of the phono-semantic character's shape (semantic) component, or the phonetic spelling of a custom Chinese character root component;
[0065] (b) if the Chinese character is a pictophonetic character, the letter combination of its root is the Chinese phonetic spelling of the main meaning character or the pronunciation of the visible meaning character, or the Chinese phonetic spelling using a custom Chinese character root component;
[0066] (c) if the Chinese character is a pictographic character, the letter combination of its root is the Chinese phonetic spelling of the character itself or the pronunciation of the visible semantic component, or the Chinese phonetic spelling using a custom Chinese character root component;
[0067] (d) If the Chinese character is a pictographic character, the letter combination of its root is the Chinese phonetic spelling of the main meaning, or the Chinese phonetic spelling of a custom Chinese character root component is adopted.
[0068] Specific examples of Chinese character roots are shown in Table 1.
[0069] Table 1
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076] Chinese character radicals refer to Chinese character components that represent the meaning of a character or the category or meaning association of a character in the formation of a Chinese character. The creation and introduction of the concept of Chinese character radicals is the core of the technology of the present invention. Considering that there are more than 400 Chinese character phonetic spellings, 400*220=88,000 combination codes can be generated after introducing 220 Chinese character radicals, and in theory, one-to-one letter encoding of all Chinese characters can be achieved. For phono-semantic characters, which account for more than 80% of the number of Chinese characters, their radicals are the radicals (shape symbols, meaning symbols); for pictographic characters, indicative characters, and ideographic characters, which account for less than 20% of the number of Chinese characters, the main ideographic components are determined by tracing the source of the character creation method or using custom components as radicals. In this way, each Chinese character can be clearly uniquely letter-encoded, thereby improving the efficiency of Chinese character input.
[0077] The difference between the Chinese character radical concept adopted in the technical solution of the present invention and the existing Chinese character radicals is as follows:
[0078] 1. Different purpose and definition: The radical is used to reveal the meaning of Chinese characters and the letter code of Chinese characters. For phono-semantic characters that account for more than 80% of the characters, the radical is the radical; the purpose of the radical concept is to classify and collect Chinese characters for easy retrieval. According to the definition of GF0011-2009 "Chinese Character Radical Table": Radical is a part of the components that can be used to form characters in batches. All characters that contain a certain component are arranged together in the character set, and the component is arranged at the beginning as the leading unit, which becomes the basis for searching characters.
[0079] 2. Different quantity: According to GF0011-2009, the total number of Chinese character radicals is 201, which will not change due to the number of characters collected in the edited dictionary; the number of Chinese character roots will increase as the number of Chinese characters to be encoded increases, and the total number of roots is more than the radicals.
[0080] 3. Different extraction rules: The extraction of radicals mainly considers the representativeness and relevance of the meaning of the word, without considering the position of the radical in the formation of Chinese characters; the extraction of radicals is mainly based on the order of position, without considering the relationship between the radical and the meaning of the word. For example, the character "案" is a phono-semantic character, "木" is shaped like "木" and pronounced like "安", so the radical is "木", indicating that the material of the item is wood, and the original meaning of the character "案" refers to a wooden plate for serving food, such as "举案齐眉". According to the GF0012-2009 Chinese character radical classification standard "first up and then down" to select the radical principle, the radical of the character "案" is "宀". This "宀" has nothing to do with the meaning of "案". Another example is the character "鹏", which is also a phono-semantic character, "鸟" is shaped like "鸟" and pronounced like "朋", so the radical is "鸟", indicating that "鹏" is a bird. According to the GF0012-2009 Chinese character radical classification standard "first left and then right" to select the radical principle, the radical is "月". This "月" has nothing to do with the meaning of "鹏". It can be seen that there is an essential difference between roots and radicals.
[0081] If there are multiple Chinese characters obtained according to the phonetic letter combination and the radical letter combination, the suffix letter combination of the Chinese character is input to determine the final Chinese character. When adding the suffix letter combination in this step, for multiple Chinese characters with the same first combination, one of the Chinese characters does not have the suffix letter combination added, and the remaining Chinese characters have the suffix letter combination added.
[0082] The suffix letters are preferably selected from the less used letters in the "Chinese Pinyin Scheme". In this embodiment, the suffix is composed of single letters or double letters (vv, rr, xx, vr, vx, rx, etc.) among v, r, and x, which is used to distinguish the situation of 2-9 pronunciations and radical combinations with repeated codes. In a specific implementation, based on the frequency of use of Chinese characters, the suffix is not added to the most frequently used Chinese character, and the remaining homophonic Chinese characters are sequentially combined with v, r, x, vv, rr, xx, vr, vx, rx, etc. as suffix letters according to the frequency of use from high to low.
[0083] In other implementations, the suffix letter combination may be a combination of any one or more letters of the 26 Latin letters to accommodate more situations.
[0084] The above input method is applicable to the basic Chinese character set consisting of 3,500 first-level Chinese characters and 3,000 second-level Chinese characters in the "General Standard Chinese Character Table", which is a total of 6,500 general standard Chinese characters. Of course, the technology of the present invention is applicable to expansion to all Chinese characters. The above input method is also applicable to all Chinese characters (70,244) included in the Unicode (ISO / IEC10646) international standard.
[0085] The relationship between some Chinese characters and Latin characters obtained based on the above method is shown in Table 2.
[0086] Table 2
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103] The above-mentioned computer Chinese character input method is illustrated as follows:
[0104] For example, for the Chinese character "村", the Chinese character is a phono-semantic character, and its root is 木, so the input of the Chinese character "村" can be realized by typing cunmu on the keyboard;
[0105] For example, for the Chinese character “歪”, the character is a pictophonetic character, and its root is “不”, so the input of the Chinese character “歪” can be realized by typing waibu on the keyboard;
[0106] For example, for the Chinese character “桑”, the character is a pictographic character, and its root is “人”, so the input of the Chinese character “桑” can be realized by typing “sanren” on the keyboard;
[0107] For example, for the Chinese character “本”, the character is a pictographic character, and its root is 木, so the input of the character “本” can be realized by typing benmu on the keyboard;
[0108] For example, for the Chinese character "妹", this Chinese character is a phono-semantic character, and its root is "女". Chinese characters with the same pronunciation and the same root as this Chinese character include "传媒" and "媚". Then it is necessary to add a suffix letter combination, that is, the input of the Chinese character "妹" can be realized by typing "meinvv" on the keyboard.
[0109] If a phrase is to be input, an input code may be obtained for each Chinese character according to the above method, and the corresponding phrase may be input based on the combination of the codes of the Chinese characters in the phrase, or each Chinese character may be input independently.
[0110] Through the above input method, Chinese character input can be achieved more accurately, avoiding the problems of using homophones incorrectly in existing pinyin input methods, and improving input accuracy.
[0111] Example 2
[0112] This embodiment provides a computer Chinese character processing method, comprising the following steps:
[0113] Acquire a natural language of Chinese characters to be processed, convert each Chinese character therein into a Latin letter code, obtain an alphabetic Chinese character language, and perform machine recognition and processing based on the alphabetic Chinese character language;
[0114] The conversion of the Chinese character Latin letter code is obtained based on a preset code mapping relationship, in which each Chinese character corresponds to each Latin letter code one by one, and the Latin letter code is obtained by the following steps:
[0115] 1) Obtaining a phonetic letter combination and a radical letter combination of a Chinese character, concatenating the phonetic letter combination and the radical letter combination to obtain a first combination, and determining whether the Chinese character corresponding to the first combination is unique; if so, using the first combination as the Latin letter code of the Chinese character; if not, executing step 2);
[0116] 2) adding the suffix letter combination of the Chinese character to each Chinese character corresponding to the first combination, concatenating the phonetic letter combination, the root letter combination and the suffix letter combination to obtain a second combination, and using the second combination as the Latin letter code of the Chinese character;
[0117] The radical letter combination is the Chinese phonetic spelling of the radical pronunciation of the Chinese character, and the radical of a certain Chinese character is determined based on the type of Chinese character creation method.
[0118] Through the above method, the computer can recognize Chinese characters using the recognition method of the 26 English letters, which greatly improves the efficiency of computer recognition and processing of square Chinese characters. As the most basic technology for Chinese character information processing, it will have a profound and positive impact on Chinese character information processing.
[0119] The rest is the same as in Example 1.
[0120] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention or the part that contributes to the prior art or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (including but not limited to a personal computer, a server, a communication terminal or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes but is not limited to: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0121] In other embodiments, an electronic device is provided, including one or more processors, a memory, and one or more programs stored in the memory, wherein the one or more programs include instructions for executing the computer Chinese character processing method as described above.
[0122] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A computer Chinese character input method, characterized in that: The following steps are involved: Obtaining a Chinese character root, the root being determined based on a Chinese character creation method type; Inputting the phonetic letter combination and the root letter combination of the Chinese character in sequence, thereby inputting the Chinese character, wherein the phonetic letter combination is the Chinese phonetic spelling of the pronunciation of the Chinese character, and the root letter combination is the Chinese phonetic spelling of the root pronunciation of the Chinese character; If there are multiple Chinese characters obtained according to the phonetic letter combination and the radical letter combination, then the suffix letter combination of the Chinese character is input to input the Chinese character.
2. The computer Chinese character input method according to claim 1, characterized in that: If there are multiple pronunciations of a certain Chinese character, the Chinese phonetic spelling is determined according to the pronunciation of the phonetic symbol of the phono-semantic character or the common pronunciation.
3. The computer Chinese character input method according to claim 1, characterized in that: If there are multiple pronunciations of a Chinese character, the Chinese phonetic spelling is determined by any existing pronunciation.
4. The computer Chinese character input method according to claim 1, characterized in that: The root of a Chinese character is determined based on the following method: If the Chinese character is a phono-semantic character, its root is the pictogram or part of the pictogram or a custom Chinese character component of the phono-semantic character; If the Chinese character is a pictophonetic character, its root is a main semantic character or a visible semantic character or a custom Chinese character component; If the Chinese character is a pictographic character, its root is a self-contained character or a visible semantic symbol or a custom Chinese character component; If the Chinese character is a pictographic character, its root is the main semantic symbol or a custom Chinese character component.
5. The computer Chinese character input method according to claim 1, characterized in that: The suffix letter combination is a combination of single or double letters from v, r, and x.
6. The computer Chinese character input method according to claim 1, characterized in that: The suffix letter combination is a combination of any one or more letters of the 26 Latin letters.
7. A computer Chinese character processing method, characterized in that: The following steps are involved: Acquire a natural language of Chinese characters to be processed, convert each Chinese character therein into a Latin letter code, obtain an alphabetic Chinese character language, and perform machine recognition and processing based on the alphabetic Chinese character language; The conversion of the Latin letter code of the Chinese character is obtained based on a preset code mapping relationship, in which each Chinese character corresponds to each Latin letter code one by one, and the Latin letter code is obtained by the following steps: 1) Obtaining a phonetic letter combination and a radical letter combination of a Chinese character, concatenating the phonetic letter combination and the radical letter combination to obtain a first combination, and determining whether the Chinese character corresponding to the first combination is unique; if so, using the first combination as the Latin letter code of the Chinese character; if not, executing step 2); 2) adding the suffix letter combination of the Chinese character to each Chinese character corresponding to the first combination, concatenating the phonetic letter combination, the root letter combination and the suffix letter combination to obtain a second combination, and using the second combination as the Latin letter code of the Chinese character; The radical letter combination is the Chinese phonetic spelling of the radical pronunciation of the Chinese character, and the radical of a certain Chinese character is determined based on the type of Chinese character creation method.
8. The computer Chinese character processing method according to claim 7, characterized in that: The phonetic letter combination is the pinyin spelling of the Chinese character pronunciation.
9. The computer Chinese character processing method according to claim 8, characterized in that: If there are multiple pronunciations of a certain Chinese character, the Chinese phonetic spelling is determined according to the pronunciation of the phonetic symbol of the phono-semantic character or the common pronunciation, or the Chinese phonetic spelling is determined according to any existing pronunciation.
10. The computer Chinese character processing method according to claim 7, characterized in that: The root of a Chinese character is determined based on the following method: If the Chinese character is a phono-semantic character, its root is the pictogram or part of the pictogram or a custom Chinese character component of the phono-semantic character; If the Chinese character is a pictophonetic character, its root is a main semantic character or a visible semantic character or a custom Chinese character component; If the Chinese character is a pictographic character, its root is a self-contained character or a visible semantic symbol or a custom Chinese character component; If the Chinese character is a pictographic character, its root is the main semantic symbol or a custom Chinese character component.