Chinese character encoding and decoding method and system based on Chinese pinyin and Chinese character characteristics
By using an encoding method based on Chinese Pinyin and Chinese character features, and combining initials, finals, stroke order, and feature encoding, the problems of homophone rate and tone input complexity in existing input methods are solved, achieving concise and efficient Chinese character input, and improving user experience and matching accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU SHENGFANGSI INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing input method products have shortcomings in terms of homophone rate, accuracy, learning cost, and adaptability to intelligent input. In particular, the homophone rate of Pinyin input method is high, the correspondence between vowels and keys in Shuangpin input method is difficult to remember, and tone input is complex and the code is longer, making it difficult to put into use.
It adopts an encoding method based on Chinese Pinyin and Chinese character features. By combining initial consonant abbreviation, final vowel abbreviation, Chinese character stroke encoding and Chinese character feature encoding, combined with dynamic encoding mechanism and auxiliary input area, it generates concise and efficient Chinese character encoding and supports multi-dimensional feature decoding.
It significantly reduces the homonym rate, improves the matching accuracy between encoding and Chinese characters, simplifies the keyboard layout, enhances the user input experience and operation smoothness, and supports direct input of various tone changes and special characters.
Smart Images

Figure CN121881982A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer and text input method technology, and in particular to a Chinese character encoding and decoding method and system based on Chinese Pinyin and Chinese character features. Background Technology
[0002] Current input method products have some shortcomings in application. For example, shape-based input methods have high homonym rates and accuracy, but their high learning cost and inability to adapt to the trend of intelligent input mean they are gradually being phased out. Pinyin input methods are simple and easy to learn, but suffer from high homonym rates and low accuracy. Double-pinyin input methods effectively shorten code length, but firstly, the correspondence between vowels and keys is irregular and difficult to remember, and secondly, they do not solve the problem of high homonym rates. Furthermore, while some input methods introduce tones, the processing is rather crude. They either directly set separate keys for tone input or set separate states for tone input, but this results in longer codes, difficulty in parsing, and an inability to directly input tone symbols such as ā, á, ǎ, and à. Therefore, input methods that support tones are often difficult to implement.
[0003] Some input methods introduce other elements to improve accuracy, but these often present some problems. For example, introducing shape codes (usually more than a hundred) on top of double pinyin increases the number of effective combinations and accuracy, but greatly increases the learning cost; some introduce multiple external elements, such as strokes and structure, which improve the accuracy of inputting a single Chinese character, but this results in excessive complexity, and for words with two or more characters, the encoding is too long and the parsing is inconvenient. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a Chinese character encoding and decoding method and system based on Chinese Pinyin and Chinese character features, which improves the matching accuracy between encoding and Chinese characters and enhances the user's input experience.
[0005] In a first aspect, embodiments of this application provide a Chinese character encoding method based on Chinese Pinyin and Chinese character features, including: When the keyboard is in the first state, the corresponding initial consonant abbreviation is generated according to the first position triggered on the keyboard. When the keyboard is in the second state, the corresponding vowel abbreviation is generated according to the second and third positions triggered on the keyboard and the triggering order. In any keyboard state, if the currently triggered fourth position belongs to the auxiliary input area of the keyboard, then the corresponding Chinese character starting stroke code is generated according to the fourth position, and the keyboard state is switched to auxiliary input state. When the keyboard is in auxiliary input mode, the corresponding Chinese character feature code is generated based on the fifth position triggered in the auxiliary input area. When the input is complete, the generated initial consonant abbreviation, final vowel abbreviation, Chinese character stroke code, and Chinese character feature code are combined to construct the Chinese character code according to the encoding generation order.
[0006] This application provides a Chinese character encoding method based on Pinyin and Chinese character features. It generates initial consonant abbreviations, final vowel abbreviations, stroke order codes, and character feature codes under different keyboard states, and combines them sequentially to form a complete Chinese character encoding. This embodiment effectively integrates the Pinyin and structural features of Chinese characters, enabling the encoding to reflect both pronunciation and character shape, thus significantly reducing the homophone rate. For example, traditional Pinyin input methods have numerous homophones, requiring users to frequently flip through pages to select the target character. This embodiment, by introducing stroke order information and structural features, and distinguishing between initial consonants and final vowels in the Pinyin encoding, can quickly filter out highly probable candidate characters in the subsequent decoding stage, greatly reducing user search time and improving the matching accuracy between the encoding and the character. Secondly, this embodiment employs a state-based dynamic encoding mechanism, allowing the same key to generate different codes under different input states, significantly reducing the space occupied by the keyboard layout and improving the simplicity and aesthetics of the input interface. Furthermore, this embodiment also includes an auxiliary input area. Users can trigger the auxiliary input area at any time to encode the structural features of the current Chinese character as needed, eliminating the need for frequent mode switching during input, resulting in smooth and natural operation and improved user input experience. Finally, because the encoding rules take into account both the naturalness of Pinyin and the structural nature of Chinese characters, the learning curve is low, allowing users to master the usage method in a short time, further enhancing the user input experience.
[0007] Furthermore, when the keyboard state is the second state, the corresponding vowel abbreviation is generated based on the triggered second and third positions on the keyboard and the triggering order, including: The second and third characters are generated according to the second and third positions on the keyboard and the triggering order; The tone and tone position are determined based on the second and third characters; Based on the tones and positions of the second and third characters, the corresponding vowel abbreviations are generated.
[0008] This application provides a method for generating abbreviated vowel codes, particularly utilizing the combination of a second and third character to construct vowel codes with tones. Specifically, this embodiment cleverly solves the problem of difficult tone input in traditional input methods. Previous methods often used individual keys or complex switching methods to mark tones, which not only increased operational complexity but also easily led to input errors. In contrast, this method embeds tones through simple combination actions, allowing users to intuitively and quickly input vowels with tones. For example, users only need to press two relevant keys in sequence, and the system can automatically synthesize the correct vowel form, avoiding cumbersome input processes. Furthermore, this design is compatible with various tone variations and supports the direct input of special characters such as ā, á, ǎ, and à, which is rare in existing input methods and demonstrates significant technological innovation. In summary, the introduction of this vowel encoding mechanism greatly enriches the functionality of input methods, improves the matching accuracy between codes and Chinese characters, and enhances the user's input experience.
[0009] Furthermore, the stroke code of the Chinese character corresponds to a stroke information of the Chinese character, which is "horizontal", "vertical", "left-falling", "dot" and "turning".
[0010] This application specifies five basic stroke types corresponding to the stroke start codes of Chinese characters: horizontal, vertical, left-falling, dot, and turning. These stroke types cover the starting stroke forms of the vast majority of Chinese characters, allowing this embodiment to add morphological features to the Chinese character encoding process without requiring a large number of keys, thus ensuring the simplicity of the encoding process. In practical applications, users can quickly select the corresponding stroke code based on the writing habits of the target Chinese character, improving the input efficiency of Chinese character encoding. It also reduces the learning difficulty for users because the stroke information aligns with people's intuitive understanding of Chinese character structure, eliminating the need to memorize complex rules or the complete strokes of the character. Furthermore, in the subsequent decoding process, the introduction of stroke codes further compresses the search space, accelerates the character matching process, and improves the matching accuracy between the code and the character.
[0011] In one possible implementation, the Chinese character feature encoding corresponds to a Chinese character feature information, which is a first feature, a second feature, a third feature, or a fourth feature; The first feature is that the target Chinese character to be encoded can be split into two or more first independent components in the left and right directions, and any one of the first independent components can be split into two or more second independent components in the top and bottom directions. The second feature is that the target Chinese character can be split into two or more first independent components in the left and right directions, and each of the first independent components cannot be split into two or more second independent components in the right and left directions. The third feature is that the target Chinese character cannot be split into two or more first independent components horizontally, and the target Chinese character can be split into two or more second independent components vertically; The fourth feature is that the target Chinese character cannot be split horizontally or vertically.
[0012] The embodiments of the present application provide four types of Chinese character feature information. These features describe the structural attributes of Chinese characters based on whether they can be split horizontally or vertically, and thus provide a retrieval basis for the subsequent Chinese character matching process. The embodiments of the present application provide a hierarchical structural framework for Chinese character encoding, enabling each Chinese character to be accurately classified. For example, the first feature is applicable to compound Chinese characters with relatively complex structures, such as "Sui", "Shao", "Shao", etc., and the fourth feature is applicable to Chinese characters with relatively simple structures, such as "Ding", "Guo", "Guo", etc. Compared with the traditional single-dimensional encoding method, the multi-dimensional feature encoding of this embodiment captures the features of Chinese characters more comprehensively, so that the target Chinese character can be located faster in the decoding stage, greatly reducing the rate of duplicate codes for users in the actual Chinese character input process, improving the matching accuracy, and also conforming to the user's memory habit of Chinese characters and improving the user's input experience.
[0013] In a possible implementation manner, the first character and the second character in the simple code of the finals are respectively generated by being triggered by different input areas, specifically: When the keyboard state is the second state, the corresponding first character is generated according to the second position triggered on the keyboard; In any keyboard state, if the currently triggered third position belongs to the auxiliary input area of the keyboard, the corresponding second character is generated according to the third position, and the keyboard state is switched to the auxiliary input state; When the keyboard state is the auxiliary input state, the starting stroke code and the Chinese character feature code of the corresponding Chinese character are sequentially generated according to the fourth position and the fifth position triggered in the auxiliary input area; The simple code of the finals is generated according to the first character and the second character.
[0014] This application embodiment defines the specific application method of the auxiliary input area. By reusing the auxiliary input area on the keyboard, the second character of the vowel abbreviation, the stroke code of the Chinese character, and the feature code of the Chinese character are input, without the need to add separate dedicated keys for different types of codes. This greatly optimizes the keyboard layout, enabling richer input functions to be carried within the limited screen (especially mobile device screens) or keyboard space, maintaining a clear and concise input interface, and avoiding the bloated interface and visual burden on users caused by too many keys. In addition, this embodiment allows users to flexibly trigger the auxiliary input area as needed during or after inputting the main code (such as the initial and the first part of the vowel) to supplement the vowel supplementary information and structural feature information (second part of the vowel, stroke, component number classification) of the input Chinese character, improving the matching accuracy between the code and the Chinese character. This design makes the encoding input process more coherent and natural, significantly improving the user's input experience and operational smoothness.
[0015] Secondly, embodiments of this application provide a Chinese character decoding method based on Pinyin and Chinese character features. This method is used to decode Chinese character encoding generated by any of the Chinese character encoding methods based on Pinyin and Chinese character features described in this application, including: Get Chinese character encoding; Based on the position of the initial consonant abbreviation in the Chinese character encoding, the Chinese character encoding is divided into several sub-encoding information, wherein there is at most one initial consonant abbreviation in the sub-encoding information, and the initial consonant abbreviation is the first code in the sub-encoding information; Based on each of the sub-encoding information, corresponding retrieval information is generated through a preset mapping table, including generating several complete initials based on the initial consonant abbreviation; if the sub-encoding information contains a vowel abbreviation, several complete vowels are generated based on the vowel abbreviation; if the sub-encoding information contains a Chinese character starting stroke code, corresponding Chinese character starting stroke information is generated based on the Chinese character starting stroke code; if the sub-encoding information contains a Chinese character feature code, corresponding Chinese character feature information is generated based on the Chinese character feature code. For each of the sub-encoding information, several candidate Chinese characters are retrieved from a preset database according to the corresponding retrieval information; If the number of sub-encoding information is 1, then the candidate Chinese characters will be displayed on the preset interface.
[0016] This application provides a Chinese character decoding method corresponding to the encoding method. First, considering the situation where a user inputs multiple Chinese characters simultaneously, the initial consonant abbreviation position splits the Chinese character encoding, which may contain multiple characters, into several sub-encoding information, and decodes each sub-encoding information separately, achieving rapid conversion from abbreviation to complete information. Specifically, during the information decoding process, the presence of final consonant abbreviations, character stroke start codes, and character feature codes in the sub-encoding information is checked sequentially, generating corresponding complete final consonant, character stroke start, and character feature information. In other words, this embodiment can still complete the Chinese character decoding process and display the Chinese character matching results even when one or more of the final consonant abbreviations, character stroke start codes, and character feature codes are missing, improving the flexibility of user Chinese character input. Second, the decoded retrieval information can cover multiple dimensions such as initial consonants, final consonants, stroke start codes, and features, making database queries more accurate, the returned results more targeted, and improving the matching accuracy between encoding and Chinese characters.
[0017] Furthermore, the step of generating several corresponding complete initials based on the initial consonant abbreviation includes: If the initial consonant abbreviation is a character in a preset first character set, then a blank placeholder is generated as the complete initial consonant in the search information; If the initial consonant abbreviation is a character in a preset second character set, then the flat tongue initial consonant and the retroflex initial consonant corresponding to the initial consonant abbreviation are generated as the several complete initial consonants.
[0018] In this embodiment, by classifying and processing the initial consonant abbreviations, the system can more flexibly adapt to the user's input habits and parse the corresponding initial consonant information. For example, in some cases, the user may not need to input a specific initial consonant or some Chinese characters may not have an initial consonant. In this case, the user inputs a character from the first character set during the encoding stage. At this time, a blank placeholder is generated during the decoding stage as the complete initial consonant in the search information, ensuring that each Chinese character has corresponding initial consonant information and ensuring that subsequent searches proceed normally. Secondly, for some easily confused alveolar and retroflex initial consonants, this embodiment can generate the corresponding alveolar and retroflex initial consonants based on the initial consonant abbreviation of a single character input by the user. This not only reduces the number of characters the user needs to input and improves input efficiency, but also significantly improves the error tolerance of Chinese character input. For example, users may have inaccurate pronunciation due to differences in dialect or accent. This design ensures that reasonable initial consonant abbreviations can be generated even with incompletely accurate input by using a preset character set mapping relationship. This enhances the robustness of the system, effectively reduces the number of keys, optimizes the keyboard layout, and improves the simplicity and aesthetics of the input interface, providing users with a more convenient and efficient interactive experience.
[0019] Furthermore, if the number of sub-encoding information is greater than 1, then according to the preset lexicon and the order of each sub-encoding information in the Chinese character encoding, the corresponding candidate Chinese characters are arranged and combined to generate several words or phrases, and each word or phrase is displayed on the preset interface.
[0020] This embodiment further considers the scenario where a user simultaneously inputs multiple Chinese characters, resulting in a combined display. Since each sub-encoding information can match multiple Chinese characters, when only one sub-encoding information exists, the matched characters can be directly displayed. However, when multiple sub-encoding information exists, the number of possible combinations of the matched characters increases significantly. Users need to select the desired character from a massive number of matching results, which greatly reduces the user experience. Therefore, this embodiment intelligently combines candidate characters into a coherent set of words or phrases by integrating a preset dictionary and the order information of each sub-encoding. This significantly reduces the number of matches, ensuring that users can quickly select the desired character, improving the matching accuracy between encoding and characters, and enhancing the user's input experience.
[0021] Thirdly, embodiments of this application provide a Chinese character encoding system based on Chinese Pinyin and Chinese character features, including a module for generating initial consonant abbreviation codes, a module for generating final vowel abbreviation codes, a module for generating Chinese character stroke start codes, a module for generating Chinese character feature codes, and a combination module; The initial consonant abbreviation generation module is used to generate the corresponding initial consonant abbreviation based on the first position triggered on the keyboard when the keyboard state is in the first state. The vowel abbreviation generation module is used to generate corresponding vowel abbreviations based on the second and third positions triggered on the keyboard and the triggering order when the keyboard is in the second state. The Chinese character starting stroke code generation module is used to generate the corresponding Chinese character starting stroke code based on the fourth position in any keyboard state if the currently triggered fourth position belongs to the auxiliary input area of the keyboard, and switch the keyboard state to auxiliary input state. The Chinese character feature encoding generation module is used to generate the corresponding Chinese character feature encoding based on the fifth position triggered in the auxiliary input area when the keyboard state is auxiliary input state. The combination module is used to construct a Chinese character code by combining the generated initial consonant abbreviation, the final vowel abbreviation, the Chinese character stroke code, and the Chinese character feature code according to the encoding generation order when the input is completed.
[0022] Furthermore, the Chinese character feature encoding corresponds to a type of Chinese character feature information, which is a first feature, a second feature, a third feature, or a fourth feature; The first feature is that the target Chinese character to be encoded can be split into two or more first independent components in the left and right directions, and any one of the first independent components can be split into two or more second independent components in the top and bottom directions. The second feature is that the target Chinese character can be split into two or more first independent components in the left and right directions, and each of the first independent components cannot be split into two or more second independent components in the right and left directions. The third feature is that the target Chinese character cannot be split horizontally into two or more first independent components, and the target Chinese character can be split vertically into two or more second independent components. The fourth feature is that the target Chinese character cannot be split horizontally or vertically.
[0023] Fourthly, the Chinese character decoding system is used to decode the Chinese character encoding generated by any of the Chinese character encoding systems based on Chinese Pinyin and Chinese character features described in this application, including an acquisition module, a splitting module, an information decoding module, a retrieval module, and a display module; The acquisition module is used to acquire Chinese character encodings; The splitting module is used to split the Chinese character encoding into several sub-encoding information according to the position of the initial consonant abbreviation in the Chinese character encoding. The sub-encoding information contains at most one initial consonant abbreviation, and the initial consonant abbreviation is the first code in the sub-encoding information. The information decoding module is used to generate corresponding retrieval information according to each of the sub-encoding information through a preset mapping table. Each retrieval information contains at least several complete initials corresponding to the initial consonant abbreviation; if the sub-encoding information contains a final vowel abbreviation, the retrieval information also contains several corresponding complete finals; if the sub-encoding information contains a Chinese character starting stroke code, the retrieval information also contains the corresponding Chinese character starting stroke information; if the sub-encoding information contains a Chinese character feature code, the retrieval information also contains the corresponding Chinese character feature information. The retrieval module is used to retrieve several candidate Chinese characters from a preset database for each sub-encoding information based on the corresponding retrieval information. The display module is used to display the candidate Chinese characters on a preset interface if the number of sub-encoding information is 1. Attached Figure Description
[0024] Figure 1 A flowchart illustrating a Chinese character encoding method based on Pinyin and Chinese character features provided in this application embodiment; Figure 2A schematic diagram of an encoding interface for applying the Chinese character encoding method based on Chinese Pinyin and Chinese character features provided in the embodiments of this application; Figure 3 A flowchart illustrating a Chinese character decoding method based on Pinyin and Chinese character features provided in this application embodiment; Figure 4 A schematic diagram of the structure of a Chinese character encoding system based on Pinyin and Chinese character features provided in this application embodiment; Figure 5 This is a schematic diagram of the structure of a Chinese character decoding system based on Chinese Pinyin and Chinese character features, provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0026] It should be noted that the step numbers in this document are only for the convenience of explaining the specific embodiments and are not intended to limit the order in which the steps are performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0027] Example 1: like Figure 1 As shown, Embodiment 1 provides a Chinese character encoding method based on Chinese Pinyin and Chinese character features, including steps S1-S5: Step S1: When the keyboard state is in the first state, generate the corresponding initial consonant abbreviation according to the first position triggered on the keyboard; Step S2: When the keyboard is in the second state, generate the corresponding vowel abbreviation based on the second and third positions triggered on the keyboard and the triggering order. Step S3: In any keyboard state, if the currently triggered fourth position belongs to the auxiliary input area of the keyboard, then generate the corresponding Chinese character starting stroke code according to the fourth position, and switch the keyboard state to auxiliary input state. Step S4: When the keyboard state is in auxiliary input state, generate the corresponding Chinese character feature code according to the fifth position triggered in the auxiliary input area; Step S5: When the input is complete, according to the encoding generation order, combine the generated initial consonant abbreviation, the final vowel abbreviation, the Chinese character stroke code, and the Chinese character feature code to construct the Chinese character code.
[0028] This application provides a Chinese character encoding method based on Pinyin and Chinese character features. It generates initial consonant abbreviations, final vowel abbreviations, stroke order codes, and character feature codes under different keyboard states, and combines them sequentially to form a complete Chinese character encoding. This embodiment effectively integrates the Pinyin and structural features of Chinese characters, enabling the encoding to reflect both pronunciation and character shape, thus significantly reducing the homophone rate. For example, traditional Pinyin input methods have numerous homophones, requiring users to frequently flip through pages to select the target character. This embodiment, by introducing stroke order information and structural features, and distinguishing between initial consonants and final vowels in the Pinyin encoding, can quickly filter out highly probable candidate characters in the subsequent decoding stage, greatly reducing user search time and improving the matching accuracy between the encoding and the character. Secondly, this embodiment employs a state-based dynamic encoding mechanism, allowing the same key to generate different codes under different input states, significantly reducing the space occupied by the keyboard layout and improving the simplicity and aesthetics of the input interface. Furthermore, this embodiment also includes an auxiliary input area. Users can trigger the auxiliary input area at any time to encode the structural features of the current Chinese character as needed, eliminating the need for frequent mode switching during input, resulting in smooth and natural operation and improved user input experience. Finally, because the encoding rules take into account both the naturalness of Pinyin and the structural nature of Chinese characters, the learning curve is low, allowing users to master the usage method in a short time, further enhancing the user input experience.
[0029] In a preferred embodiment, in step S1, the main input area of the input keyboard corresponding to this encoding method is a 3*8 layout, and the keys in the main input area are represented by 1-1 to 3-8 (first row, first column to third row, eighth column), with the bottom row being commonly used function keys, which is suitable for mobile touch screen operating systems; after adjustments according to the keyboard, this encoding method can also be applied to general-purpose keyboards.
[0030] Based on the input, the keyboard is divided into different states, and in each state, a mapping relationship is established between the key press and the code.
[0031] In the first state, the code entered by the user via the keyboard is called the "initial consonant abbreviation," which includes: Q, R, W, T, Y, P, L, S, D, F, G, H, J, K, Z, X, C, B, N, M, V, where V represents the case where the pinyin has no initial consonant. The correspondence with the keyboard layout is as follows: In this state, keys 1-8, 2-8, and 3-8 correspond to Λ, Φ, and Π, respectively. These three Greek letters are used to represent these three keys. In the no-input state, clicking these keys will input punctuation marks; in the input state, clicking these three keys will not be interpreted as initials. Furthermore, this embodiment provides corresponding initial consonant abbreviation input rules. For retroflex consonants, the 'h' is directly removed, requiring only one character to be input. For example, 'sh' requires only 'S', 'ch' only requires 'C', and 'zh' only requires 'Z'. Therefore, some alveolar and retroflex consonants in this embodiment share the same encoding. For pinyin without an initial consonant, 'v' is used to represent the initial consonant. For ease of explanation, this embodiment will use the same keyboard layout for subsequent examples, using consonant symbols to represent key positions, such as calling keys 2-1 the S key, keys 3-5 the N key, and keys 2-8 the Φ key.
[0032] Furthermore, in step S2, when the keyboard state is the second state, the corresponding vowel abbreviation is generated according to the triggered second position, third position, and triggering order on the keyboard, including: The second and third characters are generated according to the second and third positions on the keyboard and the triggering order; The tone and tone position are determined based on the second and third characters; Based on the tones and positions of the second and third characters, the corresponding vowel abbreviations are generated.
[0033] This application provides a method for generating abbreviated vowel codes, particularly utilizing the combination of a second and third character to construct vowel codes with tones. Specifically, this embodiment cleverly solves the problem of difficult tone input in traditional input methods. Previous methods often used individual keys or complex switching methods to mark tones, which not only increased operational complexity but also easily led to input errors. In contrast, this method embeds tones through simple combination actions, allowing users to intuitively and quickly input vowels with tones. For example, users only need to press two relevant keys in sequence, and the system can automatically synthesize the correct vowel form, avoiding cumbersome input processes. Furthermore, this design is compatible with various tone variations and supports the direct input of special characters such as ā, á, ǎ, and à, which is rare in existing input methods and demonstrates significant technological innovation. In summary, the introduction of this vowel encoding mechanism greatly enriches the functionality of input methods, improves the matching accuracy between codes and Chinese characters, and enhances the user's input experience.
[0034] In a preferred embodiment, before inputting Chinese characters, the user processes the finals according to the following rules and inputs the corresponding final abbreviations: a) Only the first two digits of the final are retained, and the rest are discarded; b) If the final has fewer than two digits, it is represented by 'v'; c) For vowels without tone in the finals a, e, i, o, u, ü, they are represented by â, î, ô, ê, û respectively (ü is also represented by û). The conversion process from some finals to final abbreviations is shown below: Specifically, in the second keyboard state, first input the first character of the vowel abbreviation (which can also be considered as keyboard state 2-1). Its value range is divided into two categories: 1) The tone is the first character of the vowel abbreviation (e.g., vowels starting with a, e, o and some vowels starting with ui (un, uv, in, iv)), used to represent the following vowel abbreviations: a, ā, á, ǎ, à, o, ō, ó, ǒ, ò, e, ē, é, ě, è, i, ī, í 1) ǐ, ì, u, ū, ú, ǔ, ù, ü, ǖ, ǘ, ǚ, ǜ, where a, e, i, o, u, ü represent neutral tones; 2) The tone is in the second position of the vowel (for example, the vowels ua, ui, uo, ue, io, ie, iu, ia, etc., with the tone in the second position of the vowel), which are represented by the special symbols û and î, where û represents the vowel that starts with u and has the tone in the second position, and î represents the vowel that starts with i and has the tone in the second position of the vowel.
[0035] Therefore, keyboard state 2-1 corresponds to 32 symbols. Combining u and ü, and merging neutral and first tones, reduces the number of symbols to 22. Additionally, ã, ẽ, ĩ, õ, and ũ with tildes represent tones, but not limited to specific tones. ã and ẽ correspond to one key, representing either a or e, regardless of tone; ĩ, õ, and ũ correspond to one key, representing either i, o, or u, also regardless of tone. The final correspondence between keyboard state 2-1 and the input symbols is as follows: Furthermore, the input content of keyboard state 2-2 corresponds to the second position of the vowel abbreviation code. If the input in the previous state is with a tone mark, the corresponding input range is: a, e, i, o, u, r, n, v. Since the vowel part has no tone, add the tone mark, and the corresponding input range is â, ê, û, î, ô, r, n, v. If the input in the previous state is û, î, the corresponding input range is: a, ā, á, ǎ, à, o, ō, ó, ǒ, ò, e, ē, é, ě, è, i, ī, í, ǐ, ì, u, ū, ú, ǔ, ù. There are a total of 33 codes, which is more than 24 keys. Therefore, some keys need to be merged. The neutral tone and the first tone are still merged. Considering that if the previous step input has a tone mark, the next step input must be without a tone mark; and if the previous step input has no tone mark, the next step must have a tone mark. Therefore, when the keyboard is in state 2-2, ā and â are mutually exclusive (i.e., they cannot be selected at the same time), ē and ê are mutually exclusive, ī and î are mutually exclusive, ō and ô are mutually exclusive, and ū and û are mutually exclusive. Therefore, these symbols are merged. Since the number of Chinese characters corresponding to "r" is relatively small, it is merged with "v" and they share one key. In this keyboard, keys 3-7 represent either "a" or "e," regardless of tone; keys 3-8 represent either "i," "o," or "u," also regardless of tone. Users can press these two keys if they are unsure of the tone. Therefore, when the keyboard is in its second state, the corresponding vowel abbreviations can be generated based on the second and third trigger positions and the triggering order.
[0036] Furthermore, in step S3, the Chinese character starting stroke code corresponds to a type of Chinese character starting stroke information, which is "horizontal", "vertical", "left-falling", "dot" and "turning".
[0037] This application specifies five basic stroke types corresponding to the stroke start codes of Chinese characters: horizontal, vertical, left-falling, dot, and turning. These stroke types cover the starting stroke forms of the vast majority of Chinese characters, allowing this embodiment to add morphological features to the Chinese character encoding process without requiring a large number of keys, thus ensuring the simplicity of the encoding process. In practical applications, users can quickly select the corresponding stroke code based on the writing habits of the target Chinese character, improving the input efficiency of Chinese character encoding. It also reduces the learning difficulty for users because the stroke information aligns with people's intuitive understanding of Chinese character structure, eliminating the need to memorize complex rules or the complete strokes of the character. Furthermore, in the subsequent decoding process, the introduction of stroke codes further compresses the search space, accelerates the character matching process, and improves the matching accuracy between the code and the character.
[0038] In one possible implementation, in step S4, the Chinese character feature encoding corresponds to a Chinese character feature information, which is a first feature, a second feature, a third feature, or a fourth feature; Among them, the first feature is that the target Chinese character to be encoded can be split into two or more first independent components from left to right, and any one of the first independent components can be split into two or more second independent components from top to bottom; The second feature is that the target Chinese character can be split into two or more first independent components from left to right, and each of the first independent components cannot be split into two or more second independent components from top to bottom; The third feature is that the target Chinese character cannot be split into two or more first independent components from left to right, and the target Chinese character can be split into two or more second independent components from top to bottom; The fourth feature is that the target Chinese character cannot be split either from left to right or from top to bottom.
[0039] The embodiments of the present application provide four types of Chinese character feature information. These features describe the structural attributes of Chinese characters based on whether they can be split from left to right or from top to bottom, thereby providing a retrieval basis for the subsequent Chinese character matching process. The embodiments of the present application provide a hierarchical structural framework for Chinese character encoding, enabling each Chinese character to be accurately classified. For example, the first feature is applicable to composite Chinese characters with relatively complex structures, such as "Sui", "Shao", "Shao", etc., and the fourth feature is applicable to Chinese characters with relatively simple structures, such as "Ding", "Guo", "Guo", etc. Compared with the traditional single-dimensional encoding method, the multi-dimensional feature encoding in this embodiment captures the various features of Chinese characters more comprehensively, so that the target Chinese character can be located faster in the decoding stage, significantly reducing the duplicate code rate for users in the actual Chinese character input process, improving the matching accuracy, and also conforming to the user's memory habit of Chinese characters and improving the user's input experience.
[0040] In the prior art, for new encoding types (such as tones, structures, strokes, etc.), there are usually the following methods: one is to specially set up an independent area for inputting this type of encoding (for example, many patent applications will specially set up several independent keys for inputting tones), and the other is to add a new state for inputting this type of encoding. The former requires setting independent keys, which is not applicable to mobile terminals with limited space and will result in low key utilization rate and longer encoding; the latter lacks flexibility and cannot achieve "input on demand" because it is usually difficult to skip a certain state during the input process. Therefore, the embodiments of the present application provide a dynamic input area (also called an auxiliary input area). Different from the former two, the dynamic input area does not require setting independent keys, nor does it require adding a new state for the entire keyboard. Instead, a certain area is specially designated in a certain state. This area has independent sub-states. After this area enters the input state, clicking on a key will only change the sub-state, will not change the state of the main keyboard, and will not affect the input of the main code. Clicking on a key outside the dynamic area will switch the keyboard to the next state.
[0041] Advantages of the dynamic input area: 1. There is no need to set up a separate input area outside the main keyboard, which will not make the keyboard bloated; 2. There is no need to add a dedicated keyboard state separately, which will not increase the difficulty of coding parsing; 3. High flexibility, with the main code and auxiliary code separated. The auxiliary code can be input as needed without affecting the input of the main code; 4. Overcome the thinking lag caused by "combined input". The so-called "combined input" refers to an input method that combines two different elements and maps them to the keyboard. For example, combining the second letter (without tone) of the simple code of the finals with the last stroke of the Chinese character. Clicking a key inputs both the letter and the last stroke at the same time. This input method is more efficient, but during the input process, it is necessary to combine the two elements in the mind and then correspond them to the keys, which takes time to think and the input experience is poor; 5. Support an infinite number of auxiliary codes (of course, the types of auxiliary codes should not be too many to avoid overly long codes); Steps for setting up the dynamic input area: 1) Specify the state corresponding to the dynamic input area; 2) Leave fixed blank keys (without mapping relationships) in the specified state; 3) Set up sub-states for the dynamic area. Clicking the keys in the dynamic area only changes the sub-states, but does not change the keyboard state and does not affect the normal input of other keys.
[0042] In a preferred embodiment, the keyboard state corresponding to the dynamic input area of the present application is the first state (it can also be set to any keyboard state to trigger the dynamic input area); the keys in the dynamic area are the three rightmost keys (1-8, 2-8, 3-8) in the first state, and these three keys are blank keys; the dynamic area has two sub-states: State 1-1, corresponding to the starting stroke of Chinese characters; State 1-2, corresponding to the classification by the number of components.
[0043] Regarding the starting stroke of Chinese characters, it refers to the first stroke of a Chinese character, which is divided into five categories: horizontal (一), vertical (丨), left-falling stroke (丿), dot (丶), and fold (乁乚). Chinese characters can be divided into 5 categories according to the starting stroke. For example, Chinese characters with a starting stroke of "horizontal" include "南", "苏", "下", "有", etc.; Chinese characters with a starting stroke of "vertical" include "日", "果", "上", "克", etc.; Chinese characters with a starting stroke of "left-falling stroke" include "个", "仆", "白", "句", etc.; Chinese characters with a starting stroke of "dot" include "宋", "谈", "送", "广", etc.; Chinese characters with a starting stroke of "fold" include "丝", "九", "予", "旨", "好", etc.
[0044] Regarding the classification of Chinese characters by the number of components, the number of components of a Chinese character does not refer to the total number of independent components of the entire Chinese character, but rather the number of components in two dimensions, horizontal and vertical. Classifying by the number of components means classifying Chinese characters from two dimensions: the number of horizontal components and the number of vertical components of a Chinese character, denoted as x-y, where x represents the number of horizontal components and y represents the number of vertical components. Chinese characters can be classified into the following categories: If a Chinese character can be horizontally divided into 2 or more independent components, and one of the left and right sides can be vertically divided into 2 or more independent components, it is recorded as 2-2. For example, "Sui", "Shao", "Shao", etc. If a Chinese character can be horizontally divided into 2 or more independent components, and neither the left nor the right side can be vertically divided into 2 or more independent components, it is recorded as 2-1. For example, "Chong", "Pai", "Jue", etc. If a Chinese character cannot be horizontally divided into 2 or more independent components, but can be vertically divided into 2 or more independent components, it is recorded as 1-2. For example, "Gao", "Gao", "Hua", etc. If a Chinese character cannot be horizontally divided into more than two independent components, nor can it be vertically divided into 2 or more components, it is recorded as 1-1. For example, "Ding", "Guo", "Guo", etc. There are 5 types of starting strokes for Chinese characters, while there are only 3 buttons in the dynamic area. Therefore, some strokes need to be combined. The corresponding relationships are as follows: Keys 1-8 correspond to "one" and "vertical", Keys 2-8 correspond to "slash" and "dot", and Keys 3-8 correspond to "hook" (fold), as shown below: <00001... Mapping relationship between the dynamic area and the classification by the number of components According to the above definitions, classifying Chinese characters from the number of components in two dimensions, horizontal and vertical, Chinese characters can be divided into four categories: 2-2 category, 2-1 category, 1-2 category, and 1-1 category. Since there are only three buttons in the dynamic area, some classifications need to be combined. According to the applicant's statistics on the quantities of the above four categories, it is found that the quantities of Chinese characters in the 1-2 category and the 1-1 category are relatively small, so the two are combined. Using I to represent Chinese characters in the 1-2 category and the 1-1 category; using II to represent Chinese characters in the 2-1 category; using III to represent Chinese characters in the 2-2 category, the corresponding relationships are as follows: Keys 1-8 correspond to category I Chinese characters, Keys 2-8 correspond to category II Chinese characters, and Keys 3-8 correspond to category III Chinese characters, as shown below Therefore, as shown in the following table, the complete encoding of a Chinese character consists of the initials short code, finals short code, first stroke, and the classification by the number of components. During the actual encoding process, users can omit other encodings except the initials short code as needed to improve the input efficiency, and this embodiment can still proceed with subsequent decoding and Chinese character display as normal. In a preferred embodiment, examples of Chinese character encoding are as follows: 1. Encoding method for a single Chinese character: As described above, the encoding composition of Chinese characters is: initial consonant abbreviation + final abbreviation + component number classification code + first stroke of the Chinese character. The following are some examples of Chinese character encoding: The pinyin of the character "张" is zhāng. Its initial consonant abbreviation is z, corresponding to the key Z; the final abbreviation is ān, where ā corresponds to the key Q, n corresponds to the key N, the first stroke is "折", corresponding to the key Π, and the component number classification code is 2-1, corresponding to the key Φ; so its complete key encoding is ZQNΠΦ. As Figure 2 shown, when inputting the encoding, first trigger the key 3-1 in the first keyboard state to generate the initial consonant abbreviation Z, then successively trigger the keys 1-1 and 3-5 in the second keyboard state to generate the final abbreviation an. At this time, switch the keyboard state to the auxiliary input state by triggering the key 3-8 in the auxiliary input area and generate the corresponding starting stroke encoding of the Chinese character. Finally, trigger the key 2-8 in the auxiliary input state to generate the corresponding Chinese character feature encoding, thereby completing the encoding of the Chinese character "张". At the same time, the system decodes and retrieves according to the input encoding and generates several candidate Chinese characters to be displayed at the specified position.
[0045] The pinyin of the character "黄" is huáng. Its initial consonant abbreviation is h, corresponding to the key H; the final abbreviation is uá, where u is a vowel without tone. After adding the tone mark without tone, it becomes û, corresponding to the key N, á corresponds to the key W, the first stroke is "一", corresponding to the key Λ, and the component number classification code is 1-2, corresponding to the key Λ; so its complete key encoding is HNWΛΛ; The pinyin of the character "僧" is sēng. Its initial consonant abbreviation is S, corresponding to the key S; the final abbreviation is ēn, where ē corresponds to the key Y, n corresponds to the key N; the first stroke is "丿" corresponding to the key Φ; the component number classification is 2-2, corresponding to the key Π; so its complete key encoding is SYNΦΠ; The pinyin of the character "同" is tóng. Its initial consonant abbreviation is T, corresponding to the key T; the final abbreviation is ón, where ó corresponds to the key J, n corresponds to the key N; the first stroke is "丨", corresponding to the key Λ; the component number classification is 1-1, corresponding to the key Λ; so its complete key encoding is TJNΛΛ; The pinyin of the character "福" is fú. Its initial consonant abbreviation is F, corresponding to the key F; the final abbreviation is úv, where ú corresponds to the key X, v corresponds to the key V; the first stroke is "丶", corresponding to the key Φ, and the component number classification is 2-2, corresponding to the key Π; so its complete key encoding is FXVΦΠ; The pinyin of the character "闯" is chuǎng. Its initial consonant abbreviation is C, corresponding to the key C; the final abbreviation is uǎ, where u is a vowel without tone and should be transcribed as û, corresponding to the key N, ǎ corresponds to the key R; the component number classification is 1-1, corresponding to the key Λ; the first stroke is "丶", corresponding to the key Φ, so its complete key encoding is CNRΛΦ.
[0046] The pinyin of the character "穹" is qióng. Its initial consonant simple code is Q, corresponding to the key Q; the final simple code is ió, where i is a vowel without tone and should be transcribed as î, corresponding to the key V, and ó corresponds to the key J; the first stroke is "丶", corresponding to the key Φ, and the component number classification is 1-2, corresponding to the key Λ; so its complete key code is QVJΦΛ.
[0047] The pinyin of the character "窃" is qiè. Its initial consonant simple code is Q, corresponding to the key Q; the final simple code is iè, where i is a vowel without tone and should be transcribed as î, corresponding to the key V, and è corresponds to the key Λ; the first stroke is "丶", corresponding to the key Φ, and the component number classification is 1-2, corresponding to the key Λ; so its complete key code is QVΛΦΛ.
[0048] The pinyin of the character "敌" is dí. Its initial consonant simple code is D, corresponding to the key D; the final simple code is ív, where í corresponds to the key D and v corresponds to the key V; the first stroke is "丿", corresponding to the key Φ; the component number classification is 2-1, corresponding to the key Φ; so its complete key code is DDVΦΦ.
[0049] 2. Lexical Encoding Method Use C1, C2, C3, C4, and C5 to represent the five encodings of Chinese characters respectively. For example, if the encoding of the character "敌" is DXVΦΦ, then C1 corresponds to D, C2 corresponds to X, and so on.
[0050] 2.1 Two-character Lexical Encoding Rule The encoding of a two-character lexicon is: C1C2C3C1C2C3, that is, the first three encodings of the two Chinese characters are combined together. For example, the encodings of the two Chinese characters "既是" are JXMΦΠ and SBMΛΛ respectively, then the encoding of this word is JXMSBM. Due to the strong flexibility of the dynamic input area of this input method, when higher precision is required, after inputting the first three codes, the subsequent two codes can be continued to be input, that is, JXV SBVΛΛ, or JXVΦΠ SBV, or JXVΦΠ SBVΛΛ, and the user can input the last two codes according to needs.
[0051] 2.2 Three-character Lexical Encoding Rule There are two encoding rules for three-character lexicons: C1C2C3C1C2C3C1C2C3, that is, the first three encodings of each Chinese character can be input in sequence. After each input of the first three encodings of a Chinese character, it can be determined whether to input the supplementary codes C4 and C5 according to needs; C1C2C1C2C1C2, that is, only the first two codes of the first two Chinese characters are taken. For example, the complete encodings of the three Chinese characters "幼儿园" are YΦZΠΦ, VPV ΦΛ, and YNWΛΛ respectively, then its encoding is YΦ VP YZ. Usually, as long as the first 6 encodings are input, the three-character lexicon "幼儿园" can be matched; 2.3. Four-character vocabulary encoding. There are two encoding methods for four-character vocabulary: C1C1C1C1C2C3. For the first three Chinese characters, take the first letter of each character's encoding. For the last Chinese character, take the first three letters of its encoding. And C4 and C5 can be input as needed. For example, for the first three Chinese characters of "east, south, west, north", the first letters of their encodings are D, N, X, and for the last Chinese character, it is B, L, S, Λ, Φ. Then its complete encoding is D, N, X, B, L, S, and the auxiliary codes Φ and Λ can be input as needed; C1C2C1C2C1C2C1C2. Take the first two letters of each Chinese character's encoding. For example, for each Chinese character of "as sharp as a hawk's eye", the main codes are: H, N, G, Y, R, N, R, X, J, B, V. So its encoding is H, N, Y, R, R, X, J, B; 2.4. Encoding for five-character and above vocabulary. There are two encoding methods for five-character and above vocabulary: C1C1C1C1C1. That is, take the first letter of each Chinese character's encoding. For example, the corresponding encoding for "circumstances create heroes" is S, S, Z, Y, X; C1C2C1C2C1C2C1C2.... That is, take the first two letters of each Chinese character's encoding. For example, for each Chinese character of "more haste, less speed", the main codes are: Y, B, V, S, B, V, Z, P, V, B, B, V, D, W, V. So its encoding is Y, B, S, B, Z, P, B, B, D, W.
[0052] In a possible implementation, the first character and the second character in the simple code of the final consonant are respectively generated by being triggered in different input areas. Specifically: When the keyboard state is the second state, generate the corresponding first character according to the second position triggered on the keyboard; In any keyboard state, if the currently triggered third position belongs to the auxiliary input area of the keyboard, generate the corresponding second character according to the third position, and switch the keyboard state to the auxiliary input state; When the keyboard state is the auxiliary input state, generate the corresponding starting pen code and character feature code of the Chinese character according to the fourth position and the fifth position triggered in the auxiliary input area in sequence; Generate the simple code of the final consonant according to the first character and the second character.
[0053] This application embodiment defines the specific application method of the auxiliary input area. By reusing the auxiliary input area on the keyboard, the second character of the vowel abbreviation, the stroke code of the Chinese character, and the feature code of the Chinese character are input, without the need to add separate dedicated keys for different types of codes. This greatly optimizes the keyboard layout, enabling richer input functions to be carried within the limited screen (especially mobile device screens) or keyboard space, maintaining a clear and concise input interface, and avoiding the bloated interface and visual burden on users caused by too many keys. In addition, this embodiment allows users to flexibly trigger the auxiliary input area as needed during or after inputting the main code (such as the initial and the first part of the vowel) to supplement the vowel supplementary information and structural feature information (second part of the vowel, stroke, component number classification) of the input Chinese character, improving the matching accuracy between the code and the Chinese character. This design makes the encoding input process more coherent and natural, significantly improving the user's input experience and operational smoothness.
[0054] In summary, this embodiment solves the problem of inputting characters with tone marks. By trimming and supplementing the pinyin, the pinyin encoding is made consistent in length and regular, with the tone marks fixed at the second and third positions. Existing pinyin input methods generally do not discard any encoding, resulting in inconsistent encoding lengths and a lack of in-depth exploration of the rules governing tone input in vowels. Furthermore, this embodiment introduces two Chinese character features: stroke order and structural number classification, greatly improving accuracy. Introducing these two features does not affect the input of the main code; the character features can be input as needed (pinyin as the main part, character features as the auxiliary part; during encoding parsing, the main and auxiliary parts are clearly distinguishable and will not be confused). Existing technologies lack the "dynamic input area" found in this embodiment, which may lead to confusion between the pinyin and other encodings during encoding parsing.
[0055] Example 2: like Figure 3 As shown, Embodiment 2 provides a Chinese character decoding method based on Pinyin and Chinese character features. This method is used to decode Chinese character codes generated by any of the Chinese character encoding methods based on Pinyin and Chinese character features described in this application, and includes steps S101-S501: Step S101: Obtain the Chinese character encoding; Step S201: Based on the position of the initial consonant abbreviation in the Chinese character encoding, the Chinese character encoding is split into several sub-encoding information, wherein there is at most one initial consonant abbreviation in the sub-encoding information, and the initial consonant abbreviation is the first code in the sub-encoding information; Step S301: Based on each of the sub-encoding information, generate corresponding retrieval information through a preset mapping table, including generating several complete initials based on the initial consonant abbreviation; if the sub-encoding information contains a vowel abbreviation, generate several complete vowels based on the vowel abbreviation; if the sub-encoding information contains a Chinese character starting stroke code, generate corresponding Chinese character starting stroke information based on the Chinese character starting stroke code; if the sub-encoding information contains a Chinese character feature code, generate corresponding Chinese character feature information based on the Chinese character feature code. Step S401: For each sub-encoding information, retrieve several candidate Chinese characters from a preset database according to the corresponding retrieval information; Step S501: If the number of sub-encoding information is 1, then the candidate Chinese characters are displayed on the preset interface.
[0056] This application provides a Chinese character decoding method corresponding to the encoding method. First, considering the situation where a user inputs multiple Chinese characters simultaneously, the initial consonant abbreviation position splits the Chinese character encoding, which may contain multiple characters, into several sub-encoding information, and decodes each sub-encoding information separately, achieving rapid conversion from abbreviation to complete information. Specifically, during the information decoding process, the presence of final consonant abbreviations, character stroke start codes, and character feature codes in the sub-encoding information is checked sequentially, generating corresponding complete final consonant, character stroke start, and character feature information. In other words, this embodiment can still complete the Chinese character decoding process and display the Chinese character matching results even when one or more of the final consonant abbreviations, character stroke start codes, and character feature codes are missing, improving the flexibility of user Chinese character input. Second, the decoded retrieval information can cover multiple dimensions such as initial consonants, final consonants, stroke start codes, and features, making database queries more accurate, the returned results more targeted, and improving the matching accuracy between encoding and Chinese characters.
[0057] Furthermore, in step S301, generating a plurality of corresponding complete initials based on the initial consonant abbreviation includes: If the initial consonant abbreviation is a character in a preset first character set, then a blank placeholder is generated as the complete initial consonant in the search information; If the initial consonant abbreviation is a character in a preset second character set, then the flat tongue initial consonant and the retroflex initial consonant corresponding to the initial consonant abbreviation are generated as the several complete initial consonants.
[0058] In this embodiment, by classifying and processing the initial consonant abbreviations, the system can more flexibly adapt to the user's input habits and parse the corresponding initial consonant information. For example, in some cases, the user may not need to input a specific initial consonant or some Chinese characters may not have an initial consonant. In this case, the user inputs a character from the first character set during the encoding stage. At this time, a blank placeholder is generated during the decoding stage as the complete initial consonant in the search information, ensuring that each Chinese character has corresponding initial consonant information and ensuring that subsequent searches proceed normally. Secondly, for some easily confused alveolar and retroflex initial consonants, this embodiment can generate the corresponding alveolar and retroflex initial consonants based on the initial consonant abbreviation of a single character input by the user. This not only reduces the number of characters the user needs to input and improves input efficiency, but also significantly improves the error tolerance of Chinese character input. For example, users may have inaccurate pronunciation due to differences in dialect or accent. This design ensures that reasonable initial consonant abbreviations can be generated even with incompletely accurate input by using a preset character set mapping relationship. This enhances the robustness of the system, effectively reduces the number of keys, optimizes the keyboard layout, and improves the simplicity and aesthetics of the input interface, providing users with a more convenient and efficient interactive experience.
[0059] Furthermore, in step S501, if the number of sub-encoding information is greater than 1, then according to the preset word library and the order of each sub-encoding information in the Chinese character encoding, the corresponding candidate Chinese characters are arranged and combined to generate several words or phrases, and each word or phrase is displayed on the preset interface.
[0060] This embodiment further considers the scenario where a user simultaneously inputs multiple Chinese characters, resulting in a combined display. Since each sub-encoding information can match multiple Chinese characters, when only one sub-encoding information exists, the matched characters can be directly displayed. However, when multiple sub-encoding information exists, the number of possible combinations of the matched characters increases significantly. Users need to select the desired character from a massive number of matching results, which greatly reduces the user experience. Therefore, this embodiment intelligently combines candidate characters into a coherent set of words or phrases by integrating a preset dictionary and the order information of each sub-encoding. This significantly reduces the number of matches, ensuring that users can quickly select the desired character, improving the matching accuracy between encoding and characters, and enhancing the user's input experience.
[0061] Example 3: like Figure 4 As shown, Embodiment 3 provides a Chinese character encoding system based on Chinese Pinyin and Chinese character features, including an initial consonant abbreviation generation module 10, a final vowel abbreviation generation module 20, a Chinese character stroke start encoding generation module 30, a Chinese character feature encoding generation module 40, and a combination module 50. The initial consonant abbreviation generation module 10 is used to generate the corresponding initial consonant abbreviation based on the first position triggered on the keyboard when the keyboard state is the first state. The vowel abbreviation generation module 20 is used to generate corresponding vowel abbreviations according to the second and third positions triggered on the keyboard and the triggering order when the keyboard is in the second state. The Chinese character stroke start code generation module 30 is used to generate a corresponding Chinese character stroke start code based on the fourth position in any keyboard state if the currently triggered fourth position belongs to the auxiliary input area of the keyboard, and switch the keyboard state to auxiliary input state. The Chinese character feature encoding generation module 40 is used to generate the corresponding Chinese character feature encoding based on the fifth position triggered in the auxiliary input area when the keyboard state is auxiliary input state. The combination module 50 is used to combine the generated initial consonant abbreviation, the final vowel abbreviation, the Chinese character stroke start code, and the Chinese character feature code into a Chinese character code according to the encoding generation order when the input is completed.
[0062] Furthermore, when the keyboard is in the second state, the vowel abbreviation generation module 20 generates corresponding vowel abbreviations based on the triggered second and third positions on the keyboard and the triggering order, including: Generate the corresponding second and third characters based on the second and third positions on the keyboard; Combine the second and third characters to generate a combined character; The tone and tone position of the combined characters are determined according to the triggering order, and then the corresponding vowel abbreviation is generated.
[0063] Furthermore, the stroke code of the Chinese character corresponds to a stroke information of the Chinese character, which is "horizontal", "vertical", "left-falling", "dot" and "turning".
[0064] Furthermore, the Chinese character feature encoding corresponds to a type of Chinese character feature information, which is a first feature, a second feature, a third feature, or a fourth feature; The first feature is that the target Chinese character to be encoded can be split into two or more first independent components in the left and right directions, and any one of the first independent components can be split into two or more second independent components in the top and bottom directions. The second feature is that the target Chinese character can be split into two or more first independent components in the left and right directions, and each of the first independent components cannot be split into two or more second independent components in the right and left directions. The third feature is that the target Chinese character cannot be split horizontally into two or more first independent components, and the target Chinese character can be split vertically into two or more second independent components. The fourth feature is that the target Chinese character cannot be split horizontally or vertically.
[0065] In one possible implementation, the vowel abbreviation generation module includes a first character generation unit, a second character generation unit, and a combination unit, which are respectively used to generate the first character, the second character, and the vowel abbreviation in the vowel abbreviation, specifically: When the keyboard state is the second state, the first character generation unit generates the corresponding first character according to the second position triggered on the keyboard. In any keyboard state, if the currently triggered third position belongs to the auxiliary input area of the keyboard, the second character generation unit generates the corresponding second character according to the third position and switches the keyboard state to the auxiliary input state. When the keyboard is in auxiliary input mode, the Chinese character stroke encoding generation module 30 and the Chinese character feature encoding generation module 40 generate corresponding Chinese character stroke encoding and Chinese character feature encoding based on the fourth and fifth positions triggered in the auxiliary input area, respectively. The combining unit generates the vowel abbreviation based on the first character and the second character.
[0066] This application provides a Chinese character encoding system based on Pinyin and Chinese character features. It generates initial consonant abbreviations, final vowel abbreviations, stroke order codes, and character feature codes under different keyboard states, and combines them sequentially to form a complete Chinese character encoding. This embodiment effectively integrates the Pinyin and structural features of Chinese characters, enabling the encoding to reflect both pronunciation and character shape, thus significantly reducing the homophone rate. For example, traditional Pinyin input methods have numerous homophones, requiring users to frequently flip through pages to select the target character. This embodiment, by introducing stroke order information and structural features, and distinguishing between initial consonants and final vowels in the Pinyin encoding, can quickly filter out highly probable candidate characters in the subsequent decoding stage, greatly reducing user search time and improving the matching accuracy between the encoding and the character. Secondly, this embodiment employs a state-based dynamic encoding mechanism, allowing the same key to generate different codes under different input states, significantly reducing the space occupied by the keyboard layout and improving the simplicity and aesthetics of the input interface. Furthermore, this embodiment also includes an auxiliary input area. Users can trigger the auxiliary input area at any time to encode the structural features of the current Chinese character as needed, eliminating the need for frequent mode switching during input, resulting in smooth and natural operation and improved user input experience. Finally, because the encoding rules take into account both the naturalness of Pinyin and the structural nature of Chinese characters, the learning curve is low, allowing users to master the usage method in a short time, further enhancing the user input experience.
[0067] Example 4: like Figure 5 As shown, Embodiment 4 provides a Chinese character decoding system based on Pinyin and Chinese character features. The Chinese character decoding system is used to decode the Chinese character encoding generated by any Chinese character encoding system based on Pinyin and Chinese character features as described in this application, including an acquisition module 101, a splitting module 102, an information decoding module 103, a retrieval module 104, and a display module 105. The acquisition module 101 is used to acquire Chinese character encoding; The splitting module 102 is used to split the Chinese character encoding into several sub-encoding information according to the position of the initial consonant abbreviation in the Chinese character encoding, wherein the sub-encoding information contains at most one initial consonant abbreviation, and the initial consonant abbreviation is the first code in the sub-encoding information; The information decoding module 103 is used to generate corresponding retrieval information according to each of the sub-encoding information through a preset mapping table, including generating several complete initials according to the initial consonant abbreviation; if there is a final vowel abbreviation in the sub-encoding information, generating several complete finals according to the final vowel abbreviation; if there is a Chinese character starting stroke code in the sub-encoding information, generating corresponding Chinese character starting stroke information according to the Chinese character starting stroke code; if there is a Chinese character feature code in the sub-encoding information, generating corresponding Chinese character feature information according to the Chinese character feature code. The retrieval module 104 is used to retrieve a number of candidate Chinese characters from a preset database for each sub-encoding information according to the corresponding retrieval information. The display module 105 is used to display the candidate Chinese characters on a preset interface if the number of sub-encoding information is 1.
[0068] Furthermore, the information decoding module 103 generates several corresponding complete initials based on the initial consonant abbreviation, including: If the initial consonant abbreviation is a character in a preset first character set, then a blank placeholder is generated as the complete initial consonant in the search information; If the initial consonant abbreviation is a character in a preset second character set, then the flat tongue initial consonant and the retroflex initial consonant corresponding to the initial consonant abbreviation are generated as the several complete initial consonants.
[0069] Furthermore, if the number of sub-encoding information is greater than 1, the display module 105 arranges and combines the corresponding candidate Chinese characters according to the preset word library and the order of each sub-encoding information in the Chinese character encoding, generates several words or phrases, and displays each word or phrase on the preset interface.
[0070] This application provides a Chinese character decoding system corresponding to the encoding method. Firstly, considering the scenario where a user inputs multiple Chinese characters simultaneously, the initial consonant abbreviation position splits the Chinese character encoding, which may contain multiple characters, into several sub-encoding information. Each sub-encoding information is then decoded, achieving rapid conversion from abbreviation to complete information. Specifically, during the information decoding process, the presence of a vowel abbreviation, the character's starting stroke code, and the character's feature code are checked sequentially in the sub-encoding information, generating corresponding complete vowel, starting stroke, and feature information. This means that this embodiment can still complete the Chinese character decoding process and display the matching results even when one or more of these components are missing, improving the flexibility of the user's Chinese character input. Secondly, the decoded retrieval information can cover multiple dimensions such as initial consonants, vowels, starting strokes, and features, making database queries more accurate and the returned results more targeted, thus improving the matching accuracy between the encoding and the Chinese character.
[0071] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.
Claims
1. A Chinese character encoding method based on Pinyin and Chinese character features, characterized in that, include: When the keyboard is in the first state, the corresponding initial consonant abbreviation is generated according to the first position triggered on the keyboard. When the keyboard is in the second state, the corresponding vowel abbreviation is generated according to the second and third positions triggered on the keyboard and the triggering order. In any keyboard state, if the currently triggered fourth position belongs to the auxiliary input area of the keyboard, then the corresponding Chinese character starting stroke code is generated according to the fourth position, and the keyboard state is switched to auxiliary input state. When the keyboard is in auxiliary input mode, the corresponding Chinese character feature code is generated based on the fifth position triggered in the auxiliary input area. When the input is complete, the generated initial consonant abbreviation, final vowel abbreviation, Chinese character stroke code, and Chinese character feature code are combined to construct the Chinese character code according to the encoding generation order.
2. The Chinese character encoding method based on Pinyin and Chinese character features as described in claim 1, characterized in that, When the keyboard is in the second state, the corresponding vowel abbreviation is generated based on the second and third positions triggered on the keyboard and the triggering order, including: The second and third characters are generated according to the second and third positions on the keyboard and the triggering order; The tone and tone position are determined based on the second and third characters; Based on the tones and positions of the second and third characters, the corresponding vowel abbreviations are generated.
3. The Chinese character encoding method based on Pinyin and Chinese character features as described in claim 1, characterized in that, The stroke code of the Chinese character corresponds to a stroke information of the Chinese character, which is "horizontal", "vertical", "left-falling", "dot" and "turning".
4. The Chinese character encoding method based on Pinyin and Chinese character features as described in claim 1, characterized in that, The Chinese character feature encoding corresponds to a Chinese character feature information, which is a first feature, a second feature, a third feature, or a fourth feature. The first feature is that the target Chinese character to be encoded can be split into two or more first independent components in the left and right directions, and any one of the first independent components can be split into two or more second independent components in the top and bottom directions. The second feature is that the target Chinese character can be split into two or more first independent components in the left and right directions, and each of the first independent components cannot be split into two or more second independent components in the right and left directions. The third feature is that the target Chinese character cannot be split horizontally into two or more first independent components, and the target Chinese character can be split vertically into two or more second independent components. The fourth feature is that the target Chinese character cannot be split horizontally or vertically.
5. The Chinese character encoding method based on Pinyin and Chinese character features as described in claim 1, characterized in that, The first and second characters in the abbreviated vowel codes are generated by triggering different input regions, specifically: When the keyboard is in the second state, the corresponding first character is generated according to the second position triggered on the keyboard. In any keyboard state, if the currently triggered third position belongs to the auxiliary input area of the keyboard, then the corresponding second character is generated according to the third position, and the keyboard state is switched to auxiliary input state; When the keyboard is in auxiliary input mode, the corresponding Chinese character stroke start code and Chinese character feature code are generated sequentially according to the fourth and fifth positions triggered in the auxiliary input area. The abbreviation of the vowel is generated based on the first character and the second character.
6. A Chinese character decoding method based on Pinyin and Chinese character features, characterized in that, The Chinese character decoding method is used to decode the Chinese character encoding generated by the Chinese character encoding method based on Chinese Pinyin and Chinese character features as described in any one of claims 1 to 5, including: Get Chinese character encoding; Based on the position of the initial consonant abbreviation in the Chinese character encoding, the Chinese character encoding is split into several sub-encoding information, wherein there is at most one initial consonant abbreviation in the sub-encoding information, and the initial consonant abbreviation is the first code in the sub-encoding information; Based on each of the sub-encoding information, corresponding retrieval information is generated through a preset mapping table, including generating several complete initials based on the initial consonant abbreviation; if the sub-encoding information contains a vowel abbreviation, several complete vowels are generated based on the vowel abbreviation; if the sub-encoding information contains a Chinese character starting stroke code, corresponding Chinese character starting stroke information is generated based on the Chinese character starting stroke code; if the sub-encoding information contains a Chinese character feature code, corresponding Chinese character feature information is generated based on the Chinese character feature code. For each of the sub-encoding information, several candidate Chinese characters are retrieved from a preset database according to the corresponding retrieval information; If the number of sub-encoding information is 1, then the candidate Chinese characters will be displayed on the preset interface.
7. The Chinese character decoding method based on Pinyin and Chinese character features as described in claim 6, characterized in that, The step of generating several complete initials based on the initial code includes: If the initial consonant abbreviation is a character in a preset first character set, then a blank placeholder is generated as the complete initial consonant in the search information; If the initial consonant abbreviation is a character in a preset second character set, then the flat tongue initial consonant and the retroflex initial consonant corresponding to the initial consonant abbreviation are generated as the several complete initial consonants.
8. The Chinese character decoding method based on Chinese Pinyin and Chinese character features as described in claim 6, characterized in that, If the number of sub-encoding information is greater than 1, then according to the preset dictionary and the order of each sub-encoding information in the Chinese character encoding, the corresponding candidate Chinese characters are arranged and combined to generate several words or phrases, and each word or phrase is displayed on the preset interface.
9. A Chinese character encoding system based on Pinyin and Chinese character features, characterized in that, It includes a module for generating initial consonant abbreviation codes, a module for generating final vowel abbreviation codes, a module for generating Chinese character stroke start codes, a module for generating Chinese character feature codes, and a combination module; The initial consonant abbreviation generation module is used to generate the corresponding initial consonant abbreviation based on the first position triggered on the keyboard when the keyboard state is in the first state. The vowel abbreviation generation module is used to generate corresponding vowel abbreviations based on the second and third positions triggered on the keyboard and the triggering order when the keyboard is in the second state. The Chinese character starting stroke code generation module is used to generate the corresponding Chinese character starting stroke code based on the fourth position in any keyboard state if the currently triggered fourth position belongs to the auxiliary input area of the keyboard, and switch the keyboard state to auxiliary input state. The Chinese character feature encoding generation module is used to generate the corresponding Chinese character feature encoding based on the fifth position triggered in the auxiliary input area when the keyboard state is auxiliary input state. The combination module is used to construct a Chinese character code by combining the generated initial consonant abbreviation, the final vowel abbreviation, the Chinese character stroke code, and the Chinese character feature code according to the encoding generation order when the input is completed.
10. A Chinese character decoding system based on Pinyin and Chinese character features, characterized in that, The Chinese character decoding system is used to decode the Chinese character encoding generated by the Chinese character encoding system based on Chinese Pinyin and Chinese character features as described in claim 9, and includes an acquisition module, a splitting module, an information decoding module, a retrieval module, and a display module; The acquisition module is used to acquire Chinese character encodings; The splitting module is used to split the Chinese character encoding into several sub-encoding information according to the position of the initial consonant abbreviation in the Chinese character encoding. The sub-encoding information contains at most one initial consonant abbreviation, and the initial consonant abbreviation is the first code in the sub-encoding information. The information decoding module is used to generate corresponding retrieval information based on each of the sub-encoding information through a preset mapping table, including generating several complete initials based on the initial consonant abbreviation; if the sub-encoding information contains a vowel abbreviation, generating several complete vowels based on the vowel abbreviation; if the sub-encoding information contains a Chinese character starting stroke code, generating corresponding Chinese character starting stroke information based on the Chinese character starting stroke code; if the sub-encoding information contains a Chinese character feature code, generating corresponding Chinese character feature information based on the Chinese character feature code. The retrieval module is used to retrieve a number of candidate Chinese characters from a preset database for each sub-encoding information based on the corresponding retrieval information. The display module is used to display the candidate Chinese characters on a preset interface if the number of sub-encoding information is 1.