First-right pinyin input method
By using the first right pinyin input method and auxiliary code encoding rules in the Chinese character input method, the component selection and keyboard layout are optimized, and the problems of irregular parts, excessive code length and high code repetition rate in the existing Chinese character input method are solved, and quick and simple Chinese character input is achieved.
Patent Information
- Application Number
- CN202510142296.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-29
- Publication Date
- 2025-06-13
AI Technical Summary
The existing Chinese character input method has problems such as irregular parts, unreasonable radical selection, excessive code length, high code repetition rate, and slow input speed, making it difficult to achieve simple and fast Chinese character input.
The first right pinyin input method is adopted, and the encoding is composed of the sound code and the auxiliary code. The auxiliary code includes the shape encoding method and the word head radical method. It is preferred to have 21 multi-stroke parts and five basic strokes. The components are arranged using the homophone proximity method, and quantitative calculations are performed to optimize the keyboard layout.
It realizes the standardization and intuitiveness of the selection and layout of Chinese character stroke components, simplifies the learning curve of the input method, improves the input speed and accuracy, reduces the re-code rate, and makes input Chinese characters simple and fast.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to computer Chinese character encoding input methods. After the inventor invented a new simultaneous sound and near-position Chinese character code input method that can be learned in a few minutes, in order to facilitate promotion, it was renamed the first-right pinyin input method. After receiving the statement from a certain input method company that they were not willing to learn it for even a minute, based on the first-right pinyin input method, further improvements were made, and Chinese characters can be quickly input in just one minute, so it is called the first-right pinyin input method. Of course, it also involves the keyboard for implementing this input method. Background Art
[0002] Keyboard input methods are the most widely used input methods in current Chinese character input methods. Keyboard input can be divided into three categories according to encoding: sound codes, shape codes, and sound-shape codes.
[0003] Sound codes generally use Chinese pinyin as the basis and encode using the pronunciation of Chinese characters. Shape codes encode using the glyph features of Chinese characters. Sound-shape codes encode using the phonetic features and glyph features of Chinese characters. Sound-shape codes can be further divided into two categories: sound-shape codes that only use initials and sound-shape codes that use the entire sound code of Chinese characters. If the initials and finals of Chinese characters are fully utilized and the specified sound code part is placed first and the auxiliary code is placed later, it basically does not affect thinking, the thinking is similar to that of sound codes, the homophone rate is similar to that of shape codes, and it is compatible with pinyin, increasingly showing its superiority. The sound-shape codes invented by others currently often have more Chinese character components or a higher homophone rate, while the Chinese character code invented by the inventor, on the basis of innovative encoding rules, only uses about 21 radicals and 5 basic strokes, and it only takes a few minutes to input Chinese characters simply and quickly. The 26 stroke radicals correspond one by one with the 26 English letters, which is also convenient for display on small-screen keys such as mobile phones.
[0004] When promoting the inventor's input method, it was also found that in order to avoid homophones, 10 common radicals are not encoded according to the first letter of pinyin, but are cleverly arranged according to the simultaneous sound and near-position method, which takes a few minutes to memorize. And some users are used to encoding these radicals and some other common radicals according to the first letter of pinyin. Therefore, the ease of learning still needs to be improved. Summary of the Invention
[0005] The current Chinese character input methods or Chinese character components are not standardized or the number of selected Chinese character components is not very reasonable; or the radicals, i.e., Chinese character components, do not fully consider the frequency of character formation, practical frequency, and homophone rate in commonly used Chinese characters; or the positions of the five basic strokes on the keyboard are unreasonable, which is likely to cause the same code for words and characters; or the code length is too long; or the homophone rate is too high, affecting the input speed; or only the initials or the first letters of the pinyin of Chinese characters are used; or it is not intuitive enough; or the code-taking rules are not very reasonable, which will affect the mental reaction; or when taking the code, it is necessary to constantly distinguish whether it is a left-right structure or sometimes take the code horizontally and sometimes vertically; or the arrangement of Chinese character components on the keyboard is not very regular, or even a bit far-fetched; or there is no quantitative calculation for multi-stroke components, and the selection, abandonment, and arrangement on the keyboard rely on empirical intuition; or the usage frequency of radicals is not considered enough with relatively authoritative big data statistical materials, that is, the actual frequency is not considered enough; or there is no multi-code-taking method with encoding by the first letter of the pinyin provided for the relatively common radicals at the beginning of Chinese characters. None of them can well solve the technical problem of being not fast for the simple ones and not simple for the fast ones, and it is not very convenient and fast to input Chinese characters.
[0006] The object of the present invention is to provide a computer Chinese character encoding input method with reasonable selection and layout of Chinese character strokes and components, standardized, intuitive, easy to learn, reasonable code-taking rules, and simple and fast Chinese character input, that is, the first-right pinyin input method. This input method provides a multi-code-taking method with encoding by the first letter of the pinyin for the common radicals at the beginning of Chinese characters on the basis of the original input method.
[0007] To achieve the purpose of the first-right pinyin input method, the present invention stipulates that the encoding of the first-right pinyin input method consists of two parts: the phonetic code and the auxiliary code. The auxiliary code can be before the phonetic code or after the phonetic code. Generally, it is stipulated that the phonetic code is first and the auxiliary code is after. Because with the support of artificial intelligence, big data technology, and search engine technology, pinyin has increasingly shown its superiority, but there is still the problem of troublesome selection of homophones in pinyin, and artificial intelligence, big data technology, and search engine technology cannot completely solve this problem, so the auxiliary code is still essential. The so-called auxiliary code can be direct, that is, directly connected after the phonetic code; or it can be indirect, that is, after inputting the pinyin like some input methods, press the tab key and then input the auxiliary code.
[0008] The phonetic code can use pinyin or phonetic notation. The phoneme letter initial-medial-final input method invented by the inventor himself can also be used. This input method is similar to the phonetic notation input method, but the final is expressed in phoneme form, and the initials are basically from the Latin alphabet, which is in line with international standards. Of course, the phonetic code supports full pinyin, double pinyin, phonetic notation letters, phoneme letter initial-medial-final pinyin, incomplete pinyin, and can also use techniques such as combined-stroke stenography to input the phonetic code part of Chinese characters.
[0009] There are two types of auxiliary codes in the present invention. One type is called the shape-based coding method, which occupies at most two codes. The other type is called the prefix-radical method, which occupies at most three codes. Through ingenious design, these two types of auxiliary codes can be freely used under the same input method, namely the first-right pinyin input method, and can be used in the most convenient way.
[0010] The shape-based coding method part occupies at most two codes. Generally, it consists of two letter codes. The present invention preferably selects five basic strokes and about 21 multi-stroke Chinese character components to participate in the coding. Since the National Language Commission also refers to the five basic strokes as Chinese character components, in the present invention, the five basic strokes are called single-stroke components, and the other 21 preferred Chinese character components are composed of multiple strokes and are called multi-stroke components. These multi-stroke components are all radicals, so they can also be called radicals or directly called radicals. These five basic strokes and 21 multi-stroke components are collectively referred to as basic components. When using the shape-based coding method, priority should be given to coding according to the basic components with more strokes, that is, priority should be given to coding according to multi-stroke components. Otherwise, the rule of selecting multi-stroke components will become meaningless. There are three rules for obtaining codes in the shape-based coding method:
[0011] The first rule for obtaining codes in the shape-based coding method is: For single-character Chinese characters, the corresponding codes of the first two basic components are obtained according to the writing order for coding, or the corresponding codes of the first and the last basic components of the Chinese character are obtained according to the writing order. When the Chinese character has only one basic component, only the corresponding code of this basic component is obtained or the corresponding code of this basic component is obtained twice in succession; For compound Chinese characters, the compound Chinese character is divided into two parts according to the overall structure. The part written first is the head part, and the part written later is the remaining part. The corresponding codes of the first basic component of the head part and the first basic component of the remaining part are obtained according to the writing order.
[0012] There is a weakness in this coding rule: That is, when using the shape-based coding method, after obtaining the first basic component of each Chinese character, it is necessary to consider the font type, that is, it is necessary to distinguish whether the Chinese character is a single-character Chinese character or a compound Chinese character, and then use two different code-taking rules for coding according to the two different font types. This will affect the mental reaction, and it is sometimes difficult to judge whether a Chinese character is a compound Chinese character, and sometimes it is also difficult to divide a compound Chinese character into two parts. Coding according to left-right structure Chinese characters and non-left-right structure Chinese characters is much easier, because it is very easy to distinguish whether a Chinese character is a left-right structure. There is a gap between the left part and the right part of left-right structure Chinese characters, and it is very easy to divide them into two parts according to the gap, that is, into the left and right parts. For left-middle-right structure Chinese characters, generally based on the first gap, the middle part is included in the right part, that is, the part outside the left part of left-middle-right structure Chinese characters is considered the right part.
[0013] The second coding rule of the shape component coding method is as follows: for Chinese characters with a left-right structure, the corresponding codes of the first basic component of the Chinese character and the first basic component of the right part of the Chinese character are respectively taken according to the writing order; for Chinese characters with a non-left-right structure, the corresponding codes of the first and the last basic components of the Chinese character are taken according to the writing order. If there is only one basic component, only the corresponding code of this basic component is taken or the code of this basic component is taken twice in succession. To prevent bypassing the patent, it is stipulated that for Chinese characters with a non-left-right structure, the corresponding codes of the first and the second basic components of the Chinese character are taken according to the writing order. However, such a stipulation is likely to increase a large number of homophones.
[0014] It should be particularly pointed out that the reason for not stipulating that all Chinese characters take the codes of the first two basic components or the codes of the first and the last two basic components is that such a stipulation will seemingly make the coding rules of the shape component coding method appear simple and easy to remember. In fact, it will cause a large number of homophones or pay the price of increasing a large number of components with many strokes. Why can "for Chinese characters with a left-right structure, the corresponding codes of the first basic component of the left part and the first basic component of the right part are respectively taken" reduce homophones? Because most Chinese characters are phonetic-semantic compounds, often with the radical on the left and the phonetic component on the right. The phonetic component is often a single character representing sound. If the first and the last basic components are taken according to the writing order as in a general input method, there will be a situation where the first stroke of the radical is the same as the first stroke of the phonetic component, which will bring a large number of homophones. To reduce homophones, it is necessary to select more radicals, resulting in a difficult-to-remember situation. Then why is it that "for Chinese characters with a non-left-right structure, the corresponding codes of the first and the last basic components of the Chinese character are taken according to the writing order."? The answer is also to reduce radicals. Because the first and the last strokes of the phonetic component are often different. For a certain same phonetic component, for Chinese characters with a left-right structure, the second code takes the first stroke of the phonetic component, and for Chinese characters with a non-left-right structure, the second code takes the last stroke of the phonetic component. In this way, the two second codes are different, generally able to better avoid homophones. In addition, if the first two basic components of Chinese characters with a non-left-right structure are coded according to the writing order, it is easy to cause more homophones, because the first two basic components of many up-down and enclosed structure Chinese characters are the same, while the last basic component of the Chinese character is different. Therefore, taking the last Chinese character component according to the stroke order for the second code can effectively reduce homophones. It can be seen that this coding rule can very effectively reduce homophones, making the present invention use significantly fewer radicals compared with the input methods invented by others, and not requiring any double-stroke or triple-stroke components. It is the result of repeated refinement, with a very low homophone rate not only among the 3,775 commonly used Chinese characters, but also among the 6,763 Chinese characters in the national standard and in the Xinhua Dictionary.
[0015] However, this encoding rule also requires continuous differentiation during encoding as to whether it is a left-right structure. Although it is obvious whether a Chinese character is a left-right structure, when actually inputting long texts, one has to constantly distinguish whether it is a left-right structure, which is still troublesome for thinking. Thus, when actually obtaining the code, the third rule of the shape part encoding method for code extraction needs to be used: The first code of the shape part encoding method is: regardless of anything, encode according to the writing order with the code of the first basic component of the Chinese character. The code extraction rule for the second code of the shape part encoding method: Starting from the right side of the first basic component of the Chinese character, scan from left to right or look from left to right. If a vertical line can divide the Chinese character into two parts without cutting off the strokes of the Chinese character, then the Chinese character is a left-right structure, and the part on the right side of the vertical line is the right part of the Chinese character. Then, encode according to the writing order with the code of the first basic component of the right part of the Chinese character. If a vertical line cannot divide the Chinese character into two parts without cutting off the strokes, scan the lower half layer or the lower part of the Chinese character from left to right, and then find the code of the last basic component of the Chinese character according to the writing order or encode with the corresponding code of the basic component at the lower right corner of the Chinese character. The reason for stipulating to scan the lower half layer of the Chinese character is that it is easy to find the last basic component of the Chinese character. The reason for stipulating to scan the lower half layer (or the lower layer) of the Chinese character from left to right is that it is the same as the scanning direction of left-right structure Chinese characters, both from left to right, and is consistent with the writing direction of Chinese characters. It is more convenient for thinking than finding the lower right corner of the Chinese character from top to bottom like the T-shaped Chinese character code before, and there will be no situation of obtaining codes from left to right at one time and from top to bottom at another time. The method of scanning the Chinese character from left to right twice, or scanning the lower half layer or the lower part of the Chinese character twice from left to right, is unheard of in various Chinese character input methods and is a major innovation.
[0016] Left-right structure Chinese characters often have obvious gaps and are very easy to distinguish, so it is not necessary to use a vertical line to divide them. For the second code, just start from the right side of the first basic component of the Chinese character, scan from left to right, find the gap between the left and right parts of the whole Chinese character, and the part on the right side of the gap is the right part of the Chinese character. Then, encode according to the writing order with the code of the first basic component of the right part of the Chinese character. If there is no gap between the left and right of the Chinese character, scan from left to right or look at the lower half layer (or the lower part or the lower layer part) of the Chinese character, and then find the code of the last basic component of the Chinese character according to the writing order.
[0017] Simply put, the first code of the shape coding method is: take the code code of the first basic component of the Chinese character in the order of writing, that is, take the head. When taking the second code of the shape coding method, first scan the Chinese character from left to right. If the Chinese character is a left-right structure, the right part of the Chinese character can be found, and the code code of the first basic component of the right part of the Chinese character is taken in the order of writing. If the right part of the Chinese character cannot be found, scan the lower half of the Chinese character from left to right, and find the code code of the last basic component of the Chinese character in the order of writing. There is no need to search directly in the lower right corner of the Chinese character as before, which is easy to confuse in thinking. Simply put, the method of the second code of the shape coding method is: scan from left to right, and take the code code of the first basic component of the right part of the Chinese character in the order of writing, which is abbreviated as taking the right. If the right part cannot be found, scan the Chinese character from left to right again, and find the code code of the last basic component of the Chinese character in the order of writing, which is abbreviated as taking the end if there is no right. In general, the second code of the shape coding method can be simply recorded as left-right scanning, if there is a right, take the right; left-right scanning, if there is no right, take the last; that is, if there is no right, take the last. The code selection rule of the entire shape coding method can be simply recorded as the first right, if there is no right, take the first and the last; or simply recorded as the first code takes the first, and if there is no right in the second code, take the last.
[0018] Note that when encountering some Chinese characters such as radicals for "门" or the lower half of Chinese characters for "心、灬", the first two strokes of "师", the first three strokes of "顺" and other components, they can be regarded as integral components, and it is not necessary to use vertical line segmentation. The last stroke of the Chinese character overwhelming majority is in the lower layer or the lower part of the Chinese character. When encountering some Chinese characters, the last Chinese character component in the writing order is "甫、犬、戈、弓箭" and other Chinese character components, according to the stroke order, the last basic component is not in the lower half of the Chinese character. At this time, the second code can be encoded by the code of the last stroke point according to the stroke order, or the last stroke point can be removed and then encoded, both of which are OK, anyway, almost no impact on repeated code, which is the cleverness of the fault-tolerant code of the present invention.
[0019] It can be seen from the code-taking rules of the shape-based coding method that it is slightly inconvenient for Chinese characters with non-left-right structures compared to those with left-right structures. Because for Chinese characters with left-right structures, they only need to be scanned once from left to right, while for Chinese characters with non-left-right structures, they need to be scanned again from left to right in the lower half of the Chinese character. Therefore, the present invention has made another innovation. That is, Chinese characters with non-left-right structures are given priority to take simple codes. Even if their common usage frequency is much lower than that of Chinese characters with left-right structures, it is still the case. That is, when the first code of the shape-based coding method of a Chinese character with a non-left-right structure is the same as the first code of the shape-based coding method of a Chinese character with a left-right structure, the Chinese character with a non-left-right structure is given priority to take the simple code. As long as the phonetic code of the Chinese character is input, and then the first code of the shape-based coding method is input, and the space bar is struck, the Chinese character with a left-right structure can be input. Of course, when the first codes of the shape-based coding methods of two or more Chinese characters with non-left-right structures are the same, one of the Chinese characters with non-left-right structures is designated to have a simple code, and generally the more common Chinese character with a non-left-right structure is taken as the simple code. This regulation has an advantage, that is, for Chinese characters with non-left-right structures, since they are simple codes, there is no need to scan again from the lower half of the Chinese character from left to right.
[0020] Incidentally, when taking the corresponding code of the last basic component of this Chinese character according to the writing order for coding or taking the corresponding code of the basic component where the lower right corner of the Chinese character is located for coding, the codes of most Chinese characters are the same. However, for a few Chinese characters, the last basic component is not in the lower right corner but in other positions. From the perspective of searching, it is still more convenient to take the basic component where the lower right corner is located. But for some Chinese characters, the lower right corner is not obvious. In this case, it is still better to take the corresponding code of the last basic component of this Chinese character according to the writing order for coding. My processing method is to give a fault-tolerant code, that is, either taking the code of the last basic component of this Chinese character according to the writing order for coding or taking the corresponding code of the basic component where the lower right corner of the Chinese character is located is acceptable.
[0021] My research also found that after splitting a compound Chinese character into two parts, the situation where the first stroke of the part other than the radical of compound Chinese characters with the same pronunciation and the same radical is of the same type of basic stroke is unexpectedly few, only about one or two hundred pairs. That is to say, the rate of homophonic characters will be very low. This discovery and the creative code-taking rules are the reasons for only selecting 5 basic strokes and about 21 basic components to participate in coding. Strictly speaking, there are differences between radicals and head characters, but because the radicals and head characters adopted in the present invention are all very common, they are simply referred to as head characters.
[0022] In the Chinese character code input method that I invented the earliest, which was earliest called the upper left input method, 28 basic components were selected. For the convenience of memorization, many basic components were encoded with the initial letters of the pinyin. However, when encountering several radicals with the same initial letters of the pinyin, there was no clear standard as to which one should be encoded by the initial consonant and which one should not. The homophonic radicals are mainly concentrated on "s, h, r, y, z, c". The radicals with the same pinyin initial letter of s include "氵, 扌, 山, 石, 纟", the radicals with the same pinyin initial letter of h include "火, 禾", the radicals with the same pinyin initial letter of r include "亻, 日", the radicals with the same pinyin initial letter of y include "月, 讠, 鱼", the radicals with the same pinyin initial letter of z include "竹, 足, 辶", and the radicals or multi-stroke components with the same pinyin initial letter of c include "艹, 虫". At that time, for the convenience of memory, the original Chinese character code input method did not arrange the multi-stroke components according to the number of strokes and the order of horizontal, vertical, left-falling, dot and folding, but arranged them according to pinyin or pictographic arrangement. When arranging the pinyin initials of the basic components, it was to avoid repeated codes. The remaining basic components with the same pinyin initials or initial consonants were arranged in a pictographic manner. However, the square stroke components of Chinese characters are different from Western letters after all, and it is difficult to achieve that they are very similar, which is a bit far-fetched misunderstanding. In order to avoid homophones, the pronunciation of the dot is encoded, and the F-like encoding is encoded by the shape. Other homophonic radicals also have similar far-fetched places. I have just realized this problem in the Chinese character code originally invented, but there is no good solution. After nearly ten years of hard exploration and sudden inspiration, I have finally invented a brand-new method for arranging homophonic radicals, that is, the homophonic near-position method on the keyboard. That is, when encountering several multi-stroke components with the same initial consonant or pinyin initial letter, select one of the multi-stroke components that is easy to remember and encode it according to the initial consonant or pinyin initial letter. You may call this multi-stroke component the captain, and the rest of the multi-stroke components are called team members. According to the keyboard layout, the team members are arranged next to the key position where the captain is located, generally to the left or right of the key position, usually adjacent keys. That is, the basic components of Chinese characters with the same initial consonant or the same pinyin initial letter are generally arranged side by side in the same row on the keyboard, arranged left and right. Of course, when encountering strokes or other multi-stroke components, it is also reasonable to be separated by strokes or other multi-stroke components, but just on the left or right of the key position. In this way, it is firmly positioned, which is obviously very easy to find and remember, and is easier to remember than the arrangement methods such as shape, strokes, and formulas. It is a major global first. However, there is no quantitative calculation of what is the captain, what is the team member, and how to arrange them on the keyboard, and there is no precedent for quantitative calculation of the input method at that time. The so-called homophones refer to the same pinyin initials or the same pinyin first letters, and the so-called near positions refer to the close positions on the keyboard, and many of the keys are directly adjacent.
[0023] In the original invention of the Chinese character code, since the left-falling stroke and the right-falling stroke are very commonly used, it is inconvenient to use the first letters of the pinyin to encode them. Therefore, five basic strokes, namely 丿, 丶, 一, |, 乚, are used to represent A, O, E, I, and U respectively. However, after in-depth consideration, it is indeed better to encode the horizontal, vertical, left-falling stroke, dot, and right-falling stroke with the first letters of the pinyin respectively, because it is simpler and more in line with thinking habits. More importantly, the original invention believes that the horizontal, vertical, left-falling stroke, dot (right-falling stroke), and right-falling stroke are encoded with H, S, P, D, and Z respectively, which is indeed not conducive to word duplication when the first code of the shape-part encoding method is input. However, when the second code of the shape-part encoding method is input, as long as it is handled skillfully, the word duplication rate will be much lower than encoding the horizontal, vertical, left-falling stroke, dot (right-falling stroke), and right-falling stroke with the vowels E, I, A, O, and U respectively. The reason is that the original Chinese character code, since horizontal, vertical, left-falling stroke, dot (right-falling stroke), and fold are encoded with the finals E, I, A, O, and U respectively, as a result, since the horizontal, vertical, left-falling stroke, dot (right-falling stroke), and fold appear very frequently in the second code of the shape part encoding method, a large number of words are repeated. In the new improved invention, on the basis of the arrangement of homophone proximity method, the radicals such as "氵, 纟, 扌, 月" that usually appear in the first code of the shape part encoding method of Chinese characters but rarely appear in the second code of the shape part encoding method are deliberately encoded with their finals A, I, O, and U respectively, "日" is similar to "曰", the last letter of the final of "曰" is "E", and the left part of "日" is similar to "E", which effectively avoids repeated words, thereby greatly improving and perfecting the present invention.
[0024] Quantitative analysis is a significant improvement of the first-right pinyin input method compared to the original phoneme homophone near-position Chinese character code input method. Through quantitative calculation, 21 Chinese character multi-stroke components are selected and accurately positioned on the keyboard. The following is a specific explanation: Among the 6763 Chinese characters in GB, the frequency of grouping characters in the radical prefix of 氵, 艹, 口, 木, 扌, 亻, 土, 钅 is very high, and can form more than 300 Chinese characters. If they are coded by stroke, a large number of repeated codes will be caused, so they should be selected and arranged on the key, and coded with one letter respectively. Multi-stroke components or radicals "虫, 女, 月" can also form about 250 Chinese characters. The first few strokes of "虫" are "口". In order to avoid coding "虫" and "足" as "口", resulting in a large number of repeated codes, "虫" and "足" should be selected and coded with a certain letter. If "女" and "月" are coded by stroke, they will also bring about 40 to 50 pairs of repeated codes, which should also be selected and coded with another letter respectively. The radicals 忄, 火, 讠, 纟, 石 and other character-forming abilities are slightly less, about 200 pairs. If they are encoded by strokes, the multi-stroke component 忄 can bring nearly 40 pairs of repeated codes; 火 can bring nearly 40 pairs of repeated codes; 讠 can bring about 36 pairs of repeated codes; 纟 can bring 41 pairs of repeated codes; "石" and "王" can bring 35 pairs of repeated codes; the radicals "日" and "足" can bring 40 and 36 pairs of repeated codes respectively; the radical "辶" can bring nearly 30 pairs respectively, and the repeated code 疒 can bring more than 20 pairs of repeated codes. According to the ability of avoiding repeated codes, the radicals 辶、忄、纟、日、火、讠、足、石、王、疒 are also coded with another letter respectively, wherein the ability of avoiding repeated codes of 宀 and 疒 is close, when 宀 contains acupoints, the ability of avoiding repeated codes is slightly higher than 疒, and the frequency of use of the Chinese character with the radical 宀 is higher than the frequency of use of the Chinese character with the radical 疒, at this time 宀 can be coded with a letter, 疒 is coded by strokes, and the frequency of use of the radical king is higher than the frequency of use of the radical stone. 21 radicals have all been selected in this way, and are coded with a letter respectively. The radical "fish" can bring 24 pairs of repeated codes, and the ability of avoiding repeated codes is almost the same as "宀" and "疒", because the Chinese character composed of the radical "fish" often appears in the phrase mode of "some fish", such as "carp", "silver carp", etc., so the actual frequency of use of individual characters is very low, so if only 21 radicals are selected, they are slightly abandoned. The repeated code quantity of other radicals such as , 宀 (cave), mountain is also a lot, and when the quantity of Chinese characters expands to Xinhua Dictionary, can produce 30 pairs of repeated code, 宀 together with " cave " can produce 30 pairs of repeated code, and mountain can produce nearly 30 pairs of repeated code. If at this moment, by the repeated code quantity, determine the radical selected, because the repeated code quantity of radical , 宀 (cave), mountain and foot, stone, 疒 are closer, making choices has just become a difficult problem, and the quantity of repeated code produced by radicals such as 阝, 禾, bird is less, 阝 can produce 15 pairs of repeated code, "禾" can produce 16 pairs of repeated code, "鸟" can produce 24 pairs of repeated code, generally can be discarded.
[0025] In order to solve the problem that the radicals 口, 宀 (hole), 山, 足, 石, 疒, which have about 30 pairs of repeated codes, are close and difficult to choose, I spent another year to study it intensively. Finally, I decided to use the Peking University character frequency table, the Beijing Language and Culture University word frequency collection and the Douding 6763 character frequency table (1 billion bases) as the basis, and count the radicals 口, , 宀 (hole), 山, 足, 石, 辶, 疒, 纟, 忄, 女, 王, 日, 人 (亻), 土, 讠, 月, 氵, 扌, 火, 金, 艹, 虫, 木, etc. that appear in the Chinese character encoding. When repeated codes occur, the frequency frequency of the Chinese character and the frequency frequencies of one or more Chinese characters with the repeated codes of the Chinese character are recorded, and statistical addition is performed according to the radical classification to obtain the frequency sum. Multi-stroke components with high frequency sum are selected as much as possible, which can reduce the selection of homophones. To improve the speed, take the word frequency summary of Beijing Language and Culture University, which is a statistical table of words, as an example. It is found that Chinese characters with multiple stroke components such as 口, 人 (亻), 氵, 扌, 艹, 土, 日, 讠, 月, 火, 金, 虫, 木, 纟, etc., if encoded by stroke, the frequency sum of the Chinese characters with repeated codes and the frequency sum of other Chinese characters with repeated codes are very high, so they should be selected, and 宀 (穴) The frequencies of the Chinese characters 宀 (穴 part is encoded by 宀), 辶, 竹, 火, 石, 山, etc., which generate repeated codes, are 888, 767, 191, 177, 128, 59, respectively, while the frequencies of the other Chinese characters that generate repeated codes with them are 1209, 1916, 523, 563, 363, 229, respectively. In this way, the frequency of 宀 (the hole part is encoded as 宀) and 辶 is relatively high, so it is recommended to be selected. Bamboo and fire are next. The first letter of the pinyin of the radical fire and the radical 禾 is the same, H. The frequency of the radical fire is twice that of the radical 禾 (except when the Chinese character "和" is used), so it is also recommended to be selected. The frequency of the radical 山 is relatively low, so it is not recommended to be selected. As for 石, it depends on the situation of 足. 足 is encoded as the first stroke "vertical", and the frequency of its own repeated code is about 236 thousand. The repeated code with the Chinese character containing "足" is also recommended. If the Chinese characters of "足" are included, the frequency is 316 thousand. If "足" is encoded as "口", the frequency of the Chinese characters participating in the shape coding method is 502 thousand. In general, the frequency of the radical "足" is higher, so it is recommended to select the component "足", and "石" is not selected. Of course, considering the huge difference in the frequency of the Chinese characters of "足" when they are repeated, such as the frequency of "路" and "噜" and "卢" is very different, so you can also select the component "石" instead of the component "足". Of course, you can also select "石" and "足" at the same time, but there will be two radicals on this key, and one letter does not correspond to one radical, which is not convenient to display on the mobile phone screen.
[0026] For the convenience of memory, the five basic strokes and multi-stroke components such as 王、土、钅、口、忄、疒、女、木 are all encoded by the first letter of the pinyin, and the remaining multi-stroke components are arranged and encoded according to the homophonetic method. In particular, for the convenience of memory, the concept of the first letter of the vowel is introduced on the basis of not causing a greater impact on the word re-code, that is, "月、纟、扌" are encoded by their first letters of the vowel U, I, O, and although the pronunciation of "氵" is "氵", it is often read as three drops of water, so the first letter of the vowel A of the pinyin of "三" is taken. This is easy to remember, but it is also based on quantitative calculation. The following is a specific explanation:
[0027] When arranging radicals with the same initial consonant or the same first letter of pinyin by the method of homophones and proximities, quantitative analysis and calculation of arrangement and mapping were also performed. The Chinese characters selected are all those that appear in the Xinhua Dictionary app. Some radicals have strong word-forming ability and are used frequently, but they are not evenly distributed in the syllables of the initial consonants or finals of each of the 26 letters. In the pinyin syllables where some initial consonants and finals are located, the number of Chinese characters where these radicals are located is very small. If these radicals are encoded with a specific letter, it can effectively avoid word duplication. This principle is the theoretical basis for quantitative calculation.
[0028] Said near position is in order: i.e. Q to P on the keyboard, to A to L, to Z to M, and then back to Q. The upper row of the keyboard is from left to right, from Q to P, and then to the middle row of the keyboard, from left to right, from A to L. Then to the lower row of the keyboard, from left to right, from Z to M, and then back to the Q key. Because the unique multi-stroke components of the first letters of the pinyin such as horizontal, vertical, left-falling, dot, folding and the king, earth, 钅, mouth, 忄, 疒, woman, wood are all specified to be encoded with the first letter of the pinyin, so there are only thirteen stroke radicals left arranged by the same sound near position method. These components have a high frequency of word formation, but the first letters of the pinyin are the same, resulting in a large number of repeated codes, which is that the auxiliary code of an input method such as a certain input method is difficult to improve the speed, causing the reason for failure. Among them, 亻 and 日, 讠 and 月, 横 and 火, 虫 and 艹, 折 and 辶 and 足, 竖 and 氵 and 扌 and 石 and 纟 have the same first pinyin letters, which are r, y, h, c, z, s respectively. 亻 and 日, 讠 and 月, 横 and 火, 虫 and 艹 all have the same first pinyin letters of two radicals. For the convenience of memorization, they are arranged together on the keyboard in adjacent positions on the left and right for the convenience of memorization. The following will explain the multi-stroke components with the same first pinyin letters according to the phonetic sequence of the first pinyin letters c, h, r, s, y, z.
[0029] The first letter of pinyin for 艹 and 虫 is c. According to the homophonetic method, they can only be arranged on two adjacent keys, c and v. Since v is a vowel and is rare, we only need to consider the number of Chinese characters that appear when they are the first letter of the pinyin c. 虫 appears 3 times, and the frequency is relatively low. 艹 appears 11 times, and the frequency is much higher. Since v is the first letter of the Chinese character, it is rare. The final vowel is a Chinese character with a low frequency and a low frequency. In order to avoid repeated coding of words, it is recommended to use a multi-stroke component with a relatively low frequency. Among commonly used Chinese characters, the number of Chinese characters that appear in the second code of the shape coding method is also lower for 虫, and much higher for 艹. Therefore, 虫 is coded with v, 艹 is coded with c, and can be simply recorded as 草虫. The pinyin of 艹 is cao, and the pinyin of 虫 is chong. According to the phonetic order, 卄 should be coded with c, and 虫 should be coded with v. The frequency of use of the radical 艹 is more than twice that of the radical "虫". From the perspective of brain reflection, it is more reasonable for the radical 艹 to be coded with c.
[0030] The first letter of the pinyin for the horizontal stroke and radical “火” is H. Since J has already arranged the radical 钅, “火” can only be encoded using the G adjacent to the left of the H key.
[0031] The pinyin initials of 日 and 亻 are both r, and they can only be arranged on the two adjacent keys of e and r according to the method of homophonic proximity. From the perspective of avoiding word duplication, the number and frequency sum of the basic components 日 and 亻 in the Chinese characters with the finals ue and ie are counted. Since the number of Chinese characters that the basic components "日" and 亻 appear in are only a few, the frequency sum is also very low and very close, so only the frequency of use of the radical "日" and 人 (including 亻) can be compared. According to statistics, the frequency of use of the radical 人 (including 亻) is about 2.5 times that of the radical "日". Therefore, the radical 人 (including 亻) is encoded with the pinyin initial R, and the radical "日" is arranged on the key E next to the pinyin initial R and encoded with E. The left part of the radical "日" is similar to E, and it is easier to remember to encode it with E.
[0032] The first letter of the pinyin for 竖, 纟, 扌, and 氵 is s. Many input methods use s to encode, resulting in a large number of duplicate codes, and thus encoding failure. From the keyboard layout, the basic stroke vertical is very common, so of course it is encoded with the s key. The I key, O key, and A key can be considered adjacent to the S key. 纟, 扌, and 氵 can be arranged on the I, O, and A keys. For this reason, I used operations research to perform quantitative calculations. Among the Chinese characters whose first pinyin letter is a, there is 1 Chinese character containing the component 氵, with a frequency of 5920, and there are 2 Chinese characters containing the component 扌, with a total frequency of 64779, so it is better to encode 氵 with a, while among the Chinese characters whose first pinyin letter is o, there is 1 Chinese character containing 氵, and there is no Chinese character containing 扌, so it is more appropriate to encode 扌 with o. Among the Chinese characters that begin with the first letters i, o, and a in pinyin, there is no 纟, and o and a have been used for encoding and 氵 respectively. Considering comprehensively, 纟 is encoded with the remaining i. From the frequency of the vowels i, o, and a, i is the highest, a is the second, and o is the lowest. From the frequency of radicals, 纟 is the lowest, 氵 is the second, and 扌 is the highest. From the perspective of encoding word re-coding, high-frequency vowels are suitable for matching with multi-stroke components or radicals with low frequency of use, and low-frequency vowels are suitable for matching with multi-stroke components or radicals with high frequency of use. Therefore, 纟 is suitable for encoding with i, 扌 is suitable for encoding with o, and 氵 is suitable for encoding with a. 纟, 扌, and 氵 are just encoded with the first letters i, o, and a of the vowels, respectively, which is easy to remember.
[0033] Another way to remember it is: 纟 is pronounced as si, which is two letters, 扌 is pronounced as shou, which is four letters, and 氵 is pronounced as shui, so you can start from the top row of the keyboard from left to right, and then to the middle row of the keyboard, and arrange them in order of the number of pinyin letters. If the number of pinyin letters is the same, arrange them in order of the phonetics. Put 纟, 扌, and 氵 on the i, o, and a keys respectively, and encode them with corresponding letters.
[0034] The pinyin initials of 月 and 讠 are both y, and they can only be arranged on two adjacent keys, y and u, according to the homophone and proximal method. From the perspective of avoiding word duplication, it is necessary to consider the frequency or frequency sum of the basic components "月" and 讠 appearing in Chinese characters with the finals iu or ou. The number of Chinese characters in which the basic component "月" appears at the beginning of a word is 2, while the number of Chinese characters in which 讠 appears at the beginning of a word is 8. The frequency sum (sum of usage frequency) of these Chinese characters is also relatively low for the Chinese characters in the 月 part, so it is more appropriate to encode the basic component 月 with u and 讠 with y. At this time, the input shape part encoding method will hardly cause duplication. Comparing the number of Chinese characters with the initial consonant y, the number of Chinese characters with the initial consonant "月" and "讠" is 10, and the number of Chinese characters with the initial consonant "月" is 15. The sum of the frequency, that is, the frequency of use, is higher for the radical 讠, so the basic component 讠 is encoded with y, and the basic component "月" is encoded with u, and u happens to be the first letter of the final vowel of "月", which is very easy to remember. In addition, from the perspective of phonetic sequence, the pinyin of 讠 is yan, and the pinyin of 月 is yue. If they are arranged from left to right in phonetic sequence, the basic component 讠 should be encoded with y, and the basic component "月" should be encoded with u. From the perspective of the frequency of use of radicals, the frequency of use of radical 讠 is 1.3 times that of radical "月". Therefore, the radical 讠 is encoded with Y, and the radical "月" is encoded with the key U to the right of the Y key.
[0035] The first pinyin letter of 折, 辶, 足 and 竹 are all z. The stroke 折 is very common, so of course it is represented by z. According to the homophonetic method, 辶, 足 and 竹 can only be arranged on the remaining l, Q and F keys. Among them, l and z are located on the rightmost of the second row and the leftmost of the third row of the keyboard respectively, which can be considered as proximal. The letters to the right of z in the lower row of the keyboard have already been arranged as radicals. Therefore, according to the homophonetic rule, the Q and F keys on the upper row of the keyboard can barely be considered as proximal. Since the frequency of the initial consonant L in Chinese is much more common than the initial consonants F and Q, the first pinyin letter L is arranged first. 辶 can only appear at the end of a Chinese character, and the sum of frequencies is 0. When the first pinyin letter is L, there are 7 Chinese characters with 足 at the beginning of the character, and the sum of frequencies is 109352, and there are 12 Chinese characters with 竹 at the beginning of the character, and the sum of frequencies reaches 16734. It can be seen that neither 足 nor 竹 is suitable for encoding with L. From the perspective of avoiding duplicate word codes, 辶 can be encoded with L, which can make the duplicate word codes 0. This is a very clever arrangement. There are 3 Chinese characters with the prefix "足" and the first letter of the pinyin "F", and the sum of the frequencies is 293; there are also 3 Chinese characters with the prefix "竹" and the first letter of the pinyin "F", and the sum of the frequencies is 16022. There are 5 Chinese characters with the prefix "足" and the first letter of the pinyin "Q", and the sum of the frequencies is 7626; there are 5 Chinese characters with the prefix "竹" and the first letter of the pinyin "Q", and the sum of the frequencies is 9664. From the number of characters, "足" and "竹" are both less and close in the pinyin first letter F and Q keys. From the sum of the frequencies, in the Chinese characters with the first letters of the pinyin "Q" and "F", there are more "竹" radicals, so "足" and "竹" can be selected from F and Q respectively. From the perspective of keystroke convenience, it is more appropriate to encode "竹" with F and "足" with Q. Of course, "竹" with Q and "足" with F is also OK. From the perspective of memory, the first letter of the pinyin is z, so we can only arrange them according to the strokes. The first stroke of 足 is vertical, the first stroke of 竹 is left-falling stroke, and the first stroke of 辶 is dot. The order is vertical, left-falling stroke, and dot. From left to right on the keyboard, the order is Q, F, and L. Therefore, 足, 竹, and 辶 are placed on the Q, F, and L keys respectively according to their first strokes, vertical, left-falling stroke, and dot, and are encoded with Q, F, and L. From the perspective of shape, the beginning and end of 足 are shaped like Q, the beginning of 足 is also shaped like a lowercase q, the left or right half of 竹 is shaped like F, and 辶 is also shaped like L, which is easy to remember. Of course, you can also sort them from the perspective of the number of letters and the phonetic sequence. The pinyin of "足" is zu, which only consists of two letters, so it is placed on the q on the far left of the keyboard. The pinyin of "竹" is zhu, and the pinyin of "辶" is 辶, both of which are three letters. According to the phonetic sequence, arrange "竹" and "辶" from left to right on the f and l keys respectively, and encode them with corresponding letters respectively.
[0036] It is also possible to uniformly sort the radicals from left to right according to the number of pinyins in the radicals based on homophones and arrange them in phonetic order if the number of pinyins is the same. In this way, "人" (人) is encoded with "r" and "日" (日) is encoded with "E". The rest remain unchanged, but it is not conducive to reducing the duplication of words, and "亻" is no longer encoded with the first letter of the pinyin final.
[0037] If the number of radicals is not considered on the basis of homophones and proximal positions, the characters can be arranged in phonetic order only. In this case, 亻含亻 is coded with r, 日 is coded with e, 扌 is coded with i, 氵 is coded with o, a is coded with 纟, 火 is coded with g, 艹 is coded with c, v is coded with 虫, 竹 is coded with q, 辶 is coded with f, and 足 is coded with 1.
[0038] It can be seen that in this input method, 5 strokes and 8 multi-stroke components, namely, 王、土、钅、口、忄、疒、女、木, are all coded according to the first letter of the pinyin, and 亻 (including 人), 讠, 虫 are also coded according to the first letter of the pinyin. In this way, a total of 16 stroke radicals are coded according to the first letter of the pinyin. Actually, only 10 multi-stroke radicals need to be remembered. In the Shouyou Pinyin input method, it is also possible to encode these 10 radicals according to the first letter of the pinyin, but a large number of duplicate codes will be caused. Therefore, the remaining 10 radicals are arranged according to the method of homophony (first letter of the pinyin) and proximity (adjacent position on the keyboard), and basically according to the number of letters of the pinyin of the radical. If the number of letters is the same, they are also arranged in phonetic order. Some of them also take into account strokes and similar shapes, which are very easy to remember. Among them, the five multi-stroke components, 月, 纟, 扌, 氵, are both coded by the initial letter of the vowel and arranged by homophones. In fact, only 6 are arranged by homophones. These 6 are also arranged by the number of letters and the English phonetic order based on homophones, which is very easy to remember. In order to further shorten the memory time, I have also compiled a formula, that is, the basic stroke is the captain, and it is preferentially coded with the initial letter of the pinyin. In the syllables with the same initial letter of the two pinyins, the radicals 人, 讠, 艹, and 横 are the captain, and the radicals 日, 月, 虫, 火 are teammates, and are also arranged on the left and right adjacent keys of the captain. The radicals 艹 and 虫 are homophonic to grass insects, and are arranged on the c and v keys next to each other, which is more conducive to memory. 折, 辶, 足, 竹 form a team, the captain is 折, and the radicals 辶, 足, 竹 are teammates. 竖, 纟, 扌, 氵 form a team, 竖 is the captain, 纟, 扌, 氵. The coding rule of the shape coding method is simply written as "first, no right, then last", which is very easy to remember. Most people can remember it in three to five minutes.
[0039] By optimizing about 21 multi-stroke components and five basic strokes, creatively specifying the code-taking rule of the second code of the shape coding method, creatively adopting the homophone proximity method to arrange the multi-stroke components and basic strokes and creatively performing quantitative calculation, accurate positioning, the shape coding method is simple and easy to remember, and can effectively distinguish homophones, and the repetition rate is very low in 3500 commonly used Chinese characters and 6763 commonly used Chinese characters in the national standard, and the input speed can be compared with input methods such as Wubi font. This solves the problem that any other input method has not been able to solve, and truly achieves simple and intuitive, low repetition rate, high input speed, adopts technologies such as artificial intelligence and search engines, almost no repetition code, and can be compatible with the most popular pinyin input method or phonetic input method, and is a unique ideal and perfect Chinese character input method that can be popularized to primary and secondary school students.
[0040] Some radicals are very common and frequently used, but due to their low frequency of homophones, they can only reduce a little over 10 pairs of homophones, and with only 26 key positions, they were not selected. However, if these components were selected, it would be beneficial to those who pursue typing speed. In the new invention, several selected components are double-coded, that is, they can be coded by strokes or by radical components, and they are not convenient to be displayed on the small-screen keyboards of mobile phones, etc. These components are called double components or virtual components, and can also be called double radicals or virtual radicals. The reason for calling them virtual components is that they do not appear on the letter keys of small screens such as mobile phones, but can be coded with punctuation keys. That is, it is stipulated that double components can be coded by strokes or by punctuation keys. The character-forming ability of "fish" is strong, and it can avoid 24 pairs of homophones. It is arranged on the ";" key and coded with ";". Then, classified by the usage frequency of radical components, "mountain", "left ear radical", and "rice" are respectively arranged on the ",", ".", and " / " keys, and coded with ",", ".", and " / " respectively. See the appendix Figure 2 The "." and "." are the same key and the same code. Just for clarity, "." is used.
[0041] As a variation of the present invention, "foot" or "bamboo" can also be replaced by "stone", or the component "stone" can be arranged on the L key, because the "walk radical" on the L key only appears at the end of a character. But then two characters on the L key are missing. Some radicals are very common, but due to their low frequency of homophones, they can only reduce a little over 10 pairs of homophones, and with only 26 key positions, they were not selected. However, if these components were selected, it would be beneficial to those who pursue typing speed. In the new invention, several components are double-coded, that is, they can be coded by strokes or by radical components, and they are not convenient to be displayed on the small-screen keyboards of mobile phones, etc. These components become double components or virtual components. The reason for becoming virtual components is that they do not appear on the letter keys of small screens such as mobile phones, but can be coded with punctuation keys. For example, 5 multi-stroke components can be added to the appendix Figure 1 These added multi-stroke components do not appear on the mobile phone screen and can be remembered by experts. A feasible arrangement scheme is that "stone" is coded with "1", "fish" is coded with ";", and "mountain", "left ear radical", and "rice" are coded with ",", ".", and " / " respectively. The appendix Figure 3 Appendix Figure 4 lists the arrangement and mapping methods of the strokes of some other components on the keyboard. The characteristics of these attached drawings are that some relatively infrequently used multi-stroke components with similar frequencies can be replaced. It should be noted that the appendix Figure 3 Appendix Figure 4 is only an enumeration and is a variation of the present invention.
[0042] As an auxiliary code, the shape part coding method fully considers the compatibility of words and characters at the time of invention, and adopts artificial intelligence and search engine technologies. As a direct auxiliary code, it is not necessary to press keys such as tab, which can reduce the number of keystrokes. Of course, it can also be used as an indirect auxiliary code. It is recommended to press the tab key for the indirect auxiliary code because there are many character and word homophonic codes, and the homophonic rate of single characters is also high.
[0043] After the successful invention of the first-right pinyin input method, because it can be learned in a few minutes and has a low homophonic rate, it was promoted to several major famous input method companies. The manager of a certain input method blindly admired the big data search engine technology and thought that many users were not willing to learn even for a minute. In fact, pinyin also needs to be learned. Use the v code. For people like the manager of a certain input method, it should be possible to learn the auxiliary code in one minute. However, it is very difficult to invent a new first-right pinyin input method that can be learned in one minute and is compatible with the original first-right pinyin input method. After nearly two years of painstaking research, I found that there are not many cases of hundreds of homophonic characters. Most of the numbers of homophonic characters are below thirty or forty. And I also found that the very common radicals at the beginning of homophonic characters often only appear once. Therefore, as long as the coding of this very common radical is made non-homophonic, the homophonic code can be effectively avoided. On this basis, I invented a new auxiliary code, the first radical method. Based on the original first-right pinyin input method, the first radical method occupies at most three codes and perfectly solves this problem.
[0044] Specifically, in order to be compatible and reduce homophonic codes, the five basic strokes, horizontal, vertical (including vertical hook), left-falling stroke, dot (including right-falling stroke), and fold, and 11 radicals such as "king, person (single-person radical), earth, speech radical, gold radical, mouth, heart radical, grass radical, roof radical, woman, wood" are still encoded with the first letters of their pinyin, H, S, P, D, Z, W, R, T, Y, J, K, X, C, B, N, M respectively. It is still possible to encode according to the original rules of the shape part coding method and input the second code of the shape part coding method to conveniently input Chinese characters. However, it is also possible not to input the second code of the shape part coding method. At this time, according to the prompt line, you can select using the numeric keys or the space bar.
[0045] In the original first-right pinyin input method, the radicals foot, sun, moon, silk, hand, water, bamboo, fire, walk, and insect were arranged according to the method of arranging characters with the same pronunciation and adjacent positions. Although it is very simple, it still requires memorization, which is more or less inconvenient for beginners. In the latest invention, a new coding method, namely the radical-first coding method, is adopted. The method is to code by adding the first letter of the pinyin of the radical, the first letter of the phonetic compound, and the first basic stroke of the radical. The coding length is two to three codes. Usually, only these two codes, the first letter of the pinyin and the first letter of the phonetic compound, are needed. Only when there is a situation where some radicals have the same code, the code of the first basic stroke of the radical is required as the third code. If the phonetic compound of the radical is a compound vowel, the second letter of the phonetic compound can also be used as the third code. For those with retroflex sounds, the second code can also be h. Specifically, the first letters of the pinyin and the first letters of the phonetic compounds of the radicals foot and bamboo are the same, zu. Therefore, the code of the first basic stroke of the radical needs to be added for coding. In this way, the code of the radical "foot" is zus, and the code of the radical "bamboo" is zup. Since the radical "bamboo" is more commonly used than the radical "foot", the code of the radical "bamboo" can be simplified to zu. After optimization, the code of the radical "bamboo" is zh, and the code of the radical "foot" is zu. Similarly, when coding the radicals sun, moon, silk, hand, water, fire, and insect, the first letters of the pinyin and the first letters of the phonetic compounds are ri, yu, si, so, su, ho, and co respectively. Since the radical walk usually appears at the end of a character, in the radical-first coding method, it is coded as zod and can be not coded.
[0046] Then, do the radicals other than these 21 preferred radicals need to be coded? In fact, as long as the shape radical coding method of the first-right pinyin input method is followed, they can be not coded. However, for compatibility and to meet the needs of some people, and because the coding of the radicals of some characters using the radical-first coding method is simpler and can effectively avoid homophones, coding is also provided. When coding, it is also coded by adding the first letter of the pinyin of the radical, the first letter of the phonetic compound, and the code of the first basic stroke of the radical according to the stroke order. For the radical with the syllable er, the code of the first basic stroke according to the stroke order can be directly added. The coding length is also two to three codes. The advantage of the code-taking rule of the radical-first coding method is that as long as the first letter of the pinyin of the radical located at the beginning of a character is input, the character can be found through the prompt line, thus greatly improving the universality of the present invention. There are many radicals in Chinese characters, and some input methods have hundreds of radicals. I think it is not necessary to code all hundreds of radicals. As long as the more common and easily pronounced radicals are coded, and the rest follow the shape radical coding method of the first-right pinyin input method. The following lists the codes of the more common radicals that usually appear at the beginning of Chinese characters:
[0047]
[0048]
[0049] The encoding of the radical in the above radical method is also the mapping relationship on the keyboard. It can be seen that most radicals at the beginning of a word can be obtained by taking the first letter of the pinyin plus the first letter of the final. When the first letter of the pinyin and the first letter of the final are the same, the radical with higher frequency is encoded by the first letter of the pinyin and the first letter of the final, and the rest of the radicals must also add the code of the first basic stroke of the radical in the stroke order. For example, the first letter of the pinyin of the radical "车" is c, and the first letter of the final is e, so the code of the car is ce; a few radicals at the beginning of a word also need to add the first basic stroke code of the radical, for example, the first letter of the pinyin of the radical "石" is s, the first letter of the final is i, and the first basic stroke is horizontal, which is encoded as h, so the code of the radical "石" is sih. But in this way, the code has three letters, which increases the number of keystrokes, and is often not as convenient as the original shape-part coding method, because the original shape-part coding method has a maximum of two codes, but sometimes because the pinyin comes first and the auxiliary code radical method comes later, it is only necessary to enter the first few codes of the radical method, such as entering the first one or the first two codes, so it may be convenient. It should also be pointed out that the above radicals are just examples, and other radicals that are not listed can also be encoded in the same way, but there is no need to have the original shape-part coding method rules. After entering the radical method of the Chinese character, is it necessary to enter the right part of the Chinese character or the last part of the Chinese character? The answer is that it is not necessary, because the radical method itself occupies two to three codes, and the coding length is too long. It should be pointed out that another method of obtaining the code of the radical method is to take the first letter of the pinyin of the radical plus the code of the first basic stroke of the radical according to the stroke order. If it is still not clear, then add the first letter of the vowel of the radical. The advantage of this coding rule is that it can avoid duplicate codes for words, but the disadvantage is that it is slightly inconvenient in thinking.
[0050] From the above radicals and their codes, even if some radicals are coded with three letters, it is sometimes impossible to completely avoid duplicate codes. For example, the codes of the radicals "骨、鬼、瓜" are all gup, which affects the efficiency to a certain extent. At this time, it is advisable to adopt the coding rule of fewer duplicate codes and more two codes when using three codes. That is, take the first letter of the pinyin of the radical plus the code code of the first basic stroke of the radical in the stroke order. If there are duplicate codes, then add the code code of the second basic stroke of the radical in the stroke order. The following is another more common radical code that often appears in the beginning of Chinese characters:
[0051]
[0052] In addition, the code encoding of the first, second and third basic strokes of the radical according to the stroke order is also possible. However, in this case, it is not possible to input according to the first letter of the pinyin.
[0053] When using pinyin input for the phonetic code part, there will be Chinese characters that are not recognized. For this reason, the present invention provides a fast input method based on the keyboard layout diagram of the radical coding method, that is, the keyboard layout diagram of the radical coding method can be selected from Appendix Figure 1 Appendix Figure 2 Appendix Figure 3 Appendix Figure 4 Any one of them can be selected, and once selected, it cannot be changed. Generally, Appendix Figure 1 is selected. Taking the selection of Appendix Figure 1 as an example, when inputting, for this Chinese character, according to the stroke order, combined with the basic strokes and multi-stroke components in Appendix Figure 1 , following the principle of giving priority to the radical with more strokes, the codes corresponding to the basic strokes and multi-stroke components of this Chinese character are input in sequence, and the required Chinese character is selected according to the prompt line.
[0054] Since the phonetic code part uses full pinyin, the code length is relatively long, and there is room for improvement in the final nasal sound ng of the rhyme. Therefore, the present invention creatively uses v to represent ng, because it will not affect the rhyme is represented by v. To shorten the code length and improve the input speed. At the same time, those who do not want to use v to represent ng can still use full pinyin to achieve perfect compatibility.
[0055] When inputting phrases, it is basically input according to the phonetic code. In order to solve the problem of serious homophones when inputting the phonetic code, on the basis of the first-right pinyin input method, the present invention draws on the auxiliary code. As long as after the code of a certain homophone, the auxiliary code of the first and second characters of this homophone or the first code of the radical coding method is input, the trouble of selecting homophones can be basically eliminated. Generally, as long as the first code of the auxiliary code of the first character of this homophone is input, the trouble of selecting homophones can be roughly eliminated. Among several homophones, the most commonly used homophone can be input only according to the pinyin code, and the first code of the auxiliary code of the first character is taken for the remaining several homophones, so as to more effectively distinguish homophones. If there are still duplicate codes, the first code of the auxiliary code of the second character of the homophone can be taken. Of course, sometimes, among several homophones, adding the auxiliary code sometimes changes to different pinyins. For example, the pinyin of several homophones is zhidu. If the first code of the auxiliary code of a certain homophone is i or o, and then i or o is input, it will become zhiduo or zhidui. In order to avoid syllable conflicts, a less commonly used homophone can also be selected. The selected homophone can not add the first code of the auxiliary code of its first character. The so-called homophones are generally full pinyin. After full pinyin, input the first code of the auxiliary code of the first and second characters of this homophone, which can better eliminate the selection of homophones. Of course, under the condition of simple pinyin, the first code of the auxiliary code of the first, second and even third characters of this homophone can also be input, which can also effectively eliminate the trouble of selecting homophones, but the effect is relatively weaker. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] AppendixFigure 1 One of the keyboard layout diagrams for the shape-based encoding method
[0057] Appendix Figure 2 Two of the keyboard layout diagrams for the shape-based encoding method
[0058] Appendix Figure 3 Three of the keyboard layout diagrams for the shape-based encoding method
[0059] Appendix Figure 4 Four of the keyboard layout diagrams for the shape-based encoding method
[0060] Appendix Figure 5 One of the mapping relationship diagrams of phoneme letters and finals on the keyboard
[0061] Appendix Figure 6 Two of the mapping relationship diagrams of phoneme letters and finals on the keyboard Specific implementation manners
[0062] The first right pinyin input method consists of two parts, one part is the phonetic code, i.e., the pronunciation, or the pinyin code, and the other part is the shape coding method, which is the auxiliary code in the general input method. When these two parts are composed and encoded, the phonetic code can be first and the shape coding method can be second; it can also be the shape coding method first and the phonetic code can be second. But once selected, it cannot be changed. In order to facilitate the typing and be consistent with the thinking, in order to be fully compatible with the pinyin input method, it is recommended that the pinyin be first and the shape coding method be second. This method is used in the encoding example. Pinyin can be full spelling or double spelling or simplified spelling or incomplete pinyin. Full spelling uses the standard pinyin of a Chinese character. The phonetic input method can also be used. Note that the part representing the tone in the phonetic input method should be removed, because the shape coding method and the radical method of the present invention are much higher than the tone to distinguish the heavy code ability. Double spelling has as many as 35 vowels, which are inconvenient to arrange and remember, and it has never been popularized. Therefore, in the new invention, non-professional typists do not approve of double spelling. Either Chinese Pinyin or Zhuyin is used. Zhuyin has a shorter code length, which is generally only two or three codes if the tones are not counted, while Chinese Pinyin has a maximum code length of 6 codes, so the input speed is theoretically faster than Chinese Pinyin, but Zhuyin initials are not represented in Latin, and finals are not represented in phoneme. The phoneme alphabet invented by me has the same encoding length and syllable expression method as Zhuyin, but the initials are represented in Latin, the finals are represented in phoneme, and it is easy to write and easy to display on small screens such as mobile phones. The code length is shorter than Pinyin, and the input speed is faster than Pinyin. The disadvantages of phoneme alphabet and Zhuyin alphabet are that if one phoneme alphabet is pressed, punctuation keys or number keys are used, and the keystrokes of several punctuation keys or number keys are slightly inconvenient. The initial consonants of individual letters of the phoneme letters are the same as those of the pinyin. The retroflex consonants in the pinyin can be arranged on the v, u, and i keys. Since the present invention can effectively avoid duplicate codes, if the retroflex consonants are not distinguished, the duplicate codes are still very low. In the present invention, the retroflex consonants can be treated as not being treated, that is, the retroflex consonants are not distinguished, that is, zh is encoded with z, ch is encoded with c, and sh is encoded with s. The finals of the phoneme letters are also very simple. The phonemeized finals can be easily converted to the finals in the Chinese Pinyin scheme. Just remember 一、丨、丿、丶、 フ and r represent the letters e, i, a, o, u, n, ng, and r that make up the finals in the Chinese Pinyin Scheme, and then write them in the order they were written. Or it can be represented by ∠.
[0063] A mapping relationship diagram of the letter, punctuation, and number keys on the English keyboard and the pinyin vowels and phoneme letter vowels is shown in the attached figure. Figure 5 As shown:
[0064]
[0065] The “.” in the accompanying drawing is the key where the “.” is located, that is, the key where the ">” is located.
[0066] The following is a detailed description of the shape-based encoding method.
[0067] Classify the various strokes of Chinese characters into five basic strokes: horizontal, vertical, left-falling, dot, and fold according to the regulations of the State Language Commission. A stroke is a continuous line written in one go when writing Chinese characters. The strokes can be classified into five basic strokes: horizontal, vertical, left-falling, dot, and fold. Among them, the rising stroke is incorporated into the horizontal stroke, the vertical hook is incorporated into the vertical stroke, the right-falling stroke is incorporated into the dot stroke, and the remaining various turning strokes are incorporated into the fold stroke. In the present invention, the five basic strokes of horizontal, vertical, left-falling, dot, and fold are called single-stroke components. To reduce the number of homophones, preferably about 21 Chinese character components composed of two or more strokes with high character-forming frequencies or practical frequencies are arranged on the letter keys to participate in the encoding. Since the number of strokes is two or more, they are called multi-stroke components, or root characters, or radicals in the present invention to distinguish them from single-stroke components, or basic strokes. Multi-stroke components and single-stroke components are collectively called basic components, and sometimes simply referred to as components.
[0068] The code-taking rule of the first shape-based encoding method is as follows: For a single-character, take the corresponding codes of the first two basic components according to the writing order for encoding; or it is stipulated to take the corresponding code of the first or the last basic component according to the writing order. When there is only one basic component, only take the corresponding code of this basic component for encoding; for a compound character, divide the compound character into two parts according to the overall structure. The first-written part is the head part, and the second-written part is the remaining part. Take the corresponding codes of the first basic component of the head part and the first basic component of the remaining part according to the writing order for encoding.
[0069] I have long realized in my long-term encoding research that it is obvious whether a Chinese character is a left-right structure. It is very easy to divide a Chinese character with a left-right structure into two parts at the gap. However, for Chinese characters with an up-down or enclosed structure, it is sometimes not easy to divide them into two parts, and sometimes it is even difficult to distinguish whether a character is a single-character or an up-down or enclosed structure. Classifying according to whether a Chinese character is a left-right structure is the simplest and easiest to learn. For a Chinese character with a left-middle-right structure, the middle and right parts are regarded as the right part. Of course, strictly speaking, it is better to classify according to whether the right part forms a character after division.
[0070] If all Chinese characters are divided into left-right structure and non-left-right structure, encoding can still be carried out, and the appendix can still be used. Figure 1Coding, that is to say, the selected pinyin, basic components and codes remain unchanged. Coding is also composed of pinyin and shape coding. The second shape coding method is as follows: for Chinese characters with left and right structures, the corresponding code coding of the first basic component of the writing order of the left part and the right part is taken respectively; for Chinese characters with non-left and right structures, the corresponding code coding of the first and last basic components of the Chinese character is taken in the writing order. If there is only one basic component, only the corresponding code coding of this basic component is taken or the coding of this basic component is taken twice. At this time, for Chinese characters with non-left and right structures, the corresponding code coding of the first two basic components cannot be taken in the writing order, because it will cause repeated codes, but the corresponding code coding of the first and last basic components of the Chinese character should be taken in the writing order. Since it is very clear whether a Chinese character is a left and right structure, there will be no ambiguity. Except for a few Chinese characters such as "Shun, Chuan, Zhou, Er", Chinese characters with left and right structures are easy to have gaps in the left and right parts, as long as a vertical line is used to divide the Chinese character into two according to the gap. The left-right structure of Chinese characters sometimes encounters individual Chinese characters such as "川", "顺", and "州". "川" is composed of discrete strokes and is treated as a single character. The characteristic of "顺" is that it is composed of discrete strokes plus a Chinese character component to form a Chinese character. It is generally recommended that the entire discrete stroke is counted as the left part, and the other Chinese character component is counted as the right part. For example, in the character "顺", "川" is the left part, and "页" is the right part. Of course, this input method has a large tolerance for errors, and the "顺" character 丿 is the left part, and the rest of the part is the right part. In addition, "灬" cannot be divided into two parts by a vertical line.
[0071] For reducing unnecessary repeated codes, for a few center of gravity characters, it can also be stipulated that the second code of the shape part coding method can be encoded by the code of the first or last basic component at the center of gravity, and it is suggested to be encoded by the code of the first basic component at the center of gravity by the stroke order. The so-called center of gravity character refers to the specific body Chinese character whose radical of the meaning of the word is in the middle or tail of the Chinese character, such as the characters such as "win", "carry", "Ying", "competent", and the second code of the shape part coding method can be encoded by the corresponding code of the basic component "female" at the center of gravity. For example, the characters such as "fluorescent", the second code of the shape part coding method can be encoded by the corresponding code of the basic component "fire" at the center of gravity, because the part that does not include "fire" in the "fluorescent" character is actually phonetic. The center of gravity of the Chinese character with left-center-right structure and the same left part as the right part is in the middle part, so the second code of the shape part coding method can be encoded by the code of the last basic component at the middle part. As the character "辨", the last basic component "撇" encoding of the middle part of the shape part coding method two codes can be encoded. For Chinese characters with a left-middle-right structure, the center of gravity of the shape coding method or the auxiliary code is often on the "bird" part, and the second code is encoded according to the center of gravity.
[0072] Since the last basic component of a Chinese character is generally at the lower layer of the Chinese character, usually in the lower right corner, except for Chinese character components such as "fu" and "ge", where the dot in the upper right corner is the last component according to the writing order regulations. Therefore, when encountering Chinese characters containing components such as "fu" and "ge", as a fault-tolerant code, the dot in the upper right corner can be ignored, that is, the last stroke of "fu" and "ge" is the vertical and the left-falling stroke respectively.
[0073] The optimal arrangement of 21 multi-stroke components and five basic strokes on the keyboard is shown in the appendix Figure 1 as follows. A mapping relationship between 21 multi-stroke components, five basic strokes, letters, and punctuation marks is set as:
[0074]
[0075] Encode the multi-stroke components and basic strokes with the corresponding letters according to the set relationship.
[0076] The optimal arrangement of 25 multi-stroke components and five basic strokes on the keyboard is shown in the appendix Figure 2 as follows. A mapping relationship between 25 multi-stroke components, five basic strokes, letters, and punctuation marks is set as:
[0077] The optimal arrangement of 21 multi-stroke components and five basic strokes on the keyboard is shown in the appendix Figure 3 as follows. A mapping relationship between 21 multi-stroke components, five basic strokes, letters, and punctuation marks is set as:
[0078] The optimal arrangement of 21 multi-stroke components and five basic strokes on the keyboard is shown in the appendix Figure 4 as follows. A mapping relationship between 21 multi-stroke components, five basic strokes, letters, and punctuation marks is set as:
[0079] Some radicals will change slightly in shape after forming words, and the shapes of traditional and simplified characters will also change. They must be regarded as the same basic components and encoded with the same letters. Such basic components include 亻 and 人, 讠 and 言, 钅 and 金, 氵 and 水, 氺, 扌 and 手, 忄 and 心, 纟 and 糹, 火 and 灬, etc., which are characterized by the same origin. 艹 and 廾 are similar. Some components of the original input method, such as 亻 contains 人, 忄 contains 心, 氵 contains 水, 扌 contains 手, although they are of the same origin and easy to remember, some people say that it is inconvenient to remember because of the different shapes. On the contrary, for those that only differ in traditional and simplified characters, such as 讠 and 言, 钅 and 金, which remain in the same position in Chinese characters, they are easy to remember. Therefore, two kinds of error-tolerant codes are introduced. One kind of error-tolerant code is: easy to remember is the first priority. In the new invention, the basic components that only distinguish between traditional and simplified characters are used, such as "亻" and "人" are given priority to "亻", 氵 and water, 氺 are given priority to 氵, 扌 and hand are given priority to 扌; 忄 and heart are given priority to 忄, fire and 灬 are given priority to fire, and "人" with the same origin as 亻, water and 氺 with the same origin as 氵, hand with the same origin as 扌, and 灬 with the same origin as fire will be compatible in the form of error-tolerant codes, that is, "人" can be encoded with the code of "亻", but it is considered an error-tolerant code. Similar so-called multi-stroke components of the same origin will appear in the form of error-tolerant codes. Another method is to use the same origin as the code, and use the strokes of "人" as the error-tolerant code of "亻", and so on. Basic components can also contain individual components similar to it, encoded with the same letter. For example, the component "土" can contain "士". Since the two components only differ in stroke length, coding them as the same component may be more in line with brain reaction habits. 纟 and 壹 are also very similar in shape, so 纟 can also be stipulated to contain 壹, but of course it can also be arranged in a different way.
[0080] Since the second shape coding method also has the problem of constantly distinguishing whether it is a left-right structure Chinese character, the third shape coding method is relatively simple and easy to remember. This coding rule is used in the coding example, and the attached Figure 1 The phonetic code lists pinyin and phonetic letters for selection.
[0081] Coding example: such as the coding of "Han", the initial consonant is h, the final vowel is an, and the phonetic code part is han. The first code of the shape coding method is a of the multi-stroke component 氵 according to the writing order. The second code takes the first stroke "折" of the right part of the Chinese character in the writing order. The coding of "折" is z, so the coding of "Han" is "hanaz". If the phonetic letters are adopted, it is "h勹", and the position on the corresponding keyboard is "h,". So the coding of "Han" is "h,az". Another example is the coding of "字", the phonetic code part is zi. When the shape coding method is used, the first basic component of "字" is 宀 according to the stroke order, which is encoded as b. The Chinese character is a non-left-right structure Chinese character, and the code h of the last basic component "横" of "字" is taken in the writing order again, so the coding of "字" is "zibh". If the phonetic letters are adopted, the phonetic code part is still "zi", so the coding of "字" is "zibh". For example, the word "this", the full spelling is zhe, and when the shape part coding method is used, the first basic component of the Chinese character is taken as "dot" in the order of writing, and the code is "d". The non-left-right structure Chinese character takes the code "l" of the last basic component 辶 in the order of writing. The shape part coding method of "this" is just "dl", and the coding is just "zhedl". Since whether the present invention is a retroflex consonant is not very meaningful, and the retroflex consonant southerners can not read it accurately, so the retroflex consonant can be removed, and it is also OK to be encoded as "zedl". For another example, the coding of "wood", the double spelling is mu, and the Chinese character has only one basic component "wood", and the code is m. The shape part coding method of "wood" is just "m", so the coding of wood is just mum. In order to pursue uniform code length, it can also be stipulated that the Chinese character with only one basic component can also take the code of the first or last stroke or repeat the code of the basic component as the second code of the shape part coding method. This coding example does not make such a provision.
[0082] Attached Figure 5 In Chinese, the number keys are used, and it is a bit inconvenient to type across rows. Therefore, since the w key and the y key are vacant, the frequency of p in Chinese is very low, and the finals are arranged on the p key, and there will be almost no word duplication during encoding. The same is true for the n key and the r key. Therefore, ei, en, eng, ou, and ong are arranged on the w, r, y, n, and p keys. At this time, a mapping relationship diagram of the letter punctuation number keys on the English keyboard and the pinyin finals and phoneme letter finals is shown in the attached figure. Figure 6 As shown:
[0083]
[0084] Attached Figure 6 The frequencies of the Chinese initials k and r are almost the same, so the r key in the picture can be replaced by the k key.
[0085] Attached Figure 6The arrangement is quite regular, that is, it is divided into area a, area o, and area e according to the first letter of the pinyin, and each area is arranged in the order of a, o, e, i, u, n, and ng. Area a has ao, ai, an, and ang, which are arranged on the four punctuation keys respectively. Area 0 has ou, ong, which are arranged on the n or p key, or on the k and p keys. Area e has ei, en, and eng, which are arranged on the w, r, and y keys respectively. It conforms to the keystroke rules and is convenient to type. The higher frequency finals are arranged on the keys that are convenient to type. For example, the higher frequency en and ou in Chinese are arranged on the r and n keys where the index finger is located, which are more convenient to type. Other lower frequency finals starting with e and o are arranged on other keys.
[0086] Here are some examples of encoding by radical. For example, the Chinese character "轰", which is pinyin hong, has a radical "车", and is encoded as ce. The encoding of "轰" is hongce. Another example is the Chinese character "狄", which is pinyin di, has a radical "犭", and is encoded as qu. The encoding of "狄" is diqu. It can be seen that the encoding of these Chinese characters can achieve no duplicate codes for single words, and there is no need to consider the right part or the end of the Chinese character, which is very convenient. However, sometimes duplicate codes may occur for words. For example, "狄" and the word "地区" have duplicate codes.
[0087] In order to improve the input speed, simplified codes are designed for frequently used characters. The simplified code is to take the first 1, 2 or 3 codes of the complete code of the commonly used Chinese characters, and then press a space bar to enter the Chinese character. Since the phonetic code is required to come first and the auxiliary code is required to come later, many Chinese characters need to enter the simplified code of the Chinese character. Therefore, the encoding of a single character is actually mainly the phonetic code, supplemented by the auxiliary code. The shape code plays the role of the auxiliary code. For general commonly used characters, it is enough to enter the first code of the shape code.
[0088] Since there are only about 400 pinyins of Chinese characters, there are only about 400 secondary simple codes of Chinese characters, and the coding space of the present invention has 729, therefore, for the remaining about 300 coding spaces, simple code words can also be set up to further improve the typing speed. As the pinyin of Chinese characters does not have the form of kian, the double spelling coding also does not have the form of ky, and "k", "y" are respectively the initial consonants of "can" and "with", so ky can be used as the coding of "can". Since this input method is provided with more than 300 simple code words, theoretically the phrase input speed is faster than individual characters, so this will obviously improve the input speed of Chinese characters. After tapping the key where the simple code of a certain Chinese character or phrase is located on the computer, tap the space bar again, and the corresponding Chinese character or phrase can be input.
[0089] Word input is the most commonly used method to improve the input speed of Chinese characters. Since the phonetic code is specified first and the shape-based encoding method is specified later, word input makes full use of the phonetic code for input. When using the phonetic code for word input, full spelling or double spelling can be adopted. Taking Chinese pinyin as an example, if full spelling is used, just input the Chinese pinyin of each character. Simplified spelling can also be adopted. The method is as follows:
[0090] a. For two-character words, take the initial consonant of the first character and the pinyin code of the initial consonant and final of the second character and input them in sequence; or take the initial consonant and final of the first character and the initial consonant of the second character and input their pinyin codes in sequence. For example, the simplified spelling of "coding" is bma or bianm.
[0091] b. For three-character words, take the initial consonants or the codes of the first letters of the pinyin of each character and input them in sequence, and then input a space; for example, the simplified spelling code of "computer" is "jsj". Of course, it can also be specified to take the first code of the first and second characters, that is, the code of the initial consonant, and then take the first two codes of the third character. It can also be specified to take the first two codes of the first character, and then take the first code of the second and third characters, that is, the code of the initial consonant.
[0092] c. For four-character and above words, take the codes of the initial consonants of the first three characters and the last character and input them in sequence; for example, for the four-character word "science and technology", take the simplified spelling codes of the initial consonants of each character as "kxjs". Of course, it can also be specified that for four-character and above words, take the first letters or initial consonants of the pinyin of each character in the phrase for encoding.
[0093] Using the first-right pinyin input method software, press the key corresponding to the encoding of a certain Chinese character or phrase on the computer keyboard to complete the input. Generally, it is specified that Chinese characters or phrases without homophones and reaching the specified code length will be automatically input onto the screen. For those with a code length less than the specified length, press the space bar. For single characters or phrases with homophones, select according to the prompt line. If double spelling is used for the phonetic code, the maximum code length is four keys. If full spelling is used for the phonetic code, the code length is uncertain. The present invention is compatible with words and characters.
[0094] Now many people use voice input or pinyin to input Chinese characters. Since there are many homophones in Chinese characters, homophone errors are likely to occur. This input method software provides a powerful function for modifying homophones, that is, enter the homophone modification function, move the cursor to the front or back of the incorrect homophone. Note that either uniformly specify to move the cursor to the front of the Chinese character or uniformly specify to move the cursor to the back of the Chinese character. At this time, the software automatically recognizes the pronunciation of the Chinese character, and there is no need to input the phonetic code part of the present invention. As long as the shape-based encoding method is input, it is equivalent to inputting the complete encoding of the Chinese character. Those without homophones will automatically replace the original Chinese character. For those with individual homophones, select according to the prompt line, and the selected Chinese character will automatically replace the originally input incorrect Chinese character.
[0095] An example of coding that can basically avoid the trouble of selecting homophones: When the pinyin of a phrase is the same, just enter the first code of the shape code of the first character and the second character in the phrase after the pinyin. For example, the pinyin dili has five homophones, and geography is the most commonly used. The first code of the auxiliary code of the first Chinese character can be omitted, so the code of geography is still dili, and the first code of the auxiliary code of the first Chinese character "地" must be added after the pinyin, which is t, so the code is dilit. For land power, after adding the first code t of the auxiliary code of the first Chinese character "地" after the pinyin, it will also have a duplicate code with "地利", so the first code z of the auxiliary code of the second Chinese character "力" must be added, and the code is dilitz. For sharpening, just add the first code h of the auxiliary code of the first Chinese character after the pinyin dili, and it is coded as dilih. For dilia dripping, just add the first code a of the auxiliary code of the first Chinese character after the pinyin dili, and it is coded as dilia.
[0096] In order to shorten the code length and increase the input speed, the pinyin part of ng is represented by v. For example, the phonetic code of "zhuang" is zhuang, which can be represented by zhuav.
[0097] Since the present invention has the phonetic code first, it is fully compatible with the pinyin input method and the Zhuyin input method. In order to make it more popular and compatible, the present invention also creatively adopts a two-color candidate character technology, that is, in the candidate window, after inputting letters, words will appear for selection. The words that do not adopt the shape-part encoding method are in a certain color, such as green, and the Chinese characters that adopt the shape-part encoding method, that is, the Chinese characters that adopt the Chinese character code, are in another color, such as black. After inputting black several times, the system will think that it understands the Chinese character code technology and give priority to inputting Chinese characters according to the Chinese character code to increase the speed.
[0098] An error-tolerant code is also provided, so that the Chinese characters to be input can be displayed even when some Chinese characters are incorrectly input. It should be noted that the letters in this specification, claims and drawings are not case-sensitive, and the upper and lower case letters are equivalent.
Claims
1. A computer Chinese character coding keyboard input method, namely the first-right pinyin input method, classifies the various strokes of Chinese characters into five basic strokes of horizontal, vertical, left-falling, dot and fold according to the provisions of the National Language Committee, and its characteristics are: (1) The code consists of two parts, one is the phonetic code, i.e., pinyin, or phonetic code, and the other is the auxiliary code. The auxiliary code is divided into the shape code and the radical code. The two parts of the Chinese character code can be placed in front or behind. Once selected, they cannot be changed. Generally, the phonetic code comes first and the auxiliary code comes later. (2) The phonetic code can be in full spelling, double spelling, simplified spelling or incomplete pinyin, or can be in Taiwan phonetic notation and phonetic alphabet; a mapping relationship between each letter punctuation mark number key and pinyin vowel and phonetic alphabet vowel: Another mapping relationship between the letter punctuation mark number keys and the pinyin finals and phoneme letter finals: (3) The first code-taking rule of the shape part coding is: for a single character, the corresponding code of the first two basic components is taken in the writing order, or the corresponding code of the first and last basic components is taken in the writing order. When there is only one basic component, only the corresponding code of this basic component is taken. Of course, it can also be stipulated that the single character is taken in the writing order. The corresponding code of the first and last basic components of the Chinese character; A compound character is divided into two parts according to the overall structure. The part containing the first stroke of the Chinese character in the writing order is the head part, and the part written afterwards is the remainder part. The corresponding codes of the first basic component of the head part and the first basic component of the remainder part are respectively taken in the writing order. The second code-taking rule of the shape part coding is: for Chinese characters with left-right structure, the corresponding code of the first component of the left part and the right part in writing order is taken respectively; for Chinese characters without left-right structure, the corresponding code of the first and last basic components of the Chinese character is taken in writing order, and if there is only one basic component, only the corresponding code of this basic component is taken or the code of this basic component is taken twice in succession; or it is stipulated that for Chinese characters without left-right structure, the basic component code of the first basic component of the Chinese character and the lower right corner of the Chinese character (the lower right corner of the enclosed part is taken in the case of enclosing structure) is taken in writing order; The third rule for taking the shape code: The first code of the shape code is: regardless of anything, take the code of the first basic component of the Chinese character in the order of writing; the second code of the shape code is to start from the right side of the first basic component of the Chinese character, scan or look from left to right, if the Chinese character can be divided into two by a vertical line without cutting the strokes of the Chinese character, then the Chinese character is a left-right structure, and the part to the right of the vertical line is the right part of the Chinese character, and then take the code of the first basic component of the right part of the Chinese character in the order of writing. If the Chinese character cannot be divided into two by a vertical line without cutting the strokes, scan the lower half or lower half of the Chinese character from left to right, and find the character in the order of writing. The code encoding of the last basic component in the sequence or the corresponding code of the basic component in the lower right corner of the Chinese character is taken for encoding; Chinese characters with left and right structures often have obvious gaps, which are easy to distinguish, so it is also possible not to use vertical lines to split. The second code only needs to start from the right side of the first basic component of the Chinese character, scan from left to right, find the gap between the left and right parts of the entire Chinese character, and the part on the right side of the gap is the right part of the Chinese character, and then take the code encoding of the first basic component of the right part of the Chinese character in the writing order. If there is no gap on the left and right of the Chinese character, scan from left to right or look at the lower half (or lower half or lower layer) of the Chinese character, and find the code encoding of the last basic component of the Chinese character in the writing order; In short, the first code of the shape code is: the code code of the first basic component of the Chinese character in the order of writing; the second code of the shape code is to scan the Chinese character from left to right first, if the Chinese character is a left-right structure, and the right part can be found, the code code of the first basic component of the right part of the Chinese character is taken in the order of writing; If the right part cannot be found, scan the lower half of the Chinese character from left to right, and find the code of the last basic component of the Chinese character in the writing order; (3) When encoding by shape components, it is preferred that five basic strokes and 21 multi-stroke components participate in the encoding. When selecting multi-stroke components, the character-forming frequencies and the rate of homophonic characters of these multi-stroke components in 3,755 common Chinese characters are mainly considered. The multi-stroke components are encoded by the first letter or initial consonant of their pinyin. When the first letters of the pinyin of several multi-stroke components are the same, they are arranged according to the method of homophonic and near-position, and precise positioning calculation is carried out. The method is that some radicals have a strong ability to form characters and a high usage frequency, but their distribution is not uniform in the syllables of the initial consonants or finals of the 26 letters. In the syllables of certain initial consonants and finals, the number of Chinese characters where these radicals are located is very small. If these radicals are encoded with a specific letter, it can effectively avoid character homophony. On the basis of the homophonic and near-position arrangement, radicals that usually appear in the first code of the shape component encoding of Chinese characters but rarely appear in the second code of the shape component encoding are specifically encoded with vowel letters E, I, A, O, U. According to the ability to avoid homophony, 口, 氵, 艹, 口, 木, 扌, 亻, 土, 钅 are each encoded with a letter. The radicals 辶, 忄, 纟, 日, 火, 讠, 足, 石, 王, 疒 are also each encoded with another letter. The five basic strokes of horizontal, vertical, left-falling, dot, and fold, as well as multi-stroke components such as 王, 土, 钅, 口, 忄, 宀, 女, 木 are all encoded by the first letter of their pinyin, and the remaining multi-stroke components are encoded according to the method of homophonic and near-position; Taking the word frequency总汇 of Beijing Language and Culture University statistically by words and characters as an example, through measurement and calculation, it is found that for Chinese characters encoded with multi-stroke components such as 口, 日, 人 (亻), 土, 讠, 月, 氵, 扌, 火, 金, 艹, 虫, 木, if encoded by strokes, the sum of the frequencies of homophonic characters and the sum of the frequencies of other homophonic characters are very high, so they should be selected. For 宀 (the radical 穴 is encoded as 宀), 辶, 竹, 火, 石, 山, if encoded by strokes, the sum of the frequencies of homophonic characters are 888, 767, 191, 177, 128, 59 respectively, and the sum of the frequencies of other homophonic characters are 1,209, 1,916, 523, 563, 363, 229 respectively. Thus, the sum of the frequencies of 宀 (the radical 穴 is encoded as 宀) and 辶 is relatively high, and it is recommended to select them. The frequencies of 竹 and 火 are the second, and it is also recommended to select them. The frequency of the radical 山 is relatively low, and it is not very recommended to select it. As for 石, it needs to be judged according to the situation of 足. 足 is encoded by the first stroke "vertical". The frequency of its own homophonic characters is about 236 thousand. If the homophonic characters of Chinese characters containing "足" are included, the frequency is 316 thousand. If 足 is encoded by "口", then when the Chinese characters encoded by the shape component of "足" are counted, the sum of the frequencies is 502 thousand. Generally speaking, the sum of the frequencies of the radical 足 is relatively high, and it is recommended to select the component "足", while 石 is落选. Of course, considering the large disparity in the frequencies of homophonic characters of Chinese characters with the radical 足, for example, the frequency frequencies of "路", "噜", and "卢" vary greatly, so it is also possible to select the component 石 instead of the component 足. Of course, it is also possible to select both 石 and 足 at the same time, but in this case, there will be two radicals on one key, and it is no longer one letter corresponding to one radical, which is not convenient to display on the mobile phone screen; Based on the arrangement of homophonic near-position method, the radicals that usually appear in the first code of the shape part encoding of Chinese characters but rarely appear in the second code of the shape part encoding are encoded with vowel letters E, I, A, O, U. According to the ability to avoid duplicate codes, 氵, 艹, 口, 木, 扌, 亻, 土, 钅 are each encoded with one letter, and the radicals 辶, 忄, 纟, 日, 火, 讠, 足, 石, 王, 疒 are also each encoded with another letter; the five basic strokes of horizontal, vertical, left-falling, dot, and fold, as well as multi-stroke components such as 王, 土, 钅, 口, 忄, 宀, 女, 木 are encoded according to the first letter of their pinyin, and the remaining multi-stroke components are encoded according to the homophonic near-position method; Among them, the first letters of the pinyin of 亻 and 日, 讠 and 月, horizontal and 火, 虫 and 艹, fold and 辶 and 足, vertical and 氵 and 扌 and 石 and 纟 are the same, which are r, y, h, c, z, s respectively; the first letters of the pinyin of 亻 and 日, 讠 and 月, horizontal and 火, 虫 and 艹 are the same. For the convenience of memory, they are arranged together on the keyboard in the left and right adjacent positions for easy memory; The first pinyin letters of both "艹" and "虫" are "c". According to the method of homophonic and adjacent positions, they can only be arranged on two adjacent keys, "c" and "v". Since "v" is a rare final, we only need to consider the number of Chinese characters that appear at the beginning of Chinese characters in the first pinyin letter "c". There are 3 characters with "虫", and the sum of frequencies is relatively low. There are 11 characters with "艹", and the sum of frequencies is much higher. Since "v" is used as a final, in order to avoid character-code conflicts, it is recommended to use the multi-stroke component with a relatively low sum of frequencies. Among common Chinese characters, when calculating the number of characters formed and the sum of frequencies of Chinese characters that appear in the second code of the shape-based encoding, the number is also lower for "虫", while much higher for "艹". Therefore, "虫" is encoded with "v", and "艹" is encoded with "c", which can be simply remembered as "草虫". The pinyin of "艹" is "cao", and the pinyin of "虫" is "chong". Arranged in alphabetical order, "艹" should also be encoded with "c", and "虫" with "v". The first letter of the pinyin of the horizontal stroke and the radical "火" is H. Since J has been arranged for the radical 钅, "火" can only be encoded with G adjacent to the left of the H key; The first letters of the pinyin of 日 and 亻 are both r. According to the homophonic near-position method, they can only be arranged on the two adjacent keys e and r on the left and right; from the perspective of avoiding duplicate codes of words, it is necessary to count the number and frequency sum of the basic components 日 and 亻 in Chinese characters with the finals ue and ie. Since the number of Chinese characters with the basic component "日" and the number of Chinese characters with 亻 are both relatively small, and the frequency sum is also very low and close, so only the usage frequencies of the radical "日" and 人 (including 亻) can be compared. According to statistics, the usage frequency of the radical 人 (including 亻) is about 2.5 times that of the radical "日". Therefore, the radical 人 (including 亻) is encoded with the first letter of its pinyin R, while the radical "日" is arranged on the key E next to the first letter of its pinyin R and encoded with E. The left part of the radical "日" is similar in shape to E, and it is easier to remember with E encoding; The initials of the pinyin for the radicals of vertical line, silk radical, hand radical, and water radical are all s. The basic stroke of the vertical line is very common, and of course, it is encoded with the s key. The I key, O key, and A key can be regarded as adjacent to the S key; the silk radical, hand radical, and water radical can be arranged on the I, O, and A keys. Therefore, the author has carried out quantitative calculations using operations research. Among the Chinese characters with the initial of a in pinyin, there is 1 Chinese character containing the water radical, with a frequency of 5,920, and 2 Chinese characters containing the hand radical, with a total frequency of 64,779. So it is better to encode the water radical with a. Among the Chinese characters with the initial of o in pinyin, there is 1 Chinese character containing the water radical, and there are no Chinese characters containing the hand radical. So it is more appropriate to encode the hand radical with o. Among the Chinese characters starting with the initials i, o, and a in pinyin, there are no Chinese characters with the silk radical. Since o and a have already been encoded with the hand radical and water radical respectively, considering comprehensively, the silk radical is encoded with the remaining i. From the perspective of the frequencies of the finals i, o, and a, i is the highest, a is the second, and o is the lowest. From the perspective of the usage frequencies of the radicals, the silk radical is the lowest, the water radical is the second, and the hand radical is the highest. From the perspective of the encoding word and character homophone, the finals with high frequencies are suitable for matching with the multi-stroke components or radicals with low usage frequencies, and the finals with low frequencies are suitable for matching with the multi-stroke components or radicals with high usage frequencies. Therefore, it is appropriate to encode the silk radical with i, the hand radical with o, and the water radical with a. And the silk radical, hand radical, and water radical are exactly encoded with the initials i, o, and a of the finals respectively, which is easy to remember. The pinyin of the silk radical is si, which consists of two letters. The pinyin of the hand radical is shou, which consists of four letters. The pinyin of the water radical is shui. So, starting from the upper row of the keyboard from left to right, and then to the middle row of the keyboard, according to the number of pinyin letters, when the number of pinyin letters is the same, they are arranged in the order of the phonetic sequence. The silk radical, hand radical, and water radical are respectively arranged on the i, o, and a keys, and are encoded with the corresponding letters. The initials of the pinyin for the radicals of moon and speech radical are both y. According to the method of arranging adjacent keys with the same initial sound, they can only be arranged on the two adjacent keys of y and u. From the perspective of avoiding homophone of words and characters, it is necessary to consider the frequencies or the sum of frequencies of the basic components of "moon" and speech radical appearing in Chinese characters with the finals iu or ou. The number of Chinese characters with the basic component of "moon" at the beginning of the character is 2, while the number of Chinese characters with the speech radical at the beginning of the character is 8. The sum of the frequencies (the sum of usage frequencies) of these Chinese characters is also relatively low for the Chinese characters with the moon radical. So it is more appropriate to encode the basic component of moon with u and the speech radical with y. At this time, there is almost no homophone when inputting the shape radical encoding. Also, comparing the number of Chinese characters with the initial of y in pinyin, the number of Chinese characters with the "moon" at the beginning of the character is 10, and the number of Chinese characters with the speech radical at the beginning of the character is 15. The sum of the frequencies, that is, the sum of the usage frequencies, is also higher for the speech radical. So it is more appropriate to encode the basic component of the speech radical with y and the basic component of "moon" with u. And u happens to be the initial of the pinyin of "moon", which is very easy to remember. In addition, from the perspective of the phonetic sequence, the pinyin of the speech radical is yan, and the pinyin of the moon is yue. If arranged from left to right according to the phonetic sequence, it should also be that the basic component of the speech radical is encoded with y and the basic component of "moon" is encoded with u. The pinyin initials of "zhe", "chuo", "zu", and "zhu" are all "z". The stroke "zhe" is very common and is of course represented by "z". According to the method of homophonic and adjacent positions, "chuo", "zu", and "zhu" can only be arranged on the remaining keys "l", "Q", and "F". Among them, "l" and "z" are respectively located at the far right of the second row and the far left of the third row of the keyboard, and can be considered adjacent. And the letters to the right of "z" in the lower row of the keyboard have already been arranged with radicals. Therefore, according to the rule of homophonic and adjacent positions, the "Q" key and "F" key on the upper row of the keyboard can also be勉强 considered adjacent. Since the frequency of the initial consonant "L" in Chinese is much more common than that of the initial consonants "F" and "Q", the pinyin initial "L" is arranged first. "Chuo" can only appear at the end of a Chinese character, and the sum of its frequencies is 0. When the pinyin initial is "L", there are 7 Chinese characters with "zu" at the beginning of the word, and the sum of their frequencies is 109,352. There are 12 Chinese characters with "zhu" at the beginning of the word, and the sum of their frequencies reaches 16,734. It can be seen that neither "zu" nor "zhu" is very suitable for encoding with "L". From the perspective of avoiding character and word homophones, encoding "chuo" with "L" can make the character and word homophones 0. This is a very ingenious arrangement. There are 3 Chinese characters with "zu" at the beginning of the word and the pinyin initial "F", and the sum of their frequencies is 293. There are also 3 Chinese characters with "zhu" at the beginning of the word and the pinyin initial "F", and the sum of their frequencies is 16,022. There are 5 Chinese characters with "zu" at the beginning of the word and the pinyin initial "Q", and the sum of their frequencies is 7,626. There are 5 Chinese characters with "zhu" at the beginning of the word and the pinyin initial "Q", and the sum of their frequencies is 9,664. From the perspective of the number, both "zu" and "zhu" are relatively few and close on the keys with the pinyin initials "F" and "Q". From the perspective of the sum of frequencies, among the Chinese characters with the pinyin initials "Q" and "F", there are more radicals of "zhu". Therefore, for "zu" and "zhu", one can choose either one from "F" and "Q" at will. From the perspective of convenient key pressing, it is more appropriate to encode "zhu" with "F" and "zu" with "Q". Of course, it is also possible to encode "zhu" with "Q" and "zu" with "F". From the perspective of memory, the pinyin initials are all "z", and they can only be arranged according to the strokes. The first stroke of "zu" is a vertical stroke, the first stroke of "zhu" is a left-falling stroke, and the first stroke of "chuo" is a dot. The order of arrangement is vertical stroke, left-falling stroke, dot. The order from left to right on the keyboard is "Q", "F", "L". Therefore, "zu", "zhu", and "chuo" are respectively arranged on the "Q", "F", and "L" keys from left to right according to their first strokes of vertical stroke, left-falling stroke, and dot, and are respectively encoded with "Q", "F", and "L". From the perspective of similarity in shape, the head and tail of "zu" are similar to "Q", the left or right part of "zhu" is similar to "F", and "chuo" is also similar to "L", which is easy to remember. Of course, it is also possible to sort from the perspective of the number of letters and the order of sounds. The pinyin of "zu" is "zu", which consists of only two letters, so it is arranged on the "q" at the far left of the keyboard. And the pinyin of "zhu" and "chuo" are both three letters. According to the order of sounds, "zhu" and "chuo" are respectively arranged from left to right on the "f" and "l" keys and encoded with the corresponding letters; The optimal arrangement of 21 multi-stroke components and five basic strokes on the keyboard is shown in Figure 1 of the attached drawings; A mapping relationship between 21 multi-stroke components, five basic strokes, letters, and punctuation marks is set as follows: a——Water radical b——Roof radical c——Grass radical d——Dot e——Sun radical f——Bamboo radical g——Fire radical h——Horizontal stroke i——Silk radical j——Metal radical k——Mouth radical l——Moving radical m——Wood radical n——Female radical o——Hand radical p——Left-falling stroke q——Foot radical r——Person radical or Radical of single person s——Vertical stroke t——Earth radical u——Month radical v——Insect radical w——King radical x——Heart radical y——Speech radical z——Folding stroke According to the set relationship, multi-stroke components and basic strokes are respectively encoded with corresponding letters; The preferred arrangement of 25 multi-stroke components and five basic strokes on the keyboard is shown in Figure 2; A mapping relationship between 25 multi-stroke components, five basic strokes, letters, and punctuation marks is set as: a——Water radical b——Roof radical c——Grass radical d——Dot e——Sun radical f——Bamboo radical g——Fire radical h——Horizontal stroke i——Silk radical j——Metal radical k——Mouth radical l——Moving radical m——Wood radical n——Female radical o——Hand radical p——Left-falling stroke q——Foot radical r——Person radical s——Vertical stroke t——Earth radical u——Month radical v——Insect radical w——King radical x——Heart radical y——Speech radical z——Folding stroke ;——Fish, ——Mountain.——Ear radical / ——Grain; The preferred arrangement of 21 multi-stroke components and five basic strokes on the keyboard is shown in Figure 3; A mapping relationship between 21 multi-stroke components, five basic strokes, letters, and punctuation marks is set as: a——Silk radical b——Roof radical c——Grass radical d——Dot e——Radical of single person f——Moving radical g——Fire radical h——Horizontal stroke i——Hand radical j——Metal radical k——Mouth radical l——Foot radical m——Wood radical n——Female radical o——Water radical p——Left-falling stroke q——Bamboo radical r——Sun radical s——Vertical stroke t——Earth radical u——Speech radical v——Insect radical w——King radical x——Heart radical y——Month radical z——Folding stroke The preferred arrangement of 21 multi-stroke components and five basic strokes on the keyboard is shown in Figure 4; A mapping relationship between 21 multi-stroke components, five basic strokes, letters, and punctuation marks is set as: a——Hand radical b——Roof radical c——Insect radical d——Dot e——Radical of single person f——Water radical g——Fire radical h——Horizontal stroke i——Silk radical j——Metal radical k——Mouth radical l——Grain m——Wood radical n——Female radical O——Mountain p——Left-falling stroke q——Foot radical r——Sun radical s——Vertical stroke t——Earth radical u——Speech radical v——Grass radical w——King radical x——Heart radical y——Month radical z——Folding stroke When encoding by the method of the first radical of a character, in order to be compatible and reduce the number of homophonic codes, the five basic strokes, namely horizontal stroke, vertical stroke (including vertical hook), left-falling stroke, dot (including right-falling stroke), folding stroke and 11 radicals such as "King, Person (Radical of single person), Earth, Speech (Speech radical), Metal (Metal radical), Mouth, Heart, Grass, Roof, Female, Wood", etc. are still encoded with the initial letters of their pinyin, namely H, S, P, D, Z, W, R, T, Y, J, K, X, C, B, N, M respectively. One can still encode according to the original encoding rules of the form radical and input the second code of the form radical to input Chinese characters conveniently. However, it is also possible not to input the second code of the form radical. In this case, according to the prompt line, one can use the numeric keys or the space bar to select; For the remaining radicals that appear at the beginning of a character, they are encoded according to the first letter of the pinyin of the radical, the first letter of the phonetic of the radical, and the first basic stroke of the radical in the stroke order. The encoding length is also two to three digits. The advantage of the radical-at-the-beginning method of character encoding is that once the first letter of the pinyin of the radical located at the beginning of a certain Chinese character is input, the Chinese character can be found through the prompt line, thus greatly enhancing the universality of the present invention; The following lists the encodings of the radicals that commonly appear at the beginning of Chinese characters: 犭——qu, 犬——quh, 气——qi, 其——qih; 文——we; 阝——er, 耳——erh, 二——erp; 日——ri; 田——ti; 月——yu, 鱼——yup, 羊——ya, 又——yo, 酉——yoh, 业——ye, 幺——yaz; 雨——yuh, 羽——yuz; 衤——yi, 音——yid; 片——pi; 石——sih, 饣(食)——sip, 纟——si, 氵——su, 尸——siz, 山——sa, 礻——sid, 扌——so, 鼠——sup, 身——se; 大——da, 歹——dah, 豆——do, 刀——daz; 风——fe, 父——fu, 缶——fo; 革——ge, 广——gu, 鬼——gup, 骨——gus, 谷——gup, 工——go, 瓜——gup, 弓——goh; 火——hu, 禾——he, 户——hud, 黑——hes; 巾——ji, 角——jip, 几——jip, 见——jis, 斤——jip; 立——li, 力——liz, 里——lis, 龙——lo, 老——la, 卤——lus, 鹿——lu; 竹——zu, 足——zus, 走——zo, 舟——zop, 爪——zup, 自——zi, 豸——zip, 止——zis; 西——xi, 辛——xid, 夕——xip, 穴——xu, 血——xup; 寸——cu, 车——ce, 齿——ci, 赤——cit; 虫——co, 臣——ceh, 辰——ceh; 疒——bi, 白——ba, 鼻——bi, 贝——be, 八——ba, 比——bih; 鸟——ni, 牛——nip; 目——mu, 母——muz, 门——me, 米——mi, 马——ma, 毛——map, 麻——mad, 麦——mah, 矛——maz; The encoding of the radicals in the above radical-at-the-beginning method of character encoding is also the mapping relationship on the keyboard. It can be seen that for most radicals located at the beginning of a character, taking the first letter of the pinyin plus the first letter of the phonetic is sufficient. Another method of encoding for the radical-at-the-beginning method is to take the first letter of the pinyin of the radical plus the code of the first basic stroke of the radical in the stroke order. If it is still not clear, then add the first letter of the phonetic of the radical; Another relatively common radical coding that usually appears at the beginning of Chinese characters is: taking the first letter of the pinyin of the radical plus the code of the first basic stroke of the radical in the stroke order. If there are still homophonic characters, then add the code of the second basic stroke of the radical in the stroke order. For example: 犭——qp, 犬——qh, 气——qph, 其——qhs; 文——wd; 阝——er, 耳——erh, 二——erh; 日——rs; 田——ts; 月——yp, 鱼——ypz, 衤——yd, 又——yz, 雨——yh, 酉——yhs, 业——ys, 幺——yzz; 羽——yzd; 羊——ydp, 音——ydh; 片——pp; 石——sh, 饣(食)——sp, 纟——sz, 氵——sd, 尸——szh, 山——ss, 礻——sdz, 扌——sh, 鼠——sps, 身——sps; 大——dh, 歹——dhp, 豆——dhs, 刀——dz; 风——fp, 父——fpd, 缶——fph; 革——ghs, 广——gd, 鬼——gp, 骨——gs, 谷——gpd, 工——gh, 瓜——gpp, 弓——gz; 火——hd, 禾——hp, 户——hdz, 黑——hs; 巾——js, 角——jp, 几——jpz, 见——jsz, 斤——jpp; 立——ld, 力——lz, 里——ls, 龙——lh, 老——lhs, 卤——lsh, 鹿——ldh; 竹——zp, 足——zs, 走——zh, 舟——zpp, 爪——zpp, 自——zps, 豸——zpd, 止——zsh; 西——xh, 辛——xd, 夕——xp, 穴——xdd, 血——xps; 寸——chd, 车——ch, 齿——csh, 赤——chs; 虫——cs, 臣——chs, 辰——chp; 疒——bd, 白——bp, 鼻——bps, 贝——bs, 八——bpd, 比——bh; 鸟——np, 牛——nph; 目——ms, 母——mz, 门——md, 米——mdp, 马——mz, 毛——mp, 麻——mdh, 麦——mh, 矛——mzd.
2. The first-right pinyin input method according to claim 1, wherein: When arranging several other multi-stroke components with the same first letter or initial of the pinyin as "金, 水, 火, 艹, 人, 讠", the method of homophonic and adjacent positions is adopted, that is, the other multi-stroke components with the same first letter of the pinyin are arranged near the key position where this multi-stroke component is located. Since the keyboard letter keys are divided into three rows, they are generally arranged on the left or right of the same row as this multi-stroke component.
3. The first-right pinyin input method according to claim 1 is characterized in that: Basic components such as 氵, 艹, 口, 木, 扌, 钅, 亻, 女, 讠, 忄, 月, 虫, 土, 纟, 火, 日, 石, 王, 疒, 足 are all selected from the Chinese character radicals.
4. The first-right pinyin input method according to claim 1 is characterized in that: The so-called proximal positions are in order: from Q to P on the keyboard, then to A to L, then to Z to M, and then back to Q. From left to right in the upper row of the keyboard, from Q to P, and then to the middle row of the keyboard, again from left to right, from A to L; then to the lower row of the keyboard, again from left to right, from Z to M, and then back to the Q key.
5. The first-right pinyin input method according to claim 1 is characterized in that: When the first code of the shape part encoding of a non-left-right-structured Chinese character is the same as the first code of the shape part encoding of a left-right-structured Chinese character, the non-left-right-structured Chinese character takes the simplified code first. After inputting the phonetic code of the Chinese character, then input the first code of the shape part encoding and press the space bar to input the left-right-structured Chinese character.
6. The first-right pinyin input method according to claim 1 is characterized in that: If 亻 is encoded with r, since 亻 almost only appears at the beginning of a character and rarely appears as the second code of the shape part encoding, only 5 亻 appear as the second code of the shape part encoding, which can greatly avoid character and word homophony, and is easy to remember; if 讠 is encoded with y, since 讠 almost only appears at the beginning of a character and rarely appears as the second code of the shape part encoding, only 19 of the traditional form of 讠, 言, appear as the second code of the shape part encoding, and the quantity is not large either, so it can also better avoid character and word homophony; 艹 also almost only appears at the beginning of a character and rarely appears as the second code of the shape part encoding, and there are few Chinese characters with the v (ü) final, with low frequency, so it can greatly avoid character and word homophony; Of course, if 亻 is encoded with e and 日 is encoded with r, it also conforms to the arrangement of homophonic proximal positions; if 讠 is encoded with u and 月 is encoded with y, it also conforms to the homophonic proximal positions; if 艹 is encoded with v and 虫 is encoded with c, it also conforms to the homophonic proximal positions; similarly, 纟, 扌, and 氵 can also be interchanged in the mapping on the i, o, and a key positions; but from the perspective of the homophony of the encoded characters and words, it is not very appropriate; If, on the basis of homophonic proximal positions, the number of radicals is not considered and only arranged in alphabetical order, then, 亻 including 亻 is encoded with e, 日 is encoded with r, 扌 is encoded with i, 氵 is encoded with o, a is encoded with 纟, 火 is encoded with g, 艹 is encoded with c, v is encoded with 虫, 竹 is encoded with q, 辶 is encoded with f, and 足 is encoded with l; It can be seen that in this input method, the 5 strokes and 8 multi-stroke components such as 王 (king), 土 (earth), 钅 (metal radical), 口 (mouth), 忄 (heart radical), 疒 (illness radical), 女 (female), and 木 (wood) are all encoded by the first letter of their pinyin. 人 (person), 讠 (speech radical), and 虫 (insect) are also encoded by the first letter of their pinyin. In this way, a total of 16 stroke radicals are encoded by the first letter of their pinyin. The actual number of multi-stroke radicals that need to be memorized is only 10, and they are arranged in the method of "same sound (first letter of pinyin) and adjacent position (adjacent positions on the keyboard)", and basically according to the number of letters in the pinyin of the radicals. For those with the same number of letters, they are also arranged in alphabetical order, and some also take into account the strokes and similarity in shape, which is very easy to remember. Among them, the 4 multi-stroke components 月 (moon), 纟 (silk radical), 扌 (hand radical), and 氵 (water radical) are encoded by the first letter of their韵母 (final). They are also arranged in the method of "same sound and adjacent position", and actually only 6 are arranged in the method of "same sound and adjacent position". These 6 are also arranged according to the number of letters and the English alphabetical order on the basis of "same sound and adjacent position", which is very easy to remember. To further shorten the memory time, the author has also compiled a mnemonic formula, that is, the basic strokes are the team leaders and are preferentially encoded by the first letter of their pinyin. Among the syllables with the same first two letters of pinyin, 人 (person), 讠 (speech radical), 艹 (grass radical), and 横 (horizontal stroke) are the team leaders, and 日 (sun), 月 (moon), 虫 (insect), and 火 (fire) are the teammates, and they are also arranged on the keys adjacent to the left and right of the team leaders. 艹虫 (grass and insect) is homophonic to 草虫 (grasshopper) and is arranged on the c and v keys adjacent to each other on the left and right, which is more conducive to memory. 折 (folding stroke), 辶 (walk radical), 足 (foot), and 竹 (bamboo) form a team, and the folding stroke is the team leader, and 辶, 足, and 竹 are the teammates; 竖 (vertical stroke), 纟 (silk radical), 扌 (hand radical), and 氵 (water radical) form a team, and the vertical stroke is the team leader, and 纟, 扌, and 氵; The encoding rule of the shape radical encoding is briefly remembered as "first, no right then last", which is very easy to remember; In this way, an ordinary person can remember it in three to five minutes.