Japanese full keyboard input method, device and electronic equipment

By converting Romanization to Hiragana and using a simplified Pinyin mapping method and Hiragana tree database for decoding, the problem of inaccurate output in the Japanese full keyboard input method is solved, improving input efficiency and accuracy.

CN113867546BActive Publication Date: 2025-11-11UNIV OF SCI & TECH OF CHINA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111162150.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-11-11
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing Japanese full keyboard input methods suffer from inaccurate Romanization decoding, resulting in a low rate of matching the output results with user expectations, low input efficiency, and difficulty for users to accurately input Japanese characters with similar pronunciations.

Method used

By converting the Romanized input via keystrokes into Hiragana, using a simplified mapping method to process Romanized input that cannot be directly converted, and then using a Hiragana tree database for decoding, accurate Japanese character strings are obtained.

Benefits of technology

It improved the accuracy of output results in matching user expectations, enhanced input efficiency, and enabled the function of Japanese abbreviation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113867546B_ABST
    Figure CN113867546B_ABST
Patent Text Reader

Abstract

The application discloses a Japanese full keyboard input method, device and electronic equipment, and relates to the technical field of input methods. The Japanese full keyboard input method comprises the following steps: in a Japanese full keyboard input mode, receiving key input information, wherein the key input information comprises one or more continuous Roman alphabetic information corresponding to the keys; converting the key input information into hiragana information according to a hiragana conversion table; obtaining a Japanese character string corresponding to the key input information according to the hiragana information; and outputting the Japanese character string. The application converts all Roman alphabetic information input through the keys into corresponding hiragana information, decodes all the converted hiragana information as a whole, and obtains an accurate Japanese character string, thereby improving the coincidence rate of the output result and the user's expectation, and further improving the input efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of input method technology for electronic devices, and more particularly to a Japanese full keyboard input method, device, and electronic device. Background Technology

[0002] With the continuous development of social technology, various electronic devices have become ubiquitous in people's lives and work. Input methods, as the interface between humans and electronic devices, play a crucial role in human-computer interaction. With the deepening of reform and opening up and the advancement of globalization, more and more users are pursuing multilingual input methods for electronic devices. Japanese full keyboard input, as one of the most important input methods currently available, has seen its convenience, efficiency, and accuracy continuously improved, making it a hot research topic in related fields.

[0003] In current mainstream Japanese full keyboard input methods, some Romanization characters need to be decoded into English, resulting in a low rate of matching the output results with the user's expectations, which in turn reduces input efficiency.

[0004] Furthermore, analysis of user input behavior revealed that users did not have a clear memory of the pronunciations of many Japanese characters with similar sounds. When using the full Japanese keyboard input method, users were unable to obtain the desired Japanese characters when they made input errors. Summary of the Invention

[0005] In view of the above, the present invention aims to provide a Japanese full keyboard input method, apparatus, and electronic device, and accordingly proposes a computer-readable storage medium, thereby improving the matching rate of output results with user expectations and thus improving input efficiency.

[0006] The technical solution adopted in this invention is as follows:

[0007] In a first aspect, the present invention provides a Japanese full keyboard input method, comprising:

[0008] In Japanese full keyboard input mode, it receives key input information, which includes one or more consecutive Romanized sounds corresponding to the keys;

[0009] The key input information is converted into hiragana information based on the hiragana conversion table;

[0010] Obtain the Japanese character string corresponding to the key input information based on the hiragana information;

[0011] Output Japanese text strings.

[0012] In one possible implementation, the key input information includes a direct conversion part and an indirect conversion part;

[0013] When converting key input information into hiragana information based on the hiragana conversion table, the indirect conversion part is converted using the simplified pinyin mapping method.

[0014] In one possible implementation, key input information is converted into hiragana information based on a hiragana conversion table, specifically including:

[0015] Based on the Hiragana conversion table, the Romanization of the key input information is converted, and the set of the first Hiragana after conversion is taken as the first Hiragana set, and the set of the unconverted Romanization is taken as the indirect conversion part.

[0016] The indirect conversion part is converted into the second hiragana by using the simplified pinyin mapping method. The set of the second hiragana is taken as the second hiragana set. The first hiragana and the second hiragana have position marks corresponding to the Roman syllable positions in the key input information.

[0017] The first hiragana characters in the first hiragana set and the second hiragana characters in the second hiragana set are arranged and merged according to the order of their position marks, and then used as hiragana information.

[0018] In one possible implementation, the Romanization of the key input information is converted according to the Hiragana conversion table. The set of converted first Hiragana is taken as the first Hiragana set, and the set of unconverted Romanization is taken as the indirect conversion part, specifically including:

[0019] Repeat the following steps until the updated index value equals the total length of the key input information:

[0020] Determine if the updated index value is equal to the total length;

[0021] If not, read the next Roman numeral, add the next Roman numeral to the part to be converted to update the part to be converted, and increment the cumulative length by one;

[0022] Determine if the part to be converted contains the first Romanization in the Hiragana conversion table;

[0023] If so, convert to the first Romanized pronunciation;

[0024] Determine whether the portion to be converted has been fully converted;

[0025] If not, update the first conversion result, clear the part to be converted, clear the accumulated length to zero, and update the index value;

[0026] If the updated index value is equal to the total length, then the part to be converted is placed in the position corresponding to the part to be converted in the first conversion result, and the set of hiragana in the first conversion result is taken as the first hiragana set, and the set of romanization in the first conversion result is taken as the indirect conversion part.

[0027] In one possible implementation, if the part to be converted does not contain the first Romanization in the Hiragana conversion table, then it is determined whether the length marker is negative;

[0028] If not, then determine whether the cumulative length is less than the difference between the total length and the index value;

[0029] If so, then check if the cumulative length is equal to 4;

[0030] If so, the first Roman numeral in the part to be converted is placed in the position corresponding to the first Roman numeral in the first conversion result, the first Roman numeral is removed from the part to be converted, and the cumulative length, part to be converted, and index value are updated.

[0031] In one possible implementation, if the part to be converted is completely converted, the first conversion result is updated, the part to be converted is cleared, the cumulative length is cleared to zero, and the index value is updated.

[0032] In one possible implementation, if the cumulative length is not equal to 4, then continue reading the next Roman numeral, update the part to be converted, and increment the cumulative length by one.

[0033] In one possible implementation, the indirect transformation part is transformed using a simplified mapping method, specifically including:

[0034] Each Romanization in the indirect conversion section is converted into all Hiragana characters that begin with a Romanization in the Hiragana conversion table.

[0035] In one possible implementation, the method of using abbreviated mapping to transform the indirect transformation part also includes:

[0036] If the Romanization combination formed by multiple adjacent Romanizations in the indirect conversion part is not a specified combination in the Hiragana conversion table, then the Romanization combination will be converted into the Hiragana corresponding to the specified combination.

[0037] In one possible implementation, the Japanese character string corresponding to the key input information is obtained based on the hiragana information, specifically including:

[0038] Hiragana information is divided into groups in various ways, and each group is transformed into a hiragana tree. Each division method corresponds to a tree group.

[0039] Each tree group is decoded using a decoder to obtain the decoding result corresponding to the tree group and the score of the decoding path corresponding to the decoding result, where each decoding result corresponds to a Japanese character string;

[0040] The lowest-scoring decoding result among all decoding results obtained from multiple tree groups is taken as the final decoding result.

[0041] Secondly, the present invention provides a Japanese full keyboard input device, including a receiving module, a hiragana conversion module, a character conversion module and an output module;

[0042] The receiving module is used to receive key input information in Japanese full keyboard input mode. The key input information includes one or more consecutive Romanization information corresponding to the key.

[0043] The Hiragana conversion module is used to convert key input information into Hiragana information based on the Hiragana conversion table;

[0044] The text conversion module is used to obtain the Japanese character string corresponding to the key input information based on the hiragana information;

[0045] The output module is used to output Japanese text strings.

[0046] In one possible implementation, the hiragana conversion module includes a first hiragana set acquisition module, a second hiragana set acquisition module, and a merging module;

[0047] The module for obtaining the first hiragana set is used to convert the romanization in the key input information according to the hiragana conversion table, and the set of converted first hiragana is used as the first hiragana set, while the set of unconverted romanization is used as the indirect conversion part.

[0048] The module for obtaining the second hiragana set is used to convert the indirect conversion part into the second hiragana using the simplified pinyin mapping method, and the set of the second hiragana is used as the second hiragana set. The first hiragana and the second hiragana have position marks corresponding to the Roman numeral positions in the key input information.

[0049] The merging module is used to arrange and merge the first hiragana in the first hiragana set and the second hiragana in the second hiragana set according to the order of the position markers, and then use them as hiragana information.

[0050] In one possible implementation, the text conversion module includes a tree group acquisition module, a decoding module, and a result acquisition module;

[0051] The tree group acquisition module is used to divide hiragana information into groups in various ways, and each group is converted into a hiragana tree. Each division method corresponds to a tree group.

[0052] The decoding module is used to decode each tree group using the decoder to obtain the decoding result corresponding to the tree group and the score of the decoding path corresponding to the decoding result, where each decoding result corresponds to a Japanese character string;

[0053] The result acquisition module is used to take one or more decoding results with the lowest scores among all decoding results obtained from multiple tree groups as the final decoding result.

[0054] Thirdly, the present invention provides an electronic device, comprising:

[0055] One or more processors, memory, and one or more computer programs, wherein the one or more computer programs are stored in memory, and the one or more computer programs include instructions that, when executed by an electronic device, cause the electronic device to perform the aforementioned Japanese full keyboard input method.

[0056] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to execute the aforementioned Japanese full keyboard input method.

[0057] The concept of this invention lies in converting all Romanized sounds input via keypad into corresponding Hiragana characters, decoding all the converted Hiragana characters as a whole to obtain accurate Japanese character strings, thus improving the match rate between the output and the user's expectations, and consequently improving input efficiency. Furthermore, this invention uses a simplified spelling mapping method to convert Romanized sounds that cannot be directly converted into Hiragana, thereby utilizing the full Japanese keyboard for simplified spelling and realizing a simplified Japanese spelling function. Attached Figure Description

[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings, wherein:

[0059] Figure 1 A flowchart of the Japanese full keyboard input method provided by the present invention;

[0060] Figure 2 The present invention provides a flowchart for converting key input information into hiragana information;

[0061] Figure 3 A flowchart for obtaining the first set of hiragana and the indirect conversion part provided by the present invention;

[0062] Figure 4 A flowchart for obtaining Japanese character strings provided by the present invention;

[0063] Figure 5 A schematic diagram of the hiragana information segmentation result in one embodiment of the present invention;

[0064] Figure 6 for Figure 5 The diagram shown illustrates the result of decoding the second segmentation method using a deterministic acyclic state converter in the embodiment illustrated.

[0065] Figure 7 This is a schematic diagram of the structure of the Japanese full keyboard input device provided by the present invention;

[0066] Figure 8 This is a schematic diagram of the structure of the hiragana conversion module provided by the present invention;

[0067] Figure 9 This is a schematic diagram of the structure of the text conversion module provided by the present invention;

[0068] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0069] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0070] The concept of this invention lies in converting all Romanized sounds input via keypad into corresponding Hiragana characters, decoding all the converted Hiragana characters as a whole to obtain accurate Japanese character strings, thus improving the match rate between the output and the user's expectations and improving input efficiency. Furthermore, this invention uses a simplified spelling mapping method to convert Romanized sounds that cannot be directly converted into Hiragana, thereby utilizing the full Japanese keyboard for simplified spelling and realizing a simplified Japanese spelling function.

[0071] In view of the aforementioned core concept, this invention provides at least one embodiment of a Japanese full keyboard input method, such as... Figure 1 As shown, it may include the following steps:

[0072] S110: In Japanese full keyboard input mode, it receives key input information, which includes one or more consecutive Romanized sounds corresponding to the keys.

[0073] S120: Convert key input information into hiragana information according to the hiragana conversion table.

[0074] S130: Obtain the Japanese character string corresponding to the key input information based on the hiragana information.

[0075] S140: Output Japanese text string.

[0076] Specifically, in step S110, in the Japanese full keyboard input mode, each letter key corresponds to a Roman numeral. During use, the user may intermittently press letter keys, symbol keys, newline keys, etc. It should be noted that the key input information in step S110 consists of one or more consecutive letter keys, each corresponding to one or more consecutive Roman numerals. Each Roman numeral in the key input information has a one-to-one corresponding position marker.

[0077] Understandably, if the key pressed by the user includes not only one or more consecutive letter keys, but also other non-letter keys (such as symbol keys, newline keys, etc.), the output result is the concatenation result of the non-letter output results corresponding to the non-letter keys arranged in the order of user input and the Japanese character strings corresponding to the letter keys (see step S140).

[0078] In step S120, all the Romanized information in the key input information is converted into Hiragana to obtain the Hiragana information corresponding to the Romanized information.

[0079] Understandably, considering that different users have different levels of mastery of the Japanese full keyboard input method, the key information entered by the user may not completely match the Romanization or Romanization combination in the Hiragana conversion table. Therefore, the Romanization information in step S110 includes Romanization parts that can be directly converted into Hiragana (i.e., direct conversion parts) and Romanization parts that cannot be directly converted into Hiragana (i.e., indirect conversion parts).

[0080] In one possible implementation, each Roman numeral in the indirect conversion section is converted into one of the hiragana characters in the hiragana conversion table that begins with that Roman numeral, and then merged with the hiragana characters in the direct conversion section in the order of key input information to form hiragana information.

[0081] In the above implementation, the hiragana converted from the indirect conversion part may be what the user expects or not, which reduces the matching rate between the output result and the user's expectations and reduces input efficiency.

[0082] Based on the above considerations, in a preferred embodiment, the indirect conversion part is converted into hiragana using the abbreviated mapping method.

[0083] like Figure 2 As shown, converting key input information into hiragana information involves the following steps:

[0084] S210: Determine the direct conversion part and the indirect conversion part in the key input information according to the hiragana conversion table, and convert the direct conversion part into hiragana as the first hiragana set.

[0085] In one possible implementation, if the key input information contains only one Roman syllable, then it is checked whether the Roman syllable is included in the hiragana conversion table. If so, it is converted into hiragana and used as the first hiragana set; otherwise, it is used as the indirect conversion part.

[0086] In one possible implementation, if the key input information contains multiple consecutive Roman syllables, the Roman syllable information is grouped, and a hiragana conversion table is consulted to determine whether each Roman syllable group can be directly converted. If so, the Roman syllable group is converted into hiragana; otherwise, the Roman syllable group is treated as an indirect conversion part.

[0087] In the above embodiments, there is a possibility that improper grouping may reduce the direct conversion efficiency. Therefore, in a preferred embodiment, such as... Figure 3 As shown, obtaining the first set of hiragana and the indirect conversion part specifically includes the following steps:

[0088] Initially, the index value is 0 and the cumulative length is 0.

[0089] S310: Calculate the total length of the key input information, determine whether the total length is less than 4, and if so, assign the length flag to no.

[0090] S320: Determine if the updated index value is equal to the total length. If yes, proceed to step S3130. Otherwise, proceed to step S330.

[0091] S330: Read the next Roman numeral, add the next Roman numeral to the part to be converted to update the part to be converted, and increment the cumulative length by one.

[0092] S340: Determine if the part to be converted contains the first Romanization in the Hiragana conversion table. If yes, proceed to step S350; otherwise, proceed to step S390.

[0093] S350: Convert to first Romanization.

[0094] S360: Determine whether the part to be converted has been completely converted. If yes, proceed to step S370; otherwise, proceed to step S380.

[0095] S370: If the part to be converted is completely converted, a fourth hiragana will be generated, and there are no unconverted romanizations. Then, the converted fourth hiragana will be placed in the first conversion result at the position corresponding to the completely converted romanizations to update the first conversion result. The part to be converted will be cleared, the cumulative length will be cleared to zero, the index value will be updated, and the process will return to step S320.

[0096] S380: Place the part to be converted, which contains the third hiragana to be converted, into the position corresponding to the part to be converted in the first conversion result to update the first conversion result, clear the part to be converted, clear the cumulative length to zero, update the index value, and return to step S320.

[0097] S390: If the first Roman numeral does not exist, determine whether the length marker is negative (i.e., whether the total length is less than 4). If yes, proceed to step S3110; otherwise, proceed to step S3100.

[0098] S3100: Determine whether the cumulative length is less than the difference between the total length and the index value. If yes, proceed to step S3110; otherwise, proceed to step S3120.

[0099] S3110: Determine if the cumulative length is equal to 4. If yes, proceed to step S3120; otherwise, return to step S320.

[0100] S3120: Place the first Roman numeral in the part to be converted into the position corresponding to the first Roman numeral in the first conversion result, remove the first Roman numeral from the part to be converted, update the cumulative length, part to be converted, and index value, and return to step S320.

[0101] S3130: Place the part to be converted into the position corresponding to the part to be converted in the first conversion result, and take the set of hiragana in the first conversion result as the first hiragana set, and take the set of romanization in the first conversion result as the indirect conversion part.

[0102] Taking the key input information as "hdrkary" as an example, the process of obtaining the first set of hiragana and the indirect conversion part is as follows:

[0103] P1: The total length of the key input information is 7. If the total length is less than 4 (the maximum number of Japanese hiragana characters), then the length flag is set to no. The initial index value is 0, which is not equal to the total length of 7, so P2 is executed.

[0104] P2: Read the first Roman numeral "h", update the part to be converted to "h", and increment the cumulative length by one to 1. No hiragana corresponding to the Roman numeral "h" is found in the hiragana conversion table. At this time, the length is marked as no. The cumulative length (1) is less than the difference between the total length (7) and the index value (0), and the cumulative length is not equal to 4. Then continue to judge. If the index value (0) is not equal to the total length (7), then execute step P3.

[0105] P3: Read the second Roman numeral "d", update the part to be converted to "hd", and increment the cumulative length by one to 2. No hiragana corresponding to the Roman numeral "hd" is found in the hiragana conversion table. At this time, the length is marked as no. The cumulative length (2) is less than the difference between the total length (7) and the index value (0), and the cumulative length is not equal to 4. Then continue to judge. If the index value (0) is not equal to the total length (7), then execute step P4.

[0106] P4: Read the third Roman numeral "r", update the part to be converted to "hdr", and increment the cumulative length by one to 3. No hiragana corresponding to the Roman numeral "hdr" is found in the hiragana conversion table. At this time, the length is marked as no. The cumulative length (3) is less than the difference between the total length (7) and the index value (0), and the cumulative length is not equal to 4. Then continue to judge. If the index value (0) is not equal to the total length (7), then execute step P5.

[0107] P5: Read the fourth Roman syllable "k", update the part to be converted to "hdrk", and increment the cumulative length by one to 4. No hiragana corresponding to the Roman syllable "hdrk" is found in the hiragana conversion table. At this time, the length is marked as no. The cumulative length (4) is less than the difference between the total length (7) and the index value (0), and the cumulative length is equal to 4. Then, put the first Roman syllable "h" in the part to be converted into the position corresponding to "h" in the first conversion result (the position is marked as 1) and remove "h" from the part to be converted. The part to be converted is updated to "drk", the cumulative length is updated to 3, the index value is updated to 4, and the judgment continues. If the index value (4) is not equal to the total length (7), then step P6 is executed.

[0108] P6: Read the fifth Romanization "a", update the part to be converted to "drka", and increment the cumulative length by one to 4. Find the hiragana "か" corresponding to the Romanization "ka" in the hiragana conversion table, then convert "ka" to hiragana "か", update the part to be converted to "drか". At this time, the part to be converted has not been completely converted, so put the part to be converted "drか" into the position corresponding to "drka" in the first conversion result (position marked as 2-5), clear the part to be converted, clear the cumulative length to zero, and update the index value to 5 according to the position mark (5) of the Romanization "a". Continue to judge, if the index value (5) is not equal to the total length (7), then execute step P7.

[0109] P7: Read the sixth Roman numeral "r", update the part to be converted to "r", and increment the cumulative length by one to 1. No hiragana corresponding to the Roman numeral "r" is found in the hiragana conversion table. At this time, the length is marked as no. The cumulative length (1) is less than the difference between the total length (7) and the index value (5), and the cumulative length is not equal to 4. Then continue to judge. If the index value (5) is not equal to the total length (7), then execute step P8.

[0110] P8: Read the seventh Roman syllable "y", update the part to be converted to "ry", and increment the cumulative length by one to 2. No hiragana corresponding to the Roman syllable "ry" is found in the hiragana conversion table. At this time, the length is marked as no. The cumulative length (2) equals the difference between the total length (7) and the index value (5). Then, the first Roman syllable "r" in the part to be converted is placed in the position corresponding to "r" in the first conversion result (position marked as 6), and "r" is removed from the part to be converted. The part to be converted is updated to "y", the cumulative length is updated to 1, and the index value is updated to 7 based on the position of the Roman syllable "y" (position marked as 7). Continue to judge; if the index value (7) equals the total length (7), then execute step P9.

[0111] P9: Place the part to be converted, “y,” into the position corresponding to “y” in the first conversion result (position marked as 7). The first conversion result is then “hdrかry.” Take the set of hiragana “か” in the first conversion result as the first hiragana set, and take the set of romanization “hdrry” in the first conversion result as the indirect conversion part.

[0112] It should be noted that both the first set of hiragana and the indirect conversion section have position markers corresponding to the positions of the Roman syllables in the key input information. The order of the hiragana characters in the first set of hiragana matches the order of their corresponding Roman syllables in the key input information, and the order of the Roman syllables in the indirect conversion section also matches the order of their corresponding Roman syllables in the key input information. During subsequent processing, Roman syllables that are not adjacent in the key input information cannot be converted into Roman syllable combinations in the indirect conversion section.

[0113] S220: The indirect conversion part is converted into hiragana using the simplified mapping method, which serves as the second hiragana set.

[0114] It should be noted that in the simplified mapping method, each Roman syllable is considered as a simplified spelling of a single Roman syllable or a combination of Roman syllables containing that syllable in the hiragana conversion table. Therefore, when converting to hiragana, each Roman syllable in the indirect conversion part may correspond to one or more Roman syllables or combinations of Roman syllables in the hiragana conversion table.

[0115] Specifically, when converting the indirect conversion part into hiragana, each romanization in the indirect conversion part is converted into one, several, or all hiragana characters that begin with a romanization in the hiragana conversion table.

[0116] Preferably, each Romanization is converted into all hiragana characters that begin with a Romanization.

[0117] For example, in the above example, the indirect conversion part is "hdrry". First, the Roman numeral "h" is mapped to all hiragana characters that begin with "h". Then, the Roman numeral "d" is mapped to all hiragana characters that begin with "d". Next, the Roman numeral "r" is mapped to all hiragana characters that begin with "r". Finally, the second "r" is mapped to all hiragana characters that begin with "r". This achieves the effect of inputting one Roman numeral letter being equivalent to inputting five hiragana characters, as illustrated below:

[0118]

[0119] In particular, if there are multiple Romanizations that are adjacent in position (i.e., adjacent in position markers) in the indirect conversion part, it is determined whether the Romanization combination formed by these multiple Romanizations is a specified combination in the Hiragana conversion table.

[0120] In the example above, there are two Roman syllables, “r” and “y”, which are adjacent in position (corresponding to position marks 6 and 7). The Roman syllable combination “ry” formed by them is the abbreviation of “rya”, “ryi”, “rye”, “ryu”, and “ryo”. In this case, the Roman syllable combination is the specified combination.

[0121] If the Romanization combination is a specified combination, it will be converted into the corresponding hiragana. Specifically, the Romanization that plays a decisive role in the combination will be replaced with its corresponding specified Romanization. When converting to hiragana, the replaced Romanization combination will correspond to a specific hiragana in the hiragana conversion table.

[0122] Specifically, in the example above, replacing the Romanized "y" with "Y" results in the Romanized combination "rY" corresponding to the hiragana characters "rya", "ryi", "rye", "ryu", and "ryo" in the hiragana conversion table. Therefore, the Romanized combination "ry" is now converted to the hiragana characters "rya", "ryi", "rye", "ryu", and "ryo". See the illustration below:

[0123]

[0124] If the Romanization combination is not a specified combination, then each Romanization in the combination is converted into a hiragana separately.

[0125] Therefore, in the example above, using the simplified mapping method, the number of second hiragana characters after indirect conversion of "hdrry" is 5*5*5*5=625, that is, the set of second hiragana characters contains 625 second hiragana characters.

[0126] S230: Based on the order of the Roman numerals in the key input information, arrange and merge the first hiragana in the first hiragana set and the second hiragana in the second hiragana set to obtain hiragana information.

[0127] Specifically, the hiragana characters in the first and second hiragana sets are arranged according to the order of their position markers, and all the arranged hiragana characters are merged to form hiragana information.

[0128] As can be seen from step S220, the second set of hiragana obtained by the simplified mapping method contains multiple different second hiragana. Therefore, if there is an indirect conversion part in the key input information, the obtained hiragana information includes multiple merged hiragana.

[0129] In the example above, the hiragana information includes 625 merged hiragana characters.

[0130] In step S130, since there is a correspondence between hiragana and Japanese characters, the Japanese character string corresponding to the key input information can be obtained by converting the merged hiragana in the hiragana information into a Japanese character string.

[0131] As one possible implementation, the hiragana information is divided into groups in various ways, each group is decoded into a Japanese character string using a decoder, and one or more decoding results with the highest probability are selected from all the Japanese character strings as the output.

[0132] Specifically, different segmentation methods result in different numbers of groups; it may result in one group or multiple groups.

[0133] In a preferred embodiment, such as Figure 4 As shown, obtaining the Japanese character string corresponding to the key input information includes the following steps:

[0134] S410: Divide each hiragana information into groups according to multiple methods, and transform each group into a hiragana tree. Each division method corresponds to a tree group, which is composed of all hiragana trees corresponding to the groups obtained by that division method.

[0135] Understandably, a hiragana tree corresponding to each group is obtained by matching hiragana information with a hiragana tree database. In obtaining the hiragana tree database, all hiragana pronunciations in Japanese are first collected, and then these pronunciations are stored using a tree structure to form a hiragana tree, thus obtaining the hiragana tree database.

[0136] Taking the key input message "hidari" as an example, its converted hiragana message is "ひだり". As an example, there are three ways to segment it, such as... Figure 5As shown, the root node number (such as 3537, 2244) represents the sequence number of the hiragana tree in the hiragana database. The first and third splitting methods consist of two hiragana trees, while the second splitting method consists of one hiragana tree.

[0137] S420: Decode each tree group using a decoder to obtain the decoding result corresponding to the tree group and the score of the decoding path corresponding to the decoding result, where each decoding result corresponds to a Japanese character string.

[0138] Specifically, a language model consisting of data structures such as Finite-state Transducer (FST) and Weighted Finite-state Transducer (WFST) is used to decode hiragana.

[0139] In one possible implementation, the decoding is performed using a Deterministic Acyclic Finite State Transducer data structure, which has the following characteristics:

[0140] 1. Determined: For any given state, there can only be at most one transition that can be traversed.

[0141] 2. Acyclic: It is impossible to traverse the same state repeatedly.

[0142] 3. Transition: Accepts a specific sequence, terminates in the final state, and outputs a value.

[0143] Figure 6 for Figure 5 The illustrated embodiment is a schematic diagram of the result of decoding the second segmentation method using a deterministic acyclic state converter.

[0144] like Figure 6 As shown, the language model has a total of 777,939 entries and 222,061 entry pairs. Node 0 is the initial node, and node 1 is the sentence beginning symbol. <s>Terminating nodes: Node 2 is the starting node for all terms, Node 3 is the ending node for the term "left", and Node 4 is the ending node for the term pair. <s>The terminator for "left" is node 5, which is the middle node of the term "left and right", and node 6 is the terminator for the pair "left and right". <s>The termination node of the entry "right" is node 7, and the termination node of the entry "left and right" is node 8. The line segment between two nodes represents the transfer path, and the value corresponding to the key is recorded on the transfer path. The forward transfer between nodes is triggered from node 2, and the reverse transfer returns to node 2 to achieve sustainable decoding.

[0145] The value is the penalty score. Taking the entry "-5.657657 around -0.495544" as an example, (the score in the first column)*(-10*ln10) represents the penalty score for the appearance of the entry, and (the score in the third column)*(-10*ln10) represents the penalty score for the entry to return to node 2 to form a sentence. The penalty score for the appearance of the entry "left and right" is rounded to 130, and the penalty score for returning to node 2 is 11. Corresponding Figure 6 , the forward path corresponding to the entry "left and right" is 0->1->2->5->8, and its value is "130", that is, the penalty score for the forward transfer from node 2 to node 5 "left" is 105, and the penalty score for the forward transfer from node 5 to node 8 "right" is 25. In the reverse path corresponding to the entry "left and right", the penalty score for the reverse transfer from node 8 to node 2 is 11.

[0146] Figure 6 shows Figure 5 the decoding path of the second segmentation method "hidari" in. According to the hiragana tree database, the Japanese character string corresponding to the hiragana "hidari" is "left". Figure 6 In the decoding path of, two paths 0->1->4 and 0->1->2->3 can be taken respectively, and the path with the lowest penalty score, that is, 0->1->4, is taken as the decoding result of the second segmentation method, and the penalty score of its corresponding decoding path is 85.

[0147] S430: Take one or more decoding results with the lowest scores among all the decoding results obtained from multiple tree groups as the final decoding results.

[0148] In the above implementation manner using the acyclic state transducer, the decoding result corresponding to the decoding path with the lowest penalty score is taken as the final decoding result.

[0149] In S140, all the decoding results (that is, Japanese character strings) in the final decoding results are output in ascending order of scores.

[0150] Corresponding to the above embodiments and preferred solutions, the present invention also provides an embodiment of a Japanese full keyboard input device, as Figure 7 shown, which may specifically include a receiving module 710, a hiragana conversion module 720, a character conversion module 730, and an output module 740.

[0151] The receiving module 710 is used to receive key input information in Japanese full keyboard input mode. The key input information includes one or more consecutive Romanization information corresponding to the key.

[0152] The Hiragana conversion module 720 is used to convert key input information into Hiragana information based on the Hiragana conversion table.

[0153] The text conversion module 730 is used to obtain the Japanese character string corresponding to the key input information based on the hiragana information.

[0154] Output module 740 is used to output Japanese text strings.

[0155] In one possible implementation, such as Figure 8 As shown, the hiragana conversion module 720 includes a first hiragana set acquisition module 7201, a second hiragana set acquisition module 7202, and a merging module 7203.

[0156] The first hiragana set acquisition module 7201 is used to convert the romanization in the key input information according to the hiragana conversion table, and take the converted first hiragana set as the first hiragana set, and take the unconverted romanization set as the indirect conversion part.

[0157] The second hiragana set acquisition module 7202 is used to convert the indirect conversion part into second hiragana using the simplified pinyin mapping method, and to take the set of second hiragana as the second hiragana set, wherein the first hiragana and the second hiragana have position marks corresponding to the Roman phonological positions in the key input information.

[0158] The merging module 7203 is used to arrange and merge the first hiragana in the first hiragana set and the second hiragana in the second hiragana set according to the order of the position markers, and then use them as hiragana information.

[0159] In one possible implementation, such as Figure 9 As shown, the text conversion module 730 includes a tree group acquisition module 7301, a decoding module 7302, and a result acquisition module 7303.

[0160] The tree group acquisition module 7301 is used to divide the hiragana information into groups in various ways, and each group is converted into a hiragana tree. Each division method corresponds to a tree group.

[0161] The decoding module 7302 is used to decode each tree group using the decoder to obtain the decoding result corresponding to the tree group and the score of the decoding path corresponding to the decoding result, wherein each decoding result corresponds to a Japanese character string.

[0162] The result acquisition module 7303 is used to take one or more decoding results with the lowest scores among all decoding results obtained from multiple tree groups as the final decoding result.

[0163] The above should be understood Figure 7 The division of components in the illustrated Japanese full keyboard input device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. These components can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some components can be implemented in software via processing element calls, while others are implemented in hardware. For example, a particular module can be a separate processing element or integrated into a chip within the electronic device. The implementation of other components is similar. Furthermore, these components can be fully or partially integrated together, or implemented independently. During implementation, each step of the above method or each of the above components can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0164] For example, these components can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, these components can be integrated together to form a System-On-a-Chip (SOC).

[0165] Based on the above embodiments and preferred solutions, those skilled in the art will understand that, in practice, the present invention is applicable to various implementation methods. The present invention is illustrated by the following carrier:

[0166] (1) An electronic device, which may include:

[0167] One or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by an electronic device, cause the electronic device to perform the steps / functions of the foregoing embodiments or equivalent implementations.

[0168] Figure 10 This is a schematic diagram illustrating the structure of an embodiment of the electronic device of the present invention. The device can be an electronic device or a circuit device built into the aforementioned electronic device. The electronic device can be a PC, server, smart terminal (mobile phone, tablet, watch, glasses, etc.), smart TV, remote control, smart screen, ATM, robot, drone, ICV, smart (car) vehicle, and in-vehicle equipment, etc. This embodiment does not limit the specific form of the electronic device.

[0169] Specifically, such as Figure 10 As shown, the electronic device 900 includes an input unit 960, a processor 910, and a memory 930. The processor 910 and the memory 930 can communicate with each other via an internal connection path to transmit control and / or data signals. The memory 930 stores computer programs, and the processor 910 retrieves and runs the computer programs from the memory 930. The processor 910 and the memory 930 can be combined into a single processing device, but more commonly they are independent components. The processor 910 executes the program code stored in the memory 930 to achieve the aforementioned functions. In specific implementations, the memory 930 can be integrated into the processor 910, or it can be independent of the processor 910. The input unit 960 and the processor 910 can communicate with each other via an internal connection path to transmit control and / or data signals.

[0170] In addition, to further enhance the functionality of the electronic device 900, the device 900 may also include one or more of the following: a display unit 970, an audio circuit 980, a camera 990, and a sensor 901. The audio circuit may also include a speaker 982, a microphone 984, etc. The display unit 970 may include a display screen.

[0171] Furthermore, the aforementioned electronic device 900 may also include a power supply 950 for providing electrical energy to various devices or circuits in the device 900.

[0172] It should be understood that Figure 10 The electronic device 900 shown is capable of implementing the various processes of the method provided in the foregoing embodiments. The operation and / or function of each component in the device 900 can respectively implement the corresponding processes in the above method embodiments. For details, please refer to the foregoing descriptions of the embodiments of methods, devices, etc.; to avoid repetition, detailed descriptions are appropriately omitted here.

[0173] It should be understood that Figure 10 The processor 910 in the illustrated electronic device 900 can be a system-on-a-chip (SoC). The processor 910 may include a central processing unit (CPU) and may further include other types of processors, such as a graphics processing unit (GPU), which will be described in detail below.

[0174] In summary, the various processors or processing units inside the processor 910 can work together to implement the previous method flow, and the corresponding software programs of each processor or processing unit can be stored in the memory 930.

[0175] (2) A computer-readable storage medium storing a computer program or the aforementioned apparatus on the storage medium, which, when executed, causes a computer to perform the steps / functions of the foregoing embodiments or equivalent embodiments.

[0176] In several embodiments provided by this invention, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, certain technical solutions of this invention, or the parts that contribute to the prior art, or parts of such technical solutions, can be embodied in the form of the following software products.

[0177] (3) A computer program product (which may include the above-described device), which, when run on a terminal device, causes the terminal device to execute the Japanese full keyboard input method of the foregoing embodiments or equivalent embodiments.

[0178] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the above implementation methods can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the above-mentioned computer program products may include, but are not limited to, APP; continuing from the foregoing, the above-mentioned device / terminal may be a computer device (e.g., mobile phone, PC terminal, cloud platform, server, server cluster, or network communication device such as media gateway, etc.). Furthermore, the hardware structure of the computer device may specifically include: at least one processor, at least one communication interface, at least one memory, and at least one communication bus; the processor, communication interface, and memory can all communicate with each other through the communication bus. The processor may be a central processing unit (CPU), DSP, microcontroller, or digital signal processor, and may also include a GPU, an embedded neural network processing unit (NPU), and an image signal processor (ISP). The processor may also include a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. Furthermore, the processor may have the function of operating one or more software programs, which may be stored in a storage medium such as a memory. The aforementioned memory / storage medium may include: non-volatile memory, such as a non-removable disk, USB flash drive, portable hard drive, optical disc, etc., as well as read-only memory (ROM), random access memory (RAM), etc.

[0179] In this embodiment of the invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0180] Those skilled in the art will recognize that the modules, units, and method steps described in the embodiments disclosed in this specification can be implemented using electronic hardware, computer software, and a combination of electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0181] Furthermore, the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. In particular, for embodiments such as apparatus and devices, since they are basically similar to the method embodiments, the relevant parts can be referred to the description of the method embodiments. The apparatus, devices, and other embodiments described above are merely illustrative, and the modules, units, etc., described as separate components may or may not be physically separate, that is, they may be located in one place or distributed in multiple places, such as nodes in a system network. Specifically, some or all of the modules and units can be selected according to actual needs to achieve the purpose of the above-described embodiment solutions. Those skilled in the art can understand and implement this without creative effort.

[0182] The above description of the structure, features, and effects of the present invention is based on the embodiments shown in the figures. However, the above are only preferred embodiments of the present invention. It should be noted that the technical features involved in the above embodiments and their preferred methods can be reasonably combined and matched by those skilled in the art to form a variety of equivalent solutions without departing from or changing the design concept and technical effects of the present invention. Therefore, the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.< / s> < / s> < / s>

Claims

1. A Japanese full keyboard input method, characterized in that, include: In Japanese full keyboard input mode, key input information is received, which includes one or more consecutive Romanization information corresponding to the key; the key input information includes a direct conversion part and an indirect conversion part; wherein, the direct conversion part is the Romanization part of the Romanization information that can be directly converted into Hiragana, and the indirect conversion part is the Romanization part of the Romanization information that cannot be directly converted into Hiragana; The key input information is converted into hiragana information according to the hiragana conversion table, including the conversion of the indirect conversion part by using the abbreviated mapping method. The abbreviated mapping method includes: treating each roman sound as a simplified spelling of a single roman sound or a combination of roman sounds containing that roman sound in the hiragana conversion table, and each roman sound in the indirect conversion part corresponds to one or more roman sounds or combinations of roman sounds in the hiragana conversion table. The conversion of the indirect conversion portion includes: converting each Roman numeral in the indirect conversion portion into all hiragana characters in the hiragana conversion table that begin with the Roman numeral; Based on the hiragana information, obtain the Japanese character string corresponding to the key input information; Output the Japanese text string.

2. The Japanese full keyboard input method according to claim 1, characterized in that, The process of converting the key input information into hiragana information based on the hiragana conversion table specifically includes: The Romanization of the key input information is converted according to the Hiragana conversion table. The set of converted first Hiragana is taken as the first Hiragana set, and the set of unconverted Romanization is taken as the indirect conversion part. The indirect conversion part is converted into a second hiragana using a simplified mapping method, and the set of the second hiragana is taken as the second hiragana set. The first hiragana and the second hiragana have position marks corresponding to the Roman numeral positions in the key input information. Based on the order of the Roman numerals in the key input information, the first hiragana in the first hiragana set and the second hiragana in the second hiragana set are arranged and merged to form the hiragana information.

3. The Japanese full keyboard input method according to claim 2, characterized in that, Based on the aforementioned hiragana conversion table, the Romanization of the key input information is converted. The set of converted first hiragana is taken as the first hiragana set, and the set of unconverted Romanization is taken as the indirect conversion part. Specifically, this includes: Repeat the following steps until the updated index value equals the total length of the key input information: Determine whether the updated index value is equal to the total length; If not, then read the next Roman numeral, add the next Roman numeral to the part to be converted to update the part to be converted, and increment the cumulative length by one; Determine whether the part to be converted contains the first Romanization in the Hiragana conversion table; If so, then convert the first Romanized pronunciation; Determine whether the portion to be converted has been completely converted; If not, update the first conversion result, clear the part to be converted, clear the cumulative length to zero, and update the index value; wherein updating the first conversion result includes: placing the part to be converted containing the hiragana to be converted into the position corresponding to the part to be converted in the first conversion result; If the updated index value is equal to the total length, then the part to be converted is placed in the position corresponding to the part to be converted in the first conversion result, and the set of hiragana in the first conversion result is taken as the first hiragana set, and the set of romanization in the first conversion result is taken as the indirect conversion part.

4. The Japanese full keyboard input method according to claim 3, characterized in that, If the part to be converted does not contain the first Romanization in the Hiragana conversion table, then determine whether the length marker is negative. If the total length of the key input information is less than 4, then the length marker is assigned a negative value. If not, determine whether the cumulative length is less than the difference between the total length and the index value; If so, then determine whether the cumulative length is equal to 4; If so, the first Roman numeral in the part to be converted is placed in the position corresponding to the first Roman numeral in the first conversion result, the first Roman numeral is removed from the part to be converted, and the cumulative length, the part to be converted, and the index value are updated.

5. The Japanese full keyboard input method according to claim 3, characterized in that, If the part to be converted is completely converted, then the first conversion result is updated, the part to be converted is cleared, the cumulative length is cleared to zero, and the index value is updated; wherein, updating the first conversion result includes: if the part to be converted is completely converted, then the converted hiragana is placed in the first conversion result at the position corresponding to the completely converted romanization.

6. The Japanese full keyboard input method according to claim 4, characterized in that, If the cumulative length is not equal to 4, then continue reading the next Roman numeral, update the part to be converted, and increment the cumulative length by one.

7. The Japanese full keyboard input method according to claim 1, characterized in that, The method of using the abbreviated mapping to transform the indirect transformation part further includes: If the Romanization combination formed by multiple adjacent Romanizations in the indirect conversion part is a specified combination in the Hiragana conversion table, then the Romanization combination is converted into the Hiragana corresponding to the specified combination.

8. The Japanese full keyboard input method according to claim 1, characterized in that, Based on the hiragana information, the Japanese character string corresponding to the key input information is obtained, specifically including: The hiragana information is divided into groups in various ways, and each group is transformed into a hiragana tree. Each division method corresponds to a tree group. The hiragana tree refers to the hiragana tree that stores all the hiragana pronunciations of Japanese using a tree structure. Each tree group is decoded using a decoder to obtain the decoding result corresponding to the tree group and the score of the decoding path corresponding to the decoding result, wherein each decoding result corresponds to a Japanese character string; wherein the decoder is a language model for decoding hiragana. The lowest-scoring decoding result among all decoding results obtained from multiple tree groups is taken as the final decoding result.

9. A Japanese full keyboard input device, characterized in that, It includes a receiving module, a hiragana conversion module, a text conversion module, and an output module; The receiving module is used to receive key input information in Japanese full keyboard input mode. The key input information includes one or more consecutive Romanization information corresponding to the key. The key input information includes a direct conversion part and an indirect conversion part. The direct conversion part is the Romanization part of the Romanization information that can be directly converted into Hiragana, and the indirect conversion part is the Romanization part of the Romanization information that cannot be directly converted into Hiragana. The hiragana conversion module is used to convert the key input information into hiragana information according to the hiragana conversion table. This includes converting the indirect conversion part using a simplified spelling mapping method. The simplified spelling mapping method includes treating each Roman syllable as a simplified spelling of a single Roman syllable or a combination of Roman syllables containing that Roman syllable in the hiragana conversion table, and each Roman syllable in the indirect conversion part corresponds to one or more Roman syllables or combinations of Roman syllables in the hiragana conversion table. Converting the indirect conversion part includes converting each Roman syllable in the indirect conversion part into all hiragana characters in the hiragana conversion table that begin with that Roman syllable. The text conversion module is used to obtain the Japanese character string corresponding to the key input information based on the hiragana information. The output module is used to output the Japanese text string.

10. An electronic device, characterized in that, include: One or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to perform the Japanese full keyboard input method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to perform the Japanese full keyboard input method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Japanese input method and Japanese input device of touch screen terminal

    CN103164038A