Intelligent input method and device and storage medium

By identifying and generating target character groups, combining semantic probability values and statement derivation models, the problem that existing input methods cannot meet user input habits is solved, and input efficiency and user experience are improved.

CN120295489AActive Publication Date: 2025-07-11SHENZHEN MINIPLAY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510773284.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing input method cannot meet users with distinctive input habits, resulting in users requiring additional operations during the input process to delete automatically recommended words, affecting input efficiency.

Method used

By identifying the string input by the user, a target character group is generated, and a semantic probability value and statement derivation model is used to generate initial statements that fit the user's input habits for users to choose.

Benefits of technology

It improves user input efficiency and enhances user experience, allowing users to generate input text that conforms to input habits through a small number of input operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295489A_ABST
    Figure CN120295489A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent input method and device and a storage medium, and the method comprises the steps: recognizing an obtained character string, obtaining at least one target character group, and obtaining a group serial number corresponding to each target character group; according to at least one semantic meaning corresponding to the group serial number and a reference text, calculating a plurality of semantic probability values of the character string, and taking the semantic meaning represented by the highest semantic probability value as the target semantic meaning of the character string, the reference text comprising an input record and / or a chat record; inputting the target semantics into a statement derivation model, wherein the statement derivation model generates at least one initial statement and an applicable probability value of each initial statement according to the target semantics; and sorting the initial statements according to the applicable probability values, and displaying the initial statements ranking the top K as target statements, K being a positive integer. According to the intelligent input method described by the technical scheme of the invention, the input content fitting the input habit of the user is generated while the input operation of the user is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of input methods, and particularly relates to intelligent input methods, devices, and storage media. Background Art

[0002] With the development of science and technology, an input method can perform fuzzy pronunciation or automatic correction based on the string input by a user to display the content that the user wants to input. At the same time, due to the increasing intelligent needs of users, the current input method can further generate words that match a part of the content input by the user for the user automatically based on a fixed word combination form.

[0003] However, since the current input method is only limited to correcting strings and can only recommend words for the user automatically based on a fixed word combination form, it cannot meet the needs of users with special input habits, making it impossible for users to quickly input the content they themselves want to input using the words recommended by the input method. Instead, they need to perform additional operations to delete the automatically recommended part, which hinders the input efficiency of the users.

[0004] Therefore, how to generate input content that conforms to the user's input habits while facilitating the user's input operation is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0005] The purpose of this application is to generate input content that conforms to the user's input habits while facilitating the user's input operation.

[0006] Other features and advantages of this application will become apparent through the following detailed description, or will be learned in part through the practice of this application.

[0007] According to one aspect of the embodiments of this application, an intelligent input method is provided. The method includes: Identifying the obtained string to obtain at least one target character group, and obtaining the group serial number corresponding to each target character group; Calculating multiple semantic probability values of the string according to at least one semantic corresponding to the group serial number and a reference text, and taking the semantic represented by the highest semantic probability value as the target semantic of the string. The reference text includes input records and / or chat records; Inputting the target semantic into a sentence derivation model, and the sentence derivation model generates at least one initial sentence and the applicable probability value of each initial sentence according to the target semantic; Sorting the initial sentences according to the applicable probability value, and taking the top K initial sentences as target sentences for display, where K is a positive integer.

[0008] According to one aspect of the embodiments of the present application, identifying the obtained string to obtain at least one target character group, including: In the string, obtaining a string with a target length in the order of arrangement of each character as the first character group. If the first character group can be successfully matched with the dictionary, then use the first character group as the target character group; If the first character group cannot be successfully matched with the dictionary, then gradually reduce the value of the target length to re-obtain the first character group until the first character group is successfully matched or the first character group is a single character, and use the first character group as the target character group; Continue to identify the characters that have not been divided until all characters belong to the target character group.

[0009] According to one aspect of the embodiments of the present application, identifying the obtained string to obtain at least one target character group, including: If the length of the string is greater than a preset length threshold, then divide the string into multiple substrings; In any one of the substrings, obtaining a string with a target length in the order of character arrangement as the second character group. If the second character group can be successfully matched with the dictionary, then use the second character group as the target character group; If the second character group cannot be successfully matched with the dictionary, then gradually reduce the value of the target length to re-obtain the second character group until it is successfully matched with the dictionary or the second character group is a single character, and use the second character group as the target character group; Continue to divide the characters that have not been divided in the substring until all characters belong to the target character group.

[0010] According to one aspect of the embodiments of the present application, after continuing to divide the characters that have not been divided in the substring until all characters belong to the target character group, the method further includes: For two target character groups that belong to different substrings and are adjacent, if there is a target character group that is a single character, then dissolve the two target character groups that belong to different substrings and are adjacent to obtain a compensation string; Re-obtain the target character group for each of the compensation strings.

[0011] According to one aspect of the embodiments of the present application, calculating multiple semantic probability values of the string according to at least one semantic corresponding to the group number and the reference text, including: According to at least one semantic corresponding to each group number, obtaining multiple semantic queues corresponding to the string, and the semantics of each semantic queue are different; Calculate the semantic probability values of the semantic queues according to the semantic queues and the reference text respectively.

[0012] According to one aspect of the embodiments of the present application, calculating the semantic probability values of the semantic queues according to the semantic queues and the reference text respectively includes: Calculate the sub-probability value of the current semantics according to the number of occurrences of the pre-semantics in the reference text; if there is no pre-semantics for the current semantics, use the probability of the current semantics in the reference text as the sub-probability value of the current semantics. Calculate the semantic probability values of the semantic queues according to the sub-probability values of the semantics.

[0013] According to one aspect of the embodiments of the present application, input the target semantics into a sentence derivation model, and the sentence derivation model generates at least one initial sentence and the applicable probability value of each initial sentence, including: The sentence derivation model generates a plurality of initial sentences according to the target semantics. Match each initial sentence with a control text to obtain the applicable probability values of the plurality of initial sentences.

[0014] According to one aspect of the embodiments of the present application, the sentence derivation model is a bert model; wherein, The bert model includes an input layer, a transformation layer, and an output layer. The input layer is used to vectorize the input target semantics to obtain a target vector group. The transformation layer is used to analyze the relationship between each vector in the target vector group from a set number of analysis perspectives. The output layer is used to generate a target sentence according to the analysis result of the target vector group. For the obtained open-source bert model, convert the M-bit floating-point numbers of the model weights and activation values in each layer of the bert model into N-bit integers, where M and N are positive integers and M is greater than N, and delete a plurality of transformation layers to obtain an initial sentence derivation model. Train the initial sentence derivation model to obtain a sentence derivation model.

[0015] According to one aspect of the embodiments of the present application, the present application provides an intelligent input device, including a memory, a processor, and a readable program stored on the memory. The processor executes the readable program to implement the method as described in any one of the above.

[0016] According to one aspect of the embodiments of the present application, the present application provides a readable storage medium, on which a readable program / instructions are stored. When the readable program / instructions are executed by a processor, the method as described in any one of the above is implemented.

[0017] In this application, first, the obtained string is recognized to obtain at least one target character group, and the group serial number corresponding to each target character group is obtained. Then, according to at least one semantics corresponding to the group serial number and the reference text, multiple semantic probability values of the string are calculated, and the semantics represented by the highest semantic probability value is used as the target semantics of the string. The reference text includes input records and / or chat records. Then, the target semantics is input into a statement derivation model. The statement derivation model generates at least one initial statement and the applicable probability value of the initial statement according to the target semantics, sorts the initial statements according to the applicable probability value, and displays the top K initial statements as target statements. This application accurately recognizes the input string of the user to obtain the target character group, and obtains the target semantics of the string based on the target character group and the reference text, so that the target semantics is more in line with the user's input habits. Finally, the recognized target semantics outputs target statements through the statement derivation model for the user to select. That is, when using the input method, the user only needs to input a small number of characters to generate at least one selectable target statement for the user to select. And because the target semantics conforms to the user's input habits, the obtained target semantics also conforms to the user's input habits, thereby enabling the user to generate input text that conforms to the user's input habits with very few input operations, greatly improving the user's input efficiency and enhancing the user's usage experience.

[0018] Other features and advantages of the present application will become apparent from the following detailed description, or will be learned in part from the practice of the present application.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 Shows a flowchart of an intelligent input method according to an embodiment of the present application.

[0022] Figure 2 Shows a flowchart of recognizing the obtained string to obtain at least one target character group according to an embodiment of the present application.

[0023] Figure 3The flowchart shows the recognition of the acquired string according to another embodiment of the present application to obtain at least one target character group.

[0024] Figure 4 The flowchart shows the correction of the target character group according to an embodiment of the present application.

[0025] Figure 5 The flowchart shows the calculation of multiple semantic probability values of the string according to at least one semantics corresponding to the group number and the reference text according to an embodiment of the present application.

[0026] Figure 6 The flowchart shows the calculation of the semantic probability value of each semantic queue according to each semantic queue and the reference text according to an embodiment of the present application.

[0027] Figure 7 The flowchart shows the input of the target semantics into the statement derivation model, and the statement derivation model generates at least one initial statement and the applicable probability value of each initial statement according to the target semantics.

[0028] Figure 8 The block diagram of the computer system shows the implementation of the intelligent input method according to an embodiment of the present application. Detailed implementation manners

[0029] Now, the example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and the concept of the example embodiments will be fully conveyed to those skilled in the art.

[0030] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0031] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0032] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily include all content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0033] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.

[0034] Please refer to Figure 1 , Figure 1 , which shows a flowchart of an intelligent input method according to an embodiment of the present application. The embodiments of the present application provide steps of an intelligent input method, including: Step S110, identifying the obtained string to obtain at least one target character group, and obtaining the group serial number corresponding to each target character group; Step S120, calculating multiple semantic probability values of the string according to at least one semantics corresponding to the group serial number and the reference text, and taking the semantics represented by the highest semantic probability value as the target semantics of the string, where the reference text includes input records and / or chat records; Step S130, inputting the target semantics into a statement derivation model, and the statement derivation model generates at least one initial statement and the applicable probability value of each initial statement according to the target semantics; Step S140, sorting the initial statements according to the applicable probability values, and displaying the top K initial statements as target statements, where K is a positive integer.

[0035] The above 4 steps will be described in detail below.

[0036] In step S110, a string generated based on the user's operation behavior is obtained. At this time, the string is still a series of letters or symbols. Whether it is necessary to determine Chinese according to the letters or determine other languages according to the letters, the string needs to be split to obtain at least multiple target character groups. A target character group refers to a string that can form semantics, which is convenient for subsequent generation of target statements. For example, when the user uses pinyin input and enters multiple letters to obtain a string, the target character groups determined from the string are pinyin. For example, "woaini" is split into three target character groups: "wo", "ai", and "ni".

[0037] Therefore, in step S110, the target character groups are identified from the string, at least one target character group is obtained, and the group serial numbers corresponding to the target character groups are acquired. It should be clear that each target character group can correspond to a token (analogous to the pinyin of Chinese characters). For example, "woaini" is segmented into "wo", "ai", and "ni", which respectively correspond to the group serial numbers "001", "002", and "003".

[0038] In some embodiments, each target character group can correspond to a phrase (analogous to the pinyin of Chinese words or English words). For example, "woaini" is segmented into "wo" and "aini", which respectively correspond to the group serial numbers "001" and "002".

[0039] In the above embodiments, by numbering the target character groups, the consumption of computing resources can be reduced, and the running speed of the present technical solution can be increased, that is, the output efficiency of the target statement can be increased.

[0040] In some embodiments, after the string is obtained, the string is first cleaned, such as removing punctuation marks and non-character content. After the cleaning is completed, the target character groups are obtained from the string. In some embodiments, the string is also subjected to fuzzy sound processing, and then the string is segmented to obtain the target character groups. In some embodiments, the string as a whole is also corrected, such as adjacent key correction; homophone / near-homophone correction; similar glyph correction; high-frequency character priority correction. After the correction is completed, the target character groups are obtained.

[0041] Please refer to Figure 2 , Figure 2 which shows a flowchart of identifying the obtained string to obtain at least one target character group according to an embodiment of the present application. The embodiment of the present application provides step S110 of identifying the obtained string to obtain at least one target character group, including: Step S111a: In the string, a string with a target length is obtained as the first character group according to the arrangement order of each character. If the first character group can be successfully matched with the dictionary, the first character group is used as the target character group; Step S112a: If the first character group cannot be successfully matched with the dictionary, the value of the target length is gradually decreased to re-obtain the first character group until the first character group is successfully matched or the first character group is a single character, and the first character group is used as the target character group; Step S113a: Continue to identify the characters that have not been segmented until all characters belong to the target character groups.

[0042] The following will start a detailed description of the above three steps.

[0043] In step 111a, in a string, a string of a target length is obtained as the first character group according to the arrangement order of each character. If the first character group can be successfully matched in the dictionary, the first character group is used as the target character group. It should be noted that there are multiple character groups composed of different characters in the dictionary. If the first character group can be matched with the dictionary, it means that the first character group meets the formation conditions of the target character group.

[0044] It should be noted that the target length refers to the maximum length of the character groups in the dictionary. It should be further noted that obtaining a string of the target length as the first character group according to the arrangement order of each character can be carried out in the forward order of each character (such as from left to right), or in the reverse order (from right to left).

[0045] In step S112a, if the first character group cannot be successfully matched in the dictionary, the value of the target length is gradually reduced to re-obtain the first character group until the first character group is successfully matched or the first character group is a single character, and the first character group is used as the target character group.

[0046] In step S113a, the characters that have not been divided are continuously recognized until all characters have their respective target character groups.

[0047] In the embodiment of the present application, by gradually dividing the string, the long string can be accurately divided into clear target character groups.

[0048] In some embodiments, the string is divided in the forward order of each character and the first character group is divided in the reverse order. It should be noted that the forward order means such as from left to right, that is, in the order of character input time (from the first input → the last input); the reverse order means from right to left, that is, in the reverse order of character input time (from the last input → the first input).

[0049] That is, the following steps are carried out in two directions respectively: obtaining a string of the target length as the first character group. If the first character group can be successfully matched in the dictionary, the first character group is used as the forward initial character group or the reverse initial character group according to the division direction; that is, the initial character group obtained by dividing the string in the forward order is the forward initial character group. The initial character group obtained by dividing the string in the reverse order is the reverse initial character group.

[0050] If the first character group fails to match successfully in the dictionary, the value of the target length is gradually decreased to re-obtain the first character group until the first character group matches successfully or the first character group is a single character. Then, according to the division direction, the first character group is used as the forward initial character group and the reverse initial character group. The recognition of the characters that have not been divided is continued until all characters have their respective forward initial character groups and reverse initial character groups.

[0051] In some embodiments, if the number of reverse initial character groups is less than that of the forward initial character groups, the reverse initial character group is used as the target character group. If the number of forward initial character groups is less than that of the reverse initial character groups, the forward initial character group is used as the target character group.

[0052] In other embodiments, if the number of reverse initial character groups is more than that of the forward initial character groups, the reverse initial character group is used as the target character group. If the number of forward initial character groups is more than that of the reverse initial character groups, the forward initial character group is used as the target character group.

[0053] In other embodiments, the selection rule between the reverse initial character group and the forward initial character group can be customized.

[0054] In some embodiments, the same part in the forward initial character group and the reverse initial character is used as the target character group. If the acquisition of the target character group is completely accurate, the forward initial character group and the reverse initial character will be exactly the same. However, in actual application, there will be some differences between the reverse initial character group and the forward initial character group because there are incorrect characters in the string, such as incorrect characters formed due to user input errors, which will appear in the form of "the different parts between the forward initial character group and the reverse initial character group" in this embodiment. For the above situation, the "different parts between the forward initial character group and the reverse initial character group" are re-split into characters. If there is no target character group interval between the two split characters, these two characters are added to the same deviation string, and then all the split characters are re-formed into a deviation string to correct the string. For example, adjacent key error correction; homophone / near-homophone error correction; similar glyph error correction; high-frequency word priority error correction. For the corrected string, For the corrected deviation string, the target character group obtained by using any of the division methods mentioned in the above embodiments is used. Then, all the target character groups corresponding to the string are obtained. For example, for the corrected deviation string, the forward initial character group and the reverse initial character group are continuously obtained, and according to the preset screening rule, the larger or smaller one of the forward initial character and the reverse initial character group is used as the target character group.

[0055] It should be clear that after re - splitting the "different parts in the forward initial character combination reverse initial character group" into characters for correction, there is no need to re - split the "different parts in the forward initial character combination reverse initial character group" into characters for error correction. This is because on the one hand, one - time cleaning is sufficient to complete the error correction of characters. On the other hand, the string cannot be corrected multiple times, as multiple corrections of the string will lead to string distortion.

[0056] Please refer to Figure 3 , Figure 3 FIG. shows a flowchart of identifying the obtained string to obtain at least one target character group according to another embodiment of the present application. The embodiments of the present application provide step S110 of identifying the obtained string to obtain at least one target character group, including: Step S111b, if the length of the string is greater than a preset length threshold, then divide the string into multiple substrings; Step S112b, in any one of the substrings, obtain a string of a target length as the second character group in the order of character arrangement. If the second character group can be successfully matched with the dictionary, then use the second character group as the target character group; Step S113b, if the second character group cannot be successfully matched with the dictionary, then gradually reduce the value of the target length to re - obtain the second character group until it is successfully matched with the dictionary or the second character group is a single character, and use the second character group as the target character group; Step S114b, continue to divide the characters in the substring that have not been divided until all characters belong to a target character group.

[0057] The above 4 steps are described in detail below.

[0058] In step S111b, if the length of the string is greater than a preset length threshold, then divide the obtained string to obtain multiple substrings.

[0059] In some embodiments, when dividing the string, the string can be divided into substrings of a set length. In some embodiments, the string is divided into a set number of substrings, which can be evenly divided or not evenly divided.

[0060] In step S112b, for any one of the substrings, obtain a string of a second length as the second character group in the order of character arrangement. If the second character group can be successfully matched in the dictionary, then use the second character group as the target character group. It should be clear that for the substring, it can be divided in the forward order of each character and the second character group can be divided in the reverse order.

[0061] In step S113b, if the second character group fails to match successfully in the dictionary, the value of the second length is gradually decreased to re-obtain the second character group until a successful match is achieved or the second character group is a single character, and the second character group is used as the target character group.

[0062] In step S114b, the characters in the substring that have not been divided are continuously divided until all characters belong to a target character group.

[0063] It should be clear that for each substring, the technical means described in steps S112b and S114b are used to obtain the target character group.

[0064] In some embodiments, to determine the target character group in the sub-characters, the target character group can be determined respectively using the methods described in the above embodiments.

[0065] In the embodiments of the present application, if the string length is too large, it will cause too long a division time for the string, affecting the input efficiency of the input method. Therefore, the string is divided into multiple substrings to respectively perform the recognition of the target character group, thereby accelerating the acquisition of the target character group.

[0066] Please refer to Figure 4 , Figure 4 which shows a flowchart of correcting the target character group according to an embodiment of the present application. The embodiments of the present application provide steps for correcting the target character group, including: Step S201, for two target character groups that belong to different substrings and are adjacent, if there is a target character group that is a single character, then the two target character groups that belong to different substrings and are adjacent are disassembled to obtain a compensation string; Step S202, re-obtain the target character group for each compensation string.

[0067] The above two steps are described in detail below.

[0068] In step S201, for two target character groups that belong to different substrings and are adjacent, if there is a target character group that is a single character, then the two target character groups that belong to different substrings and are adjacent are respectively disassembled, and then the characters obtained by disassembling these two target character groups are used as the compensation string.

[0069] In step S202, re-obtain the target character group for each compensation string. It should be clear that the method for obtaining the target character group for the compensation string can use the technical means for obtaining the target character group described in any of the above embodiments.

[0070] In the embodiments of the present application, dividing a string into multiple substrings may cause the loss of some target character groups. For example, adjacent characters that should form a target character group are respectively divided into two adjacent substrings. In view of the above situation, for two adjacent target character groups belonging to different substrings, if there is a target character group that is a single character, the two adjacent target character groups belonging to different substrings are respectively disassembled to obtain a compensation string. The characters that should originally belong to the same target character group are reassigned to the same compensation string, and the acquisition of the target character group is performed again. Through the above technical means, while efficiently obtaining the target character group of the string, the accuracy of obtaining the target character is ensured.

[0071] In step S120, each group number corresponds to a target character group, and the possible semantics of each target character group, that is, the possible semantics corresponding to the group number of the target character group. For example, if the target character group is pinyin, the semantics corresponding to the target character group is the Chinese character.

[0072] In some embodiments, the respective semantics corresponding to the group numbers are also represented by semantic numbers. Each semantic number represents a kind of semantics.

[0073] For example, for the target character group "qin", the group number is "001", the semantics are "qin, qin, qin", etc., and the semantic numbers are "qin (001001), qin (001002), qin (001003)", etc. In this way, the characters can also be digitized, thereby improving the running efficiency of the input method and reducing the occupation of computing resources by the input method. Each group number corresponds to at least one semantics. By combining the semantics of each group number, the semantics corresponding to the string can be obtained, that is, there is at least one semantics corresponding to the string. Selecting one semantics corresponding to each group number respectively can obtain a semantic queue. Calculate the semantic probability values in the semantic queue respectively. The semantic queue with the highest semantic probability value is used as the semantics of the string.

[0074] Please refer to Figure 5 , Figure 5 shows a flowchart of calculating multiple semantic probability values of a string according to at least one semantics corresponding to a group number and a reference text according to an embodiment of the present application. The embodiments of the present application provide step S120 of calculating multiple semantic probability values of a string according to at least one semantics corresponding to a group number and a reference text, including: Step S121, according to at least one semantics corresponding to each group number, obtain multiple semantic queues corresponding to the string, and the semantics of each semantic queue are different; Step S122, calculate the semantic probability values of each semantic queue according to each semantic queue and the reference text.

[0075] The above two steps will be described in detail below.

[0076] In step S121, each group number corresponds to at least one semantics. By combining the semantics corresponding to each group number, the semantics corresponding to the string can be obtained. The more the number of semantics corresponding to the group number, the more semantics queues corresponding to the string. By selecting one semantics corresponding to each group number respectively, a semantics queue of the string can be obtained. In this way, multiple semantics queues corresponding to the string can be obtained, and the semantics of each semantics queue are different.

[0077] In step S122, according to the semantics corresponding to each group number in the semantics queue and the reference text, the semantic probability value of each semantics queue is calculated.

[0078] In some embodiments, the reference text can be various texts (such as some articles with writing specifications) for calculating the matching degree of each semantics queue with various texts, so as to calculate the semantic probability value of the semantics queue.

[0079] Please refer to Figure 6 , Figure 6 , which shows a flowchart for calculating the semantic probability value of each semantics queue according to each semantics queue and the reference text according to an embodiment of the present application. The embodiment of the present application provides step S122 for calculating the semantic probability value of each semantics queue according to each semantics queue and the reference text, including: Step S1221, calculate the sub-probability value of the current semantics according to the number of times the pre-semantics appears in the reference text; if the current semantics has no pre-semantics, then use the probability of the current semantics appearing in the reference text as the sub-probability value of the current semantics; Step S1222, calculate the semantic probability value of each semantics queue according to the sub-probability values of each semantics.

[0080] The above two steps will be described in detail below.

[0081] It should be clear that each group number corresponds to at least one semantics, and each semantics queue includes a certain semantics of each group number. The arrangement order of the group numbers is the same as the arrangement order of the target character groups corresponding to the group numbers, and the arrangement order of the semantics corresponding to each group number is the same as the arrangement order of the group numbers, that is, the arrangement order of the semantics in each semantics queue is the same as the arrangement order of the group numbers.

[0082] In step S1221, according to the number of occurrences of the pre-semantics in the reference text, calculate the sub-probability value of the current semantics; if the current semantics has no pre-semantics, then use the probability of the current semantics appearing in the reference text (i.e., the ratio of the current semantics in the reference text) as the sub-probability value of the current semantics. It should be clear that the current semantics refers to the semantics for which the sub-probability value needs to be calculated currently, and the pre-semantics refers to the semantics that are located before the current semantics in the semantic arrangement order. In some other embodiments, the pre-semantics refers to the semantics that are located before the current semantics in the semantic arrangement order and are adjacent to the current semantics.

[0083] In some embodiments, use the number of occurrences of the pre-semantics in the reference text as the denominator, and use the number of adjacent occurrences of the pre-semantics and the current semantics in the reference text as the numerator. The resulting ratio is the sub-probability value of the current semantics. If the current semantics has no pre-semantics, then use the number of occurrences of the current semantics in the reference text as the numerator and the total number of semantics in the reference text as the denominator. The resulting ratio is the sub-probability value of the current semantics.

[0084] In some embodiments, each semantics in the reference text can also be digitized. To facilitate the calculation of semantic probability values, the same semantics have the same semantic number. The same semantic number can be quickly searched in the reference text. At the same time, digitizing the reference text greatly reduces the memory occupancy of the intelligent input method of the present application and saves storage resources.

[0085] In step S1222, calculate the semantic probability value of the semantic queue according to the sub-probability values of each semantics.

[0086] In some implementations, use the sum value of the sub-probability values of each semantics as the semantic probability value of the semantic queue. In some embodiments, use the product obtained by multiplying the sub-probability values of each semantics as the semantic probability value of the semantic queue. In some embodiments, substitute the sub-probability values of each semantics into a set formula for calculation to obtain the semantic probability value of the semantic queue.

[0087] In some embodiments, the calculation of the probability of various semantics of a string is performed through three aspects: the rationality of semantic splicing of adjacent group numbers, the usage record of the input method used by the user, and the chat record of the user. Furthermore, it makes the target semantics of the obtained string logically smooth while conforming to the user's input habits and the user's current chat scenario, ensuring the optimal semantic effect of the string.

[0088] In some embodiments, the first text may be a reference text. The first text refers to a conventional text without grammar errors. The first text may be various texts (such as some articles that comply with writing norms) used to calculate the matching degree between each semantic queue and various texts, so as to calculate the reasonable probability value of the semantic queue. The higher the reasonable probability value, the more the semantic queue complies with the grammar rules and the more reasonable it is. Finally, the semantic probability value of the semantic queue calculated by steps S1221 - S1222 is used.

[0089] In some embodiments, the second text may also be a reference text, and the second text is used to record the user's input record.

[0090] First, take the first text as the reference text, and use the individual sub - probability values in the semantic queue calculated by the technical means described in steps S1221 - S1222 as the first probability values, and calculate the reasonable probability value of the semantic queue according to each first probability value. In some embodiments, the product obtained by multiplying the first probability values of each semantics is used as the reasonable probability value of the semantic queue. In some embodiments, substitute the first probability values of each semantics into a set formula for calculation to obtain the reasonable probability value of the semantic queue.

[0091] Then, take the second text as the reference text, and use the individual sub - probability values in the semantic queue calculated by the technical means described in steps S1221 - S1222 as the second probability values, and calculate the historical probability value of the semantic queue according to each second probability value. In some embodiments, the product obtained by multiplying the second probability values of each semantics is used as the historical probability value of the semantic queue. In some embodiments, substitute the second probability values of each semantics into a set formula for calculation to obtain the historical probability value of the semantic queue. The higher the historical probability value, the more the description complies with the user's input habits.

[0092] Finally, calculate the semantic probability value of the semantic queue according to the reasonable probability value and the historical probability value. Exemplarily, use the sum value of the reasonable probability value and the historical probability value as the semantic probability value of the semantic queue. Or, calculate the sum value of the reasonable probability value and the historical probability value according to a preset weight as the semantic probability value of the semantic queue. Or, substitute the reasonable probability value and the historical probability value into a preset formula to calculate the semantic probability value of the semantic queue.

[0093] In some other embodiments, the third text may be a reference text, which refers to the chat records generated under the current chat box or chat event (it should be clear that the chat event refers to a certain event targeted in the chat box). After calculating the historical probability value of the semantic queue, the third text is used as the reference text, and the sub-probability values in the semantic queue calculated by the technical means described in steps S1221 - S1222 are used as the third probability values, and the scenario probability value of the semantic queue is calculated based on each third probability value. In some embodiments, the product obtained by multiplying the third probability values of each semantics is used as the scenario probability value of the semantic queue. In some embodiments, the third probability values of each semantics are substituted into a set formula for calculation to obtain the scenario probability value of the semantic queue. The higher the historical probability value, the more the manual complies with the user's current chat scenario.

[0094] Finally, the semantic probability value of the semantic queue is calculated based on the reasonable probability value, the historical probability value, and the scenario probability value. Exemplarily, the sum value of the reasonable probability value and the historical probability value is used as the semantic probability value of the semantic queue. Or, the sum value of the reasonable probability value and the historical probability value is calculated according to a preset weight as the semantic probability value of the semantic queue. Or, the reasonable probability value and the historical probability value are substituted into a preset formula to calculate the semantic probability value of the semantic queue.

[0095] In some embodiments, the second text and the third text need to be updated frequently because the second text and the third text have timeliness, and the update rules for the second text and the third text can be set by the developer themselves. Exemplarily, the second text and / or the third text are updated every set time interval. For example, in some embodiments, whenever the semantic probability value is obtained, the chat records of a set time period under the current chat box or chat event are used as the third text.

[0096] In step S130, the target semantics of the string are input into the statement derivation model, and the statement derivation model generates at least one initial statement and the applicable probability value of the initial statement according to the target semantics. The statement derivation model is used to extend and derive statements based on a small number of prompt words (such as the target semantics) as the initial statements.

[0097] For the calculation of the applicable probability value of each initial statement, the following technical means can be used. Please refer to Figure 7 , Figure 7 shows a flowchart of inputting the target semantics into the statement derivation model, and the statement derivation model generating at least one initial statement and the applicable probability value of each initial statement according to the target semantics. The embodiment of the present application provides step S130 of inputting the target semantics into the statement derivation model, and the statement derivation model generating at least one initial statement and the applicable probability value of each initial statement according to the target semantics, including: Step S131: The statement derivation model generates multiple initial statements according to the target semantics. Step S132: Match each initial statement with the reference text to obtain the applicable probability values of multiple initial statements.

[0098] The above two steps are described in detail below.

[0099] In step S131, the statement derivation model generates multiple initial statements according to the target semantics.

[0100] In step S132, match each initial statement with the chat record to obtain the applicable probability values of multiple initial statements. Specifically, divide the initial statement to obtain multiple tokens (equivalent to the semantics above). When calculating the applicable probability value of each initial statement, first calculate the probability value of each token in the initial statement as the initial probability value, and predict the initial probability value of each token in sequence according to the position order of the tokens in the initial statement; among them, For the current token, calculate the initial probability value of the current token according to the number of times the previous token appears in the reference text; if the current token has no previous token, then use the probability of the current token appearing in the reference text (that is, the proportion of the semantics corresponding to the current token in the reference text) as the initial probability value of the current token.

[0101] In some embodiments, use the number of times the previous token appears in the reference text as the denominator, and use the number of times the previous token and the current token appear adjacent to each other in the reference text as the numerator. The obtained ratio is the initial probability value of the current token. If the current token has no previous token (that is, the current token is the first token of the initial statement), then use the number of times the current token appears in the reference text as the numerator, and use the total number of tokens in the reference text as the denominator. The obtained ratio is the initial probability value of the current token. Calculate the applicable probability value of the initial statement according to the initial probability values of each token.

[0102] In some implementations, use the sum value of the initial probability values of each token as the applicable probability value of the initial statement. In some embodiments, use the product obtained by multiplying the initial probability values of each token as the applicable probability value of the initial statement. In some embodiments, substitute the initial probability values of each token into a set formula for calculation to obtain the applicable probability value of the initial statement.

[0103] In some embodiments, use the second text as the reference text, calculate the initial probability values of each token by the above technical means, and then calculate the historical applicable probability value of the initial statement according to the initial probability values of each token, and directly use the historical applicable probability value as the applicable probability value of the initial statement.

[0104] In some implementations, first, the second text is used as the reference text, and the initial probability values of each token calculated by the above technical means are used as the first initial probability values. Then, the historical applicability probability value of the initial statement is calculated based on the first initial probability values of each token. Exemplarily, the sum value of the first initial probability values of each token is used as the historical applicability probability value of the initial statement. In some embodiments, the product obtained by multiplying the first initial probability values of each token is used as the historical applicability probability value of the initial statement. In some embodiments, the first initial probability values of each token are substituted into a set formula for calculation to obtain the historical applicability probability value of the initial statement.

[0105] Then, the third text is used as the reference text, and the initial probability values of each token calculated by the above technical means are used as the second initial probability values. Then, the scenario applicability probability value of the initial statement is calculated based on the second initial probability values of each token. In some implementations, the sum value of the second initial probability values of each token is used as the scenario applicability probability value of the initial statement. In some embodiments, the product obtained by multiplying the second initial probability values of each token is used as the scenario applicability probability value of the initial statement. In some embodiments, the second initial probability values of each token are substituted into a set formula for calculation to obtain the scenario applicability probability value of the initial statement.

[0106] In some embodiments, the applicability probability value of the initial statement is calculated based on the historical applicability probability value and the scenario applicability probability value. Exemplarily, the sum value of the historical applicability probability value and the scenario applicability probability value is used as the applicability probability value of the initial statement. Or, the sum value of the historical applicability probability value and the scenario applicability probability value is used as the applicability probability value of the initial statement. Or, the historical applicability probability value and the scenario applicability probability value are substituted into a preset formula for calculation to obtain the applicability probability value of the initial statement.

[0107] The obtained initial statement is further screened in the above manner, ensuring that the initial statement can better conform to the user's input habits and the current chat scenario.

[0108] In step S140, the initial statements are sorted according to the applicability probability value, and the top K initial statements are used as the target statements for display for the user to select.

[0109] This application accurately identifies the target character group from the string input by the user, and obtains the target semantics of the string based on the target character group and the reference text, making the target semantics more in line with the user's input habits. Finally, the identified target semantics are output as target statements through a statement derivation model for the user to select. That is, when using the input method, the user only needs to input a small number of characters to generate at least one candidate target statement for the user to select. And because the target semantics conform to the user's input habits, the obtained target semantics also conform to the user's input habits, thereby enabling the user to generate input text that conforms to the user's input habits with very few input operations, greatly improving the user's input efficiency and enhancing the user's experience.

[0110] In some embodiments, the statement derivation model is a BERT model (Bidirectional Encoder Representation from Transformers); wherein, the BERT model includes an input layer, a transformation layer, and an output layer. The input layer is used to vectorize the input target semantics to obtain a target vector group. The transformation layer is used to analyze the relationships between the vectors in the target vector group from a set number of analysis perspectives. The output layer is used to generate target statements based on the analysis results of the target vector group. For the obtained open-source BERT model, convert the M-bit floating-point numbers of the model weights and activation values in each layer of the BERT model into N-bit integers, where M and N are positive integers and M > N (for example, compress the model weights from 32-bit floating-point (FP32) to 8-bit integer (INT8), reducing the volume by 75% while maintaining an inference accuracy of more than 95%). And delete multiple transformation layers to obtain an initial statement derivation model.

[0111] For example, use TensorFlow Lite or ONNX to perform model conversion on the BERT model. Through the application of INT8 quantization (volume reduction of 75%, CPU inference acceleration of 3 - 4 times), structured pruning can also be performed. By improving the BERT model, the input method relying on the BERT model can run conveniently on mobile terminals (such as mobile phones). Finally, train the initial statement derivation model to obtain the statement derivation model.

[0112] Figure 8 Shows a block diagram of a computer system for implementing an intelligent input method according to an embodiment of the present application.

[0113] It should be noted that Figure 8 The shown computer system 800 is only an example and should not impose any restrictions on the functions and usage scope of the embodiments of the present application.

[0114] Such as Figure 8As shown, computer system 800 includes a central processing unit 801 (CPU), which can perform various appropriate actions and processes according to programs stored in read-only memory 802 (ROM) or programs loaded from storage section 808 into random access memory 803 (RAM). In random access memory 803, various programs and data required for system operation are also stored. The central processing unit 801, read-only memory 802, and random access memory 803 are connected to each other via bus 804. Input / output interface 805 (Input / Output interface, i.e., I / O interface) is also connected to bus 804.

[0115] The following components are connected to input / output interface 805: input section 806 including a keyboard, a mouse, etc.; output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; storage section 808 including a hard disk, etc.; and communication section 809 including a network interface card such as a local area network card, a modem, etc. Communication section 809 performs communication processing via a network such as the Internet. Drive 810 is also connected to input / output interface 805 as needed. Removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on drive 810 as needed so that a computer program read from it can be installed into storage section 808 as needed.

[0116] In particular, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit 801, various functions defined in the system of the present application are executed.

[0117] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0119] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0120] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0121] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0122] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An intelligent input method, characterized in that The method includes: Identifying the obtained string to obtain at least one target character group, and obtaining the group serial number corresponding to each of the target character groups; Calculating multiple semantic probability values of the string according to at least one semantics corresponding to the group serial number and a reference text, and taking the semantics represented by the highest semantic probability value as the target semantics of the string, where the reference text includes input records and / or chat records; Inputting the target semantics into a statement derivation model, where the statement derivation model generates at least one initial statement and the applicable probability value of each of the initial statements according to the target semantics; Sorting the initial statements according to the applicable probability values, and displaying the top K initial statements as target statements, where K is a positive integer.

2. The method according to claim 1, characterized in that Identifying the obtained string to obtain at least one target character group, including: In the string, obtaining a string with a target length in the arrangement order of each character as the first character group. If the first character group can be successfully matched with a dictionary, the first character group is used as the target character group; If the first character group cannot be successfully matched with the dictionary, the value of the target length is gradually reduced to re-obtain the first character group until the first character group is successfully matched or the first character group is a single character, and the first character group is used as the target character group; Continuing to identify the characters that have not been divided until all characters belong to the target character groups.

3. The method according to claim 1, wherein Identifying the obtained string to obtain at least one target character group, including: If the length of the string is greater than a preset length threshold, the string is divided into multiple substrings; In any one of the substrings, obtaining a string with a target length in the character arrangement order as the second character group. If the second character group can be successfully matched with a dictionary, the second character group is used as the target character group; If the second character group cannot be successfully matched with the dictionary, the value of the target length is gradually reduced to re-obtain the second character group until it is successfully matched with the dictionary or the second character group is a single character, and the second character group is used as the target character group; Continuing to divide the characters in the substring that have not been divided until all characters belong to the target character groups.

4. The method according to claim 3, wherein After continuing to divide the characters in the substring that have not been divided until all characters belong to the target character groups, the method further includes: For two target character groups that belong to different substrings and are adjacent, if there is a target character group that is a single character, the two target character groups that belong to different substrings and are adjacent are dissolved to obtain a compensation string; Re-obtaining the target character groups for each of the compensation strings.

5. The method according to claim 1, characterized in that Calculating multiple semantic probability values of the string according to at least one semantics corresponding to the group serial number and a reference text, including: According to at least one semantics corresponding to each of the group serial numbers, obtaining multiple semantic queues corresponding to the string, where the semantics of each semantic queue are different; Calculating the semantic probability value of each semantic queue according to each semantic queue and the reference text.

6. The method according to claim 5, wherein Calculate the semantic probability values of the semantic queues according to each of the semantic queues and the reference text, including: Calculate the sub-probability value of the current semantics according to the number of occurrences of the pre-semantics in the reference text; if there is no pre-semantics for the current semantics, use the probability of the current semantics appearing in the reference text as the sub-probability value of the current semantics; Calculate the semantic probability values of the semantic queues according to the sub-probability values of the semantics.

7. The method according to claim 1, wherein Input the target semantics into the sentence derivation model, and the sentence derivation model generates at least one initial sentence and the applicable probability value of each initial sentence according to the target semantics, including: The sentence derivation model generates a plurality of initial sentences according to the target semantics; Match each of the initial sentences with the reference text to obtain the applicable probability values of the initial sentences.

8. The method according to claim 1, wherein The sentence derivation model is a bert model; wherein, The bert model includes an input layer, a transformation layer and an output layer. The input layer is used to vectorize the input target semantics to obtain a target vector group. The transformation layer is used to analyze the relationship between the vectors in the target vector group from a set number of analysis perspectives. The output layer is used to generate a target sentence according to the analysis result of the target vector group; For the obtained open-source bert model, convert the M-bit floating-point numbers of the model weights and activation values in each layer of the bert model into N-bit integers, where M and N are positive integers and M is greater than N, and obtain an initial sentence derivation model by deleting a plurality of transformation layers; Train the initial sentence derivation model to obtain a sentence derivation model.

9. An intelligent input device, comprising a memory, a processor, and a readable program stored on the memory, characterized in that, The processor executes the readable program to implement the method according to any one of claims 1 to 8.

10. A readable storage medium, characterized in that, Stored thereon are readable programs / instructions which, when executed by a processor, implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Statement recognition method and device, equipment and computer storage medium

    CN113064497A

  • Method and system for converting pinyin into Chinese characters

    CN117032469A

  • Text inputting method, apparatus and system

    US20140136970A1

  • Electronic device and control method therefor

    US20200058298A1