Text input information processing method and device and storage medium
By generating semantic features that characterize text input information, combining vocabulary-level and character-level position information, the problems of single functions and poor applicability of the existing input method prediction methods are solved, and high accuracy prediction in complete vocabulary and incomplete vocabulary scenarios are achieved.
Patent Information
- Application Number
- CN202510803900.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing input method to predict user input intentions has a single function, poor applicability and low accuracy, making it difficult to effectively predict vocabulary in different scenarios.
By obtaining the user's text input information, semantic features representing the text input information are generated, combined with vocabulary level and character level position information, and using the vocabulary prediction layer to generate predicted vocabulary, which is suitable for complete vocabulary and non-complete vocabulary scenarios.
It improves the accuracy and applicability of predicted vocabulary, can compatible with word completion and word prediction in different scenarios, and enhances the generalization ability of the input method.
Smart Images

Figure CN120335625A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of input method information processing, and particularly to a method, apparatus, and storage medium for processing text input information. Background Art
[0002] During the process of a user gradually inputting characters or words, a Latin input method program needs to predict the user's input intention. For example, after the user inputs a character, predict the complete words that the user may input, or after the user inputs a complete word, predict the words that the user may input next.
[0003] In the prior art, a statistical language modeling method is usually used to implement the prediction of the user's input intention, including the n-gram model and the tree matching method.
[0004] The tree matching method is mainly used to predict the complete words that the user may input according to the input characters of the incomplete word that the user is inputting (referred to as the word completion scenario). In this method, a dictionary is constructed based on a prefix tree (Trie), and the input character sequence of the user is mapped to all possible matching paths in the dictionary. For example, when inputting "r", the system will retrieve all words starting with "r", such as "row", "river", "right", etc., and filter out the most relevant word through a sorting algorithm (such as word frequency, context probability).
[0005] The n-gram model is mainly used to predict the next word that the user may input through several complete words that the user has already input (referred to as the word prediction scenario). This method realizes the prediction by statistically calculating the occurrence probability of phrases with a fixed length in the training corpus. For example, when the user inputs "ducks in", if the "triple" form of "ducks in a" frequently appears in the training corpus, then it is inclined to predict that the next word the user needs to input is "a".
[0006] The existing methods for predicting the user's input intention have the following problems: 1) The functions of the prediction methods are single. For example, the tree matching method can only generate predicted words when the user's text input information includes input characters corresponding to incomplete words; while the n-gram model can only generate predicted words when all the user's text input information is complete words. Therefore, if you want to satisfy the generation of predicted words under the text input information in two different situations, you need to deploy the programs of these two methods at the same time, which brings inconvenience in application.
[0007] 2) The tree matching method highly depends on the coverage of the dictionary, making it difficult to reasonably predict words or newly formed words that do not appear in the dictionary. Moreover, when the user inputs quickly and continuously or the language context changes, the prediction effect is poor. Therefore, the applicability of this method is poor and the prediction accuracy is low.
[0008] 3) When the n-gram model makes predictions, it calculates the probability of the occurrence of word sequences with a fixed number of words in the training corpus. Since it only depends on the context of a limited window, it is difficult to capture long-distance semantic associations. In scenarios where the training data is sparse or the combination diversity is strong, the n-gram model is easily restricted by the data coverage and has low generalization ability.
[0009] In view of the technical problems in the above-mentioned prior art, such as the single function of the method for predicting the user's input intention in the input method, the poor applicability and low accuracy of the word completion method, and the poor generalization ability of the word prediction method, no effective solution has been proposed yet. Summary of the Invention
[0010] Embodiments of the present disclosure provide a method, an apparatus, and a storage medium for processing text input information, so as to at least solve the technical problem in the prior art that the methods for predicting the user's input intention in the input method generally have poor applicability and low accuracy.
[0011] According to one aspect of the embodiments of the present disclosure, a method for processing text input information is provided, including: obtaining the user's text input information, where the text input information includes the complete words that the user has already input and / or the input characters of the incomplete words that the user is inputting; generating a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is also used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to the words and character-level position information corresponding to the characters; generating output information corresponding to the text input information by using a vocabulary prediction layer according to the first semantic feature, where the output information includes predicted words corresponding to the text input information.
[0012] According to another aspect of the embodiments of the present disclosure, a storage medium is further provided, where the storage medium includes a stored program, and when the program runs, the method described in any one of the above is executed by a processor.
[0013] According to another aspect of the embodiments of the present disclosure, there is also provided a text input information processing device, including: an acquisition module, configured to acquire the text input information of a user, where the text input information includes complete words that the user has input and / or input characters of incomplete words that the user is inputting; a feature generation module, configured to generate a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is further used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to words and character-level position information corresponding to characters; an information generation module, configured to generate output information corresponding to the text input information according to the first semantic feature by using a word prediction layer, where the output information includes predicted words corresponding to the text input information.
[0014] According to another aspect of the embodiments of the present disclosure, there is also provided a text input information processing device, including: a processor; and a memory, connected to the processor, configured to provide instructions for the processor to perform the following processing steps: acquire the text input information of a user, where the text input information includes complete words that the user has input and / or input characters of incomplete words that the user is inputting; generate a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is further used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to words and / or character-level position information corresponding to characters; generate output information corresponding to the text input information according to the first semantic feature by using a word prediction layer, where the output information includes predicted words corresponding to the text input information.
[0015] In the embodiments of the present disclosure, the method can distinguish the text input information into each input character in the complete words and incomplete words for feature extraction. When performing feature extraction, not only the semantics of the complete words and each input character in the text input information are considered, but also the dual position information (word-level position information and character-level position information) of the complete words and input characters is considered, so as to predict the information that the user needs to input next in this way, which can improve the prediction effect. Moreover, compared with the tree matching and n-gram methods, it can be compatible with the word completion scenario and the word prediction scenario, and has stronger applicability. Description of the Drawings
[0016] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure. In the drawings: Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present disclosure; Figure 2 It is a schematic flowchart of the text input information processing method according to the first aspect of Embodiment 1 of the present disclosure; Figure 3A It is a schematic structural diagram of a vocabulary prediction model for processing text input information provided by Embodiment 1 of the present disclosure; Figure 3B is Figure 3A A schematic structural diagram of the joint encoding layer in the vocabulary prediction model shown; Figure 4 It is a schematic diagram of the text input information processing device according to the first aspect of Embodiment 2 of the present disclosure; and Figure 5 It is a schematic diagram of the text input information processing device according to the first aspect of Embodiment 3 of the present disclosure. Detailed implementation manners
[0017] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0019] Embodiment 1 According to this embodiment, a method embodiment for processing text input information is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.
[0020] The method embodiments provided in this embodiment may be executed on a mobile terminal, a computer terminal, a server, or similar computing devices. Figure 1 shows a hardware structure block diagram of a computing device for implementing a text input information processing method. As Figure 1 shown, the computing device may include one or more processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory for storing data, a transmission device for communication functions, and an input / output interface. Among them, the memory, the transmission device, and the input / output interface are connected to the processor through a bus. In addition, it may further include: a display, a keyboard, and a cursor control device connected to the input / output interface. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computing device may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0021] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computing device. As involved in the embodiments of the present disclosure, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0022] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the text input information processing method in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the text input information processing method of the above application program. The memory may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the computing device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0023] The transmission device is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computing device. In one example, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computing device.
[0025] It should be noted here that in some alternative embodiments, the above Figure 1 shown computing device may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific concrete example and is intended to show the types of components that may exist in the above-mentioned computing device.
[0026] Under the above operating environment, according to the first aspect of this embodiment, a method for processing text input information is provided. This method can be implemented by Figure 1 the shown computing device. Figure 2 shows a schematic flowchart of the method. Referring to Figure 2 shown, the method includes: S202: Obtain the user's text input information, where the text input information includes the complete words that the user has already input and / or the input characters of the incomplete words that the user is inputting; S204: Generate a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is also used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to the words and character-level position information corresponding to the characters; and S206: Generate output information corresponding to the text input information by using a vocabulary prediction layer according to the first semantic feature, where the output information includes predicted words corresponding to the text input information.
[0027] Specifically, the computing device can obtain the user's text input information, which includes the complete words that the user has already input and / or the input characters of the incomplete words that the user is inputting (S202).
[0028] The text input information mentioned here may only include the complete words that the user has already entered; or only include the input characters of the incomplete words that the user is currently entering; or include both the complete words that the user has already entered and the input characters of the incomplete words that the user is currently entering. Herein, an input character refers to a single character that can form a complete word.
[0029] The input method mentioned here may be a Latin input method. For the convenience of description, the following examples will all be given using the Latin input method. Of course, this method can also be applied to input methods in other language systems, such as the English input method, etc.
[0030] For example, if the user finally needs to enter "Ducks in a row", when the user enters "Du", then "Du" is the text input information, and this text input information only includes the input characters "D" and "u" of the incomplete word "Du". When the user enters "Ducks in a", the text input information "Ducks in a" only includes complete words, that is, the three complete words "Ducks", "in", and "a". When the user enters "Ducks in a r", the text input information "Ducks in a r" includes both the complete words "Ducks", "in", and "a", and the input character "r" of the incomplete word.
[0031] The computing device can determine which are complete words and which are input characters corresponding to incomplete words in the text input information through the input method. For a number of characters in the text input information, if the input method determines that the number of characters is a word with input completed, the computing device determines that the number of characters is a complete word; otherwise, if the input method has not determined that the number of characters is a word with input completed, the computing device determines that the number of characters is an input character corresponding to an incomplete word.
[0032] For example, in the Latin input method, if the user presses the space bar after entering a number of characters, the input method can determine that the user has completed the input of a word, so that the computing device can recognize that the number of characters is a complete word that the user has already entered. In addition, if the user does not press the space bar after entering a number of characters, the input method cannot determine that the user has completed the input of a word, so that the computing device can recognize that the number of characters is an input character of an incomplete word that the user is currently entering.
[0033] The computing device can generate a first semantic feature (S204) for characterizing the semantics of the text input information according to the text input information, wherein the first semantic feature is also used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to words and / or character-level position information corresponding to characters.
[0034] Among them, the lexical-level position information is used to indicate which word in the text input information a complete word and / or an input character in the text input information belongs to. For example, in the case where the text input information is "Ducks in a ro", the lexical-level position information is used to indicate that "Ducks" is the 1st word in the text input information, "in" is the 2nd word in the text input information, "a" is the 3rd word in the text input information, and "r" and "o" belong to the 4th word in the text input information.
[0035] In addition, the character-level position information is used to indicate which character a complete word and / or an input character in the text input information is relative to an incomplete word that the user is inputting. For example, in the case where the text input information is "Ducks in aro". Since the incomplete word is "ro", the character-level position information is used to indicate that "Ducks", "in", and "a" do not belong to the incomplete word, and to indicate that "r" is the 1st character of the incomplete word and "o" is the 2nd character of the incomplete word.
[0036] That is to say, the first semantic feature can not only characterize the semantic features of the text input information, but also characterize the position features of the complete words and input characters included in the text input information in the text input information. The position features reflect the lexical-level position information and character-level position information of the complete words and input characters. As for the specific details of the first semantic feature, they will be described in detail later.
[0037] Finally, the computing device can generate output information corresponding to the text input information according to the first semantic feature, where the output information includes a predicted word corresponding to the text input information (S206).
[0038] Among them, in the case where all the text input information of the user consists of complete words, the predicted word is generated based on the complete words in the text input information. For example, in the case where the text input information is "Ducks in a" (only including complete words), the computing device can generate the predicted word "row" according to the text input information, and the output information can be "row". This situation corresponds to the word prediction scenario described in the background art.
[0039] In addition, in the case where the user's text input information includes both complete words and input characters of incomplete words, the predicted word is generated based on both the complete words input by the user and the input characters. For example, in the case where the text input information is "Ducks in a ro", the computing device can generate the predicted word "row" based on this text input information, so that the output information can be "row". This case corresponds to the word completion scenario described in the background art.
[0040] It should be noted that although in the above example, the output information is the same as the predicted word, in some cases, the output information can also be different from the predicted word. Of course, the output information can include not only the predicted word, but can be a phrase or short sentence including the predicted word. For example, in the case where the text input information is "Ducks i", the predicted word is "in", and the output information can be "Ducks in".
[0041] In addition, although the above is described by taking the word prediction scenario and the word completion scenario as examples, in this embodiment, it is also possible to satisfy both scenarios at the same time. For example, in the case where the text input information is "Ducks i", the predicted word can also be "in a row", so that the output information can also be "Ducks in a row". Thus, this embodiment can further complete word prediction on the basis of achieving word completion.
[0042] As described in the background art, the existing input methods are single in function in the user prediction method. For example, the tree matching method can only generate predicted words when the user's text input information includes input characters corresponding to incomplete words; while the n-gram model can only generate predicted words when the user's text input information consists of all complete words. Therefore, if one wants to satisfy the generation of predicted words for text input information in two different cases, the programs of these two methods need to be deployed simultaneously, which brings inconvenience in application.
[0043] In view of this, in the technical solution of this application, after extracting features from the user's input text input information, the generated semantic features are not only used to represent the semantics of the complete words and each input character in the text input information, but also can represent the dual position information (word-level position information and character-level position information) of the complete words and input characters. Thus, in this way, the technical solution of this application can perform word prediction for both the case where the text input information includes input characters corresponding to incomplete words and the case where the text input information consists of all complete words by using only a single word prediction model. Thereby, the technical problem in the prior art that the existing input method for predicting the user input intention is single in function is solved.
[0044] Optionally, when the text input information only includes complete words, the predicted word includes the next complete word relative to the last complete word in the text input information; and when the text input information includes input characters of incomplete words, the predicted word includes the complete word corresponding to the input characters.
[0045] Therefore, the word prediction model described in this application can be applicable to both the word completion scenario corresponding to the tree matching method and the word prediction scenario corresponding to the n-gram model. Compared with the tree matching method and the n-gram model, which can be compatible with the word completion scenario and the word prediction scenario, it has stronger applicability, thus solving the technical problem of the single application scenario of the tree matching method and the n-gram model in the prior art.
[0046] Optionally, the operation of generating the first semantic feature for characterizing the semantics of the text input information according to the text input information includes: for each semantic unit in the text input information, generating a corresponding second semantic feature, where a semantic unit is a single complete word and / or a single input character of the text input information; determining the position features corresponding to each semantic unit, where the position features are used to characterize the word-level position information and character-level position information of the corresponding semantic unit in the text input information; fusing the second semantic feature with the corresponding position feature to obtain the corresponding third semantic feature; and determining the first semantic feature according to the third semantic feature.
[0047] Specifically, Figure 3A FIG. shows a schematic diagram of a word prediction model deployed on a computing device to process text input information. The word prediction model is, for example, a program module that can be executed by the computing device.
[0048] Referring to Figure 3A As shown, the structure of the word prediction model can specifically include three layers, namely: a joint encoding layer for generating the third semantic feature, a context modeling model for context modeling, and a word prediction layer for generating predicted words. Figure 3A The joint encoding layer in mainly generates the third semantic feature of the text input information, that is, the joint encoding layer is mainly used to determine the second semantic feature and position features (including the word position encoding and character position encoding mentioned below) of the semantic units in the text input information, and fuse them to obtain the third semantic feature. Then, the third semantic feature of each semantic unit can be input into the context modeling model to obtain the first semantic feature. In addition, referring to Figure 3B As shown, the joint encoding layer further includes: a first embedding layer, a second embedding layer, a third embedding layer, and a fusion layer.
[0049] Therefore, in the process of generating the first semantic feature for characterizing the semantics of the text input information according to the text input information, the computing device inputs the text deviceFigure 3A The vocabulary prediction model shown. Thus, the joint encoding layer of the vocabulary prediction model receives the text input information.
[0050] Thus, referring to Figure 3B As shown, when the first embedding layer of the joint encoding layer extracts the semantic features of the text input information, it can use a complete vocabulary in the text input information as the minimum unit for feature extraction (i.e., the semantic unit mentioned above), or use an input character in an incomplete vocabulary in the text input information as the minimum unit for feature extraction (i.e., the semantic unit), and extract semantic features (second semantic features) for each minimum unit respectively.
[0051] Among them, the second semantic feature is a feature uniquely corresponding to different complete vocabularies or different input characters. For example, the text input information is represented as: a set of complete vocabularies W = w 1, w 2, ..., w t-1 , and a set of input characters of incomplete vocabularies C t = c 1, c 2, ..., c k . Therefore, the first embedding layer can generate second semantic features corresponding to complete vocabularies E w ( w i ) ( i = 1 to t -1), and generate second semantic features corresponding to input characters E c ( c j ) ( j = 1 to k ). Thus, the second semantic features corresponding to complete vocabularies E w ( w i ) and the second semantic features corresponding to input characters E c ( c j ) together constitute the second semantic features corresponding to the text input information E ( x m ) ( m = 1 to t + k -1). The specific generation method of the second semantic features will be described in detail later.
[0052] Then, with further reference to Figure 3B as shown, the second embedding layer and the third embedding layer of the joint encoding layer can be used to determine the position features corresponding to each semantic unit P m (the position features P m can include Figure 3B the lexical position features in P w ( x m ) and the character position features P c ( x m )), where the position features P m can be used to represent the lexical-level position information and character-level position information of the corresponding semantic unit in the text input information (corresponding to the lexical position encoding and the character position encoding). Regarding the specific determination method of the position features P m , it will be described in detail later.
[0053] Then, the fusion layer fuses the second semantic feature E ( x m ) with the corresponding position feature P m to obtain the corresponding third semantic feature. Among them, in this embodiment, for example, the fusion can be achieved by summing the second semantic feature E ( x m ) and the corresponding position feature P m to generate the corresponding third semantic feature X ( x m ). However, the fusion method of the second semantic feature and the position feature of the same semantic unit is not limited to this.
[0054] Then, with reference to Figure 3A as shown, the context modeling model then determines the corresponding first semantic feature according to the third semantic feature X ( x m ). The specific operation of determining the first semantic feature will be described in detail later.
[0055] Thus, in the technical solution of the present application, for each semantic unit, corresponding semantic features (i.e., the second semantic features) and position features are respectively generated. Among them, the second semantic features can represent the semantics of the semantic unit, and the position features can represent the lexical-level position information and character-level position information of the semantic unit. Thus, the third semantic features generated by fusing the second semantic features and the position features can represent both the semantic information of the corresponding semantic unit and the lexical-level position information and character-level position information of the corresponding semantic unit. Furthermore, the first semantic features determined according to the third semantic features can represent both the semantic information of the corresponding semantic unit and the lexical-level position information and character-level position information of the corresponding semantic unit. Thus, in this way, the vocabulary prediction model can perform more accurate vocabulary prediction for the text input information based on the lexical-level position information and character-level position information of each semantic unit.
[0056] Optionally, the operation of generating corresponding second semantic features for each semantic unit in the text input information includes: in the case where the semantic unit is a single complete word, taking the word embedding vector of the single complete word as the second semantic feature corresponding to the semantic unit; and in the case where the semantic unit is a single input character, taking the character embedding vector of the single input character as the second semantic feature corresponding to the semantic unit.
[0057] Specifically, referring to Figure 3B as shown, in the process of generating the second semantic features E ( x m ), the first embedding layer performs an embedding operation on the semantic unit. When a semantic unit is a complete word, the second semantic feature of the semantic unit is the semantic feature generated by embedding (embedding) the complete word itself. When a semantic unit is an input character in an incomplete word, the second semantic feature of the semantic unit is the semantic feature generated by embedding (embedding) the input character itself.
[0058] Specifically, the first embedding layer can obtain the word embedding vector or character embedding vector of a semantic unit by looking up a table. When the semantic unit x m is a complete word w i , the index of the complete word can be queried in the preset vocabulary v w . Thus, the word embedding vector of the complete word w i can be queried according to the index w i E w (w i ) ∈ R d , as the second semantic feature corresponding to this semantic unit E ( x m ), where d represents the dimension of the vector.
[0059] Similarly, when the semantic unit x m is the input character c j , the first embedding layer can query the index of this input character in the preset character table v c , and thus query the character embedding vector of this input character c j c j ( E c ) ∈ R c j as the second semantic feature corresponding to this semantic unit d E (( x m )).
[0060] Thus, when the text input information includes the set of complete vocabulary W = w 1, w 2,..., w t-1 and the set of input characters of incomplete vocabulary C t = c 1, c 2,..., c k , the second semantic feature sequence of each semantic unit corresponding to the text input information is: E = E w ( w 1)),..., E w ( w t-1 ), E c ( c 1)),..., E c ( c 1)] ∈ R (t-1+k)×d .
[0061] Preferably, the vocabulary v w and the character table v c can be mapped to the same vector space. Thus, in this case, the computing device can calculate the distance between the word vectors of the complete vocabulary and the character vectors of the characters. Thus, in this case, the same model can be used to implement both the word prediction function and the word completion function. For example, when the text input information only contains complete vocabulary, the complete vocabulary can be mapped to this vector space, and then subsequent operations corresponding to the word prediction function can be further performed; when the text input information includes both complete vocabulary and input characters of incomplete vocabulary, the complete vocabulary and the input characters can be mapped to this vector space respectively, and then subsequent operations corresponding to the word completion function can be further performed.
[0062] Moreover, the computing device can pre-determine the word embedding vectors of each vocabulary and the character embedding vectors of each character, so as to construct the vocabulary and the character table. Thus, when the first embedding layer determines the word embedding vectors of each vocabulary in the vocabulary and the character embedding vectors of the characters in the character table, it can query and determine in the corresponding vocabulary and character table, which will not be elaborated here.
[0063] Optionally, the operation of determining the position feature corresponding to each semantic unit includes: determining the word position encoding corresponding to the semantic unit according to the position of the vocabulary corresponding to the semantic unit in the text input information; determining the character position encoding corresponding to the semantic unit, where, when the semantic unit is a character in an incomplete vocabulary, the corresponding character position encoding indicates the character position of the semantic unit in the corresponding vocabulary, and when the semantic unit is a complete vocabulary, the corresponding character position encoding can be used to determine that the semantic unit is a complete vocabulary; and determining the position feature corresponding to the semantic unit according to the word position encoding and the character position encoding corresponding to the semantic unit.
[0064] That is to say, the position feature corresponding to the semantic unit contains two layers of position information. One layer represents the position information at the word level, and the position information at the word level of the semantic unit is reflected through the corresponding word position encoding; one layer represents the position information at the character level, and the position information at the character level of the semantic unit is reflected through the corresponding character position encoding, so that the position features of each semantic unit can completely and hierarchically represent the position situation of each semantic unit in the entire text input information.
[0065] For example, when the text input information is "Ducks in a row" (assuming that the user has not pressed the space bar to indicate that "r", "o", and "w" form a complete word after entering these three characters), the text input information includes three complete words, namely "Ducks", "in", and "a", as well as three input characters, "r", "o", and "w". That is, it contains six semantic units, namely "Ducks", "in", "a", "r", "o", and "w". Table 1 below shows the word position encodings and character position encodings corresponding to each semantic unit of this text input information by way of example.
[0066] Table 1
[0067] The word position encodings of these six semantic units can be [1, 2, 3, 4, 4, 4], indicating that "ducks" is the 1st word in the text input information, "in" is the 2nd word in the text input information, "a" is the 3rd word in the text input information, and "r", "o", and "w" all belong to the 4th word in the text input information (i.e., "r", "o", and "w" are the respective characters of "row").
[0068] The character position encodings of these six semantic units can be [0, 0, 0, 1, 2, 3]. That is, the character position encodings of the complete words "Ducks", "in", and "a" can be set to 0, indicating that this semantic unit is a complete word and not an input character in an incomplete word (of course, this character position encoding can also be a code such as "a" or "b", etc., used to indicate that this semantic unit is a complete word and does not belong to the input characters of an incomplete word). The input characters in the incomplete word are encoded according to their positions in the word. The "1, 2, 3" in [0, 0, 0, 1, 2, 3] indicates that "r", "o", and "w" correspond to the 1st, 2nd, and 3rd characters of the word "row" respectively.
[0069] Then, for the convenience of the neural network model to extract features of the word-level position information and character-level position information of the semantic units, the second embedding layer can perform embedding on the word position encodings of the semantic units to obtain word position features P w ( x m ) ∈ R d , and the third embedding layer can perform embedding on the character position encodings of the semantic units to obtain character position features P c ( x m ) ∈R d Then, based on the character position feature and the lexical position feature of the semantic unit, the corresponding position feature of the semantic unit can be determined. For example, the sum of the character position feature and the lexical position feature of the semantic unit can be directly used as the corresponding position feature of the semantic unit. P m ∈ R d Among them, for the second embedding layer and the third embedding layer, for example, the word embedding method can be referred to for the embedding operations of the lexical position encoding and the character position encoding.
[0070] Then, after the fusion layer determines the second semantic feature, the lexical position feature, and the character position feature of each semantic unit, it can fuse the second semantic feature, the lexical position feature, and the character position feature of the semantic unit to obtain the corresponding third semantic feature. Among them, the second semantic feature, the lexical position feature, and the character position feature of the same semantic unit can be summed to obtain a summation result, and this summation result can be used as the third semantic feature corresponding to this semantic unit. The formula can be specifically expressed as follows: X ( x m ) = E ( x m ) + P m , where P m = P w ( x m ) + P c ( x m ) Among them, X ( x m ) represents the third semantic feature of the semantic unit x m , E ( x m ) represents the second semantic feature of the semantic unit x m , P w ( x m ) represents the lexical position feature of the semantic unit x m , where P w ( x m ) ∈R d , P c ( x m ) represents the character position feature of the semantic unit x m . Among them, P c ( x m ) ∈ R d , P m represents the position feature corresponding to the semantic unit x m . The above formula represents the third semantic feature of a semantic unit, which is the sum result of the corresponding second semantic feature, lexical position feature, and character position feature. And among them, X ( x m ) ∈ R d .
[0071] Thus, the sequence X formed by the third semantic features corresponding to each semantic unit can be shown as follows: X = X ( x 1), X ( x 2), ..., X ( x t-1+k )] ∈ R (t-1+k)*d .
[0072] In subsequent steps, this sequence X is input into the context modeling model of the vocabulary prediction model to obtain the first semantic feature.
[0073] Thus, in this way, the second semantic feature can not only represent the semantic features of each semantic unit, but also represent which word each semantic unit is in the text input information, and represent which character in which word the semantic unit is as a character. Thus, the second semantic feature contains more comprehensive and detailed position information corresponding to the semantic unit. Thus, the vocabulary prediction model can perform more accurate vocabulary prediction based on the second semantic feature.
[0074] Optionally, the operation of determining the first semantic feature according to the third semantic feature includes: inputting the third semantic feature into the context modeling model based on the self-attention mechanism to determine the first semantic feature.
[0075] That is to say, after the computing device determines the third semantic feature that contains both semantic information and multi-level location information, it can input the third semantic feature into the context modeling model based on the self-attention mechanism to determine the first semantic feature for context modeling of the third semantic feature.
[0076] Specifically, the sequence composed of the third semantic features corresponding to the above-mentioned semantic units can be X , input into the context modeling model to obtain the above-mentioned first semantic feature. Among them, the context modeling model can be implemented by a N -layer Transformer network. N can be a numerically preset value by humans. The first semantic feature output by the context modeling model can be expressed as: H = Transformer( X ) ∈ R (t+k)×d .
[0077] Among them, one layer in the used Transformer network can be composed of a masked multi-head self-attention module and a feed-forward neural network. The calculation process of one layer is shown in the following formula:
[0078] Among them, Q = XW Q ; K = XW K ; V = XW V ; M is a mask used to ensure that the prediction tasks completed by the context modeling model and the vocabulary prediction layer are autoregressive tasks.
[0079] As described in the background art, when the n-gram model makes predictions, it calculates the probability of the occurrence of a word sequence with a fixed number of words in the training corpus. Since it only depends on the context of a limited window, it is difficult to capture long-distance semantic associations. According to the technical solution of the present application, the first semantic feature generated by the context modeling model based on the multi-head self-attention mechanism can extract context information within the range of the entire text input information, so that it can more effectively capture long-distance semantic features, and thus has stronger generalization ability. Therefore, making predictions based on the first semantic feature that implies context association information can further improve the accuracy of predictions.
[0080] Optionally, according to the first semantic feature, the operation of generating output information corresponding to the text input information by using the vocabulary prediction layer includes: using the vocabulary prediction layer to determine the probability distribution of candidate words according to the first semantic feature; determining the predicted vocabulary according to the determined probability distribution; and generating output information based on the predicted vocabulary.
[0081] That is to say, the vocabulary prediction model can input the first semantic feature into the vocabulary prediction layer to obtain the predicted vocabulary, and generate output information based on the predicted vocabulary. For example, the output information can directly be the predicted vocabulary.
[0082] Specifically, the vocabulary prediction layer can first determine the fourth semantic feature used to represent the semantic feature of the predicted vocabulary according to the first semantic feature H t+k ∈ R d , where H t+k is the first semantic feature H in the t + k th token corresponding feature. Then, the vocabulary prediction layer can project the fourth semantic feature H t+k to obtain the probability distribution of candidate words P ( y t ), (where the probability distribution can be the probability distribution projected in the vocabulary), as shown in the following formula: P ( y t ) = softmax( H t+k W o + b o ) where W ∈ R d*|v| is the output projection matrix for generating the probability distribution of candidate words. Thus, the vocabulary prediction layer can determine the predicted vocabulary according to P ( y t ). The predicted vocabulary can be the candidate word with the highest probability value, or the candidate words with the probability values ranked in the top n can all be used as the predicted vocabulary for the user to select, n The specific value of can be preset artificially.
[0083] It should be noted that the method for processing text input information in this specification can support three prediction scenarios of the input method, including the word completion scenario mentioned above, the scenario of predicting the next word that the user may input based on several complete words already input by the user (hereinafter referred to as the word prediction scenario), and can also include the scenario of correcting spelling errors for incomplete words that the user is inputting and have spelling mistakes (hereinafter referred to as the spelling correction scenario). Specifically, through three types of training samples, the tasks of the three prediction scenarios can be completed by the same model.
[0084] Optionally, the method further includes: obtaining training samples, where the training samples include text input information samples and annotation information, and the annotation information includes predicted words corresponding to the respective text input information samples. Among them, the training samples include the first type of training samples, the second type of training samples, and the third type of training samples. The text input information samples of the first type of training samples only include complete words, the text input information samples of the second type of training samples contain incomplete words, and the text input information samples of the third type of training samples contain input characters with spelling mistakes; and training the context modeling model and the vocabulary prediction layer according to the training samples.
[0085] The predicted words mentioned here corresponding to the respective text input information samples are: the words that the user actually wants to input marked corresponding to the text input information samples. The training samples include the first type of training samples, the second type of training samples, and the third type of training samples. The first type of training samples correspond to the word prediction scenario, the second type of training samples correspond to the word completion scenario, and the third type of training samples correspond to the spelling correction scenario.
[0086] The following uses examples to illustrate the three prediction scenarios.
[0087] The word prediction scenario is: after the user inputs an incomplete word that is being input, automatically recommend the next most likely word that the user will input based on the context of the user's text input information. For example, when inputting "Li Bai was a poet of the", the system predicts the target word "Tang". This word prediction scenario corresponds to the first type of training samples. If the text input information sample in the first type of training samples is "Li Bai was a poet of the", the corresponding annotation information is "Tang".
[0088] Therefore, whether in the training stage or in the prediction stage, the text input information only contains a sequence of complete words W = w 1, ..., w t , and the goal is to predict w t+1, the training objective is equivalent to maximizing the likelihood of obtaining the annotation information, as shown in the following formula:
[0089] where is the annotation information.
[0090] The word completion scenario is as follows: when the text input information contains an incomplete word that the user is typing, based on the context of the user's input and the partial input characters in the incomplete word, predict the most likely complete word corresponding to the incomplete word. Let the text input information be "You are so b", W = ["You", "are", "so"], and the user is typing the character "b". At this time, C = [b], and the computing device needs to predict the target word "beautiful".
[0091] This word completion scenario corresponds to the second type of training samples. If the text input information sample in the second type of training samples is "You are so b", then the corresponding annotation information is "beautiful", and the training objective is equivalent to maximizing the likelihood of obtaining the annotation information, as shown in the following formula.
[0092]
[0093] where is the annotation information.
[0094] The spelling correction scenario is as follows: when the user makes a spelling mistake in the input (such as the user types "raow" but actually should be "row"), this method can use the context and character distribution to correct the spelling. At this time, the incomplete word input by the user is incorrect, and this method needs to guide towards the most likely target word during the process of generating the prediction result. The spelling correction scenario corresponds to the third type of training samples. If the text input information sample in the first type of training samples is "Ducks in a raow", then the corresponding annotation information is "row".
[0095] This method uses context semantics and deep modeling to automatically correct the text input information with spelling mistakes, and the training objective can be expressed as:
[0096] where is the annotation information, is the set of input characters in the incomplete word with spelling mistakes.
[0097] In addition, the first semantic feature generated by the above context modeling method can also be used to push auxiliary functions to the user. Optionally, the method further includes: according to the first semantic feature, using an auxiliary function judgment model to determine whether to push a preset auxiliary function to the user.
[0098] The above auxiliary function judgment model can be a classification model, specifically a binary classification model. The output of the auxiliary function judgment model can be whether to push a preset auxiliary function to the user. Of course, since there can be multiple types of auxiliary functions, the auxiliary function judgment model can also be a multi-classification model, and the output can include the type of auxiliary function to be pushed to the user and not to push an auxiliary function to the user. Training samples for training the auxiliary function judgment model can be obtained through manual annotation or weak supervision, so as to pre-train the auxiliary function judgment model.
[0099] The pushed auxiliary function can refer to the function that the user needs determined according to the user's text input information. The auxiliary functions mentioned here can include functions such as search intent recognition, intelligent assistant call, and knowledge Q&A jump. For example: Search intent recognition: When the input is like "population of Rome", the computing device can recognize the user's potential intention to query information related to Rome and push the function of searching for "Rome" to the user.
[0100] Intelligent assistant call: When the input is like "remind me to call John at 8", the computing device can push the function of "create a reminder" to the user. If the user agrees, the computing device can call the system schedule interface to create a corresponding reminder for the user.
[0101] Knowledge Q&A jump: When the input is like "who is the president of France", the computing device can push the function of "query the answer using the intelligent assistant" to the user.
[0102] Among them, when the text input information is w 1, w 2, ..., w t-1 , c 1, c 2, ..., c k , the computing device can determine the first semantic feature corresponding to the text input information through the text input information processing method provided in this specification H ∈ R (t+k)*d , and determine the feature representation of the last semantic unit from it h =H t+k Then, input the feature representation into the auxiliary function judgment model to obtain a judgment result (certainly, the first semantic feature corresponding to the text input information can also be input into the auxiliary function judgment model to obtain a judgment result). In the training stage, the auxiliary function judgment model can be trained by minimizing the cross-entropy loss.
[0103] In addition, referring to Figure 1 As shown, according to the second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program runs, the method described in any one of the above is executed by a processor.
[0104] Therefore, according to this embodiment, the accuracy of generating a prediction result for text input information can be improved, and the compatibility with different types of prediction tasks can be improved, thereby improving the applicability.
[0105] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0107] Embodiment 2 Figure 4 The text input information processing device 400 according to the first aspect of this embodiment is shown. The device 400 corresponds to the method according to the first aspect of Embodiment 1. Referring to Figure 4 As shown, the device 400 includes: an acquisition module 410, configured to acquire the text input information of the user, where the text input information includes the complete vocabulary that the user has input and / or the input characters of the incomplete vocabulary that the user is inputting; A feature generation module 420, configured to generate a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is further used to characterize the complete vocabulary and / or the position information of the input characters in the text input information, and the position information includes vocabulary-level position information corresponding to the vocabulary and / or character-level position information corresponding to the characters; An information generation module 430, configured to generate output information corresponding to the text input information according to the first semantic feature by using a vocabulary prediction layer, where the output information includes predicted vocabulary corresponding to the text input information.
[0108] Optionally, when the text input information only includes complete vocabulary, the predicted vocabulary includes the next complete vocabulary relative to the last complete vocabulary in the text input information; and when the text input information includes input characters of an incomplete vocabulary, the predicted vocabulary includes the complete vocabulary corresponding to the input characters, and where the apparatus further includes: an auxiliary function determination module 440, configured to determine whether to push a preset auxiliary function to the user according to the first semantic feature by using an auxiliary function determination model.
[0109] Optionally, the feature generation module 420 is specifically configured to generate a corresponding second semantic feature for each semantic unit in the text input information, where the semantic unit is a single complete vocabulary and / or a single input character in the text input information; determine position features corresponding to each semantic unit, where the position features are used to characterize the vocabulary-level position information and character-level position information of the corresponding semantic unit in the text input information; fuse the second semantic feature with the corresponding position feature to obtain a corresponding third semantic feature; and determine the first semantic feature according to the third semantic feature.
[0110] Optionally, when the semantic unit is a single complete vocabulary, the feature generation module 420 is specifically configured to use the word embedding vector of the single complete vocabulary as the second semantic feature corresponding to the semantic unit; and when the semantic unit is a single input character, use the character embedding vector of the single input character as the second semantic feature corresponding to the semantic unit.
[0111] Optionally, the feature generation module 420 is specifically configured to determine a vocabulary position encoding corresponding to the semantic unit according to the position of the vocabulary corresponding to the semantic unit in the text input information; determine a character position encoding corresponding to the semantic unit, where, when the semantic unit is a character in an incomplete vocabulary, the corresponding character position encoding indicates the character position of the semantic unit in the corresponding vocabulary, and when the semantic unit is a complete vocabulary, the corresponding character position encoding can be used to determine that the semantic unit is a complete vocabulary; and determine the position feature corresponding to the semantic unit according to the vocabulary position encoding and the character position encoding corresponding to the semantic unit.
[0112] Optionally, the feature generation module 420 is specifically configured to input the third semantic feature into a context modeling model based on a self-attention mechanism to determine the first semantic feature. Further, the apparatus further includes: a training module 450, configured to obtain training samples, where the training samples include text input information samples and annotation information, and the annotation information includes predicted words corresponding to the respective text input information samples. The training samples include first-class training samples, second-class training samples, and third-class training samples. The text input information samples of the first-class training samples include only complete words, the text input information samples of the second-class training samples include incomplete words, and the text input information samples of the third-class training samples include input characters with spelling mistakes. And train the context modeling model and the vocabulary prediction layer according to the training samples.
[0113] Optionally, the information generation module 430 is specifically configured to use the vocabulary prediction layer to determine the probability distribution of candidate words according to the first semantic feature; determine the predicted words according to the determined probability distribution; and generate output information based on the predicted words.
[0114] Therefore, according to this embodiment, the accuracy of the predicted result corresponding to the generated text input information can be improved, and the compatibility with different types of prediction tasks can be improved, thereby improving the applicability.
[0115] Embodiment 3 Figure 5 Fig. shows a text input information processing apparatus 500 according to the first aspect of this embodiment. The apparatus 500 corresponds to the method according to the first aspect of Embodiment 1. Refer to Figure 5 As shown, the apparatus 500 includes: a processor 510; and a memory 520, connected to the processor 510, for providing instructions for the processor to perform the following processing steps: obtaining the user's text input information, where the text input information includes complete words that the user has already input and / or input characters of incomplete words that the user is inputting; generating a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is further used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to the words and / or character-level position information corresponding to the characters; generating output information corresponding to the text input information by using the vocabulary prediction layer according to the first semantic feature, where the output information includes predicted words corresponding to the text input information.
[0116] Optionally, when the text input information only includes complete words, the predicted word includes the next complete word relative to the last complete word in the text input information; and when the text input information includes input characters of an incomplete word, the predicted word includes the complete word corresponding to the input characters.
[0117] Optionally, the operation of generating a first semantic feature for characterizing the semantics of the text input information according to the text input information includes: for each semantic unit in the text input information, generating a corresponding second semantic feature, where the semantic unit is a single complete word and / or a single input character in the text input information; determining a position feature corresponding to each semantic unit, where the position feature is used to characterize the word-level position information and character-level position information of the corresponding semantic unit in the text input information; fusing the second semantic feature with the corresponding position feature to obtain a corresponding third semantic feature; and determining the first semantic feature according to the third semantic feature.
[0118] Optionally, the operation of generating a corresponding second semantic feature for each semantic unit in the text input information includes: when the semantic unit is a single complete word, using the word embedding vector of the single complete word as the second semantic feature corresponding to the semantic unit; and when the semantic unit is a single input character, using the character embedding vector of the single input character as the second semantic feature corresponding to the semantic unit.
[0119] Optionally, the operation of determining the position feature corresponding to each semantic unit specifically includes: according to the position of the word corresponding to the semantic unit in the text input information, determining the word position encoding corresponding to the semantic unit; determining the character position encoding corresponding to the semantic unit, where, when the semantic unit is a character in an incomplete word, the corresponding character position encoding indicates the character position of the semantic unit in the corresponding word, and when the semantic unit is a complete word, the corresponding character position encoding can be used to determine that the semantic unit is a complete word; and determining the position feature corresponding to the semantic unit according to the word position encoding and character position encoding corresponding to the semantic unit.
[0120] Optionally, the operation of determining the first semantic feature according to the third semantic feature includes: inputting the third semantic feature into a context modeling model based on a self-attention mechanism to determine the first semantic feature.
[0121] Optionally, the operation of determining the output information corresponding to the text input information using a vocabulary prediction layer according to the first semantic feature specifically includes: using the vocabulary prediction layer to determine the probability distribution of candidate words according to the first semantic feature; determining the predicted word according to the determined probability distribution; and generating the output information based on the predicted word.
[0122] Optionally, the method further includes: obtaining training samples, where the training samples include text input information samples and annotation information, and the annotation information includes predicted words corresponding to the respective text input information samples. Among them, the training samples include first-class training samples, second-class training samples, and third-class training samples. The text input information samples of the first-class training samples only include complete words, the text input information samples of the second-class training samples contain incomplete words, and the text input information samples of the third-class training samples contain input characters with spelling mistakes; and training the context modeling model and the vocabulary prediction layer according to the training samples.
[0123] Optionally, the method further includes: determining, according to the first semantic feature and using the auxiliary function judgment model, whether to push a preset auxiliary function to the user.
[0124] Therefore, according to this embodiment, the accuracy of the predicted result corresponding to the generated text input information can be improved, and the compatibility with different types of prediction tasks can be improved, thereby improving the applicability.
[0125] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0126] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0127] In the several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0128] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0129] In addition, in each embodiment of the present invention, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0130] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.
[0131] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for processing text input information, characterized in that Including: Obtaining the text input information of the user, where the text input information includes the complete words that the user has input and / or the input characters of the incomplete words that the user is inputting; Generating a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is further used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to the words and character-level position information corresponding to the characters; And Generating output information corresponding to the text input information by using a vocabulary prediction layer according to the first semantic feature, where the output information includes predicted words corresponding to the text input information.
2. The method according to claim 1, wherein In the case where the text input information only includes complete words, the predicted words include the next complete word relative to the last complete word in the text input information; and in the case where the text input information includes the input characters of the incomplete words, the predicted words include the complete words corresponding to the input characters, and where The method further includes: determining whether to push a preset auxiliary function to the user according to the first semantic feature by using an auxiliary function judgment model.
3. The method according to claim 1, characterized in that The operation of generating a first semantic feature for characterizing the semantics of the text input information according to the text input information includes: Generating corresponding second semantic features for each semantic unit in the text input information, where the semantic unit is a single complete word and / or a single input character of the text input information; Determining position features corresponding to each semantic unit, where the position features are used to characterize the word-level position information and character-level position information of the corresponding semantic unit in the text input information; Fusing the second semantic feature with the corresponding position feature to obtain a corresponding third semantic feature; and Determining the first semantic feature according to the third semantic feature.
4. The method according to claim 3, characterized in that, The operation of generating corresponding second semantic features for each semantic unit in the text input information includes: In the case where the semantic unit is a single complete word, using the word embedding vector of the single complete word as the second semantic feature corresponding to the semantic unit; and In the case where the semantic unit is a single input character, using the character embedding vector of the single input character as the second semantic feature corresponding to the semantic unit.
5. The method according to claim 3, characterized in that, The operation of determining position features corresponding to each semantic unit specifically includes: Determining a word position encoding corresponding to the semantic unit according to the position of the word corresponding to the semantic unit in the text input information; Determining a character position encoding corresponding to the semantic unit, where in the case where the semantic unit is a character in an incomplete word, the corresponding character position encoding indicates the character position of the semantic unit in the corresponding word, and in the case where the semantic unit is a complete word, the corresponding character position encoding can be used to determine that the semantic unit is a complete word; and Determine the position feature corresponding to the semantic unit according to the lexical position encoding and character position encoding corresponding to the semantic unit.
6. The method according to claim 3, characterized in that, The operation of determining the first semantic feature according to the third semantic feature includes: Input the third semantic feature into a context modeling model based on a self-attention mechanism to determine the first semantic feature, and wherein the method further includes: Obtain training samples, where the training samples include text input information samples and annotation information, and the annotation information includes predicted words corresponding to the respective text input information samples, where The training samples include first-class training samples, second-class training samples, and third-class training samples. The text input information samples of the first-class training samples only include complete words, the text input information samples of the second-class training samples contain incomplete words, and the text input information samples of the third-class training samples contain input characters with spelling mistakes; and Train the context modeling model and the vocabulary prediction layer according to the training samples.
7. The method according to claim 1, characterized in that, The operation of generating output information corresponding to the text input information by using the vocabulary prediction layer according to the first semantic feature specifically includes: Use the vocabulary prediction layer to determine the probability distribution of candidate words according to the first semantic feature; Determine the predicted word according to the determined probability distribution; and Generate the output information based on the predicted word.
8. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, the method according to any one of claims 1 to 7 is executed by a processor.
9. A text input information processing device, characterized in that, Comprising: An acquisition module, configured to acquire the user's text input information, where the text input information includes the complete words already input by the user and / or the input characters of the incomplete words being input by the user; A feature generation module, configured to generate a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is further used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to words and / or character-level position information corresponding to characters; And An information generation module, configured to generate output information corresponding to the text input information by using a vocabulary prediction layer according to the first semantic feature, where the output information includes predicted words corresponding to the text input information.
10. A text input information processing device, characterized in that, Comprising: A processor; And A memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: Acquire the user's text input information, where the text input information includes the complete words already input by the user and / or the input characters of the incomplete words being input by the user; Generate a first semantic feature for characterizing the semantics of the text input information according to the text input information, where the first semantic feature is further used to characterize the position information of the complete words and / or input characters in the text input information, and the position information includes word-level position information corresponding to words and / or character-level position information corresponding to characters; According to the first semantic feature, an output information corresponding to the text input information is generated by using a lexical prediction layer, where the output information includes predicted vocabulary corresponding to the text input information.
Citation Information
Patent Citations
Efficient input prediction method and device
CN104102720A
Smart association method and device for input method
CN104133855A
Text rewriting method, electronic equipment and storage device
CN112668343A
Character input method and device, electronic equipment and medium
CN114356118A
Chinese address element analysis method and device based on vocabulary enhancement and storage medium
CN114792091A