Information completion method and information completion model training method and device

By using the two-way decoded double tower information completion deep learning model, the prefix and suffix information are encoded and candidate tokens in two directions are generated, and the direction with higher certainty is selected for information completion, which solves the problem of poor semantic prediction performance in the existing technology and improves the accuracy of information completion.

CN120295613APending Publication Date: 2025-07-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337491.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing deep learning models have poor semantic prediction performance in information completion tasks, especially the one-way decoding model cannot effectively utilize contextual semantics, resulting in insufficient completion accuracy.

Method used

The two-way decoding double tower information completion deep learning model is adopted. By encoding and decoding the prefix and suffix information separately, candidate tokens and their probability are generated in two directions, and directions with higher certainty are selected for information completion.

Benefits of technology

The semantic prediction accuracy of the information completion model is improved and the applicability in the information completion scenario is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295613A_ABST
    Figure CN120295613A_ABST
Patent Text Reader

Abstract

The invention discloses an information completion method and an information completion model training method and device, relates to the technical field of artificial intelligence such as natural language processing, large models and deep learning, and can be applied to scenes such as code completion. The specific implementation scheme is as follows: obtaining prefix information and suffix information of to-be-complemented information from original input information; based on the prefix information and the suffix information, performing completion prediction on the to-be-completed information by adopting a preset large model to obtain a first candidate token completed in a first direction and a probability thereof, and a second candidate token completed in a second direction and a probability thereof; determining a target complementation direction in the first direction and the second direction according to the probability of the first candidate token and the probability of the second candidate token; and filling the candidate token corresponding to the target completion direction in the first candidate token and the second candidate token along the target completion direction. According to the method, the model performance of information completion can be improved, so that the information completion model has better adaptability and higher semantic prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technologies such as natural language processing, large models, and deep learning, and particularly relates to an information completion method, an information completion model training method, and a device. Background Art

[0002] Information completion is an important task in natural language processing. Its purpose is to predict the next unappeared information part based on the existing information content, so as to obtain some key information from the existing information content to complete the current information content.

[0003] In related technologies, deep learning models are usually used for information completion. However, the deep learning models for information completion in related technologies are usually unidirectional decoding, resulting in poor semantic prediction performance of the models. Summary of the Invention

[0004] The present disclosure provides an information completion method, an information completion model training method, a device, an intelligent agent, an electronic device, and a storage medium.

[0005] According to a first aspect of the present disclosure, an information completion method is provided, including:

[0006] Obtain prefix information and suffix information of the information to be completed from the original input information;

[0007] Based on the prefix information and the suffix information, use a preset information completion model to perform completion prediction on the information to be completed, and obtain a first candidate token and the probability of the first candidate token for completion in a first direction, a second candidate token and the probability of the second candidate token for completion in a second direction; the structure of the information completion model is a two-tower information completion deep learning model structure with bidirectional decoding;

[0008] Determine a target completion direction between the first direction and the second direction according to the probability of the first candidate token and the probability of the second candidate token;

[0009] Fill the candidate token corresponding to the target completion direction among the first candidate token and the second candidate token along the target completion direction.

[0010] According to a second aspect of the present disclosure, an information completion model training method is provided, including:

[0011] Obtain training sample data;

[0012] Obtain prefix information and suffix information from the training sample data;

[0013] Input the prefix information and the suffix information into the information completion model to perform completion prediction on the training sample data, and obtain the first candidate token completed in the first direction and the second candidate token completed in the second direction;

[0014] Determine the model loss value according to the true token behind the prefix information, the true token in front of the suffix information, the first candidate token, and the second candidate token in the training sample data;

[0015] Train the information completion model according to the model loss value.

[0016] According to the third aspect of the present disclosure, there is provided an information completion device, including:

[0017] An acquisition module, configured to acquire the prefix information and the suffix information of the information to be completed from the original input information;

[0018] A prediction module, configured to perform completion prediction on the information to be completed based on the prefix information and the suffix information by using a preset information completion model, and obtain the first candidate token completed in the first direction, the probability of the first candidate token, the second candidate token completed in the second direction, and the probability of the second candidate token; the structure of the information completion model is a two-tower information completion deep learning model structure with bidirectional decoding;

[0019] A selection module, configured to determine the target completion direction in the first direction and the second direction according to the probability of the first candidate token and the probability of the second candidate token;

[0020] A filling module, configured to fill the candidate token corresponding to the target completion direction among the first candidate token and the second candidate token along the target completion direction.

[0021] According to the fourth aspect of the present disclosure, there is provided an information completion model training device, including:

[0022] A first acquisition module, configured to acquire training sample data;

[0023] A second acquisition module, configured to acquire the prefix information and the suffix information from the training sample data;

[0024] A prediction module, configured to input the prefix information and the suffix information into the information completion model to perform completion prediction on the training sample data, and obtain the first candidate token completed in the first direction and the second candidate token completed in the second direction; the structure of the information completion model is a two-tower information completion deep learning model structure with bidirectional decoding;

[0025] A determination module, configured to determine a model loss value according to real tokens behind the prefix information, real tokens in front of the suffix information, a first candidate token, and a second candidate token in training sample data;

[0026] A training module, configured to train an information completion model according to the model loss value.

[0027] According to a fifth aspect of the present disclosure, there is provided an agent, including:

[0028] An input module, configured to receive input information;

[0029] A processing module, configured to determine a target task based on the input information received by the input module, determine an information completion model based on the target task, and execute the methods described in the foregoing first aspect and second aspect by calling the information completion model to obtain output information;

[0030] An output module, configured to output the output information obtained by the processing module.

[0031] According to a sixth aspect of the present disclosure, there is provided an electronic device, including:

[0032] At least one processor; and

[0033] A memory communicatively connected to the at least one processor; wherein,

[0034] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the methods described in the foregoing first aspect and second aspect.

[0035] According to a seventh aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the methods described in the foregoing first aspect and second aspect.

[0036] According to an eighth aspect of the present disclosure, there is provided a program product, including at least one of a program and instructions, wherein when the at least one of the program and instructions is executed by a processor, steps of the methods described in the foregoing first aspect and second aspect are implemented.

[0037] The technology according to the present application improves the semantic prediction accuracy of the model, so that it can have higher applicability in the information completion scenario.

[0038] It should be understood that the content described in this part is not intended to identify key or important features of embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0039] The drawings are used to better understand the present solution and do not constitute a limitation to this application. Among them:

[0040] Figure 1 is a flowchart of the information completion method provided by an embodiment of the present disclosure;

[0041] Figure 2 is a schematic structural diagram of the information completion model provided by an embodiment of the present disclosure;

[0042] Figure 3 is a flowchart of the information completion method provided by an embodiment of the present disclosure;

[0043] Figure 4 is a flowchart of the information completion model training method provided by an embodiment of the present disclosure;

[0044] Figure 5 is a block diagram of the information completion device provided by an embodiment of the present disclosure;

[0045] Figure 6 is a block diagram of the information completion model training device provided by an embodiment of the present disclosure;

[0046] Figure 7 is a block diagram of the intelligent agent provided by an embodiment of the present disclosure;

[0047] Figure 8 is a block diagram of the electronic device used to implement the embodiments of this application. Detailed Embodiments

[0048] The following describes exemplary embodiments of the present application in conjunction with the drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0049] Embodiments of the present disclosure relate to the fields of artificial intelligence technologies such as natural language processing, large models, and deep learning.

[0050] Artificial Intelligence, abbreviated as AI in English, is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0051] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. It is a discipline that takes language as the object and uses computer technology to analyze, understand, and process natural language, that is, using the computer as a powerful tool for language research, quantitatively studying language information with the support of the computer, and providing a language description that can be commonly used between humans and computers.

[0052] Large Language Model (LLM for short, also called large model) refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. Large Language Models can handle various natural language tasks, such as text classification, question answering, dialogue, etc., and are an important way to artificial intelligence.

[0053] Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds.

[0054] An Agent is an entity that can perceive the environment and take actions to achieve specific goals. It can be software, hardware, or a system, and has autonomy, adaptability, and interaction capabilities. The Agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then executes actions to affect the environment or achieve a predetermined goal.

[0055] It should be noted that in the technical solutions of this disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0056] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0057] It should also be noted that in the embodiments of the present disclosure, some existing industry solutions such as certain software, components, models, etc. may be mentioned. They should be considered exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present disclosure, but it does not mean that the applicant has already or necessarily used this solution.

[0058] In the model information completion technology in the related art, the prefix and suffix are mainly distinguished by flags, and the prefix and suffix are uniformly spliced together to unidirectionally predict the information content in the middle part from front to back. This model structure has the following problems: 1) unified feature encoding without distinguishing the semantic differences of the context; 2) unidirectional decoding, which can only infer the completed information from front to back. In fact, the certainty of some semantics is higher when supplemented from back to front, and the effect is better.

[0059] Based on this, the embodiments of the present disclosure provide an information completion method, an information completion model training method, and a device. By using a two-tower information completion deep learning model with bidirectional decoding to perform completion prediction on the input, output results of completion in two directions can be obtained. By selecting the direction with higher certainty based on the certainty of the semantics in the two directions for information completion, the semantic prediction accuracy of the model can be improved, so that it can have higher applicability in the information completion scenario.

[0060] Among them, it should be noted that the execution subject of the information completion method in the embodiments of the present disclosure can be an information completion device. This device can be implemented in software and / or hardware, and this device can be configured in an electronic device, and the electronic device can include but are not limited to terminals, server sides, etc.

[0061] The information completion method, the information completion model training method, and the device in the embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0062] Figure 1 It is a flowchart of the information completion method provided by the embodiments of the present disclosure. As Figure 1 shown, the information completion method may include but is not limited to the following steps.

[0063] In step 101, the prefix information and the suffix information of the information to be completed are obtained from the original input information.

[0064] In some embodiments, the original input information may be information input by the user through a terminal. Exemplarily, a user interaction interface is provided for the user, and there is an input box in the user interaction interface, and the user can perform input operations in the input box.

[0065] In some embodiments, the information to be completed may be the code to be completed, or may also be the text to be completed, but is not limited thereto. That is to say, the technical solution provided by the embodiments of the present application can be applied to the code completion scenario, or can also be applied to the text completion scenario, but is not limited thereto. For example, it can also be applied to other completion scenarios where the middle content is completed with known prefixes and suffixes.

[0066] In some embodiments, the original input information may include a first flag for distinguishing the prefix and a second flag for distinguishing the suffix. In the embodiments of the present application, the prefix information of the information to be completed can be obtained from the original input information through the first flag, and the suffix information of the information to be completed can be obtained from the original input information through the second flag.

[0067] In some embodiments, the information completion method involved in the embodiments of the present application can be implemented based on an information completion model. Exemplarily, the information completion model can be a large model, but is not limited thereto. For example, taking the information completion model as a large model and the completion service type to which the information to be completed belongs as the text completion service as an example, assuming that the user input information is "Please complete the text between the two words 'X city' and'scenic spot' to make it a complete sentence", when receiving this input information, the information to be completed can be determined from this input information based on semantic analysis as the text content between "X city" and "scenic spot", the prefix information is "X city", and the suffix information is "scenic spot".

[0068] Another example is that taking the completion service type to which the information to be completed belongs as the code completion service as an example, assuming that the user input information is a code snippet, the prefix code information can be obtained from the input code snippet based on the code prefix flag, and the suffix code information can be obtained from the input code snippet based on the code suffix flag.

[0069] In step 102, based on the prefix information and the suffix information, the information to be completed is complemented and predicted by using a preset information completion model to obtain the first candidate token and the probability of the first candidate token for the first-direction completion, and the second candidate token and the probability of the second candidate token for the second-direction completion.

[0070] In some embodiments, the network structure of the above information completion model may be a two-tower information completion deep learning model structure with bidirectional decoding. In some embodiments, as Figure 2 shown, the information completion model may include, but is not limited to, a prefix encoding unit, a suffix encoding unit, and a bidirectional decoding unit. Among them, the network structures of the prefix encoding unit and the suffix encoding unit can both be the encoders in the Transformer model, and the bidirectional decoding unit can be the decoder in the Transformer model, and the output layer of the decoder is changed to a bidirectional output layer.

[0071] Exemplarily, the bidirectional output layer can be understood as including two output layers, such as a first output layer and a second output layer. The first output layer is used to output the first candidate token completed from the first direction and the probability of the first candidate token. The second output layer is used to output the second candidate token completed from the second direction and the probability of the second candidate token. For example, the bidirectional output layer can include two softmax layers. One softmax layer is used to calculate the probability of the first candidate token, such as expressed as probabilities prefix = softmax(output prefix ); Another softmax layer is used to calculate the probability of the second candidate token, such as expressed as probabilities suffix = softmax(output suffix ).

[0072] In some embodiments, the above-mentioned first direction can be understood as the direction of completing from front to back; the second direction can be understood as the direction of completing from back to front. Alternatively, in some embodiments, the first direction can be understood as the direction of completing from back to front, and the second direction can be understood as the direction of completing from front to back.

[0073] In the embodiments of the present application, the prefix information and the suffix information can be input into the information completion model. The dual-tower structure in the information completion model performs forward and backward differential encoding on the input prefix information and suffix information, that is, encodes the prefix information and the suffix information separately. Exemplarily, forward encoding is performed on the prefix information to obtain prefix features, backward encoding is performed on the suffix information to obtain suffix features, and the prefix features and the suffix features are decoded through the bidirectional decoding mechanism in the information completion model and the first candidate token completed in the first direction and its probability, and the second candidate token completed in the second direction and its probability are output. Exemplarily, the forward encoding can be used to encode the prefix information and calculate the prefix depth parsing into prefix features; the backward encoding can be used to encode the suffix information and calculate the suffix depth parsing into suffix features. The bidirectional decoding refers to the depth decoding of comprehensively calculating the prefix features and the suffix features and outputting the probabilities of the two tokens completed in the forward and backward directions.

[0074] In step 103, according to the probability of the first candidate token and the probability of the second candidate token, determine the target completion direction in the first direction and the second direction.

[0075] In an embodiment of the present application, the completion direction can be selected from the first direction and the second direction according to the magnitudes of the probabilities of the first candidate token and the second candidate token, and the selected completion direction is determined as the target completion direction.

[0076] In some embodiments, the magnitudes of the probabilities of the first candidate token and the second candidate token can be compared; in the case where the probability of the first candidate token is greater than the probability of the second candidate token, the above-mentioned first direction is determined as the target completion direction; or, in the case where the probability of the second candidate token is greater than the probability of the first candidate token, the above-mentioned second direction is determined as the target completion direction. For example, taking the first direction as the direction of completion from front to back and the second direction as the direction of completion from back to front as an example, if the probability of the first candidate token is greater than the probability of the second candidate token, the direction of completion from front to back (which can also be called the forward completion direction) can be determined as the target completion direction; if the probability of the second candidate token is greater than the probability of the first candidate token, the direction of completion from back to front (which can also be called the backward completion direction) can be determined as the target completion direction.

[0077] That is to say, the direction with higher certainty among the certainties of the two-way semantics can be selected as the target completion direction according to the magnitudes of the probability of the first candidate token completed in the first direction and the probability of the second candidate token completed in the second direction.

[0078] Optionally, in some embodiments, if the probability of the first candidate token is equal to the probability of the second candidate token, the first direction and / or the second direction can be determined as the target completion direction. Exemplarily, if the probability of the first candidate token is equal to the probability of the second candidate token, any one of the first direction and the second direction can be determined as the target completion direction, such as determining the first direction as the target completion direction, or alternatively, the second direction can also be determined as the target completion direction. Or, exemplarily, if the probability of the first candidate token is equal to the probability of the second candidate token, both the first direction and the second direction can be determined as the target completion direction. This can further improve the prediction efficiency of the model.

[0079] In step 104, the candidate token corresponding to the target completion direction among the first candidate token and the second candidate token is filled along the target completion direction.

[0080] In an embodiment of the present application, after determining the target completion direction, the candidate token corresponding to the target completion direction can be filled along the target completion direction.

[0081] In the above embodiments, by using the information completion deep learning model with two-way decoding, i.e., the information completion model, to perform completion prediction on the input, output results of completion in two directions can be obtained. By selecting the direction with higher certainty based on the certainty of semantics in the two directions for information completion, the semantic prediction accuracy of the model can be improved, thereby enabling higher applicability in the information completion scenario.

[0082] In some embodiments, when the target completion direction is the first direction, the first candidate token can be filled along the first direction. Exemplarily, taking the first direction as the direction of completion from front to back as an example, if the target completion direction is the first direction, the first candidate token can be filled in the to-be-filled token position behind the prefix information along the first direction. For example, taking the prefix information as "X City" and the suffix information as "scenic spots" as an example, assuming the first candidate token is "There are", the second candidate token is "a", and the target completion direction is the first direction (such as the direction of completion from front to back), the first candidate token "There are" can be filled behind the prefix information "X City" to obtain the information "X City There are".

[0083] In some embodiments, when the target completion direction is the second direction, the second candidate token can be filled along the second direction. Exemplarily, taking the second direction as the direction of completion from back to front as an example, if the target completion direction is the second direction, the second candidate token can be filled in the to-be-filled token position in front of the suffix information along the second direction. For example, taking the prefix information as "X City" and the suffix information as "scenic spots" as an example, assuming the first candidate token is "There are", the second candidate token is "natural wonders", and the target completion direction is the second direction (such as the direction of completion from back to front), the second candidate token "natural wonders" can be filled in front of the suffix information "scenic spots" to obtain the information "natural wonders scenic spots".

[0084] In some embodiments, when the target completion direction is the first direction and the second direction, the first candidate token can be filled along the first direction, and the second candidate token can be filled along the second direction. Exemplarily, taking the first direction as the direction of forward completion and the second direction as the direction of backward completion as an example, if the target completion direction is the first direction and the second direction, the first candidate token can be filled in the token position to be filled behind the prefix information along the first direction, and the second candidate token can be filled in the token position to be filled in front of the suffix information along the second direction. For example, taking the prefix information as "X City" and the suffix information as "scenic spots" as an example, assuming the first candidate token is "There is" and the second candidate token is "natural wonders", and the target completion direction is the first direction (such as the direction of forward completion) and the second direction (such as the direction of backward completion), the first candidate token "There is" can be filled behind the prefix information "X City", and the second candidate token "natural wonders" can be filled in front of the suffix information "scenic spots", and the obtained information is "X City There is natural wonders scenic spots".

[0085] It should be noted that in the process of information completion based on the information completion model, if the completion prediction end condition is met, the information completion model can stop generating more information. If the completion prediction end condition is not met, the information completion model can continue the next prediction, that is, predict the next token position to be filled to generate more content.

[0086] It is worth noting that when making the next prediction, it is necessary to update the prefix information and / or suffix information input to the model to facilitate the completion prediction based on the updated input. In some embodiments, when the target completion direction is the first direction, after filling the first candidate token into the corresponding token position to be filled, the suffix information can be kept unchanged, and the combination of the prefix information and the filled first candidate token can be used as the new prefix information, and the next token position to be filled is predicted based on the new prefix information and the suffix information. Exemplarily, the new prefix information and the suffix information are input into the information completion model for completion prediction.

[0087] In some embodiments, when the target completion direction is the second direction, after filling the second candidate token into the corresponding token position to be filled, the prefix information can be kept unchanged, and the combination of the suffix information and the filled second candidate token can be used as the new suffix information, and the next token position to be filled is predicted based on the prefix information and the new suffix information. Exemplarily, the new suffix information and the prefix information are input into the information completion model for completion prediction.

[0088] In some embodiments, when the target completion directions are the first direction and the second direction, after the first candidate token and the second candidate token are respectively filled into the corresponding tokens to be filled, the combination of the prefix information and the filled first candidate token can be used as the new prefix information, and the combination of the suffix information and the filled second candidate token can be used as the new suffix information. Then, based on the new prefix information and the new suffix information, the prediction of the next token to be filled is performed. Exemplarily, the new prefix information and the new suffix information are input into the information completion model for completion prediction.

[0089] If the completion prediction end condition is not met, the information completion model can continue the next prediction. Until the completion prediction end condition is met, the information completion model can stop generating more information. Exemplarily, the completion prediction end condition can include at least one of the following: the generated information to be completed reaches a preset maximum length; a specific end symbol is encountered; a preset number of generation rounds is reached; a reasonable end condition is detected (for example, in code completion, when the information completion model determines that the definition of a function or method has been completed, it will stop generating more code).

[0090] Optionally, in some embodiments, after the tokens to be filled between the prefix information and the suffix information are filled, the tokens in the filled information to be completed that meet the elimination condition are eliminated; wherein, the elimination condition is determined based on pre-configuration and / or the original input information. Exemplarily, when the user inputs information, the elimination condition can be set in the original input information, or the elimination condition can also be pre-configured, such as characters or symbols to be eliminated. Based on this, after the tokens to be filled between the prefix information and the suffix information are filled, the tokens in the filled information to be completed that meet the elimination condition can be eliminated. This can improve the information completion effect and the accuracy of information completion.

[0091] Figure 3 This is a flowchart of the information completion method provided by the embodiments of the present disclosure. On the basis of Figure 1 as shown in Figure 3 the above-mentioned method for performing completion prediction on the information to be completed by using a preset information completion model based on prefix information and suffix information to obtain the first candidate token and the probability of the first candidate token for completion in the first direction, and the second candidate token and the probability of the second candidate token for completion in the second direction, a possible implementation manner includes the following steps.

[0092] In step 301, the prefix information is encoded by the prefix encoding unit in the information completion model to obtain a prefix feature.

[0093] Exemplarily, the network structure of the prefix encoding unit can be the encoder in the Transformer model. This encoder can process the prefix information to obtain the embedded representation of the prefix information, and use the self-attention mechanism to extract features from the embedded representation of the prefix information, thereby obtaining prefix features. Among them, this encoder can be composed of multiple layers, and each layer has two main parts: one is the multi-head self-attention mechanism, which allows the model to consider other words in the sequence while processing a word and capture the relationships between them; the other is the feed forward neural network, which can further process the output of the attention layer.

[0094] In step 302, the suffix information is encoded based on the suffix encoding unit in the information completion model to obtain suffix features.

[0095] Exemplarily, the network structure of the suffix encoding unit can be the encoder in the Transformer model. This encoder can process the suffix information to obtain the embedded representation of the suffix information, and use the self-attention mechanism to extract features from the embedded representation of the suffix information, thereby obtaining suffix features. The structure of this encoder is the same as that of the encoder described in step 301 above and will not be elaborated here.

[0096] In step 303, the prefix features and suffix features are fused based on the bidirectional decoding unit in the information completion model to obtain prefix-suffix fusion features.

[0097] Exemplarily, the bidirectional decoding unit can be the decoder in the Transformer model, and the output layer of this decoder is changed to a bidirectional output layer. This bidirectional decoding unit can fuse the prefix features and suffix features.

[0098] In step 304, the prefix-suffix fusion features are input into the bidirectional output layer in the bidirectional decoding unit to obtain the first candidate token and its probability for the first direction completion, and the second candidate token and its probability for the second direction completion.

[0099] Exemplarily, this bidirectional output layer can include two softmax layers. One softmax layer is used to calculate the probability of the first candidate token, such as expressed as probabilities prefix = softmax(output prefix ); the other softmax layer is used to calculate the probability of the second candidate token, such as expressed as probabilitiessuffix = softmax(output suffix ), two output results can be obtained using this bidirectional output layer. One output result is the first candidate token completed in the first direction and the probability of the first candidate token, and the other output result is the second candidate token completed in the second direction and the probability of the second candidate token.

[0100] Thus, by using a dual-tower information completion deep learning model with bidirectional decoding to perform completion prediction on the input, feature encoding can be separately performed on the prefix information and the suffix information, and two output results completed in two directions can be obtained in the decoding stage. By selecting the direction with higher certainty through the certainty of semantics in two directions for information completion, the semantic prediction accuracy of the model can be improved.

[0101] Figure 4 is a flowchart of the information completion model training method provided by an embodiment of the present disclosure. As Figure 4 shown, the information completion model training method may include but is not limited to the following steps.

[0102] In step 401, training sample data is obtained.

[0103] Exemplarily, the training sample data may be text samples, or may also be code samples, but is not limited thereto.

[0104] In step 402, prefix information and suffix information are obtained from the training sample data.

[0105] Exemplarily, for the optional implementation manner of obtaining prefix information and suffix information from the training sample data, reference may be made to the description of the relevant embodiments in step 101 in Figure 1 , which will not be elaborated here.

[0106] In step 403, the prefix information and the suffix information are input into the information completion model to perform completion prediction on the training sample data, and a first candidate token completed in the first direction and a second candidate token completed in the second direction are obtained.

[0107] Exemplarily, the structure and function of the information completion model can be referred to the description of the relevant embodiments in step 102 in the above Figure 1 , which will not be elaborated here.

[0108] In step 404, according to the real token behind the prefix information, the real token in front of the suffix information, the first candidate token, and the second candidate token in the training sample data, the model loss value is determined.

[0109] In some embodiments, the first loss value may be determined based on the true tokens and the first candidate tokens following the prefix information in the training sample data; the second loss value may be determined based on the true tokens and the second candidate tokens preceding the suffix information in the training sample data; and the model loss value may be determined based on the first loss value and the second loss value. Exemplarily, a preset loss function may be used to calculate the model loss value. For example, the loss function may be a cross-entropy loss function, but it is not limited thereto. Exemplarily, when the first loss value and the second loss value are obtained, the first loss value and the second loss value may be added or weighted and summed to obtain the model loss value.

[0110] In step 405, the information completion model is trained according to the model loss value.

[0111] Exemplarily, the parameters of the information completion model may be adjusted according to the model loss value until the training end condition is satisfied. The training end condition includes that the model loss value is less than a preset loss threshold, or it may also include that the number of training iterations reaches a preset number, etc.

[0112] In the above embodiments, by training the dual-tower information completion deep learning model with bidirectional decoding, the model can separately perform feature encoding on the prefix information and the suffix information, and two-directionally completed output results can be obtained in the decoding stage. By using the certainty of the semantics in two directions to select the direction with higher certainty for information completion, the semantic prediction accuracy of the model can be improved, so that it can have higher applicability in the information completion scenario.

[0113] Figure 5 is a block diagram of the information completion device provided by the embodiments of the present disclosure. As Figure 5 shown, the information completion device may include an acquisition module 501, a prediction module 502, a selection module 503, and a filling module 504.

[0114] Among them, the acquisition module 501 is configured to acquire the prefix information and the suffix information of the information to be completed from the original input information.

[0115] The prediction module 502 is configured to perform completion prediction on the information to be completed based on the prefix information and the suffix information by using a preset information completion model, and obtain the first candidate token and the probability of the first candidate token for the first-direction completion, and the second candidate token and the probability of the second candidate token for the second-direction completion; the structure of the information completion model is a dual-tower information completion deep learning model structure with bidirectional decoding.

[0116] A selection module 503, configured to determine a target completion direction between a first direction and a second direction according to the probability of a first candidate token and the probability of a second candidate token.

[0117] A padding module 504, configured to pad the candidate token corresponding to the target completion direction among the first candidate token and the second candidate token along the target completion direction.

[0118] In some embodiments, the prediction module 502 is configured to: encode prefix information by a prefix encoding unit in the information completion model to obtain a prefix feature; encode suffix information by a suffix encoding unit in the information completion model to obtain a suffix feature; fuse the prefix feature and the suffix feature by a bidirectional decoding unit in the information completion model to obtain a prefix-suffix fusion feature; input the prefix-suffix fusion feature into a bidirectional output layer in the bidirectional decoding unit to obtain a first candidate token and the probability of the first candidate token for completion in the first direction, and a second candidate token and the probability of the second candidate token for completion in the second direction.

[0119] In some embodiments, the selection module 503 is configured to: compare the probability of the first candidate token with the probability of the second candidate token; when the probability of the first candidate token is greater than the probability of the second candidate token, determine the first direction as the target completion direction; or, when the probability of the second candidate token is greater than the probability of the first candidate token, determine the second direction as the target completion direction; or, when the probability of the first candidate token is equal to the probability of the second candidate token, determine the first direction and / or the second direction as the target completion direction.

[0120] In some embodiments, the padding module 504 is configured to: when the target completion direction is the first direction, pad the token position to be padded behind the prefix information with the first candidate token along the first direction; or, when the target completion direction is the second direction, pad the token position to be padded in front of the suffix information with the second candidate token along the second direction; or, when the target completion direction is the first direction and the second direction, pad the token position to be padded behind the prefix information with the first candidate token along the first direction, and pad the token position to be padded in front of the suffix information with the second candidate token along the second direction.

[0121] In some embodiments, the prediction module 502 is further configured to: when the target completion direction is the first direction, after filling the first candidate token into the corresponding token position to be filled, keep the suffix information unchanged, and use the combination of the prefix information and the filled first candidate token as the new prefix information, and predict the next token position to be filled based on the new prefix information and the suffix information; or, when the target completion direction is the second direction, after filling the second candidate token into the corresponding token position to be filled, keep the prefix information unchanged, and use the combination of the suffix information and the filled second candidate token as the new suffix information, and predict the next token position to be filled based on the prefix information and the new suffix information; or, when the target completion direction is the first direction and the second direction, after filling the first candidate token and the second candidate token into the corresponding token positions to be filled respectively, use the combination of the prefix information and the filled first candidate token as the new prefix information, and use the combination of the suffix information and the filled second candidate token as the new suffix information, and predict the next token position to be filled based on the new prefix information and the new suffix information.

[0122] In some embodiments, the information completion device may further include a deletion module. The deletion module is configured to: after filling the token positions to be filled between the prefix information and the suffix information, delete the tokens in the completed information to be completed that meet the deletion conditions; where the deletion conditions are determined based on pre-configuration and / or the original input information.

[0123] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0124] Figure 6 is a block diagram of an information completion model training device provided by an embodiment of the present disclosure. As Figure 6 shown, the information completion model training device may include a first acquisition module 601, a second acquisition module 602, a prediction module 603, a determination module 604, and a training module 605.

[0125] Among them, the first acquisition module 601 is configured to acquire training sample data.

[0126] The second acquisition module 602 is configured to acquire prefix information and suffix information from the training sample data.

[0127] A prediction module 603 is configured to input the prefix information and the suffix information into an information completion model to perform completion prediction on the training sample data, and obtain a first candidate token for completion in a first direction and a second candidate token for completion in a second direction; the structure of the information completion model is a dual - tower information completion deep - learning model structure with bidirectional decoding.

[0128] A determination module 604 is configured to determine a model loss value according to the real token behind the prefix information, the real token in front of the suffix information, the first candidate token, and the second candidate token in the training sample data.

[0129] A training module 605 is configured to train the information completion model according to the model loss value.

[0130] In some embodiments, the determination module 604 is configured to: determine a first loss value according to the real token behind the prefix information and the first candidate token in the training sample data; determine a second loss value according to the real token in front of the suffix information and the second candidate token in the training sample data; and determine the model loss value according to the first loss value and the second loss value.

[0131] In some embodiments, the information completion model at least includes a prefix encoding unit, a suffix encoding unit, and a bidirectional decoding unit. Among them, the network structures of the prefix encoding unit and the suffix encoding unit are both encoders in the Transformer model, the bidirectional decoding unit is a decoder in the Transformer model, and the output layer of the decoder is changed to a bidirectional output layer.

[0132] Regarding the device in the above - mentioned embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0133] Figure 7 is a block diagram of an agent provided by an embodiment of the present disclosure. As Figure 7 shown, the agent may include an input module 701, a processing module 702, and an output module 703. Among them, the input module 701 is configured to receive input information; the processing module 702 is configured to determine a target task based on the input information received by the input module, determine an information completion model based on the target task, and execute the information completion method or the information completion model training method in the above - mentioned method embodiments by calling the information completion model to obtain output information; the output module 703 is configured to output the output information obtained by the processing module. Exemplarily, the information completion model may be a large model, but is not limited thereto.

[0134] According to the embodiments of the present application, the present application also provides an electronic device and a readable storage medium.

[0135] As shown Figure 8 in the figure, it is a block diagram of an electronic device according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0136] As shown Figure 8 in the figure, the electronic device includes: one or more processors 801, a memory 802, and interfaces for connecting the various components, including a high-speed interface and a low-speed interface. The various components are interconnected using different buses and can be mounted on a common motherboard or otherwise mounted as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, each device providing part of the necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 8 Here, one processor 801 is taken as an example.

[0137] The memory 802 is the non-transitory computer-readable storage medium provided by the present application. Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the information completion method or the information completion model training method provided by the present application. The non-transitory computer-readable storage medium of the present application stores computer instructions, and the computer instructions are used to make a computer execute the information completion method or the information completion model training method provided by the present application.

[0138] The memory 802, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the information completion method or the information completion model training method in the embodiments of the present application (for example, Figure 5 the acquisition module 501, the prediction module 502, the selection module 503, and the filling module 504 shown in the appendix, Figure 6The first acquisition module 601, the second acquisition module 602, the prediction module 603, the determination module 604, and the training module 605 shown). The processor 801 executes various functional applications and data processing of the server by running non-transitory software programs, instructions, and modules stored in the memory 802, that is, implements the information completion method or the information completion model training method in the above method embodiments.

[0139] The memory 802 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device and the like. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 802 may optionally include a memory remotely provided relative to the processor 801, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0140] The electronic device may further include: an input device 803 and an output device 804. The processor 801, the memory 802, the input device 803, and the output device 804 may be connected through a bus or other means. Figure 8 Taking the connection through the bus as an example.

[0141] The input device 803 may receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, and a joystick. The output device 804 may include a display device, an auxiliary lighting device (for example, an LED), and a tactile feedback device (for example, a vibration motor), etc. The display device may include but is not limited to a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0142] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0143] These computing programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0144] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0145] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain network.

[0146] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0147] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitation is made herein.

[0148] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. An information completion method, comprising: Obtaining prefix information and suffix information of the information to be completed from the original input information; Based on the prefix information and the suffix information, using a preset information completion model to perform completion prediction on the information to be completed, obtaining a first candidate token for completion in a first direction and the probability of the first candidate token, a second candidate token for completion in a second direction and the probability of the second candidate token; the structure of the information completion model is a dual - tower information completion deep - learning model structure with bidirectional decoding; Determining a target completion direction between the first direction and the second direction according to the probability of the first candidate token and the probability of the second candidate token; Filling the candidate token corresponding to the target completion direction among the first candidate token and the second candidate token along the target completion direction.

2. The method according to claim 1, wherein, The step of using a preset information completion model to perform completion prediction on the information to be completed based on the prefix information and the suffix information, obtaining a first candidate token for completion in a first direction and the probability of the first candidate token, a second candidate token for completion in a second direction and the probability of the second candidate token, includes: Encoding the prefix information based on a prefix encoding unit in the information completion model to obtain a prefix feature; Encoding the suffix information based on a suffix encoding unit in the information completion model to obtain a suffix feature; Fusing the prefix feature and the suffix feature based on a bidirectional decoding unit in the information completion model to obtain a prefix - suffix fusion feature; Inputting the prefix - suffix fusion feature into a bidirectional output layer in the bidirectional decoding unit to obtain a first candidate token for completion in a first direction and the probability of the first candidate token, a second candidate token for completion in a second direction and the probability of the second candidate token.

3. The method according to claim 1, wherein, The step of determining a target completion direction between the first direction and the second direction according to the probability of the first candidate token and the probability of the second candidate token, includes: Comparing the probabilities of the first candidate token and the second candidate token; When the probability of the first candidate token is greater than the probability of the second candidate token, determining the first direction as the target completion direction; or, When the probability of the second candidate token is greater than the probability of the first candidate token, determining the second direction as the target completion direction; or, When the probability of the first candidate token is equal to the probability of the second candidate token, determining the first direction and / or the second direction as the target completion direction.

4. The method according to claim 1, wherein, The step of filling the candidate token corresponding to the target completion direction among the first candidate token and the second candidate token along the target completion direction, includes: When the target completion direction is the first direction, fill the to-be-filled token position behind the prefix information with the first candidate token along the first direction; or, When the target completion direction is the second direction, fill the to-be-filled token position in front of the suffix information with the second candidate token along the second direction; or, When the target completion direction is the first direction and the second direction, fill the to-be-filled token position behind the prefix information with the first candidate token along the first direction, and fill the to-be-filled token position in front of the suffix information with the second candidate token along the second direction.

5. The method according to claim 4, further comprising: When the target completion direction is the first direction, after filling the first candidate token into the corresponding to-be-filled token position, keep the suffix information unchanged, and use the combination of the prefix information and the filled first candidate token as the new prefix information, and perform prediction of the next to-be-filled token position based on the new prefix information and the suffix information; Or, When the target completion direction is the second direction, after filling the second candidate token into the corresponding to-be-filled token position, keep the prefix information unchanged, and use the combination of the suffix information and the filled second candidate token as the new suffix information, and perform prediction of the next to-be-filled token position based on the prefix information and the new suffix information; Or, When the target completion direction is the first direction and the second direction, after filling the first candidate token and the second candidate token into the corresponding to-be-filled token positions respectively, use the combination of the prefix information and the filled first candidate token as the new prefix information, use the combination of the suffix information and the filled second candidate token as the new suffix information, and perform prediction of the next to-be-filled token position based on the new prefix information and the new suffix information.

6. The method according to any one of claims 1-5, further comprising: After filling the to-be-filled token positions between the prefix information and the suffix information, perform elimination processing on the tokens in the filled to-be-completed information that meet the elimination conditions; wherein, the elimination conditions are determined based on pre-configuration and / or the original input information.

7. An information completion model training method, comprising: Obtain training sample data; Obtain prefix information and suffix information from the training sample data; Input the prefix information and the suffix information into an information completion model to perform completion prediction on the training sample data, and obtain a first candidate token for completion in the first direction and a second candidate token for completion in the second direction; The structure of the information completion model is a two-tower information completion deep learning model structure with bidirectional decoding; Determine a model loss value based on the true tokens after the prefix information, the true tokens before the suffix information, the first candidate token, and the second candidate token in the training sample data; Train the information completion model according to the model loss value.

8. The method according to claim 7, wherein The determining the model loss value based on the true tokens after the prefix information, the true tokens before the suffix information, the first candidate token, and the second candidate token in the training sample data includes: Determine a first loss value based on the true tokens after the prefix information and the first candidate token in the training sample data; Determine a second loss value based on the true tokens before the suffix information and the second candidate token in the training sample data; Determine the model loss value according to the first loss value and the second loss value.

9. The method according to claim 7 or 8, wherein The information completion model at least includes a prefix encoding unit, a suffix encoding unit, and a bidirectional decoding unit. Among them, the network structures of the prefix encoding unit and the suffix encoding unit are both encoders in the Transformer model, the bidirectional decoding unit is the decoder in the Transformer model, and the output layer of the decoder is changed to a bidirectional output layer.

10. An information completion device, including: An acquisition module, configured to acquire prefix information and suffix information of information to be completed from the original input information; A prediction module, configured to, based on the prefix information and the suffix information, use a preset information completion model to perform completion prediction on the information to be completed, and obtain a first candidate token for the first-direction completion and the probability of the first candidate token, a second candidate token for the second-direction completion, and the probability of the second candidate token; the structure of the information completion model is a dual-tower information completion deep learning model structure with bidirectional decoding; A selection module, configured to determine a target completion direction between the first direction and the second direction according to the probability of the first candidate token and the probability of the second candidate token; A filling module, configured to fill the candidate token corresponding to the target completion direction among the first candidate token and the second candidate token along the target completion direction.

11. An information completion model training device, including: A first acquisition module, configured to acquire training sample data; A second acquisition module, configured to acquire prefix information and suffix information from the training sample data; A prediction module, configured to input the prefix information and the suffix information into the information completion model to perform completion prediction on the training sample data, and obtain a first candidate token for the first-direction completion and a second candidate token for the second-direction completion; The structure of the information completion model is a dual-tower information completion deep learning model structure with bidirectional decoding; A determination module, configured to determine a model loss value according to the true tokens behind the prefix information, the true tokens in front of the suffix information, the first candidate token, and the second candidate token in the training sample data; A training module, configured to train the information completion model according to the model loss value.

12. An agent, comprising: An input module, configured to receive input information; A processing module, configured to determine a target task based on the input information received by the input module, determine an information completion model based on the target task, and execute the method according to any one of claims 1-9 by invoking the information completion model to obtain output information; An output module, configured to output the output information obtained by the processing module.

13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

15. A program product, including at least one of a program and instructions, wherein, When at least one of the program and the instructions is executed by a processor, the steps of the method according to any one of claims 1-9 are implemented.