Text Prediction Method, Apparatus, Device, and Storage Medium

By training the text prediction model, iterative training is used to determine the loss value using position probability and answer text, which solves the accuracy problem of existing text prediction methods in knowledge-based answer prediction, and achieves higher prediction accuracy and applicability.

CN113704391BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110380399.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-08
Publication Date
2025-07-18
Estimated Expiration
2041-04-08

AI Technical Summary

Technical Problem

Existing text prediction methods have poor results when predicting highly knowledgeable answers and are prone to predict fixed combinations in languages, resulting in significant differences between the results and the actual results.

Method used

By obtaining multiple training text groups, the position probability of each word in each result text corresponds to the position where the answer text is located is determined, the first training loss value is determined based on the position probability and the answer text, and the text prediction model iteratively trains the text prediction model to improve prediction accuracy.

Benefits of technology

It improves the accuracy and applicability of text prediction, can better identify the position of the answer text in the result text, and reduces the difference between the predicted results and the actual results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113704391B_ABST
    Figure CN113704391B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a text prediction method, apparatus, device, and storage medium, which can be applied to fields such as artificial intelligence and computers. The method includes: obtaining a target query text and a target result text corresponding to the target query text; and determining a target answer text from the target result text through a text prediction model based on the target query text and the target result text. By adopting the embodiments of the present application, the efficiency and accuracy of text prediction can be improved, and the applicability is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a text prediction method, apparatus, device, and storage medium. Background Art

[0002] With the continuous development of computer technology and artificial intelligence technology, text prediction tasks such as answer prediction have been gradually applied to various industries.

[0003] On the one hand, existing text prediction methods are based on the way humans speak for prediction, and often have poor effects when predicting answers with strong knowledge. On the other hand, since existing text prediction methods often construct training samples by masking word strings or by actively constructing contexts to obtain training samples, the resulting text prediction methods often tend to predict fixed collocations in language, and the prediction results often deviate greatly from the actual results.

[0004] Therefore, how to improve the accuracy of text prediction has become an urgent problem to be solved. Summary of the Invention

[0005] Embodiments of this application provide a text prediction method, apparatus, device, and storage medium, which can improve the accuracy of text prediction and have high applicability.

[0006] On the one hand, embodiments of this application provide a text prediction method, which includes:

[0007] Obtain a target query text and a target result text corresponding to the target query text;

[0008] Based on the target query text and the target result text, determine a target answer text from the target result text through a text prediction model;

[0009] Wherein, the text prediction model is trained based on the following method:

[0010] Obtain a plurality of training text groups, each of the training text groups includes a query text, a result text of the query text, and an answer text corresponding to the query text, and each of the result texts includes a corresponding answer text;

[0011] Determine the encoding features corresponding to each result text and the corresponding query text, input the encoding features into an initial model, and through the initial model, determine the position probability of each word in each result text corresponding to the position of the answer text in the corresponding result text;

[0012] For each of the above result texts, based on the position probability corresponding to the result text and the answer text, determine the first training loss value corresponding to the result text, and determine the total training loss value based on the above first training loss value;

[0013] Iteratively train the above initial model based on the above total training loss value until the above total training loss value meets the preset training end condition, and determine the above text prediction model based on the model after training ends.

[0014] On the other hand, an embodiment of the present application provides a text prediction device, and the device includes:

[0015] An acquisition module, configured to acquire a target query text and a target result text corresponding to the target query text;

[0016] A prediction module, configured to determine a target answer text from the target result text through a text prediction model based on the target query text and the target result text;

[0017] The above device includes a training module, and the training module includes:

[0018] An acquisition unit, configured to acquire a plurality of training text groups, each of the above training text groups includes a query text, a result text of the query text, and an answer text corresponding to the query text, and each of the above result texts includes a corresponding answer text;

[0019] A prediction unit, configured to determine the encoded features corresponding to each of the above result texts and the corresponding query text, input the above encoded features into an initial model, and through the initial model, determine the position probability of each word in each of the above result texts corresponding to the position where the answer text is located in the corresponding result text;

[0020] A determination unit, configured to, for each of the above result texts, determine the first training loss value corresponding to the result text based on the position probability corresponding to the result text and the answer text, and determine the total training loss value based on the above first training loss value;

[0021] A training unit, configured to iteratively train the above initial model based on the above total training loss value until the above total training loss value meets the preset training end condition, and determine the above text prediction model based on the model after training ends.

[0022] In some feasible implementation manners, for each of the above result texts, the position probability corresponding to the result text includes a first prediction probability that each word in the result text is the starting position of the answer text in the result text, and a second prediction probability that it is the ending position;

[0023] The above determination unit is configured to:

[0024] Determine the starting position and the ending position of the answer text in the result text;

[0025] Based on the first prediction probability corresponding to the first target word and the second prediction probability corresponding to the second target word in the result text, determine the first training loss value corresponding to the result text;

[0026] Wherein, the above-mentioned first target word corresponds to the starting position of the answer text in the result text, and the above-mentioned second target word corresponds to the ending position of the answer text in the result text.

[0027] In some feasible implementation manners, for each of the above-mentioned result texts, the above-mentioned determining unit is configured to:

[0028] Determine the first true probability that each word in the result text belongs to the target text in the result text;

[0029] Through the above-mentioned initial model, based on the encoded feature corresponding to the result text, determine the third prediction probability that each word in the result text belongs to the target text in the result text, wherein the target text in the result text is the text associated with the corresponding answer text;

[0030] Based on the third prediction probability and the first true probability corresponding to the result text, determine the text relevance between the result text and the corresponding query text;

[0031] Based on the text relevance corresponding to the result text and the above-mentioned first training loss value, determine the total training loss value.

[0032] In some feasible implementation manners, for each of the above-mentioned result texts, the above-mentioned determining unit is configured to:

[0033] Determine the starting position and the ending position of the answer text in the result text;

[0034] For each word in the result text, based on the position of the word in the result text and the starting position and the ending position of the answer text in the result text, determine the first true probability that the word belongs to the target text in the result text.

[0035] In some feasible implementation manners, for each of the above-mentioned result texts, the above-mentioned determining unit is configured to:

[0036] For each word in the result text, based on the encoded feature corresponding to the result text, determine the fourth prediction probability that the word is the starting position of the target text in the result text and the fifth prediction probability that the word is the ending position of the target text in the result text;

[0037] Based on the fourth prediction probability and the fifth prediction probability corresponding to each word in the result text, determine the third prediction probability that each word in the result text belongs to the target text in the result text.

[0038] In some feasible embodiments, for each of the above result texts, the determining unit is configured to:

[0039] For each word in the result text, determine the third target word before the word in the result text and the fourth target word after the word in the result text;

[0040] For each word in the result text, based on the fourth prediction probability corresponding to the word, the fourth prediction probability corresponding to the third target word corresponding to the word, the fifth prediction probability corresponding to the word, and the fifth prediction probability corresponding to the fourth target word corresponding to the word, determine the third prediction probability that the word belongs to the target text in the result text.

[0041] In some feasible embodiments, for each of the above result texts, the determining unit is configured to:

[0042] Based on the first true probability corresponding to the result text, determine the second true probability that each word in the result text does not belong to the target text in the result text;

[0043] Based on the third prediction probability corresponding to the result text, determine the sixth prediction probability that each word in the result text does not belong to the target text in the result text;

[0044] Based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text, determine the text relevance between the result text and the corresponding query text.

[0045] In some feasible embodiments, for each of the above result texts, the determining unit is configured to:

[0046] Based on the text relevance corresponding to each of the above result texts, determine a text relevance threshold;

[0047] If the text relevance corresponding to the result text is greater than or equal to the above relevance threshold, then determine a second training loss value based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text, and determine a total training loss value based on the above first training loss value and the above second training loss value;

[0048] If the text relevance corresponding to the result text is less than the above relevance threshold, then determine the above first training loss value as the above total training loss value.

[0049] In some feasible embodiments, for each of the above result texts, the determining unit is configured to:

[0050] Determine the first weight corresponding to the first training loss value and the second weight corresponding to the second training loss value;

[0051] Based on the first training loss value and the corresponding first weight, and the second training loss value and the corresponding second weight, determine the total training loss value.

[0052] In some feasible embodiments, for each of the above result texts, the determining unit is configured to:

[0053] Determine the sum of the text relevance degrees corresponding to the respective result texts;

[0054] Determine the number of each of the above result texts, and based on the number of each of the above result texts and the sum of the text relevance degrees, determine the text relevance threshold.

[0055] In some feasible embodiments, the obtaining unit is further configured to:

[0056] Determine a plurality of texts to be processed. For each of the above texts to be processed, determine the corresponding training text group based on the following method:

[0057] Determine the text corresponding to any text interval in the text to be processed as the answer text, and based on the other texts in the text to be processed except the answer text, determine the query text corresponding to the answer text;

[0058] Determine a plurality of retrieval texts for the query text, determine the text similarity between the query text and each of the above retrieval texts, and determine the retrieval text with the highest text similarity as the result text corresponding to the query text;

[0059] Determine the query text, the result text, and the answer text corresponding to the text to be processed as a training text group.

[0060] In some feasible embodiments, the prediction module is configured to:

[0061] Determine the target encoding features corresponding to the target query text and the target result text;

[0062] Input the target encoding features into a text prediction model to obtain the first target probability that each word in the target result text is the starting position of the target answer text corresponding to the target query text, and the second target probability of the ending position;

[0063] Based on the first target probability and the second target probability corresponding to each word in the target result text, determine the target answer text from the target result text.

[0064] On the other hand, an embodiment of the present application provides an electronic device, including a processor and a memory, which are interconnected;

[0065] The above-mentioned memory is used to store a computer program;

[0066] The above-mentioned processor is configured to execute the text prediction method provided by the embodiment of the present application when calling the above-mentioned computer program.

[0067] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the text prediction method provided by the embodiment of the present application.

[0068] On the other hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text prediction method provided by the embodiment of the present application.

[0069] In the embodiment of the present application, during the model training process, by determining the position probability of each word in the result text corresponding to the position in the answer text in each training text group, it can be ensured that determining the first training loss value based on the position probability and the answer text can better represent the difference between the predicted position of each word in the result text corresponding to the answer text and the position of the answer text in the result text, thereby improving the training effect. The text prediction model trained based on the embodiment of the present application can improve the accuracy of text prediction and has high applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0071] Figure 1 is a flowchart of the text prediction method provided by the embodiment of the present application;

[0072] Figure 2 is a flowchart of the training method of the text prediction model provided by the embodiment of the present application;

[0073] Figure 3 is a schematic diagram of the scenario for determining the training text group provided by the embodiment of the present application;

[0074] Figure 4It is a schematic diagram of the scenario of the training method of the text prediction model provided by the embodiment of the present application;

[0075] Figure 5 It is a schematic flowchart of the method for determining the total training loss value provided by the embodiment of the present application;

[0076] Figure 6 It is a schematic diagram of a scenario of the training method of the text model provided by the embodiment of the present application;

[0077] Figure 7 It is a schematic diagram of another scenario of the training method of the text model provided by the embodiment of the present application;

[0078] Figure 8 It is a schematic diagram of the structure of the text prediction device provided by the embodiment of the present application;

[0079] Figure 9 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0080] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0081] See Figure 1 , Figure 1 It is a schematic flowchart of the text prediction method provided by the embodiment of the present application. As Figure 1 shown, the text prediction method provided by the embodiment of the present application may include the following steps:

[0082] Step S11: Obtain the target query text and the target result text corresponding to the target query text.

[0083] In some feasible implementation manners, the above-mentioned target query text is a text for querying information, which may be a query statement input by a user, or a text obtained by semantic recognition of a query voice, and can be specifically determined based on the requirements of the actual application scenario, and is not limited herein.

[0084] Among them, the above-mentioned target result text is a relevant text including the target answer text corresponding to the target query text, which may be a result text input by a user, or a text including the target answer corresponding to the target query text obtained by searching based on big data, search algorithms, search engines, etc., and can be specifically determined based on the requirements of the actual application scenario, and is not limited herein.

[0085] For example, in response to a confirmation query operation on a search interface, the text corresponding to the confirmation query operation is obtained and determined as the target query text. Further, based on a retrieval model, a search engine, etc., multiple result texts related to the target query text are retrieved, and a target result text including the target answer text corresponding to the target query text is obtained therefrom.

[0086] As another example, the user inputs the target query text and the target result text on a question query interface, and triggers a confirmation query operation to determine the target answer text corresponding to the target query text included in the target result text. Further, in response to the confirmation query operation on the question query interface, the target query text corresponding to the confirmation query and the corresponding target result text can be obtained.

[0087] Step S12: Based on the target query text and the target result text, determine the target answer text from the target result text through a question-answering model.

[0088] In some feasible implementation manners, after the target query text and the corresponding target result text are obtained, the target coding features corresponding to the target query text and the target result text can be determined. Specifically, the word sequence of the target query text and the word sequence of the target result text can be determined, and then the word sequences of the target query text and the target result text are concatenated, and the concatenated word sequence is encoded to obtain the target coding features.

[0089] Further, the above target coding features are input into a text prediction model, and through the text prediction model, the target probability (hereinafter referred to as the first target probability for convenience of description) that each word in the target result text is the starting position of the target answer text corresponding to the target query text, and the target probability (hereinafter referred to as the second target probability for convenience of description) that each word in the target result text is the ending position of the target answer text are obtained.

[0090] In other words, the text prediction model can determine two probabilities corresponding to each word in the target result text based on the target coding features. For each word in the target result text, the first target probability that the word is the target answer text in the target result text and the second target probability that the word is the target answer text in the target result text can be obtained through the text prediction model.

[0091] Further, the highest probability among the first target probabilities corresponding to each word in the target result text can be determined, and the word corresponding to the highest probability is determined as the starting position of the target answer text in the target result text, that is, the first word of the target answer text. Similarly, the highest probability among the second target probabilities corresponding to each word in the target result text can be determined, and the word corresponding to the highest probability is determined as the ending position of the target answer text in the target result text, that is, the last word of the target answer text. Thus, the target answer text in the target result text can be determined based on the starting position and the ending position of the target answer (the first word and the last word).

[0092] In some feasible implementation manners, the above text prediction model can be obtained by training with a training text group. The text prediction method and the training method of the text model provided by the embodiments of the present application can be executed by any electronic device or server.

[0093] Among them, the server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server or a server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.

[0094] For the specific training manner of the training method of the text model provided by the embodiments of the present application, reference can be made to Figure 2 。 Figure 2 is a schematic flowchart of the training method of the text prediction model provided by the embodiments of the present application. As Figure 2 shown, the training method of the text prediction model provided by the embodiments of the present application may include the following steps:

[0095] Step S21: Obtain a plurality of training text groups, each training text group includes a query text, a result text of the query text, and an answer text corresponding to the query text, and each result text includes the corresponding answer text.

[0096] In some feasible implementation manners, before training the text prediction model, it is necessary to obtain a plurality of training text groups, and use the obtained training text groups as training data to train the initial model. Among them, the above initial model can be obtained by combining one or more neural networks and existing text processing models based on neural networks.

[0097] Among them, the neural network includes but is not limited to Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory artificial neural network (LSTM), Gated Recurrent Unit (GRU), etc. The above text processing models based on neural networks include but are not limited to the BERT (Bidirectional Encoder Representations from Transformers) model and the feature extraction model based on the XLnet (Extra-Long Net) model, etc. Specifically, it can be configured and selected based on the actual application scenario requirements, and no limitation is made here.

[0098] In some feasible implementation manners, for each training text group, the training text group includes a query text, a result text of the query text, and an answer text corresponding to the query text.

[0099] Among them, the result text of the query text includes the answer text corresponding to the query text, that is, for each query text, the answer text corresponding to the query text can be determined from the result text corresponding to the query text.

[0100] When obtaining the training text group, multiple texts to be processed can be obtained, and each text to be processed is processed to obtain a training text group. Among them, the above texts to be processed can be relevant text paragraphs in any field (such as medical, sports, education, and biology, etc.) obtained based on multiple text sources, such as text paragraphs obtained from news, textbooks, blogs, and papers in relevant fields, or can also be manually written text paragraphs. The specific obtaining method can be determined based on the actual application scenario requirements, and no limitation is made here.

[0101] In the embodiments of the present application, when obtaining the text to be processed, it can be obtained from a database or a blockchain in different fields and different application scenarios, or can be obtained from the network based on technologies such as big data. Specifically, it can be determined based on the actual application scenario requirements, and no limitation is made here.

[0102] Specifically, for each text to be processed, the text corresponding to any text interval in the text to be processed can be determined as the answer text corresponding to the text to be processed. For example, names, data, etc. in the text to be processed are used as the answer text. For example, if the text to be processed is "The champion of the A League in the 2021 - 2022 season is the Tigers team", then "the Tigers team" can be used as the answer text.

[0103] Further, for the other texts in the processed text except the answer text, they can be used as the query text corresponding to the answer text. For example, if the text to be processed is "The champion of the A League in the 2021-2022 season is the Tigers", and "the Tigers" is used as the answer text, then the text "The champion of the A League in the 2021-2022 season is" can be determined as the query text.

[0104] Optionally, for the convenience of model training, the position of the answer text can be filled with a preset symbol, that is, the answer text in the text to be processed is replaced with a preset symbol to obtain the query text. Among them, the above preset symbol can be specifically determined based on the actual application scenario requirements and is not limited here. For example, if the text to be processed is "The champion of the A League in the 2021-2022 season is the Tigers", and "the Tigers" is used as the answer text, and the preset symbol is [BLANC], then the text "The champion of the A League in the 2021-2022 season is [BLANC]" can be determined as the query text.

[0105] Based on the above implementation method, the query text and the answer text in a training text group can be determined. For the result text in the training text group, multiple retrieval texts corresponding to the query text can be determined first. For example, multiple retrieval texts related to the query text can be determined through a retrieval algorithm, a search engine, etc., or multiple paragraphs written manually related to the query text can be determined as the retrieval texts corresponding to the query text, which can be specifically determined based on the actual application scenario requirements and is not limited here.

[0106] Further, for the multiple retrieval texts corresponding to the query text, the retrieval texts can be initially screened based on the answer text corresponding to the query text, and the retrieval texts containing the answer text can be screened out. For this query text, even if the retrieved text after screening contains the answer text, it is very likely that there are texts that include the answer text but have little relevance to the query text. For example, if the query text is "In a game where Zhang San faced [BLANC] in the playoffs of the second season, he even scored 63 points wildly", and the answer text is "the Tigers", if a retrieved text is "...Li Si even scored 39 points in the game against the Tigers and refreshed his personal scoring record...". Even though the retrieved text contains the answer text "the Tigers", it is obviously very different from the actual content described in the query text.

[0107] Based on this, for the filtered retrieval texts, the text similarity between each retrieval text and the query text can be determined, and the retrieval text with the highest text similarity is determined as the result text corresponding to the query text. Furthermore, the query text, result text, and answer text obtained based on a text to be processed can be determined as a training text group. Among them, the calculation methods of the above text similarity include but are not limited to the corresponding calculation methods such as Euclidean distance, cosine similarity, Jaccard distance, and Manhattan distance, etc., which can be specifically determined based on the actual application scenario requirements and are not limited here.

[0108] See Figure 3 , Figure 3 is a schematic diagram of the scenario for determining the training text group provided by an embodiment of the present application. In Figure 3 the text to be processed is "In the ninth year of Kangxi (1670), the Lufeng Circuit was established, stationed in Fengyang Prefecture (the governing seat is now in Fengyang County, Anhui), and governing Luzhou Prefecture and Fengyang Prefecture". Taking the "Fengyang County" in this text to be processed as the answer text, and replacing the "Fengyang County" in the text to be processed with a preset symbol "BLANC", the query text "In the ninth year of Kangxi (1670), the Lufeng Circuit was established, stationed in Fengyang Prefecture (the governing seat is now in Anhui [BLANC]), and governing Luzhou Prefecture and Fengyang Prefecture" is obtained.

[0109] Furthermore, the query text is retrieved and the answer text is screened to obtain the corresponding retrieval texts (such as Retrieval Text 1, Retrieval Text 2, and Retrieval Text 3), and the text similarity between each retrieval text and the query text is determined. The retrieval text with the highest text similarity, "In the ninth year of Kangxi (1760), the Lufeng Circuit was established, stationed in Fengyang. Both Fengyang Prefecture and County belonged to the Lufeng Circuit. In the twentieth year of Qianlong (1755), Linhuai County was incorporated into Fengyang County...", is determined as the result text.

[0110] At this time, the query text, result text, and answer text obtained based on this text to be processed can be determined as a training text group.

[0111] Among them, for each text to be processed, different texts in this processed text can be determined as different answer texts, and then multiple groups of training text groups can be determined based on the same text to be processed and different answer texts. The specific determination method is not elaborated here.

[0112] Based on the above method, if the texts to be processed obtained are all from the same field, the text prediction model trained based on the training text group determined from the text to be processed can be applicable to the text prediction tasks in this field. If the texts to be processed obtained come from multiple fields, when the number of training text groups obtained based on the text to be processed is large enough, the text prediction model trained based on the training text group can be applicable to the text prediction tasks in various fields.

[0113] For example, see Figure 4 ,Figure 4 This is a schematic diagram of the scenario of the training method of the text prediction model provided by the embodiments of the present application. In Figure 4 , by obtaining the text to be processed corresponding to domain 1 and performing model training on the training text group obtained from the text to be processed corresponding to domain 1, a text prediction model for domain 1 can be obtained. By obtaining the text to be processed corresponding to domain 2 and performing model training on the training text group obtained from the text to be processed corresponding to domain 2, a text prediction model for domain 2 can be obtained. By obtaining the text to be processed corresponding to domain f and performing model training on the training text group obtained from the text to be processed corresponding to domain f, a text prediction model for domain f can be obtained. In the search scenario, according to the factual questions in different domains input by the user, the target result text corresponding to the factual question (target query text) in this domain can be determined, and the text prediction model corresponding to this domain can be called to predict the answer from the target result text. Optionally, the text prediction models for each domain described above can be integrated into a text prediction module, so as to obtain a text prediction module applicable to multiple domains.

[0114] The training method of the cross-domain text classification model provided by the embodiments of the present application is applicable to the field of machine learning (ML) in artificial intelligence (AI), as well as the fields of cloud computing and artificial intelligence cloud services in cloud technology, and a text prediction model can be trained.

[0115] Among them, artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence.

[0116] Machine learning (ML) is a special study on how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. In the embodiments of the present application, through the training method of the cross-domain text classification model provided by this embodiment, the machine can be equipped with cross-domain text prediction capabilities.

[0117] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. The training method of the cross-domain text prediction model provided by the embodiments of this application can be implemented based on cloud computing in cloud technology.

[0118] Artificial intelligence cloud services, generally also known as AIaaS (AI as a service). This is currently the mainstream service method of an artificial intelligence platform. Specifically, the AIaaS platform will split several common artificial intelligence services and provide independent or packaged services in the cloud, such as cross-domain text prediction services, etc.

[0119] Step S22: Determine the encoding features corresponding to each result text and the corresponding query text, input the encoding features into the initial model, and through the initial model, determine the position probability of each word in each result text corresponding to the position of the answer text in the result text.

[0120] In some feasible implementation manners, for each training text group, the encoding features corresponding to the result text and the corresponding query text in the training text group can be determined, and then the encoding features can be used as the model input of the initial model.

[0121] Specifically, for each training text group, the word sequence of the result text and the word sequence of the query text in the training text group can be determined, and then the word sequence of the query text and the word sequence of the result text are concatenated to obtain a concatenated word sequence corresponding to the query text and the result text.

[0122] For example, represents the query text, represents the result text, k is the index of the training text group, then the query text in the kth training text group has a word sequence of The result text in the kth training text group has a word sequence of

[0123] where M is the number of words in the query text and represents the Mth word in the query text N is the number of words in the result text and represents the Nth word in the result text . Concatenate the word sequences and to obtain the concatenated word sequence Among them, [CLS] represents a preset symbol for the training task, and [SEP] represents a preset symbol for the end of a sequence of conjunctions or a sequence of words. The specific symbol representation can be determined based on the requirements of the actual application scenario and will not be limited here.

[0124] Among them, the words in the query text and the result text can be a single character and / or a word, such as a Chinese character, an English word, or a Chinese phrase, etc. The specific can be determined based on the requirements of the actual application scenario and will not be limited here.

[0125] Furthermore, each word in the concatenated word sequence is encoded to obtain the word vector corresponding to each word. Then, the word vectors corresponding to each word in the concatenated word sequence are input into the initial model, and the initial model encodes the word vectors corresponding to each word in the concatenated word sequence to obtain encoded features. Alternatively, the concatenated word sequence is input into an existing encoding model (such as the BERT model), and the encoding model is used to obtain the encoded features corresponding to the concatenated word sequence.

[0126] For example, based on the concatenated word sequence, the corresponding encoded feature H∈R can be obtained (M+N+3)×d , where R is the set of real numbers, M is the number of words in the query text, N is the number of words in the result text, and d is the encoding dimension corresponding to the word encoding. The specific can be determined based on the requirements of the actual application scenario and the model structure and will not be limited here.

[0127] In some feasible implementation manners, the initial model can determine the position probability of each word in each result text corresponding to the position of the answer text in the result text based on the above encoded features. Among them, for the result text in each training text group, the position probability corresponding to the result text includes the first prediction probability that each word in the result text is the starting position of the answer text in the result text, and the second prediction probability that each word in the result text is the ending position of the answer text in the result text.

[0128] In other words, for each encoded feature, the initial model can determine, based on the encoded feature, the first prediction probability that each word in the result text corresponding to the encoded feature is the first word of the answer text in the result text, and the second prediction probability that it is the last word of the answer text in the result text.

[0129] Among them, the initial model can determine the first prediction probability that each word in the result text is the starting position of the answer text and the second prediction probability that it is the ending position of the answer text based on the encoded feature in the following manner:

[0130]

[0131] Among them, represents the starting position of the answer text in the result text corresponding to the k-th training text group, The first prediction probability that the i-th word in the result text is the starting position of the answer text in the result text, represents the ending position of the answer text in the result text corresponding to the k-th training text group, The second prediction probability that the i-th word in the result text is the ending position of the answer text in the result text.

[0132] where H is the encoded feature, and are the relevant parameters for the initial model to determine the first prediction probability corresponding to each word in the result text, and are the relevant parameters for the initial model to determine the second prediction probability corresponding to each word in the result text. and Specifically, it can be determined based on the requirements of the actual application scenario and the model structure of the initial model, and there is no limitation here.

[0133] Step S23: For each result text, based on the position probability corresponding to the result text and the answer text, determine the first training loss value corresponding to the result text, and determine the total training loss value based on the first training loss value.

[0134] In some feasible implementation manners, after the initial model determines the first prediction probability that each word in the corresponding result text is the starting position of the answer text, and the second prediction probability that each word is the ending position of the answer text, based on the encoded feature corresponding to the training text group, the corresponding first training loss value can be determined.

[0135] Specifically, the starting position and the ending position of the answer text in the result text of the training text group can be determined, and then the first target word corresponding to the starting position of the answer text in the result text and the second target word corresponding to the ending position can be determined. Further, based on the first prediction probability corresponding to the first target word and the second prediction probability corresponding to the second target word, the first training loss value can be determined.

[0136] Among them, the first training loss value can be determined based on the following expression:

[0137]

[0138] where represents the starting position of the answer text in the result text corresponding to the k-th training text group, The first prediction probability that the i-th word in the result text is the starting position of the answer text in the result text, represents the starting position of the answer text in the result text corresponding to the k-th training text group The second predicted probability indicating that the i-th word in the result text is the end position of the answer text in the result text, where N is the number of words in the result text in the result text

[0139] Among them is an indicator function. If the i-th word in the result text is the start position of the answer text in the result text, the value of the indicator function is 1, that is, the probability that this word is the start position of the answer text in the result text is 1; if the i-th word in the result text is not the start position of the answer text in the result text, the value of the indicator function is 0, that is, the probability that this word is the start position of the answer text in the result text is 0

[0140] Among them is an indicator function. If the i-th word in the result text is the end position of the answer text in the result text, the value of the indicator function is 1, that is, the probability that this word is the end position of the answer text in the result text is 1; if the i-th word in the result text is not the end position of the answer text in the result text, the value of the indicator function is 0, that is, the probability that this word is the end position of the answer text in the result text is 0

[0141] Among them, the first training loss value characterizes the difference between the predicted probability and the true probability that each word in the result text is the start position of the answer text in the result text, and the difference between the predicted probability and the true probability that each word is the end position of the answer text in the result text. The smaller the first training loss value, the smaller the difference between the predicted probability and the true probability that each word is the start position of the answer text in the result text, and the difference between the predicted probability and the true probability that each word is the end position of the answer text in the result text

[0142] In some feasible embodiments, for each result text, the first training loss value corresponding to the result text can be determined as the total training loss value. That is, when training the initial model based on each training text group, the first training loss value corresponding to the training text group can be determined as the total training loss value for this model training. Furthermore, the final model trained based on the total training loss value has the ability to predict the answer text included in the result text

[0143] Optionally, in addition to the ability to train the initial model to determine the first predicted probability that each word in the result text is the start position of the answer text and the second predicted probability that it is the end position of the answer text based on the training text group, that is, the ability to train the initial model to predict the answer text in the result text, the initial model can also be trained to predict the context related to the answer text in the result text (for the convenience of description, hereinafter referred to as the target text). Furthermore, the total training loss value can be determined based on the above two training tasks

[0144] Specifically, reference can be made toFigure 5 , Figure 5 is a schematic flowchart of the method for determining the total training loss value provided by an embodiment of the present application. Among them, Figure 5 the method for determining the total training loss shown is for the result text in each training text group. The method for determining the total training loss value provided by an embodiment of the present application may include the following steps:

[0145] Step S231: Determine the first true probability that each word in the result text belongs to the target text in the result text.

[0146] In some feasible embodiments, the target text in the result text is the text (context) associated with the corresponding answer text, that is, the semantics of the target text in the result text is similar to the semantics of the corresponding answer text, such as the descriptive text related to the answer text, the paragraph text with the answer text as the main content, etc.

[0147] When determining the first true probability that each word in the result text belongs to the target text in the result text, the start position and end position of the answer text in the result text can be determined. Since the start position and end position of the target text associated with the answer text in the result text cannot be clearly defined, the first true probability that each word in the result text belongs to the target text in the result text can be determined based on the distance between each word in the result text and the start position and end position of the answer text in the result text. Specifically, it can be determined based on the following expression:

[0148]

[0149] Among them, represents the start position of the answer text in the result text corresponding to the kth training text group, represents the end position of the answer text in the result text corresponding to the kth training text group i is the index of the word in the result text, represents the distance between the ith word in the result text and the distance, represents the distance between the ith word in the result text and the distance, represents the target text in the result text corresponding to the kth training text group, represents the ith word in the result text corresponding to the kth training text group. q is a hyperparameter less than 1 and greater than 0, and δ is a fixed-size word window range, which can be specifically determined based on the requirements of the actual application scenario and is not limited here.

[0150] As can be seen from the above formula, for any word in the result text, if the word is between the start position and the end position of the answer text in the result text, the first true probability that the word belongs to the target text in the result text is 1; if the word is within the word window before the start position of the answer text in the result text, the first true probability that the word belongs to the target text in the result text is If the word is within the word window after the end position of the answer text in the result text, the first true probability that the word belongs to the target text in the result text is

[0151] Among them, for a word located outside any of the above word windows and the range between the start position and the end position, it is determined that the word does not belong to the target text in the result text, and the corresponding first true probability of the word is 0.

[0152] Step S232: Through the initial model, based on the encoding features corresponding to the result text, determine the third prediction probability that each word in the result text belongs to the target text in the result text.

[0153] In some feasible implementation manners, after determining the first true probability that each word in the result text belongs to the target text in the result text, the third prediction probability that each word in the result text belongs to the target text in the result text can be determined, and then based on the first true probability and the third prediction probability corresponding to each word in the result text, the second training loss value corresponding to the training task is determined.

[0154] Specifically, based on the encoding features corresponding to the result text (that is, the encoding features determined based on the word sequences of the result text and the corresponding query text), the fourth prediction probability of the start position of the target text in the result text and the fifth prediction probability of the end position of the target text in the result text can be determined. The specific determination process is determined by the following expression:

[0155]

[0156] Among them, represents the start position of the target text in the result text corresponding to the kth training text group, represents the fourth prediction probability that the ith word in the result text is the start position of the target text in the result text, represents the end position of the target text in the result text corresponding to the kth training text group, represents the fifth prediction probability that the ith word in the result text is the end position of the target text in the result text.

[0157] Among them, H is the encoding feature, and are the relevant parameters of the initial model for determining the fourth prediction probability corresponding to each word in the result text. and are the relevant parameters of the initial model for determining the fifth prediction probability corresponding to each word in the result text. and Specifically, it can be determined based on the requirements of the actual application scenario and the model structure of the initial model, and there is no limitation here.

[0158] Furthermore, based on the fourth prediction probability and the fifth prediction probability corresponding to each word in the result text, determine the third prediction probability that each word in the result text belongs to the target text in the result text. For each word in the result text, the third target word before the word in the result text and the fourth target word after the word in the result text can be determined. Then, based on the fourth prediction probability corresponding to the word, the fourth prediction probability corresponding to the third target word corresponding to the word, the fifth prediction probability corresponding to the word, and the fifth prediction probability corresponding to the fourth target word corresponding to the word, determine the third prediction probability that the word belongs to the target text in the result text.

[0159] That is, for each word in the result text, based on the fourth prediction probability that the word and the words before the word in the result text correspond to the start position of the target text, and the fifth prediction probability that the word and the words after the word in the result text correspond to the end position of the target text, determine the third prediction probability that the word belongs to the target text in the result text. The specific determination method can be determined based on the following expression:

[0160]

[0161] where represents the start position of the target text in the result text corresponding to the k-th training text group, i and j are the indexes of the words in the result text, represents the fourth prediction probability that the i-th word in the result text is the start position of the target text in the result text, represents the end position of the target text in the result text corresponding to the k-th training text group, represents the fifth prediction probability that the i-th word in the result text is the end position of the target text in the result text. represents the target text in the result text corresponding to the k-th training text group, represents the i-th word in the result text corresponding to the k-th training text group.

[0162] Step S233: Based on the third prediction probability and the first true probability corresponding to the result text, determine the text relevance between the result text and the corresponding query text.

[0163] In some feasible embodiments, after determining the third prediction probability and the first true probability that each word in the result text belongs to the target text in the result text, the text relevance between the result text and the corresponding query text can be determined.

[0164] Specifically, based on the first true probability corresponding to the result text, the second true probability that each word in the result text does not belong to the target text in the result text can be determined. Based on the third prediction probability corresponding to the result text, the sixth prediction probability that each word in the result text does not belong to the target text in the result text can be determined. Further, based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text, the text relevance between the result text and the corresponding query text is determined. Specifically, it can be determined through the following expression:

[0165]

[0166] Wherein, represents the query text, represents the result text, k is the index of the training text group, N is the number of words in the result text and M is the number of words in the query text . represents the i-th word in the result text corresponding to the k-th training text group.

[0167] Wherein, represents the text relevance between the result text and the query text in the k-th training text group. represents the third prediction probability that the i-th word belongs to the target text in the result text, represents the first true probability that the i-th word belongs to the target text in the result text.

[0168] Wherein, represents the second true probability that the i-th word does not belong to the target text in the result text, represents the sixth prediction probability that the i-th word does not belong to the target text in the result text,

[0169] Wherein, for any result text, if the text relevance corresponding to the result text is higher, it indicates that the text content of the result text is more similar to the text content of the corresponding query text. Furthermore, it can indicate that the relevance between the target text in the result text and the corresponding answer text is higher, and the probability that the result text is the context of the corresponding query text is higher.

[0170] Step S234: Determine the total training loss value based on the text relevance corresponding to the result text and the first training loss value.

[0171] In some feasible embodiments, after determining the text relevance between the result text and the corresponding query text in each training text group, a relevance threshold can be determined based on the text relevance corresponding to each result text, so as to determine the total training loss value based on the relevance threshold and the corresponding first training loss value.

[0172] Among them, the above text relevance threshold can be a preset value, can also be determined based on actual model parameters, and can also be determined based on the relevant probability obtained during model training, which is not limited here.

[0173] Optionally, the above relevance threshold can be determined based on the text relevance corresponding to each result text. Specifically, the sum of the text relevance corresponding to each result text can be determined, and further the number of each result text can be determined. Based on the number of each result text and the sum of the text relevance, the text relevance threshold can be determined.

[0174] As an example, the above relevance threshold can be the average relevance of the text relevance corresponding to each of the above result texts where, |D A | represents the number of training text groups, that is, the number of text relevance.

[0175] In some feasible embodiments, for each result text, if the text relevance corresponding to the result text is greater than or equal to the above relevance threshold, it means that the text content of the result text has a high similarity to the text content of the corresponding query text during this training process. Furthermore, it can be shown that the relevance between the target text and the corresponding answer text in the result text is high. Therefore, it shows that the training task for context prediction based on the training text group corresponding to the result text is an effective training task, and the corresponding training text group is non-noise data. Furthermore, the second training loss value can be determined based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text.

[0176] In other words, the second training loss value can be determined based on the first true probability that each word in the result text belongs to the target text, the second true probability that each word does not belong to the target text, the third prediction probability that each word belongs to the target text, and the sixth prediction probability that each word does not belong to the target text. Specifically, it can be determined based on the following expression:

[0177]

[0178] where, k is the index of the training text group, N is the number of words in the result text and M is the number of words in the query text represents the i-th word in the result text corresponding to the k-th training text group.

[0179] Among them, represents the third prediction probability that the i-th word belongs to the target text in the result text, represents the first true probability that the i-th word belongs to the target text in the result text. represents the second true probability that the i-th word does not belong to the target text in the result text, represents the sixth prediction probability that the i-th word does not belong to the target text in the result text,

[0180] Among them, the above-mentioned second training loss value characterizes the differences between the true probabilities and prediction probabilities of each word in the result text belonging to the target text, and the differences between the true probabilities and prediction probabilities of each word not belonging to the target text. The smaller the second training loss value, the smaller the differences between the true probabilities and prediction probabilities of each word in the result text belonging to the target text, and the smaller the differences between the true probabilities and prediction probabilities of each word not belonging to the target text. Furthermore, it shows that the difference between the target text in the result text predicted by the initial model and the target text in the actual result text is smaller.

[0181] Furthermore, after determining the second training loss value corresponding to the context prediction training task based on the training text group, the total training loss value can be determined based on the first training loss value and the second training loss value. Specifically, the first weight corresponding to the first training loss value and the second weight corresponding to the second training loss value can be determined, and then the total training loss value can be determined based on the first training loss and the corresponding first weight, as well as the second weight corresponding to the second training loss value. Specifically, it can be determined based on the following expression:

[0182]

[0183] Among them, is the first training loss value, (1 - λ) is the first weight corresponding to the first training loss value, is the second training loss value, γ is the second weight corresponding to the second training loss value, is the total training loss value. And the above-mentioned first weight and second weight can be specifically determined based on the requirements of the actual application scenario, and are not limited here.

[0184] Among them, represents the query text, represents the result text, k is the index of the training text group, represents the k-th training text group, D A,1 represents non-noise data represents that the corresponding text relevance is greater than or equal to the relevance threshold.

[0185] On the other hand, for each result text, if the text relevance corresponding to the result text is less than the above-mentioned relevance threshold, it indicates that the similarity between the text content of the result text and the text content of the corresponding query text is relatively low during this training process. Furthermore, it can be shown that the relevance between the target text in the result text and the corresponding answer text is relatively low. Therefore, it is demonstrated that the training task of context prediction based on the training text group corresponding to the result text is an ineffective training task, and the corresponding training text group is noise data. Consequently, the first training loss value can be determined as the final total training loss value. Specifically, it can be determined based on the following expression:

[0186]

[0187] where, is the first training loss value, is the total training loss value. Among them, represents the query text, represents the result text, k is the index of the training text group, represents the k-th training text group, D A,2 represents the noise data, represents that the corresponding text relevance is less than the relevance threshold.

[0188] Step S24: Iteratively train the initial model based on the total training loss value until the total training loss value meets the preset training end condition, and then determine the text prediction model based on the model after training ends.

[0189] In some feasible implementation manners, after determining the total training loss value corresponding to the training of the initial model, the initial model can be iteratively trained based on the total training loss value until the training stops when the preset training end condition is met.

[0190] Among them, the above-mentioned training end loss can be that the total training loss value reaches convergence, or the total training loss values of a continuous preset number are less than the preset threshold, which can be specifically determined based on the requirements of the actual application scenario and will not be limited here.

[0191] For each training, when the total training loss value meets the preset training end condition, the training can be stopped and the final text prediction model can be determined based on the model at the time of stopping training. If the total training loss value does not meet the preset training end condition, the relevant parameters of the model can be adjusted, and the adjusted model can be trained again until the total training loss value meets the preset training end condition and the training stops.

[0192] Among them, when determining the final text prediction model based on the model at the time of stopping training, relevant parameters of the model at the time of stopping training can be fine-tuned to adapt to different application scenarios. For example, corresponding adjustments are made for search scenarios, question-and-answer scenarios, etc. respectively to obtain the final text prediction model.

[0193] In some feasible implementation manners, when training the initial model based on the training text group, two initial models with different initial model parameters can be alternately trained simultaneously, and then based on the model with a higher text prediction effect after each training ends, the text prediction model is obtained.

[0194] See Figure 6 , Figure 6 is a schematic diagram of a scenario of the training method of the text model provided in the embodiments of the present application. As Figure 6 shown, the initial model A and the initial model B are respectively models with different initial model parameters. When performing model training, multiple training text groups can be divided into two training text group sets. For the training text group set 1, based on this set, a training task of context prediction is performed on the initial model A to train the initial model A to identify noise data (that is, to judge the relevance between the result text and the corresponding query text in each training text group). Further, based on each training text group in the training text group set 1, a training task of answer prediction is performed on the initial model B, and the total training loss value is determined based on the noise data and non-noise data identified by the initial model A, and the model parameters of the initial model A are updated based on the total training loss value.

[0195] At the same time, for the training text group set 2, based on this set, a training task of context prediction is performed on the initial model B to train the initial model B to identify noise data (that is, to judge the relevance between the result text and the corresponding query text in each training text group). Further, based on each training text group in the training text group set 2, a training task of answer prediction is performed on the initial model A, and the total training loss value is determined based on the noise data and non-noise data identified by the initial model B, and the model parameters of the initial model B are updated based on the total training loss value.

[0196] Based on the above method, the model training of the initial model A and the initial model B can be completed, and then the obtained final models are tested to determine the final text prediction model based on the model with a better text prediction effect.

[0197] In some feasible implementation manners, the initial model can also be trained for a training task based on the training text group first, and then trained for another training task after obtaining a stable model, and the final text prediction model is obtained based on the model at the end of the training.

[0198] Participate in Figure 7 ,Figure 7 This is another schematic diagram of the scenario of the training method of the text model provided by the embodiments of the present application. In Figure 7 , the initial model can be first trained for the context prediction task based on the training text group, and the model parameters of the initial model can be updated based on the second training loss value generated in this training task, and an intermediate model can be obtained when the second training loss value meets the corresponding training end condition. The intermediate model is trained for the answer prediction task using the training text group, and the model parameters of the intermediate model are updated based on the first training loss value generated in this training task, and the training is stopped when the first training loss value meets the corresponding training end condition, and the final text prediction model is obtained based on the model at the time of stopping training.

[0199] In some feasible embodiments, the calculation methods involved in the first training loss value, the second training loss value, the total training loss value, and the text relevance corresponding to the result text in the embodiments of the present application can be performed based on cloud computing and other means.

[0200] Among them, cloud computing refers to obtaining the required resources in a demand-based and easily expandable manner through the network, and is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance.

[0201] In the embodiments of the present application, the position probability of each word in the result text corresponding to the position of the answer text in each training text group can be determined through the answer prediction training task in the model training process, so that determining the first training loss value based on the position probability and the answer text can better characterize the difference between the predicted position of each word in the result text corresponding to the answer text and the position of the answer text in the result text, thereby improving the training effect. On the other hand, through the context prediction training task in the model training process, the model can be enabled to have the ability to determine whether the result text contains context related to the answer text. Furthermore, when the model is iteratively trained based on the loss value corresponding to the context prediction training task and the training loss value in the answer prediction training task, the model adjustment effect can be further improved, so that when text prediction is performed based on the trained text prediction model, it has high text prediction accuracy and high applicability.

[0202] In the embodiments of the application, the text prediction effect of the text prediction model obtained in the embodiments of the present application and the text prediction effects of other models can be seen in Table 1.

[0203] Table 1: Comparison of text prediction effects of different models

[0204]

[0205] In Table 1, F1 represents the harmonic mean of the precision and recall corresponding to the model. Precision represents the proportion of words in the answer text predicted by the model that belong to the standard answer, and recall represents the proportion of words in the standard answer that are in the answer text predicted by the model. EM (Exact Match) represents the proportion of samples in the test set where the answer text predicted by the text is exactly the same as the standard answer. Among them, the higher the F1 value and EM value, the better the text prediction effect of the model.

[0206] As can be seen from Table 1, the text prediction model in this application has a better text prediction effect compared to the model trained based on BERT, the model obtained by using the same training method as this application but taking the second training loss value as the total training loss value, the model obtained by using the same training method as this application but taking the first training loss value as the total training loss value, and the text obtained by using the same training method as this application but not determining the total training loss value based on whether the result text in the training text group is noise text.

[0207] See Figure 8 , Figure 8 is a schematic structural diagram of the text prediction device provided by the embodiment of the present application. The text prediction device 1 provided by the embodiment of the present application includes:

[0208] An acquisition module 11, configured to acquire a target query text and a target result text corresponding to the target query text;

[0209] A prediction module 12, configured to determine a target answer text from the target result text through a text prediction model based on the target query text and the target result text;

[0210] The device includes a training module 13, and the training module 13 includes:

[0211] An acquisition unit 131, configured to acquire a plurality of training text groups, each of the training text groups including a query text, a result text of the query text, and an answer text corresponding to the query text, and each of the result texts including a corresponding answer text;

[0212] A prediction unit 132 is configured to determine the encoded features corresponding to each of the above result texts and the corresponding query texts, input the encoded features into the initial model, and through the initial model, determine the position probabilities of each word in each of the above result texts corresponding to the position of the answer text in the corresponding result text;

[0213] A determination unit 133 is configured to, for each of the above result texts, determine a first training loss value corresponding to the result text based on the position probability and the answer text corresponding to the result text, and determine a total training loss value based on the first training loss value;

[0214] A training unit 134 is configured to perform iterative training on the initial model based on the total training loss value until the total training loss value meets a preset training end condition, and determine the text prediction model based on the model after training ends.

[0215] In some feasible embodiments, for each of the above result texts, the position probability corresponding to the result text includes a first prediction probability that each word in the result text is the starting position of the answer text in the result text and a second prediction probability that is the ending position.

[0216] The above determination unit 133 is configured to:

[0217] Determine the starting position and the ending position of the answer text in the result text;

[0218] Based on the first prediction probability corresponding to the first target word and the second prediction probability corresponding to the second target word in the result text, determine the first training loss value corresponding to the result text;

[0219] Wherein, the first target word corresponds to the starting position of the answer text in the result text, and the second target word corresponds to the ending position of the answer text in the result text.

[0220] In some feasible embodiments, for each of the above result texts, the above determination unit 133 is configured to:

[0221] Determine a first true probability that each word in the result text belongs to the target text in the result text;

[0222] Through the initial model, determine a third prediction probability that each word in the result text belongs to the target text in the result text based on the encoded features corresponding to the result text, wherein the target text in the result text is the text associated with the corresponding answer text;

[0223] Based on the third prediction probability and the first true probability corresponding to the result text, determine the text relevance between the result text and the corresponding query text;

[0224] Determine the total training loss value based on the text relevance corresponding to the result text and the above first training loss value.

[0225] In some feasible embodiments, for each of the above result texts, the determining unit 133 is configured to:

[0226] Determine the starting position and the ending position of the answer text in the result text;

[0227] For each word in the result text, based on the position of the word in the result text, as well as the starting position and the ending position of the answer text in the result text, determine a first true probability that the word belongs to the target text in the result text.

[0228] In some feasible embodiments, for each of the above result texts, the determining unit 133 is configured to:

[0229] For each word in the result text, based on the encoded features corresponding to the result text, determine a fourth prediction probability that the word is the starting position of the target text in the result text, and a fifth prediction probability that it is the ending position.

[0230] Based on the fourth prediction probabilities and the fifth prediction probabilities corresponding to the words in the result text, determine a third prediction probability that each word in the result text belongs to the target text in the result text.

[0231] In some feasible embodiments, for each of the above result texts, the determining unit 133 is configured to:

[0232] For each word in the result text, determine a third target word whose position in the result text is before the word, and a fourth target word whose position is after the word;

[0233] For each word in the result text, based on the fourth prediction probability corresponding to the word, the fourth prediction probability corresponding to the third target word corresponding to the word, the fifth prediction probability corresponding to the word, and the fifth prediction probability corresponding to the fourth target word corresponding to the word, determine a third prediction probability that the word belongs to the target text in the result text.

[0234] In some feasible embodiments, for each of the above result texts, the determining unit 133 is configured to:

[0235] Based on the first true probability corresponding to the result text, determine a second true probability that each word in the result text does not belong to the target text in the result text;

[0236] Based on the third prediction probability corresponding to the result text, determine a sixth prediction probability that each word in the result text does not belong to the target text in the result text;

[0237] Determine the text relevance between the result text and the corresponding query text based on the first true probability, the third predicted probability, the second true probability, and the sixth predicted probability corresponding to the result text.

[0238] In some feasible embodiments, for each of the above result texts, the determining unit 133 is configured to:

[0239] Determine a text relevance threshold based on the text relevance corresponding to each of the above result texts;

[0240] If the text relevance corresponding to the result text is greater than or equal to the above relevance threshold, determine a second training loss value based on the first true probability, the third predicted probability, the second true probability, and the sixth predicted probability corresponding to the result text, and determine a total training loss value based on the above first training loss value and the above second training loss value;

[0241] If the text relevance corresponding to the result text is less than the above relevance threshold, determine the above first training loss value as the above total training loss value.

[0242] In some feasible embodiments, for each of the above result texts, the determining unit 133 is configured to:

[0243] Determine a first weight corresponding to the above first training loss value and a second weight corresponding to the above second training loss value;

[0244] Determine a total training loss value based on the above first training loss value and the corresponding first weight, and the above second training loss value and the corresponding second weight.

[0245] In some feasible embodiments, for each of the above result texts, the determining unit 133 is configured to:

[0246] Determine the sum of the text relevances of the text relevances corresponding to each of the above result texts;

[0247] Determine the number of each of the above result texts, and determine a text relevance threshold based on the number of each of the above result texts and the above sum of text relevances.

[0248] In some feasible embodiments, the obtaining unit 131 is further configured to:

[0249] Determine a plurality of texts to be processed, and for each of the above texts to be processed, determine a training text group corresponding to the text to be processed based on the following method:

[0250] Determine the text corresponding to any text interval in the text to be processed as the answer text, and determine the query text corresponding to the answer text based on the other texts in the text to be processed except the answer text;

[0251] Determine multiple retrieval texts for the query text, determine the text similarity between the query text and each of the above-mentioned retrieval texts, and determine the retrieval text with the highest text similarity as the result text corresponding to the query text;

[0252] Determine the query text, result text, and answer text corresponding to the text to be processed as a training text group.

[0253] In some feasible implementation manners, the above prediction module 12 is used for:

[0254] Determine the target encoding features corresponding to the above target query text and the above target result text;

[0255] Input the above target encoding features into a text prediction model to obtain a first target probability that each word in the above target result text is the starting position of the target answer text corresponding to the above target query text, and a second target probability of the ending position;

[0256] Based on the first target probability and the second target probability corresponding to each word in the above target result text, determine the above target answer text from the above target result text.

[0257] The above text prediction device may be a computer program (including program code) running in a computer device, and can execute the implementation manners provided in each of the above Figure 1 、 Figure 2 and / or Figure 5 by each of the functional modules built therein. For specific implementation manners, reference may be made to the implementation manners provided in each of the above steps, which will not be elaborated herein.

[0258] In some feasible implementation manners, the text prediction device provided in the embodiments of the present application may be implemented in a software manner. The text prediction device provided in the embodiments of the present application may be software in the form of a program and a plug-in, etc., and includes a series of modules, including an acquisition module 11, a prediction module 12, and a text prediction module 13. Among them, the acquisition module 11, the prediction module 12, and the text prediction module 13 are used to implement the text prediction method provided in the embodiments of the present application.

[0259] In some embodiments, the text prediction device provided by the embodiments of the present application can be implemented in a combination of software and hardware. As an example, the text prediction device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the text prediction method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0260] See Figure 9 , Figure 9 is a schematic structural diagram of an electronic device provided by the embodiments of the present application. As Figure 9 shown, the electronic device 1000 in this embodiment may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above-mentioned electronic device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to implement connection communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1004 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally be at least one storage device located far from the aforementioned processor 1001. As Figure 9 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0261] In Figure 9 the electronic device 1000 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to achieve:

[0262] Obtain a target query text and a target result text corresponding to the target query text;

[0263] Based on the above-mentioned target query text and the above-mentioned target result text, determine the target answer text from the above-mentioned target result text through a text prediction model;

[0264] Obtain multiple training text groups, each of the above-mentioned training text groups includes a query text, the result text of the above-mentioned query text, and the answer text corresponding to the above-mentioned query text, and each of the above-mentioned result texts includes the corresponding answer text;

[0265] Determine the encoding features corresponding to each of the above-mentioned result texts and the corresponding query texts, input the above-mentioned encoding features into the initial model, and through the above-mentioned initial model, determine the position probability of each word in each of the above-mentioned result texts corresponding to the position of the answer text in the corresponding result text;

[0266] For each of the above-mentioned result texts, based on the position probability and the answer text corresponding to the result text, determine the first training loss value corresponding to the result text, and determine the total training loss value based on the above-mentioned first training loss value;

[0267] Iteratively train the above-mentioned initial model based on the above-mentioned total training loss value until the above-mentioned total training loss value meets the preset training end condition, and determine the above-mentioned text prediction model based on the model after training ends.

[0268] In some feasible implementation manners, for each of the above-mentioned result texts, the position probability corresponding to the result text includes the first prediction probability that each word in the result text is the starting position of the answer text in the result text, and the second prediction probability of the ending position;

[0269] The above-mentioned processor 1001 is used for:

[0270] Determine the starting position and the ending position of the answer text in the result text;

[0271] Based on the first prediction probability corresponding to the first target word in the result text and the second prediction probability corresponding to the second target word, determine the first training loss value corresponding to the result text;

[0272] Wherein, the above-mentioned first target word corresponds to the starting position of the answer text in the result text, and the above-mentioned second target word corresponds to the ending position of the answer text in the result text.

[0273] In some feasible implementation manners, for each of the above-mentioned result texts, the above-mentioned processor 1001 is used for:

[0274] Determine the first true probability that each word in the result text belongs to the target text in the result text;

[0275] Based on the above initial model, determine the third prediction probability that each word in the result text belongs to the target text in the result text based on the encoding features corresponding to the result text, where the target text in the result text is the text associated with the corresponding answer text;

[0276] Based on the third prediction probability and the first true probability corresponding to the result text, determine the text relevance between the result text and the corresponding query text;

[0277] Based on the text relevance corresponding to the result text and the above first training loss value, determine the total training loss value.

[0278] In some feasible embodiments, for each of the above result texts, the processor 1001 is configured to:

[0279] Determine the start position and end position of the answer text in the result text;

[0280] For each word in the result text, based on the position of the word in the result text, and the start position and end position of the answer text in the result text, determine the first true probability that the word belongs to the target text in the result text.

[0281] In some feasible embodiments, for each of the above result texts, the processor 1001 is configured to:

[0282] For each word in the result text, based on the encoding features corresponding to the result text, determine the fourth prediction probability that the word is the start position of the target text in the result text, and the fifth prediction probability that the word is the end position of the target text in the result text;

[0283] Based on the fourth prediction probability and the fifth prediction probability corresponding to each word in the result text, determine the third prediction probability that each word in the result text belongs to the target text in the result text.

[0284] In some feasible embodiments, for each of the above result texts, the processor 1001 is configured to:

[0285] For each word in the result text, determine the third target word whose position in the result text is before the word, and the fourth target word whose position in the result text is after the word;

[0286] For each word in the result text, based on the fourth prediction probability corresponding to the word, the fourth prediction probability corresponding to the third target word corresponding to the word, the fifth prediction probability corresponding to the word, and the fifth prediction probability corresponding to the fourth target word corresponding to the word, determine the third prediction probability that the word belongs to the target text in the result text.

[0287] In some feasible embodiments, for each of the above result texts, the above processor 1001 is configured to:

[0288] Based on the first true probability corresponding to the result text, determine the second true probability that each word in the result text does not belong to the target text in the result text;

[0289] Based on the third prediction probability corresponding to the result text, determine the sixth prediction probability that each word in the result text does not belong to the target text in the result text;

[0290] Based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text, determine the text relevance between the result text and the corresponding query text.

[0291] In some feasible embodiments, for each of the above result texts, the above processor 1001 is configured to:

[0292] Based on the text relevance corresponding to each of the above result texts, determine a text relevance threshold;

[0293] If the text relevance corresponding to the result text is greater than or equal to the above relevance threshold, then determine a second training loss value based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text, and determine a total training loss value based on the above first training loss value and the above second training loss value;

[0294] If the text relevance corresponding to the result text is less than the above relevance threshold, then determine the above first training loss value as the above total training loss value.

[0295] In some feasible embodiments, the above processor 1001 is configured to:

[0296] Determine a first weight corresponding to the above first training loss value and a second weight corresponding to the above second training loss value;

[0297] Based on the above first training loss value and the corresponding first weight, and the above second training loss value and the corresponding second weight, determine a total training loss value.

[0298] In some feasible embodiments, the above processor 1001 is configured to:

[0299] Determine the sum of the text relevance of the text relevance corresponding to each of the above result texts;

[0300] Determine the number of each of the above result texts, and based on the number of each of the above result texts and the above sum of text relevance, determine a text relevance threshold.

[0301] In some feasible embodiments, the above-mentioned processor 1001 is further configured to:

[0302] Determine a plurality of texts to be processed. For each of the above-mentioned texts to be processed, determine a training text group corresponding to the text to be processed based on the following method:

[0303] Determine the text corresponding to any text interval in the text to be processed as the answer text, and based on the other texts in the text to be processed except the answer text, determine the query text corresponding to the answer text;

[0304] Determine a plurality of retrieval texts for the query text, determine the text similarity between the query text and each of the above-mentioned retrieval texts, and determine the retrieval text with the highest text similarity as the result text corresponding to the query text;

[0305] Determine the query text, the result text, and the answer text corresponding to the text to be processed as a training text group.

[0306] In some feasible embodiments, the above-mentioned processor 1001 is configured to:

[0307] Determine the target encoding features corresponding to the above-mentioned target query text and the above-mentioned target result text;

[0308] Input the above-mentioned target encoding features into the text prediction model to obtain the first target probability that each word in the above-mentioned target result text is the starting position of the target answer text corresponding to the above-mentioned target query text, and the second target probability of the ending position;

[0309] Based on the first target probability and the second target probability corresponding to each word in the above-mentioned target result text, determine the above-mentioned target answer text from the above-mentioned target result text.

[0310] It should be understood that in some feasible embodiments, the above-mentioned processor 1001 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0311] In a specific implementation, the above-mentioned electronic device 1000 may execute the implementation manners provided in each of the steps as described above through its built-in various functional modules. Figure 2 For details, reference may be made to the implementation manners provided in each of the above steps, which will not be elaborated herein.

[0312] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program that is executed by a processor to implement Figure 1 、 Figure 2 and / or Figure 5 the methods provided in each of the steps, and for details, reference may be made to the implementation manners provided in each of the above steps, which will not be elaborated herein.

[0313] The above computer-readable storage medium may be an internal storage unit of any of the foregoing text prediction devices or electronic devices, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. The above computer-readable storage medium may further include magnetic disks, optical discs, read-only memory (ROM), or random access memory (RAM), etc. Further, the computer-readable storage medium may include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0314] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 , Figure 2 and / or Figure 5 the methods provided by each step in

[0315] The terms "first", "second", etc. in the claims, the description, and the drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or electronic device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or electronic devices. Referring to "embodiment" in this article means that a specific feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. Displaying this phrase at various positions in the description does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments. The term "and / or" used in the description and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0316] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0317] The above-disclosed are only the preferred embodiments of this application, and thus cannot be used to limit the scope of rights of this application. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.

Claims

1. A text prediction method, characterized in that, The method includes: Obtaining a target query text and a target result text corresponding to the target query text; Based on the target query text and the target result text, determining a target answer text from the target result text through a text prediction model; Wherein, the text prediction model is trained based on the following method: Obtaining a plurality of training text groups, each of the training text groups including a query text, a result text of the query text, and an answer text corresponding to the query text, and each of the result texts including a corresponding answer text; Determining encoding features corresponding to each of the result texts and the corresponding query texts, inputting the encoding features into an initial model, and through the initial model, determining the position probability of each word in each of the result texts corresponding to the position where the answer text is located in the corresponding result text; For each of the result texts, based on the position probability corresponding to the result text and the true probability that each word in the result text is the start position and the end position of the answer text in the result text, determining a first training loss value corresponding to the result text, wherein the position probability corresponding to each of the result texts includes a first prediction probability that each word in the result text is the start position of the answer text in the result text and a second prediction probability that it is the end position; For each of the result texts, determining a first true probability that each word in the result text belongs to the target text in the result text; through the initial model, based on the encoding features corresponding to the result text, determining a third prediction probability that each word in the result text belongs to the target text in the result text, wherein the target text in the result text is a text associated with the corresponding answer text; based on the third prediction probability and the first true probability, determining the text relevance between the result text and the corresponding query text; based on the text relevance corresponding to the result text and the first training loss value, determining a total training loss value; Performing iterative training on the initial model based on the total training loss value until the total training loss value meets a preset training end condition, and determining the text prediction model based on the model after the training ends.

2. The method according to claim 1, wherein For each of the result texts, the determining a first training loss value corresponding to the result text based on the position probability corresponding to the result text and the true probability that each word in the result text is the start position and the end position of the answer text in the result text includes: Determining the start position and the end position of the answer text in the result text; Based on the first prediction probability corresponding to a first target word in the result text, the second prediction probability corresponding to a second target word, the true probability that the first target word is the start position of the answer text in the result text, and the true probability that the second target word is the end position of the answer text in the result text, determining a first training loss value corresponding to the result text; Wherein, the first target word corresponds to the start position of the answer text in the result text, and the second target word corresponds to the end position of the answer text in the result text.

3. The method according to claim 1, wherein For each of the result texts, determining the first true probability that each word in the result text belongs to the target text in the result text includes: Determining the start position and end position of the answer text in the result text; For each word in the result text, based on the position of the word in the result text, and the start position and end position of the answer text in the result text, determining the first true probability that the word belongs to the target text in the result text.

4. The method according to claim 1, characterized in that, For each of the result texts, determining the third predicted probability that each word in the result text belongs to the target text in the result text based on the encoded features corresponding to the result text includes: For each word in the result text, based on the encoded features corresponding to the result text, determining the fourth predicted probability that the word is the start position of the target text in the result text, and the fifth predicted probability that it is the end position; Based on the fourth predicted probability and the fifth predicted probability corresponding to each word in the result text, determining the third predicted probability that each word in the result text belongs to the target text in the result text.

5. The method according to claim 4, wherein For each of the result texts, determining the third predicted probability that each word in the result text belongs to the target text in the result text based on the fourth predicted probability and the fifth predicted probability corresponding to each word in the result text includes: For each word in the result text, determining the third target word before the word in the result text, and the fourth target word after the word; For each word in the result text, based on the fourth predicted probability corresponding to the word, the fourth predicted probability corresponding to the third target word corresponding to the word, the fifth predicted probability corresponding to the word, and the fifth predicted probability corresponding to the fourth target word corresponding to the word, determining the third predicted probability that the word belongs to the target text in the result text.

6. The method according to claim 1, wherein For each of the result texts, determining the text relevance between the result text and the corresponding query text based on the third predicted probability and the first true probability corresponding to the result text includes: Based on the first true probability corresponding to the result text, determining the second true probability that each word in the result text does not belong to the target text in the result text; Based on the third predicted probability corresponding to the result text, determining the sixth predicted probability that each word in the result text does not belong to the target text in the result text; Based on the first true probability, the third predicted probability, the second true probability, and the sixth predicted probability corresponding to the result text, determining the text relevance between the result text and the corresponding query text.

7. The method according to claim 6, characterized in that For each of the result texts, determining the total training loss value based on the text relevance corresponding to the result text and the first training loss value includes: Based on the text relevance corresponding to each of the result texts, determining a text relevance threshold; If the text relevance corresponding to the result text is greater than or equal to the relevance threshold, then based on the first true probability, the third predicted probability, the second true probability, and the sixth predicted probability corresponding to the result text, determining a second training loss value, and based on the first training loss value and the second training loss value, determining the total training loss value; If the text relevance corresponding to the result text is less than the relevance threshold, the first training loss value is determined as the total training loss value.

8. The method according to claim 7, wherein The determining based on the first training loss value and the second training loss value includes: Determining a first weight corresponding to the first training loss value and a second weight corresponding to the second training loss value; Based on the first training loss value and the corresponding first weight, and the second training loss value and the corresponding second weight, determining the total training loss value.

9. The method according to claim 7, characterized in that, Determining the text relevance threshold based on the text relevance corresponding to each result text includes: Determining the sum of the text relevance of the text relevance corresponding to each result text; Determining the number of each result text, and based on the number of each result text and the sum of the text relevance, determining the text relevance threshold.

10. The method according to claim 1, wherein The method further includes: Determining a plurality of texts to be processed, and for each text to be processed, determining a training text group corresponding to the text to be processed based on the following method: Determining the text corresponding to any text interval in the text to be processed as the answer text, and based on the other texts in the text to be processed except the answer text, determining the query text corresponding to the answer text; Determining a plurality of retrieval texts of the query text, determining the text similarity between the query text and each retrieval text, and determining the retrieval text with the highest text similarity as the result text corresponding to the query text; Determining the query text, the result text, and the answer text corresponding to the text to be processed as a training text group.

11. The method according to claim 1, characterized in that, The determining the target answer text from the target result text through the text prediction model based on the target query text and the target result text includes: Determining the target encoding features corresponding to the target query text and the target result text; Inputting the target encoding features into the text prediction model to obtain a first target probability that each word in the target result text is the starting position of the target answer text corresponding to the target query text, and a second target probability of the ending position; Based on the first target probability and the second target probability corresponding to each word in the target result text, determining the target answer text from the target result text.

12. A text prediction device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a target query text and a target result text corresponding to the target query text; A prediction module, configured to determine a target answer text from the target result text through a text prediction model based on the target query text and the target result text; The apparatus includes a training module, and the training module includes: An acquisition unit, configured to acquire a plurality of training text groups, each training text group including a query text, a result text of the query text, and an answer text corresponding to the query text, and each result text including a corresponding answer text; A prediction unit, configured to determine the encoding features corresponding to each of the result texts and the corresponding query texts, input the encoding features into an initial model, and through the initial model, determine the position probabilities of each word in each of the result texts corresponding to the position of the answer text in the corresponding result text; for the position probabilities corresponding to each of the result texts, it includes the first prediction probability that each word in the result text is the starting position of the answer text in the result text, and the second prediction probability that it is the ending position. A determination unit, configured to, for each of the result texts, based on the position probabilities corresponding to the result text and the true probabilities that each word in the result text is the starting position and the ending position of the answer text in the result text, determine the first training loss value corresponding to the result text, where the position probabilities corresponding to each of the result texts include the first prediction probability that each word in the result text is the starting position of the answer text in the result text, and the second prediction probability that it is the ending position. The determination unit is configured to, for each of the result texts, determine the first true probability that each word in the result text belongs to the target text in the result text; through the initial model, based on the encoding features corresponding to the result text, determine the third prediction probability that each word in the result text belongs to the target text in the result text, where the target text in the result text is the text associated with the corresponding answer text; based on the third prediction probability and the first true probability, determine the text relevance between the result text and the corresponding query text; based on the text relevance corresponding to the result text and the first training loss value, determine the total training loss value. A training unit, configured to perform iterative training on the initial model based on the total training loss value, and when the total training loss value meets the preset training end condition, determine the text prediction model based on the model after training ends.

13. The device according to claim 12, wherein, For each of the result texts, when the determination unit determines the first training loss value corresponding to the result text based on the position probabilities corresponding to the result text and the true probabilities that each word in the result text is the starting position and the ending position of the answer text in the result text, it is configured to: Determine the starting position and the ending position of the answer text in the result text. Based on the first prediction probability corresponding to the first target word in the result text, the second prediction probability corresponding to the second target word, the true probability that the first target word is the starting position of the answer text in the result text, and the true probability that the second target word is the ending position of the answer text in the result text, determine the first training loss value corresponding to the result text. Wherein, the first target word corresponds to the starting position of the answer text in the result text, and the second target word corresponds to the ending position of the answer text in the result text.

14. The device according to claim 12, characterized in that, For each of the result texts, when the determination unit determines the first true probability that each word in the result text belongs to the target text in the result text, it is configured to: Determine the starting position and the ending position of the answer text in the result text. For each word in the result text, based on the position of the word in the result text and the start and end positions of the answer text in the result text, determine the first true probability that the word belongs to the target text in the result text.

15. The device according to claim 12, characterized in that, For each of the result texts, when the determining unit determines the third prediction probability that each word in the result text belongs to the target text in the result text based on the encoding feature corresponding to the result text, it is used for: For each word in the result text, based on the encoding feature corresponding to the result text, determine the fourth prediction probability that the word is the start position of the target text in the result text and the fifth prediction probability that it is the end position. Based on the fourth prediction probability and the fifth prediction probability corresponding to each word in the result text, determine the third prediction probability that each word in the result text belongs to the target text in the result text.

16. The device according to claim 15, characterized in that, For each of the result texts, when the determining unit determines the third prediction probability that each word in the result text belongs to the target text in the result text based on the fourth prediction probability and the fifth prediction probability corresponding to each word in the result text, it is used for: For each word in the result text, determine the third target word before the word in the result text and the fourth target word after the word. For each word in the result text, based on the fourth prediction probability corresponding to the word, the fourth prediction probability corresponding to the third target word corresponding to the word, the fifth prediction probability corresponding to the word, and the fifth prediction probability corresponding to the fourth target word corresponding to the word, determine the third prediction probability that the word belongs to the target text in the result text.

17. The device according to claim 12, characterized in that, For each of the result texts, when the determining unit determines the text relevance between the result text and the corresponding query text based on the third prediction probability and the first true probability corresponding to the result text, it is used for: Based on the first true probability corresponding to the result text, determine the second true probability that each word in the result text does not belong to the target text in the result text. Based on the third prediction probability corresponding to the result text, determine the sixth prediction probability that each word in the result text does not belong to the target text in the result text. Based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text, determine the text relevance between the result text and the corresponding query text.

18. The device according to claim 17, characterized in that, For each of the result texts, when the determining unit determines the total training loss value based on the text relevance corresponding to the result text and the first training loss value, it is used for: Based on the text relevance corresponding to each of the result texts, determine the text relevance threshold. If the text relevance corresponding to the result text is greater than or equal to the relevance threshold, then determine the second training loss value based on the first true probability, the third prediction probability, the second true probability, and the sixth prediction probability corresponding to the result text, and determine the total training loss value based on the first training loss value and the second training loss value. If the text relevance corresponding to the result text is less than the relevance threshold, then determine the first training loss value as the total training loss value.

19. The device according to claim 18, wherein, When the determining unit is based on the first training loss value and the second training loss value, it is used to: Determine a first weight corresponding to the first training loss value and a second weight corresponding to the second training loss value; Based on the first training loss value and the corresponding first weight, and the second training loss value and the corresponding second weight, determine the total training loss value.

20. The device according to claim 18, wherein, When the determining unit determines the text relevance threshold based on the text relevance corresponding to each of the result texts, it is used to: Determine the sum of the text relevance of the text relevance corresponding to each of the result texts; Determine the number of each of the result texts, and based on the number of each of the result texts and the sum of the text relevance, determine the text relevance threshold.

21. The device according to claim 12, characterized in that The obtaining unit is further used to: Determine a plurality of texts to be processed, and for each of the texts to be processed, determine a training text group corresponding to the text to be processed based on the following method: Determine the text corresponding to any text interval in the text to be processed as the answer text, and based on the other texts in the text to be processed except the answer text, determine the query text corresponding to the answer text; Determine a plurality of retrieval texts for the query text, determine the text similarity between the query text and each of the retrieval texts, and determine the retrieval text with the highest text similarity as the result text corresponding to the query text; Determine the query text, the result text, and the answer text corresponding to the text to be processed as a training text group.

22. The device according to claim 12, characterized in that, When the prediction module determines the target answer text from the target result text through the text prediction model based on the target query text and the target result text, it is used to: Determine the target encoding features corresponding to the target query text and the target result text; Input the target encoding features into the text prediction model to obtain a first target probability that each word in the target result text is the starting position of the target answer text corresponding to the target query text, and a second target probability of the ending position; Based on the first target probability and the second target probability corresponding to each word in the target result text, determine the target answer text from the target result text.

23. An electronic device, characterized in that, Comprising a processor and a memory, the processor and the memory are interconnected; The memory is used to store a computer program; The processor is configured to execute the method according to any one of claims 1 to 11 when calling the computer program.

24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 11.

25. A computer program product, the computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Machine reading model training method and device, question and answer method and device

    CN108959396A

  • Question and answer method, computing equipment and storage medium

    CN112417126A