Text processing methods, human-computer interaction methods, devices and storage media
By acquiring the prompt vector and prompt text of the text to be processed from the interactive device, and inputting them into the question-and-answer model to generate accurate response text, the problem of inaccurate responses from interactive devices in complex scenarios is solved, and efficient processing of various question-and-answer types is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2023-04-19
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, interactive devices struggle to accurately respond to user questions during human-computer dialogue, especially in scenarios with complex logic where the output responses are often inaccurate.
By obtaining the prompt vector and prompt text of the text to be processed, the prompt is input into the question-answering model to generate accurate response text. The prompt information in vector form and prompt information in text form are used to supplement the knowledge of the question-answering model and improve the accuracy of the response.
It improves the accuracy of the question-answering model in generating response text, and can handle different types of question-answering scenarios, including extractive, multiple-choice, summary, and judgmental question-answering, making up for the open-world problem caused by insufficient training samples.
Smart Images

Figure CN116450792B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a text processing method, a human-computer interaction method, a device, and a storage medium. Background Technology
[0002] Artificial intelligence technology has enabled human-computer dialogue, which can be applied to various scenarios. For example, interactive devices such as service robots, smart speakers, and mobile terminals can facilitate content searching and information retrieval. Service robots, for instance, can be deployed in shopping malls and restaurants as guide robots. In this scenario, the user and the interactive device typically engage in multi-turn question-and-answer dialogues, and the content output by the device is relatively simple. Another example is interactive devices with pre-installed software that can be used for tasks such as article writing and test preparation. In this scenario, the interactive devices can be various mobile terminals used by the user, and the content output is more complex.
[0003] In all the scenarios described above, how to enable interactive devices to accurately respond to user questions becomes an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a text processing method, a human-computer interaction method, a device, and a storage medium, for enabling the interactive device to accurately respond to user questions.
[0005] In a first aspect, embodiments of the present invention provide a text processing method, including:
[0006] Get the text to be processed;
[0007] Determine the prompt vector and prompt text corresponding to the text to be processed;
[0008] The text to be processed, the prompt vector, and the prompt text are input into the question-answering model, so that the question-answering model outputs the response text corresponding to the text to be processed.
[0009] Secondly, embodiments of the present invention provide a human-computer interaction method, including:
[0010] In response to interactive operations on the interactive device, acquire the text to be processed;
[0011] Determine the prompt vector and prompt text corresponding to the text to be processed;
[0012] The text to be processed, the prompt vector, and the prompt text are input into the question-answering model so that the question-answering model can determine the response text corresponding to the text to be processed.
[0013] Output the response text.
[0014] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the text processing method in the first aspect or the human-computer interaction method in the second aspect. The electronic device may also include a communication interface for communicating with other devices or communication systems.
[0015] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, wherein when the executable code is executed by a processor of an electronic device, the processor is able to implement at least the text processing method as described in the first aspect above, or the human-computer interaction method as described in the second aspect above.
[0016] In the text processing method provided by this invention, the interactive device acquires the text to be processed and determines the corresponding prompt vector and prompt text. Finally, the text to be processed, along with its corresponding prompt vector and prompt text, is input into a question-answering model, which then outputs the correct response text. As the names suggest, the prompt vector is the prompt information in vector form, and the prompt text is the prompt information in text form. Using this method, on the one hand, vector prompts can serve as supplementary parameters to the question-answering model, improving the accuracy of the response text generated by the model; on the other hand, text prompts can supplement the background knowledge of the text to be processed, enabling the question-answering model to generate response text more accurately based on more background knowledge. In other words, using different forms of prompt information can supplement the question-answering model from different perspectives, thereby improving the accuracy of the response text generated by the model. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart of a text processing method provided in an embodiment of the present invention;
[0019] Figure 2 A flowchart illustrating a method for determining a prompt vector according to an embodiment of the present invention;
[0020] Figure 3 A flowchart illustrating a method for determining prompt text provided in an embodiment of the present invention;
[0021] Figure 4 A flowchart illustrating another method for determining prompt text provided in an embodiment of the present invention;
[0022] Figure 5a This is a schematic diagram illustrating the determination of response text in an academic setting, as provided in an embodiment of the present invention.
[0023] Figure 5b A schematic diagram of the display interface of a terminal device provided in an embodiment of the present invention;
[0024] Figure 6 This is a schematic diagram illustrating the determination of response text in a home setting, as provided in an embodiment of the present invention.
[0025] Figure 7 A flowchart of a model training method provided in an embodiment of the present invention;
[0026] Figure 8 A flowchart of another model training method provided in an embodiment of the present invention;
[0027] Figure 9 A flowchart illustrating yet another model training method provided in this embodiment of the invention;
[0028] Figure 10 A flowchart illustrating yet another model training method provided in this embodiment of the invention;
[0029] Figure 11 A flowchart illustrating a method for determining performance indicators provided in an embodiment of the present invention;
[0030] Figure 12 A flowchart of a human-computer interaction method provided in an embodiment of the present invention;
[0031] Figure 13 This is a schematic diagram of the structure of a text processing device provided in an embodiment of the present invention;
[0032] Figure 14 This is a schematic diagram of the structure of a human-computer interaction device provided in an embodiment of the present invention;
[0033] Figure 15 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention;
[0034] Figure 16 This is a schematic diagram of another electronic device provided in an embodiment of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” used in the embodiments of this invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one.
[0037] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0038] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to identification.” Similarly, depending on the context, the phrases “if determination” or “if identification (of the condition or event of the statement)” can be interpreted as “when determination” or “in response to determination” or “when identification (of the condition or event of the statement)” or “in response to identification (of the condition or event of the statement).”
[0039] It should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.
[0040] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0041] Before describing in detail the text processing method and human-computer interaction method provided in the following embodiments of the present invention, the relevant concepts involved in the following embodiments can also be explained:
[0042] A pre-trained model is a model trained on and stored using a large amount of data. By executing historical tasks, this model acquires some historical functionality. When new functionality is introduced, instead of training a new model from scratch, the model can be directly built upon this pre-trained model to perform incremental tasks corresponding to the new functionality, thus enabling the model to acquire the new functionality.
[0043] Prompt information: Used to guide the pre-trained model in its recognition direction regarding the original input data. The prompt information, along with the original input data, can be fed together with the latest input data into the pre-trained model for recognition. The prompt information guides the pre-trained model in its recognition direction.
[0044] Text to be processed: Text received by the interactive device and generated by the user.
[0045] For the text to be processed, one scenario is that the text may include a question. The user can input this question into an interactive device that provides human-computer dialogue services, and the device can output a corresponding response. This is common in question-and-answer scenarios (i.e., the user asks, the interactive device answers), and the response text output by the device is usually short and logically simple. For example, in a home setting, the interactive device could be a smart speaker. The user could ask the device, "What's the weather like today?", and the smart speaker could directly respond by retrieving the weather information and broadcasting it to the user. Another example is in an online customer service scenario, where the interactive device could be an online service robot. The user could ask the robot, "Is this garment prone to fading?", and the robot could directly respond by retrieving the material of the garment and answering whether it is prone to fading.
[0046] In another scenario, the text to be processed may include not only the question text but also the corresponding stem text. The stem text can be understood as background text related to the question text. Users can input both the question text and the stem text into the interactive device, which will then understand the semantics of the two texts and output the corresponding response text. In this case, the interactive device can be any mobile terminal device used by the user, and the response text output by the interactive device is usually quite long and logically complex. For example, in an academic setting, a user can input the stem text into a mobile phone with the corresponding software installed: "The development of artificial intelligence can be divided into five stages: (1) Initial development stage: 1943-1960s; (2) Reflective development stage: 1970s; (3) Application development stage: 1980s; (4) Stable development stage: 1990s-2010; (5) Flourishing development stage: 2011 to present," and also input the question text: "Use these materials to write a 1000-word report on the development of artificial intelligence technology." Ultimately, the phone can generate a report for the user about the development of artificial intelligence technology.
[0047] It should be noted that, in practice, regardless of whether the text to be processed generated by the user contains the question text, the interactive device can execute the methods provided in the following embodiments of the present invention to realize human-computer interaction, more specifically, human-computer dialogue.
[0048] Based on the above description, some embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0049] Figure 1 This is a flowchart illustrating a text processing method provided in an embodiment of the present invention. The text processing method provided in this embodiment can be executed by an interactive device with data processing capabilities. The interactive device can be, as mentioned in the background art, an online service robot, a smart speaker, a mobile terminal device, etc. Figure 1 As shown, the method may include the following steps:
[0050] S101, Obtain the text to be processed.
[0051] In response to user interaction with the interactive device, the interactive device can acquire the text to be processed generated by the user. Optionally, the interaction can be an input operation triggered by the user's interface on the interactive device, or it can be voice input by the user into the interactive device. Corresponding to the interaction, the text to be processed can be text collected by the interactive device from the user's input, or it can be text converted from the user's voice. Optionally, as described above, the text to be processed acquired by the interactive device can be of two types: the text to be processed can include the question text, or it can include both the question text and the question stem text. Furthermore, users can generate text to be processed with different content in different scenarios.
[0052] S102, determine the prompt vector and prompt text corresponding to the text to be processed.
[0053] Next, the interactive device can determine the corresponding prompt vector and prompt text based on the received text to be processed. The prompt vector is the prompt information in vector form, and the prompt text is the prompt information in text form.
[0054] For determining the cue vector, one alternative approach is for the interactive device to first encode the text to be processed to obtain an encoded result in vector form, and then select cue vectors similar to the encoded result from a pre-created cue pool. Optionally, the similarity between the cue vector and the encoded result can be specifically represented by cosine distance, Euclidean distance, Chebyshev distance, etc.
[0055] One possible approach to determining the prompt text is for the interactive device to first select semantically related knowledge from a pre-established knowledge graph based on the semantics of the text to be processed. This selected knowledge becomes the prompt text. The process of filtering prompt text from the knowledge graph can also be considered as knowledge mining from the knowledge graph.
[0056] S103, input the text to be processed, the prompt vector, and the prompt text into the question-answering model so that the question-answering model can output the response text corresponding to the text to be processed.
[0057] Finally, the interactive device can fuse the text to be processed, the prompt vector, and the prompt text, and input the fusion result into the question-answering model, so that the question-answering model can output the response text corresponding to the text to be processed.
[0058] Considering the different representations of the text to be processed, the prompt vector, and the prompt text, one optional fusion method is to convert the text to be processed and the prompt text into vectors respectively and then fuse them with the prompt vector. Considering the information loss that occurs after converting the text to be processed and the prompt text into vectors, another fusion method is proposed: the interactive device can first fuse the encoded result of the text to be processed with the prompt vector to obtain a first fusion result; then, it can fuse the text to be processed and the prompt text to obtain a second fusion result. Finally, the first fusion result and the second fusion result are jointly input into the question answering model.
[0059] Optionally, vector fusion can include direct addition or direct concatenation. Alternatively, an attention mechanism can be used for addition or concatenation.
[0060] Optionally, the question-answering model can be a Generalized Linear Model (GLM) with strong text generation capabilities and a large model size, or a Generative Pre-training Transformer (GPT) model, etc. However, in practice, considering that the function of the question-answering model is relatively simple, namely generating response text, the question-answering model can also be a Text-to-Text Transformer (T5) model with moderate text generation capabilities and a small model size, or a Bidirectional Encoder Representation from Transformers (BERT) model, etc.
[0061] It is worth noting that both the prompt vector and the prompt text in this invention are used to guide the direction in which the interactive device generates the response text. The prompt vector shares the same properties as the question-answering model—that is, both are trainable and adjustable—so the prompt vector can serve as a supplementary parameter to the question-answering model, guiding the direction in which it generates the response text. Simultaneously, the prompt text obtained through knowledge mining can serve as supplementary background knowledge for the text to be processed, further guiding the direction of the question-answering model's generation.
[0062] In this embodiment, the interactive device acquires the text to be processed and determines the corresponding prompt vector and prompt text. Finally, the text to be processed, along with its corresponding prompt vector and prompt text, is input into the question-answering model, which then outputs the correct response text. As the names suggest, the prompt vector is the prompt information in vector form, and the prompt text is the prompt information in text form. Using this method, on the one hand, vector prompts can serve as supplementary parameters to the question-answering model, improving the accuracy of the response text generated by the model; on the other hand, text prompts can supplement the background knowledge of the text to be processed, enabling the question-answering model to generate response text more accurately based on more background knowledge. In other words, using different forms of prompt information can supplement the question-answering model from different perspectives, thereby improving the accuracy of the response text generated by the model.
[0063] Furthermore, the beneficial effects of the text processing methods provided in the various embodiments of the present invention can also be understood from the following aspects:
[0064] In practice, the same question-answering model can support different types of questions and answers, such as extractive question answering, multiple-choice question answering, summary question answering, and judgment question answering, etc.
[0065] Extractive question answering involves finding the corresponding response text from a given question text. For example, if the question text is "How long does primary school last?" and the question text is "Primary education is generally six years," then the response text would be "six years."
[0066] Multiple-choice question answering: This involves selecting the correct answer from multiple options. For example, if the question text is "Which of the following model structures performs best in question answering?" and the four options are "A, MLP, B, CNN, C, LSTM, D, Transformer", then the answer text would be "D". Multiple-choice question answering can be generated in the academic scenario described above.
[0067] Summary-based question and answer: This generates corresponding response text for the question text, which may not appear in the question or question stem text. Summary-based question and answer can be generated in the aforementioned home, customer service, and academic scenarios.
[0068] For judgment-based questions and answers, the corresponding answer is either "true" or "false". For example, if the question text is "Is the sky blue?", the answer text is "true"; if the question text is "Is the sky green?", the answer text is "false". Judgment-based questions and answers can be generated in the aforementioned home and customer service scenarios.
[0069] However, in practice, the number of training samples used to train a question-answering model to achieve any of the aforementioned question-answering methods is often limited. Therefore, new questions that were not learned during the training phase may arise during the use of the question-answering model, resulting in an open-world problem. When using the text processing methods provided in the above and following embodiments, since the question-answering model can use prompt text as supplementary knowledge to the text to be processed, this prompt text can compensate for the knowledge not learned during the training phase. This makes it easier for the question-answering model to understand the text to be processed and ultimately outputs a more accurate response text. In other words, prompt text can improve the open-world problem caused by the limited number of training samples.
[0070] For determining the cue vector corresponding to the text to be processed, optionally, when the cue pool contains only a cue vector, then the following can be used: Figure 1 The method described in the illustrated embodiment. Optionally, when the hint pool stores the hint vector in the form of key-value pairs, i.e., the index is the key and the hint vector is the value, the following can be used. Figure 2 The method shown determines the hint vector. An index and the hint vector that index points to can form a key-value pair stored in the hint pool. Figure 2 This is a flowchart illustrating a method for determining a prompt vector according to an embodiment of the present invention. Figure 2 As shown, the following steps may be included:
[0071] S201, Encode the text to be processed.
[0072] S202, determine a preset number of target indexes based on the similarity between the encoding result and the indexes in the prompt pool.
[0073] S203, determine the prompt vector pointed to by the target index in the prompt pool as the prompt vector corresponding to the text to be processed.
[0074] The interactive device can first acquire the text to be processed and then encode the text using its own configured encoder to obtain the encoded result. The method for acquiring the text to be processed can be found in [link to relevant documentation]. Figure 1 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0075] Then, the interactive device can determine a preset number of target indices based on the similarity between the encoded result and each index in the cue pool. Specifically, the interactive device can select a preset number of target indices from the cue pool that are most similar to the encoded result. Optionally, the similarity between the encoded result and the index can be expressed as cosine distance, Euclidean distance, Chebyshev distance, etc. Finally, the interactive device can concatenate the cue vectors pointed to by the target indices in the cue pool to obtain the cue vector corresponding to the text to be processed.
[0076] It should be noted that, optionally, each index in the hint pool can also be represented as a vector. An index can contain part of the data in the hint vector it points to. That is, although the hint vectors in the hint pool and the indexes pointing to the hint vectors can both be represented as vectors, the vector used to describe the index contains less data.
[0077] In this embodiment, when the prompt vectors are stored in the prompt pool as key-value pairs, the interactive device can first encode the text to be processed. Then, based on the encoding result, it determines a preset number of target indices in the prompt pool that are similar to the encoded result. Finally, the concatenation result of the prompt vectors pointed to by the target indices is used as the prompt vector corresponding to the text to be processed. Furthermore, in this embodiment, since the vector used to describe the index contains less data, the interactive device can calculate the similarity between the encoding result and the index more quickly, meaning the prompt vector determination is faster. The vector prompts obtained according to the above method can also serve as supplementary parameters for the question-answering model, improving the accuracy of the response text generated by the question-answering model.
[0078] Furthermore, the user-generated text to be processed can belong to any of the above-mentioned question-and-answer formats, so regardless of whether this embodiment or... Figure 1 The suggestion vector determination method provided in the illustrated embodiment allows interactive devices to determine the suggestion vector corresponding to the text to be processed from the suggestion pool based on similarity, thus achieving text-level suggestion vector determination. Compared to using question-and-answer type-level suggestion vectors, which uses the same suggestion vector for different texts belonging to the same question-and-answer category, this method... Figure 2 The illustrated embodiment can determine a more granular cue vector, which is more targeted and can more accurately guide the question-answering model to output the response text corresponding to the text to be processed.
[0079] Regarding the determination of the prompt text, in addition to Figure 1 The method provided in the illustrated embodiments may optionally be... Figure 3 This is a flowchart illustrating a method for determining prompt text according to an embodiment of the present invention. Figure 3 As shown, the following steps may be included:
[0080] S301, Based on the first similarity between the historical text and the text to be processed, determine the first target text in the historical text.
[0081] The historical text can be texts to be processed generated by different users within a historical time period and received by the same interactive device, or it can be various texts collected through the Internet. When the text to be processed includes question text, the historical text can include historical question texts and the corresponding reference answer texts. When the text to be processed includes both question stem text and question text, the historical text can include historical question stem texts, the corresponding historical question texts, and the corresponding reference answer texts.
[0082] Optionally, the interactive device can use a preset algorithm to calculate the first similarity between the historical text and the text to be processed. Optionally, the exchanging device can also input the historical text and the text to be processed into a first filtering model, so that the model outputs the first similarity between the historical text and the text to be processed. Optionally, the first filtering model can be a computationally efficient two-head BERT model. The first similarity can also be represented as cosine distance, Euclidean distance, Markov distance, etc. It should be noted that since there are usually multiple historical texts, the first similarity can be specifically represented as a similarity distribution, which consists of multiple similarity values, the same number as the number of historical texts.
[0083] Then, the interactive device can filter out a preset number of first target texts with the highest similarity from the historical text based on the calculated first similarity. Among them, the historical question texts in the filtered first target texts have a high similarity to the question texts in the text to be processed.
[0084] S302, input the text to be processed and the first target text into the pre-trained model so that the pre-trained model can generate the first prompt text corresponding to the first target text.
[0085] Based on the first target text obtained in step S301, the interactive device can also input this first target text and the text to be processed into a pre-trained model, so that the pre-trained model can generate first prompt texts corresponding to each of the first target texts. The first prompt text corresponding to any first target text can be considered as the knowledge contained in that first target text. Optionally, the pre-trained model can be a GLM model or a GPT model.
[0086] S303, determine the prompt text corresponding to the text to be processed based on the first prompt text.
[0087] Finally, the interactive device can directly concatenate the first prompt text generated in step S302 and use it as the prompt text corresponding to the text to be processed.
[0088] In this embodiment, the interactive device can filter out a first target text from the historical text based on a first similarity between the historical text and the text to be processed. Then, a pre-trained model generates a first prompt text corresponding to each of the first target texts, further identifying this first prompt text as the prompt text corresponding to the text to be processed. Since the question text contained in the first target text filtered according to the first similarity has a high similarity to the question text contained in the text to be processed, inputting the first target text and the text to be processed together into the pre-trained model can make the prompt text output by the pre-trained model more suitable for the question text in the text to be processed. This prompt text can further accurately guide the question-answering model, ultimately enabling the question-answering model to output a more accurate response message.
[0089] To further improve the efficiency of response text generation, Figure 4 A flowchart illustrating another method for determining prompt text provided in an embodiment of the present invention. For example... Figure 4 As shown, the steps may include the following:
[0090] S401, Based on the first similarity between the historical text and the text to be processed, determine the first target text in the historical text.
[0091] S402, input the text to be processed and the first target text into the pre-trained model so that the pre-trained model can generate the first prompt text corresponding to the first target text.
[0092] For the specific implementation process of steps S401 to 402 above, please refer to [link / reference]. Figure 3 The specific descriptions of the relevant steps in the illustrated embodiments will not be repeated here.
[0093] S403, based on the text to be processed, the first target text, and the first prompt text, determine the second similarity between the fused text and the text to be processed, and the fused text is the fusion result of the first target text and the first prompt text.
[0094] S404, Based on the second similarity, determine the second target text in the first target text.
[0095] Optionally, the interactive device can fuse the first target text and the first prompt text, and use a preset algorithm to calculate a second similarity between the fused result and the text to be processed. Optionally, the interactive device can also input the text to be processed, the first prompt text, and the first target text into a second filtering model, so that the model outputs a second similarity between the first target text and the text to be processed.
[0096] Then, based on the calculated second similarity, the interactive device can further filter out a preset number of second target texts with the highest similarity from the first target text. Optionally, the second filtering model can be a computationally sophisticated and effective cross-beta BERT model. It should be noted that since there are usually multiple second target texts, the second similarity, like the first similarity, can also be represented as a similarity distribution, consisting of multiple similarity values, with the same number of similarity values as the first target text. The second similarity can also be expressed as cosine distance, Euclidean distance, Markov distance, etc.
[0097] Since the second target text is selected from the first target text, the number of first target texts is greater than the number of second target texts; specifically, the number of first target texts is much greater than the number of second target texts.
[0098] As described above, the second filtering model can fine-tune the first target text to obtain the second target text. Furthermore, since the second filtering model also considers the first prompt text during text filtering, the ranking result of the second filtering model on the first target text differs from that of the first filtering model. Therefore, while fine-tuning the first target text, the second filtering model also reorders the first target text based on its corresponding first prompt text.
[0099] S405, determine the prompt text corresponding to each of the second target texts as the prompt text corresponding to the text to be processed.
[0100] Finally, the interactive device can concatenate the prompt texts corresponding to the second target texts that have been finely filtered out, and determine the concatenated result as the prompt text corresponding to the text to be processed.
[0101] In this embodiment, after using the first screening model to coarsely screen out the first target text from the historical text, the second screening model can be used to finely screen and rearrange the first target text to obtain a second target text that has a high similarity to the problem text in the text to be processed and is less numerous.
[0102] On the one hand, the second filtering model filters and rearranges the first target text, enabling the selection of second target texts with high similarity to the question text in the text to be processed. Furthermore, the prompt text corresponding to this second target text is used as additional knowledge for the text to be processed, making the guidance of the question-answering model's generation direction more accurate. On the other hand, the question-answering model can generate response text using only a small number of prompt texts corresponding to each of the second target texts, which reduces the computational load in the response text generation process and improves the efficiency of response text generation.
[0103] To facilitate understanding, when the text to be processed contains both the question stem and the question text, the specific implementation process of the text processing methods provided in the above embodiments can be illustrated below using an academic scenario. The following process can also be combined with... Figure 5a and 5b understand.
[0104] In academic settings, interactive devices can be various mobile terminal devices used by users, such as mobile phones. In this scenario, users can input the text to be processed into the interactive interface provided by the question-and-answer software installed on their mobile phones: input(c,q): "[c] The development of artificial intelligence can be divided into five stages: (1) Initial development stage: 1943-1960s; (2) Reflection development stage: 1970s; (3) Application development stage: 1980s; (4) Stable development stage: 1990s-2010; (5) Flourishing development stage: 2011 to present", "[q] Use these materials to help me write a report on the development of artificial intelligence technology, 1000 words." Among them, [c] is the question text, and [q] is the problem text.
[0105] After that, the phone can be used as described above. Figure 2 The method in the illustrated embodiment determines the cue vector Pm(c,q). Specifically, the text to be processed, input(c,q), is first encoded using an encoder to obtain the encoding result x. Then, the similarity between the encoding result x and indices 1 to 6 in the cue pool is calculated. Based on the calculated similarity, the three target indices with the highest similarity are determined from the six indices, namely indices 1, 3, and 6. Finally, the cue vectors pointed to by the three target indices in the cue pool are... Concatenate the vectors to obtain the prompt vector corresponding to the text input(c,q) to be processed. Furthermore, the cue vector Pm(c,q) is fine-grained and text-level.
[0106] At the same time, the mobile phone can follow the above Figure 4 The illustrated embodiment determines the prompt text Pk(c,q) in this way. That is, the mobile phone can use the coarse screening function of the first screening model, the fine screening and rearrangement function of the second screening model, and the prompt generation function of the pre-trained model to select the second target text with high similarity to the text input(c,q) to be processed, as well as the prompt text corresponding to the second target text, from the historical text.
[0107] Historical text can be text entered by the user into the question-and-answer software installed on the phone within a historical time period, and may include, for example:
[0108] Historical text 1: "[c] The development of artificial intelligence can be divided into 5 stages. [q] Please write a report on the development of artificial intelligence technology. [a] The development of artificial intelligence can be divided into 5 stages: the initial development stage, the reflective development stage, the application development stage, the stable development stage, and the vigorous development stage."
[0109] Historical text 2: "[c] The development of artificial intelligence has gone through five periods from 1943 to the present. [q] From the perspective of the time of technological development, please write a report on the development of artificial intelligence technology. [a] The development of artificial intelligence can be divided into five stages: (1) Initial development period: 1943-1960s; (2) Reflection development period: 1970s; (3) Application development period: 1980s; (4) Stable development period: 1990s-2010; (5) Vigorous development period: 2011 to the present."
[0110] Historical text 3: "[c] The development of neural networks has gone through three periods from 1943 to the present. [q] Please write a report on the development of neural networks. [a] One way to implement artificial intelligence technology is through neural networks. A neural network is a computational model based on biological neural networks, which mainly includes three development stages: (1) The birth of neural networks: 1943-1969; (2) The revival of neural networks: 1980s-1990s; (3) Deep learning: 2006 to the present."
[0111] Where [c] represents the question stem text, [q] represents the question text, and [a] represents the response text corresponding to the question text.
[0112] At this point, by using the first screening model to perform a coarse screening of the three historical texts, we can obtain historical text 1 and historical text 2 as the first target text. Then, by using the second screening model to perform a fine screening and rearrangement of historical text 1 and historical text 2, we obtain historical text 2 as the second target text.
[0113] Specifically, to determine the prompt text corresponding to the second target text, the text to be processed, input(c,q), historical text 1, and historical text 2 can be input into the pre-trained model so that the pre-trained model can generate the prompt texts corresponding to historical text 1 and historical text 2 respectively. That is, the prompt text 1 corresponding to historical text 1 is "The development process of artificial intelligence can be divided into 5 stages: the initial development stage, the reflection development stage, the application development stage, the stable development stage, and the vigorous development stage." and the prompt text 2 corresponding to historical text 2 is "The development process of artificial intelligence can be divided into 5 stages: (1) the initial development stage: 1943-1960s; (2) the reflection development stage: 1970s; (3) the application development stage: 1980s; (4) the stable development stage: 1990s-2010; (5) the vigorous development stage: 2011 to present."
[0114] At this point, the mobile phone can input the text to be processed (input(c,q), historical text 1, historical text 2, prompt text 1, and prompt text 2 into the second filtering model. The model will then filter and rearrange the historical text 2 with the highest similarity from historical text 1 and historical text 2, and use the prompt text 2 corresponding to historical text 2 as the prompt text Pk(c,q) corresponding to the text to be processed (input(c,q)).
[0115] Finally, the text to be processed, input(c,q), the prompt vector, Pm(c,q), and the prompt text, Pk(c,q), are input into the question-and-answer software installed on the mobile phone, so that the question-and-answer software can output the response text corresponding to the text to be processed, input(c,q). That is, "The development history of artificial intelligence can be divided into five stages: (1) the initial development stage: 1943-1960s; (2) the reflection development stage: 1970s; (3) the application development stage: 1980s; (4) the stable development stage: 1990s-2010; (5) the vigorous development stage: 2011 to present."
[0116] Question-and-answer apps installed on mobile phones can also provide... Figure 5b The interface shown can display the user-inputted text input(c,q) and the response text output by the question-and-answer software.
[0117] When the text to be processed contains problematic text, the following provides an exemplary description of the specific implementation process of the text processing methods provided in the above embodiments, using a home setting as an example. The following process can also be combined with... Figure 6 understand.
[0118] In a home setting, the interaction device can be a smart speaker. A user can ask the smart speaker, "What's the weather like today?" The smart speaker can then convert the user's speech into text, obtaining the text to be processed, input(q), "What's the weather like today?". The smart speaker can then further determine the corresponding prompt vector Pm(q) and prompt text Pk(q) for this text input(q).
[0119] Next, regarding the determination of the cue vector Pm(q),
[0120] Smart speakers can be configured as described above. Figure 2 as well as Figure 5a The method of the illustrated embodiment determines the prompt vector Pm(q). The specific process will not be described in detail here.
[0121] At the same time, smart speakers can follow the above... Figure 4 as well as Figure 5a The method shown in the embodiment determines the prompt text Pk(q).
[0122] In this scenario, historical text can be the text that the user inputs into the smart speaker within a historical time period, such as:
[0123] Historical text 1: "[q] How's the weather today? [a] Sunny"; Historical text 2: "[q] What's the wind chill today? [a] 8℃"; Historical text 3: "[q] What's the wind force today? [a] No wind". Here, q represents the question text, and a represents the corresponding response text.
[0124] At this point, by using the first filtering model to perform a coarse screening of the three historical texts, we can obtain historical text 1 and historical text 2. Then, by using the second filtering model to perform a fine screening and rearrangement of historical text 1 and historical text 2, we obtain historical text 2. Among them, historical text 1 and historical text 2 obtained by the first filtering model are the first target texts in the above embodiments, and historical text 2 obtained by the second filtering model is the second target text in the above embodiments.
[0125] Specifically, to determine the prompt text corresponding to the second target text, the smart speaker can input the text to be processed (input(q), historical text 1, and historical text 2) into a pre-trained model. The pre-trained model will then generate prompt texts corresponding to historical text 1 and historical text 2, respectively. Specifically, prompt text 1 for historical text 1 is "Weather; Sunny," and prompt text 2 for historical text 2 is "Feeling temperature; 8℃." Then, the smart speaker can input the text to be processed (input(q), historical text 1, historical text 2, prompt text 1, and prompt text 2) into a second filtering model. This model will then select historical text 2 with the highest similarity from historical text 1 and historical text 2, and use the prompt text 2 corresponding to historical text 2 as the prompt text Pk(q) corresponding to the text to be processed (input(q)).
[0126] Finally, the text to be processed, input(q), the prompt vector, Pm(q), and the prompt text, Pk(q), are input into the question-answering model in the smart speaker so that the question-answering model can output the response text corresponding to the text to be processed, input(q), namely, "Today's perceived temperature is 8℃".
[0127] In determining the response text in the aforementioned scenarios, vector hints can serve as supplementary parameters to the question-answering model, improving the accuracy of the generated response text. Conversely, textual hints can supplement the background knowledge of the text being processed, enabling the question-answering model to generate more accurate response text based on broader contextual knowledge. In other words, using different forms of hints can supplement the question-answering model from various perspectives, thereby improving the accuracy of the generated response text. Furthermore, filtering models can be used to filter and rearrange the text to obtain more accurate hint text.
[0128] According to the above Figures 1 to 6 As can be seen from the usage process of the interactive device described in the illustrated embodiment, when generating response text, a prompt pool, a question-answering model, a first filtering model, a second filtering model, and a pre-trained model are used sequentially. It should also be noted that at least one of the above models and the prompt pool can be deployed on the interactive device or on the cloud. In this case, the interactive device can send the text to be processed to the cloud and then receive the response text from the cloud.
[0129] The training process for each model is described below. The following embodiments can be executed by a training device. Optionally, the training device can be the interactive device mentioned in the above embodiments, or it can be another device besides the interactive device, which can be deployed in the cloud. The training text set used in the following embodiments can be a set of files corresponding to the same or different question-and-answer types.
[0130] To ensure the question-answering model outputs more accurate response text, it can be trained. Furthermore, to improve the accuracy of the prompt vectors and ensure they guide the question-answering model more effectively, a prompt pool can optionally be trained simultaneously with the question-answering model. Figure 7 This is a flowchart illustrating a model training method provided in an embodiment of the present invention. This method can be used to simultaneously train a question-answering model and a prompt pool. Figure 7 As shown, the method may include the following steps:
[0131] S501, Obtain training text. The training text includes training question text and corresponding reference answer text, or it includes training question stem text, corresponding training question text, and corresponding reference answer text.
[0132] Optionally, the training device can acquire training texts from a training text set, and the training text set can be a pre-collected text set applicable to any type of question and answer. The training texts may include training question texts and corresponding reference answer texts, or training question stem texts, corresponding training question texts, and corresponding reference answer texts.
[0133] S502, determine the prompt vector and prompt text corresponding to the training text, wherein the prompt vector corresponding to the training text is included in the prompt pool.
[0134] Based on the training text obtained in step S501, the training device can use the methods in the above embodiments to determine the prompt vector and prompt text corresponding to the training text, which will not be repeated here.
[0135] The cue vectors corresponding to the training text are contained in the cue pool. Optionally, the cue pool can store cue vectors in key-value pairs, where the index is the key and the cue vector is the value. An index and the cue vector it points to can form a key-value pair stored in the cue pool.
[0136] S503, using the training text, the prompt vector corresponding to the training text, and the prompt text corresponding to the training text as training data, and the reference answer text corresponding to the training text as supervision information, train the question answering model, the prompt vector corresponding to the training text, and the index pointing to the prompt vector corresponding to the training text.
[0137] The training device can input the prompt vector and prompt text corresponding to the determined training text, as well as the text in the training text other than the reference answer text, as training data into the question answering model, and use the reference answer text in this training text as supervision information to train the question answering model, and at the same time train the prompt vector corresponding to the text and the index pointing to the indicator vector corresponding to the training text.
[0138] The training process for the question-answering model and the hint pool can involve calculating the loss between the predicted answer text output by the question-answering model and the reference answer text in the training text, and then adjusting the model parameters of the question-answering model using gradient descent based on the calculated loss value. This adjustment also includes adjusting the hint vector corresponding to the text and the index pointing to the indicator vector corresponding to the training text, so that the final loss value meets preset requirements. Optionally, relative entropy (Kullback-Leibler Divergence, or KL divergence) or cross-entropy loss functions can be used for loss calculation.
[0139] When training a cue pool using a training text, what is trained is the cue vector corresponding to that training text, and the index pointing to that cue vector.
[0140] In this embodiment, by training the question-answering model and the prompt pool simultaneously, training efficiency can be improved, the prompt vectors in the prompt pool can be made more accurate, and the training effect of the question-answering model can be improved.
[0141] Figure 8 A flowchart of another model training method provided in an embodiment of the present invention, which can be used to train a first screening model. Figure 8 As shown, the method may include the following steps:
[0142] S601, In the training text set, select other training texts whose similarity to any training text in the training text set meets the preset requirements.
[0143] Optionally, the training text set can be a pre-collected text set applicable to any question-and-answer format. Any training text in the training text set can include training question text, or training question stem text and the corresponding training question text.
[0144] The training device can select other training texts from the training text set whose similarity to any training text in the training text set meets preset requirements. Optionally, the training device can use the Best Matching25 (BM25) algorithm to calculate the similarity between any training text and the remaining training texts in the training text set, and then select other training texts that meet preset requirements from the remaining training texts based on the calculated similarity. Here, the similarity is actually the similarity between the training question text in any training text and the training question text in the remaining training texts.
[0145] Optionally, the similarity between any training text and the remaining training texts can also be expressed as cosine distance, Euclidean distance, Chebyshev distance, etc. Optionally, the preset requirement may include a similarity greater than a preset threshold, that is, the remaining training texts with a similarity greater than the preset threshold are identified as other training texts, or the preset requirement may also include a preset number, that is, the remaining training texts with the highest similarity and a preset number are identified as other training texts.
[0146] S602, input any training text and other training texts into the first screening model, so that the first screening model outputs a first similarity distribution, wherein the first similarity distribution reflects the similarity between any training text and other training texts.
[0147] Then, the training device can input any training text and other training texts into the first screening model, which will output a first similarity distribution. This first similarity distribution can be composed of the similarity scores between the question text in any training text and other training texts. The first screening model can be a computationally efficient dual-head BERT model to quickly output the similarity distribution between the question text in any training text and other training texts. The similarity score can also be considered as the model's score on whether any training text is similar to other training texts.
[0148] S603, input any training text and other training texts into the pre-trained model, so that the pre-trained model outputs the prompt text corresponding to the other training texts, and the first probability distribution of the prompt texts, wherein the first probability distribution reflects the confidence that the prompt text is the reference answer text corresponding to any training text.
[0149] The training device can then input any training text along with other training texts into a pre-trained model with fixed parameters. The pre-trained model will output the prompt texts corresponding to each of the other training texts, as well as a first probability distribution of the prompt texts. The first probability distribution can consist of multiple probability values, each representing the confidence level that the prompt text corresponding to one of the other training texts is the reference answer text corresponding to any training text. This confidence level can be considered as the pre-trained model's score of the similarity between the prompt text and the reference answer text.
[0150] S604. Train the first screening model based on the loss calculation results between the first similarity distribution and the first probability distribution.
[0151] Furthermore, the training device can calculate the loss on the first similarity distribution and the first probability distribution mentioned above, and use the calculated loss value to adjust the first screening model, that is, train the first screening model. Optionally, relative entropy loss or cross-entropy loss can be used for loss calculation.
[0152] Optionally, in practice, the training device may use text sets applicable to different types of questions and answers to train the first screening model multiple times to improve the text screening capability of the first screening model.
[0153] In this embodiment, since the existing pre-trained model has good performance, training the first screening model based on the loss between the first similarity distribution and the first probability distribution output by the pre-trained model can improve the learning effect of the first screening model.
[0154] Figure 9 This is a flowchart illustrating another model training method provided in an embodiment of the present invention, which can be used to train a second screening model. Figure 9 As shown, the method may include the following steps:
[0155] S701, In the training text set, select other training texts that meet the preset requirements in terms of similarity to any training text in the training text set.
[0156] The execution process of step S701 described above is similar to the corresponding steps in the aforementioned embodiments, and can be found in the following example. Figure 8 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0157] S702, input any training text and other training texts into the pre-trained model, so that the pre-trained model outputs the prompt text corresponding to the other training texts and the second probability distribution of the prompt texts, wherein the second probability distribution reflects the confidence that the prompt text is the reference answer text corresponding to any training text.
[0158] It should be noted that the second probability distribution in this embodiment is the same as the first probability distribution in the above embodiments, and the generation process can be found in the description in the above embodiments. Since probability distributions are used in the training process of different models, they are distinguished by their names.
[0159] S703, input any training text, other training texts, and prompt texts corresponding to other training texts output by the pre-trained model into the second filtering model, so that the second filtering model outputs a second similarity distribution, wherein the second similarity distribution reflects the similarity between any training text and other training texts.
[0160] The generation process of the second similarity distribution is similar to that of the first similarity distribution, except that, compared to the first screening model, the input of the second screening model includes prompt texts corresponding to other training texts.
[0161] S704, Train the second screening model based on the loss calculation results between the second probability distribution and the second similarity distribution.
[0162] The execution process of step S704 described above is similar to the corresponding steps in the aforementioned embodiments, and can be found in the following example: Figure 8 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0163] In this embodiment, the training device can train the second screening model based on the loss value calculated by the second similarity distribution and the second probability distribution, and adds prompt texts corresponding to other training texts to the training data. Therefore, it can better improve the learning effect of the second screening model.
[0164] Furthermore, the beneficial effects achieved by the text processing methods provided in the various embodiments of the present invention can also be understood from the following perspectives:
[0165] In practice, the same question-answering model can support different types of questions and answers. However, the number of training texts used to train different types of questions and answers is different, meaning that the training texts have a long-tail distribution. This will cause the question-answering model to have different response text generation capabilities for different types of questions and answers. That is, for question-answering types with a large number of training texts, the question-answering model can generate response text more accurately when it receives the text to be processed belonging to that type of question and answer; otherwise, it cannot generate response text accurately.
[0166] To mitigate the impact of long-tail distribution on question-answering model performance, the prompt pool can be initially trained using a training sample set corresponding to one type of question-answering. This initial training results in a prompt pool containing knowledge from that single training sample set. Further training can be conducted using training sample sets from other question-answering types. This retraining process allows for knowledge sharing among the training sample sets corresponding to different types of questions and answers. In other words, the knowledge contained in any prompt vector within the prompt pool is a fusion of knowledge from the training sample sets of different question-answering types.
[0167] Since the suggestion vectors in the suggestion pool are obtained through knowledge sharing, the suggestion vectors, as a parameter supplement to the question answering model, can compensate for the question answering model's ability to generate response texts for question types with a small number of training texts.
[0168] Figures 7-9 The illustrated embodiments provide detailed descriptions of the training processes for the prompt pool, question-answering model, first selection model, and second selection model. Optionally, further optimization can be performed on each model to further improve its performance. Figure 10 A flowchart illustrating yet another model training method provided in an embodiment of the present invention. For example... Figure 10 As shown, the method may include the following steps:
[0169] S801, determine the performance metrics for the question-answering model, the first screening model, and the second screening model.
[0170] Optionally, the training device can calculate the loss values of the question-answering model, the first selection model, and the second selection model separately, and determine the performance index of each of the three models based on the loss calculation results. The specific process for determining the performance index can be found below. Figure 11 The description in the illustrated embodiment is provided. Optionally, performance metrics may include the similarity between the model's predicted output and the reference result. The higher the similarity, the better the model's performance.
[0171] S802, in the training text set, identify candidate training texts whose similarity to the target training text meets the preset requirements.
[0172] Optionally, the training text set can be a pre-collected text set applicable to any type of question and answer. The training text may include training question texts and corresponding reference answer texts, or it may include training question stem texts, corresponding training question texts, and corresponding reference answer texts.
[0173] The training device can first use the BM25 algorithm to coarsely screen out candidate training texts from the training sample set that meet the preset similarity requirements with the target training text. The number of candidate training texts is at least one, and the similarity can also be expressed as cosine distance, Euclidean distance, Markov distance, etc.
[0174] It should be noted that, as described above, although the first and second screening models also have text screening capabilities, in order to ensure the accuracy of subsequent model optimization, a third-party algorithm, namely the BM25 algorithm, is used here to determine the candidate training texts.
[0175] Optionally, the preset requirements may include a similarity greater than a preset threshold, that is, the remaining training texts with a similarity greater than the preset threshold are determined as candidate training texts, or the preset requirements may also include a preset number, that is, the remaining training texts with the highest similarity and a preset number are determined as candidate training texts.
[0176] S803, input the candidate training text, the prompt vector corresponding to the candidate training text, and the prompt text corresponding to the candidate training text into the question answering model, so that the question answering model outputs the predicted answer text of the candidate training text.
[0177] S804, determine the loss value distribution between the predicted answer text of the candidate training text and the reference answer text corresponding to the candidate training text.
[0178] The training device can input the prompt vector and prompt text corresponding to the determined candidate training text, as well as the text in the candidate training text other than the reference answer text, as training data into the question answering model, so that the question answering model can output the predicted answer text of the candidate training text.
[0179] Then, the training device can perform loss calculations on the predicted answer text output by the question-answering model and the reference answer text corresponding to the candidate training texts to obtain the loss value distribution of the question-answering model. The loss value distribution may include the loss value corresponding to at least one first candidate training text. Optionally, the loss calculation method may include cross-entropy loss, or it may include KL divergence, etc.
[0180] S805, the target training text and candidate training text are input into the first screening model, so that the first screening model outputs a third similarity distribution, which reflects the similarity between the target training text and the candidate training text.
[0181] The training device can input the target training text and candidate training texts into a first screening model to obtain a third similarity distribution output by the first screening model. This third similarity distribution can be composed of the similarity scores between the question text in the target training text and the candidate training texts. The similarity score can also be considered as the model's scoring of whether the target training text and the candidate training texts are similar.
[0182] S806, the target training text, the candidate training text, and the prompt text corresponding to the candidate training text output by the pre-trained model are input into the second filtering model, so that the second filtering model outputs the fourth similarity distribution, wherein the fourth similarity distribution reflects the confidence of the similarity between the target training text and the candidate training text.
[0183] The training device can input the target training text, candidate training text, and the prompt text corresponding to the candidate training text output by the pre-trained model as training data into the second screening model to obtain the fourth similarity distribution. The generation process of the fourth similarity distribution is similar to that of the third similarity distribution, except that, compared to the first screening model, the input of the second screening model includes the prompt text corresponding to the candidate training text.
[0184] S807 uses the distribution of models whose performance metrics meet preset requirements as supervision information to train the remaining models.
[0185] Ultimately, the training equipment can use the distribution corresponding to the model whose performance metrics meet the preset requirements as supervision information to train the remaining models based on their performance metrics.
[0186] Optionally, if the model whose performance meets the preset requirements can be the model with the best performance, then the distribution corresponding to the model with the best performance can be used as supervision information to train the remaining models with poor performance. That is, the parameters of the model with the best performance are fixed, and the parameters of the remaining models are optimized.
[0187] For example, assuming that the model whose performance meets the preset requirements is a question-answering model, the loss value distribution corresponding to the question-answering model is used as supervision information to train the first screening model and the second screening model, so that the third similarity distribution output by the first screening model and the fourth similarity distribution output by the second screening model are close to the loss value distribution output by the question-answering model.
[0188] In this embodiment, based on the performance metrics of the question-answering model, the first screening model, and the second screening model, the training device can use the distribution corresponding to the models whose performance metrics meet preset requirements as supervision information, while simultaneously training the remaining models with poor performance metrics. It can be seen that the above method is actually a dynamic knowledge inter-distillation process, that is, distilling the knowledge of high-performing models into other low-performing models, so that the other models can learn the knowledge from the high-performing models. Furthermore, the above knowledge inter-distillation process can also be dynamically performed according to actual circumstances.
[0189] Figure 10 As mentioned in the illustrated embodiment, the training device can determine the performance metrics of the question-answering model, the first selection model, and the second selection model. Optionally, the process for determining the performance metrics of each model can be as follows: Figure 11 As shown. Then Figure 11 The flowchart of a method for determining performance indicators provided in an embodiment of the present invention may specifically include the following steps:
[0190] S901, Obtain the predicted answer text corresponding to each training text in the training text set output by the question answering model.
[0191] S902, Based on the similarity between the reference answer texts corresponding to the predicted answer texts and the training texts, determine the first set of training texts in the training text set.
[0192] By leveraging the response text generation capability of the question-answering model, we can obtain the predicted answer text corresponding to each training text in the training text set output by the question-answering model. Based on the similarity between this predicted answer text and its corresponding reference answer text, we determine the first group of training texts in the training text set. In the first group of training texts, the predicted reference answer output by the question-answering model has a high similarity to the reference answer text of that training text; that is, for the first group of training texts, the question-answering model can accurately output its answer text.
[0193] S903, based on the similarity between the second target training text in the training text set output by the first screening model and the second candidate training text in the training text set, determine the second set of training texts from the second candidate training texts.
[0194] S904, based on the similarity between the second target training text and the second candidate training text output by the second screening model, a third group of training texts is determined from the second candidate training texts, wherein each of the three groups of training texts contains the same number of texts.
[0195] The first screening model can be based on... Figure 3 The illustrated embodiment involves selecting training texts similar to the target training text from among the candidate training texts in the training text set, excluding the target training text, to form a second set of training texts. Furthermore, in order to... Figure 10 In the illustrated embodiment, the target training text and the alternative training text are distinguished. In this embodiment, the target training text is referred to as the second target training text, and the alternative training text is referred to as the second alternative training text. The second target training text can be any training text in the training text set.
[0196] Similar to the first screening model, the second screening model can also determine the third set of training texts. In steps S901 to S904 above, the same set of training texts is actually used. The question-answering model can screen out the training texts that can accurately output the response text, and the two screening models respectively screen out the training texts similar to the target training text from the second candidate training texts.
[0197] S905, obtain the prompt text and prompt vector corresponding to each of the three sets of training texts. The prompt text is output by the pre-trained model, and the prompt vector is included in the prompt pool.
[0198] S906, input each group of training texts, the corresponding prompt texts and prompt vectors of each group of training texts into the question answering model, so that the question answering model outputs the predicted answer texts of each training text in each group of training texts.
[0199] S907, determine the performance metrics of the question-answering model, the first screening model, and the second screening model based on the loss calculation results corresponding to the predicted answer texts for each group of training texts output by the question-answering model.
[0200] The training device can acquire the corresponding prompt texts and prompt vectors for each of the three sets of training texts. It then inputs each set of training texts, along with their corresponding prompt texts and prompts, into a question-answering model. The model outputs the predicted answer text for each training text in each set. The prompt texts are output by the pre-trained model, and the prompt vectors are contained in a prompt pool. The process for determining the prompt texts and prompt vectors can be found in the descriptions of the above embodiments and will not be repeated here.
[0201] Finally, based on the loss calculation results between the predicted answer text and the reference answer text for each group of training texts output by the question-answering model, the performance metrics of the question-answering model, the first selection model, and the second selection model can be determined. Here, the loss calculation result refers to the loss value corresponding to each group of training texts. Optionally, the loss calculation method may include cross-entropy loss, or it may include KL divergence, etc.
[0202] In this embodiment, the training device can concatenate the three sets of training texts, along with their corresponding prompt texts and prompt vectors, and input the concatenation result into the question-answering model. The question-answering model then outputs the loss calculation results for each set of training texts. Finally, based on these loss calculation results, the performance metrics of the question-answering model, the first selection model, and the second selection model can be determined. In other words, by inputting the selection results of the question-answering model, the first selection model, and the second selection model into the question-answering model, the question-answering model can more accurately output the performance metrics of each model, thus accurately determining the training effect of each model.
[0203] Figures 1 to 11 The illustrated embodiment details the processing of user-generated text data by the interactive device. Based on the user's interaction flow, a human-computer interaction method can also be proposed. Figure 12 This is a flowchart illustrating a human-computer interaction method provided in an embodiment of the present invention. The executing entity of this method can also be an interactive device. For example... Figure 12 As shown, the method may include the following steps:
[0204] S1001, in response to an interactive operation on the interactive device, acquires the text to be processed.
[0205] The interactive device can respond to a user's touch operation to acquire text to be processed. Optionally, the interactive device can also respond to a user's voice input operation to acquire speech to be processed. Furthermore, the interactive device can use its built-in processing module to convert the speech to be processed into text to be processed.
[0206] S1002, Determine the prompt vector and prompt text corresponding to the text to be processed.
[0207] Based on the text to be processed obtained in step S1001, the interactive device can determine the prompt vector and prompt text corresponding to the text to be processed. For details on determining the prompt vector and prompt text, please refer to the specific descriptions of the relevant steps in the above embodiments, which will not be repeated here.
[0208] S1003, Input the text to be processed, the prompt vector, and the prompt text into the question-answering model so that the question-answering model can determine the response text corresponding to the text to be processed.
[0209] S1004, Output the response text.
[0210] Optionally, the interactive device can directly input the text to be processed, the prompt vector, and the prompt text together into the question-answering model, or it can input the concatenated result of the text to be processed, the prompt vector, and the prompt text into the question-answering model. The question-answering model then determines the corresponding response text for the text to be processed and outputs this response text. Optionally, the output response text can be in text format or can be played back via audio.
[0211] Furthermore, the process for determining the prompt vector and prompt text, the use of the question-answering model, and the training process can be found in the detailed descriptions of the relevant steps in the above embodiments, and will not be repeated here.
[0212] In this embodiment, in response to an interactive operation on the interactive device, the text to be processed is acquired. Then, the corresponding prompt vector and prompt text for the text to be processed can be determined. The text to be processed, the prompt vector, and the prompt text are then input into the question-answering model, which determines the corresponding response text. Finally, the response text is output. In the above process, the prompt vector and prompt text are used together to guide the question-answering model in responding to the text to be processed. On the one hand, the vector prompts can supplement the question-answering model as parameters to improve the accuracy of the response text generated by the model; on the other hand, the text prompts can supplement the text to be processed, enabling the question-answering model to generate response text more accurately based on more background knowledge. That is, using different forms of prompt information can supplement the question-answering model from different perspectives to improve the accuracy of the response text generated by the model.
[0213] The following describes in detail one or more embodiments of a text processing apparatus according to the present invention. Those skilled in the art will understand that these text processing apparatuses can all be configured using commercially available hardware components through the steps taught in this invention.
[0214] Figure 13 This is a schematic diagram of the structure of a text processing device provided in an embodiment of the present invention, as shown below. Figure 13 As shown, the device includes:
[0215] The text acquisition module 11 is used to acquire the text to be processed.
[0216] The prompt information determination module 12 is used to determine the prompt vector and prompt text corresponding to the text to be processed.
[0217] The response text output module 13 is used to input the text to be processed, the prompt vector, and the prompt text into the question-and-answer model so that the question-and-answer model can output the response text corresponding to the text to be processed.
[0218] The text to be processed includes the text of the question to be processed, or it includes the text of the question stem to be processed and the text of the question stem to be processed corresponding to the text of the question to be processed.
[0219] Optionally, the prompt information determination module 12 is used to encode the text pair to be processed; determine a preset number of target indices based on the similarity between the encoding result and the indices in the prompt pool; and determine the prompt vector pointed to by the target index in the prompt pool as the prompt vector corresponding to the text to be processed.
[0220] Optionally, the apparatus further includes: a first training module 14, configured to acquire training text, the training text including training question text and reference answer text corresponding to the training question text, or including training question stem text, training question text corresponding to the training question stem text, and reference answer text corresponding to the training question text; determine the hint vector and hint text corresponding to the training text, the hint vector corresponding to the training text being included in the hint pool; use the training text, the hint vector corresponding to the training text, and the hint text corresponding to the training text as training data, and use the reference answer text corresponding to the training text as supervision information, to train the question answering model, the hint vector corresponding to the training text, and the index pointing to the hint vector corresponding to the training text.
[0221] Optionally, the prompt information determination module 12 is configured to determine a first target text in the historical text based on a first similarity between the historical text and the text to be processed; input the text to be processed and the first target text into a pre-trained model to generate a first prompt text corresponding to each of the first target texts by the pre-trained model; and determine the prompt text corresponding to the text to be processed based on the first prompt text.
[0222] The historical text includes the historical question text and the corresponding reference answer text, or it includes the historical question stem text, the historical question text corresponding to the historical question stem text, and the corresponding reference answer text.
[0223] Optionally, the prompt information determination module 12 is configured to determine a second similarity between the fused text and the text to be processed based on the text to be processed, the first target text, and the first prompt text, wherein the fused text is the fusion result of the first target text and the first prompt text; determine a second target text in the first target text based on the second similarity; and determine the prompt text corresponding to each of the second target texts as the prompt text corresponding to the text to be processed, wherein the number of the first target texts is greater than the number of the second target texts.
[0224] Optionally, the prompt information determination module 12 is further configured to input the historical text and the text to be processed into a first filtering model, so that the first filtering model outputs the first similarity; and determine the first target text based on the first similarity.
[0225] Optionally, the device further includes: a second training module 15, configured to: select other training texts in the training text set whose similarity to any training text in the training text set meets a preset requirement; input the any training text and the other training texts into the first filtering model, so that the first filtering model outputs a first similarity distribution, the first similarity distribution reflecting the similarity between the any training text and the other training texts; input the any training text and the other training texts into the pre-training model, so that the pre-training model outputs prompt texts corresponding to the other training texts, and a first probability distribution of the prompt texts, the first probability distribution reflecting the confidence that the prompt texts are reference answer texts corresponding to the any training texts; and train the first filtering model based on the loss calculation result between the first similarity distribution and the first probability distribution.
[0226] Optionally, the prompt information determination module 12 is used to input the text to be processed, the first prompt text, and the first target text into the second filtering model, so that the second filtering model outputs the second similarity.
[0227] Optionally, the device further includes a performance index determination module 16, used to determine the performance indexes of the question-answering model, the first screening model, and the second screening model.
[0228] The third training module 17 is used to: determine candidate training texts in the training text set whose similarity to the target training text meets preset requirements; input the candidate training texts, the corresponding prompt vectors, and the corresponding prompt texts into the question answering model, so that the question answering model outputs the predicted answer texts corresponding to the candidate training texts; determine the loss value distribution between the predicted answer texts and the corresponding reference answer texts of the candidate training texts; input the target training text and the candidate training texts into the first filtering model, so that the first filtering model outputs a third similarity distribution, which reflects the similarity between the target training text and the candidate training texts; input the target training text, the candidate training texts, and the prompt texts corresponding to the candidate training texts output by the pre-trained model into the second filtering model, so that the second filtering model outputs a fourth similarity distribution, which reflects the confidence level of the similarity between the target training text and the candidate training texts; and use the distributions corresponding to models whose performance indicators meet preset requirements as supervision information to train the remaining models.
[0229] Figure 13 The device shown can perform Figures 1 to 11 For the methods shown in the embodiments, the parts not described in detail in this embodiment can be referred to the following: Figures 1 to 11 The relevant descriptions of the illustrated embodiments are provided below. For the execution process and technical effects of this technical solution, please refer to [link / reference]. Figures 1 to 11 The descriptions in the illustrated embodiments will not be repeated here.
[0230] Figure 14 This is a schematic diagram of the structure of a human-computer interaction device provided in an embodiment of the present invention, as shown below. Figure 14 As shown, the device includes:
[0231] The text acquisition module 21 is used to acquire the text to be processed in response to interactive operations on the interactive device.
[0232] The comprehensive prompt information determination module 22 is used to determine the prompt vector and prompt text corresponding to the text to be processed.
[0233] The response text determination module 23 is used to input the text to be processed, the prompt vector, and the prompt text into the question-and-answer model so that the question-and-answer model can determine the response text corresponding to the text to be processed.
[0234] The text output module 24 is used to output the response text.
[0235] Optionally, the text acquisition module 21 is configured to acquire the voice to be processed in response to a voice input operation; convert the voice to be processed into the text to be processed; and output the response text, including playing the response text.
[0236] In one possible design, the text processing methods provided in the above embodiments can be applied to an electronic device, such as... Figure 15 As shown, the electronic device may include a first processor 31 and a first memory 32. The first memory 32 is used to store data supporting the electronic device in performing the above-described actions. Figures 1 to 11 In the text processing method program provided in the illustrated embodiment, the first processor 31 is configured to execute the program stored in the first memory 32.
[0237] The program includes one or more computer instructions, wherein when executed by the first processor 31, the one or more computer instructions can perform the following steps:
[0238] Get the text to be processed;
[0239] Determine the prompt vector and prompt text corresponding to the text to be processed;
[0240] The text to be processed, the prompt vector, and the prompt text are input into the question-answering model, so that the question-answering model outputs the response text corresponding to the text to be processed.
[0241] Optionally, the first processor 31 is also used to perform the aforementioned Figures 1 to 11 All or part of the steps in the illustrated embodiments.
[0242] The structure of the electronic device may also include a first communication interface 33 for the electronic device to communicate with other devices or communication systems.
[0243] In addition, embodiments of the present invention provide a computer storage medium for storing computer software instructions used by the aforementioned electronic device, which includes instructions for executing the above-mentioned... Figures 1 to 11 The text processing method shown involves the following procedures.
[0244] In one possible design, the human-computer interaction methods provided in the above embodiments can be applied to an electronic device, such as... Figure 16 As shown, the electronic device may include a second processor 41 and a second memory 42. The second memory 42 is used to store data supporting the electronic device in performing the above-described actions. Figure 12 In the human-computer interaction method program provided in the illustrated embodiment, the second processor 41 is configured to execute the program stored in the second memory 42.
[0245] The program includes one or more computer instructions, wherein the one or more computer instructions, when executed by the second processor 41, can perform the following steps:
[0246] In response to interactive operations on the interactive device, acquire the text to be processed;
[0247] Determine the prompt vector and prompt text corresponding to the text to be processed;
[0248] The text to be processed, the prompt vector, and the prompt text are input into the question-answering model so that the question-answering model can determine the response text corresponding to the text to be processed.
[0249] Output the response text.
[0250] Optionally, the second processor 41 is also used to perform the aforementioned Figure 12 All or part of the steps in the illustrated embodiments.
[0251] The structure of the electronic device may also include a second communication interface 43 for the electronic device to communicate with other devices or communication systems.
[0252] In addition, embodiments of the present invention provide a computer storage medium for storing computer software instructions used by the aforementioned electronic device, which includes instructions for executing the above-mentioned... Figure 12 The program involved in the human-computer interaction method shown.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text processing method, characterized in that, include: Get the text to be processed; Determine the prompt vector and prompt text corresponding to the text to be processed. The prompt vector serves as a supplementary parameter of the question-answering model, and the prompt text is semantically related to the text to be processed. The text to be processed, the prompt vector, and the prompt text are input into the question-and-answer model so that the question-and-answer model can output the response text corresponding to the text to be processed. Determining the prompt vector corresponding to the text to be processed includes: The text pairs to be processed are encoded; Based on the similarity between the encoding result and the index in the hint pool, a preset number of target indexes are determined; The prompt vector pointed to by the target index in the prompt pool is determined as the prompt vector corresponding to the text to be processed.
2. The method according to claim 1, characterized in that, The text to be processed includes the text of the question to be processed, or it includes the text of the question stem to be processed and the text of the question stem to be processed corresponding to the text of the question to be processed.
3. The method according to claim 1, characterized in that, The method further includes: Obtain training text, which includes training question text and corresponding reference answer text, or training question text, training question text, and corresponding reference answer text. Determine the cue vector and cue text corresponding to the training text, wherein the cue vector corresponding to the training text is included in the cue pool; The training text, the corresponding prompt vector, and the corresponding prompt text are used as training data, and the reference answer text corresponding to the training text is used as supervision information to train the question-answering model, the prompt vector corresponding to the training text, and the index pointing to the prompt vector corresponding to the training text.
4. The method according to claim 2, characterized in that, Historical texts include historical question texts and corresponding reference answer texts, or they may include historical question texts, historical question texts corresponding to the historical question texts, and corresponding reference answer texts. The step of determining the prompt text corresponding to the text to be processed includes: Based on the first similarity between the historical text and the text to be processed, a first target text is determined in the historical text; The text to be processed and the first target text are input into the pre-trained model so that the pre-trained model can generate the first prompt text corresponding to the first target text. The prompt text corresponding to the text to be processed is determined based on the first prompt text.
5. The method according to claim 4, characterized in that, The step of determining the prompt text corresponding to the text to be processed based on the first prompt text includes: Based on the text to be processed, the first target text, and the first prompt text, a second similarity is determined between the fused text and the text to be processed, wherein the fused text is the fusion result of the first target text and the first prompt text; Based on the second similarity, a second target text is determined in the first target text; The prompt text corresponding to each of the second target texts is determined as the prompt text corresponding to the text to be processed, and the number of the first target texts is greater than the number of the second target texts.
6. The method according to claim 4, characterized in that, The step of determining the first target text in the historical text based on the first similarity between the historical text and the text to be processed includes: The historical text and the text to be processed are input into the first filtering model, so that the first filtering model outputs the first similarity. The first target text is determined based on the first similarity.
7. The method according to claim 6, characterized in that, The method further includes: In the training text set, select other training texts whose similarity to any training text in the training text set meets the preset requirements; The first screening model is input into any training text and the other training texts, and the first screening model outputs a first similarity distribution, which reflects the similarity between any training text and the other training texts. The pre-trained model is input with any training text and the other training texts, and the pre-trained model outputs the prompt text corresponding to the other training texts, and the first probability distribution of the prompt texts, wherein the first probability distribution reflects the confidence that the prompt texts are the reference answer texts corresponding to any training texts. The first screening model is trained based on the loss calculation result between the first similarity distribution and the first probability distribution.
8. The method according to claim 6, characterized in that, The step of determining a second similarity between the first target text and the text to be processed based on the text to be processed, the first prompt text, and the first target text includes: The text to be processed, the first prompt text, and the first target text are input into the second filtering model, so that the second filtering model outputs the second similarity.
9. The method according to claim 8, characterized in that, The method further includes: Determine the performance metrics for the question-answering model, the first filtering model, and the second filtering model; In the training text set, select alternative training texts that meet the preset requirements for similarity with the target training text; The candidate training text, the prompt vector corresponding to the candidate training text, and the prompt text corresponding to the candidate training text are input into the question answering model, so that the question answering model outputs the predicted answer text corresponding to the candidate training text; Determine the loss value distribution between the predicted answer text of the candidate training text and the reference answer text corresponding to the candidate training text; The target training text and the candidate training text are input into the first screening model, so that the first screening model outputs a third similarity distribution, which reflects the similarity between the target training text and the candidate training text. The target training text, the candidate training text, and the prompt text corresponding to the candidate training text output by the pre-trained model are input into the second filtering model, so that the second filtering model outputs a fourth similarity distribution, which reflects the confidence level of the similarity between the target training text and the candidate training text. The distribution corresponding to the model that meets the performance requirements is used as supervision information to train the remaining models.
10. A human-computer interaction method, characterized in that, include: In response to interactive operations on the interactive device, acquire the text to be processed; Determine the prompt vector and prompt text corresponding to the text to be processed. The prompt vector serves as a supplementary parameter of the question-answering model, and the prompt text is semantically related to the text to be processed. The text to be processed, the prompt vector, and the prompt text are input into the question-and-answer model so that the question-and-answer model can determine the response text corresponding to the text to be processed. Output the response text; Determining the prompt vector corresponding to the text to be processed includes: The text pairs to be processed are encoded; Based on the similarity between the encoding result and the index in the hint pool, a preset number of target indexes are determined; The prompt vector pointed to by the target index in the prompt pool is determined as the prompt vector corresponding to the text to be processed.
11. The method according to claim 10, characterized in that, The step of obtaining the text to be processed in response to an interactive operation on the interactive device includes: In response to voice input, acquire the voice to be processed; Convert the speech to be processed into the text to be processed; The output of the response text includes: Play the response text.
12. An electronic device, characterized in that, include: A memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor performs a text processing method as described in any one of claims 1 to 9, or performs a human-computer interaction method as described in any one of claims 10 to 11.
13. A non-transitory machine-readable storage medium, characterized in that, The non-transitory machine-readable storage medium stores executable code that, when executed by a processor of an electronic device, causes the processor to perform a text processing method as described in any one of claims 1 to 9, or causes the processor to perform a human-computer interaction method as described in any one of claims 10 to 11.