Data processing method, question and answer method, electronic device, storage medium and product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]现有的答案生成方案中,在上下文内容较多时,模型的计算量大,导致答案的生成速度慢
首先,通过第一损失驱动预设模型学习从历史对话、第i轮问题到第i轮答案之间的关联关系,保证预设模型生成的预测答案与参考答案的一致性;其次,第二字符作为预测答案的摘要,其与参考答案对比形成的第二损失,能促使预设模型在学习生成准确答案的同时,同步学习对预测答案的压缩能力,使预设模型生成的第二字符可精准表征预测答案的核心信息;最终,通过第一损失和第二损失协同优化预设模型的参数,可在不损失预设模型答案生成准确性的前提下,让预设模型同时具备高准确性答案生成与答案压缩的双重能力,后续推理时仅需调用对应的压缩字符即可替代原始多轮对话内容,减少输入token数量,有效提高了答案生成速度。
Smart Images

Figure CN121681734B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a data processing method, a question-answering method, an electronic device, a storage medium, and a product. Background Technology
[0002] With the rapid development of natural language processing technology, the demand for answer generation in fields such as intelligent customer service, dialogue systems, and virtual assistants is increasing.
[0003] In existing answer generation schemes, when there is a lot of context, the computational load of the model is large, resulting in a slow answer generation speed. Summary of the Invention
[0004] To address the aforementioned technical problems, embodiments of this application provide a data processing method, a question-and-answer method, a data processing apparatus, a question-and-answer apparatus, an electronic device, a storage medium, and a product, which can improve the speed of model-generated answers.
[0005] Firstly, a data processing method is provided, including the following steps: Set a first character for the reference answer to the i-th round question, where the first character is defined as a summary of the reference answer, and i is a natural number; Using a preset model, based on the previous i-1 rounds of dialogue, the i-th round question, the reference answer to the i-th round question, and the first character, the predicted answer to the i-th round question and the second character corresponding to the predicted answer are predicted. The second character is used to indicate the summary generated by the preset model for the predicted answer to the i-th round question. Based on the reference answer and the predicted answer of the i-th round question, determine the first loss; Based on the second character and the reference answer to the i-th round question, determine the second loss; Based on the first loss and the second loss, the parameters of the preset model are updated.
[0006] Secondly, a question-and-answer method is also provided, which includes: Using a pre-defined model, based on the fifth character corresponding to each question in the first k-1 rounds of dialogue, the sixth character corresponding to the predicted answer to each question in the first k-1 rounds of dialogue, and the question in the kth round, the predicted answer to the question in the kth round, the seventh character corresponding to the question in the kth round, and the eighth character corresponding to the predicted answer to the question in the kth round are predicted, where k is a natural number; The fifth character is a summary generated by the preset model for the question in the dialogue when predicting the answer in the corresponding round of dialogue; The sixth character is a summary generated by the preset model for the predicted answer in the corresponding round of dialogue. The seventh character is used to indicate the summary generated by the preset model for the k-th round problem; The eighth character is used to indicate the summary generated by the preset model for the predicted answer to the question in the k-th round.
[0007] Thirdly, a data processing apparatus is provided, the data processing apparatus comprising: The setting module is used to set a first character for the reference answer to the i-th round question, where the first character is defined as a summary of the reference answer, and i is a natural number; The first prediction module is used to predict the predicted answer to the i-th round question and the second character corresponding to the predicted answer by using a preset model based on the previous i-1 rounds of dialogue, the i-th round question, the reference answer to the i-th round question, and the first character. The second character is used to indicate the summary generated by the preset model for the predicted answer to the i-th round question. The first determining module is used to determine the first loss based on the reference answer and the predicted answer of the i-th round question; The second determining module is used to determine the second loss based on the second character and the reference answer to the i-th round question; An update module is used to update the parameters of the preset model based on the first loss and the second loss.
[0008] Fourthly, a question-and-answer device is also provided, the question-and-answer device comprising: The second prediction module is used to predict the predicted answer to the k-th round question, the seventh character of the k-th round question, and the eighth character of the predicted answer to the k-th round question based on the fifth character of each round question in the first k-1 rounds of dialogue, the sixth character of the predicted answer to each round question, and the k-th round question, where k is a natural number. The fifth character is a summary generated by the preset model for the question in the dialogue when predicting the answer in the corresponding round of dialogue; The sixth character is a summary generated by the preset model for the predicted answer in the corresponding round of dialogue. The seventh character is used to indicate the summary generated by the preset model for the k-th round problem; The eighth character is used to indicate the summary generated by the preset model for the predicted answer to the question in the k-th round.
[0009] Fifthly, embodiments of this application also provide an electronic device, including a processor and a memory, wherein the memory stores multiple instructions; the processor loads instructions from the memory to execute the steps of any data processing method provided in embodiments of this application, or to execute the steps of any question-and-answer method provided in embodiments of this application.
[0010] Sixthly, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform the steps of any data processing method provided in embodiments of this application, or to perform the steps of any question-and-answer method provided in embodiments of this application.
[0011] In a seventh aspect, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of any data processing method provided in embodiments of this application, or execute the steps of any question-and-answer method provided in embodiments of this application.
[0012] Beneficial effects: First, the first loss drives the pre-defined model to learn the relationships between historical dialogues, the i-th round question, and the i-th round answer, ensuring the consistency between the predicted answer generated by the pre-defined model and the reference answer. Second, the second character serves as a summary of the predicted answer, and its comparison with the reference answer forms the second loss, enabling the pre-defined model to learn to compress the predicted answer while simultaneously learning to generate accurate answers. This allows the second character generated by the pre-defined model to accurately represent the core information of the predicted answer. Finally, by collaboratively optimizing the parameters of the pre-defined model through the first and second losses, the model can simultaneously possess the dual capabilities of high-accuracy answer generation and answer compression without sacrificing the accuracy of answer generation. In subsequent inference, only the corresponding compressed character needs to be called to replace the original multi-round dialogue content, reducing the number of input tokens and effectively improving the answer generation speed. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram illustrating the application environment of the data processing method and / or question-and-answer method provided in some embodiments of this application; Figure 2 This is a schematic flowchart of a data processing method provided in some embodiments of this application; Figure 3This is a flowchart illustrating data processing methods in existing technologies; Figure 4 This is another schematic flowchart of a data processing method provided in some embodiments of this application; Figure 5 This is a flowchart illustrating the process of determining the first loss and the second loss provided in some embodiments of this application; Figure 6 This is a flowchart illustrating the question-and-answer method in existing technologies; Figure 7 This is a flowchart illustrating the question-and-answer method provided in some embodiments of this application; Figure 8 This is another flowchart illustrating the question-and-answer method provided in some embodiments of this application; Figure 9 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a question-and-answer device provided in one embodiment of this application; Figure 11 This is a schematic diagram of the internal structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0016] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0017] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0018] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated.
[0019] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0020] The following describes the relevant content, terms, meanings, technical issues, technical solutions, and beneficial effects involved in the embodiments of this application.
[0021] Terminology Explanation: LLM (Large Language Model) is a deep learning model built on the Transformer architecture. Its core features are that it has a massive number of parameters (billions or trillions) and is pre-trained on ultra-large-scale text datasets through self-supervised learning.
[0022] TTFT (Time to First Token) refers to the time from the user input prompt to the model outputting the first token during LLM generation.
[0023] TPOT (Time Per Output Token): refers to the average generation time of each subsequent token after the first token is output during LLM generation.
[0024] Streaming Generation: Streaming generation for large language models is a technique for real-time text output, allowing users to receive results word by word or sentence by sentence without waiting for the entire sequence to be generated. This technology is crucial for interactive scenarios such as chatbots, real-time translation, and code completion.
[0025] NTP (Next Token Prediction) is a core task primarily performed during the pre-training of large-scale language models. It predicts the most likely next word or symbol in a sequence based on the given context.
[0026] GQA (Grouped-Query Attention): By grouping the multiple heads in the multi-head attention mechanism, each group shares the same key and value, significantly reducing the number of key-value calculations.
[0027] Supervised Fine-Tuning (SFT) is a technique for further training large-scale language models (such as GPT) using labeled task-specific data to adapt them for downstream tasks (such as text classification and question answering). Its core function is to adjust model parameters to improve their performance and accuracy in the target domain. It is widely used in vertical domains (such as finance) or specific tasks (such as dialogue generation) to further optimize the output style, security, or business requirements of existing large models. It is particularly suitable for tasks that require efficient use of small data to quickly customize and adjust the base model.
[0028] Intelligent dialogue systems centered around large language models are replacing traditional human agents in various sectors of the financial industry, such as marketing, customer service, debt collection, and asset management, providing businesses with cost-effective, efficient, and secure solutions. In practical large language model-based generative intelligent dialogue systems, especially those using voice as the interaction method, the faster generation of the next round of dialogue plays a crucial role in the overall system experience, particularly in terms of fluency and human-likeness.
[0029] Optimizing the generation speed of LLM in intelligent dialogue mainly revolves around two core metrics: TTFT and TPOT, with TTFT being the more critical one. This is because the system cannot output any generation information to the other party before the first token is generated. After the first token is generated, the system can use a streaming generation approach, similar to large language models, to generate and output subsequent tokens simultaneously. Therefore, TPOT is not a key indicator of generation speed in LLM-based generative intelligent dialogue systems.
[0030] The majority of time in TTFT is consumed in computation. During the prefill phase at the start of each generation process, we need to calculate the key-value pair for each input token. Longer hints contain more input tokens, leading to a greater amount of computation. Existing technologies typically optimize the overall computational load by reducing the number of tokens that need to be computed or improving the efficiency of key-value calculation for a single token. One approach is to reduce the number of tokens that need to be computed through dynamic token pruning. This method dynamically selects a subset of important tokens at each level based on attention scores, and then selectively computes only the key-value pairs of tokens crucial to the NTP (Next Token Prediction) prediction task, delaying or discarding the computation of other tokens. The proportion of pruned tokens changes with the task's evolution cycle; more tokens are retained in the early stages to balance efficiency and accuracy, while a higher pruning ratio is used in later stages. Another approach is the prefix caching scheme, which improves the efficiency of key-value calculation for a single token by caching. By caching the key-value values of Prompt with the same prefix, new requests with matching prefixes can directly reuse the cache, skipping pre-filling calculations. This can produce good results in scenarios with repeated prefixes, such as multi-turn dialogues and prompts.
[0031] The two methods mentioned above have some drawbacks: 1. Both dynamic token pruning and prefix caching schemes require specific adaptation modifications for different model attention architectures, such as GQA (Grouped-Query Attention). In particular, for some general architectures that are universally applicable to different attention mechanisms, and new attention architectures are constantly emerging, directly applying old pruning schemes may lead to a significant reduction in existing optimization performance. Therefore, these schemes are not independent of the model's attention mechanism and lack universal applicability and high generalization.
[0032] 2. Both of the above-mentioned approaches have a significant negative impact on the overall output quality (dialogue effect) of generative intelligent dialogue systems, especially the token pruning approach. Its core is to first predict the importance of tokens based on heuristic algorithms, and then dynamically remove input tokens that are not important to the current prediction during the calculation process. For example, it uses the distribution entropy of a token among multiple attention heads to represent its importance; the smaller the entropy value, the fewer different attention heads are focusing on that token, and the more likely that token should be pruned. However, this mechanism does not consider its own semantics and contextual information. Pruning may lead to highly inaccurate context due to the loss of key information. Furthermore, pruning is irreversible; tokens pruned early on may become crucial in subsequent generation, all of which will affect the generation quality.
[0033] 3. Although the prefix caching scheme does not affect the generation quality, it is essentially a retrieval method that sacrifices space for time. This increases storage costs and affects lightweight deployment. Furthermore, in scenarios with high dynamic repetition rates and low repetition rates, the probability of accessing the cache is low, resulting in a lack of significant optimization effect on the system's TTFT.
[0034] Therefore, existing TTFT optimization schemes (which optimize the overall computation by reducing the number of tokens that need to be computed or improving the key-value computation efficiency of a single token) are more suitable for tasks with less stringent output quality requirements, such as large-scale data preprocessing, but are extremely risky in scenarios with high precision requirements.
[0035] To address the aforementioned issues, this application provides a data processing method, a question-answering method, a data processing apparatus, a question-answering apparatus, an electronic device, a computer-readable storage medium, and a computer program product. In this application, firstly, a first loss is used to drive a preset model to learn the correlation between historical dialogues, the i-th round question, and the i-th round answer, ensuring the consistency between the predicted answer generated by the preset model and the reference answer. Secondly, a second character serves as a summary of the predicted answer; its comparison with the reference answer forms a second loss, which enables the preset model to learn to generate accurate answers while simultaneously learning to compress the predicted answers, allowing the second character generated by the preset model to accurately represent the core information of the predicted answer. Finally, by collaboratively optimizing the parameters of the preset model through the first and second losses, the preset model can simultaneously possess the dual capabilities of high-accuracy answer generation and answer compression without sacrificing the accuracy of answer generation. Subsequent inference only requires calling the corresponding compressed character to replace the original multi-round dialogue content, reducing the number of input tokens and effectively improving the answer generation speed.
[0036] To better understand the data processing method, question-and-answer method, data processing apparatus, question-and-answer apparatus, electronic device, computer-readable storage medium, and computer program product provided in the embodiments of this application, the application environment applicable to the embodiments of this application is described below.
[0037] Please see Figure 1 , Figure 1 This diagram illustrates an application environment for a data processing method and / or question-answering method provided in an embodiment of this application. As one implementation, the data processing method and / or question-answering method provided in this embodiment can be applied to an electronic device. This electronic device can be, for example,... Figure 1 The server 110 shown can be connected to the terminal device 120 via a network. The network serves as a medium for providing a communication link between the server 110 and the terminal device 120. The network can include various connection types, such as wired communication links, wireless communication links, etc., and this embodiment is not limited thereto. Optionally, in other embodiments, the electronic device can also be a smartphone, laptop, etc.
[0038] It should be understood that Figure 1 The server 110, network, and terminal device 120 shown are merely illustrative. Depending on the implementation requirements, any number of servers, networks, and terminal devices can be included. For example, server 110 can be a physical server or a server cluster consisting of multiple servers, and terminal device 120 can be a mobile phone, tablet, desktop computer, laptop computer, smart speaker, smart wearable device, etc. It is understood that embodiments of this application can also allow multiple terminal devices 120 to access server 110 simultaneously.
[0039] In some embodiments, the terminal device 120 may send a model training request or a question-and-answer request to the server. After receiving the model training request or question-and-answer request, the server 110 may train the model or generate answers using the data processing method or question-and-answer method described in the embodiments of this application.
[0040] As another implementation, the server 110 and the terminal device 120 described in this application embodiment can be integrated. For example, the server 110 or the terminal device 120 can directly receive the model training request or question answering request input by the user and train the model or generate answers.
[0041] The following will separately describe in detail the data processing method, question-and-answer method, data processing device, question-and-answer device, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present application with reference to the accompanying drawings. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from that shown in the accompanying drawings.
[0042] On the one hand, this embodiment provides a data processing method. As Figure 2 shown, the data processing method includes steps S21 - S25: S21. Set a first character for the reference answer of the i-th round of questions, where the first character is defined as the abstract of the reference answer, and i is a natural number.
[0043] Specifically, first prepare the samples for model training, obtain a conversation including m rounds of interactions between users and agents (which can be a voice conversation or a text conversation. If it is a voice conversation, convert it to text), and perform tokenization on it, that is: cut the content of the conversation into basic units that the model can understand and process. A basic unit is a token. The tokenized representation is as follows: The first round of user questions: token<1, user>_1, token<1, user>_2, ……, token<1, user>_n11; The first round - agent answer: token<1, agent>_1, token<1, agent>_2, ……, token<1, agent>_n12; The second round - user questions: token<2, user>_1, token<2, user>_2, ……, token<1, user>_n21; The second round - agent answer: token<2, agent>_1, token<2, agent>_2, ……, token<2, agent>_n22; …… The i-th round - user questions: token<i, user>_1, token<2, user>_2, ……, token<1, user>_ni1; The i-th round - agent answer: token<2, agent>_1, token<2, agent>_2, ……, token<2, agent>_ni2; …… The m-th round - user questions: token<m, user>_1, token<m, user>_2, ……, token<m, user>_nm1; The m-th round - seat answer: token<m, seat>_1, token<m, seat>_2, ……, token<m, seat>_nm2; Among them, ni1 represents the token length of the user's question in the i-th round of conversation, ni2 represents the token length of the seat answer in the i-th round of conversation, nm1 represents the token length of the user's question in the m-th round of conversation, and nm2 represents the token length of the seat answer in the m-th round of conversation.
[0044] In this embodiment, the model training can be SFT (Supervised Fine-Tuning) training. Multiple samples can be generated according to the above conversation. Each sample takes the seat answer of a certain round as the sample target. Thus, a conversation containing m rounds can generate m SFT training samples. Each sample contains two parts: input X and target Y. For example, the sample with the seat answer of the i-th round as the target is: Input X = { token<1, user>_1, token<1, user>_2, ……, token<1, user>_n11, token<1, seat>_1, token<1, seat>_2, …, token<1, seat>_n12, ……, token<i, user>_1, token<i, user>_2, ……, token<i, user>_ni1}; Target Y = { token<i, seat>_1, token<i, seat>_2, ……, token<i, seat>_ni2}.
[0045] It can be seen that in the sample with the seat answer of the i-th round as the target, the input X is the first i - 1 rounds of conversations (the questions and answers of each round in the first i - 1 rounds of conversations) and the question of the i-th round, and the target Y is the seat answer of the i-th round. After generating the samples, the samples can be used for training.
[0046] In this embodiment, after generating the sample with the seat answer of the i-th round as the target, a first character is also set for the reference answer of the i-th round question (i.e., the seat answer of the i-th round). The first character is defined as the summary of the reference answer, where i is a natural number. The first character is a preset character. For example, it can be agent_sum_i. The first character represents the summary label of the reference answer of the i-th round question.
[0047] A summary of the reference answer refers to a textual summary obtained by compressing and refining the core semantic information or key content of the reference answer. It aims to capture the semantic essence of the reference answer. The first character is the symbolic form of this summary. For example, a specific character (such as "agent_sum_i") can be assigned to the reference answer for the i-th round of questions; this character is defined as the summary of that reference answer. For instance, if the reference answer for the i-th round of questions is "We recommend that you prioritize using fixed-term deposits to ensure the safety of your funds," then its summary might be compressed to "fixed-term deposits ensure safety," and the first character "agent_sum_i" represents this summary.
[0048] S22. Using a preset model, based on the previous (i-1) rounds of dialogue, the i-th round question, the reference answer to the i-th round question, and the first character, predict the predicted answer to the i-th round question and the second character corresponding to the predicted answer. The second character is used to indicate the summary generated by the preset model for the predicted answer to the i-th round question.
[0049] Specifically, the step of predicting the predicted answer to the i-th round question and the corresponding second character based on the first i-1 rounds of dialogue, the i-th round question, the reference answer, and the first character using a preset model includes: concatenating the first i-1 rounds of dialogue, the i-th round question, and the reference answer to the i-th round question to obtain concatenated text; performing word segmentation on the concatenated text to obtain a word sequence; encoding the words in the word sequence to obtain a first feature vector corresponding to the word sequence; concatenating the second feature vector set for the first character with the first feature vector to obtain a third feature vector; and predicting the predicted answer to the i-th round question and the corresponding second character based on the third feature vector using the preset model.
[0050] The preset model can be a large language model, such as an autoregressive decoding architecture based on a multi-layer Transformer module self-attention mechanism. The process of the preset model generating the predicted answer and the second character is as follows: The preset model generates a vector at each output position j based on the third feature vector. Then, a fully connected layer is used to calculate the vector at each position j to determine the probability of each word in the known vocabulary appearing at that position j. Finally, a sampling method generates the answer token at that position j. The answer tokens at each output position are concatenated to obtain the predicted answer for the i-th round of questions. Then, the preset model generates the second character based on the predicted answer. The second character indicates the summary generated by the preset model for the predicted answer and represents the compressed information of the predicted answer.
[0051] The second feature vector is a feature vector randomly generated for the first character. Initially, it has no semantic information. However, when the preset model generates each word of the predicted answer, the attention mechanism of the preset model will pay attention to all parts of the third feature vector throughout the process. This includes the second feature vector randomly generated for the first character. The second feature vector participates in the calculation of the preset model when generating the predicted word at each step. The preset model will learn to use the second feature vector to guide the generation of the predicted answer, thereby establishing the association between the second feature vector and the predicted answer.
[0052] The summary of the predicted answer is a core text summary obtained by the pre-defined model after extracting the content of the complete predicted answer it generates. It aims to capture the semantic essence of the predicted answer in a concise textual form. The second character is a compressed representation of this summary; this second character is an encoded form of the summary of the predicted answer. Essentially, they are different representations of the same information—the summary is readable text, and the second character is its corresponding encoding.
[0053] In existing technologies, the predicted answer to the i-th round question is usually generated only based on the previous i-1 rounds of dialogue, the i-th round question, and the reference answer. However, no character is defined to indicate the summary of the reference answer, and the preset model does not generate a character representing the summary of the predicted answer. As a result, in each reasoning process, the existing technology needs to input all contextual information into the preset model, resulting in a huge amount of input, a serious overload of computation, and a low speed in generating the predicted answer.
[0054] S23. Based on the reference answer and the predicted answer of the i-th round question, determine the first loss.
[0055] Specifically, the first loss can be the cross-entropy loss between the reference answer and the predicted answer.
[0056] S24. Based on the second character and the reference answer to the i-th round question, determine the second loss.
[0057] In some embodiments, determining the second loss based on the second character and the reference answer to the i-th round question includes: decompressing the second character to obtain a first text; and determining the second loss based on the reference answer to the i-th round question and the first text.
[0058] Specifically, the first text obtained after decompressing the second character is compared with the reference answer by calculating the cross-entropy of the tokens at each position to obtain the cross-entropy loss corresponding to each position. The cross-entropy losses of all positions are summed to obtain the second loss.
[0059] S25. Based on the first loss and the second loss, update the parameters of the preset model.
[0060] In existing technologies, the model parameters are typically updated only using the first loss, such as... Figure 3 The diagram illustrates the data processing method in the prior art. First, the first i-1 rounds of dialogue, the i-th round question, and the reference answer are concatenated into a complete text. Then, a token segmenter converts it into a token sequence. The token sequence is encoded to obtain an encoded feature vector. Next, the model (an autoregressive decoding architecture based on a multi-layer Transformer module self-attention mechanism) generates a vector at each output position based on the encoded feature vector. A fully connected layer is then used to calculate the vector at each position to determine the probability of each word in the known vocabulary at that position. Finally, a sampling method generates the answer token at that position. The training objective is to align the model's generated predicted answer with the reference answer. The loss function calculates the cross-entropy loss (first loss) between the tokens of the output predicted answer and the corresponding tokens of the reference answer. The loss function only calculates the loss between the predicted answer and the reference answer, with the aim of teaching the model to generate accurate predicted answers.
[0061] In this embodiment, a first loss is determined based on the difference between the reference answer and the predicted answer to optimize the accuracy of answer generation of the preset model. At the same time, a second loss is determined based on the association between the second character and the reference answer to enable the preset model to learn to generate accurate compressed representations. Finally, by combining the first loss and the second loss to update the model parameters, the preset model not only learns to accurately generate predicted answers during training, but also learns to generate high-quality compressed summaries for the predicted answers. Thus, during the inference stage, the characters corresponding to these summaries can replace the original long text input, significantly reducing the number of input tokens, significantly reducing the first token generation time (TTFT), and improving the speed of answer generation.
[0062] In some embodiments, the method further includes: setting a third character for the i-th round question, wherein the third character is defined as a summary of the i-th round question; The step of predicting the predicted answer to the i-th round question and the corresponding second character based on the first i-1 rounds of dialogue, the i-th round question, the reference answer to the i-th round question, and the first character using a preset model includes: Using the preset model, based on the first i-1 rounds of dialogue, the i-th round question, the third character, the reference answer to the i-th round question, and the first character, the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question are predicted. The fourth character is used to indicate the summary generated by the preset model for the i-th round question. After obtaining the fourth character corresponding to the i-th round question, the method further includes: determining a third loss based on the fourth character and the i-th round question; The step of updating the parameters of the preset model based on the first loss and the second loss includes: updating the parameters of the preset model based on the first loss, the second loss and the third loss.
[0063] Specifically, in this embodiment, not only is a first character set for the reference answer of the i-th round, but a third character is also set for the question of the i-th round to represent its summary. Based on the previous i-1 rounds of dialogue, the question of the i-th round, the third character, the reference answer, and the first character, the preset model also generates a fourth character indicating the summary of the question of the i-th round, which represents the compressed information of the question of the i-th round. By decompressing the fourth character, the decompressed text is obtained, and the cross-entropy loss between the decompressed text and the question of the i-th round is calculated to obtain the third loss. Based on the first loss, the second loss, and the third loss, the total loss is determined (for example, by weighted summation of the three to obtain the total loss), and the parameters of the preset model are updated according to the total loss.
[0064] In some embodiments, the step of predicting the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question based on the first i-1 rounds of dialogue, the i-th round question, the third character, the reference answer to the i-th round question, and the first character using a preset model includes steps S261-S264: S261. Concatenate the first i-1 rounds of dialogue, the i-th round question, and the reference answer to the i-th round question to obtain the second text; S262. Perform word segmentation on the second text to obtain a word sequence; S263. Encode the words in the word sequence to obtain a first vector; S264. Using the preset model, based on the first vector, the second vector set for the first character, and the third vector set for the third character, predict the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question.
[0065] Similar to the second vector, the third vector is a vector randomly initialized for the third character. The third vector participates in the calculation of the preset model in the process of generating the second character. The preset model learns to use the third feature vector to guide the generation of the second character, thereby establishing the association between the third feature vector and the second character.
[0066] In some embodiments, the step of predicting the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question based on the preset model, using the first vector, the second vector set for the first character, and the third vector set for the third character, includes steps S2641-S2643: S2641. Insert the second vector after the vector corresponding to the reference answer of the i-th round question in the first vector to obtain the fourth vector; S2642. Insert the third vector into the fourth vector after the vector corresponding to the i-th round question and before the vector corresponding to the reference answer of the i-th round question to obtain the fifth vector; S2643. Using the preset model and based on the fifth vector, predict the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question.
[0067] Specifically, the second vector is located at the end of the first vector; that is, the first and second vectors are concatenated to obtain the fourth vector. The third vector is located after the vector of the last word in the i-th round of questions in the fourth vector and before the vector of the first word in the reference answer.
[0068] In some embodiments, predicting the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question based on the fifth vector using the preset model includes: predicting the predicted answer to the i-th round question and the fourth character corresponding to the i-th round question based on the fifth vector using the preset model; and predicting the second character corresponding to the predicted answer based on the predicted answer to the i-th round question using the preset model.
[0069] Specifically, in this embodiment, firstly, the preset model generates the predicted answer to the i-th round question based on the fifth vector, and generates the fourth character representing the summary of the i-th round question; then, the preset model generates the second character representing the summary of the predicted answer based on the predicted answer to the i-th round question.
[0070] In some other embodiments, the preset model simultaneously generates the predicted answer to the i-th round question, a fourth character representing the summary of the i-th round question, and a second character representing the summary of the predicted answer based on the fifth vector.
[0071] like Figure 4 The diagram shown is another schematic flowchart of a data processing method provided in some embodiments of this application. This data processing method includes: The third text is obtained by concatenating the (i-1)th round dialogue, the i-th round question, and the reference answer. The third text is then segmented to obtain a token sequence. A special token1 (first character) is inserted at the end of the token corresponding to the reference answer in this sequence. A special token2 (third character) is inserted at the end of the token corresponding to the i-th round question and before the token corresponding to the reference answer. Each token in the sequence except for special token1 and special token2 is encoded, and a random vector is initialized for special token1 and special token2 respectively, resulting in a feature vector for the sequence. Based on this feature vector, a pre-defined model (an autoregressive decoding architecture based on a multi-layer Transformer module self-attention mechanism) generates the predicted answer to the i-th round question, a special token3 (second character) representing a summary of the predicted answer, and a special token4 (fourth character) representing a summary of the i-th round question. The first loss is determined by the predicted answer and the reference answer. The second loss is determined by special token3 and the reference answer. The third loss is determined by special token4 and the i-th round question. The model parameters of the pre-defined model are updated based on the first, second, and third losses.
[0072] In the process of generating the special token 4, the attention mask only focuses on the user's question in the i-th round, ignoring the previous i-1 rounds of dialogue. This aligns with the principle that prediction depends only on the preceding context, ensuring that the special token 4 contains only the information that needs to be compressed. When generating the predicted answer, the attention mask does not focus on the special token 4; however, when generating the special token 3, the attention mask only focuses on the predicted answer.
[0073] like Figure 5 The diagram shown is a flowchart illustrating the determination process of the first loss and the second loss provided in some embodiments of this application. For the first loss, the cross-entropy loss is calculated for each token of the predicted answer output by the preset model and the reference answer, and the cross-entropy loss values of each token are summed to obtain the first loss. For the second loss, the special token3 output by the preset model is decompressed, and the cross-entropy loss is calculated for each token of the reference answer and the cross-entropy loss values of each token are summed to obtain the second loss.
[0074] In this embodiment, the total loss (obtained by weighted summation of the first loss, second loss, and third loss) is backpropagated through gradient optimization to all model parameters, including the embedding layer and attention layer, involved in generating the predicted answer token. This drives the preset model to learn the mapping relationship from the input to the expected output, thereby generating accurate predicted answers and a compressed representation (summary) of the user input question in the i-th round and the machine-generated answer in the i-th round.
[0075] like Figure 6 The diagram shown is a flowchart of a question-and-answer method in the prior art, which includes: The questions and answers from each round of dialogue in the first i-1 rounds, along with the question from the i-th round, are concatenated into a fourth text. The fourth text is then segmented to obtain a token sequence corresponding to the fourth text. This token sequence is then encoded to obtain a feature vector corresponding to the fourth text. Based on this feature vector, the model predicts the answer to the question from the i-th round.
[0076] Specifically, this question-answering method is applied to an intelligent question-answering system. This system simulates a dialogue between an agent and a user. After the user asks a question, the model (an autoregressive decoding architecture based on a multi-layer Transformer module self-attention mechanism) generates a predicted answer to the question in the i-th round based on contextual information (the first i-1 rounds of dialogue and the question in the i-th round). If the user's question in the first round has n11 tokens, the agent's answer in the first round has n12 tokens, the user's question in the second round has n21 tokens, the agent's answer in the first round has n22 tokens, ..., and the user's question in the i-th round has ni1 tokens, then the model performs calculations on a total of n11+n12+n21+n22+...+ni1 tokens, resulting in a huge computational load and a long time spent generating the first token of the predicted answer.
[0077] like Figure 7 The diagram shown is a flowchart illustrating a question-and-answer method provided in some embodiments of this application. The question-and-answer method includes step S71: S71. Using a preset model, based on the fifth character corresponding to each question in the first k-1 rounds of dialogue, the sixth character corresponding to the predicted answer to each question in the first k-1 rounds of dialogue, and the question in the kth round, predict the predicted answer to the question in the kth round, the seventh character corresponding to the question in the kth round, and the eighth character corresponding to the predicted answer to the question in the kth round, where k is a natural number; Wherein, the fifth character is the summary generated by the preset model for the question in the dialogue when predicting the answer in the corresponding round of dialogue; the sixth character is the summary generated by the preset model for the predicted answer in the dialogue when predicting the answer in the corresponding round of dialogue; the seventh character is used to indicate the summary generated by the preset model for the question in the kth round; and the eighth character is used to indicate the summary generated by the preset model for the predicted answer to the question in the kth round.
[0078] Specifically, this question-answering method can be applied to intelligent question-answering systems. These systems simulate an agent engaging in dialogue with a user using a pre-set model. After the user asks a question in the first round, the pre-set model generates a first-round predicted answer based on that question, along with characters indicating a summary of the first-round predicted answer (representing compressed information of the first-round predicted answer) and characters indicating a summary of the first-round question (representing compressed information of the first-round question). After the user asks a second-round question, the pre-set model generates a second-round predicted answer based on the characters indicating the first-round question summary, the characters indicating the first-round predicted answer summary, and the second-round question. It also generates characters indicating the first-round predicted answer summary and characters indicating the first-round question summary. Therefore, when predicting the answer to the second round of questions, compared to the prior art which requires all tokens of the first round of questions and answers to participate in the calculation, in this embodiment, only one token indicating the summary of the first round of questions represents all tokens of the first round of questions, and only one token indicating the summary of the predicted answer of the first round represents all tokens of the first round of answers. This greatly reduces the amount of computation of the preset model, thereby improving the generation speed of the first token of the predicted answer of the second round.
[0079] Similarly, when generating answers for the corresponding questions in subsequent rounds of dialogue, the questions and answers from previous rounds of dialogue are replaced with the corresponding characters. Therefore, each round of dialogue can be represented by two characters (one character represents the question for that round, and the other character represents the answer for that round), which greatly reduces the computational load of the preset model and reduces the time to generate the first token of the predicted answer. The more rounds, the better the effect. Moreover, the reference information of the preset model is not reduced throughout the process because the question and answer information of each round of dialogue is still there.
[0080] like Figure 8 The diagram shown is another flowchart illustrating a question-and-answer method provided in some embodiments of this application. This question-and-answer method includes: Receive the user's question in the i-th round, and obtain the special tokens corresponding to the questions and answers in each of the previous i-1 rounds of dialogue. Encode the token sequence corresponding to the question in the i-th round, the special tokens corresponding to the questions and answers in each of the previous i-1 rounds of dialogue, and the special tokens corresponding to the answers in each of the previous i-1 rounds of dialogue to obtain the encoded feature vector. Using a preset model (a model based on an autoregressive decoding architecture with a multi-layer Transformer module self-attention mechanism), generate the answer in the i-th round, a special token representing the summary of the answer in the i-th round, and a special token representing the summary of the question in the i-th round based on the encoded feature vector.
[0081] In this embodiment, the input for reasoning in the first i-1 rounds of dialogue uses a representation of each previous round—an input-specific token—instead of the tokens from the original dialogue. This significantly reduces the number of input tokens, from n11 + n12 + n21 + n22 + … + n(i-1)1 + n(i-1)2 tokens to 1 + 1 + … + 1 = 2*(i-1) tokens. If we calculate based on an average of 20 tokens per round of dialogue, this reduces the number of input tokens in each round of reasoning by a factor of 20, thereby greatly reducing the TTFT (Time to TFT) (generation time of the first token) when generating the predicted answer in each round of dialogue.
[0082] In the process of generating the answer for the i-th round, in addition to generating the token for the answer for the i-th round, two special tokens are also generated. These tokens are used to represent the user input part (the question raised by the user) in the i-th round of dialogue and the model generation part (the predicted answer generated by the model) in the i-th round of dialogue. These two representations will be directly used as input tokens in the reasoning process of generating the answers for the (i+1), (i+2), ... rounds.
[0083] The question-answering method provided in this application is universally applicable to various transformer architectures and does not require special modification or optimization. While ensuring minimal loss of input information used in the inference process and minimal impact on the accuracy and effectiveness of the inference results, it significantly reduces the TTFT index of each round of dialogue, thereby making the overall intelligent dialogue interaction process smoother. At the same time, the saved inference time can be applied to other areas that require more computation for optimization, thereby improving the overall effect of the intelligent dialogue system.
[0084] like Figure 9 The diagram shown is a structural schematic of a data processing apparatus according to an embodiment of this application. The data processing apparatus includes: Setting module 91 is used to set a first character for the reference answer to the i-th round question, where the first character is defined as a summary of the reference answer, and i is a natural number; The first prediction module 92 is used to predict the predicted answer to the i-th round question and the second character corresponding to the predicted answer by using a preset model based on the previous i-1 rounds of dialogue, the i-th round question, the reference answer to the i-th round question and the first character. The second character is used to indicate the summary generated by the preset model for the predicted answer to the i-th round question. The first determining module 93 is used to determine the first loss based on the reference answer and the predicted answer of the i-th round question; The second determining module 94 is used to determine the second loss based on the second character and the reference answer to the i-th round question; The update module 95 is used to update the parameters of the preset model based on the first loss and the second loss.
[0085] In some embodiments, the second determining module 94 is further configured to: decompress the second character to obtain the first text; and determine the second loss based on the reference answer to the i-th round question and the first text.
[0086] In some embodiments, the setting module 91 is further configured to: set a third character for the i-th round question, wherein the third character is defined as a summary of the i-th round question; The first prediction module 92 is also used for: Using the preset model, based on the first i-1 rounds of dialogue, the i-th round question, the third character, the reference answer to the i-th round question, and the first character, the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question are predicted. The fourth character is used to indicate the summary generated by the preset model for the i-th round question. After obtaining the fourth character corresponding to the i-th round question, the second determining module 94 is further configured to: determine the third loss based on the fourth character and the i-th round question; The update module 95 is further configured to: update the parameters of the preset model based on the first loss, the second loss, and the third loss.
[0087] In some embodiments, the step of predicting the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question based on a preset model, using the first i-1 rounds of dialogue, the i-th round question, the third character, the reference answer to the i-th round question, and the first character, includes: By concatenating the first i-1 rounds of dialogue, the i-th round question, and the reference answer to the i-th round question, a second text is obtained; The second text is segmented to obtain a word sequence; The words in the word sequence are encoded to obtain a first vector; Using the preset model, based on the first vector, the second vector set for the first character, and the third vector set for the third character, the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question are predicted.
[0088] In some embodiments, the first prediction module 92 is further configured to: The second vector is inserted after the vector corresponding to the reference answer to the i-th round question in the first vector to obtain the fourth vector; The third vector is inserted into the fourth vector after the vector corresponding to the i-th round question and before the vector corresponding to the reference answer of the i-th round question to obtain the fifth vector; Based on the preset model and the fifth vector, the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question are predicted.
[0089] In some embodiments, the first prediction module 92 is further configured to: Based on the preset model and the fifth vector, the predicted answer to the i-th round question and the fourth character corresponding to the i-th round question are predicted. Based on the predicted answer to the i-th round question, the second character corresponding to the predicted answer is predicted using the preset model.
[0090] The data processing apparatus provided in this application firstly drives a preset model to learn the correlation between historical dialogues, the i-th round question, and the i-th round answer through a first loss, ensuring the consistency between the predicted answer generated by the preset model and the reference answer. Secondly, the second character serves as a summary of the predicted answer, and the second loss formed by comparing it with the reference answer enables the preset model to learn the ability to compress the predicted answer while learning to generate accurate answers, so that the second character generated by the preset model can accurately represent the core information of the predicted answer. Finally, by co-optimizing the parameters of the preset model through the first and second losses, the preset model can simultaneously possess the dual capabilities of high-accuracy answer generation and answer compression without sacrificing the accuracy of answer generation. In subsequent inference, only the corresponding compressed character needs to be called to replace the original multi-round dialogue content, reducing the number of input tokens and effectively improving the answer generation speed.
[0091] like Figure 10 The diagram shown is a structural schematic of a question-and-answer device according to an embodiment of this application. The question-and-answer device includes: The second prediction module 1001 is used to predict the predicted answer to the k-th round question, the seventh character corresponding to the k-th round question, and the eighth character corresponding to the predicted answer to the k-th round question, based on the fifth character corresponding to the question in each round of the first k-1 rounds of dialogue, the sixth character corresponding to the predicted answer to the question in each round, and the question in the k-th round, where k is a natural number. Wherein, the fifth character is the summary generated by the preset model for the question in the dialogue when predicting the answer in the corresponding round of dialogue; the sixth character is the summary generated by the preset model for the predicted answer in the dialogue when predicting the answer in the corresponding round of dialogue; the seventh character is used to indicate the summary generated by the preset model for the question in the kth round; and the eighth character is used to indicate the summary generated by the preset model for the predicted answer to the question in the kth round.
[0092] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0093] In one embodiment, the internal structure diagram of the electronic device can be as follows: Figure 11 As shown, the electronic device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a data processing method or a question-and-answer method. The display unit of the electronic device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the electronic device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the electronic device, or external keyboards, touchpads, or mice, etc.
[0094] Those skilled in the art will understand that Figure 11The structure shown is only a block diagram of a part of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0095] Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0096] Since the computer program stored in the computer-readable storage medium can execute any of the data processing methods or question-and-answer methods provided in the embodiments of this application, it can achieve the beneficial effects that any of the data processing methods or question-and-answer methods provided in the embodiments of this application can achieve. For details, please refer to the previous embodiments, which will not be repeated here.
[0097] Based on the same inventive concept, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the above embodiments.
[0098] It should be noted that the object data (including but not limited to user device information, user personal information, etc.) and dialogue data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of relevant countries and regions. Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods.
[0099] Any reference to memory, database, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0100] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0101] In the above embodiments of the data processing apparatus, question-and-answer apparatus, computer-readable storage medium, electronic device, and computer program product, the descriptions of each embodiment have different focuses. Parts not described in detail in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes and beneficial effects of the data processing apparatus, question-and-answer apparatus, computer-readable storage medium, computer program product, electronic device, and their corresponding units described above can be referred to the descriptions of the data processing methods and question-and-answer methods in the above embodiments, and will not be repeated here.
[0102] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0103] The foregoing has provided a detailed description of a data processing method, question-and-answer method, data processing apparatus, question-and-answer apparatus, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A data processing method, characterized by, Includes the following steps: Set a first character for the reference answer to the i-th round question, where the first character is defined as a summary of the reference answer, and i is a natural number; Using a preset model, based on the previous i-1 rounds of dialogue, the i-th round question, the reference answer to the i-th round question, and the first character, the predicted answer to the i-th round question and the second character corresponding to the predicted answer are predicted. The second character is used to indicate the summary generated by the preset model for the predicted answer to the i-th round question. Based on the reference answer and the predicted answer of the i-th round question, determine the first loss; Based on the second character and the reference answer to the i-th round question, determine the second loss; Based on the first loss and the second loss, the parameters of the preset model are updated.
2. The data processing method according to claim 1, characterized in that, The determination of the second loss based on the second character and the reference answer to the i-th round question includes: The second character is decompressed to obtain the first text; Based on the reference answer to the i-th round question and the first text, the second loss is determined.
3. The data processing method of claim 1, wherein, The method further includes: A third character is set for the i-th round question, and the third character is defined as a summary of the i-th round question; The step of predicting the predicted answer to the i-th round question and the corresponding second character based on the first i-1 rounds of dialogue, the i-th round question, the reference answer to the i-th round question, and the first character using a preset model includes: Using the preset model, based on the first i-1 rounds of dialogue, the i-th round question, the third character, the reference answer to the i-th round question, and the first character, the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question are predicted. The fourth character is used to indicate the summary generated by the preset model for the i-th round question. After obtaining the fourth character corresponding to the i-th round question, the method further includes: Based on the fourth character and the i-th round question, determine the third loss; The step of updating the parameters of the preset model based on the first loss and the second loss includes: The parameters of the preset model are updated based on the first loss, the second loss, and the third loss.
4. The data processing method according to claim 3, characterized in that, The step of predicting the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question based on the first i-1 rounds of dialogue, the i-th round question, the third character, the reference answer to the i-th round question, and the first character using a preset model includes: By concatenating the first i-1 rounds of dialogue, the i-th round question, and the reference answer to the i-th round question, a second text is obtained; The second text is segmented to obtain a word sequence; The words in the word sequence are encoded to obtain a first vector; Using the preset model, based on the first vector, the second vector set for the first character, and the third vector set for the third character, the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question are predicted.
5. The data processing method according to claim 4, characterized in that, The step of predicting the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question based on the preset model, using the first vector, the second vector set for the first character, and the third vector set for the third character, includes: The second vector is inserted after the vector corresponding to the reference answer to the i-th round question in the first vector to obtain the fourth vector; The third vector is inserted into the fourth vector after the vector corresponding to the i-th round question and before the vector corresponding to the reference answer of the i-th round question to obtain the fifth vector; Based on the preset model and the fifth vector, the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question are predicted.
6. The data processing method according to claim 5, characterized in that, The step of predicting the predicted answer to the i-th round question, the second character corresponding to the predicted answer, and the fourth character corresponding to the i-th round question based on the fifth vector using the preset model includes: Based on the preset model and the fifth vector, the predicted answer to the i-th round question and the fourth character corresponding to the i-th round question are predicted. Based on the predicted answer to the i-th round question, the second character corresponding to the predicted answer is predicted using the preset model.
7. A question and answer method characterized by, The method includes: Using a pre-defined model, based on the fifth character corresponding to each question in the first k-1 rounds of dialogue, the sixth character corresponding to the predicted answer to each question in the first k-1 rounds of dialogue, and the question in the kth round, the predicted answer to the question in the kth round, the seventh character corresponding to the question in the kth round, and the eighth character corresponding to the predicted answer to the question in the kth round are predicted, where k is a natural number; The fifth character is a summary generated by the preset model for the question in the dialogue when predicting the answer in the corresponding round of dialogue; The sixth character is a summary generated by the preset model for the predicted answer in the corresponding round of dialogue. The seventh character is used to indicate the summary generated by the preset model for the k-th round problem; The eighth character is used to indicate the summary generated by the preset model for the predicted answer to the question in the k-th round.
8. An electronic device, comprising: It includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps of the data processing method as described in any one of claims 1 to 6, or the question-and-answer method as described in claim 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the data processing method as described in any one of claims 1 to 6, or the question-and-answer method as described in claim 7.
10. A computer program product, characterised in that, The computer program product includes a computer program or instructions, which are executed by a processor using the steps of the data processing method as described in any one of claims 1 to 6, or the question-and-answer method as described in claim 7.
Citation Information
Patent Citations
Question answer generation method and device based on large language model, equipment and medium
CN118093830A
Business database dialogue method and device, electronic equipment and medium
CN119378676A