Method and system for improving accuracy of reply knowledge generated by dialogue system

By adopting search-enhanced teacher model training, knowledge injection and distillation learning in the dialogue system, as well as comparative learning optimization, the problems of knowledge accuracy and reasoning efficiency of the dialogue system generated replies are solved, and efficient generation results are achieved.

CN120296125APending Publication Date: 2025-07-11SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510358775.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When generating replies, there is a trade-off between knowledge accuracy and reasoning efficiency in existing dialogue systems, which cannot simultaneously improve the knowledge accuracy and reasoning efficiency of generating replies.

Method used

Retrieval-enhanced teacher model is used for training, knowledge-intensive dialogue responses are generated through maximum likelihood estimation, and knowledge injection and sentence particle size distillation learning are performed on the student model, and the generation quality of the student model is optimized in combination with comparative learning.

Benefits of technology

Without relying on external knowledge retrieval, the knowledge accuracy and reasoning efficiency of the dialogue system generating replies is improved, and efficient generation effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296125A_ABST
    Figure CN120296125A_ABST
Patent Text Reader

Abstract

The invention provides a method and system for improving the accuracy of reply knowledge generated by a dialogue system, and the method comprises the steps: a teacher model employs a retrieval enhancement generation model, carries out training through maximum likelihood estimation, and generates a dialogue response with dense knowledge; knowledge injection is carried out on the student model, and related knowledge is injected into model parameters through external knowledge such as FAQ in training data; the student model learns from various knowledge-rich responses generated by the teacher model through sentence granularity distillation learning so as to improve the knowledge accuracy of reply generation; and through comparative learning optimization, the student model selects a better reply from multiple labels generated by the teacher model for learning, so that the generation quality is further improved. According to the method, through distillation learning of the teacher and student models, the knowledge-intensive dialogue labels generated by the teacher model are transmitted to the student models, so that replies can be generated with high accuracy and high reasoning efficiency under the condition of no retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing. Specifically, it relates to a method and system for improving the knowledge accuracy of generated responses in a dialogue system. In particular, it relates to a method for improving the knowledge accuracy of generated responses in a dialogue system based on knowledge distillation and multi-label contrast learning. Background Art

[0002] Existing dialogue systems have improved the knowledge accuracy of dialogue generation through the Retrieval-Augmented Generation (RAG) method, but their inference efficiency is low; while the Retrieval-Free Generation (RFG) method has a faster inference speed, but the generated responses perform poorly in terms of knowledge accuracy and are prone to the "hallucination" phenomenon. Existing solutions involve a trade-off between efficiency and the knowledge accuracy of generated responses. Therefore, how to improve the inference efficiency while having a high knowledge accuracy in the generated responses has become an urgent problem to be solved.

[0003] After retrieval, the patent application document CN117669693A, application number CN202311422140.3, discloses a knowledge distillation method and system based on a multi-teacher multi-modal model, belonging to the field of natural language processing. Through multiple teacher models jointly performing multi-modal knowledge distillation into a student model, these teacher models have different architectures, initializations, training data, or tasks, and this diversity helps to extract different perspectives and types of knowledge, thereby improving the robustness of the student model and its understanding ability of images, texts, and image-text multi-modalities, and enhancing the accuracy of image recognition, the accuracy of text understanding, and the recall rate and accuracy of multi-modal retrieval. However, it does not consider how to apply these methods to generation tasks, which will result in the inability of this method to improve the knowledge accuracy of dialogue generation in a dialogue system. Summary of the Invention

[0004] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a method and system for improving the knowledge accuracy of generated responses in a dialogue system.

[0005] According to the method for improving the knowledge accuracy of generated responses in a dialogue system provided by the present invention, it includes:

[0006] Step S1: Adopt a retrieval-augmented teacher model and train it through maximum likelihood estimation to generate knowledge-intensive dialogue responses;

[0007] Step S2: Inject knowledge into the student model, and inject relevant knowledge into the parameters of the student model through external knowledge in the training data;

[0008] Step S3: The student model learns through sentence-level distillation learning from multiple responses generated by the teacher model to improve the knowledge accuracy of the generated responses;

[0009] Step S4: Through comparative learning optimization, the student model selects better responses from the multi-label generated by the teacher model for learning to further improve its generation quality.

[0010] Preferably, step S1 includes: concatenating the writing dialogue history text and the external knowledge required for the current user query and inputting them into the language model, thereby training a dialogue model based on retrieval enhancement, the formula of which is:

[0011]

[0012] in, is the loss function of maximum likelihood estimation, U t represents the historical conversation text, K represents the external knowledge required by the current user’s query, |u t+1 | indicates the length of the next user input, w i is the teacher model learning object u t+1 The i-th word of , θ represents the teacher model with θ as parameter; p θ (w i ∣w <i ,U t ,K) represents the given historical dialogue text U t , the external knowledge K required by the current user query, and all previous word units w <i Under the condition of i probability;

[0013] After training with the maximum likelihood estimation loss function, the teacher model for generating knowledge-intensive responses is obtained, and its generation is expressed by the following formula:

[0014]

[0015] in, is the response predicted by the teacher model, is the i-th token in the response.

[0016] Preferably, step S2 comprises:

[0017] For external knowledge K, split it into the problem part K of knowledge Q ={q1,…,q i} and the answer part K A ={a1,…,a j} and then inject it into the student model parameters, using the maximum likelihood estimation loss function to inject knowledge into the student model, the formula is:

[0018]

[0019] Among them, represents the loss function for knowledge injection, φ are the parameters of the student model, and j is the number of answers; p φ (a t ∣a <t ,K Q ) is the probability of generating the t-th answer a <t given all the previous answers a Q and the question part K t .

[0020] Preferably, the step S3 includes:

[0021] To further improve the knowledge accuracy of the system response generated by the dialogue model based on parameterized knowledge, the predicted response generated by the teacher model is used as the training label, and knowledge distillation training is performed on the student model with φ as the model parameters. The formula for training using negative log-likelihood loss is:

[0022]

[0023] Among them, represents the negative log-likelihood loss function, is the length of the response predicted by the teacher model; represents the probability of generating the i-th token given all the previous tokens t and the historical dialogue text U .

[0024] Preferably, the step S4 includes:

[0025] To further improve the knowledge accuracy of the response generated by the dialogue system based on parametric knowledge, multiple predicted responses generated by the teacher model are used as the training supervision signals for knowledge distillation. The multiple predicted responses generated by the teacher model are written as: These M predicted responses have been sorted from good to bad according to the fluency and knowledge accuracy of the responses, so as to calculate the log-likelihood of the student model predicting these M labels. The formula is:

[0026]

[0027] Among them, l i (φ) represents the log-likelihood loss of the i-th predicted response, L i is the length of the i-th predicted response, is the length of the i-th predicted response, is the i-th predicted response generated by the teacher model, is the j-th token in this response, is the sequence of all previous tokens, U t is the historical dialogue text; is given all the previous tokens and the historical dialogue text U t the probability of generating the j-th token ;

[0028] To further improve the knowledge accuracy of the responses generated by the student dialogue model based on parametric knowledge, contrastive learning is used to encourage the student model to learn from the objects with better fluency and knowledge accuracy among these M labels. The loss function of this contrastive learning is:

[0029]

[0030] max{0, ρ - (l m (φ) - l n (φ))}

[0031] where represents the loss function of contrastive learning, M is the number of predicted responses generated by the teacher model, l m (φ), l n (φ) are the log-likelihood losses of the m-th and n-th predicted responses respectively, φ is the parameter of the student model, and ρ is a predefined boundary. Thus, the overall training loss is:

[0032]

[0033] where α is a hyperparameter used to balance the weights between the negative log-likelihood loss and the contrastive learning loss.

[0034] According to the system for improving the knowledge accuracy of responses generated by a dialogue system provided by the present invention, it includes:

[0035] Module M1: Adopt a retrieval-enhanced teacher model, which is trained by maximum likelihood estimation to generate knowledge-intensive dialogue responses;

[0036] Module M2: Inject knowledge into the student model, and inject relevant knowledge into the parameters of the student model through external knowledge in the training data;

[0037] Module M3: The student model learns through sentence-level distillation from multiple responses generated by the teacher model to improve the knowledge accuracy of the generated responses;

[0038] Module M4: Through contrastive learning optimization, the student model selects better responses from the multiple labels generated by the teacher model for learning to further improve its generation quality.

[0039] Preferably, the module M1 includes: concatenating the writing dialogue history text and the external knowledge required for the current user's question and inputting them into a language model to train a retrieval-enhanced dialogue model, and its formula is:

[0040]

[0041] Wherein, is the loss function of maximum likelihood estimation, U t represents the historical dialogue text, K represents the external knowledge required for the current user's question, |u t+1 | represents the length of the next user input, w i is the i-th token of the teacher model learning object u t+1 , θ represents the teacher model with θ as the parameter; p θ (w i ∣w <i , U t , K) represents the probability of generating the i-th token w t under the conditions of the given historical dialogue text U <i , the external knowledge K required for the current user's question, and all the previous tokens w i ;

[0042] After being trained by the maximum likelihood estimation loss function, a teacher model for generating knowledge-intensive responses is obtained, and its generation is represented by the following formula:

[0043]

[0044] Wherein, is the response predicted by the teacher model, is the i-th token in this response.

[0045] Preferably, the module M2 includes:

[0046] For the external knowledge K, the process of splitting it into the question part K Q ={q1,...,q i} and the answer part K A ={a1,...,a j} and then injecting it into the parameters of the student model, using the maximum likelihood estimation loss function to perform knowledge injection on the student model, and its formula is:

[0047]

[0048] Wherein, represents the loss function of knowledge injection, φ is the parameter of the student model, j is the number of answers; p φ (a t ∣a <t,K q ) is the probability of generating the t-th answer a <t under the conditions of all the previous answers a Q and the question part K t .

[0049] Preferably, the module M3 includes:

[0050] To further improve the knowledge accuracy of the system responses generated by the dialogue model based on parameterized knowledge, the predicted responses generated by the teacher model are used as training labels, and knowledge distillation training is performed on the student model with φ as the model parameters. The formula for training using the negative log-likelihood loss is:

[0051]

[0052] where represents the negative log-likelihood loss function, is the length of the response predicted by the teacher model; represents the probability of generating the i-th token under the conditions of all the previous tokens t and the historical dialogue text U .

[0053] Preferably, the module M4 includes:

[0054] To further improve the knowledge accuracy of the responses generated by the dialogue system based on parameter knowledge, multiple predicted responses generated by the teacher model are used as training supervision signals for knowledge distillation. The multiple predicted responses generated by the teacher model are written as: These M predicted responses have been sorted from good to bad according to the fluency and knowledge accuracy of the responses, so as to calculate the log-likelihood of the student model predicting these M labels. The formula is:

[0055]

[0056] where l i (φ) represents the log-likelihood loss of the i-th predicted response, L i is the length of the i-th predicted response, is the length of the i-th predicted response, is the i-th predicted response generated by the teacher model, is the j-th token in this response, is the sequence of all the previous tokens, U t is the historical dialogue text; is the probability of generating the i-th token under the conditions of all the previous tokens tUnder the condition of, generate the j-th token probability;

[0057] To further improve the knowledge accuracy of the generated responses of the student dialogue model based on parametric knowledge, contrastive learning is used to encourage the student model to learn the objects with better fluency and knowledge accuracy among these M labels. The loss function of this contrastive learning is as follows:

[0058]

[0059] max{0, ρ - (l m (φ) - l n (φ))}

[0060] Among them, represents the loss function of contrastive learning, M is the number of predicted responses generated by the teacher model, l m (φ), l n (φ) are the logarithmic likelihood losses of the m-th and n-th predicted responses respectively, φ is the parameter of the student model, and ρ is a predefined boundary. Thus, the overall training loss is:

[0061]

[0062] Among them, α is a hyperparameter used to balance the weights between the negative log-likelihood loss and the contrastive learning loss.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] The present invention proposes a method for improving the knowledge accuracy of the generated responses of a dialogue system based on knowledge distillation and multi-label contrastive learning. Through the distillation learning of the teacher-student model, the knowledge-intensive dialogue labels generated by the teacher model are transmitted to the student model for training, enabling the student model to achieve high knowledge accuracy and high reasoning efficiency without retrieving external knowledge. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objectives, and advantages of the present invention will become more apparent:

[0066] Figure 1 is a flowchart of a method for improving the knowledge accuracy of the generated responses of a dialogue system based on knowledge distillation and multi-label contrastive learning provided by the present invention;

[0067] Figure 2 is a schematic structural diagram of a system for improving the knowledge accuracy of the generated responses of a dialogue system based on knowledge distillation and multi-label contrastive learning provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0068] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.

[0069] Embodiment 1

[0070] The present invention provides an external knowledge extraction method in a spoken dialogue scenario with contrastive learning and attention mechanism, as Figure 1 shown, including:

[0071] Step 1: Train a teacher model enhanced by retrieval, and train it through maximum likelihood estimation to generate knowledge-intensive dialogue responses;

[0072] Step 2: Inject knowledge into the student model, and inject relevant knowledge into the model parameters through external knowledge such as FAQs in the training data;

[0073] Step 3: The student model learns through sentence-level distillation learning from a variety of knowledge-rich responses generated by the teacher model to improve the knowledge accuracy of the generated responses;

[0074] Step 4: Optimize through contrastive learning, and let the student model select a better response from the multi-labels generated by the teacher model for learning to further improve its generation quality.

[0075] Specifically, first concatenate the written dialogue history text and the external knowledge required for the current user's query and input them into the language model, so as to train a dialogue model enhanced by retrieval, and its formula can be written as:

[0076]

[0077] where U t represents the written dialogue history text, K represents the external knowledge required for the current user's query, and u t+1 represents the accurate label of the response generated by the dialogue system. w i is the i-th token of the teacher model's learning object u t+1 . θ represents the teacher model with θ as the parameter. After training with the maximum likelihood estimation loss function, a teacher model for generating knowledge-intensive responses can be obtained, and its generation can be expressed by the following formula. Among them is the response predicted by the teacher model, is the i-th token in this response.

[0078]

[0079] Next, it is necessary to first inject external knowledge into the target student model. This external knowledge can be knowledge in a specific domain or general knowledge similar to Wikipedia. For external knowledge K, it is split into the question part K Q ={q1,…,q i} and the answer part K A ={a1,…,a j} of the knowledge, and then the process of injecting it into the student model parameters. Specifically, the maximum likelihood estimation loss function can be used to inject knowledge into the student model, and its formula can be written as:

[0080]

[0081] where φ represents the student model with φ as the model parameters.

[0082] To further improve the knowledge accuracy of the system responses generated by the student dialogue model based on parameterized knowledge, the predicted responses generated by the teacher model are used as training labels to perform knowledge distillation training on the student model with φ as the model parameters. The formula for training using the negative log-likelihood loss can be written as:

[0083]

[0084] To further improve the knowledge accuracy of the responses generated by the dialogue system based on parameter knowledge, multiple predicted responses generated by the teacher model are used as training supervision signals for knowledge distillation. The multiple predicted responses generated by the teacher model can be written as: These M predicted responses have been sorted from good to bad according to the fluency and knowledge accuracy of the responses, so that the log-likelihood of the student model predicting these M labels can be calculated, and its formula can be written as:

[0085]

[0086] To further improve the knowledge accuracy of the responses generated by the student dialogue model based on parameter knowledge, this method uses contrastive learning to encourage the student model to learn the objects with better fluency and knowledge accuracy among these M labels. The loss function of this contrastive learning method can be written as:

[0087]

[0088] max{0,ρ-(l m (φ)-l n (φ))}

[0089] where ρ is a predefined boundary. Thus, the overall training loss of this method can be written as:

[0090]

[0091] Among them, α is a hyperparameter used to balance the weights between the negative log-likelihood loss and the contrastive learning loss.

[0092] Example 2

[0093] The present invention also provides a system for improving the knowledge accuracy of the generated responses of a dialogue system. The system for improving the knowledge accuracy of the generated responses of a dialogue system can be implemented by executing the process steps of the method for improving the knowledge accuracy of the generated responses of a dialogue system. That is, those skilled in the art can understand the method for improving the knowledge accuracy of the generated responses of a dialogue system as a preferred implementation manner of the system for improving the knowledge accuracy of the generated responses of a dialogue system.

[0094] Such as Figure 2 , according to the system for improving the knowledge accuracy of the generated responses of a dialogue system provided by the present invention, it includes: Module M1: adopting a retrieval-enhanced teacher model, training through maximum likelihood estimation, and generating knowledge-intensive dialogue responses; Module M2: injecting knowledge into the student model, and injecting relevant knowledge into the parameters of the student model through the external knowledge in the training data; Module M3: the student model learns through sentence-level distillation learning from multiple responses generated by the teacher model to improve the knowledge accuracy of the generated responses; Module M4: through contrastive learning optimization, enabling the student model to select a better response from the multi-labels generated by the teacher model for learning to further improve its generation quality.

[0095] The Module M1 includes: concatenating the written dialogue history text and the external knowledge required for the current user's query and inputting them into the language model, thereby training a retrieval-enhanced dialogue model, and its formula is:

[0096]

[0097] Among them, is the loss function of maximum likelihood estimation, U t represents the historical dialogue text, K represents the external knowledge required for the current user's query, |u t+1 | represents the length of the next user input, w i is the i-th token of the teacher model's learning object u t+1 , θ represents the teacher model with θ as the parameter; p θ (w i ∣w <i ,U t ,K) represents that given the historical dialogue text U t , the external knowledge K required for the current user's query, and all the previous tokens w <iUnder the condition of, generate the i-th token w i The probability of;

[0098] After being trained by the maximum likelihood estimation loss function, a teacher model for generating knowledge-intensive responses is obtained, and its generation is represented by the following formula:

[0099]

[0100] Among them, Is the response predicted by the teacher model, Is the i-th token in this response.

[0101] The module M2 includes: For the external knowledge K, splitting it into the question part K Q ={q1,…,q i} and the answer part K A ={a1,…,a j} and then injecting it into the parameters of the student model. The maximum likelihood estimation loss function is used to inject knowledge into the student model, and its formula is:

[0102]

[0103] Among them, Represents the loss function of knowledge injection, φ is the parameter of the student model, and j is the number of answers; p φ (a t ∣a <t ,K Q ) is the probability of generating the t-th answer a <t given all the previous answers a Q and the question part K t .

[0104] The module M3 includes: To further improve the knowledge accuracy of the system response generated by the dialogue model based on parameterized knowledge, using the predicted response generated by the teacher model as the training label, and performing knowledge distillation training on the student model with φ as the model parameter. The formula for training using the negative log-likelihood loss is:

[0105]

[0106] Among them, Represents the negative log-likelihood loss function, is the length of the response predicted by the teacher model; Represents the probability of generating the i-th token given all the previous tokens t and the historical dialogue text U .

[0107] The module M4 includes: To further improve the knowledge accuracy of the generated responses of the dialogue system based on parametric knowledge, multiple predicted responses generated by the teacher model are used as the training supervision signals for knowledge distillation. The multiple predicted responses generated by the teacher model are written as: The M predicted responses have been sorted from good to bad according to the fluency and knowledge accuracy of the responses, so as to calculate the log-likelihood of the student model predicting these M labels. The formula is:

[0108]

[0109] where l i (φ) represents the log-likelihood loss of the i-th predicted response, and L i is the length of the i-th predicted response, is the length of the i-th predicted response, is the i-th predicted response generated by the teacher model, is the j-th token in this response, is the sequence of all previous tokens, U t is the historical dialogue text; is given all the previous tokens and the historical dialogue text U t under the condition of generating the j-th token probability;

[0110] To further improve the knowledge accuracy of the generated responses of the student dialogue model based on parametric knowledge, contrastive learning is used to encourage the student model to learn the objects with better fluency and knowledge accuracy among these M labels. The loss function of this contrastive learning is:

[0111]

[0112] max{0, ρ - (l m (φ) - l n (φ))}

[0113] where, represents the loss function of contrastive learning, M is the number of predicted responses generated by the teacher model, l m (φ), l n (φ) are the log-likelihood losses of the m-th and n-th predicted responses respectively, φ is the parameter of the student model, and ρ is a predefined boundary. Thus, the overall training loss is:

[0114]

[0115] where α is a hyperparameter used to balance the weights between the negative log-likelihood loss and the contrastive learning loss.

[0116] Those skilled in the art know that in addition to implementing the systems, devices and their respective modules provided by the present invention in the form of pure computer-readable program codes, it is entirely possible to make the systems, devices and their respective modules provided by the present invention be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, the systems, devices and their respective modules provided by the present invention can be considered as a kind of hardware components, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware components; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware components.

[0117] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A method for improving the knowledge accuracy of responses generated by a dialogue system, characterized in that, Including: Step S1: Adopt a retrieval-enhanced teacher model and train it through maximum likelihood estimation to generate knowledge-intensive dialogue responses; Step S2: Inject knowledge into the student model, and inject relevant knowledge into the parameters of the student model through the external knowledge in the training data; Step S3: The student model learns through sentence-level distillation learning from multiple responses generated by the teacher model to improve the knowledge accuracy of the generated responses; Step S4: Through contrastive learning optimization, let the student model select better responses from the multi-labels generated by the teacher model for learning to further improve its generation quality.

2. The method for improving the accuracy of knowledge in generating responses of a dialogue system according to claim 1, wherein, The said Step S1 includes: Concatenate the written dialogue history text and the external knowledge required for the current user's query and input them into the language model, thereby training a retrieval-enhanced dialogue model, and its formula is: Among them, is the loss function of maximum likelihood estimation, U t represents the historical dialogue text, K represents the external knowledge required by the current user's query, |u t+1 | represents the length of the next user input, w i is the i-th token of the teacher model learning object u t+1 , θ represents the teacher model with θ as the parameter; p θ (w i ∣w <i ,U t ,K) represents the probability of generating the i-th token w t under the conditions of the given historical dialogue text U <i , the external knowledge K required by the current user's query, and all the previous tokens w i ; After being trained by the maximum likelihood estimation loss function, a teacher model for generating knowledge-intensive responses is obtained, and its generation is represented by the following formula: Among them, is the response predicted by the teacher model, is the i-th token in this response.

3. The method for improving the accuracy of response knowledge generated by a dialogue system according to claim 2, wherein The said Step S2 includes: For external knowledge K, the process of splitting it into the problem part K Q ={q1,…,q i} and the answer part K A ={a1,…,a j} and then injecting it into the student model parameters. The student model is injected with knowledge using the maximum likelihood estimation loss function, and its formula is: Among them, represents the loss function of knowledge injection, φ is the parameter of the student model, and j is the number of answers; p φ (a t ∣a <t ,K Q 0 is the probability of generating the t-th answer a <t given all the previous answers a Q and the question part K t .

4. The method for improving the accuracy of knowledge for generating responses in a dialogue system according to claim 3, wherein The said Step S3 includes: To further improve the knowledge accuracy of the system responses generated by the dialogue model based on parametric knowledge, the predicted responses generated by the teacher model are used as training labels to perform knowledge distillation training on the student model with φ as the model parameter. The formula for training using the negative log-likelihood loss is as follows: Among them, represents the negative log-likelihood loss function, is the reply length predicted by the teacher model; represents the probability of generating the i-th token given all the previous tokens t and the historical dialogue text U of.

5. The method for improving the accuracy of the knowledge of the generated response of the dialogue system according to claim 4, characterized in that, The said Step S4 includes: To further improve the knowledge accuracy of the responses generated by the dialogue system based on parametric knowledge, multiple predicted responses generated by the teacher model are used as the training supervision signals for knowledge distillation. The multiple predicted responses generated by the teacher model are written as: The M predicted responses have been sorted from good to bad according to the fluency and knowledge accuracy of the responses, so as to calculate the log-likelihood of the student model predicting these M labels. The formula is: where l i (φ) represents the log-likelihood loss of the i-th predicted response, L i is the length of the i-th predicted response, is the length of the i-th predicted response, is the i-th predicted response generated by the teacher model, is the j-th token in this response, is the sequence of all previous tokens, U t is the historical dialogue text; is given all previous tokens and the historical dialogue text U t under the condition of generating the j-th token probability; To further improve the knowledge accuracy of the responses generated by the student dialogue model based on parameter knowledge, contrastive learning is used to encourage the student model to learn the object with better fluency and knowledge accuracy among these M labels, and the loss function of this contrastive learning is: Among them, represents the loss function of contrastive learning, M is the number of predicted responses generated by the teacher model, and l m (φ), l n (φ) are the log-likelihood losses of the m-th and n-th predicted responses respectively, φ is the parameter of the student model, ρ is a predefined boundary, so the overall training loss is: Among them, α is a hyperparameter used to balance the weights between the negative log-likelihood loss and the contrastive learning loss.

6. A system for improving the knowledge accuracy of generated responses in a dialogue system, characterized in that, Including: Module M1: Adopt a retrieval-enhanced teacher model and train it through maximum likelihood estimation to generate knowledge-intensive dialogue responses; Module M2: Inject knowledge into the student model, and inject relevant knowledge into the parameters of the student model through the external knowledge in the training data; Module M3: The student model learns through sentence-level distillation learning from multiple responses generated by the teacher model to improve the knowledge accuracy of the generated responses; Module M4: Through contrastive learning optimization, let the student model select better responses from the multi-labels generated by the teacher model for learning to further improve its generation quality.

7. The system for improving the knowledge accuracy of the generated responses of the dialogue system according to claim 6, characterized in that, The said Module M1 includes: Concatenate the written dialogue history text and the external knowledge required for the current user's query and input them into the language model, thereby training a retrieval-enhanced dialogue model, and its formula is: Among them, is the loss function of maximum likelihood estimation, U t represents the historical dialogue text, K represents the external knowledge required for the current user's question, |u t+1 | represents the length of the next user input, w i is the i-th token of the teacher model learning object u t+1 and θ represents the teacher model with θ as the parameter; p θ (w i ∣w <i ,U t ,K) represents the probability of generating the i-th token w t under the condition of the given historical dialogue text U <i , the external knowledge K required for the current user's question, and all the previous tokens w i ; After being trained by the maximum likelihood estimation loss function, a teacher model for generating knowledge-intensive responses is obtained, and its generation is represented by the following formula: Among them, is the response predicted by the teacher model, is the i-th token in this response.

8. The system for improving the knowledge accuracy of the generated responses of the dialogue system according to claim 7, wherein, The said Module M2 includes: For external knowledge K, the process of splitting it into the problem part K Q ={q1,…,q i} and the answer part K A ={a1,…,a j} and then injecting it into the student model parameters. The student model is injected with knowledge using the maximum likelihood estimation loss function, and its formula is: Among them, represents the loss function of knowledge injection, φ is the parameter of the student model, and j is the number of answers; p φ (a t ∣a <t ,K Q ) is the probability of generating the t-th answer a <t given all the previous answers a Q and the question part K t .

9. The system for improving the accuracy of response knowledge generated by a dialogue system according to claim 8, characterized in that, The said Module M3 includes: To further improve the knowledge accuracy of the system responses generated by the dialogue model based on parametric knowledge, the predicted responses generated by the teacher model are used as training labels to perform knowledge distillation training on the student model with φ as the model parameter. The formula for training using the negative log-likelihood loss is as follows: Among them, represents the negative log-likelihood loss function, is the reply length predicted by the teacher model; represents the probability of generating the i-th token and the historical dialogue text U t under the condition of all previous tokens is the probability.

10. The system for improving the accuracy of reply knowledge generated by a dialogue system according to claim 9, wherein The said Module M4 includes: To further improve the knowledge accuracy of the generated responses of the dialogue system based on parametric knowledge, multiple predicted responses generated by the teacher model are used as the training supervision signals for knowledge distillation. The multiple predicted responses generated by the teacher model are written as: The M predicted responses have been sorted from good to bad according to the fluency and knowledge accuracy of the responses, so as to calculate the log-likelihood of the student model predicting these M labels. The formula is: where, l i (φ) represents the log-likelihood loss of the i-th predicted response, L i is the length of the i-th predicted response, is the length of the i-th predicted response, is the i-th predicted response generated by the teacher model, is the j-th token in this response, is the sequence of all previous tokens, U t is the historical dialogue text; is given all previous tokens and the historical dialogue text U t under the condition, the probability of generating the j-th token ; To further improve the knowledge accuracy of the responses generated by the student dialogue model based on parameter knowledge, contrastive learning is used to encourage the student model to learn the object with better fluency and knowledge accuracy among these M labels, and the loss function of this contrastive learning is: Among them, represents the loss function of contrastive learning, M is the number of predicted responses generated by the teacher model, and l m (φ), l n (φ) are the log-likelihood losses of the m-th and n-th predicted responses respectively, φ is the parameter of the student model, and ρ is a predefined boundary, so the overall training loss is: Among them, α is a hyperparameter used to balance the weights between the negative log-likelihood loss and the contrastive learning loss.

Citation Information

Patent Citations

  • Knowledge distillation method and system based on multi-teacher multi-modal model

    CN117669693A

Cited By

  • Intelligent customer service base large model distillation method and system and storage medium

    CN122065914A