Dialogue generation method based on exogenous knowledge generation potential sorting enhancement

Through the enhanced method of generating potential of exogenous knowledge, the quality of dialogue responses caused by search and generation differences is solved, and more efficient and accurate dialogue generation results are achieved.

CN120179780APending Publication Date: 2025-06-20UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510250971.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing dialogue generation method ignores the differences between retrieval and generation in multiple rounds of dialogue scenarios, resulting in the knowledge sorting results of the searcher recall being detrimental to the generation task and affecting the quality of the reply.

Method used

A dialogue generation method based on the potential sorting enhancement of exogenous knowledge generation potential is adopted. By generating a latent force function, the potential of different exogenous knowledge fragments for generating reply tags is evaluated, the data-enhanced positive and negative sample pairs are constructed, and the searcher and generator are fine-tuned to improve retrieval accuracy and generation quality.

Benefits of technology

Effectively bridge the differences between retrieval and generation, improve the quality and generalization capabilities of dialogue responses, and provide more accurate, reliable and efficient dialogue generation technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179780A_ABST
    Figure CN120179780A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language models, and discloses a dialogue generation method based on exogenous knowledge generation potential sorting enhancement, which comprises the following steps: quantifying generation potentials of different knowledge fragments, and performing correction and normalization; constructing positive and negative sample pairs to perform data enhancement; a retriever is finely adjusted on the data enhanced training set; splicing a dialogue history and a previous set knowledge fragment with the highest correlation after a searcher sorts, and finely adjusting a generator; and the generator performs reasoning on the test set and returns replies generated according to the dialogue history and the recalled knowledge fragments. At present, a mainstream retriever more pays more attention to semantic correlation between dialogue history and exogenous knowledge, but does not have the ability of generating reply labels by combining the dialogue history with the exogenous knowledge. According to the method, the difference is considered in the searcher training, so that the knowledge set recalled by the searcher can better promote the generation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and particularly relates to a dialogue generation method based on enhanced potential ranking of exogenous knowledge generation. Background Art

[0002] Building a dialogue system that can communicate with humans naturally and coherently has always been a long-term goal of natural language processing tasks. In the information age, intelligent chatbots have broad application scenarios in various fields. For example, in the customer service field, they can be used to automatically answer common questions, provide product information, handle complaints, etc.; in the healthcare field, they can be used for case recording, health consultation, medical diagnosis assistance, etc.; in the tourism and catering fields, they can provide travel suggestions, book flights and hotels, recommend restaurants and cuisine, etc. Generally speaking, intelligent chatbots can play an important role in improving service efficiency and enhancing user experience in multiple fields.

[0003] With the rapid development of large language models, great success has been achieved in the dialogue generation task. However, researchers have found that traditional large language models still have many deficiencies in the dialogue generation task, such as hallucinations, insufficient background knowledge, lack of informative and boring responses, inconsistent context in the generated content, etc. Knowledge retrieval enhanced generation technology has been proven effective for many knowledge-intensive tasks. Providing relevant exogenous knowledge to the model helps alleviate the hallucination problem, enhance the informativeness of responses, reduce templated and boring responses, and enhance the interpretability and generalization ability of the model.

[0004] The current mainstream research direction in multi-turn dialogue scenarios is knowledge retrieval enhanced text generation, which mainly includes two key components, a knowledge retriever and a dialogue generator. Among them, the knowledge retriever aims to select relevant knowledge from an exogenous knowledge base. The knowledge recalled by the retriever after sorting is crucial for the next step of dialogue generation and directly affects the response quality of the generator. A large number of researchers have explored and improved the knowledge retriever, mainly divided into two categories: one is to utilize the posterior information of labels in the dataset, such as using the dialogue history and response labels as the input of the retrieval model to train a knowledge retriever. This model has high retrieval accuracy and is used as a teacher model. Subsequently, the performance of the student model in the retrieval task is enhanced through knowledge distillation. This type of method focuses on the posterior information of the model encoder part and lacks a clear standard to measure the ability of knowledge to autoregressively generate response labels; the other is to train a re-ranker through methods such as attention distribution calculation, multi-task learning, and hard sample mining. This type of method has the problem of error accumulation. When the training samples contain a large amount of non-label knowledge that is easily confused with label knowledge, the retriever is difficult to distinguish, thereby damaging the performance of the re-ranking model.

[0005] In the era of the rapid development of large language models and information technology, researchers have made a great deal of efforts in developing intelligent chatbots and achieved remarkable success. However, existing chatbots cannot fully meet society's pursuit of service efficiency and quality. Facing the demand for intelligent chat systems from a large user group in various fields, dialogue generation technology that is more accurate, reliable, efficient, and provides useful information is of great significance and huge commercial value. Summary of the Invention

[0006] The key issue that this invention focuses on is that mainstream retrieval-augmented generation methods ignore the differences between retrieval and generation. The retriever pays more attention to the semantic relevance between the dialogue history and different knowledge fragments during training, rather than the promotion effect of knowledge fragments on generating reply labels. This will result in the ranking result of the knowledge retrieved by the retriever being suboptimal for the next generation task, thereby damaging the quality of the generated dialogue replies.

[0007] To solve the above technical problems, this invention provides a dialogue generation method based on enhancing the ranking of the generation potential of external knowledge.

[0008] To solve the above technical problems, this invention adopts the following technical solutions:

[0009] A dialogue generation method based on enhancing the ranking of the generation potential of external knowledge. The dialogue generation model used includes a retriever and a generator. The retriever retrieves relevant knowledge fragments according to the dialogue history, and the generator generates replies based on the dialogue history and the retrieved relevant knowledge fragments. The method specifically includes:

[0010] Obtain the dialogue history. Since the dialogue involves a specific topic, external knowledge is required to enhance generation.

[0011] Successively splice different external knowledge fragments with the obtained dialogue history, input them into the retriever fine-tuned with a data-augmented training set, output the relevance between the dialogue history and different knowledge fragments, and obtain the corresponding knowledge fragment ranking result after sorting the relevance values.

[0012] Splice the dialogue history with the top set number of knowledge fragments with the highest relevance after being sorted by the retriever, input them into the generator fine-tuned through training, and output the reply generated based on the dialogue history and the retrieved knowledge fragments.

[0013] In one embodiment, the dialogue is unstructured natural language text and involves a specific topic.

[0014] In one embodiment, the training and inference processes of the retriever and generator of the dialogue generation model include:

[0015] Construct training sample A, where each training sample A includes a dialogue history, multiple exogenous knowledge segments, a knowledge relevance label, and a response label; quantify the generation potential of the exogenous knowledge segments in training sample A to obtain training sample B;

[0016] In training sample B, splice the dialogue history onto the labeled knowledge and unlabeled knowledge respectively to construct positive and negative sample pairs; according to the generation potential of the exogenous knowledge segments in training sample B, screen the unlabeled knowledge with a generation potential higher than the set value in the exogenous knowledge segments as pseudo-labeled knowledge to construct data-augmented positive and negative sample pairs; correct and normalize the generation potential of the exogenous knowledge segments in training sample B; each positive and negative sample pair and the corresponding knowledge generation potential, dialogue history, and response label constitute training sample C;

[0017] Input the positive and negative sample pairs and the corresponding dialogue history in training sample C into the retriever, fine-tune the retriever through contrastive learning, and output the relevance between the dialogue history and different knowledge segments; for different knowledge segments, sort them according to the relevance to obtain the sorted set of exogenous knowledge segments;

[0018] Input the dialogue history, the sorted set of exogenous knowledge segments, and the response label into the BART model to perform sequence-to-sequence generation task training to obtain a fine-tuned generator, and then the generator performs inference;

[0019] Use the retriever to sort the exogenous knowledge to obtain the top set number of exogenous knowledge segments, form the context with the dialogue history of the test sample, input it into the fine-tuned generator for inference, and output the response based on the context.

[0020] In one embodiment, the quantification of the generation potential of the exogenous knowledge segments in training sample A specifically includes:

[0021] The generation potential is used to evaluate the potential of different exogenous knowledge segments for generating the response label; input the dialogue history into the generation model to obtain the probabilities of different tokens in the generated response label, multiply the probabilities of all tokens to obtain probability one P(y|x,θ); splice the dialogue history with a specific exogenous knowledge segment and obtain probability two P(y|x,k,θ) of generating the response label in the same way; divide probability two by probability one and take the logarithm to obtain the generation potential quantification result; the generation potential of all exogenous knowledge segments of the same dialogue sample is quantified through the generation potential quantification function.

[0022] In one embodiment, the generation potential quantification function specifically includes:

[0023]

[0024] Among them, \(x\) represents the dialogue history, \(y\) represents the response label, \(k\) represents an external knowledge fragment, \(\theta\) represents the parameters of the generation model, and \(Q(x, y, k, \theta)\) represents the quantification result of the generation potential of the external knowledge fragment \(k\).

[0025] In one embodiment, all external knowledge fragments of the same dialogue sample are quantified for generation potential through a generation potential quantification function, specifically including:

[0026] Decompose the generation process into multiple time steps \(t\), \(n\) represents the length of the response label \(y\), and \(y\) t represents the token generated at time step \(t\); for each time step \(t\), the conditional probability \(p(y\) t |y\) <t , x, \(\theta)\) represents the probability that the dialogue generation model generates the current token given the tokens before time step \(t\); the probability of generating the token sequence is obtained by multiplying; the probability \(p(y\) t |x, \(\theta)\) represents the probability that the dialogue history itself generates the response label, and the probability \(p(y\) t |x, k, \(\theta)\) represents the probability that the dialogue history generates the response label after combining with a specific external knowledge fragment; the logarithm of the quotient of \(P(y|x, k, \theta)\) and \(P(y|x, \theta)\) represents the logarithmic magnification of the probability of generating the response label after the dialogue history is concatenated with the specific knowledge compared to the probability before concatenation.

[0027] In one embodiment, in the training sample \(B\), the dialogue history is respectively concatenated on the labeled knowledge and unlabeled knowledge to construct positive and negative sample pairs; according to the generation potential of the external knowledge fragments in the training sample \(B\), the unlabeled knowledge with a generation potential higher than the set value in the external knowledge fragments is screened as pseudo-labeled knowledge, and data-augmented positive and negative sample pairs are constructed, specifically including the following steps:

[0028] S21: Determine whether the external knowledge fragment is labeled knowledge according to the annotation information of the external knowledge fragment. If so, concatenate the dialogue history on the labeled knowledge as the positive sample; if not, proceed to step S22:

[0029] S22: Set a threshold \(\epsilon\), calculate the product of the generation potential of the labeled knowledge and the threshold \(\epsilon\). If the generation potential of the unlabeled knowledge is greater than the product, use the current unlabeled knowledge as pseudo-labeled knowledge and concatenate the dialogue history on the pseudo-labeled knowledge as the positive sample; if the generation potential of the unlabeled knowledge is less than or equal to the product, the current unlabeled knowledge is not used as pseudo-labeled knowledge, and the dialogue history is concatenated on the unlabeled knowledge as the negative sample; the positive sample and the negative sample form data-augmented positive and negative sample pairs.

[0030] In one embodiment, the correction and normalization of the generation potential of the exogenous knowledge fragments in the training sample B specifically include: when performing normalization, the maximum-minimum normalization is adopted.

[0031] In one embodiment, the positive and negative sample pairs and the corresponding dialogue history in the training sample C are input into the retriever, and the retriever is fine-tuned through contrastive learning, and the output is the correlation between the dialogue history and different knowledge fragments; for different knowledge fragments, they are sorted according to the correlation to obtain the sorted set of exogenous knowledge fragments, which specifically includes:

[0032] The fine-tuning process of the retriever is as follows:

[0033] The retriever selects any pre-trained language model containing an encoder structure;

[0034] The loss function in the retriever fine-tuning process selects the Margin Ranking loss, which is used to ensure that the score of the positive sample is higher than that of the negative sample, and the score gap between the two is at least a given margin;

[0035] During fine-tuning, the positive and negative sample pairs, the corresponding generation potential, and the dialogue history in the training sample C are input into the retriever, and the parameters are updated through backpropagation of the loss function;

[0036] The inference process of the retriever is as follows:

[0037] For the test sample, the following operations are iteratively performed until the correlations of all exogenous knowledge are output:

[0038] The dialogue history is concatenated with a single exogenous knowledge fragment as the context information;

[0039] The context information is input into the fine-tuned retriever, and the correlation between the dialogue history and the current exogenous knowledge fragment is output;

[0040] The exogenous knowledge fragments are sorted according to the correlation values. The larger the correlation value, the higher the corresponding knowledge correlation and the higher the ranking. The sorted result of the exogenous knowledge fragments is output.

[0041] In one embodiment, the dialogue history, the sorted set of exogenous knowledge fragments, and the reply labels are input into the BART model for sequence-to-sequence generation task training to obtain a fine-tuned generator, and then the generator performs inference, which specifically includes:

[0042] The fine-tuning process specifically includes:

[0043] The generator fine-tuning process uses the cross-entropy loss function to minimize the negative log-likelihood of the generated sequence by the BART model, making the generated sequence as close as possible to the response label sequence. The dialogue history in the training samples, the top set number of external knowledge segments recalled by the retriever through sorting, and the response labels are input into the generator, and the parameters are updated through backpropagation of the cross-entropy loss function.

[0044] The inference process specifically includes:

[0045] Use the retriever to sort the external knowledge to obtain the top set number of external knowledge segments, form the context with the dialogue history of the test sample, input it into the generator for inference, and output the response based on the context.

[0046] Compared with the prior art, the beneficial technical effects of the present invention are:

[0047] The present invention defines an interpretable generation potential quantification function, enabling the gain of different external knowledge for generating response labels to be used for numerical comparison and calculation, which is the key point of the present invention.

[0048] The motivation of the present invention is to bridge the gap between retrieval and generation. Currently, the retriever pays more attention to the semantic relevance between the dialogue history and external knowledge, rather than the ability to generate response labels by combining the dialogue history and external knowledge. The present invention takes this difference into account in the retriever training, enabling the knowledge set recalled by the retriever to better promote the generation task, which is the innovation point of the present invention.

[0049] The present invention designs and implements a retriever based on enhancing the generation potential of external knowledge. By setting a threshold for screening according to the generation potential, it fully mines the labeled and unlabeled knowledge in the training samples to form positive and negative sample pairs for data augmentation, effectively improving the retrieval accuracy and generalization of the retriever, which is one of the core innovation points of the present invention.

[0050] The present invention designs and implements a general framework for knowledge retrieval-enhanced generation. In addition to knowledge retrieval-enhanced dialogue generation, it can also be migrated to various knowledge-intensive natural language tasks, such as slot filling, open-domain question answering, and other tasks that require obtaining a large amount of external knowledge. This method can be applied to tasks in different fields and styles, improving the practicality and adaptability of the method, which is one of the core innovation points of the present invention. Description of the Drawings

[0051] Figure 1 It is a flowchart of the dialogue generation method in the embodiment of the present invention.

[0052] Figure 2 It is a flowchart of the training and inference of the dialogue generation model in the embodiment of the present invention.

[0053] Figure 3It is a flowchart for constructing positive and negative samples in an embodiment of the present invention.

[0054] Figure 4 It is a schematic flowchart for fine-tuning a retriever in an embodiment of the present invention.

[0055] Figure 5 It is a schematic flowchart for the retriever to perform inference in an embodiment of the present invention.

[0056] Figure 6 It is a schematic flowchart for fine-tuning a generator in an embodiment of the present invention.

[0057] Figure 7 It is a schematic flowchart for the generator to perform inference in an embodiment of the present invention. Detailed implementation manners

[0058] A preferred implementation manner of the present invention will be described in detail below with reference to the accompanying drawings.

[0059] The present invention mainly aims at the problem that in the multi-turn dialogue scenario of the current mainstream retrieval-enhanced generation method, there are differences in retrieval and generation, and proposes a dialogue generation method based on the enhancement of the sorting of the generation potential of external knowledge. This method obtains the potential of knowledge fragments to generate reply labels through a well-explained scheme, quantifies the size of the generation potential in numerical form, and constructs paired positive and negative samples based on the generation potential to train the retriever. This method can make full use of limited labeled data for data augmentation, thereby improving the knowledge retrieval accuracy and generation quality of the model.

[0060] On the one hand, the present invention quantifies the generation potential and uses it as the data for training the retriever, which bridges the difference between retrieval and generation to a certain extent; on the other hand, the present invention performs data augmentation on the basis of quantifying the generation potential. Although the labeled knowledge has a high generation potential, there is still a considerable part of the unlabeled knowledge that also has a very high generation potential. These unlabeled knowledge form positive and negative sample pairs based on the generation potential as pseudo-labels, improving the accuracy of the retriever.

[0061] Among them, labeled knowledge refers to the correct knowledge directly related to the current task and already labeled (such as the knowledge required to generate reply labels); unlabeled knowledge is external knowledge that has not been labeled but may be superficially related to the problem (actually irrelevant or easily confused).

[0062] This method is a general data augmentation method and can be extended to various knowledge-intensive tasks of retrieval-enhanced generation.

[0063] The present invention provides a dialogue generation method based on enhancing the potential ranking of external knowledge. The dialogue generation model used includes a retriever and a generator. The retriever retrieves relevant knowledge fragments according to the dialogue history, and the generator generates responses according to the dialogue history and the retrieved relevant knowledge fragments. The method specifically includes the following steps:

[0064] S1, obtain the dialogue history;

[0065] S2, sequentially splice different external knowledge fragments with the obtained dialogue history, input them into the retriever fine-tuned with a data-augmented training set, output the relevance between the dialogue history and different knowledge fragments, and obtain the corresponding knowledge fragment ranking result after sorting the relevance values;

[0066] S3, splice the dialogue history with the top set number of knowledge fragments with the highest relevance after sorting by the retriever, input them into the generator fine-tuned through training, and output the response generated according to the dialogue history and the recalled knowledge fragments.

[0067] Specifically, after receiving a dialogue, the dialogue system will retrieve relevant knowledge in the manner as Figure 1 shown, generate a response according to the dialogue history and the recalled knowledge fragments, and return the generated content. The form of the dialogue input by the user is flexible and diverse unstructured natural language text, generally involving certain specific topics, such that the dialogue generation model needs to supplement relevant external knowledge. Sequentially splice different knowledge fragments in the external knowledge base with the obtained dialogue history, input them into the retriever fine-tuned with a data-augmented training set, output the relevance between the dialogue history and different knowledge fragments, and obtain the corresponding knowledge fragment ranking result after sorting according to the relevance values. In a preferred embodiment, splice the dialogue history with the top 5 knowledge fragments with high relevance after sorting and recalling by the retriever, input them into the generator model fine-tuned through training, and output the response content generated according to the dialogue history and the recalled knowledge fragments.

[0068] In one embodiment, the training and inference processes of the retriever and generator of the dialogue generation model include:

[0069] Construct training sample A. Each training sample A includes a dialogue history, multiple external knowledge fragments, a knowledge relevance label, and a response label; quantify the generation potential of the external knowledge fragments in training sample A to obtain training sample B;

[0070] In training sample B, the dialogue history is concatenated with labeled knowledge and unlabeled knowledge respectively to construct positive and negative sample pairs; according to the generation potential of exogenous knowledge fragments in training sample B, unlabeled knowledge with a generation potential higher than a set value is selected from the exogenous knowledge fragments as pseudo-labeled knowledge to construct data-augmented positive and negative sample pairs; the generation potential of exogenous knowledge fragments in training sample B is corrected and normalized; each positive and negative sample pair, along with the corresponding knowledge generation potential, dialogue history, and response label, constitutes training sample C.

[0071] The positive and negative sample pairs and the corresponding dialogue history in training sample C are input into the retriever, and the retriever is fine-tuned through contrastive learning to output the relevance between the dialogue history and different knowledge fragments; for different knowledge fragments, they are sorted according to the relevance to obtain the sorted set of exogenous knowledge fragments.

[0072] The dialogue history, the sorted set of exogenous knowledge fragments, and the response label are input into the BART model for sequence-to-sequence generation task training to obtain a fine-tuned generator, and then the generator performs inference.

[0073] The retriever is used to sort the exogenous knowledge to obtain the top set number of exogenous knowledge fragments, which form the context with the dialogue history of the test sample and are input into the fine-tuned generator for inference to output a context-based response.

[0074] Specifically, as Figure 2 shown, the training process and inference process of the dialogue generation model of the present invention are as follows:

[0075] Training sample A and test sample are constructed. Each sample includes four parts: dialogue history, multiple exogenous knowledge fragments, knowledge relevance label, and response label.

[0076] The generation potential of the training sample is quantified to obtain the training sample B after generation potential quantification. The generation potential values of the knowledge fragments are output, and this is iteratively executed for all exogenous knowledge until all knowledge generation potential quantification results are obtained. On training sample B, the dialogue history is concatenated with labeled knowledge and unlabeled knowledge respectively to construct positive and negative sample pairs. According to the generation potential, a threshold ∈ is set to select unlabeled knowledge with a relatively high generation potential as pseudo-labeled knowledge to construct data-augmented positive and negative sample pairs. In addition, the generation potential quantification results of training sample B are corrected and normalized. Each positive and negative sample pair, the corresponding knowledge generation potential, the corresponding dialogue history, and the response label constitute the augmented training sample C.

[0077] Using training sample C, the retriever is fine-tuned: the positive and negative sample pairs and the corresponding dialogue history are input for contrastive learning, and the output is the relevance value between the dialogue history and the knowledge fragments. For different knowledge fragments, they are sorted according to the relevance value to obtain the knowledge set ranked among the top after retrieval, that is, the sorted set of exogenous knowledge.

[0078] Fine-tuning the generator: Input the dialogue history, the set of external knowledge of the training samples sorted by the retriever, and the response label into the BART model to perform sequence-to-sequence generation task training, and obtain the fine-tuned generator.

[0079] Input the dialogue history and external knowledge of the test samples into the fine-tuned retriever, output the correlation values of each piece of external knowledge, sort all the external knowledge, and obtain the sorted set of external knowledge.

[0080] Input the dialogue history and the sorted set of external knowledge of the test samples into the fine-tuned generator, and the output is the response of the model to this dialogue history.

[0081] In one embodiment, the quantification of the generation potential of the external knowledge fragments in training sample A specifically includes:

[0082] The generation potential is used to evaluate the potential of different external knowledge fragments for generating response labels; input the dialogue history into the generation model to obtain the probabilities of different tokens in the generated response label, multiply the probabilities of all tokens to obtain probability one P(y|x,θ); then, concatenate the dialogue history with a specific external knowledge fragment and obtain probability two P(y|x,k,θ) of generating the response label in the same way; divide probability two by probability one and take the logarithm to obtain the quantification result of the generation potential; the generation potential of all external knowledge fragments of the same dialogue sample is quantified through the generation potential quantification function.

[0083] In one embodiment, the generation potential quantification function specifically includes:

[0084]

[0085] Among them, x represents the dialogue history, y represents the response label, k represents a piece of external knowledge fragment, θ represents the parameters of the generation model, and Q(x,y,k,θ) represents the quantification result of the generation potential of the external knowledge fragment k.

[0086] Specifically, the present invention introduces a concept called "generation potential quantification function" to evaluate the potential of different knowledge for generating response labels. The generation potential quantification function takes into account the generation capabilities of two key parts in the generation process: one is the ability of the dialogue history to generate response labels. The other is the ability of the dialogue history to generate response labels when providing external knowledge. In a preferred embodiment, the generation model selects the BART-Large generation model.

[0087] In one embodiment, the quantification of the generation potential of all external knowledge fragments of the same dialogue sample through the generation potential quantification function specifically includes:

[0088] Decompose the generation process into multiple time steps \(t\), \(n\) represents the length of the response label \(y\), and \(y\) t represents the token generated at time step \(t\); for each time step \(t\), the conditional probability \(p(y\) t |y <t , x, \(\theta\)) represents the probability that the dialogue generation model generates the current token given the tokens before time step \(t\); the probability of generating a token sequence is obtained by multiplying continuously; the probability \(p(y\) t |x, \(\theta\)) represents the probability that the dialogue history itself generates the response label, and the probability \(p(y\) t |x, k, \(\theta\)) represents the probability that the dialogue history generates the response label after combining a specific external knowledge fragment; the logarithm of the quotient of \(P(y|x,k,\theta)\) and \(P(y|x,\theta)\) represents the logarithm magnification of the probability of generating the response label after the dialogue history is concatenated with specific knowledge compared to the probability before concatenation, which can also be called the gain magnification. Since there are often extremely high gains, the logarithmic operation makes the data smoother. The present invention refers to this gain as the "generation potential" of the corresponding knowledge, and \(Q(x,y,k,\theta)\) represents the quantization result of the generation potential for knowledge \(k\).

[0089] The generation potential takes into account the relationship between the dialogue history, knowledge, and response label at the same time. It is a relative quantization standard and is meaningful only when all three correspond. If any one of the three changes, the generation potential will also change.

[0090] In one of the embodiments, correct and normalize the generation potential of the external knowledge fragment in the training sample B, specifically including: when normalizing, use the maximum-minimum normalization.

[0091] After defining the generation potential, the quantization results of the generation potential of all external knowledge fragments can be output by calculating the training sample A through the above formula. The quantization results of the generation potential are floating-point decimals, so there is a basis for numerical correction, normalization, and setting thresholds to screen out false labels.

[0092] After quantization, the label knowledge may not have the maximum generation potential. The present invention believes that there is a deviation between the generation model and the dataset label when quantifying the generation potential, resulting in this problem. The correction of the label knowledge can alleviate this problem. In the specific implementation of the present invention, the generation potential of the label knowledge will be corrected so that its generation potential is the largest among all external knowledge of the corresponding training sample. Then, normalize the generation potential of all knowledge. Normalization can improve the stability and convergence speed of data training. The normalization process uses the maximum-minimum normalization. The correction and normalization formulas are defined as follows:

[0093]

[0094] k = m·g+(1 - m)·k;

[0095]

[0096] Among them, m represents the correction coefficient, ensuring that the generation potential of the labeled knowledge after correction is the largest among all the exogenous knowledge of the current training samples. k max represents the maximum value in the set of generation potentials corresponding to the unlabeled knowledge. k represents the generation potential corresponding to a specific knowledge, and g is an indicator function. If the current k is labeled knowledge, then g = 1; otherwise, g = 0. X represents any one in a set of data, and X max represents the maximum value in a set of data, and X min represents the minimum value in a set of data.

[0097] Next, the present invention will demonstrate the quantization process of the knowledge generation potential with an example, showing the quantization of the generation potential, label correction, and normalization. In this example, the present invention reduces the scale of the out-of-sample exogenous knowledge base to 3 pieces of knowledge for subsequent demonstration. Table 1 shows the dialogue history and reply label information of the samples, and Table 2 shows the exogenous knowledge of the samples. Among them, Knowledge 1 is labeled knowledge, and the rest are unlabeled knowledge. The present invention hopes to intuitively feel the calculation process of the generation potential by showing the quantized generation potential values.

[0098] Table 1 Dialogue History and Reply Label Information of the Samples

[0099]

[0100]

[0101] Table 2 Exogenous Knowledge of the Samples

[0102]

[0103] The present invention splices the reply label and the dialogue history before and after Knowledge 1 and inputs them into the BART generation model respectively. In the two inputs, whether the dialogue history is spliced with Knowledge 1 or not is the input sequence of the generation model, and the reply label is the target sequence of the generation model. By recording the generation probabilities of all the tokens in the generated reply label before and after the generation model utilizes the exogenous knowledge, the quantization result of the generation potential is calculated and output, which is 29.31094213861547. Similarly, the quantization result of Knowledge 2 is 10.304353492513857, and that of Knowledge 3 is 4.884158780328512.

[0104] The maximum value of the non-label knowledge generation potential corresponds to Knowledge 2, which is 10.304353492513856. Substituting it into the above m calculation formula, the correction coefficient is 0.9115385058807445. Substituting it into the above k calculation formula, the corrected generation potentials corresponding to Knowledge 1, Knowledge 2, and Knowledge 3 are 3.504428241, 0.911538505, and 0.432059983 respectively. After performing maximum-minimum normalization, the final values are 1.0, 0.15606154, and 0.0 respectively.

[0105] In one embodiment, before the generation potential correction and normalization of the training sample B, the dialogue history is respectively concatenated on the label knowledge and the non-label knowledge to construct positive and negative sample pairs; according to the generation potential of the external knowledge fragments in the training sample B, non-label knowledge with a generation potential higher than a set value is screened out from the external knowledge fragments as pseudo-label knowledge, and data-augmented positive and negative sample pairs are constructed, which specifically include the following steps:

[0106] S21: Determine whether the external knowledge fragment is label knowledge according to the annotation information of the external knowledge fragment. If so, concatenate the dialogue history on the label knowledge as the positive sample; if not, proceed to step S22:

[0107] S22: Set a threshold ∈, calculate the product of the generation potential of the label knowledge and the threshold ∈. If the generation potential of the non-label knowledge is greater than the product, use the current non-label knowledge as pseudo-label knowledge and concatenate the dialogue history on the pseudo-label knowledge as the positive sample; if the generation potential of the non-label knowledge is less than or equal to the product, the current non-label knowledge is not used as pseudo-label knowledge, and the dialogue history is concatenated on the non-label knowledge as the negative sample; the positive sample and the negative sample form data-augmented positive and negative sample pairs.

[0108] Specifically, the construction of positive and negative sample pairs, as well as setting the threshold ∈, screening non-label knowledge with relatively high generation potential as pseudo-label knowledge, and constructing data-augmented positive and negative sample pairs are mentioned above. Conventional contrastive learning training only uses label knowledge as the positive sample and non-label knowledge as the negative sample to construct positive and negative sample pairs. The present invention considers screening negative sample knowledge fragments with relatively high generation potential after quantifying the generation potential, which can be used as pseudo-label knowledge and regarded as positive samples to construct data-augmented positive and negative sample pairs. This method utilizes the advantages of quantification, expands the limited training data, and can improve the generalization of the retrieval accuracy of the retriever. The specific process of constructing positive and negative sample pairs is as Figure 3 shown.

[0109] Among them, setting a threshold ∈ is required to determine whether it is pseudo-label knowledge, and the specific processing process includes:

[0110] Calculating the product of the generation potential of the label knowledge and the threshold ∈;

[0111] If the calculated product is less than the potential of generating non-label knowledge, then the non-label knowledge meets the pseudo-label standard and is used as a positive sample; if the product is greater than or equal to the potential of generating non-label knowledge, then it is used as a negative sample.

[0112] Return the pair of positive and negative samples.

[0113] The selection of the threshold is very important. Because if the threshold is too large, there will be too little pseudo-label knowledge, and the gain for training the retriever is not obvious. If the threshold is too small, there will be too much low-quality pseudo-label data, which will damage the performance of the retriever and increase the computational overhead. In addition, the threshold selection has a strong correlation with the dialogue and external knowledge, and it should be different for different data and scenarios. It is recommended that the selection range is between [0.8, 1.6].

[0114] The following example shows the specific process of constructing the pair of positive and negative samples. In the previous text, the generation potentials of Knowledge 1, Knowledge 2, and Knowledge 3 before correction and normalization are 29.31094213861547, 10.304353492513857, and 4.884158780328512 respectively. Then, the pairs of positive and negative samples (Knowledge 1, Knowledge 2) and (Knowledge 1, Knowledge 3) can be constructed. If the threshold ∈ = 1.2, the product of it and the generation potential of Knowledge 1 is greater than the generation potentials of Knowledge 2 and Knowledge 3. Therefore, it is impossible to construct the pair of positive and negative samples using the pseudo-label. If the threshold ∈ = 0.2, the product of it and the generation potential of Knowledge 1 is less than the generation potential of Knowledge 2 and greater than the generation potential of Knowledge 3. Therefore, the pseudo-label Knowledge 2 can be obtained to expand the pair of positive and negative samples for data augmentation (Knowledge 2, Knowledge 3).

[0115] In one of the embodiments, input the pair of positive and negative samples in the training sample C and the corresponding dialogue history into the retriever, and fine-tune the retriever through contrastive learning. The output is the correlation between the dialogue history and different knowledge fragments. For different knowledge fragments, sort them according to the correlation to obtain the sorted set of external knowledge fragments, specifically including:

[0116] The fine-tuning process of the retriever is as follows:

[0117] The retriever selects any pre-trained language model containing an encoder structure;

[0118] The loss function in the retriever fine-tuning process selects the Margin Ranking loss, which is used to ensure that the score of the positive sample is higher than that of the negative sample, and the score gap between the two is at least a given margin;

[0119] During fine-tuning, input the pair of positive and negative samples, the corresponding generation potential, and the dialogue history in the training sample C into the retriever, and update the parameters through backpropagation of the loss function.

[0120] The inference process of the retriever is as follows:

[0121] For the test samples, the following operations are iteratively performed until the relevance of all external knowledge is output:

[0122] Concatenate the conversation history with a single piece of external knowledge segment as the context information;

[0123] Input the context information into the fine-tuned retriever to output the relevance between the conversation history and the current external knowledge segment;

[0124] Sort the external knowledge segments according to the relevance values. The larger the relevance value, the higher the corresponding knowledge relevance and the higher the ranking. Output the sorted result of the external knowledge segments.

[0125] Specifically, in the multi-turn dialogue generation task with external knowledge enhancement, the user's utterances generally involve certain specific topics. The internal knowledge of the generator model often cannot well guide the generation model to generate informative responses. In addition, the generation may face problems such as hallucination and templatized and boring responses. Therefore, the generator needs to supplement relevant external knowledge. The retriever aims to rank and recall the most relevant knowledge to the current dialogue from the external knowledge base, and the knowledge ranked at the top will be input to the generator subsequently to address the above problems and challenges.

[0126] There are many basic models that the retriever can choose from. Any pre-trained language model containing an Encoder structure can be used. Here, the open-source model Electra-Large-Discriminator is selected for fine-tuning.

[0127] Facing the training samples and test samples, the retriever performs fine-tuning and inference respectively, where the fine-tuning process is as Figure 4 shown, and the inference process is as Figure 5 shown.

[0128] (I) Fine-tuning process:

[0129] The loss function selects Margin Ranking Loss, also known as Pairwise Ranking Loss, which is commonly used in ranking tasks such as recommendation systems and information retrieval. The goal of this loss function is to ensure that the score of the positive sample is higher than that of the negative sample, and the score difference between the two is at least a given margin. The loss function is defined as follows:

[0130] Loss(x1,x2)=max(0,-y*(x1-x2)+margin)

[0131] Among them, if the potential values of the positive and negative sample pairs corresponding to x1 and x2 respectively, and if x1 should be ranked before x2 according to the value, then y = 1, otherwise y = -1. In specific implementation, the present invention selects a smaller margin value, and the recommended range is between [0.1, 0.2].

[0132] During fine-tuning, the positive and negative sample pairs constructed from the training samples and the potential values are input into the retriever, and the parameters are updated through loss backpropagation.

[0133] (II) Inference process:

[0134] For the test samples, the present invention iteratively performs the following operations:

[0135] Concatenate the conversation history with a single piece of external knowledge as the context information;

[0136] Input the obtained context information into the fine-tuned retriever;

[0137] Output the relevance predicted by the retriever, that is, the relevance between the conversation history and the current external knowledge;

[0138] If there is still external knowledge for which the predicted relevance value has not been obtained, repeat the above steps. If the relevance values of all external knowledge have been obtained, continue to the next step;

[0139] Sort according to the relevance values. The larger the value, the higher the corresponding knowledge relevance and the higher the ranking. Output the result of the sorted knowledge.

[0140] In a preferred embodiment, the present invention selects the top 5 pieces of knowledge and inputs them to the generator for fine-tuning and inference of the generation task.

[0141] In one embodiment, the conversation history, the sorted set of external knowledge fragments, and the reply label are input into the BART model for sequence-to-sequence generation task training to obtain a fine-tuned generator. Subsequently, the generator performs inference, which specifically includes:

[0142] The fine-tuning process specifically includes:

[0143] The cross-entropy loss function is used in the generator fine-tuning process to minimize the negative log-likelihood of the sequence generated by the BART model, so that the generated sequence is as close as possible to the reply label sequence; the conversation history in the training samples, the top set number of external knowledge fragments recalled by the retriever after sorting, and the reply label are input into the generator, and the parameters are updated through the backpropagation of the cross-entropy loss function.

[0144] The inference process specifically includes:

[0145] The retriever is used to sort the exogenous knowledge to obtain a set number of exogenous knowledge fragments before sorting, which are combined with the conversation history of the test sample to form a context, which is input into the generator for reasoning and outputs a response based on the context.

[0146] Specifically, there are many basic models that the generator can choose from. Here, the present invention selects the open source Encoder-Decoder model BART-Large for fine-tuning. Faced with training samples and test samples, the generator performs fine-tuning and reasoning respectively. The specific process is as follows Figure 6 and Figure 7 shown.

[0147] (I) Fine-tuning process:

[0148] The BART model usually uses the cross-entropy loss function in sequence-to-sequence generation tasks. The goal of the Cross-Entropy loss function is to minimize the negative log-likelihood of the model-generated sequence so that the generated sequence is as close to the true target sequence as possible. The loss function is defined as follows:

[0149] The probability distribution of the sequence generated by the model is P, and the probability distribution of the target sequence is Q. N is the sequence length, V is the vocabulary size, and Q ij is the probability of the jth word in the vocabulary at the i-th position in the target sequence, P ij is the probability of the jth word in the vocabulary at the i-th position in the sequence generated by the model.

[0150] In a preferred embodiment, during fine-tuning, the conversation history in the training sample, the top 5 pieces of knowledge sorted and recalled by the retriever, and the reply label are input into the generator, and the parameters are updated through loss back propagation.

[0151] (II) Reasoning process:

[0152] First, the top 5 sets of exogenous knowledge are obtained by using the retriever to rank, and the context is formed with the conversation history of the test sample. Then, it is input into the generator for reasoning and outputs a response based on the context.

[0153] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.

[0154] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for dialogue generation based on enhanced ranking of exogenous knowledge generation potential, characterized in that: The adopted dialogue generation model includes a retriever and a generator. The retriever retrieves relevant knowledge fragments according to the dialogue history, and the generator generates responses according to the dialogue history and the retrieved relevant knowledge fragments. The method specifically includes: Obtaining a history of conversations involving a specific topic; Different exogenous knowledge fragments are concatenated with the acquired conversation history one by one, and input into the retriever fine-tuned with the data-enhanced training set. The correlation between the conversation history and different knowledge fragments is output, and the corresponding knowledge fragment ranking result is obtained after sorting the correlation values. The conversation history is concatenated with a set number of knowledge fragments with the highest relevance after sorting by the retriever, and the input is input into the trained and fine-tuned generator, which outputs the response generated based on the conversation history and the recalled knowledge fragments.

2. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 1 is characterized in that: The conversation is unstructured natural language text and involves a specific topic.

3. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 1 is characterized in that: The retriever and generator training and reasoning process of the dialogue generation model includes: Construct training samples A, each of which includes a conversation history, multiple exogenous knowledge fragments, knowledge relevance labels, and reply labels; quantify the generation potential of the exogenous knowledge fragments in training sample A to obtain training samples B; In training sample B, the conversation history is spliced ​​onto the labeled knowledge and unlabeled knowledge respectively to construct positive and negative sample pairs; according to the generation potential of the exogenous knowledge fragments in training sample B, the unlabeled knowledge with generation potential higher than the set value in the exogenous knowledge fragments is selected as pseudo-labeled knowledge to construct data-enhanced positive and negative sample pairs; the generation potential of the exogenous knowledge fragments in training sample B is corrected and normalized; each positive and negative sample pair and the corresponding knowledge generation potential, conversation history, and reply label constitute training sample C; The positive and negative sample pairs and the corresponding dialogue history in the training sample C are input into the retriever, and the retriever is fine-tuned through contrastive learning. The output is the correlation between the dialogue history and different knowledge fragments. Different knowledge fragments are sorted according to the correlation to obtain a sorted set of exogenous knowledge fragments. The conversation history, the sorted set of external knowledge fragments, and the reply labels are input into the BART model for sequence-to-sequence generation task training to obtain a fine-tuned generator, which is then used for reasoning. The retriever is used to sort the exogenous knowledge to obtain a pre-set number of exogenous knowledge fragments, which are combined with the conversation history of the test sample to form a context. The fragments are input into the fine-tuned generator for reasoning and output a context-based response.

4. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 3 is characterized in that: The quantification of the generation potential of the exogenous knowledge fragments in the training sample A specifically includes: The generation potential is used to evaluate the potential of different exogenous knowledge fragments for generating reply labels. The conversation history is input into the generation model to obtain the probabilities of different tags in the generated reply labels, and the probabilities of all tags are multiplied to obtain the probability one P(y|x,θ). The conversation history is spliced ​​with specific exogenous knowledge fragments, and the probability two P(y|x,k,θ) of generating reply labels is obtained in the same way. The probability two is divided by the probability one and the logarithm is taken to obtain the generation potential quantification result. The generation potential of all exogenous knowledge fragments in the same conversation sample is quantified through the generation potential quantification function.

5. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 4 is characterized in that: The generation potential quantification function specifically includes: Among them, x represents the conversation history, y represents the reply label, k represents an exogenous knowledge fragment, θ represents the parameters of the generation model, and Q(x, y, k, θ) represents the quantitative result of the generation potential of the exogenous knowledge fragment k.

6. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 4, characterized in that: All the exogenous knowledge fragments of the same dialogue sample are quantified in terms of generation potential through a generation potential quantification function, specifically including: The generation process is decomposed into multiple time steps t, n represents the length of the reply label y, y t represents the tag generated at time step t; for each time step t, the conditional probability p(y t |y <t ,x,θ) represents the probability of the dialogue generation model generating the current token given the token before time step t; the probability of generating a token sequence is obtained by multiplication; the probability p(y t |x,θ) represents the probability of generating a reply tag in the conversation history itself, and the probability p(y t |x,k,θ) represents the probability of generating a reply tag after the dialogue history is combined with a specific exogenous knowledge fragment; the logarithm of the quotient of P(y|x,k,θ) and P(y|x,θ) represents the logarithmic multiple of the probability of generating a reply tag after the dialogue history is spliced ​​with specific knowledge compared to the probability before splicing.

7. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 3 is characterized in that: In the training sample B, the conversation history is spliced ​​onto the label knowledge and the non-label knowledge to construct positive and negative sample pairs; According to the generation potential of the exogenous knowledge fragments in the training sample B, the non-labeled knowledge with the generation potential higher than the set value in the exogenous knowledge fragments is selected as the pseudo-labeled knowledge, and the positive and negative sample pairs of data enhancement are constructed, which specifically includes the following steps: S21: Determine whether the exogenous knowledge fragment is labeled knowledge according to the annotation information of the exogenous knowledge fragment. If yes, splice the conversation history onto the labeled knowledge as a positive sample; if no, proceed to step S22: S22: Set a threshold ∈, calculate the product of the potential for generating label knowledge and the threshold ∈, and if the potential for generating non-label knowledge is greater than the product, use the current non-label knowledge as pseudo-label knowledge, and splice the conversation history onto the pseudo-label knowledge as a positive sample; If the generation potential of the unlabeled knowledge is less than or equal to the product, the current unlabeled knowledge is not used as pseudo-labeled knowledge, and the conversation history is spliced ​​onto the unlabeled knowledge as a negative sample; the positive sample and the negative sample constitute a data-enhanced positive-negative sample pair.

8. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 3 is characterized in that: The correction and normalization of the generation potential of the exogenous knowledge fragments in the training sample B specifically includes: during normalization, maximum and minimum value normalization is adopted.

9. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 3 is characterized in that: The positive and negative sample pairs in the training sample C and the corresponding dialogue history are input into the retriever, and the retriever is fine-tuned through contrastive learning, and the output is the correlation between the dialogue history and different knowledge fragments; for different knowledge fragments, they are sorted according to the correlation to obtain a sorted set of exogenous knowledge fragments, which specifically includes: The process of fine-tuning the retriever is as follows: The retriever selects any pre-trained language model including an encoder structure; The loss function used in the retriever fine-tuning process is the Margin Ranking loss, which is used to ensure that the score of the positive sample is higher than the score of the negative sample, and the difference between the scores of the two is at least a given margin; During fine-tuning, the positive and negative sample pairs in the training sample C, the corresponding generation potential, and the dialogue history are input into the retriever, and the parameters are updated through the back propagation of the loss function; The reasoning process of the retriever is as follows: For the test sample, the following operations are iteratively performed until the relevance of all exogenous knowledge is output: Splice the conversation history with a single piece of external knowledge as context information; Input the context information into the fine-tuned retriever, and output the relevance between the conversation history and the current exogenous knowledge fragment; The exogenous knowledge fragments are sorted according to the relevance value. The larger the relevance value is, the higher the corresponding knowledge relevance is and the higher the ranking is. The sorted exogenous knowledge fragments are output as the sorting result.

10. The method for generating dialogues based on the enhanced ranking of exogenous knowledge generation potential according to claim 3, characterized in that: The conversation history, the sorted set of external knowledge fragments, and the reply label are input into the BART model, and sequence-to-sequence generation task training is performed to obtain a fine-tuned generator, and then the generator performs reasoning, specifically including: The fine-tuning process specifically includes: The generator fine-tuning process uses the cross-entropy loss function to minimize the negative log-likelihood of the sequence generated by the BART model, making the generated sequence as close as possible to the reply label sequence; the conversation history in the training sample, a pre-set number of exogenous knowledge fragments sorted and recalled by the retriever, and the reply label are input into the generator, and the parameters are updated through the back propagation of the cross-entropy loss function; The reasoning process specifically includes: The retriever is used to sort the exogenous knowledge to obtain a pre-set number of exogenous knowledge fragments, which are combined with the conversation history of the test sample to form a context, which is input into the generator for reasoning and outputs a context-based response.

Citation Information

Cited By

  • Model training method, dialogue processing method and dialogue system

    CN121960792A