Multi-round question and answer data generation method and device, equipment and storage medium

Through adversarial training and reinforcement learning optimization generators, combined with generative adversarial networks and reinforcement learning methods, the problem of lack of representation and diversity of multiple rounds of question-answer test data in the existing technology is solved, and efficient and high-quality multi-round question-answer data generation is achieved.

CN119940413APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510067964.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing intelligent Q&A system relies on manual design of test data sets in multiple rounds of Q&A tests, which is inefficient and difficult to simulate user interaction in a comprehensive and realistic manner, resulting in a lack of representation and diversity in the test data.

Method used

The initial generator is trained and optimized through the preset strategy gradient method and the reward function of the generative adversarial network structure and the reinforcement learning method, and generates target question-and-answer pairs, and uses the preset evaluation algorithm to evaluate the matching degree to judge the qualified conditions of the question-and-answer pairs.

Benefits of technology

The efficiency and quality of multiple rounds of Q&A data generation is significantly improved, avoiding the limitations of manual design, and ensuring extensive coverage and sense of realism of generated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940413A_ABST
    Figure CN119940413A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-round question and answer data generation method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the adversarial training and reinforcement learning optimization of an initial generator through a preset strategy gradient method, and employing a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method, obtaining a generator after learning optimization; the generative adversarial network structure comprises a generator and a discriminator; generating a target number of target question and answer pairs for the input user intention and dialogue context based on the learning optimization generator, and performing matching degree evaluation on the target question and answer pairs and the target dialogue data set by using a preset evaluation algorithm to obtain an evaluation result; and judging whether the target question and answer pair meets a preset qualified condition based on a matching degree evaluation result, and if the target question and answer pair meets the preset qualified condition, outputting the target question and answer pair meeting the preset qualified condition. Multi-round question and answer test data can be efficiently, comprehensively and automatically generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for generating multi-round question and answer data. Background Art

[0002] Existing intelligent question-answering systems usually rely on manually designed test data sets in multi-round question-answering tests. This manually designed method is not only inefficient, but also difficult to fully and realistically simulate the diverse interactions of users, resulting in a lack of representativeness and diversity in the test data. Especially when faced with complex multi-round question-answering scenarios, manually designed test data often cannot cover the user's multiple possible intentions, questions and answers of different difficulty levels, and interactions in different situations, and thus cannot truly reflect the performance and robustness of the system.

[0003] From the above, it can be seen that how to efficiently and comprehensively automatically generate multi-round question-answering test data is a problem that needs to be solved urgently. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for generating multi-round question and answer data, which can efficiently and comprehensively automatically generate multi-round question and answer test data. The specific scheme is as follows:

[0005] In a first aspect, the present application provides a method for generating multi-round question-answering data, comprising:

[0006] The initial generator is respectively subjected to adversarial training and reinforcement learning optimization by using a preset policy gradient method and a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method to obtain a generator after learning optimization; the generative adversarial network structure includes a generator and a discriminator;

[0007] Generate a target number of target question-answer pairs based on the input user intent and conversation context by the learned optimized generator, and use a preset evaluation algorithm to evaluate the matching degree of the target question-answer pairs and the target conversation dataset to obtain corresponding evaluation results;

[0008] Based on the matching evaluation result, it is determined whether the target question-answer pair meets the preset qualification conditions. If the target question-answer pair meets the preset qualification conditions, the target question-answer pair that meets the preset qualification conditions is output.

[0009] Optionally, the method of performing adversarial training and reinforcement learning optimization on the initial generator by presetting a policy gradient method and using a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method to obtain a generator after learning optimization includes:

[0010] Generate several initial question-answer pairs based on the input user intent and conversation context through the initial generator of the generative adversarial network structure;

[0011] The initial generator is adversarially trained by a preset policy gradient method and based on a discriminator of a generative adversarial network structure and the initial question-answer pair to obtain an adversarially trained generator, and the adversarially trained generator is reinforced learned and optimized using a reward function of a reinforcement learning method to obtain a learned optimized generator.

[0012] Optionally, the initial generator is subjected to adversarial training by a preset policy gradient method and based on a discriminator of a generative adversarial network structure and the initial question-answer pair to obtain an adversarially trained generator, and the adversarially trained generator is subjected to reinforcement learning optimization using a reward function of a reinforcement learning method to obtain a learning-optimized generator, including:

[0013] Using the discriminator of the generative adversarial network structure to perform authenticity scoring on the initial question-answer pair to obtain a corresponding authenticity scoring result;

[0014] Determining a discriminator loss function value for characterizing the performance of the initial generator based on the discriminator loss function and the authenticity score result;

[0015] If the discriminator loss function value is greater than the preset loss threshold, it indicates that the performance of the initial generator is poor. Based on the discriminator loss function value and the preset policy gradient method and using the reward function of the reinforcement learning method, the initial generator is subjected to reinforcement learning optimization to obtain a generator after learning optimization.

[0016] Optionally, the performing reinforcement learning optimization on the initial generator based on the discriminator loss function value and a preset policy gradient method and using a reward function of a reinforcement learning method to obtain a generator after learning optimization includes:

[0017] Obtaining a context score for the initial question-answer pair by calculating the cosine similarity between the initial question-answer pair and the conversation context;

[0018] Obtaining a diversity score of the initial question-answer pair by calculating the semantic similarity between any two of the initial question-answer pairs;

[0019] Determine the text complexity of the initial question-answer pair by statistically analyzing different text syntactic structures in the initial question-answer pair, then obtain the lexical richness of the initial question-answer pair using a preset lexical richness index, and determine the complexity score of the initial question-answer pair based on the text length of the initial question-answer pair, the text complexity, the lexical richness, and the corresponding weighting coefficients;

[0020] The reward function value is determined based on the context score, the diversity score, the complexity score and the corresponding weighting coefficients, and the initial generator is subjected to reinforcement learning optimization through the reward function value and the preset policy gradient method to obtain a generator after learning optimization.

[0021] Optionally, the using a preset evaluation algorithm to perform a matching evaluation on the target question-answer pair and the target dialogue dataset to obtain a corresponding evaluation result includes:

[0022] We filter the data features of manually annotated standard data sets and historical real conversation data, and use the filtered data to build the target conversation data set.

[0023] A preset evaluation algorithm is used to evaluate the matching degree of the target question-answer pair and the target dialogue data set to obtain a corresponding evaluation result.

[0024] Optionally, the using a preset evaluation algorithm to perform a matching evaluation on the target question-answer pair and the target dialogue dataset to obtain a corresponding evaluation result includes:

[0025] The BLEU evaluation algorithm, the ROUGE evaluation algorithm, and the cosine similarity evaluation algorithm are used to evaluate the matching degree of the target question-answer pair and the target dialogue data set to obtain corresponding evaluation results.

[0026] Optionally, after judging whether the target question-answer pair meets a preset qualification condition based on the matching evaluation result, the method further includes:

[0027] If the target question-answer pair does not meet the preset qualification conditions, jump to the step of performing adversarial training and reinforcement learning optimization on the initial generator respectively by using the preset policy gradient method and utilizing the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method.

[0028] In a second aspect, the present application provides a multi-round question and answer data generating device, comprising:

[0029] A generator optimization module, which is used to perform adversarial training and reinforcement learning optimization on the initial generator by using a preset policy gradient method and a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method, so as to obtain a generator after learning optimization; the generative adversarial network structure includes a generator and a discriminator;

[0030] A matching evaluation module is used to generate a target number of target question-answer pairs based on the input user intent and dialogue context by the generator after learning optimization, and to perform matching evaluation on the target question-answer pairs and the target dialogue data set using a preset evaluation algorithm to obtain corresponding evaluation results;

[0031] The question-answer pair judgment module is used to judge whether the target question-answer pair meets the preset qualification conditions based on the matching evaluation result. If the target question-answer pair meets the preset qualification conditions, the question-answer pair that meets the preset qualification conditions is output.

[0032] In a third aspect, the present application provides an electronic device, including:

[0033] Memory, used to store computer programs;

[0034] A processor is used to execute the computer program to implement the aforementioned method for generating multi-round question and answer data.

[0035] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned multi-round question and answer data generation method.

[0036] The present application uses a preset policy gradient method and utilizes a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method to perform adversarial training and reinforcement learning optimization on an initial generator respectively, so as to obtain a generator after learning optimization; the generative adversarial network structure includes a generator and a discriminator; based on the generator after learning optimization, a target number of question-answer pairs are generated for the input user intention and dialogue context, and a preset evaluation algorithm is used to perform a matching evaluation on the question-answer pairs and a target dialogue data set to obtain a corresponding evaluation result; based on the matching evaluation result, it is determined whether the question-answer pairs meet a preset qualification condition, and if the question-answer pairs meet the preset qualification condition, the question-answer pairs that meet the preset qualification condition are output.

[0037] As can be seen from the above, the present application uses the discriminator of the generative adversarial network structure to perform adversarial training on the generator, so that the initial generator can generate more realistic and natural target question and answer pairs, and then introduces the reward function of the reinforcement learning method to further optimize the initial generator. Then, the generator optimized by adversarial training and reinforcement learning is used to generate question and answer pairs based on the input user intention and dialogue context, and then a preset evaluation algorithm is used to evaluate the matching degree of the target question and answer pairs with the target dialogue data set. In this way, based on the matching evaluation results, the target question and answer pairs that meet the preset qualification conditions are output, which can significantly improve the efficiency and quality of multi-round question and answer data generation, avoid the limitations of artificial design, and ensure the wide coverage and realism of the generated data. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0039] Figure 1 A flow chart of a method for generating multi-round question-answering data disclosed in this application;

[0040] Figure 2 A flowchart of a specific method for generating multi-round question-answering data disclosed in this application;

[0041] Figure 3 This is a schematic diagram of the structure of a multi-round question-answering data generating device disclosed in this application;

[0042] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] At present, intelligent question-and-answer systems usually rely on manually designed test data sets in multi-round question-and-answer tests. Not only is it inefficient, but it is also difficult to fully and realistically simulate the diverse interactions of users, resulting in a lack of representativeness and diversity in the test data. Especially when faced with complex multi-round question-and-answer scenarios, manually designed test data often cannot cover the user's multiple possible intentions, questions and answers of different difficulty levels, and interactions in different situations, and cannot truly reflect the performance of the system. To this end, the present application provides a multi-round question-and-answer data generation method, which outputs target question-and-answer pairs that meet preset qualification conditions based on the matching evaluation results, which can significantly improve the efficiency and quality of multi-round question-and-answer data generation, avoid the limitations of manual design, and ensure the wide coverage and realism of the generated data.

[0045] See also Figure 1 As shown, an embodiment of the present invention discloses a method for generating multi-round question-answering data, including:

[0046] Step S11, by presetting the policy gradient method and using the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method, the initial generator is respectively subjected to adversarial training and reinforcement learning optimization to obtain a generator after learning optimization; the generative adversarial network structure includes a generator and a discriminator.

[0047] In this embodiment, the input dialogue context and user intention are obtained, and then a number of initial question-answer pairs are generated through the dialogue context and the user intention and using the initial generator of the generative adversarial network structure; the initial generator is a generator built based on a pre-trained Transformer (i.e., a neural network model) model; the initial generator is adversarially trained through the discriminator of the generative adversarial network structure and the initial question-answer pairs to obtain an adversarially trained generator, and then the adversarially trained generator is subjected to reinforcement learning optimization based on the reward function of the reinforcement learning method to obtain a learning optimized generator.

[0048] Specifically, the initial generator is subjected to adversarial training and reinforcement learning optimization respectively by using a preset policy gradient method and utilizing a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method to obtain a generator after learning optimization, including: generating a number of initial question-answer pairs based on the input user intention and conversation context by using the initial generator of the generative adversarial network structure; performing adversarial training on the initial generator by using a preset policy gradient method and based on the discriminator of the generative adversarial network structure and the initial question-answer pairs to obtain a generator after adversarial training, and performing reinforcement learning optimization on the generator after adversarial training by utilizing a reward function of a reinforcement learning method to obtain a generator after learning optimization.

[0049] In a specific implementation, if the user intent is to query the order status, and the conversation context is that the user recently has an order number 12345, and the order status is "shipped", the initial question-answer pair generated by the initial generator based on the input user intent and conversation context and using the generative adversarial network structure can be: User question: What is the status of my order 12345? Customer service answer: Your order 12345 has been shipped and is expected to be delivered within 3 working days.

[0050] In this embodiment, after obtaining the initial question-answer pair, the discriminator of the generative adversarial network structure is used to perform authenticity evaluation on the initial question-answer pair to obtain an authenticity score, and then the discriminator loss function value used to characterize the performance of the initial generator is determined by the authenticity score and the loss function of the discriminator. If the discriminator loss function value is greater than a preset loss threshold, it indicates that the quality of the initial question-answer pair generated by the initial generator is poor. Based on the discriminator loss function value and the reward function of the reinforcement learning method, the initial generator is optimized by reinforcement learning to obtain an optimized generator.

[0051] Specifically, the initial generator is adversarially trained by a preset policy gradient method and based on the discriminator of the generative adversarial network structure and the initial question and answer pair to obtain a generator after adversarial training, and the generator after adversarial training is optimized by reinforcement learning using a reward function of a reinforcement learning method to obtain a generator after learning optimization, including: using the discriminator of the generative adversarial network structure to perform authenticity scoring on the initial question and answer pair to obtain a corresponding authenticity scoring result; determining a discriminator loss function value used to characterize the performance of the initial generator based on the loss function of the discriminator and the authenticity scoring result; if the discriminator loss function value is greater than a preset loss threshold, it characterizes that the performance of the initial generator is poor, and the initial generator is optimized by reinforcement learning based on the discriminator loss function value and the preset policy gradient method and using the reward function of the reinforcement learning method to obtain a generator after learning optimization.

[0052] It can be understood that the loss function of the discriminator is used to quantify its ability to distinguish between real data and the initial question-answer pair, and the loss function of the discriminator is:

[0053] ;

[0054] Among them, x is the real data; is the distribution of real data x Expected value in is the true probability distribution of the real data x; is the authenticity score of the discriminator for the real data x, that is, the probability that the discriminator considers x to be real data; a is the initial question-answer pair generated by the initial generator; is the authenticity score output by the discriminator for the initial question and answer pair a, that is, the probability that the discriminator considers a to be real data. If the discriminator loss function is greater than the preset loss threshold, it means that the initial question and answer generated by the initial generator are very different from the real data, that is, it indicates that the performance of the initial generator is poor. The loss function of the discriminator makes the authenticity score of the discriminator for real data as close to 1 as possible, and the authenticity score of the generated false data initial question and answer pair as close to 0 as possible. It is worth mentioning that the preset loss threshold can be adjusted accordingly according to the actual situation, and no specific limitation is made here.

[0055] In this embodiment, when the initial generator is optimized by reinforcement learning using the discriminator loss function value and the reward function of the reinforcement learning method, the cosine similarity between the initial question and answer pair and the dialogue context is calculated to obtain the context score of the initial question and answer pair, the semantic similarity between any two initial question and answer pairs is calculated to obtain the diversity score of the initial question and answer pair, and the weighted average of the text length, text complexity and vocabulary richness in the initial question and answer pair is calculated to obtain the complexity score of the initial question and answer pair; then the reward function value is determined based on the context score, the diversity score, the complexity score and the corresponding weighting coefficients, and the initial generator is optimized by reinforcement learning using the reward function value and the preset policy gradient method to obtain a generator after learning optimization.

[0056] Specifically, the reinforcement learning optimization of the initial generator based on the discriminator loss function value and the preset policy gradient method and using the reward function of the reinforcement learning method to obtain the generator after learning optimization includes: obtaining the context score of the initial question and answer pair by calculating the cosine similarity between the initial question and answer pair and the dialogue context; obtaining the diversity score of the initial question and answer pair by calculating the semantic similarity between any two of the initial question and answer pairs; determining the text complexity of the initial question and answer pair by statistically analyzing different text syntactic structures in the initial question and answer pair, and then obtaining the lexical richness of the initial question and answer pair by using a preset lexical richness index, and determining the complexity score of the initial question and answer pair based on the text length of the initial question and answer pair, the text complexity, the lexical richness and the corresponding weighting coefficients; determining the reward function value based on the context score, the diversity score, the complexity score and the corresponding weighting coefficients, and performing reinforcement learning optimization on the initial generator through the reward function value and the preset policy gradient method to obtain the generator after learning optimization.

[0057] It can be understood that the cosine similarity between the initial question-answer pair and the conversation context is calculated to obtain the context score of the initial question-answer pair, and the calculation formula is:

[0058] ;

[0059] in, represents the i-th information unit in the dialogue context; a is the initial question-answer pair; is the cosine similarity between the conversation context and the initial question-answer pair. The context score is a score that measures the consistency between the initial question-answer pair and the conversation context.

[0060] In a specific embodiment, if is a set of question-answer pairs corresponding to the initial question-answer pairs, For i initial question-answer pairs, the semantic similarity between any two of the initial question-answer pairs is calculated to obtain the diversity score of the initial question-answer pairs. The calculation formula is as follows:

[0061] ;

[0062] in, is the size of the question-answer pair set A, that is, the total number of initial question-answer pairs generated; is the cosine similarity between the i-th and j-th initial question-answer pairs, to measure the semantic similarity between the two initial question-answer pairs; when the cosine similarity is smaller, it proves that the initial question-answer pairs are less similar, and the diversity score is higher.

[0063] Furthermore, the weighted average of the text length, text complexity and vocabulary richness in the initial question-answer pair is calculated to obtain the complexity score of the initial question-answer pair. The calculation formula is as follows:

[0064] ;

[0065] in, is the initial question-answer pair The length of the text can be measured using the number of words; is the initial question-answer pair The text complexity can be quantified by counting different text syntactic structures in the initial question-answer pair or calculating the dependency tree depth, the level of the syntactic tree, etc. of the initial question-answer pair; is the initial question-answer pair The vocabulary richness can be obtained by using a preset vocabulary richness index or calculating the number of different types of vocabulary in the initial question-answer pair. α, β, γ are weighting coefficients corresponding to the text length of the initial question-answer pair, the text complexity, and the vocabulary richness, respectively, and are used to indicate the importance of different complexity indicators. The weighting coefficients can be adjusted accordingly according to actual conditions and are not specifically limited here.

[0066] Further, after obtaining the context score, the diversity score, and the complexity score, the reward function value of the question-answer pair is determined using the reward function, and the formula of the reward function is:

[0067] ;

[0068] in, assigning a score to the context; assigning a score to said diversity; is the complexity score, α, β, γ and the weighted coefficients corresponding to the context score, the complexity score and the complexity score of the initial question-answer pair, respectively. After obtaining the reward function value, the initial generator is optimized by reinforcement learning based on the reward function value to obtain a generator after learning optimization.

[0069] In this embodiment, the total loss function of the initial generator is a weighted combination of the GAN loss function (i.e., the adversarial loss function of the generator) and the reinforcement learning loss function. The total loss function is as follows:

[0070] ;

[0071] ;

[0072] ;

[0073] in, is the total loss function, is the GAN loss function; is the reinforcement learning loss function; a is the initial question-answer pair generated by the initial generator; Score the authenticity of the discriminator's output for the initial question-answer pair a; is the generation strategy of the initial generator; S is the state space, including all initial question-answer pairs and dialogue context of the current dialogue); is the reward function value, which indicates the quality of generating the initial question-answer pair. The initial generator needs to be optimized under the guidance of the total loss function.

[0074] It can be understood that the preset policy gradient method is to calculate the gradient of the generator strategy with respect to its parameters to update the parameters of the initial generator, and the calculation formula is as follows:

[0075] ;

[0076] in, are the parameters of the initial generator; Probability distribution of generated strategies for the generator; is the reward function value; is the gradient of the generator strategy; through the reward function value and the update strategy Make the generated question-answer pairs optimal in terms of quality, complexity, and diversity.

[0077] Step S12: Based on the learned optimized generator, a target number of target question-answer pairs are generated for the input user intent and dialogue context, and a preset evaluation algorithm is used to perform a matching evaluation on the target question-answer pairs and the target dialogue data set to obtain a corresponding evaluation result.

[0078] In this embodiment, after the generator after learning optimization generates the target number of target question-answer pairs based on the input user intent and dialogue context, the target dialogue dataset is constructed using the manually annotated standard dataset and historical real dialogue data, and the target question-answer pairs and the target dialogue dataset are evaluated for matching using a preset evaluation algorithm to obtain corresponding evaluation results. Specifically, the use of a preset evaluation algorithm to evaluate the matching of the target question-answer pairs and the target dialogue dataset to obtain corresponding evaluation results includes: performing data feature screening on the manually annotated standard dataset and historical real dialogue data, and using the screened data to construct the target dialogue dataset; using a preset evaluation algorithm to evaluate the matching of the target question-answer pairs and the target dialogue dataset to obtain corresponding evaluation results. In a specific implementation, if the number of initial question-answer pairs generated is greater than the target number, the initial question-answer pairs are discarded.

[0079] It can be understood that the target question-answer pair and the target dialogue data set are evaluated for matching using the BLEU evaluation algorithm, the ROUGE evaluation algorithm, and the cosine similarity evaluation algorithm. Specifically, the target question-answer pair and the target dialogue data set are evaluated for matching using a preset evaluation algorithm to obtain a corresponding evaluation result, including: the target question-answer pair and the target dialogue data set are evaluated for matching using the BLEU evaluation algorithm, the ROUGE evaluation algorithm, and the cosine similarity evaluation algorithm to obtain a corresponding evaluation result.

[0080] Furthermore, the BLEU evaluation algorithm is used to evaluate the similarity of the target question-answer pair and the target dialogue dataset. The calculation formula is as follows:

[0081] ;

[0082] Where a is the target question-answer pair, and a* is the target dialogue dataset. The ROUGE evaluation algorithm is used to evaluate the similarity of the target question-answer pair and the target dialogue dataset. The calculation formula is as follows:

[0083] ;

[0084] Where a is the target question-answer pair, and a* is the target conversation dataset. The cosine similarity evaluation algorithm is used to evaluate the similarity of the target question-answer pair and the target conversation dataset. The calculation formula is as follows:

[0085] ;

[0086] Among them, a is the target question-answer pair, and a* is the target dialogue dataset.

[0087] Step S13: determine whether the target question-answer pair meets the preset qualification conditions based on the matching evaluation result; if the target question-answer pair meets the preset qualification conditions, output the target question-answer pair that meets the preset qualification conditions.

[0088] In this embodiment, based on the matching evaluation result, it is determined whether the target question-answer pair meets the preset qualification conditions. If the target question-answer pair meets the preset qualification conditions, the target question-answer pair that meets the preset qualification conditions is output; if the target question-answer pair does not meet the preset qualification conditions, it jumps to the step of performing adversarial training and reinforcement learning optimization on the initial generator respectively through the preset policy gradient method and the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method. Specifically, after determining whether the target question-answer pair meets the preset qualification conditions based on the matching evaluation result, it also includes: if the target question-answer pair does not meet the preset qualification conditions, it jumps to the step of performing adversarial training and reinforcement learning optimization on the initial generator respectively through the preset policy gradient method and the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method.

[0089] It can be understood that whether the target question and answer pair meets the quality standard is judged based on the matching evaluation result. If the matching evaluation result indicates that the quality of the target question and answer pair meets the quality standard, the target question and answer pair that meets the quality standard is output; if the matching evaluation result indicates that the quality of the target question and answer pair does not meet the quality standard, the method jumps to the steps of performing adversarial training and reinforcement learning optimization on the initial generator respectively by using the preset policy gradient method and the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method, so as to further optimize the generator.

[0090] In a specific implementation, the output of the target question-answer pair that meets the preset qualification conditions may be as follows: User Question 1: What is the status of my order 12345? Customer Service Answer 1: Your order 12345 has been shipped and is expected to be delivered within 3 working days. User Question 2: When is my order expected to be delivered? Customer Service Answer 2: Your order is expected to be delivered within 3 working days. User Question 3: What if I don’t receive the goods? Customer Service Answer 3: If you don’t receive the goods within the scheduled time, please contact customer service in time and we will help you find and solve the problem.

[0091] As can be seen from the above, the present application uses the discriminator of the generative adversarial network structure to perform adversarial training on the generator, so that the initial generator can generate more realistic and natural target question and answer pairs, and then introduces the reward function of the reinforcement learning method to further optimize the initial generator. Then, the generator optimized by adversarial training and reinforcement learning is used to generate question and answer pairs based on the input user intention and dialogue context, and then a preset evaluation algorithm is used to evaluate the matching degree of the target question and answer pairs with the target dialogue data set. In this way, based on the matching evaluation results, the target question and answer pairs that meet the preset qualification conditions are output, which can significantly improve the efficiency and quality of multi-round question and answer data generation, avoid the limitations of artificial design, and ensure the wide coverage and realism of the generated data.

[0092] It can be seen from the above embodiments that the present application generates a target number of target question-answer pairs based on the learning optimized generator after adversarial training and reinforcement learning optimization. Therefore, the process of generating a target number of target question-answer pairs based on the learning optimized generator after adversarial training and reinforcement learning optimization is described.

[0093] See also Figure 2 As shown, the embodiment of the present invention discloses a specific method for generating multi-round question-answering data, including:

[0094] Obtain the user intention and conversation context provided by the user end, generate a series of initial question-answer pairs consisting of user questions and corresponding customer service answers based on the user intention and the conversation context using an initial generator, and evaluate the initial question-answer pairs using a discriminator of a generative adversarial network structure, that is, compare the initial question-answer pairs with real data; if the authenticity score obtained by comparison does not meet the preset authenticity condition, use the loss function of the discriminator to feed back the initial generator and perform adversarial training to obtain an adversarially trained generator, and then use the reward function of the reinforcement learning method to perform reinforcement learning optimization on the adversarially trained generator to obtain a learning-optimized generator;

[0095] If the authenticity score obtained by comparison meets the preset authenticity condition, a quality assessment is performed based on the context score, diversity score, and complexity score of the initial question and answer pair to obtain a corresponding quality assessment result. If the quality assessment result is that the initial question and answer pair meets the preset qualification condition, the initial question and answer pair with the preset qualification condition is output as the target question and answer pair; if the quality assessment result is that the initial question and answer pair does not meet the preset qualification condition, the adversarial training generator is subjected to reinforcement learning optimization through a preset policy gradient method and using the reward function of the reinforcement learning method to obtain a learning optimized generator.

[0096] Then, the generator after learning optimization is used to regenerate the target number of initial question-answer pairs based on user intent and conversation context, and jumps to the step of evaluating the initial question-answer pairs using the discriminator of the generative adversarial network structure until the target question-answer pairs that meet the preset qualification conditions are output. In the whole process, the loss function based on the discriminator and the loss function of reinforcement learning constitute the total loss function of the initial generator, and the initial generator is iteratively optimized based on the total loss function.

[0097] As can be seen from the above, this application constructs the total loss function of the initial generator based on the loss function of the discriminator and the loss function of reinforcement learning, and iteratively optimizes the initial generator based on the total loss function, and then iterates the optimized generator and generates a target question-answer pair that meets the preset qualification conditions based on the input user intent and dialogue context. In this way, the output target question-answer pair is high-quality multi-round question-answer test data, which not only improves the ability of multi-round question-answer generation, but also improves the efficiency and quality of multi-round question-answer data generation, and can cover various user intentions and dialogue scenarios.

[0098] Accordingly, see Figure 3 As shown, the present application also provides a multi-round question and answer data generation device, comprising:

[0099] A generator optimization module 11 is used to perform adversarial training and reinforcement learning optimization on the initial generator by using a preset policy gradient method and a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method, so as to obtain a generator after learning optimization; the generative adversarial network structure includes a generator and a discriminator;

[0100] A matching evaluation module 12 is used to generate a target number of target question-answer pairs based on the input user intent and dialogue context by the generator after learning optimization, and to perform matching evaluation on the target question-answer pairs and the target dialogue data set using a preset evaluation algorithm to obtain a corresponding evaluation result;

[0101] The question-answer pair judgment module 13 is used to judge whether the target question-answer pair meets the preset qualification conditions based on the matching evaluation result. If the target question-answer pair meets the preset qualification conditions, the question-answer pair that meets the preset qualification conditions is output.

[0102] As can be seen from the above, the present application uses the discriminator of the generative adversarial network structure to perform adversarial training on the generator, so that the initial generator can generate more realistic and natural target question and answer pairs, and then introduces the reward function of the reinforcement learning method to further optimize the initial generator. Then, the generator optimized by adversarial training and reinforcement learning is used to generate question and answer pairs based on the input user intention and dialogue context, and then a preset evaluation algorithm is used to evaluate the matching degree of the target question and answer pairs with the target dialogue data set. In this way, based on the matching evaluation results, the target question and answer pairs that meet the preset qualification conditions are output, which can significantly improve the efficiency and quality of multi-round question and answer data generation, avoid the limitations of artificial design, and ensure the wide coverage and realism of the generated data.

[0103] In some specific implementations, the generator optimization module 11 may specifically include:

[0104] An initial question-answer pair generating unit, configured to generate a plurality of initial question-answer pairs based on the input user intent and dialogue context by using the initial generator of the generative adversarial network structure;

[0105] An adversarial training unit is used to perform adversarial training on the initial generator through a preset policy gradient method and based on a discriminator of a generative adversarial network structure and the initial question-answer pair to obtain an adversarially trained generator, and to perform reinforcement learning optimization on the adversarially trained generator using a reward function of a reinforcement learning method to obtain a learning-optimized generator.

[0106] In some specific implementations, the generator optimization module 11 may specifically include:

[0107] An authenticity scoring unit, used to perform an authenticity scoring on the initial question-answer pair using the discriminator of the generative adversarial network structure to obtain a corresponding authenticity scoring result;

[0108] A loss function value determining unit, configured to determine a discriminator loss function value for characterizing the performance of the initial generator based on the loss function of the discriminator and a authenticity scoring result;

[0109] The generator optimization completion unit is used to characterize that the performance of the initial generator is poor if the discriminator loss function value is greater than a preset loss threshold, and to perform reinforcement learning optimization on the initial generator based on the discriminator loss function value and a preset policy gradient method and using a reward function of a reinforcement learning method to obtain a generator after learning optimization.

[0110] In some specific implementations, the generator optimization module 11 may specifically include:

[0111] a cosine similarity calculation unit, configured to obtain a context score of the initial question-answer pair by calculating the cosine similarity between the initial question-answer pair and the conversation context;

[0112] A semantic similarity calculation unit, configured to obtain a diversity score of the initial question-answer pairs by calculating the semantic similarity between any two of the initial question-answer pairs;

[0113] a text complexity determination unit, configured to determine the text complexity of the initial question-answer pair by statistically analyzing different text syntactic structures in the initial question-answer pair, and then obtain the vocabulary richness of the initial question-answer pair using a preset vocabulary richness index, and determine the complexity score of the initial question-answer pair based on the text length of the initial question-answer pair, the text complexity, the vocabulary richness, and the corresponding weighting coefficients;

[0114] A reward function value determination unit is used to determine the reward function value based on the context score, the diversity score, the complexity score and the corresponding weighting coefficients, and to perform reinforcement learning optimization on the initial generator through the reward function value and the preset policy gradient method to obtain a generator after learning optimization.

[0115] In some specific implementations, the matching evaluation module 12 may specifically include:

[0116] A feature screening unit is used to screen data features of manually annotated standard data sets and historical real dialogue data, and use the screened data to construct a target dialogue data set;

[0117] The matching evaluation unit is used to perform matching evaluation on the target question-answer pair and the target dialogue data set using a preset evaluation algorithm to obtain a corresponding evaluation result.

[0118] In some specific implementations, the matching evaluation module 12 may specifically include:

[0119] The matching evaluation completion unit is used to perform matching evaluation on the target question-answer pair and the target dialogue data set using a BLEU evaluation algorithm, a ROUGE evaluation algorithm, and a cosine similarity evaluation algorithm to obtain a corresponding evaluation result.

[0120] In some specific implementations, the multi-round question and answer data generating device may further include:

[0121] The target question-answer pair judgment unit is used to jump to the steps of performing adversarial training and reinforcement learning optimization on the initial generator respectively by using the preset policy gradient method and the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method if the target question-answer pair does not meet the preset qualification conditions.

[0122] Furthermore, the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the multi-round question and answer data generation method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0123] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0124] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0125] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the multi-round question and answer data generation method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0126] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the multi-round question-answering data generation method disclosed above is implemented. For the specific steps of the method, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.

[0127] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0128] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0129] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0130] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0131] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for generating multi-round question-answering data, characterized in that: include: By presetting the policy gradient method and using the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method, the initial generator is respectively subjected to adversarial training and reinforcement learning optimization to obtain the generator after learning optimization; The generative adversarial network structure includes a generator and a discriminator; Generate a target number of target question-answer pairs based on the input user intent and conversation context by the learned optimized generator, and use a preset evaluation algorithm to evaluate the matching degree of the target question-answer pairs and the target conversation dataset to obtain corresponding evaluation results; Based on the matching evaluation result, it is determined whether the target question-answer pair meets the preset qualification conditions. If the target question-answer pair meets the preset qualification conditions, the target question-answer pair that meets the preset qualification conditions is output.

2. The method for generating multi-round question-answer data according to claim 1, characterized in that: The method uses a preset policy gradient method and utilizes a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method to perform adversarial training and reinforcement learning optimization on the initial generator to obtain a generator after learning optimization, including: Generate several initial question-answer pairs based on the input user intent and conversation context through the initial generator of the generative adversarial network structure; The initial generator is adversarially trained by a preset policy gradient method and based on a discriminator of a generative adversarial network structure and the initial question-answer pair to obtain an adversarially trained generator, and the adversarially trained generator is reinforced learned and optimized using a reward function of a reinforcement learning method to obtain a learned optimized generator.

3. The method for generating multi-round question-answer data according to claim 2, characterized in that: The method performs adversarial training on the initial generator by using a preset policy gradient method and a discriminator based on a generative adversarial network structure and the initial question-answer pair to obtain an adversarially trained generator, and performs reinforcement learning optimization on the adversarially trained generator using a reward function of a reinforcement learning method to obtain a learning optimized generator, including: Using the discriminator of the generative adversarial network structure to perform authenticity scoring on the initial question-answer pair to obtain a corresponding authenticity scoring result; Determining a discriminator loss function value for characterizing the performance of the initial generator based on the discriminator loss function and the authenticity score result; If the discriminator loss function value is greater than the preset loss threshold, it indicates that the performance of the initial generator is poor. Based on the discriminator loss function value and the preset policy gradient method and using the reward function of the reinforcement learning method, the initial generator is subjected to reinforcement learning optimization to obtain a generator after learning optimization.

4. The method for generating multi-round question-answer data according to claim 3, characterized in that: The method of performing reinforcement learning optimization on the initial generator based on the discriminator loss function value and the preset policy gradient method and using the reward function of the reinforcement learning method to obtain a generator after learning optimization includes: Obtaining a context score for the initial question-answer pair by calculating the cosine similarity between the initial question-answer pair and the conversation context; Obtaining a diversity score of the initial question-answer pair by calculating the semantic similarity between any two of the initial question-answer pairs; Determine the text complexity of the initial question-answer pair by statistically analyzing different text syntactic structures in the initial question-answer pair, then obtain the lexical richness of the initial question-answer pair using a preset lexical richness index, and determine the complexity score of the initial question-answer pair based on the text length of the initial question-answer pair, the text complexity, the lexical richness, and the corresponding weighting coefficients; The reward function value is determined based on the context score, the diversity score, the complexity score and the corresponding weighting coefficients, and the initial generator is subjected to reinforcement learning optimization through the reward function value and the preset policy gradient method to obtain a generator after learning optimization.

5. The method for generating multi-round question-answer data according to claim 1, characterized in that: The using of a preset evaluation algorithm to perform a matching evaluation on the target question-answer pair and the target dialogue data set to obtain a corresponding evaluation result includes: We filter the data features of manually annotated standard data sets and historical real conversation data, and use the filtered data to build the target conversation data set. A preset evaluation algorithm is used to evaluate the matching degree of the target question-answer pair and the target dialogue data set to obtain a corresponding evaluation result.

6. The method for generating multi-round question-answer data according to claim 5, characterized in that: The using of a preset evaluation algorithm to perform a matching evaluation on the target question-answer pair and the target dialogue data set to obtain a corresponding evaluation result includes: The BLEU evaluation algorithm, the ROUGE evaluation algorithm, and the cosine similarity evaluation algorithm are used to evaluate the matching degree of the target question-answer pair and the target dialogue data set to obtain corresponding evaluation results.

7. The method for generating multi-round question-answer data according to any one of claims 1 to 6, characterized in that: After judging whether the target question-answer pair meets a preset qualification condition based on the matching evaluation result, the method further includes: If the target question-answer pair does not meet the preset qualification conditions, jump to the step of performing adversarial training and reinforcement learning optimization on the initial generator respectively by using the preset policy gradient method and utilizing the discriminator of the generative adversarial network structure and the reward function of the reinforcement learning method.

8. A multi-round question-answering data generating device, characterized in that: include: The generator optimization module is used to perform adversarial training and reinforcement learning optimization on the initial generator by using a preset policy gradient method and a discriminator of a generative adversarial network structure and a reward function of a reinforcement learning method, so as to obtain a generator after learning optimization; The generative adversarial network structure includes a generator and a discriminator; A matching evaluation module is used to generate a target number of target question-answer pairs based on the input user intent and dialogue context by the generator after learning optimization, and to perform matching evaluation on the target question-answer pairs and the target dialogue data set using a preset evaluation algorithm to obtain corresponding evaluation results; The question-answer pair judgment module is used to judge whether the target question-answer pair meets the preset qualification conditions based on the matching evaluation result. If the target question-answer pair meets the preset qualification conditions, the question-answer pair that meets the preset qualification conditions is output.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method for generating multi-round question and answer data as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the method for generating multi-round question and answer data as described in any one of claims 1 to 7 is implemented.