Multi-round dialogue synthesis method using real data

Through multiple rounds of dialogue backbone abstract technology, user intentions and context context are extracted from real data and high-quality multi-round dialogue data are generated, which solves the problems of data quality, diversity and context coherence in the existing technology, and achieves more natural and logically consistent dialogue data generation.

CN120144705AActive Publication Date: 2025-06-13INST OF SOFTWARE - CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510204012.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing multi-round dialogue data synthesis methods have data quality problems, insufficient dialogue diversity and context inconsistency, making it difficult to generate high-quality, natural and logically consistent dialogue data.

Method used

Using a method based on multi-round dialogue backbone abstraction, user intentions, key issues and context are extracted from real data to generate high-quality multi-round dialogue data. The method includes identifying the core information of the conversation, abstraction and transformation mechanisms, synthesis methods based on real usage scenarios, and constructing high-quality prompt words.

Benefits of technology

The data quality, consistency and nature in multiple rounds of dialogue generation have been significantly improved. The generated data is more natural and the subject is more consistent, and it can more accurately identify and respond to complex user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144705A_ABST
    Figure CN120144705A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-round dialogue synthesis method using real data. The method comprises the following steps: 1) designing a group of initial cue words; 2) inputting the selected dialogue data and the initial cue word into a large language model, and outputting a dialogue intention and interaction logic; (3) detecting whether the dialogue intention and the interaction logic meet set conditions or not, if not, adjusting the cue word, and inputting the cue word and the selected dialogue data into a large language model; (4) repeating the step (3) to obtain refined final prompt words meeting set conditions; 5) classifying dialogue intentions in the final cue word, extracting K types of dialogue intentions with the highest intention category, and generating a dialogue information flow and scene for each type of dialogue intentions; 6) generating a corresponding backbone template according to the dialogue information flow and scene of each type of dialogue intention; 7) filling a context corresponding to the backbone template; and 8) filling the specific content of each backbone template by using a large language model, generating a dialogue stream, verifying the dialogue stream, and carrying out self-consistency examination on the dialogue stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of data synthesis and relates to a multi-round dialogue synthesis method using real data. Background Art

[0002] In recent years, multi-round dialogue technology, as one of the key technologies in the field of large language models, has received widespread attention from the industry and academia. In advanced dialogue systems, whether in intelligent customer service, chatbots, virtual assistants, or professional scenarios such as education, games, and medical care, multi-round dialogue plays a core role, aiming to achieve a more natural and smoother interactive experience. However, training high-quality multi-round dialogue models essentially relies on a large amount of real user intentions, diverse and coherent dialogue data. The synthesis of multi-round dialogue data is not only to solve the problem of high cost of obtaining dialogue data, but its core lies in the following key significance:

[0003] 1. Improve the capabilities of the dialogue system

[0004] Multi-round conversation data can help the model better understand contextual information, form global cognition and understanding capabilities, and enhance the ability to handle complex conversation situations. Real user intention modeling and scenario simulation data can enrich the model's performance in different industries and scenarios.

[0005] 2. Diversity and generalization

[0006] Synthetic data can avoid possible biases in real data, expand the contextual diversity of the dialogue model, and improve the generalization ability to unseen scenarios.

[0007] 3. Adapt to the characteristics of different fields

[0008] In specific industries (such as medical consultation, educational counseling, etc.), multiple rounds of dialogue need to cover domain-specific knowledge. Synthetic data can generate relevant scenario dialogues in a targeted manner, reducing the cost of relying on manual labeling.

[0009] 4. Save data collection costs

[0010] Traditional methods of collecting user conversation data usually require a lot of manpower and time. Synthesizing multiple rounds of data can significantly reduce data collection and annotation costs and improve R&D efficiency.

[0011] The existing large language model mainly includes the following three technical solutions in the multi-round dialogue data collection stage:

[0012] 1. User-generated conversation data

[0013] This type of data comes from the queries raised by users and is gradually collected in multiple rounds of interactions after the system gives responses. The data synthesized by this method is a true reflection of real user intentions, especially rich, natural, and diverse query methods and problem scenarios. However, the data collection depends on user interactions, with limited scale and diversity. Moreover, due to the uncontrollability of actual user behaviors, the generated dialogue data may be mixed with a large amount of noise and is difficult to directly use for training models.

[0014] 2. Synthetic high-quality data

[0015] Manually write a small amount of high-quality question-and-answer data to fine-tune the model, and use the fine-tuned model to generate a batch of relatively high-quality dialogue data. This method enhances the model's dialogue ability through a small amount of high-quality dialogue data. This approach relies on manual work and is difficult to scale up; although the synthesized question-and-answer pairs are of relatively high quality, the generated multi-turn dialogues have poor controllability and may not reflect the complexity and diversity brought by real user intentions.

[0016] 3. Generate data based on context

[0017] Some dialogue frameworks will, based on a given initial question, predict the questions that users may ask in future interactions through the model and continue to generate responses for multiple rounds. This can better simulate the context continuity of multi-turn dialogues in actual applications; however, the global quality control of the generated results is weak, and there may be problems such as logical inconsistency or derailment in the dialogue flow. The dialogue topics generated by the model may lack diversity and are prone to repetition or being patternized.

[0018] Through the above analysis, the main problems existing in the existing solutions can be determined as follows:

[0019] 1. Data quality problems

[0020] When maintaining naturalness and high quality, the existing generation methods will inevitably introduce noisy data. Some data generation methods rely too much on the existing capabilities of the model, which may lead to problems such as illogical output and unnatural expressions, ultimately affecting the effective learning of the model.

[0021] 2. Insufficient dialogue diversity

[0022] The synthesized data often has the problem of insufficient diversity. Especially in multi-turn dialogues, the generated dialogues are prone to being rigid or patternized and are difficult to meet the diverse problems and behaviors of users in actual scenarios.

[0023] 3. Context incoherence

[0024] Multi-turn conversations need to ensure the consistency of context logic and the persistence of the topic. However, existing methods often fail to accurately retain the conversation topic, and the generated content may deviate or stray from the initial question. When dealing with complex intention switching and non-linear conversations, the model lacks the ability of global understanding, resulting in poor conversation continuity. Summary of the Invention

[0025] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a multi-turn conversation synthesis method using real data. The present invention is a new multi-turn conversation data synthesis method, which utilizes the multi-turn conversation backbone abstraction technology extracted from real data to enhance the response ability of the model in multi-turn conversation scenarios.

[0026] The key points of the present invention are as follows:

[0027] 1. Propose a method based on multi-turn conversation backbone abstraction. Generate high-quality conversation data with different intention classifications from a batch of unlabeled data. In the collection of general multi-turn conversation data, it mainly relies on user-generated data or reverse-generated data, which brings problems of noise or unnaturalness. The backbone abstraction method includes identifying core information in the conversation such as user intention, key questions, and context, so as to generate coherent multi-turn data suitable for high-quality conversation generation.

[0028] 2. Propose an abstraction and transformation mechanism for conversation backbones. By abstracting the user intention, system response, and implicit context in multi-turn conversations, each turn of the conversation can be simplified into a basic structure, thereby reasonably retaining key information points and ensuring the compactness and coherence of the entire multi-turn conversation.

[0029] 3. A multi-turn conversation synthesis method based on real usage scenarios. Through backbone abstraction of real conversations, such as the lmsys-chat-1m dataset, and combining with the actual scenario to synthesize conversation data, simulating complex and real multi-turn interactions of users. This step is not limited to the generation of high-quality responses, but also will simulate continuous questions from real users, making the prompts more in line with the behavior habits of users in reality.

[0030] 4. Construct high-quality prompt words. Use a large language model with sufficient capabilities as a referee to analyze the intentions and patterns in multi-turn conversations in the dataset in the early stage.

[0031] The technical solution of the present invention is as follows:

[0032] A multi-turn conversation synthesis method using real data, the steps of which include:

[0033] 1) Design a set of initial prompt words for the large language model to analyze multi-turn conversation intentions and interaction logic;

[0034] 2) inputting the selected real user conversation data and the initial prompt words into the large language model, and outputting the conversation intention and interaction logic of the selected real user conversation data;

[0035] 3) Detect whether the currently outputted dialogue intent and interaction logic meet the set conditions. If not, adjust the prompt words; then input the adjusted prompt words and the selected real user dialogue data into the large language model, and output the dialogue intent and interaction logic of the selected real user dialogue data;

[0036] 4) Repeat step 3) until the dialogue intent and interaction logic output by the large language model meet the set conditions, and obtain the refined final prompt words;

[0037] 5) Classify the conversation intents in the refined final prompt words, extract the K conversation intents with the highest intent categories, and generate a corresponding conversation information flow and random scene for each of the K conversation intents;

[0038] 6) Generate a backbone template of the corresponding dialogue intent according to the dialogue information flow and random scene of each dialogue intent in the K types of dialogue intents;

[0039] 7) Fill in the context corresponding to each backbone template using the data in the selected knowledge base;

[0040] 8) Use the large language model to fill in the specific content of each backbone template to generate a dialogue flow; verify the generated content based on the contextual logic consistency and thematic consistency; after the verification, use the large language model as a referee to conduct a self-consistency review on each generated dialogue flow; if the self-consistency review passes, it will be used as the dialogue flow of the corresponding category dialogue intent of the backbone template, otherwise it will be iteratively revised.

[0041] Furthermore, interaction backbone information is extracted from the dialogue information flow and random scenarios of each type of dialogue intention as a backbone template; the interaction backbone information includes the user's main intention, key issues, typical interaction patterns and necessary context.

[0042] Furthermore, the setting condition is that the dialogue intention and interaction logic can stably and high-quality describe the characteristics of multi-round dialogues.

[0043] Furthermore, the categories of conversation intentions include: problem-solving interaction, educational interaction, health interaction, exploratory interaction, entertainment interaction, simulation interaction, emotional support interaction, information retrieval interaction and transactional interaction.

[0044] Furthermore, the selected knowledge base is a Wiki knowledge base.

[0045] Further, the initial prompt is: Please analyze the main intention, interaction structure, and logical clarity of the following conversation,..., and output a summary: [Conversation text].

[0046] Further, use the large language model to extract the conversation information flow and the backbone template of the random scenario for each type of conversation intention in the K types of conversation intentions.

[0047] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the above method.

[0048] A computer-readable storage medium, on which a computer program is stored, characterized in that the computer program realizes the above method when executed by a processor.

[0049] The advantages of the present invention are as follows:

[0050] The present invention combines the innovative multi-turn conversation backbone abstraction technology, significantly improving the data quality, coherence, and naturalness of the large language model in multi-turn conversation generation. By extracting the conversation backbone, the generated data not only maintains the context structure of the conversation but also highlights the core information of the user's intention. The trained model can more accurately identify and respond to complex user needs, thereby generating responses more suitable for actual application scenarios.

[0051] The present invention reduces the gap between synthetic conversation data and real user interaction data, and the generated synthetic conversation data is more natural and has higher topic consistency. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] The present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0054] I. Cleaning and preprocessing real user conversation data

[0055] Collect multi-turn conversation data of real users, such as lmsys-chat-1m, wildchat. First, preprocess the data to clean possible noises. The specific method is as follows:

[0056] 1. Remove obviously invalid or low-quality data

[0057] Delete some conversation data with non-standard and incomplete formats, such as those with only single-round interactions or discontinuous conversation flows, and data that only contains user inputs but no model responses. Exclude blank, garbled, overly repetitive, or obviously meaningless text content, such as pure symbol inputs, no keyboard strokes, and obvious noise.

[0058] 2. Filter inappropriate content

[0059] Through keyword filtering and natural language discrimination methods, clean the content in the conversation that contains discriminatory, violent, offensive language, or sensitive information (such as account numbers, passwords, names, etc.) to ensure the ethics and generality of the dataset.

[0060] 3. Remove conversations with disjointed context or chaotic logic

[0061] For multi-round conversations, by detecting content with discontinuous and incoherent context relationships, clean conversations with obviously contradictory semantic effects or chaotic logic. For example, the user's question and the model's answer are obviously irrelevant or off-topic.

[0062] 4. Standardize the format and correct potential errors

[0063] Standardize the basic format of the data (such as usernames, reply paragraph separation, punctuation) to ensure the clarity and consistent readability of the conversation text. Use spelling and grammar correction tools (such as model-based or rule-based language tools) to correct spelling mistakes and language problems in the conversation.

[0064] II. Construct prompt words and use large language models for analysis

[0065] Input data: Cleaned real user conversation data.

[0066] Specific process:

[0067] 1. Initial prompt word design

[0068] Design a set of initial prompt words for analyzing the intent and interaction logic of multi-round conversations through manual or model-based methods. For example: Please analyze the main intent, interaction structure, and logical clarity of the following conversation,..., and output a summary in the required format (such as a structured json format, containing the values corresponding to the current conversation intent, interaction structure, and logical clarity, for subsequent analysis and clustering): [Conversation text]. The goal of the prompt words is to enable the large language model to accurately summarize the intent, logic, and key issues of multi-round conversations.

[0069] 2. Iterative process of improving prompt words

[0070] According to the analysis results of the dialogue intent and interaction logic output in the previous round, the prompt words are adjusted manually or through the model. If the previous round is not comprehensive enough, more explicit analysis instructions can be added. If the model in the previous round did not recognize the context information, the prompt words can be adjusted to describe the context dependency requirements.

[0071] The prompt words are iterated until the result can cover most common conversation intentions, and the conversation intention analysis generated by the model output is logically self-consistent and clearly expressed.

[0072] 3. Stop Condition

[0073] When determining the stability and high-quality description capabilities of the model output, we not only adopted quantitative evaluation criteria (intent accuracy, interaction logic consistency, and stability), but also combined the research results in a systematic review paper in the field of human-computer interaction (Rapp et al., 2021, The human side of human-chatbot interaction: A systematic literature review of ten years of research on text-based chatbots). The paper summarizes the common conversation intentions of humans when interacting with chatbots. Starting from a systematic review of human-computer interaction, the present invention uses a large model to analyze the conversation intentions of multiple multi-round conversation data sets, performs cluster analysis, and analyzes several intentions with the highest statistical proportion, and matches them with the intentions mentioned in the paper, that is, the present invention pre-sets the conversation intentions of real user conversation data, and after finally determining multiple conversation intentions, designs the interaction logic (conversation information flow) that needs to be met in multi-round conversations for each conversation intention.

[0074] The specific optimization process of the present invention is as follows:

[0075] Comprehensive multi-round dialogue dataset: Analyze multiple multi-round dialogue datasets, filter dialogue instances with high intent, and use a large model for intent labeling and logical reasoning analysis.

[0076] Cluster analysis and manual screening: Clustering is performed based on multiple analysis results, and domain experts manually screen high-quality conversation intents that are common to humans and account for a large proportion.

[0077] Comparison with intents in review papers: Compare and analyze the selected intent categories with the research results in the paper by Rapp et al. (2021) to ensure high coverage, and finally summarize nine common human conversation intents as intent labels corresponding to multi-round conversations.

[0078] Information flow design and rule formulation: Based on the common intents summarized, design a reasonable dialogue information flow for each intent to conform to the natural interaction logic, and use it to guide the optimization of the prompt words of the subsequent large model and the formulation of distillation rules. That is, after obtaining nine user intents, let the large model generate the scenario topics that each dialogue intent may contain, such as generating 500 dialogue scenarios for each dialogue intent. The dialogue intent and the dialogue information flow are the same inputs each time the distillation model is performed, and at the same time, the generated scenario topics are provided. For each dialogue intent, more than 500 rounds of dialogue scenarios can be obtained.

[0079] When all of the following conditions are met, terminate the optimization iteration: intent recognition accuracy ≥ 90%; interactive logic consistency score ≥ 85%; output stability (the model can be analyzed normally and output results as required) ≥ 85%. All nine common human dialogue intents are effectively covered, and a reasonable dialogue information flow is designed for each intent.

[0080] The finally output prompt words can not only ensure high-accuracy intent recognition and logical analysis capabilities, but also comprehensively cover common human dialogue intents, providing key support for formulating rules for subsequent distillation of large models.

[0081] When the output dialogue intent and logical analysis results can stably and qualitatively describe the characteristics of multi-round dialogues, stop the iteration. The finally output prompt words clearly include the requirements for consistency checks of multiple intent classifications and context logic.

[0082] Output results: refined finally output prompt words and the intent and logical analysis results of multi-round dialogues, providing key inputs for the backbone abstraction method.

[0083] III. Summarize high-frequency dialogue intents and extract interaction backbones

[0084] Input data: multi-round dialogue intents, corresponding dialogue information flows, and random scenarios that conform to the dialogue intents in the analysis results.

[0085] Specific steps:

[0086] 1. The model counts high-frequency intents

[0087] Classify the analyzed multi-round dialogues according to intent labels, and extract the frequently occurring intent categories (such as consulting questions, requirement confirmation, error handling, etc.). Use the large language model assisted method to summarize the commonalities of each intent category.

[0088] 2. Extract the interaction backbone

[0089] Extract the interactive backbone information based on the statistically identified high-frequency intent classifications, the dialogue information flow within each intent category, and random scenarios. The backbone includes the user's main intent, key questions, typical interaction patterns, and necessary context. For example: User asks a question → System confirms the question scope → User clarifies the intent → System provides a specific answer.

[0090] Output: A backbone template containing high-frequency dialogue intent classifications and interaction logic, providing a core example for the backbone abstraction method.

[0091] IV. Generate multi-turn dialogues by combining the backbone abstraction method with external data such as Wiki

[0092] Input data: The dialogue backbone template in the analysis results, and external knowledge base expansion data such as Wiki.

[0093] Specific steps:

[0094] 1. Integrate external knowledge data

[0095] Use Wiki data to fill in the specific context in the backbone template, such as generating answer elements for specific questions in the intent classification.

[0096] 2. Generate multi-turn dialogue data based on the backbone

[0097] For each backbone template, use a large model to fill in the specific content. Verify the generated content based on the logical consistency and theme consistency of the context to ensure that the generated dialogue flow remains coherent and conceptually consistent in multiple turns.

[0098] 3. Automatically adjust and optimize

[0099] Use a large language model as a referee to conduct self-consistency review on each generated dialogue flow to ensure no logical contradictions. Automatically screen or mark the dialogues that do not conform to the definition of the logical backbone and incorporate them into subsequent correction and iteration.

[0100] Output: A batch of generated multi-turn dialogue data with high theme consistency and coherent dialogue logic.

[0101] According to the above experimental steps, nine common dialogue intents were collected and the corresponding dialogue information flows were compiled. The information flow reflects the development trend and direction of each multi-turn dialogue. Finally, according to different interactive intent categories, different information flows were matched as the Prompt templates for constructing multi-turn dialogue instructions in the future.

[0102]

[0103]

[0104] V. Output a multi-turn dialogue dataset with high theme consistency

[0105] Final result: Combine the cleaned real user data, the backbone template, and the multi-round conversations generated by the large model to form a high-quality conversation dataset.

[0106] This dataset has the following characteristics:

[0107] 1. The multi-round conversation topics are continuous and logically consistent.

[0108] 2. It can effectively cover various intent classifications and has scalability.

[0109] 3. The data has been reviewed and screened, and has high naturalness and context coherence.

[0110] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.

Claims

1. A method for synthesizing multi-round dialogues using real data, the steps comprising: 1) Design a set of initial prompt words for the large language model to analyze multi-round dialogue intent and interaction logic; 2) inputting the selected real user conversation data and the initial prompt words into the large language model, and outputting the conversation intention and interaction logic of the selected real user conversation data; 3) Check whether the currently output dialogue intention and interaction logic meet the set conditions. If not, adjust the prompt words; Then the adjusted prompt words and the selected real user conversation data are input into the large language model, and the conversation intention and interaction logic of the selected real user conversation data are output; 4) Repeat step 3) until the dialogue intent and interaction logic output by the large language model meet the set conditions, and obtain the refined final prompt words; 5) Classify the conversation intent in the refined final prompt words and extract the K conversation intents with the highest intent categories. And generate a corresponding dialogue information flow and random scene for each of the K types of dialogue intentions; 6) Generate a backbone template of the corresponding dialogue intent according to the dialogue information flow and random scene of each dialogue intent in the K types of dialogue intents; 7) Fill in the context corresponding to each backbone template using the data in the selected knowledge base; 8) Use the large language model to fill in the specific content of each backbone template to generate a dialogue flow; verify the generated content based on the contextual logic consistency and thematic consistency; after the verification, use the large language model as a referee to conduct a self-consistency review on each generated dialogue flow; if the self-consistency review passes, it will be used as the dialogue flow of the corresponding category dialogue intent of the backbone template, otherwise it will be iteratively revised.

2. The method according to claim 1, characterized in that Interaction backbone information is extracted from the dialogue information flow and random scenes of each type of dialogue intention as a backbone template; the interaction backbone information includes the user's main intention, key questions, typical interaction patterns and necessary context.

3. The method according to claim 1, characterized in that The setting condition is that the dialogue intention and interaction logic can stably and high-quality describe the characteristics of multi-round dialogues.

4. The method according to claim 1, 2 or 3, characterized in that: The categories of conversation intentions include: problem-solving interaction, educational interaction, health interaction, exploratory interaction, entertainment interaction, simulation interaction, emotional support interaction, information retrieval interaction and transactional interaction.

5. The method according to claim 1, 2 or 3, characterized in that: The selected knowledge base is a Wiki knowledge base.

6. The method according to claim 1, 2 or 3, characterized in that: The initial prompt words are: Please analyze the main intention, interaction structure and logical clarity of the following dialogue, ..., and output a summary: [Conversation text].

7. The method according to claim 1, 2 or 3, characterized in that: The large language model is used to extract the dialogue information flow and backbone templates of random scenes for each of the K types of dialogue intentions.

8. A server, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 8.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Prompt word updating method, device and system based on multi-dimensional large language model

    CN118132716A

  • Prompt word generation method of large language model and question answering method of large language model

    CN119168062A

  • Computing process flow control via determination of dialogue context between a user and an artificial intelligence assistant

    US12141539B1

  • Method and system for generating intent responses through virtual agents

    US20230350929A1

Cited By

  • Database natural language query autoregression intention classification method and system

    CN120995178A