Training dataset generation, dialogue processing method, device, medium and product

CN122132833APending Publication Date: 2026-06-02BEIKE TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIKE TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-06-02

Smart Images

  • Figure CN122132833A_ABST
    Figure CN122132833A_ABST
Patent Text Reader

Abstract

This disclosure relates to a training dataset generation, dialogue processing method, device, medium, and product. The method includes: acquiring multiple first dialogue data samples from a target application domain; each first dialogue data sample includes input content and its corresponding response content; inputting the first dialogue data samples into a target generative model to obtain corresponding inference steps; the inference steps characterize the inference process of generating response content from the input content; integrating the first dialogue data samples and their corresponding inference steps to obtain a thought chain data sample corresponding to the first dialogue data sample; the thought chain data sample is data with a triple structure including sequentially arranged input content, inference steps, and response content; and constructing a training dataset corresponding to the target application domain based on each first dialogue data sample and each thought chain data sample. This improves the efficiency of acquiring thought chain data samples for the target application domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a training dataset generation, dialogue processing method, device, medium, and product. Background Technology

[0002] Many existing applications or agents based on large-scale machine learning models (referred to as large models) involve multi-step reasoning processes when dealing with some complex tasks, which is called the thought chain technology.

[0003] Current thought chain technology is mainly applied to general domains. Large models based on it are prone to logical errors (such as omitting reasoning factors), illusion problems (such as referencing non-existent factors), and difficulties in context tracking (such as changes in requirements during multi-turn dialogues) when handling more complex tasks. These problems become increasingly severe when applied to specialized domains due to the dynamic nature and diverse requirements of those domains. Therefore, there is an urgent need for large models based on thought chain technology that can be better applied to specialized domains. However, training such large models requires collecting a large amount of manually labeled thought chain data, making the training process time-consuming and labor-intensive. Summary of the Invention

[0004] To address the aforementioned technical issues, this disclosure provides a training dataset generation method, dialogue processing method, device, medium, and product.

[0005] In a first aspect, embodiments of this disclosure provide a method for generating a training dataset, the method comprising: Acquire multiple first dialogue data samples in the target application domain; wherein, the first dialogue data sample includes input content and response content corresponding to the input content; The first dialogue data sample is input into the target generative model to obtain the reasoning steps corresponding to the first dialogue data sample; wherein, the reasoning steps are used to characterize the reasoning process of generating the response content from the input content; the target generative model includes a model that has the ability to generate reasoning steps in thought chain data, obtained by fine-tuning a pre-trained initial generative model. By integrating the first dialogue data sample and the reasoning steps corresponding to the first dialogue data sample, a thought chain data sample corresponding to the first dialogue data sample is obtained; wherein, the thought chain data sample is data comprising a triple structure of the input content, the reasoning steps and the response content arranged in sequence; Based on each of the first dialogue data samples and each of the thought chain data samples, a training dataset corresponding to the target application domain is constructed.

[0006] Secondly, embodiments of this disclosure provide a dialogue processing method, the method comprising: Receive dialogue input; Based on the dialogue input content, a dialogue processing model is invoked to obtain the dialogue response content corresponding to the dialogue input content, or to obtain the dialogue inference steps and the dialogue response content corresponding to the dialogue input content; wherein, the dialogue processing model is obtained by fine-tuning a preset generative model using a training dataset, and the training dataset is obtained using the training dataset generation method described in any embodiment of this disclosure.

[0007] Thirdly, embodiments of this disclosure also provide an apparatus for generating a training dataset, the apparatus comprising: The first dialogue data sample acquisition module is used to acquire multiple first dialogue data samples in the target application domain; wherein, the first dialogue data sample includes input content and response content corresponding to the input content; The reasoning step acquisition module is used to input the first dialogue data sample into the target generative model to obtain the reasoning steps corresponding to the first dialogue data sample; wherein, the reasoning steps are used to characterize the reasoning process of generating the response content from the input content; the target generative model includes a model that has the ability to generate reasoning steps in thought chain data, obtained by fine-tuning the pre-trained initial generative model. The thought chain data sample acquisition module is used to integrate the first dialogue data sample and the reasoning steps corresponding to the first dialogue data sample to obtain the thought chain data sample corresponding to the first dialogue data sample; wherein, the thought chain data sample is data comprising a triple structure of the input content, the reasoning steps and the response content arranged in sequence; The training dataset construction module is used to construct a training dataset corresponding to the target application domain based on each of the first dialogue data samples and each of the thought chain data samples.

[0008] Fourthly, embodiments of this disclosure also provide a dialogue processing apparatus, the apparatus comprising: The dialogue input content receiving module is used to receive dialogue input content; The dialogue response content acquisition module is used to call the dialogue processing model based on the dialogue input content to obtain the dialogue response content corresponding to the dialogue input content, or to obtain the dialogue inference steps and the dialogue response content corresponding to the dialogue input content; wherein, the dialogue processing model is obtained by fine-tuning a preset generative model using a training dataset, and the training dataset is obtained using the training dataset generation method described in any embodiment of this disclosure.

[0009] Fifthly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: Processor and memory; The processor executes the training dataset generation method described in any embodiment of this disclosure, or the dialogue processing method described in any embodiment of this disclosure, by calling the program or instructions stored in the memory.

[0010] In a sixth aspect, embodiments of this disclosure also provide a computer-readable storage medium storing a program or instructions that cause a computer to perform a training dataset generation method described in any embodiment of this disclosure, or a dialogue processing method described in any embodiment of this disclosure.

[0011] In a seventh aspect, embodiments of this disclosure also provide a computer program product for executing the training dataset generation method described in any embodiment of this disclosure, or for executing the dialogue processing method described in any embodiment of this disclosure.

[0012] The training dataset generation method, device, storage medium, and program product provided in this disclosure can acquire multiple first dialogue data samples, including input content and response content, in a target application domain; input the first dialogue data samples into a target generative model to obtain inference steps corresponding to the first dialogue data samples; these inference steps characterize the inference process of generating the response content from the input content; integrate the first dialogue data samples and the inference steps corresponding to the first dialogue data samples to obtain thought chain data samples corresponding to the first dialogue data samples; and construct a training dataset corresponding to the target application domain based on each of the first dialogue data samples and each of the thought chain data samples. This achieves the automatic generation of a large number of proprietary thought chain data samples for the target application domain, greatly reducing the manpower and time consumption in the process of collecting training samples, thereby improving the efficiency of acquiring thought chain data samples and providing a data foundation for obtaining a large model with proprietary logical reasoning capabilities for the target application domain. Furthermore, by obtaining a training dataset that mixes first dialogue data samples and thought chain data samples, the diversity of the training dataset for the large model can be further improved, thereby further providing a data foundation for the large model to effectively balance the model's reasoning ability and illusion suppression ability.

[0013] The dialogue processing method, device, storage medium, and program product provided in this disclosure can use the aforementioned generated hybrid training dataset to fine-tune a pre-trained generative model to obtain a dialogue processing model. This model is then used to process received dialogue input to obtain either dialogue response content or dialogue reasoning steps and the corresponding dialogue response content. This allows the dialogue processing model to additionally learn logical reasoning patterns specific to the target application domain, effectively balancing the model's reasoning ability and illusion suppression to achieve hybrid reasoning capabilities. Specifically, when handling complex tasks, it can trigger reasoning capabilities to generate thought chain data, improving the logical rationality, accuracy, and interpretability of the response content. Conversely, when handling simple tasks, it can directly respond / reply, preserving the fluency of the dialogue. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0015] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart illustrating a method for generating a training dataset according to an embodiment of this disclosure; Figure 2 A flowchart illustrating a dialogue processing method provided in an embodiment of this disclosure; Figure 3 A schematic diagram of the structure of a training dataset generation device provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of a dialogue processing device provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0017] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be described in further detail below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0018] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0019] The training dataset generation method provided in this disclosure is mainly applicable to scenarios where a training dataset for large model training is to be obtained, which contains at least some thought chain data samples within a specific domain. This training dataset generation method can be executed by a training dataset generation device, which can be implemented in software and / or hardware. This device can be integrated into an electronic device with certain data processing capabilities, such as a laptop, desktop computer, server, or server cluster.

[0020] Figure 1 This is a flowchart illustrating a method for generating a training dataset according to an embodiment of this disclosure. See also... Figure 1 The specific methods for generating this training dataset include: S110. Obtain multiple first dialogue data samples in the target application domain.

[0021] The target application domain is the application domain for which a large model supporting the thought chain technology is to be acquired. For example, the target application domain could be one characterized by policy / market dynamism, diverse demands, high risk, and complex processing tasks. The first dialogue data sample is dialogue data obtained from the complete set of information generated from a single dialogue session (referred to as dialogue session data), used for model training. It can be single-turn or multi-turn dialogue data. In this embodiment, the first dialogue data sample includes input content and corresponding response content. For example, in a question-and-answer scenario, the first dialogue data sample includes the question and response content from at least one round of dialogue.

[0022] Specifically, the electronic device first collects raw, relevant data from the target application domain as the foundation for automatically generating thought chain data. Given that the large model's functionality is primarily implemented through dialogue, the electronic device collects dialogue data from various business scenarios within the target application domain, especially dialogue data from business scenarios with high task complexity. Then, the electronic device can clean the collected dialogue data, removing irrelevant or low-quality data and retaining relatively complete, high-value dialogue data as the first dialogue data sample.

[0023] In some embodiments, if the target application domain includes the real estate domain, then S110 includes: extracting multiple raw dialogue data from the dialogue session data corresponding to the real estate business scenario in the real estate domain; cleaning each raw dialogue data to obtain each first dialogue data sample.

[0024] Among them, the real estate business scenarios include at least one of the following: property introduction scenario (introducing properties to users through dialogue), property recommendation scenario (recommending properties that meet users' needs through dialogue), tax prediction scenario (predicting the taxes and fees payable for property transactions through dialogue), and loan consultation scenario.

[0025] Specifically, the real estate sector is characterized by deep integration of policies, high dynamism in policy and market changes, highly diverse user needs, and high risk. Many of its business scenarios are highly complex and domain-specific. For example, in tax prediction, large models need to consider multiple factors such as property price, size, and local policies; similarly, property recommendation needs to match user budget, property location, floor, and surrounding amenities. These highly complex business scenarios are not suitable for large models to directly output responses. Instead, multi-step reasoning is required to guide the model to gradually decompose and solve complex tasks, reducing the risk of model jumps and illusions, and ensuring the accuracy and interpretability of the output. Meanwhile, general domain thinking chain technologies (such as mathematical reasoning and logical question answering) lack the ability to deeply integrate real estate policies, market changes, and diverse needs. Therefore, electronic devices can automatically generate and process dedicated thinking chain data for the real estate sector.

[0026] Based on the foregoing explanation, electronic devices can extract at least one round of raw dialogue data from various dialogue sessions corresponding to highly complex real estate business scenarios. Then, this raw dialogue data is cleaned to obtain first dialogue data samples within the real estate sector. This provides a data foundation for the subsequent automatic generation of thought chain data within the real estate sector.

[0027] S120. Input the first dialogue data sample into the target generative model to obtain the reasoning steps corresponding to the first dialogue data sample.

[0028] The target generative model is a generative model obtained by fine-tuning a pre-trained initial generative model, possessing the ability to generate reasoning steps from thought chain data. For example, the target generative model is a generative model capable of generating reasoning steps from thought chain data in a target application domain. The initial generative model here can be, for example, a large language model whose model performance evaluation metrics (such as accuracy, recall, etc.) meet the corresponding threshold. The thought chain data here refers to data containing a complete chain of thought processes, including input content, reasoning steps, and output content (also known as response content).

[0029] The above reasoning steps are used to characterize the reasoning process of generating response content from input content. They can include the sequential execution of operation processes such as feature parsing based on input content, logical association derivation, and response content generation and verification, and are used to reproduce the complete reasoning process of the target generative model from input content to response content. For example, the reasoning steps are text processing logic adapted to the large language model, used to concretely characterize the step-by-step execution operation of its reasoning process of obtaining response content from input content. For example, it can include at least the following operation steps: (1) Input parsing operation - extracting and quantifying semantic features, core demands and contextual association information of input content; (2) Logical derivation operation - constructing a logical chain from input content to response content according to preset language reasoning rules based on the extracted information; (3) Content generation operation - generating response content that conforms to the output specifications of the large language model according to the constructed logical chain.

[0030] Specifically, in this embodiment of the disclosure, a target generative model can be obtained in advance. Then, each of the aforementioned first dialogue data samples is input into the target generative model, and the model processes the data to output the reasoning steps corresponding to each first dialogue data sample.

[0031] In some embodiments, S120 includes: inputting a first dialogue data sample into a target generative model to obtain an initial inference step corresponding to the first dialogue data sample; performing logical compliance filtering and / or domain relevance filtering on each initial inference step to obtain an inference step corresponding to the first dialogue data sample.

[0032] The initial inference steps are the inference steps directly output by the target generative model. Logical compliance filtering filters out inference steps with logical errors. Domain relevance filtering filters out inference steps irrelevant to the target application domain.

[0033] Specifically, electronic devices can obtain the initial inference steps corresponding to each first dialogue data sample using a target generative model. However, these initial inference steps may contain logical errors or domain irrelevance issues due to the uncertainty of the target generative model. Therefore, electronic devices can utilize filtering methods from related technologies to perform logical compliance filtering and / or domain relevance filtering on each initial inference step, filtering out initial inference steps that do not meet the filtering rules and obtaining high-quality inference steps. In this way, some first dialogue data samples will have corresponding inference steps, while others will not. This improves the quality and accuracy of the inference steps, thereby improving the accuracy of the subsequently obtained training dataset.

[0034] In some embodiments, if the first dialogue data sample includes multi-turn dialogue data in a dialogue session, then S120 includes: inputting the first dialogue data sample into the target generative model to obtain the local inference steps corresponding to each turn of dialogue data; cleaning each local inference step to obtain the inference steps corresponding to the first dialogue data sample.

[0035] The cleaning process includes at least one of deduplication, logical compliance filtering, and domain relevance filtering.

[0036] Specifically, considering that single-turn dialogues have certain logical inconsistencies and cannot accurately reflect user intent and its changes, and that actual dialogue processes are mainly logically coherent multi-turn dialogues, this embodiment of the disclosure can obtain a first dialogue data sample containing multi-turn dialogue data. That is, each first dialogue data sample contains multi-turn dialogue data of a certain number of rounds, rather than single-turn dialogue data. Thus, when cleaning the original dialogue data in S110, a filtering condition for the number of dialogue rounds can be added to filter out dialogue data with fewer than this filtering condition. This allows the large model to learn the logical structure and rhythm of the dialogue through multi-turn dialogue data, and enhances the large model's memory of the context, thereby further improving the logical rationality of the reasoning steps.

[0037] Building upon the above, for each round of dialogue data in each first dialogue data sample, the electronic device can generate its corresponding inference steps, called local inference steps, using a target generative model. The number of local inference steps generated corresponds to the number of rounds of dialogue data contained in a first dialogue data sample. Then, considering the potential for repetition or partial overlap between multiple rounds within the same dialogue session, the electronic device can perform deduplication on the multiple local inference steps corresponding to each first dialogue data sample to filter out those reaching a certain degree of repetition. Next, for the remaining local inference steps, the electronic device can further perform logic compliance filtering and / or domain relevance filtering to filter out local inference steps with logical errors and those irrelevant to the target application domain. Finally, the multi-round dialogue data and the filtered local inference steps together constitute the inference steps corresponding to the first dialogue data sample.

[0038] In some embodiments, prior to S120, the method further includes: acquiring multiple second dialogue data samples and reference inference steps corresponding to the second dialogue data samples; and fine-tuning the initial generative model based on the second dialogue data samples and the reference inference steps to obtain the target generative model.

[0039] The second dialogue data sample is high-quality dialogue data obtained from the dialogue session data. For example, the second dialogue data sample can be a small number of first dialogue data samples extracted from each of the first dialogue data samples. The reference inference steps are high-quality inference steps that are manually annotated.

[0040] Specifically, the initial generative model is a pre-trained large model with a certain ability to generate inference steps, but it may have some output bias when applied to the target application domain. Therefore, this embodiment can obtain a small number of inference step examples from the target application domain to fine-tune the initial generative model. In this way, the electronic device can obtain second dialogue data samples and their corresponding reference inference steps, and then use these data to fine-tune the initial generative model to obtain the target generative model. This improves the applicability of the target generative model to the target application domain, thereby enhancing the rationality and accuracy of the generated inference steps.

[0041] S130. Integrate the first dialogue data sample and the reasoning steps corresponding to the first dialogue data sample to obtain the thought chain data sample corresponding to the first dialogue data sample.

[0042] The thought chain data sample consists of a triple structure consisting of sequentially arranged input content, reasoning steps, and response content.

[0043] Specifically, for each first dialogue data sample, the electronic device can associate and integrate the first dialogue data sample with the aforementioned inference steps corresponding to the first dialogue data sample to obtain a thought chain data sample containing input content, inference steps, and response content. For example, a predefined triple structure corresponds to a preset data structure, which includes sequentially arranged input content fields, inference step fields, and response content fields. Then, the electronic device can fill the input content and response content from the first dialogue data sample into the "input content field" and "response content field" of the preset data structure, respectively, and fill the corresponding inference steps into the "inference step field" of the preset data structure to obtain a thought chain data sample containing sequentially arranged input content, inference steps, and response content.

[0044] If a first dialogue data sample does not have a corresponding reasoning step, then that first dialogue data sample does not have a corresponding thought chain data sample.

[0045] S140. Based on the first dialogue data samples and the thought chain data samples, construct the training dataset corresponding to the target application domain.

[0046] Specifically, during the training of a large model, if the training dataset contains an excessive amount (e.g., around 70%, or a relatively large proportion) of thought chain data samples, the large model may focus too much on the thought chain content and ignore the actual response content, leading to increased model illusions. Conversely, if the training dataset contains too few thought chain data samples (e.g., around 3%, or a relatively small proportion), the model's reasoning ability will be limited, resulting in insufficient reasoning capabilities. Therefore, in this embodiment of the disclosure, a portion of data samples can be extracted from each of the first dialogue data samples and each thought chain data sample according to a certain proportion to form a training dataset corresponding to the target application domain, which mixes thought chain data and ordinary dialogue data.

[0047] It should be noted that, according to the descriptions of the foregoing embodiments, there is some overlap between the first dialogue data sample and the thought chain data sample. For example, a thought chain data sample always corresponds to a first dialogue data sample. Therefore, during the generation of the training dataset, if a certain thought chain data sample is selected, the first dialogue data sample corresponding to that thought chain data sample will be removed to ensure the diversity of the training dataset.

[0048] The method for generating training datasets provided in the above embodiments of this disclosure can acquire multiple first dialogue data samples, including input content and response content, in a target application domain; input the first dialogue data samples into a target generative model to obtain the inference steps corresponding to the first dialogue data samples; these inference steps are used to characterize the inference process of generating response content from input content; integrate the first dialogue data samples and the inference steps corresponding to the first dialogue data samples to obtain thought chain data samples corresponding to the first dialogue data samples; and construct a training dataset corresponding to the target application domain based on each first dialogue data sample and each thought chain data sample. This achieves the automatic generation of a large number of proprietary thought chain data samples for the target application domain, greatly reducing the manpower and time consumption in the process of collecting training samples, thereby improving the efficiency of acquiring thought chain data samples and providing a data foundation for obtaining a large model with proprietary logical reasoning capabilities for the target application domain. Furthermore, by obtaining a training dataset that mixes first dialogue data samples and thought chain data samples, the diversity of the training dataset of the large model can be further improved, thereby further providing a data foundation for the large model to effectively balance the model's reasoning ability and illusion suppression ability.

[0049] In some embodiments, S140 includes: extracting samples from each first dialogue data sample and each thought chain data sample according to the target mixing ratio to construct a training dataset corresponding to the target application domain.

[0050] The target blending ratio is a predetermined and reasonable proportion of the mixing of thought chain data samples and ordinary dialogue data samples. Different application domains have different domain characteristics and reasoning ability requirements, so the target blending ratio is adaptable to different application domains. The target blending ratio can be an empirically determined ratio within the target application domain, or it can be determined by using training datasets with different ratios for model training and performance evaluation. For example, in the real estate domain, the target blending ratio can be set at around 16%.

[0051] Specifically, the electronic device can calculate the first data volume corresponding to the thought chain data sample based on the total data volume of each first dialogue data sample and each thought chain data sample, as well as the target mixing ratio. Then, it can calculate the data volume corresponding to the first dialogue data sample from the total data volume and the first data volume. Next, it extracts thought chain data samples of the first data volume from each thought chain data sample. Then, it extracts first dialogue data samples of the second data volume, which are different from the extracted thought chain data samples, from each first dialogue data sample. These extracted data samples constitute a mixed training dataset for the target application domain. This training dataset can effectively balance the model's reasoning ability and illusion suppression ability, thus it can be applied to large-scale model training in various scenarios within the target application domain.

[0052] In some embodiments, the thought chain data samples can be further filtered based on the value of the reasoning steps to construct a hybrid training dataset. Specifically, S140 includes steps A through C.

[0053] Step A: Input the thought chain data sample into the value assessment model to determine the reasoning value corresponding to the thought chain data sample.

[0054] The value assessment model is a pre-trained machine learning model that evaluates the value of inference steps. For example, it could be a large model with good performance in the relevant technology. Inference value is used to characterize the contribution of the inference step to the generated response content. For example, if the inference step can better support the accuracy, coherence and compliance of the response content with the scenario requirements, then the value of the inference step is relatively high and the numerical value of the inference value is large.

[0055] Specifically, the electronic device can input the aforementioned thought chain data samples into the value assessment model and output the reasoning value corresponding to each thought chain data sample.

[0056] Step B: Based on the target value threshold and each reasoning value, filter the data samples of each thought chain to determine the filtered thought chain data samples.

[0057] Among them, the target value threshold is a pre-set critical value of reasoning value, which is a condition for data screening / filtering.

[0058] Specifically, the electronic device compares each inference value with a target value threshold. If the inference value is less than the target value threshold, the thought chain data sample corresponding to that inference value is discarded; if the inference value is greater than or equal to the target value threshold, the thought chain data sample corresponding to that inference value is retained. In this way, some thought chain data samples with low inference values ​​can be filtered out, while those with high inference values ​​are retained as the filtered thought chain data samples.

[0059] Step C: Based on the filtered thought chain data samples and the remaining first dialogue data samples, construct the training dataset corresponding to the target application domain.

[0060] The remaining first dialogue data samples include the first dialogue data samples other than the first dialogue data samples corresponding to the filtered thought chain data samples in each first dialogue data sample.

[0061] Specifically, the electronic device removes the first dialogue data samples corresponding to the filtered thought chain data samples from all the first dialogue data samples, obtaining the remaining first dialogue data samples. Then, the training dataset corresponding to the target application domain is constructed from the filtered thought chain data samples and the remaining first dialogue data samples.

[0062] In the above embodiments, the proportion of thought chain data samples in the training dataset can be controlled by the target value threshold, so that the training dataset contains high-value thought chain data samples, thereby further improving the reasoning ability and reasoning rationality of the large model trained using the training dataset.

[0063] In some embodiments, based on the above, the electronic device can obtain training datasets with different mixing ratios (e.g., 70%, 16%, 14%, 3.7%, 0%, etc.) by adjusting the target value threshold. Then, the electronic device can train the model using the training datasets with different mixing ratios under the same computing environment and initial generative model, obtaining the trained model. Afterwards, the trained models are invoked using the test input content from the test dialogue data to obtain test response content, or test response content and test inference steps. The degree of fit between each test response content (and test inference step) and business requirements is evaluated using evaluation algorithms, either manually or through related technologies, or the similarity between each test response content (and test inference step) and the response content (and inference step) in the test dialogue data is evaluated. By comparing the evaluation results, the test response content (and test inference step) with the best evaluation result is determined, and its corresponding target value threshold is determined as the final value threshold in the target application domain. Its corresponding mixing ratio (e.g., 16%) is also a relatively reasonable mixing ratio, and this mixing ratio can be determined as the target mixing ratio. The proportion of the filtered thought chain data samples in each first dialogue data sample is determined as the target mixing ratio. This method of determining the target mixing ratio through experiments with different mixing ratios improves its rationality. This ensures that the training dataset contains high-value thought chain data samples, further enhancing the reasoning ability and rationality of the large model trained on the training dataset.

[0064] The dialogue processing method provided in this disclosure is mainly applicable to scenarios where large models are used for dialogue responses, and is particularly suitable for scenarios where dialogue responses are performed in a target application domain. This dialogue processing method can be executed by a dialogue processing device, which can be implemented in software and / or hardware. The device can be integrated into an electronic device with certain data processing capabilities, such as a smartphone, laptop, desktop computer, server, or server cluster.

[0065] Figure 2 This is a flowchart of a dialogue processing method provided in an embodiment of this disclosure. See also... Figure 2 The methods for generating this training dataset include: S210, Receive dialogue input.

[0066] Specifically, electronic devices can obtain dialogue input content through human-computer interaction interfaces or upstream and downstream communication services. This dialogue input content can be real-time dialogue input obtained during instant communication, or questions received by intelligent agents with various functions.

[0067] S220. Based on the dialogue input content, call the dialogue processing model to obtain the dialogue response content corresponding to the dialogue input content, or obtain the dialogue reasoning steps and dialogue response content corresponding to the dialogue input content.

[0068] The dialogue processing model is obtained by fine-tuning a preset generative model using a training dataset generated by the training dataset generation method described in any embodiment of this disclosure. The preset generative model is a generative model with dialogue processing capabilities, such as a large language model.

[0069] Specifically, in this embodiment, the hybrid training dataset obtained in the foregoing embodiments can be used to fine-tune the preset generative model to obtain a dialogue processing model suitable for the target application domain. Because the training dataset is a mixture of high-quality thought chain data samples and first dialogue data samples, the trained dialogue processing model has the ability to recognize complex and simple tasks, and has a hybrid reasoning ability to generate reasoning steps and response content for complex tasks and directly generate response content for simple tasks.

[0070] Building upon the aforementioned fine-tuning training, electronic devices can further configure reward functions and their operational methods during the reinforcement learning phase of model training. For complex, labeled tasks (such as tax prediction and property recommendation), if the model output is found to lack pre-defined thought chain labels (e.g., structured labels like <|thought_start|>...<|thought_end|>), a penalty is applied; conversely, if the model output contains pre-defined thought chain labels, a reward is applied. For simple, labeled tasks (such as greetings and basic information queries), thought chains are not mandatory, maintaining the conciseness of the response. This further enhances the logicality and interpretability of the dialogue processing model.

[0071] Based on the above explanation, after receiving dialogue input, the electronic device can trigger the invocation of the dialogue processing model to perform dialogue processing within the model content. When the dialogue input corresponds to a complex task, it outputs dialogue reasoning steps and dialogue response content; when the dialogue input corresponds to a simple task, it outputs dialogue response content.

[0072] In some embodiments, electronic devices can guide the dialogue processing model to actively generate thought chains in complex tasks by adding constraints to the model prompts corresponding to the dialogue processing model to generate inference steps. Specifically, the constraints may include: if the dialogue input corresponds to a complex task, then generate inference steps and response content; if the dialogue input corresponds to a simple task, then generate response content.

[0073] In other embodiments, conditions for triggering the generation of thought chains can be set outside the dialogue processing model. In this case, "calling the dialogue processing model based on the dialogue input content to obtain the dialogue reasoning steps and dialogue response content corresponding to the dialogue input content" in S220 includes: if a preset keyword is detected or the number of dialogue rounds exceeds a preset number, a reasoning instruction is obtained; based on the dialogue input content and the reasoning instruction, the dialogue processing model is called to obtain the dialogue reasoning steps and dialogue response content.

[0074] Specifically, during dialogue processing, the electronic device can monitor whether the dialogue session data contains preset keywords corresponding to complex tasks, or whether the number of dialogue session rounds has exceeded the preset number corresponding to complex tasks. If both monitoring results are negative, the current dialogue process can be considered to still fall under the category of simple tasks. In this case, the electronic device can input the dialogue input content into the dialogue processing model to trigger the dialogue processing model to directly output the dialogue response content. If at least one of the above monitoring results is positive, it indicates that the current dialogue process has fallen under the category of complex tasks, and then a reasoning instruction is generated. Then, the electronic device inputs the dialogue input content and the reasoning instruction together into the dialogue processing model to trigger the dialogue processing model to simultaneously run the reasoning step generation function, outputting the dialogue response content and the dialogue reasoning steps.

[0075] The dialogue processing method provided in the above embodiments of this disclosure can use the aforementioned generated hybrid training dataset to fine-tune a pre-trained large model to obtain a dialogue processing model. This model is then used to process received dialogue input to obtain dialogue response content corresponding to the input, or dialogue reasoning steps and response content corresponding to the input. This allows the dialogue processing model to additionally learn logical reasoning patterns specific to the target application domain, effectively balancing the model's reasoning ability and illusion suppression to achieve hybrid reasoning capabilities. Specifically, when handling complex tasks, it can trigger reasoning capabilities to generate thought chain data, improving the logical rationality, accuracy, and interpretability of the response content. Conversely, when handling simple tasks, it can directly respond / reply, preserving the fluency of the dialogue.

[0076] Figure 3 This is a schematic diagram of a device for generating a training dataset according to an embodiment of this disclosure. Figure 3 As shown, the training dataset generation device 300 includes: The first dialogue data sample acquisition module 310 is used to acquire multiple first dialogue data samples in the target application domain; wherein, the first dialogue data sample includes input content and response content corresponding to the input content; The reasoning step acquisition module 320 is used to input the first dialogue data sample into the target generative model and obtain the reasoning steps corresponding to the first dialogue data sample; wherein, the reasoning step is used to characterize the reasoning process of generating response content from input content; the target generative model includes a model that has the ability to generate reasoning steps in thought chain data, obtained by fine-tuning the pre-trained initial generative model. The thought chain data sample acquisition module 330 is used to integrate the first dialogue data sample and the reasoning steps corresponding to the first dialogue data sample to obtain the thought chain data sample corresponding to the first dialogue data sample; wherein, the thought chain data sample is data with a triple structure including input content, reasoning steps and response content arranged in sequence. The training dataset construction module 340 is used to construct a training dataset corresponding to the target application domain based on each first dialogue data sample and each thought chain data sample.

[0077] The training dataset generation apparatus provided in this embodiment can acquire multiple first dialogue data samples, including input content and response content, in a target application domain; input the first dialogue data samples into a target generative model to obtain the inference steps corresponding to the first dialogue data samples; the inference steps are used to characterize the reasoning process of generating response content from input content; integrate the first dialogue data samples and the inference steps corresponding to the first dialogue data samples to obtain the thought chain data samples corresponding to the first dialogue data samples; and construct a training dataset corresponding to the target application domain based on each first dialogue data sample and each thought chain data sample. This achieves the automatic generation of a large number of proprietary thought chain data samples for the target application domain, greatly reducing the manpower and time consumption in the process of collecting training samples, thereby improving the efficiency of acquiring thought chain data samples and providing a data foundation for obtaining a large model with proprietary logical reasoning capabilities for the target application domain. Furthermore, by obtaining a training dataset that mixes first dialogue data samples and thought chain data samples, the diversity of the training dataset of the large model can be further improved, thereby further providing a data foundation for the large model to effectively balance the model's reasoning ability and illusion suppression ability.

[0078] In some embodiments, the training dataset generation apparatus 300 further includes a target generative model acquisition module, used for: Before inputting each first dialogue data sample into the target generative model to generate the inference steps corresponding to the first dialogue data sample, multiple second dialogue data samples and reference inference steps corresponding to the second dialogue data samples are obtained. Based on the second dialogue data sample and reference reasoning steps, the initial generative model is fine-tuned and trained to obtain the target generative model.

[0079] In some embodiments, the reasoning step obtaining module 320 is specifically used for: If the first dialogue data sample includes multiple rounds of dialogue data in a dialogue session, then the first dialogue data sample is input into the target generative model to obtain the local inference steps corresponding to each round of dialogue data. Each local reasoning step is cleaned to obtain the reasoning steps corresponding to the first dialogue data sample; wherein, the cleaning process includes at least one of deduplication, logical compliance filtering and domain relevance filtering.

[0080] In some embodiments, the training dataset construction module 340 is specifically used for: Based on the target mixing ratio, samples are extracted from each first dialogue data sample and each thought chain data sample to construct a training dataset corresponding to the target application domain.

[0081] In other embodiments, the training dataset construction module 340 is specifically used for: Input the thought chain data sample into the value assessment model to determine the reasoning value corresponding to the thought chain data sample; whereby the reasoning value is used to characterize the contribution of the reasoning step to the generated response content; Based on the target value threshold and each reasoning value, the data samples of each thinking chain are filtered to determine the filtered thinking chain data samples. Based on the filtered thought chain data samples and the remaining first dialogue data samples, a training dataset corresponding to the target application domain is constructed; wherein, the remaining first dialogue data samples include the first dialogue data samples other than the first dialogue data samples corresponding to the filtered thought chain data samples.

[0082] In some embodiments, the first dialogue data sample acquisition module 310 is specifically used for: If the target application area includes the real estate sector, then multiple raw dialogue data are extracted from the dialogue session data corresponding to the real estate business scenarios in the real estate sector; among them, the real estate business scenarios include at least one of the following: property introduction scenario, property recommendation scenario, tax and fee prediction scenario, and loan consultation scenario; The original dialogue data is cleaned and processed to obtain the first dialogue data samples.

[0083] The training dataset generation apparatus provided in this disclosure can execute the training dataset generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0084] Figure 4 This is a schematic diagram of the structure of a dialogue processing device provided in an embodiment of this disclosure. Figure 4 As shown, the dialogue processing device 400 includes: Dialogue input content receiving module 410 is used to receive dialogue input content; The dialogue response content acquisition module 420 is used to call the dialogue processing model based on the dialogue input content to obtain the dialogue response content corresponding to the dialogue input content, or to obtain the dialogue reasoning steps and dialogue response content corresponding to the dialogue input content; wherein, the dialogue processing model is obtained by fine-tuning the preset generative model using a training dataset, and the training dataset is obtained using the training dataset generation method described in any embodiment of this disclosure.

[0085] The dialogue processing apparatus provided in this embodiment can use the aforementioned generated hybrid training dataset to fine-tune a pre-trained large model to obtain a dialogue processing model. This model is then used to process received dialogue input to obtain either dialogue response content or dialogue reasoning steps and response content. This allows the dialogue processing model to additionally learn logical reasoning patterns specific to the target application domain, effectively balancing the model's reasoning ability and illusion suppression to achieve hybrid reasoning capabilities. Specifically, when handling complex tasks, it can trigger reasoning capabilities to generate thought chain data, improving the logical rationality, accuracy, and interpretability of the response content. Conversely, when handling simple tasks, it can directly respond / reply, preserving the fluency of the dialogue.

[0086] In some embodiments, the model prompts corresponding to the dialogue processing model include constraints for generating inference steps. The constraints include: If the dialogue input corresponds to a complex task, then reasoning steps and response content are generated; If the dialogue input corresponds to a simple task, a response is generated.

[0087] In some embodiments, the dialogue response content acquisition module 420 is specifically used for: If a preset keyword or the number of dialogue rounds exceeds a preset number, a reasoning instruction is obtained; Based on the dialogue input and reasoning instructions, the dialogue processing model is invoked to obtain the dialogue reasoning steps and dialogue response content.

[0088] The dialogue processing apparatus provided in this disclosure can execute the dialogue processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0089] It is worth noting that in any embodiment of the above-mentioned device, the modules included are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional module are only for easy differentiation and are not used to limit the scope of protection of this disclosure.

[0090] This disclosure also provides an electronic device comprising: one or more processors and a memory; wherein the memory is used to store one or more programs or instructions. When the one or more programs or instructions are executed by the one or more processors, the one or more processors cause the one or more processors to implement the training dataset generation method or dialogue processing method provided in any embodiment of this disclosure.

[0091] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Figure 5 As shown, the electronic device 500 includes a processor 501, a memory 502, an input device 503, and an output device 504, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown). The number of processors 501 and memory 502 can be one or more. Figure 5 The example below uses a processor 501 and a memory 502.

[0092] The processor 501 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.

[0093] Memory 502 may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. In some embodiments, memory 502 may further include memory remotely located relative to processor 501, which can be connected to electronic device 500 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. One or more computer programs or instructions may be stored in memory 502, which processor 501 can execute to implement the training dataset generation method, dialogue processing method, and / or other desired functions described in any embodiment of this disclosure. Various content such as dialogue data and training datasets may also be stored in memory 502.

[0094] Input device 503 may include, for example, a keyboard, a mouse, etc. Output device 504 may output various information to the outside, including training datasets, dialogue processing models, dialogue response content, or dialogue reasoning steps, etc. Output device 504 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0095] Understandably, for the sake of simplification, Figure 5 Only some of the components of the electronic device 500 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 500 may include any other suitable components depending on the specific application.

[0096] In addition to the methods and apparatus described above, the training dataset generation method or dialogue processing method in any embodiment of this disclosure can also be implemented as a computer software program. For example, embodiments of this disclosure also include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from memory. When the computer program is run by a processor, it causes the processor to execute the training dataset generation method or dialogue processing method provided in any embodiment of this disclosure.

[0097] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0098] Furthermore, embodiments of this disclosure also provide a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, causes the processor to perform the training dataset generation method or dialogue processing method provided in embodiments of this disclosure.

[0099] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0100] It should be noted that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of this disclosure. As shown in this specification and claims, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. The term "and / or" includes any one and all combinations of one or more of the associated listed items. Relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another entity or operation and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element.

[0101] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating a training dataset, characterized in that, include: Acquire multiple first dialogue data samples in the target application domain; wherein, the first dialogue data sample includes input content and response content corresponding to the input content; The first dialogue data sample is input into the target generative model to obtain the reasoning steps corresponding to the first dialogue data sample; wherein, the reasoning steps are used to characterize the reasoning process of generating the response content from the input content; the target generative model includes a model that has the ability to generate reasoning steps in thought chain data, obtained by fine-tuning a pre-trained initial generative model. By integrating the first dialogue data sample and the reasoning steps corresponding to the first dialogue data sample, a thought chain data sample corresponding to the first dialogue data sample is obtained; wherein, the thought chain data sample is data comprising a triple structure of the input content, the reasoning steps and the response content arranged in sequence; Based on each of the first dialogue data samples and each of the thought chain data samples, a training dataset corresponding to the target application domain is constructed.

2. The method according to claim 1, characterized in that, Before inputting each of the first dialogue data samples into the target generative model to obtain the inference step corresponding to the first dialogue data sample, the method further includes: Acquire multiple second dialogue data samples and the corresponding reference inference steps for the second dialogue data samples; Based on the second dialogue data sample and the reference inference steps, the initial generative model is fine-tuned and trained to obtain the target generative model.

3. The method according to claim 1, characterized in that, If the first dialogue data sample includes multi-turn dialogue data from a dialogue session, then the step of inputting the first dialogue data sample into the target generative model to obtain the inference step corresponding to the first dialogue data sample includes: Input the first dialogue data sample into the target generative model to obtain the local inference steps corresponding to each round of dialogue data; Each of the aforementioned local reasoning steps is cleaned to obtain the reasoning steps corresponding to the first dialogue data sample; wherein, the cleaning process includes at least one of deduplication, logical compliance filtering, and domain relevance filtering.

4. The method according to claim 1, characterized in that, The step of constructing a training dataset corresponding to the target application domain based on each of the first dialogue data samples and each of the thought chain data samples includes: According to the target mixing ratio, samples are extracted from each of the first dialogue data samples and each of the thought chain data samples to construct a training dataset corresponding to the target application domain.

5. The method according to claim 1, characterized in that, The step of constructing a training dataset corresponding to the target application domain based on each of the first dialogue data samples and each of the thought chain data samples includes: The thought chain data sample is input into the value assessment model to determine the reasoning value corresponding to the thought chain data sample; wherein, the reasoning value is used to characterize the contribution of the reasoning step to the generation of the response content; Based on the target value threshold and the inference value of each, the data samples of each thought chain are filtered to determine the filtered thought chain data samples. Based on the filtered thought chain data samples and the remaining first dialogue data samples, a training dataset corresponding to the target application domain is constructed; wherein, the remaining first dialogue data samples include the first dialogue data samples other than the first dialogue data samples corresponding to the filtered thought chain data samples.

6. The method according to claim 1, characterized in that, If the target application domain includes the real estate domain, then obtaining multiple first dialogue data samples from the target application domain includes: Multiple raw dialogue data are extracted from the dialogue session data corresponding to real estate business scenarios in the real estate field; wherein, the real estate business scenarios include at least one of the following: property introduction scenario, property recommendation scenario, tax and fee prediction scenario, and loan consultation scenario; The original dialogue data is cleaned to obtain the first dialogue data samples.

7. A dialogue processing method, characterized in that, include: Receive dialogue input; Based on the dialogue input content, a dialogue processing model is invoked to obtain the dialogue response content corresponding to the dialogue input content, or to obtain the dialogue inference steps and the dialogue response content corresponding to the dialogue input content; wherein, the dialogue processing model is obtained by fine-tuning a preset generative model using a training dataset, and the training dataset is obtained using the training dataset generation method as described in any one of claims 1 to 6.

8. The method according to claim 7, characterized in that, The model prompts corresponding to the dialogue processing model contain constraints for generating inference steps. The constraints include: If the dialogue input corresponds to a complex task, then reasoning steps and response content are generated; If the dialogue input corresponds to a simple task, then a response is generated.

9. The method according to claim 7, characterized in that, Based on the dialogue input content, the dialogue processing model is invoked to obtain the dialogue inference steps corresponding to the dialogue input content and the dialogue response content, including: If a preset keyword or the number of dialogue rounds exceeds a preset number, a reasoning instruction is obtained; Based on the dialogue input content and the reasoning instruction, the dialogue processing model is invoked to obtain the dialogue reasoning steps and the dialogue response content.

10. An electronic device, characterized in that, The electronic device includes: Processor and memory; The processor executes the training dataset generation method as described in any one of claims 1 to 6, or the dialogue processing method as described in any one of claims 7 to 9, by calling programs or instructions stored in the memory.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that cause a computer to perform the method for generating a training dataset as described in any one of claims 1 to 6, or the dialogue processing method as described in any one of claims 7 to 9.

12. A computer program product, characterized in that, The computer program product is used to implement the method for generating the training dataset as described in any one of claims 1 to 6, or to perform the dialogue processing method as described in any one of claims 7 to 9.