Conversation preference data construction method and device, computer equipment and storage medium
By collaboratively generating and evaluating candidate dialogue responses using multiple models, the problems of low efficiency and high subjectivity in constructing dialogue preference data are solved, achieving efficient and diverse generation of dialogue preference data and improving the quality and coverage of training data for large language models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU QUWAN NETWORK TECH CO LTD
- Filing Date
- 2025-12-05
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies suffer from low efficiency and high cost in constructing dialogue preference data, and manual annotation is subjective, making it difficult to cover diverse dialogue scenarios and affecting model alignment performance.
By leveraging the collaborative efforts of multiple large-scale dialogue generation models and multiple large-scale dialogue evaluation models, candidate dialogue responses are generated and evaluated. Dialogue preference data is generated using multi-model voting, reducing the subjectivity of human evaluation and improving the objectivity and robustness of the evaluation.
It achieves end-to-end automated synthesis of dialogue preference data, reduces the cost and time of manual annotation, generates diverse and high-quality reinforcement learning training data, and improves the efficiency and reliability of model dialogue preference alignment.
Smart Images

Figure CN121833934A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a dialogue preference data construction method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] With the development of artificial intelligence technology, a technology of generating dialogue data by using a large language model has appeared. In the technology, the generated dialogue reply needs to align with human preferences, that is, the content generated by the model needs to conform to human values, emotional needs and dialogue habits, such as naturalness, humor, emotional resonance, rather than just pursuing grammatical correctness.
[0003] In the traditional technology, the large language model used to generate dialogue data is usually trained based on a reinforcement learning method of human feedback. By labeling positive and negative samples, that is, dialogue preference data, a reward model is trained to guide the reinforcement learning training.
[0004] However, in the above method, the dialogue preference data is usually obtained by manual labeling. However, manual labeling has problems such as high cost, low efficiency, limited scale, and is difficult to cover diversified dialogue scenarios. At the same time, the subjectivity of manual labeling easily leads to data bias, affecting the alignment effect of the model. Therefore, the current dialogue preference data construction is not convenient. SUMMARY
[0005] Therefore, it is necessary to provide a dialogue preference data construction method, device, computer equipment, computer readable storage medium and computer program product capable of improving the efficiency of dialogue preference data construction.
[0006] In a first aspect, the present application provides a dialogue preference data construction method, comprising:
[0007] obtaining original dialogue data and an original dialogue reply corresponding to the original dialogue data;
[0008] constructing dialogue data generation requests respectively from the original dialogue data, the original dialogue reply and a plurality of prompt words corresponding to different prompt word types and pre-constructed, inputting each dialogue data generation request into a plurality of dialogue generation large models corresponding to different architectures and pre-constructed, and generating candidate dialogue replies matched with each dialogue data generation request by each dialogue generation large model;
[0009] inputting each candidate dialogue reply and the original dialogue reply as a to-be-evaluated dialogue reply corresponding to the original dialogue data into a plurality of dialogue evaluation large models corresponding to different architectures and pre-constructed, and outputting dialogue preference evaluation results for each to-be-evaluated dialogue reply by each dialogue evaluation large model respectively.
[0010] Based on the dialogue preference evaluation results respectively output by the large model according to each of the dialogues, a dialogue preference data of the original dialogue data is generated.
[0011] In one of the embodiments, the prompt word types include a first prompt word type, a second prompt word type, and a third prompt word type; the prompt words of the first prompt word type and the second prompt word type are used to prompt generation of high-quality dialogue replies, and the prompt words of the third prompt word type are used to prompt generation of low-quality dialogue replies; the dialogue data generation requests include a first dialogue data generation request, a second dialogue data generation request, and a third dialogue data generation request, the first dialogue data generation request is generated based on the prompt words of the first prompt word type, the second dialogue data generation request is generated based on the prompt words of the second prompt word type, and the third dialogue data generation request is generated based on the prompt words of the third prompt word type.
[0012] The dialogue data generation requests are input into a plurality of dialogue generation large models corresponding to different architectures which are constructed in advance, and candidate dialogue replies matched with the dialogue data generation requests are generated by the dialogue generation large models, including:
[0013] The first dialogue data generation request is input into a current dialogue generation large model, and a first candidate dialogue reply is generated by the current dialogue generation large model; the first candidate dialogue reply is a candidate dialogue reply with high emotional intensity, and the current dialogue generation large model is any one of the dialogue generation large models corresponding to different architectures;
[0014] The second dialogue data generation request is input into the current dialogue generation large model, and a second candidate dialogue reply is generated by the current dialogue generation large model; the second candidate dialogue reply is a candidate dialogue reply with high naturalness;
[0015] The third dialogue data generation request is input into the current dialogue generation large model, and a third candidate dialogue reply is generated by the current dialogue generation large model; the third candidate dialogue reply is a low-quality candidate dialogue reply.
[0016] In one of the embodiments, the original dialogue data and the original dialogue reply corresponding to the original dialogue data are obtained, including:
[0017] Sample dialogue data containing multiple rounds of dialogues are obtained;
[0018] The last round of dialogue data in the sample dialogue data is taken as the original dialogue reply, and the dialogue data in the sample dialogue data except the last round of dialogue data is taken as the original dialogue data.
[0019] In one of the embodiments, the inputting the candidate dialogue replies and the original dialogue reply corresponding to the original dialogue data as the to-be-evaluated dialogue replies into the pre-constructed dialogue evaluation large models corresponding to different architectures respectively comprises:
[0020] The inputting the to-be-evaluated dialogue replies into the current dialogue evaluation large model, and the obtaining of the target to-be-evaluated dialogue reply from the to-be-evaluated dialogue replies by the current dialogue evaluation large model, wherein the current dialogue evaluation large model is any one of the dialogue evaluation large models;
[0021] The target to-be-evaluated dialogue reply is the preferred dialogue reply selected from the to-be-evaluated dialogue replies by the current dialogue evaluation large model, and the target to-be-evaluated dialogue reply is the dialogue preference evaluation result output by the current dialogue evaluation large model for the to-be-evaluated dialogue replies;
[0022] The generating of the dialogue preference data of the original dialogue data based on the dialogue preference evaluation results output by the dialogue evaluation large models respectively comprises:
[0023] The generating of the dialogue preference data of the original dialogue data based on the target to-be-evaluated dialogue replies selected by the dialogue evaluation large models respectively.
[0024] In one of the embodiments, the dialogue preference data is used to represent that the to-be-evaluated dialogue reply is the dialogue preference positive sample of the original dialogue data, or the dialogue preference negative sample of the original dialogue data;
[0025] The generating of the dialogue preference data of the original dialogue data based on the target to-be-evaluated dialogue replies selected by the dialogue evaluation large models respectively comprises:
[0026] If the current to-be-evaluated dialogue reply is selected as the target to-be-evaluated dialogue reply by any one of the dialogue evaluation large models, the current to-be-evaluated dialogue reply is taken as the dialogue preference positive sample of the original dialogue data, wherein the current to-be-evaluated dialogue reply is any one of the to-be-evaluated dialogue replies;
[0027] If the current to-be-evaluated dialogue reply is not selected as the target to-be-evaluated dialogue reply by any one of the dialogue evaluation large models, the current to-be-evaluated dialogue reply is taken as the dialogue preference negative sample of the original dialogue data;
[0028] The dialogue preference positive sample and the dialogue preference negative sample of the original dialogue data are taken as the dialogue preference data of the original dialogue data.
[0029] In one of the embodiments, the inputting each of the to-be-evaluated dialogue reply into the current dialogue evaluation large model comprises:
[0030] The current dialogue evaluation large model is inputted with each of the to-be-evaluated dialogue reply and a pre-constructed dialogue evaluation prompt word, wherein the dialogue evaluation prompt word comprises role information, dialogue reply selection standard information and dialogue reply selection rule information; the dialogue reply selection rule information is used to indicate that the original dialogue reply is taken as the target to-be-evaluated dialogue reply in a case where each of the candidate dialogue reply does not meet the dialogue reply selection standard information.
[0031] The current dialogue evaluation large model is used to perform a task of selecting a dialogue reply according to the dialogue reply selection standard information and the dialogue reply selection rule information by using a role represented by the role information, and a to-be-evaluated dialogue reply selected by the current dialogue evaluation large model is taken as the target to-be-evaluated dialogue reply.
[0032] In one of the embodiments, after the dialogue preference data of the original dialogue data is generated, the method further comprises:
[0033] The original dialogue data and the dialogue preference data as a reward signal are inputted into a dialogue reply generation model to be trained, so as to perform reinforcement learning training on the dialogue reply generation model.
[0034] In a second aspect, the present application further provides a dialogue preference data construction device, comprising:
[0035] An original dialogue acquisition module is configured to acquire original dialogue data and original dialogue replies corresponding to the original dialogue data;
[0036] A candidate reply generation module is configured to construct dialogue data generation requests respectively by inputting the original dialogue data, the original dialogue replies and a plurality of prompt words corresponding to different prompt word types and pre-constructed into the dialogue data generation requests, input the dialogue data generation requests into a plurality of dialogue generation large models corresponding to different architectures and pre-constructed, and generate candidate dialogue replies matched with the dialogue data generation requests by the dialogue generation large models.
[0037] A dialogue preference evaluation module is configured to input each of the candidate dialogue replies and the original dialogue replies as to-be-evaluated dialogue replies corresponding to the original dialogue data into a plurality of dialogue evaluation large models corresponding to different architectures and pre-constructed, and output dialogue preference evaluation results for each of the to-be-evaluated dialogue replies by the dialogue evaluation large models.
[0038] The dialogue preference generation module is configured to generate dialogue preference data of the original dialogue data based on dialogue preference evaluation results respectively output by the dialogue evaluation large models.
[0039] In a third aspect, the present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the dialogue preference data construction method when executing the computer program.
[0040] In a fourth aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the dialogue preference data construction method when executed by a processor.
[0041] In a fifth aspect, the present application also provides a computer program product comprising a computer program, and the computer program implements the steps of the dialogue preference data construction method when executed by a processor.
[0042] The dialogue preference data construction method, device, computer device, computer readable storage medium and computer program product can generate candidate dialogue replies with different quality standards based on the dialogue generation large models corresponding to different architectures and the prompt words corresponding to different prompt word types, implement a diversified, multi-scene dialogue reply generation strategy based on multiple models and multiple prompt words, and then evaluate and select the dialogue reply to be evaluated composed of each candidate dialogue reply and the original dialogue reply based on the dialogue evaluation large models corresponding to different architectures, so that objective dialogue preference evaluation results can be obtained through the objective selection of each dialogue evaluation large model, and dialogue preference data can be finally generated, the subjectivity of human evaluation is reduced, the objectivity and robustness of preference evaluation and judgment are improved, and the reliability of training data is improved. Therefore, the present application can realize end-to-end automatic synthesis from the original dialogue to dialogue preference data for reinforcement learning training through the two linking stages of diversified reply generation and preference selection and data construction, reduces the large amount of manpower expenditure caused by manual labeling, improves the construction efficiency of reinforcement learning training data, and thus provides diversified, high-quality and scalable reinforcement learning training data for preference alignment of large language models. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can be obtained without creative labor based on these drawings.
[0044] Figure 1A flowchart of a method for constructing conversation preference data in an embodiment;
[0045] Figure 2 A storage format example of original conversation data in an embodiment;
[0046] Figure 3 A flowchart of a method for constructing conversation preference data in another embodiment;
[0047] Figure 4 A block diagram of a construction device for conversation preference data in an embodiment;
[0048] Figure 5 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0049] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0050] It should be noted that the terms "first", "second", etc. used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" used in the present application and any variations thereof are intended to cover non-exclusive inclusion. The term "a plurality of" used in the present application means two or more. The term "and / or" used in the present application means one of the options or any combination of multiple options.
[0051] With the rapid improvement of the capabilities of large language models (LLM), the "alignment with human preferences" has become a core requirement, i.e. to make the content generated by the model consistent with human values, emotional needs and conversation habits (such as natural, humorous, and emotionally resonant), rather than just pursuing grammatical correctness.
[0052] Currently, reinforcement learning from human feedback (RLHF) is the mainstream technical framework to achieve this goal: the core logic is to "label positive and negative samples (preference data) by humans → train reward model (RM) → use RM to guide RL (Reinforcement Learning) training", so that the large language model learns "what is a good reply".
[0053] However, the traditional RLHF human-assisted data generation technology has the following shortcomings:
[0054] Dialogue data generation is inefficient and costly: it relies on manual annotation—annotators sort candidate responses according to their "degree of conformity to human preferences" (e.g., 1-5 points), or directly filter "chosen responses" and "rejected responses." Annotating a single dialogue is time-consuming (requiring reading the context and judging preferences), and large-scale annotation requires a lot of manpower, making it difficult to support the "massive amount of preference data" required for LLM training.
[0055] Insufficient diversity of candidate responses: It relies heavily on "a single large language model + general prompts" to generate candidate responses. For example, it uses only a single model and generates a small number of candidates through general instructions such as "continue the dialogue" or "optimize the response". It does not systematically distinguish between "high-quality" and "low-quality" samples. It is limited by the "inherent bias" of the model (such as a single language style and fixed emotional expression) and cannot cover the diverse human dialogue habits, which can easily lead to model overfitting.
[0056] The generation of positive and negative samples is unsystematic: There is a lack of clear "high-quality sample guidance" and "low-quality sample guidance" design. High-quality samples rely solely on the model's "default optimization," while low-quality samples are mostly randomly generated (such as grammatical errors). They do not accurately hit the low-quality characteristics of real-world scenarios such as "lack of emotional resonance, perfunctory responses, and dialogue stagnation," and a systematic positive and negative sample generation mechanism has not been formed.
[0057] The subjective nature of preference assessment is high: manual annotation is affected by the emotions and cognitive differences of the annotators, single model evaluation is affected by the bias of the model itself, and there is a lack of multi-model consensus verification. All of these cannot objectively reflect "group human preferences", resulting in low data reliability.
[0058] Low level of automation in the process: The generation, annotation, and evaluation stages are mostly independent steps that require manual coordination (such as exporting the generated candidates to the annotation team), which makes it impossible to form an end-to-end pipeline and further reduces efficiency.
[0059] It is evident that the bottleneck of traditional RLHF (Reinforcement Learning Based on Human Feedback) lies in the construction of preference data: high-quality human annotations are the core source of data, but this process suffers from high costs, low efficiency, and limited scale, and it is difficult to cover diverse dialogue scenarios (such as different emotional tendencies and topics); at the same time, the subjectivity of human annotations can easily lead to data bias, affecting the model alignment effect. Therefore, "automated and large-scale generation of high-quality preference data" has become a key technical pain point in the LLM alignment field.
[0060] Therefore, this application proposes a method for constructing dialogue preference data, which aims to synthesize reinforcement learning data through the synergy of "diversified response generation" and "multi-model preference selection" in order to provide high-quality, scalable reinforcement learning training data for the dialogue data preference alignment of large language models.
[0061] It should be noted that the dialogue preference data construction method provided in this application is mainly applied to large-scale language models that require precise alignment with complex human preferences, including but not limited to: dialogue interaction models (such as emotion assistants and chatbots); intelligent question-answering models (such as customer service robots and knowledge question-answering systems); and generative content models (such as dialogue continuation and personalized reply generation tools). The reinforcement learning preference data generated can be directly used for RL (reinforcement learning) training of the above models to achieve alignment between model behavior and human preferences.
[0062] In one exemplary embodiment, such as Figure 1 As shown, a method for constructing dialogue preference data is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps S102 to S108. Wherein:
[0063] Step S102: Obtain the original dialogue data and the original dialogue responses corresponding to the original dialogue data.
[0064] It should be noted that the original dialogue data and original dialogue responses can be obtained by extracting core information from the original sample dialogue data. The original dialogue data and original dialogue responses can be collectively referred to as the dialogue context.
[0065] The sample dialogue data can be text, voice, and other information generated during multi-turn interactions between multiple entities, including dialogue content between humans and between humans and AI. It can be high-quality multi-turn dialogue fragments that have undergone preliminary cleaning and filtering, such as publicly available dialogue datasets or manually collected dialogues. For example, the sample dialogue data can be formatted as follows: Figure 2 Store the data in the format shown.
[0066] In practical applications, the terminal can extract core information from the original sample dialogue data based on the high-quality multi-turn dialogue sample dialogue data that has undergone preliminary cleaning and screening, thereby obtaining the original dialogue data and the original dialogue responses corresponding to the original dialogue data.
[0067] Step S104: Take the original dialogue data, the original dialogue response, and a number of pre-built prompt words corresponding to different prompt word types to construct dialogue data generation requests. Input each dialogue data generation request into a number of pre-built dialogue generation models corresponding to different architectures. Generate candidate dialogue responses that match each dialogue data generation request through each dialogue generation model.
[0068] The prompt word type can correspond to different response quality standards, which are used to prompt various dialogue generation models to generate dialogue responses with different response quality standards.
[0069] The large dialogue generation model can be a large language model (LLM) used to generate dialogues. Since different large dialogue generation models have differences in "model personality" (such as language style and emotional expression tendency), in this embodiment of the application, multiple large language models corresponding to different architectures can be selected as the large dialogue generation model. For example, multiple cutting-edge LLMs with different architectures or training backgrounds can be selected to cover a wider range of response types while avoiding the dominance of bias from a single model.
[0070] In practice, the terminal can combine the original dialogue data, the original dialogue response, and multiple pre-built prompt words corresponding to different prompt word types to construct dialogue data generation requests corresponding to different prompt word types for each dialogue generation big model. Through multiple pre-built dialogue generation big model APIs (Application Programming Interfaces) corresponding to different architectures, each dialogue data generation request is input to the dialogue generation big model in an asynchronous and concurrent manner. Each dialogue generation big model generates candidate dialogue responses that match each dialogue data generation request.
[0071] Thus, the first stage of the embodiments of this application is realized, namely the stage of generating diversified responses based on multiple models and multiple prompts, which provides rich samples for the selection of dialogue preferences in subsequent large dialogue evaluation models.
[0072] Step S106: Input each candidate dialogue response and the original dialogue response as the dialogue response to be evaluated corresponding to the original dialogue data into multiple pre-built large dialogue evaluation models corresponding to different architectures, and output the dialogue preference evaluation results for each dialogue response to be evaluated through each large dialogue evaluation model.
[0073] The large dialogue evaluation model can be a large language model (LLM) used to evaluate each candidate dialogue response and the original dialogue response; it can also be called an evaluator model. For example, multiple LLMs with different architectures or training backgrounds can be selected as the large dialogue evaluation model. They can overlap with the large dialogue generation model, and their number can be greater than the number of large dialogue generation models. The purpose is to simulate "consensus and disagreement among human evaluators" and reduce the subjectivity of a single model through the diversity of the large dialogue evaluation model.
[0074] The dialogue responses to be evaluated can be formed by combining the candidate dialogue responses generated by each large dialogue generation model with the original dialogue responses containing the last round of dialogue data from the sample dialogue data.
[0075] In practical applications, the terminal can combine the candidate dialogue responses generated by the large dialogue generation model in the first stage with the original dialogue responses containing the last round of dialogue data from the sample dialogue data to form the dialogue response to be evaluated corresponding to the original dialogue data. The terminal can then input multiple pre-built large dialogue evaluation models corresponding to different architectures and, based on pre-built dialogue evaluation prompts, output dialogue preference evaluation results for each dialogue response to be evaluated through each large dialogue evaluation model.
[0076] Step S108: Based on the dialogue preference evaluation results output by each dialogue evaluation model, generate dialogue preference data from the original dialogue data.
[0077] In practice, the terminal can summarize the dialogue preference evaluation results output by each dialogue evaluation model to generate dialogue preference data from the original dialogue data.
[0078] Thus, the second stage of the embodiment of this application is realized, namely the preference selection and data construction stage based on multi-model voting. By simulating "human group evaluation" through multiple large dialogue evaluation models, the candidate dialogue responses in the first stage are ranked according to preferences, and RL (reinforcement learning) training data (i.e. dialogue preference data) containing "preference information" is generated.
[0079] The aforementioned dialogue preference data construction method, based on multiple large-scale dialogue generation models corresponding to different architectures and multiple prompt words corresponding to different prompt word types, can generate candidate dialogue responses with different quality standards. This realizes a diversified and multi-scenario dialogue response generation strategy based on multiple models and multiple prompt words. Furthermore, based on multiple large-scale dialogue evaluation models corresponding to different architectures, the method evaluates and selects the dialogue response to be evaluated, composed of each candidate dialogue response and the original dialogue response. This allows for objective selection by each large-scale dialogue evaluation model, resulting in objective dialogue preference evaluation results and ultimately generating dialogue preference data. This reduces the subjectivity of human evaluation, improves the objectivity and robustness of preference evaluation judgments, and enhances the reliability of training data. Therefore, this solution achieves end-to-end automated synthesis from "original dialogue" to "dialogue preference data" used for reinforcement learning training through two interconnected stages: "diversified response generation" and "preference selection and data construction." This reduces the significant manpower expenditure associated with manual annotation, improves the efficiency of reinforcement learning training data construction, and provides diversified, high-quality, and scalable reinforcement learning training data for preference alignment of large language models.
[0080] In a possible implementation, three types of prompt words can be designed for each dialogue generation model; that is, the prompt word types can include: a first prompt word type, a second prompt word type, and a third prompt word type. For each type of prompt word, the "response quality standard" for the dialogue response generated by the dialogue generation model needs to be clearly defined. The prompt words of the first and second types are used to prompt the generation of high-quality dialogue responses, while the prompt words of the third type are used to prompt the generation of low-quality dialogue responses.
[0081] Optionally, the three types of prompt words and their corresponding response quality standards are shown in Table 1 below:
[0082] Table 1. Criteria for Prompt Design
[0083]
[0084] Therefore, based on the three types of prompt words mentioned above, the dialogue data generation request can include: a first dialogue data generation request, a second dialogue data generation request, and a third dialogue data generation request. Understandably, the first dialogue data generation request is generated based on prompt words of the first prompt word type, the second dialogue data generation request is generated based on prompt words of the second prompt word type, and the third dialogue data generation request is generated based on prompt words of the third prompt word type.
[0085] Therefore, in an exemplary embodiment, each dialogue data generation request is input into a pre-built set of large dialogue generation models corresponding to different architectures, and candidate dialogue responses matching each dialogue data generation request are generated through each large dialogue generation model. This includes: inputting a first dialogue data generation request into the current large dialogue generation model, and generating a first candidate dialogue response through the current large dialogue generation model; inputting a second dialogue data generation request into the current large dialogue generation model, and generating a second candidate dialogue response through the current large dialogue generation model; and inputting a third dialogue data generation request into the current large dialogue generation model, and generating a third candidate dialogue response through the current large dialogue generation model.
[0086] The current large-scale dialogue generation model can be any one of the large-scale dialogue generation models corresponding to different architectures.
[0087] Based on the response quality criteria shown in the table above, it can be understood that the first candidate dialogue response is a candidate dialogue response with high emotionality; the second candidate dialogue response is a candidate dialogue response with high naturalness; and the third candidate dialogue response is a candidate dialogue response with low quality.
[0088] In practical applications, the three types of prompt words (i.e., the first prompt word type, the second prompt word type, and the third prompt word type) are combined with the dialogue context (i.e., the original dialogue data (dialogue history) and the original dialogue responses (the last round of responses)) to obtain three types of dialogue data generation requests corresponding to the three types of prompt words (i.e., the first dialogue data generation request, the second dialogue data generation request, and the third dialogue data generation request). The three types of dialogue data generation requests are then input into the current large dialogue generation model in an asynchronous and concurrent manner. The current large dialogue generation model generates three types of candidate dialogue responses (i.e., the first candidate dialogue response, the second candidate dialogue response, and the third candidate dialogue response). The three types of dialogue data generation requests are then input into other large dialogue generation models, so that each large dialogue generation model generates its own three types of candidate dialogue responses.
[0089] For example, if three large dialogue generation models are selected, then the "dialogue history + last round role / content" (i.e., original dialogue data + original dialogue response, i.e., dialogue context) is combined with three types of prompt words to generate three dialogue data generation requests for each of the three large dialogue generation models, for a total of 3×3=9 requests; these nine requests are then asynchronously and concurrently processed through the large dialogue generation model API to improve generation efficiency and avoid sequential waiting; finally, each of the three large dialogue generation models can generate three types of candidate dialogue responses, covering high quality (including high sentimentality and high naturalness) and low quality, for a total of nine candidate dialogue responses.
[0090] Furthermore, after each large dialogue generation model generates three types of candidate dialogue responses, post-processing can be performed on each candidate dialogue response. By cleaning the data, meta-information such as original dialogue data, "role prefixes" (such as "assistant:"), and explanatory text in the candidate dialogue responses can be removed to obtain clean candidate dialogue responses, so as to retain "clean response content".
[0091] Therefore, the data structure output by the first stage of "Diverse Response Generation Based on Multiple Models and Multiple Prompts" can be "Dialogue Context + Candidate Dialogue Response Mapping" (e.g., "LLM 1_Prompt A"), where the "key" of the mapping is "Dialogue Generation Model + Prompt Word Type", and the "value" is the clean response content after post-processing of the candidate dialogue response.
[0092] In this embodiment, the technical solution constructs multiple types of prompt words with various response quality standards, generates dialogue data generation requests corresponding to different types of prompt words, and inputs these requests into multiple large-scale dialogue generation models to generate candidate dialogue responses covering different quality standards, including "high-quality positive samples + low-quality negative samples." This allows the generated dialogue responses to cover diverse human dialogue habits across various scenarios. By combining "model diversity" with "prompt guidance," and by constructing a systematic positive and negative sample generation mechanism, the "emotion / style standards" of high-quality samples and the "negative features" of low-quality samples are accurately defined. This enables the large-scale dialogue generation models to clearly learn "preference boundaries," thereby improving the diversity and representativeness of candidate dialogue responses. It ensures the richness and relevance of candidate dialogue responses, providing abundant samples for subsequent preference evaluation and selection. This effectively solves the problems of insufficient candidate response diversity and ambiguous sample guidance in traditional solutions.
[0093] In an exemplary embodiment, obtaining the original dialogue data and the original dialogue response corresponding to the original dialogue data includes: obtaining sample dialogue data containing multiple rounds of dialogue; using the last round of dialogue data in the sample dialogue data as the original dialogue response, and using the dialogue data in the sample dialogue data other than the last round of dialogue data as the original dialogue data.
[0094] Based on the above description, the sample dialogue data can be high-quality multi-turn dialogue fragments that have undergone preliminary cleaning and screening.
[0095] The original dialogue data can be the dialogue data obtained after extracting the core information from the sample dialogue data, excluding the last round of dialogue data, such as all historical interaction content other than the last round of dialogue data, which can also be called dialogue history.
[0096] The original dialogue response can be the last round of dialogue data after extracting the core information from the sample dialogue data. It can also be called the last round response or the last round role / content, such as the role (e.g., "user" or "assistant") and content (original response text) of the last round of response during the interaction.
[0097] In practical applications, the terminal can acquire sample dialogue data containing multiple rounds of dialogue, extract core information from the sample dialogue data, take the last round of dialogue data in the sample dialogue data as the original dialogue response, and take the dialogue data in the sample dialogue data other than the last round of dialogue data as the original dialogue data.
[0098] The technical solution of this embodiment extracts core information from the sample dialogue data, uses the last round of dialogue data in the sample dialogue data as the original dialogue response, and uses the dialogue data in the sample dialogue data other than the last round of dialogue data as the original dialogue data, thus clarifying the specific content of the original dialogue data and the original dialogue response, thereby providing high-quality input data for the subsequent generation of dialogue responses from multiple models.
[0099] In an exemplary embodiment, each candidate dialogue response and the original dialogue response are used as the dialogue responses to be evaluated corresponding to the original dialogue data to input multiple pre-built large dialogue evaluation models corresponding to different architectures. Each large dialogue evaluation model outputs a dialogue preference evaluation result for each dialogue response to be evaluated. This includes: inputting each dialogue response to be evaluated into the current large dialogue evaluation model; obtaining the target dialogue response to be evaluated from each dialogue response to be evaluated through the current large dialogue evaluation model; and using the target dialogue response to be evaluated as the dialogue preference evaluation result output by the current large dialogue evaluation model for each dialogue response to be evaluated.
[0100] Based on the dialogue preference evaluation results output by each dialogue evaluation model, dialogue preference data of the original dialogue data is generated, including: dialogue preference data of the original dialogue data generated based on the target dialogue responses selected by each dialogue evaluation model.
[0101] The current dialogue evaluation model is any one of the various dialogue evaluation models.
[0102] The target dialogue response to be evaluated can be the preferred dialogue response selected by the current dialogue evaluation model from the dialogue responses to be evaluated.
[0103] In practice, the terminal can combine the candidate dialogue responses and the original dialogue responses to form dialogue responses to be evaluated. Then, the dialogue responses to be evaluated are input into the current dialogue evaluation model of any one of the large dialogue evaluation models. Based on the pre-built dialogue evaluation prompts, the current dialogue evaluation model evaluates and selects each dialogue response to be evaluated to obtain the target dialogue response to be evaluated. Finally, the target dialogue response to be evaluated is used as the dialogue preference evaluation result output by the current dialogue evaluation model for each dialogue response to be evaluated.
[0104] Furthermore, each dialogue evaluation model evaluates and selects the responses to each dialogue to be evaluated, obtaining the target responses to be evaluated selected by each dialogue evaluation model (i.e., dialogue preference evaluation results). The target responses to be evaluated by each dialogue evaluation model are then summarized to generate dialogue preference data, which includes the dialogue context (i.e., the original dialogue data and the original dialogue responses), each dialogue response to be evaluated, and the dialogue preference evaluation results.
[0105] For example, continuing with the example of three large dialogue generation models, in the second stage of preference selection and data construction based on multi-model voting, four large dialogue evaluation models are selected, which can overlap with the large dialogue generation models.
[0106] After obtaining 9 candidate dialogue responses in the first stage, the original dialogue responses (last-round responses) from the last round of dialogue data in the sample dialogue data can be added to form a list of 10 candidate responses, which serve as the dialogue responses to be evaluated corresponding to the original dialogue data. Then, the 10 candidate response lists (each dialogue response to be evaluated) are numbered in a fixed order (0-8 are generated candidate dialogue responses, 9 is the original dialogue response) to format the candidate response lists, forming a candidate response list index. The formatted candidate response lists, along with the dialogue context, are then input into each evaluator model (the large dialogue evaluation model).
[0107] Four evaluator models (the large dialogue evaluation model) can evaluate and vote on a list of 10 candidate responses (each dialogue response to be evaluated) based on pre-built dialogue evaluation prompts. Each evaluator model outputs one target dialogue response to be evaluated (i.e., the dialogue preference evaluation result) and can output a selection index (for example, if evaluator model 1 selects dialogue response 3 to be evaluated, then evaluator model 1 outputs "Selection: 3"; if evaluator model 2 selects dialogue response 5 to be evaluated, then evaluator model 2 outputs "Selection: 5").
[0108] Furthermore, based on the target dialogue responses selected by each evaluator model, the selection indices of the four evaluator models are summarized to form an "evaluator selection index list" (which is also the dialogue preference evaluation result obtained by summarizing the target dialogue responses of each large dialogue evaluation model), for example, [3,5,3,4].
[0109] Then, the dialogue context (including the original dialogue data and the original dialogue responses), the list of 10 candidate responses (i.e., each dialogue response to be evaluated, with an index number consistent with the evaluation), and the evaluator selection index list (i.e., the dialogue preference evaluation results, recording the voting evaluation results of each evaluator model) are combined to generate the dialogue preference data of the original dialogue data.
[0110] The technical solution of this embodiment employs multiple large-scale dialogue evaluation models to evaluate and select responses to each dialogue to be evaluated, obtaining the target dialogue responses selected by each large-scale dialogue evaluation model. This multi-model evaluation strategy reduces the subjectivity of a single subject (human or model) and improves the objectivity and robustness of preference judgments. As a result, the final output of the second stage of "preference selection and data construction based on multi-model voting" is obtained, namely, the generated dialogue preference data. This provides an important data foundation for the subsequent extraction of positive and negative samples of dialogue preferences based on the dialogue preference data, so that it can be directly applied to the reinforcement learning training of LLM.
[0111] It should be noted that dialogue preference data can be used to characterize positive samples of dialogue preferences that are responses to the original dialogue data, or negative samples of dialogue preferences that are responses to the original dialogue data.
[0112] Therefore, in an exemplary embodiment, based on the target dialogue responses selected by each large dialogue evaluation model, dialogue preference data of the original dialogue data is generated, including: if the current dialogue response to be evaluated is selected as the target dialogue response to be evaluated by any large dialogue evaluation model, then the current dialogue response to be evaluated is used as a positive sample of dialogue preference in the original dialogue data; if the current dialogue response to be evaluated is not selected as the target dialogue response to be evaluated by any large dialogue evaluation model, then the current dialogue response to be evaluated is used as a negative sample of dialogue preference in the original dialogue data; and the positive and negative samples of dialogue preference in the original dialogue data are used as the dialogue preference data of the original dialogue data.
[0113] The current dialogue response to be evaluated is any one of the dialogue responses to be evaluated.
[0114] For example, the terminal can extract "preference pairs (yw, yl)" from the "evaluator selection index list" (i.e., dialogue preference evaluation results). If the current dialogue response to be evaluated is selected as the target dialogue response by any large dialogue evaluation model, then the current dialogue response to be evaluated is taken as the positive dialogue preference sample yw of the original dialogue data, representing the "preference response" selected by the evaluator model (such as the responses corresponding to indices 3 and 5); if the current dialogue response to be evaluated is not selected as the target dialogue response by any large dialogue evaluation model, then the current dialogue response to be evaluated is taken as the negative dialogue preference sample yl of the original dialogue data, representing the "non-preference response" that was not selected (such as the responses corresponding to indices 0 and 1).
[0115] Then, the "preference pair (yw, yl)" consisting of the positive dialogue preference sample yw and the negative dialogue preference sample yl can be used as the "reward signal" of the original dialogue data to form the final dialogue preference data used as input to LLM for RL training.
[0116] Furthermore, in an exemplary embodiment, after generating the dialogue preference data of the original dialogue data, the method further includes: inputting the original dialogue data and the dialogue preference data as a reward signal into the dialogue response generation model to be trained, so as to perform reinforcement learning training on the dialogue response generation model.
[0117] In practice, the terminal can use the extracted preference pair (yw, yl) as the "reward signal" of the dialogue preference data, and input the original dialogue data and the dialogue preference data as the reward signal into the dialogue response generation model to be trained, so as to perform reinforcement learning training on the dialogue response generation model, thereby guiding the RL algorithm of the LLM model to generate responses that are more in line with human preferences.
[0118] The technical solution of this embodiment extracts positive and negative dialogue preference samples from the target dialogue responses selected by each large dialogue evaluation model. These positive and negative samples are then used as reward signals and input into the dialogue response generation model to be trained for reinforcement learning training. This eliminates the need for additional training of a separate RM (reward model), thus simplifying the RL (reinforcement learning) process.
[0119] In an exemplary embodiment, each dialogue response to be evaluated is input into the current dialogue evaluation model. The target dialogue response to be evaluated is obtained from each dialogue response to be evaluated through the current dialogue evaluation model. This includes: inputting each dialogue response to be evaluated, as well as pre-built dialogue evaluation prompts, into the current dialogue evaluation model; using the role represented by the role information in the current dialogue evaluation model, performing the task of selecting a dialogue response according to the dialogue response selection criteria information and dialogue response selection rule information, and using the dialogue response to be evaluated selected by the current dialogue evaluation model as the target dialogue response to be evaluated.
[0120] The dialogue evaluation prompts can include role information, dialogue response selection criteria information, and dialogue response selection rules information, and can also be called "professional dialogue quality evaluation instructions" or "dialogue selection instructions". The dialogue response selection rules information can be used to indicate that if each candidate dialogue response does not meet the dialogue response selection criteria information, the original dialogue response will be used as the target dialogue response to be evaluated.
[0121] For example, dialogue evaluation prompts can be designed as follows: Role definition (role information): "Professional dialogue quality assessment expert (doctoral supervisor level)"; Selection criteria (dialogue response selection criteria information): Immerse yourself in the dialogue context as a "high EQ chat expert" and select a response that is "concise and clear, humorous / witty, and without machine translation feel"; Special rules (dialogue response selection rule information): If all candidate dialogue responses are not suitable for the dialogue response selection criteria, select the "original dialogue response" (default is set to option 9, i.e., the last round of response); Output format: The selected candidate response list index can be output, i.e., the selection index (0-9), such as "Selection: 5", without any explanation.
[0122] In practical applications, the terminal can input each dialogue response to be evaluated, as well as pre-built dialogue evaluation prompts, into the current dialogue evaluation model. Through the current dialogue evaluation model, the role represented by the role information of the dialogue evaluation prompts is used to evaluate and select each dialogue response to be evaluated according to the dialogue response selection criteria information and dialogue response selection rule information. The selection index can be output according to the output format specified by the dialogue evaluation prompts, and the dialogue response to be evaluated selected by the current dialogue evaluation model can be used as the target dialogue response to be evaluated.
[0123] The technical solution of this embodiment, based on pre-constructed dialogue evaluation prompts containing role information, dialogue response selection criteria information, and dialogue response selection rule information, can enable each large dialogue evaluation model to use the role represented by the role information to accurately perform the task of selecting dialogue responses according to the dialogue response selection criteria information and dialogue response selection rule information. This allows the large dialogue evaluation model to simulate human group evaluation, accurately judge and select the dialogue responses to be evaluated, and thus obtain objective preference selection.
[0124] In one possible implementation, specifically, such as Figure 3 As shown, this application also provides another method for constructing dialogue preference data, including:
[0125] Phase 1: Diverse response generation.
[0126] Extracting core dialogue information: Based on the high-quality multi-turn dialogues (sample dialogue data) that have undergone preliminary cleaning and filtering, the core information in the sample dialogue data is extracted to obtain the original dialogue data and the original dialogue responses.
[0127] Multi-model and multi-coma design: The original dialogue data and original dialogue responses (i.e., "dialogue history + last round response") are combined with three pre-built coma types corresponding to different coma types to generate three types of dialogue data generation requests for each of the three dialogue generation models, resulting in a total of 9 requests.
[0128] Candidate dialogue responses are generated concurrently: These 9 requests are asynchronously and concurrently sent through the dialogue generation big model API. Each dialogue generation big model generates three types of candidate dialogue responses covering "high quality (including high sentiment and high naturalness) and low quality", for a total of 9 candidate dialogue responses.
[0129] Post-processing yields clean candidates: Post-processing of each candidate dialogue response can be performed by data cleaning to remove meta-information such as original dialogue data, "role prefixes" (such as "assistant:"), and explanatory text that may exist in the candidate dialogue responses, so as to retain the clean response content and obtain clean candidate dialogue responses.
[0130] Thus, the first stage outputs "dialogue context + 9 candidate dialogue response mappings", where the "key" of the mapping is "dialogue generation big model + prompt word type", and the "value" is the clean response content after post-processing of the candidate dialogue responses.
[0131] Phase Two: Preference Selection and Data Construction.
[0132] Merging candidate dialogue responses: Based on the 9 candidate dialogue responses, supplement them with the original dialogue responses (i.e., the last round responses) that include the last round of dialogue data from the sample dialogue data, resulting in a list of 10 candidate responses, which will serve as the dialogue responses to be evaluated corresponding to the original dialogue data.
[0133] Multiple evaluator model selection: Four LLMs with different architectures or training backgrounds are preferred as the large dialogue evaluation model (evaluator model), which can overlap with the large dialogue generation model.
[0134] Evaluation prompt loading and voting: Dialogue evaluation prompts can be pre-built, containing role information, dialogue response selection criteria, and dialogue response selection rules. Then, each dialogue response to be evaluated (a list of 10 candidate responses) and the pre-built evaluation prompts can be input into the current large-scale dialogue evaluation model. Each large-scale model uses the role information represented by the dialogue evaluation prompts to evaluate and select responses according to the dialogue response selection criteria and rules. Each large-scale model can output a selection index; for example, evaluator model 1 selects dialogue response 3, and evaluator model 2 selects dialogue response 5.
[0135] Collect evaluator model selection indices: Based on the target dialogue responses selected by each evaluator model, the selection indices of the four evaluator models are summarized to form an "evaluator selection index list" (i.e., dialogue preference evaluation results). At the same time, dialogue responses selected as target responses by any large dialogue evaluation model (i.e., positive dialogue preference samples yw) and dialogue responses not selected as target responses by any large dialogue evaluation model (i.e., negative dialogue preference samples yl) can be extracted from the evaluator selection index list. The positive dialogue preference samples yw and the negative dialogue preference samples yl are combined to form a "preference pair (yw, yl)". Thus, the preference pair (yw, yl) can be used as a "reward signal" for dialogue preference data to guide the RL algorithm of the LLM model to train and generate responses that are more in line with human preferences.
[0136] Finally, the dialogue context (including the original dialogue data and the original dialogue responses), the list of 10 candidate responses (i.e., each dialogue response to be evaluated), and the evaluator selection index list (i.e., the dialogue preference evaluation results) are combined to generate the dialogue preference data of the original dialogue data.
[0137] Therefore, the second stage outputs complete dialogue preference data.
[0138] Based on the first stage of generating diverse responses and the second stage of preference selection and data construction, this application finally outputs dialogue preference data (including dialogue context, responses to be evaluated, and dialogue preference evaluation results) for reinforcement learning training.
[0139] Compared with the "traditional RLHF manual-assisted data generation scheme", this application has the following significant advantages as shown in Table 2:
[0140] Table 2 Comparison between this application and conventional technologies
[0141]
[0142] It should be noted that the specific limitations of the above steps can be found in the specific limitations of a dialogue preference data construction method described above.
[0143] In addition, this application also provides alternative solutions to some embodiments of this application described above, as follows:
[0144] For example, in the first stage of generating diverse responses, a single model plus more diverse prompts can be used: for instance, a high-performance LLM can be used, but more numerous and dimensional prompts (e.g., 5-10 types) can be designed for it. The diversity of prompts can make up for the lack of model uniformity, but the purpose of expanding the diversity of candidate set can still be achieved.
[0145] Different prompts can be used instead of different ones by controlling the diversity of sampling parameters: the same type and prompt can be used, but the diverse dialogue responses can be generated by significantly adjusting the decoding parameters (such as temperature, top-p) (high temperature may generate random creative responses of poor quality, while low temperature may generate deterministic responses of high quality), thereby obtaining a candidate set with a wide quality distribution.
[0146] For the second stage of preference selection evaluation, a dedicated evaluation model can be trained instead of directly using multiple LLMs for voting. This involves first training a dedicated "evaluator" model (a classifier or reward model) with a small amount of manually labeled data, and then using this single trained model to score or rank all candidate dialogue responses. This ensures high consistency in evaluation criteria but increases the model training cost.
[0147] Candidate dialogue responses can be compared pairwise: instead of a "multiple-choice" approach, the evaluator model compares each candidate response pairwise, determining which is better in each pair, and finally ranking all candidates by statistical win rate. This method better fits common human annotation methods, but the computational cost increases exponentially (possibly quadratically) with the number of candidates.
[0148] For the overall process of "diverse response generation - preference selection evaluation - training data construction", end-to-end reward modeling can be used: instead of explicitly generating preference pairs, one or more models can directly output continuous reward scores based on the dialogue context and candidate dialogue responses; and in the second stage, one or more models score each candidate dialogue response based on its reward score and take the average score, and finally select high-scoring responses as positive samples and low-scoring responses as negative samples.
[0149] It should be noted that while these alternatives may be effective in specific scenarios, they typically come at the cost of sacrificing some advantage of the embodiments described in this application (e.g., a single model may lack diversity; training a dedicated evaluator increases complexity and cost; pairwise comparisons are computationally expensive). The combination of "multi-model + multi-hint" generation and "multi-model voting" evaluation in this application achieves a balance between efficiency, cost, and effectiveness.
[0150] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0151] Based on the same inventive concept, this application also provides a dialogue preference data construction apparatus for implementing the dialogue preference data construction method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the dialogue preference data construction apparatus provided below can be found in the limitations of the dialogue preference data construction method described above, and will not be repeated here.
[0152] In one exemplary embodiment, such as Figure 4 As shown, a device for constructing dialogue preference data is provided, comprising:
[0153] The original dialogue acquisition module 410 is used to acquire the original dialogue data and the original dialogue responses corresponding to the original dialogue data.
[0154] The candidate response generation module 420 is used to construct dialogue data generation requests from the original dialogue data, the original dialogue response, and a number of pre-built prompt words corresponding to different prompt word types. Each dialogue data generation request is input into a number of pre-built dialogue generation models corresponding to different architectures, and candidate dialogue responses that match each dialogue data generation request are generated through each dialogue generation model.
[0155] The dialogue preference evaluation module 430 is used to input each candidate dialogue response and the original dialogue response as the dialogue response to be evaluated corresponding to the original dialogue data into multiple pre-built large dialogue evaluation models corresponding to different architectures, and output the dialogue preference evaluation results for each dialogue response to be evaluated through each large dialogue evaluation model.
[0156] The dialogue preference generation module 440 is used to generate dialogue preference data of the original dialogue data based on the dialogue preference evaluation results output by each dialogue evaluation model.
[0157] In an exemplary embodiment, the candidate response generation module 420 is further configured to input a first dialogue data generation request into the current dialogue generation model and generate a first candidate dialogue response through the current dialogue generation model; input a second dialogue data generation request into the current dialogue generation model and generate a second candidate dialogue response through the current dialogue generation model; and input a third dialogue data generation request into the current dialogue generation model and generate a third candidate dialogue response through the current dialogue generation model.
[0158] In an exemplary embodiment, the candidate response generation module 420 is further configured to acquire sample dialogue data containing multiple rounds of dialogue; take the last round of dialogue data in the sample dialogue data as the original dialogue response, and take the dialogue data in the sample dialogue data other than the last round of dialogue data as the original dialogue data.
[0159] In an exemplary embodiment, the dialogue preference evaluation module 430 is specifically used to input each dialogue response to be evaluated into the current dialogue evaluation model, obtain the target dialogue response to be evaluated from each dialogue response to be evaluated through the current dialogue evaluation model, and use the target dialogue response to be evaluated as the dialogue preference evaluation result output by the current dialogue evaluation model for each dialogue response to be evaluated; the target dialogue response to be evaluated is the preferred dialogue response selected by the current dialogue evaluation model from the dialogue responses to be evaluated.
[0160] In an exemplary embodiment, the dialogue preference generation module 440 is specifically used to generate dialogue preference data of the original dialogue data based on the target dialogue responses selected by each dialogue evaluation big model.
[0161] In an exemplary embodiment, the dialogue preference generation module 440 is further configured to: if the current dialogue response to be evaluated is selected as the target dialogue response to be evaluated by any large dialogue evaluation model, then use the current dialogue response to be evaluated as a positive sample of dialogue preferences in the original dialogue data; if the current dialogue response to be evaluated is not selected as the target dialogue response to be evaluated by any large dialogue evaluation model, then use the current dialogue response to be evaluated as a negative sample of dialogue preferences in the original dialogue data; and use the positive and negative samples of dialogue preferences in the original dialogue data as the dialogue preference data of the original dialogue data.
[0162] In an exemplary embodiment, the dialogue preference evaluation module 430 is further configured to input each dialogue response to be evaluated and pre-built dialogue evaluation prompts into the current dialogue evaluation model; the dialogue evaluation prompts include role information, dialogue response selection criteria information, and dialogue response selection rule information; the dialogue response selection rule information is used to indicate that if each candidate dialogue response does not meet the dialogue response selection criteria information, the original dialogue response will be used as the target dialogue response to be evaluated; the current dialogue evaluation model uses the role represented by the role information to perform the task of selecting a dialogue response according to the dialogue response selection criteria information and the dialogue response selection rule information, and uses the dialogue response to be evaluated selected by the current dialogue evaluation model as the target dialogue response to be evaluated.
[0163] Each module in the aforementioned dialogue preference data construction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0164] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data for a dialogue preference data construction method. The input / output interface allows the processor to exchange information with external devices. The communication interface allows wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a dialogue preference data construction method. The display unit of the computer device forms a visually visible image and can be a display screen, projection device, or virtual reality imaging device.
[0165] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0166] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the various embodiments of the dialogue preference data construction method described above.
[0167] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the various embodiments of the dialogue preference data construction method described above.
[0168] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the embodiments of the dialogue preference data construction method described above.
[0169] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0170] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0171] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0172] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for constructing dialogue preference data, characterized in that, The method includes: Obtain the original dialogue data and the original dialogue responses corresponding to the original dialogue data; The original dialogue data, the original dialogue response, and a number of pre-built prompt words corresponding to different prompt word types are used to construct dialogue data generation requests. Each dialogue data generation request is input into a number of pre-built large dialogue generation models corresponding to different architectures. Candidate dialogue responses that match each dialogue data generation request are generated through each of the large dialogue generation models. Each candidate dialogue response and the original dialogue response are used as inputs to multiple pre-constructed large dialogue evaluation models corresponding to different architectures, with each of the candidate dialogue responses and the original dialogue response being used as inputs to the dialogue response to be evaluated. The dialogue preference evaluation results for each dialogue response to be evaluated are then output by each of the large dialogue evaluation models. Based on the dialogue preference evaluation results output by each of the aforementioned dialogue evaluation models, dialogue preference data of the original dialogue data is generated.
2. The method according to claim 1, characterized in that, The prompt word types include: a first prompt word type, a second prompt word type, and a third prompt word type; wherein, the prompt words of the first and second prompt word types are used to prompt the generation of high-quality dialogue responses, and the prompt words of the third prompt word type are used to prompt the generation of low-quality dialogue responses; the dialogue data generation request includes: a first dialogue data generation request, a second dialogue data generation request, and a third dialogue data generation request, wherein the first dialogue data generation request is generated based on the prompt words of the first prompt word type, the second dialogue data generation request is generated based on the prompt words of the second prompt word type, and the third dialogue data generation request is generated based on the prompt words of the third prompt word type; The step of inputting each of the dialogue data generation requests into multiple pre-built large dialogue generation models corresponding to different architectures, and generating candidate dialogue responses that match each of the dialogue data generation requests through each of the large dialogue generation models, includes: The first dialogue data generation request is input into the current dialogue generation model, and the current dialogue generation model generates a first candidate dialogue response; the first candidate dialogue response is a candidate dialogue response with high sentiment, and the current dialogue generation model is any one of the dialogue generation models corresponding to different architectures; The second dialogue data generation request is input into the current dialogue generation model, and the current dialogue generation model generates a second candidate dialogue response; the second candidate dialogue response is a candidate dialogue response with high naturalness. The request to generate the third dialogue data is input into the current dialogue generation model, and the current dialogue generation model generates a third candidate dialogue response; the third candidate dialogue response is a low-quality candidate dialogue response.
3. The method according to claim 2, characterized in that, The acquisition of the original dialogue data and the corresponding original dialogue responses includes: Obtain sample dialogue data containing multiple rounds of dialogue; The last round of dialogue data in the sample dialogue data is used as the original dialogue response, and the dialogue data in the sample dialogue data other than the last round of dialogue data is used as the original dialogue data.
4. The method according to claim 1, characterized in that, The process involves inputting multiple pre-constructed large-scale dialogue evaluation models, each corresponding to a different architecture, with the candidate dialogue responses and the original dialogue responses as inputs to the dialogue responses to be evaluated corresponding to the original dialogue data. Each of these large-scale dialogue evaluation models outputs dialogue preference evaluation results for each of the dialogue responses to be evaluated, including: Each of the dialogue responses to be evaluated is input into the current dialogue evaluation model, and the target dialogue response to be evaluated is obtained from each of the dialogue responses to be evaluated through the current dialogue evaluation model; the current dialogue evaluation model can be any one of the dialogue evaluation models. The target dialogue response to be evaluated is used as the dialogue preference evaluation result output by the current dialogue evaluation model for each of the dialogue responses to be evaluated; the target dialogue response to be evaluated is the preferred dialogue response selected by the current dialogue evaluation model from the dialogue responses to be evaluated; The dialogue preference data generated from the original dialogue data, based on the dialogue preference evaluation results output by each of the aforementioned dialogue evaluation models, includes: Based on the target dialogue responses selected by each of the aforementioned dialogue evaluation models, dialogue preference data of the original dialogue data is generated.
5. The method according to claim 4, characterized in that, The dialogue preference data is used to characterize whether the dialogue response to be evaluated is a positive sample of the dialogue preference in the original dialogue data, or a negative sample of the dialogue preference in the original dialogue data. The process of generating dialogue preference data from the original dialogue data based on the target dialogue responses selected by each of the aforementioned large dialogue evaluation models includes: If the current dialogue response to be evaluated is selected as the target dialogue response to be evaluated by any of the large dialogue evaluation models, then the current dialogue response to be evaluated is used as a positive sample of dialogue preference in the original dialogue data; the current dialogue response to be evaluated is any one of the dialogue responses to be evaluated. If the current dialogue response to be evaluated is not selected as the target dialogue response to be evaluated by any of the large dialogue evaluation models, then the current dialogue response to be evaluated is used as a negative sample of the dialogue preference in the original dialogue data. The positive and negative samples of dialogue preferences from the original dialogue data are used as the dialogue preference data of the original dialogue data.
6. The method according to claim 4, characterized in that, The step of inputting each of the dialogue responses to be evaluated into the current dialogue evaluation model, and obtaining the target dialogue response to be evaluated from each of the dialogue responses to be evaluated through the current dialogue evaluation model, includes: Each of the dialogue responses to be evaluated, along with pre-constructed dialogue evaluation prompts, are input into the current dialogue evaluation model. The dialogue evaluation prompts include role information, dialogue response selection criteria information, and dialogue response selection rule information. The dialogue response selection rule information is used to indicate that if each of the candidate dialogue responses does not meet the dialogue response selection criteria information, the original dialogue response will be used as the target dialogue response to be evaluated. The current dialogue evaluation model uses the role represented by the role information to perform the task of selecting a dialogue response according to the dialogue response selection criteria information and the dialogue response selection rule information, and takes the dialogue response to be evaluated selected by the current dialogue evaluation model as the target dialogue response to be evaluated.
7. The method according to any one of claims 1 to 6, characterized in that, After generating the dialogue preference data from the original dialogue data, the method further includes: The original dialogue data, along with the dialogue preference data serving as a reward signal, are input into the dialogue response generation model to be trained, so as to perform reinforcement learning training on the dialogue response generation model.
8. A device for constructing dialogue preference data, characterized in that, The device includes: The original dialogue acquisition module is used to acquire the original dialogue data and the original dialogue responses corresponding to the original dialogue data. The candidate response generation module is used to construct dialogue data generation requests from the original dialogue data, the original dialogue response, and a number of pre-built prompt words corresponding to different prompt word types, and input each of the dialogue data generation requests into a number of pre-built dialogue generation models corresponding to different architectures, and generate candidate dialogue responses that match each of the dialogue data generation requests through each of the dialogue generation models. The dialogue preference evaluation module is used to input multiple pre-built large dialogue evaluation models corresponding to different architectures into the candidate dialogue responses and the original dialogue responses as the dialogue responses to be evaluated corresponding to the original dialogue data, and output dialogue preference evaluation results for each of the dialogue responses to be evaluated through each of the large dialogue evaluation models. The dialogue preference generation module is used to generate dialogue preference data of the original dialogue data based on the dialogue preference evaluation results output by each of the large dialogue evaluation models.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.