Preference data sample set construction method and device, equipment, medium and product
By using automated annotation technology and multiple question-answering models and evaluation dimensions, the problem of low efficiency in constructing question-answering preference data sample sets was solved, and efficient data sample set generation was achieved.
Patent Information
- Application Number
- CN202411150868.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies for constructing question-answer preference data sample sets are inefficient, mainly relying on manual annotation, which leads to resource waste and low efficiency.
By acquiring answers from multiple question-answering models, combining annotation instructions and evaluation dimensions, and using the target model for automated annotation, a preference data sample set is constructed, including evaluations of authenticity, relevance, usefulness, harmlessness, and conciseness, generating triplet samples.
It enables automated annotation of question-answer preference data samples, saving annotation resources and improving the efficiency of data sample set construction.
Smart Images

Figure CN121597787A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of neural network technology, and in particular relates to a method, apparatus, device, medium and product for constructing a preference data sample set. Background Technology
[0002] A data sample set is a collection of data samples used to train a neural network model. By training the neural network model using data samples from the data sample set, a model that meets the user's needs can be obtained, such as a face recognition model, a speech recognition model, and a question-answering preference model.
[0003] In related technologies, when constructing the preference data sample set for a question-answer preference model, manual annotation is used to label the answers to questions to obtain the question-answer preference data sample set. However, constructing the question-answer preference data sample set using manual annotation is inefficient. Summary of the Invention
[0004] This application provides a method, apparatus, device, medium, and product for constructing a preference data sample set, which can solve the problem of low efficiency in constructing preference data sample sets.
[0005] In a first aspect, embodiments of this application provide a method for constructing a preference data sample set, including:
[0006] Obtain at least two answers from at least two question-answering models, each addressing the target question.
[0007] The target question and at least two answers, along with annotation instructions, are input into the target model to obtain the annotation results of the target model annotating at least two answers respectively;
[0008] Based on the annotation results, at least two answers, and the target question, construct a sample set of preference data.
[0009] Secondly, embodiments of this application provide a preference data sample set construction apparatus, comprising:
[0010] The acquisition module is used to acquire at least two answers obtained by at least two question-answering models respectively answering the target question;
[0011] The annotation module is used to input the target question and at least two answers, along with annotation instructions, into the target model, and obtain the annotation results of the target model annotating at least two answers respectively;
[0012] The building module is used to construct a sample set of preference data based on the annotation results, at least two answers, and the target question.
[0013] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the steps of the preference data sample set construction method provided in embodiments of this application.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the steps of the preference data sample set construction method provided in embodiments of this application.
[0015] Fifthly, embodiments of this application provide a computer program product, the computer program product including computer program instructions, which, when executed by a processor, implement the steps of the preference data sample set construction method provided in embodiments of this application.
[0016] In this embodiment, at least two answers are obtained by at least two question-answering models respectively answering the target question; the target question and at least two answers, combined with annotation instructions, are input into the target model to obtain annotation results of the target model annotating the at least two answers respectively; based on the annotation results, the at least two answers, and the target question, a preference data sample set is constructed. Compared with the manual annotation method used in related technologies to construct the preference data sample set, this method can automatically annotate the question-answer preference data samples, saving annotation resources and improving the efficiency of preference data sample set construction. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the method for constructing a preference data sample set provided in an embodiment of this application;
[0019] Figure 2 This is a schematic diagram of the architecture for constructing a sample set of preference data provided in an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of the structure of the preference data sample set construction device provided in the embodiments of this application;
[0021] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0022] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.
[0024] The following description, in conjunction with the accompanying drawings, details the method, apparatus, device, medium, and product for constructing a preference data sample set provided in this application, through specific embodiments and application scenarios.
[0025] In some possible implementations of the embodiments of this application, the preference data sample set construction method and apparatus provided in the embodiments of this application can be applied to electronic devices. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc., and can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM or self-service machine, etc. The embodiments of this application do not specifically limit the scope.
[0026] Figure 1 This is a flowchart illustrating the method for constructing a preference data sample set provided in an embodiment of this application. Figure 1 As shown, methods for constructing preference data sample sets may include:
[0027] Step 101: Obtain at least two answers from at least two question-answering models, each answering the target question;
[0028] In some possible implementations of the embodiments of this application, the target question can be input into at least two question-answering models, and the at least two question-answering models can output answers to the target question respectively.
[0029] Step 102: Input the target question and at least two answers, along with the annotation instructions, into the target model to obtain the annotation results of the target model annotating at least two answers respectively;
[0030] In some possible implementations of the embodiments of this application, the annotation instructions may include role instructions, evaluation criteria for at least two evaluation dimensions, input instructions, and output criteria for annotation results; wherein, the role instructions are used to prompt the target model for its role and at least two evaluation dimensions for evaluating the answer, the input instructions are used to prompt the target model for recognizing the question and the answer, and the output criteria are used to specify the content and format of the target model's response.
[0031] For example, the role instruction is: This is an evaluation assistant used to evaluate the quality of the question answering model's answers to questions. The answers will be evaluated from five dimensions: authenticity, relevance, usefulness, harmlessness, and conciseness.
[0032] In some possible implementations of the embodiments of this application, at least two evaluation dimensions include at least two of the following: authenticity, relevance, usefulness, harmlessness, and simplicity.
[0033] In some possible implementations of the embodiments of this application, the authenticity of the answer means that the information provided in the answer should be accurate and must be based on credible facts and data. This information can be verified through certain authoritative websites, and the answer should avoid content that is obviously inconsistent with objective facts.
[0034] The criteria for assessing authenticity may include: a first score for a completely wrong answer; a second score for a partially wrong answer; and a third score for a completely correct answer. The first score is lower than the second score, and the second score is lower than the third score. For example, the first score might be 0, the second score 1, and the third score 2.
[0035] Specifically, if the part of the answer that is relevant to the question contains obvious factual errors that may seriously mislead the user, the answer is considered completely wrong and receives a score of 0 for authenticity. If the part of the answer that is relevant to the question is correct, but some details are incorrect, such as the city's area in a response to a question about "introducing a city," but the discrepancy is within an acceptable range, the answer is considered partially wrong and receives a score of 1 for authenticity. If the answer contains no errors, it is considered completely correct and receives a score of 2 for authenticity.
[0036] In some possible implementations of the embodiments of this application, the relevance of the answer refers to whether the answer directly and accurately addresses the question, and the answer should closely revolve around the question and not deviate from the topic. Relevance can be evaluated from two perspectives: semantic relevance and template consistency. Semantic relevance refers to whether the answer directly and accurately solves the problem; the answer should closely revolve around the question, not deviate from the topic, and provide the information that the user truly needs. Template consistency means that if a template is specified for the answer, the question must be answered according to the corresponding template format to ensure that the presentation of the answer conforms to the template format requirements.
[0037] The criteria for evaluating relevance may include: a first score when the answer does not address the question at all; a second score when the answer does not directly address the question; and a third score when the answer directly addresses the question.
[0038] Specifically, when an answer is completely unrelated to the question, the relevance score is 0; when an answer does not directly answer the question but touches on it in the middle of the answer, deviating somewhat from the topic; or when it does not fully follow the template requirements in the instructions but the main content still meets the requirements, the relevance score is 1; when an answer completely, clearly, and accurately answers the question and fully meets the given template requirements, such an answer is highly relevant to the question in terms of content and strictly follows the requirements in terms of format, and is the highest quality answer, with a relevance score of 2.
[0039] In some possible implementations of the embodiments of this application, the usefulness of an answer refers to its ability to provide a relevant, helpful, complete, and logically consistent answer to the question. Usefulness can be evaluated from three perspectives: information completeness, logical coherence, and content richness. Information completeness refers to whether the answer covers all the important information required for the question and whether any key aspects are omitted or not mentioned. Logical coherence refers to whether the information in the answer maintains logical consistency and whether there are reasonable connections and transitions between different parts. Content richness refers to the answer's ability to fully address the user's question while providing additional relevant information to increase the depth and breadth of the answer.
[0040] The evaluation criteria for usefulness may include: when the information in the answer is incomplete, the usefulness score is the first score; when the information in the answer is partially complete, the usefulness score is the second score; and when the information in the answer is completely complete, the usefulness score is the third score.
[0041] In this evaluation, answers are considered incomplete and receive a usefulness score of 0 if they are too brief, lack important information, are logically inconsistent, contain contradictions or irrelevant content that severely hinders user comprehension. Answers that cover the essential information required for the question, have no obvious logical errors, and maintain a certain degree of coherence are considered partially complete and receive a usefulness score of 1. Answers that comprehensively address the question, are logically consistent, and provide rich relevant information such as explanations and examples to help users gain a more comprehensive understanding are considered fully complete and receive a usefulness score of 2.
[0042] In some possible implementations of this application's embodiments, the harmlessness of an answer refers to the ability to politely refuse information involving specific, harmful, or sensitive content, and to refrain from giving offensive or discriminatory responses. Harmlessness primarily considers whether the content of the answer has a negative impact or causes harm to the user. An answer that conforms to the principle of harmlessness should avoid sensitive and inappropriate topics, ensuring that the content is healthy, positive, and constructive.
[0043] The criteria for assessing harmfulness may include: when the answer contains harmful content, the harmlessness score is the first score; when the answer does not contain harmful content, the harmlessness score is the second score.
[0044] For example, when an answer contains harmful content, the answer's harmlessness score is 0; when an answer does not contain harmful content, the answer's harmlessness score is 1.
[0045] In some possible implementations of the embodiments of this application, in terms of the conciseness of the answer, conciseness means that the answer should provide a direct and clear response to the question, focusing on the core of the question and avoiding lengthy explanations or unnecessary details.
[0046] The criteria for assessing conciseness may include: a first score for a severely non-concise answer; a second score for a slightly non-concise answer; and a third score for a concise answer.
[0047] For example, when an answer affects the user's efficiency in obtaining information, such as by repeating large sections of content, excessively elaborating on topics that are weakly or unrelated to the question, or using a complex list structure to answer the question, such as nesting multiple levels of items or having too many items, the answer is considered severely unconcise and receives a conciseness score of 0. When an answer has a minor impact on the user, such as by repeating a few sentences, the answer is considered slightly unconcise and receives a conciseness score of 1. When an answer is direct, relevant, clearly structured, well-controlled in terms of items, complete in information, and concise in language, the answer is considered concise and receives a conciseness score of 2.
[0048] In some possible implementations of this application's embodiments, the input instruction may include parameters corresponding to the target question and parameters corresponding to at least two answers. The parameters corresponding to the target question and the parameters corresponding to at least two answers may be arranged sequentially. The range of each parameter is marked using start and end markers.
[0049] For example, taking two answers as an example, the input instructions can be as follows:
[0050] [Target Question]
[0051] {Qurey}
[0052] [First answer begins]
[0053] {responseA}
[0054] [End of first answer]
[0055] [Start of the second answer]
[0056] {responseB}
[0057] [End of second answer]
[0058] In the input instructions shown above, Qurey represents the parameters of the target question, responseA represents the parameters of the first response, and responseB represents the parameters of the second response.
[0059] In some possible implementations of the embodiments of this application, step 102 may include: filling the target question and at least two answers into the parameter content corresponding to the input instruction in the annotation instruction to obtain input prompt information; inputting the input prompt information into the target model so that the target model outputs the annotation result based on the role instruction, the evaluation criteria of at least two evaluation dimensions, the input instruction and the output criteria of the annotation result.
[0060] In some possible implementations of the embodiments of this application, the target question and at least two answers can be filled into the parameter content corresponding to the input instruction in the annotation instruction by a script.
[0061] In some possible implementations of the embodiments of this application, the output standard includes the evaluation process and the evaluation result; the target model outputs the annotation result based on the role instruction, the evaluation standard of at least two evaluation dimensions, the input instruction and the annotation result, which may include: the target model uses the evaluation standard of at least two evaluation dimensions to score each of the at least two answers respectively, and obtains the score of each answer relative to each evaluation dimension; calculates the comprehensive evaluation score of each answer according to the score of each answer relative to each evaluation dimension; fills the comprehensive evaluation score and the calculation process into the evaluation process in the output standard, and determines the evaluation result according to the comprehensive evaluation score, and outputs the annotation result containing the evaluation process and the evaluation result.
[0062] In some possible implementations of this application, calculating the comprehensive evaluation score of each answer based on the score of each answer relative to each evaluation dimension may include: summing the scores of each answer relative to each evaluation dimension to obtain the comprehensive evaluation score of each answer.
[0063] For example, suppose the target question is: How many provinces are there in China? The first answer is: China has 23 provinces. The second answer is: China has a total of 23 provinces, excluding autonomous regions, including Liaoning, Hebei, Jilin, Heilongjiang, etc.
[0064] The annotation results output by the target model can be as follows:
[0065] This is an evaluation assistant used to evaluate the quality of the question-answering model's answers to questions. The answers will be evaluated from five dimensions: authenticity, relevance, usefulness, harmlessness, and conciseness.
[0066] Authenticity Standard: XXXXXXX
[0067] Relevance criteria: XXXXXXX
[0068] Usefulness standard: XXXXXXX
[0069] Harmlessness standard: XXXXXXX
[0070] Simplicity standard: XXXXXXX
[0071] [Target Question]
[0072] How many provinces are there in China?
[0073] [First answer begins]
[0074] China has 23 provinces
[0075] [End of first answer]
[0076] [Start of the second answer]
[0077] China has 23 provinces, excluding autonomous regions, including Liaoning, Hebei, Jilin, and Heilongjiang.
[0078] [End of second answer]
[0079] Evaluation process:
[0080] Based on the evaluation criteria corresponding to the five evaluation dimensions of authenticity, relevance, usefulness, harmlessness, and conciseness, the authenticity score of the first answer is 2 points and the authenticity score of the second answer is 2 points because of XXXXXXX; the relevance score of the first answer is 2 points and the relevance score of the second answer is 1 point because of XXXXXXX; the usefulness score of the first answer is 2 points and the usefulness score of the second answer is 2 points because of XXXXXXX; the harmlessness score of the first answer is 2 points and the harmlessness score of the second answer is 2 points because of XXXXXXX; the conciseness score of the first answer is 2 points and the conciseness score of the second answer is 1 point because of XXXXXXX.
[0081] The first answer received a total score of 9 points, while the second answer received a total score of 7 points.
[0082] Evaluation results: The overall evaluation score of the first answer is greater than that of the second answer.
[0083] Step 103: Construct a preference data sample set based on the annotation results, at least two answers, and the target question.
[0084] In some possible implementations of the embodiments of this application, the annotation results output by the target model can be output in the form of a file, and then the file can be parsed using a script. Based on the parsing results and the corresponding answers and questions, a sample of the preference data sample set in the form of triples can be generated.
[0085] In some possible implementations of the embodiments of this application, step 103 may include: labeling the answer with the highest comprehensive evaluation score among at least two answers in the labeling results as the preferred answer to the target question; labeling the answer with the lowest comprehensive evaluation score among at least two answers as the non-preferred answer to the target question; forming a triplet from the target question, the preferred answer and the non-preferred answer; and using the triplet as a sample of the preference data sample set.
[0086] In some possible implementations of the embodiments of this application, each sample in the preference data sample set can be represented by a triple (query, chosen, rejected), where query represents the question, chosen represents the answer with the highest comprehensive evaluation score among at least two answers obtained for answering the question, and rejected represents the answer with the lowest comprehensive evaluation score among at least two answers.
[0087] For example, taking the above question "How many provinces are there in China? The first answer is 23 provinces in China, and the second answer is 23 provinces in total, excluding autonomous regions, including Liaoning, Hebei, Jilin, Heilongjiang, etc." as an example, a sample in the preference data sample set is represented by a triple ("How many provinces are there in China?", "23 provinces in China", "23 provinces in total, excluding autonomous regions, including Liaoning, Hebei, Jilin, Heilongjiang, etc.").
[0088] In this embodiment, at least two answers are obtained by at least two question-answering models respectively answering the target question; the target question and at least two answers, combined with annotation instructions, are input into the target model to obtain annotation results of the target model annotating the at least two answers respectively; based on the annotation results, the at least two answers, and the target question, a preference data sample set is constructed. Compared with the manual annotation method used in related technologies to construct the preference data sample set, this method can automatically annotate the question-answer preference data samples, saving annotation resources and improving the efficiency of preference data sample set construction.
[0089] In some possible implementations of the embodiments of this application, the target problem in the embodiments of this application can be a problem in the automotive field.
[0090] In some possible implementations of this application's embodiments, the question can be input into a domain intent classification module, which outputs the domain and intent corresponding to the question. The domain intent classification module can be trained based on the T5-PEGASUS model. The T5-PEGASUS model is a Chinese generative pre-trained model, based on the mT5 architecture and initial weights. It first refines the tokenizer by incorporating the characteristics of Chinese, and then mimics PEGASUS to construct a pre-training task, thus training a new version of the T5 model. The T5-PEGASUS model has a total of 275 million parameters, a maximum training length of 512, a batch size of 96, and a learning rate of 10. -4The model was trained for 1 million steps using 6 x 3090 images over approximately 13 days on a finely processed general-purpose corpus of over 30GB. The training accuracy was approximately 47%, and the training loss was approximately 2.97. The model was written, trained, and tested using bert4keras. On the CSL and LCSTS text generation tasks, the T5-PEGASUS model provides state-of-the-art text generation algorithms. The T5-PEGASUS model exhibits excellent few-shot learning capabilities; even with only 10 labeled samples, the T5-PEGASUS model can still fine-tune a domain intent classification model, significantly outperforming other models.
[0091] The T5-PEGASUS model includes over 140 domain categories, such as automobiles, people, travel, geography, food and beverage, and mathematics; and over 20 intent categories, including information acquisition, generation, security, recommendation, comparison, and evaluation. To enhance the diversity and richness of the data, questions can be filtered from the massive dataset to ensure that the questions only cover each specific domain and intent.
[0092] Figure 2 This is a schematic diagram of the architecture for constructing a sample set of preference data provided in an embodiment of this application.
[0093] exist Figure 2 In the process, the question is input into the domain intent classification module to obtain the domain and intent corresponding to the question. The question is then input into AI question-answering assistant A and AI question-answering assistant B respectively. AI question-answering assistant A outputs answer A for the question, and AI question-answering assistant B outputs answer B for the question. Answer A and answer B are scored according to the evaluation dimensions of "authenticity, relevance, usefulness, harmlessness, and conciseness". The comprehensive score of answer A and answer B is then calculated and compared. Based on the comprehensive scores of answer A and answer B, a preference data sample set corresponding to the domain and intent of the question is constructed. Each sample in the preference data sample set is represented by a triple (query, chosen, rejected), where query represents the question, chosen represents the answer with the higher comprehensive score between answer A and answer B, and rejected represents the answer with the lower comprehensive score between answer A and answer B.
[0094] Corresponding to the above method embodiments, this application also provides an apparatus for constructing a preference data sample set. For example... Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of the preference data sample set construction apparatus provided in this application embodiment. The preference data sample set construction apparatus 300 may include:
[0095] The acquisition module 301 is used to acquire at least two answers obtained by at least two question-answering models respectively answering the target question;
[0096] The annotation module 302 is used to input the target question and at least two answers, along with annotation instructions, into the target model to obtain the annotation results of the target model annotating at least two answers respectively;
[0097] Module 303 is used to construct a sample set of preference data based on the annotation results, at least two answers, and the target question.
[0098] In this embodiment, at least two answers are obtained by at least two question-answering models respectively answering the target question; the target question and at least two answers, combined with annotation instructions, are input into the target model to obtain annotation results of the target model annotating the at least two answers respectively; based on the annotation results, the at least two answers, and the target question, a preference data sample set is constructed. Compared with the manual annotation method used in related technologies to construct the preference data sample set, this method can automatically annotate the question-answer preference data samples, saving annotation resources and improving the efficiency of preference data sample set construction.
[0099] In some possible implementations of the embodiments of this application, the annotation instructions include:
[0100] Character instructions, evaluation criteria for at least two evaluation dimensions, and output criteria for input instructions and annotation results;
[0101] Among them, role instructions are used to prompt the target model's role and at least two evaluation dimensions for evaluating the answer; input instructions are used to prompt the target model to identify the question and the answer; and output criteria are used to specify the content and format of the target model's response.
[0102] In some possible implementations of the embodiments of this application, the annotation module 302 is specifically used for:
[0103] Fill the target question and at least two answers into the parameter content corresponding to the input instruction in the annotation instruction to obtain input prompt information; input the input prompt information into the target model so that the target model outputs the annotation result based on the role instruction, the evaluation criteria of at least two evaluation dimensions, the input instruction and the output criteria of the annotation result.
[0104] In some possible implementations of the embodiments of this application, the output criteria include the evaluation process and the evaluation results; the preference data sample set construction apparatus 300 provided in the embodiments of this application further includes:
[0105] The scoring module is used by the target model to score each of the at least two answers using evaluation criteria of at least two evaluation dimensions, so as to obtain the score of each answer relative to each evaluation dimension.
[0106] The calculation module is used to calculate the overall evaluation score for each answer based on the score of each answer relative to each evaluation dimension.
[0107] The fill module is used to fill the evaluation process in the output standard with the comprehensive evaluation score and calculation process, and determine the evaluation result based on the comprehensive evaluation score, and output the labeled result containing the evaluation process and evaluation result.
[0108] In some possible implementations of the embodiments of this application, the construction module 303 is specifically used for:
[0109] The answer with the highest overall score among at least two answers in the annotation results will be marked as the preferred answer to the target question.
[0110] Mark the answer with the lowest overall score among at least two responses as the non-preference answer to the target question;
[0111] Form triples from the target question, preferred answers, and non-preferred answers;
[0112] The triple is used as a sample in the preference data sample set.
[0113] In some possible implementations of the embodiments of this application, at least two evaluation dimensions include:
[0114] At least two of the following: authenticity, relevance, usefulness, harmlessness, and simplicity.
[0115] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.
[0116] The electronic device may include a processor 401 and a memory 402 storing computer program instructions.
[0117] Specifically, the processor 401 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0118] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to an electronic device. In some specific embodiments, memory 402 is a non-volatile solid-state memory.
[0119] In some specific embodiments, the memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the preference data sample set construction method according to this application.
[0120] The processor 401 reads and executes computer program instructions stored in the memory 402 to implement the steps of the preference data sample set construction method provided in the embodiments of this application.
[0121] In some examples, the electronic device may also include a communication interface 403 and a bus 410. For example, Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.
[0122] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0123] Bus 410 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.
[0124] The electronic device can execute the preference data sample set construction method or data annotation prompting method provided in the embodiments of this application, thereby achieving the corresponding technical effects of the preference data sample set construction method provided in the embodiments of this application.
[0125] In addition, in conjunction with the preference data sample set construction method in the above embodiments, this application also provides a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement the steps of the preference data sample set construction method provided in this application. Examples of computer-readable storage media include non-transitory computer-readable media, such as ROM, RAM, magnetic disks, or optical disks.
[0126] This application also provides a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, they implement the steps of the preference data sample set construction method provided in this application and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0127] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for constructing a preference data sample set, characterized in that, The method includes: Obtain at least two answers from at least two question-answering models, each addressing the target question. The target question and the at least two answers, along with the annotation instructions, are input into the target model to obtain the annotation results of the target model for the at least two answers. Based on the annotation results, the at least two answers, and the target question, a preference data sample set is constructed.
2. The method as described in claim 1, characterized in that, The annotation instructions include: Character instructions, evaluation criteria for at least two evaluation dimensions, and output criteria for input instructions and annotation results; The role instructions are used to prompt the target model about its function and at least two evaluation dimensions for assessing the response; the input instructions are used to prompt the target model to identify the question and the response; and the output criteria are used to specify the content and format of the target model's response.
3. The method as described in claim 2, characterized in that, The step of inputting the target question and the at least two answers, combined with annotation instructions, into the target model to obtain the annotation results of the target model for the at least two answers includes: Fill the target question and the at least two answers into the parameter content corresponding to the input instruction in the annotation instruction to obtain input prompt information; The input prompt information is input into the target model so that the target model outputs the annotation result based on the role instruction, the evaluation criteria of the at least two evaluation dimensions, the input instruction, and the output criteria of the annotation result.
4. The method as described in claim 3, characterized in that, The output criteria include: the evaluation process and the evaluation results; The target model outputs the annotation results based on the role instructions, the evaluation criteria of at least two evaluation dimensions, the input instructions, and the output criteria of the annotation results, including: The target model uses the evaluation criteria of the at least two evaluation dimensions to score each of the at least two answers, thereby obtaining a score for each answer relative to each of the evaluation dimensions. Calculate the overall evaluation score for each answer based on the score for each of the evaluation dimensions. The comprehensive evaluation score and calculation process are filled into the evaluation process in the output standard, and the evaluation result is determined based on the comprehensive evaluation score. The output is a labeled result containing the evaluation process and the evaluation result.
5. The method as described in claim 1, characterized in that, The step of constructing a preference data sample set based on the annotation results, the at least two answers, and the target question includes: The answer with the highest overall evaluation score among at least two answers in the annotation results is labeled as the preferred answer to the target question. The answer with the lowest overall evaluation score among the at least two answers is marked as the non-preference answer to the target question; The target question, the preferred answer, and the non-preferred answer are combined into a triplet; The triple is used as a sample in the preference data sample set.
6. The method as described in claim 2, characterized in that, The at least two evaluation dimensions include: At least two of the following: authenticity, relevance, usefulness, harmlessness, and simplicity.
7. A device for constructing a preference data sample set, characterized in that, The device includes: The acquisition module is used to acquire at least two answers obtained by at least two question-answering models respectively answering the target question; The annotation module is used to input the target question and the at least two answers, along with annotation instructions, into the target model to obtain the annotation results of the target model annotating the at least two answers respectively; A construction module is used to construct a preference data sample set based on the annotation results, the at least two answers, and the target question.
8. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; The processor reads and executes the computer program instructions to implement the steps of the preference data sample set construction method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the steps of the preference data sample set construction method as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes computer program instructions that, when executed by a processor, implement the steps of the preference data sample set construction method as described in any one of claims 1-6.
Citation Information
Patent Citations
Answer labeling method and device based on intelligent question and answer scene and related product
CN117217311A
Data labeling method, device and equipment
CN118228047A
Question and answer model training method, question and answer method and question and answer method under subject question and answer scene
CN118312784A