Preference data set making method, equipment and medium
By designing prompts and extracting target problems, generating preference datasets and updating models, the problem of large language models avoiding answers in edge cases is solved, improving their processing power and user satisfaction.
Patent Information
- Application Number
- CN202510094606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
Large language models will experience cognitive errors and avoid answers when facing problems with future predictions, sensitive topics or edge cases, and the dataset construction cost is high, limiting their application in AI agents.
By designing the first and second prompts, the language model's reply is generated, the target question is extracted and inputted to the second language model with high resource occupancy, the preference data set is generated, and the model is updated through audit and training.
It improves the ability of large language models to handle edge problems, reduces the cost of data set production, broadens the scope of answers, and significantly improves user satisfaction and model performance.
Smart Images

Figure CN120011513A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence and machine learning, and in particular to a method, device and medium for producing a preference data set. Background Art
[0002] With the development of science and technology, various large models have been produced for various aspects of reality. Among them, large language models (large models) have shown extraordinary potential in dealing with complex problems and providing high-quality text generation. The current large language models have certain basic cognitive capabilities after completing pre-training.
[0003] However, without involving external knowledge bases, these models often have cognitive errors when faced with questions involving future predictions, sensitive topics or other edge cases, and choose not to answer due to preset security strategies or uncertainty avoidance mechanisms. This limitation is particularly evident in application scenarios such as AI agents. Although human feedback reinforcement learning is widely regarded as an effective way to optimize the responses of such models, it relies on high-cost dataset construction, which has become a major obstacle to large-scale implementation. Summary of the invention
[0004] The embodiments of the present application provide a method, device and medium for preparing a preference dataset, which are used to solve the problem that a language model avoids answering certain questions due to limitations and the high cost of dataset construction.
[0005] The present application embodiment adopts the following technical solutions:
[0006] On the one hand, an embodiment of the present application provides a method for producing a preference data set, the method comprising: generating a first answer language for a first question type according to a preset first prompt language of a first question type; the first question type is a question that cannot be answered by the first language model; extracting target questions that meet the second question type from all questions in the first question type; the second question type is a question that can be answered in theory; according to the second prompt language of each target question, matching the answer language for each target question to obtain a designated target question with a second answer language; inputting each non-designated target question in the target question into the second language model respectively to obtain a third answer language for each non-designated target question; the second language model occupies more resources than the first language model; generating a preference data set according to the designated target question, the first answer language and the second answer language of the designated target question, the non-designated target question, and the first answer language and the third answer language of the non-designated target question.
[0007] In one example, based on a preset first prompt for a first question type, a first reply for the first question type is generated by a first language model, specifically including: matching prompts for the first question type in a preset prompt library to determine the first prompt; matching replies for the first question type in a preset reply library to generate a first reply for the first question type using the first language model.
[0008] In one example, based on the second prompt of each target question, the answer language is matched for each target question to obtain a designated target question with a second answer language, specifically including: in a preset prompt library, according to the content of each target question, the second prompt of each target question is determined; in a preset answer language library, the answer language is matched for each target question to obtain a designated target question with a second answer language.
[0009] In one example, after generating a preference data set based on a specified target question, a first answer and a second answer to the target question, a non-specified target question, and a first answer and a third answer to the non-specified target question, the method further includes: reviewing the preference data set, and training the first language model based on the reviewed preference data set to update the first language model.
[0010] In one example, a preference data set is reviewed, a first language model is trained based on the reviewed preference data set, and the first language model is updated, specifically including: inputting the preference data set into a pre-built review model to obtain a first review data set; sending the first review data set to a client to obtain a second review data set after manual review; and training the first language model based on the second review data set to update the first language model.
[0011] In one example, each non-specified target question in the target question is input into the second language model respectively, and after obtaining the third answer language for each non-specified target question, the method further includes: determining a non-specified target question table in which the third answer language is empty, and sending the non-specified target question table to the client.
[0012] In one example, after determining that the third reply is an empty non-specified target question table and sending the non-specified target question table to the client, the method further includes: sending the non-specified target question table manually written on the client to the first language model, training the first language model according to the non-specified target question table to update the first language model.
[0013] In one example, before generating a first answer to the first question type from a first language model based on a preset first prompt for the first question type, the method further includes: collecting an operation log of the first language model and extracting the operation log to obtain the first question type.
[0014] On the other hand, an embodiment of the present application provides a preference data set production device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above-mentioned preference data set production methods.
[0015] On the other hand, an embodiment of the present application provides a non-volatile computer storage medium for producing a preference data set, which stores computer executable instructions, and the computer executable instructions can execute any of the above-mentioned methods for producing a preference data set.
[0016] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:
[0017] This application obtains the first answer by designing the first prompt for the question, ensuring a good user experience during the operation of the online service even in abnormal situations, and helps to collect abnormal data. By innovatively combining large model prompt technology and secondary feedback mechanism, the ability of large models to handle marginal problems is improved, the cost of data set production is reduced, and the scope of answers is broadened, significantly improving user satisfaction and model performance. In addition, the specified target questions, non-specified target questions, and their first answers, second answers, or third answers are integrated into a preference data set, which can be used to further train the model and provide the model with rich and multi-dimensional learning materials, which is conducive to the optimization and upgrading of the model's long-term performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solution of the present application, some embodiments of the present application will be described in detail below in conjunction with the accompanying drawings, in which:
[0019] Figure 1 A flowchart of a method for preparing a preference data set provided in an embodiment of the present application;
[0020] Figure 2 A schematic diagram of a device flow of a method for preparing a preference data set provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of the structure of a preference data set production device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0023] Some embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0024] Figure 1 A flowchart of a method for preparing a preference data set provided in an embodiment of the present application. The method can be applied to different business fields. Certain input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.
[0025] The analysis method involved in the embodiments of the present application can be implemented by a terminal device or a server, and the present application does not impose any special restrictions on this. For the convenience of understanding and description, the following embodiments are described in detail by taking a server as an example.
[0026] It should be noted that in the process of localized deployment, in order to balance GPU resources, a large model with limited parameters is generally selected. This type of large model with small parameters (first language model) runs faster and takes up less resources, but the disadvantage is that the ability is limited and it is often unable to align with user questions. Therefore, we will choose the knowledge of the large model (second language model) to generate a more suitable data set to train and align the large model with small parameters. For example, a large language model with 72B requires at least two A800s with 80G video memory, which is difficult to meet in most localized deployment scenarios, and GPUs are very expensive.
[0027] Based on this, Figure 1 The process in may include the following steps:
[0028] S101: Generate a first answer language of a first language model for a first question type according to a preset first prompt language of the first question type; the first question type is a question that the first language model cannot answer.
[0029] It should be noted that in some embodiments of the present application, before designing a prompt based on the first question type, the first question type must first be obtained; the first language model is trained using the acquired model training data, and then the operation log of the first language model is collected, and the format of unanswerable questions is extracted from the operation log to determine it as the first question type.
[0030] It should also be noted that the prompt is used to guide the language model to generate specific outputs or perform specific tasks. It is like an instruction or question to the model. The model will generate a response based on the information and semantics contained in the prompt, using the knowledge and language patterns learned during the training process.
[0031] In some embodiments of the present application, the prompt is set by adding specific instruction words to design specific prompts. First, a prompt library containing many specific instruction words is constructed, and specific instruction words are added thereto according to task requirements. The prompt library contains prompts, the relationship between prompts and questions, and the relationship between prompts and answers. Then, an answer library corresponding to the prompt library is constructed;
[0032] In the prompt library, the first prompt is determined according to the first question type, and the answer is matched for the first question type to generate the first answer for the first question type by the first language model; for example: for questions that cause the model to give an "unable to answer" response or silence, prompts will be designed so that the model can return a more user-friendly message, such as "Please wait, we cannot give you an answer for the time being."
[0033] This step ensures that a good user experience is maintained during the operation of online services even when encountering abnormal situations, and helps collect abnormal data.
[0034] S102: Extracting target questions that meet the second question type from all questions in the first question type; the second question type is a question that can be answered theoretically.
[0035] It should be noted that although the first type of questions are questions that the model "cannot answer", there are many questions that can be answered by theory, such as those about events that will happen in the future. Although there is no relevant training in the training set, prompts can be used to make inferences.
[0036] Furthermore, the system automatically collects the operation logs of all questions of the first question type run by the first language model, and records all questions that fail to receive appropriate answers in detail for subsequent analysis and processing.
[0037] All questions that are legal and compliant and can be answered by theory but cannot be correctly answered by the model are screened out from the logs and classified to obtain various target questions of the second question type. By eliminating illegal and irregular questions, the purity and legality of the preferred data set are increased.
[0038] S103: According to the second prompt language of each target question, answer language matching is performed on each target question to obtain a designated target question with a second answer language.
[0039] It should be noted that in some embodiments of the present application, in a preset prompt library, keywords are extracted from the question content according to the question content of each target question, and the second prompt for each target question is determined based on the keywords. The second prompt adopts an innovative prompt retry strategy, specifically a diversified prompt strategy is designed to stimulate the large model to generate different answer possibilities for unanswered questions. According to the prompt strategy, the answer language is matched for each target question to obtain the second answer language for each target question. The second answer language is the correct answer corresponding to each target question.
[0040] Through the secondary feedback mechanism, diversified prompt strategies are designed to stimulate different answer possibilities, improve the success rate of answer acquisition, and ensure the accuracy and applicability of the data.
[0041] S104: Input each non-specified target question in the target question into a second language model to obtain a third answer to each non-specified target question; the second language model occupies more resources than the first language model.
[0042] It should be noted that the second language model is a model with larger parameters, looser strategies, and higher resource usage than the first language model.
[0043] In some embodiments of the present application, in the process of matching the reply language for each target question according to the second prompt language of each target question, there may be target questions that cannot be matched with corresponding reply languages. The target questions that cannot be matched with reply languages are defined as non-specified target questions. The large model strategy is used to input the non-specified target questions into the second language model, and the third reply language for each non-specified target question is obtained according to the third prompt language set in the second language model.
[0044] Furthermore, among all non-specified target questions, non-specified target questions with empty third answer words are screened out, and summarized to obtain a list of non-specified target questions with empty third answer words, and the list of non-specified target questions with empty third answer words is sent to the client for manual answering, and then the list of non-specified target questions with empty third answer words that have completed manual answering on the client is sent to the first language model, and the first language model is trained and fine-tuned according to the list of non-specified target questions.
[0045] By combining large model prompt technology, the answer range of the first language model is broadened, while the cost-effectiveness ratio is optimized, significantly improving user satisfaction and model performance.
[0046] S105: Generate a preference data set according to the designated target question, the first answer language and the second answer language of the designated target question, the non-designated target question, the first answer language and the third answer language of the non-designated target question.
[0047] It should be noted that after the preference data set is generated, it is necessary to review the preference data set. First, the preference data set is input into a pre-built review model to obtain a first review data set; the first review data set is a machine review process based on keyword review; then the first review data set is sent to the client to obtain a second review data set after manual review; the second review data set is a data set obtained by manual review based on the first review data set; the second review data set is input into the first language model, and the first language model is trained and fine-tuned according to the second review data set.
[0048] By implementing a small proportion of manual review before fine-tuning the preference dataset, the quality and applicability of the dataset is ensured. Although this step increases the labor cost slightly, it greatly improves the content of the dataset and the effectiveness of model training. Based on the verified preference dataset, the large model is fine-tuned in a targeted manner, with the goal of adjusting the decision boundary of the model so that it can handle previously rejected questions more flexibly while maintaining security and robustness, and optimize its handling strategy and answering ability on marginal questions.
[0049] It should be noted that although the embodiments of the present application are based on Figure 1 Steps S101 to S105 are described in sequence, but this does not mean that steps S101 to S105 must be performed in a strict order. Figure 1 The order shown in the figure introduces and explains step S101 to step S105 in sequence to facilitate those skilled in the art to understand the technical solution of the embodiment of the present application. In other words, in the embodiment of the present application, the order between step S101 to step S105 can be appropriately adjusted according to actual needs.
[0050] pass Figure 1 The method ensures that a good user experience can be maintained even when anomalies occur during the operation of online services, and helps collect abnormal data. By innovatively combining large model prompt technology and secondary feedback mechanism, the large model's ability to handle edge problems is improved, the cost of data set production is reduced, and the scope of answers is broadened, significantly improving user satisfaction and model performance. In addition, the specified target questions, non-specified target questions, and their first answers, second answers, or third answers are integrated into a preference dataset, which can be used to further train the model, providing the model with rich and multi-dimensional learning materials, which is conducive to the optimization and upgrading of the model's long-term performance.
[0051] Figure 2 A schematic diagram of a device flow of a method for preparing a preference dataset provided in an embodiment of the present application.
[0052] exist Figure 2 The upper part is a question collection process of a method for making a preference dataset, including: judging from the large model (first language model) whether it is a format that cannot be answered, and if so, collecting a log dataset; the lower part is a process for making a preference dataset, including: inputting the questions in the log dataset (non-specified target questions in the target questions) into a model with larger parameters (second language model) to obtain a preference dataset.
[0053] Figure 3 A schematic diagram of a preference data set preparation device provided in an embodiment of the present application includes:
[0054] at least one processor; and,
[0055] a memory communicatively connected to at least one processor; wherein,
[0056] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can execute any one of the above-mentioned preference data set preparation methods.
[0057] Some embodiments of the present application provide a non-volatile computer storage medium for producing a preference data set, which stores computer executable instructions. The computer executable instructions can execute any one of the above-mentioned methods for producing a preference data set.
[0058] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0059] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0060] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0061] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0062] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0064] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0065] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM), and non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0066] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0067] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0068] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the technical principle of the present application should fall within the protection scope of the present application.
Claims
1. A method for preparing a preference data set, characterized in that: The method comprises: Generate a first answer language of a first language model for a first question type according to a preset first prompt language of a first question type; the first question type is a question that cannot be answered by the first language model; Extracting target questions that meet the second question type from all questions in the first question type; the second question type is a question that can be answered theoretically; According to the second prompt language of each target question, matching the answer language of each target question to obtain the designated target question with the second answer language; Inputting each non-specified target question in the target question into a second language model respectively to obtain a third answer language for each non-specified target question; the second language model occupies more resources than the first language model; A preference data set is generated according to the designated target question, the first answer language and the second answer language of the designated target question, the non-designated target question, the first answer language and the third answer language of the non-designated target question.
2. The method according to claim 1, characterized in that The step of generating a first answer language of the first language model for the first question type according to a preset first prompt language of the first question type specifically includes: Matching prompts for the first question type in a preset prompt library to determine the first prompt; In a preset answer language library, answer language matching is performed on the first question type to generate the first answer language of the first language model for the first question type.
3. The method according to claim 1, characterized in that According to the second prompt of each target question, matching the answer of each target question to obtain the designated target question with the second answer specifically includes: Determine, in a preset prompt library, a second prompt for each target question according to the content of each target question; In the preset answer language library, answer language matching is performed on each target question to obtain a designated target question with a second answer language.
4. The method according to claim 1, characterized in that: After generating the preference data set according to the designated target question, the first answer language and the second answer language of the designated target question, the non-designated target question, the first answer language and the third answer language of the non-designated target question, the method further includes: The preference data set is reviewed, and the first language model is trained according to the reviewed preference data set to update the first language model.
5. The method according to claim 4, characterized in that The reviewing of the preference data set, training the first language model according to the reviewed preference data set, and updating the first language model specifically includes: Inputting the preference data set into a pre-built audit model to obtain a first audit data set; Sending the first audit data set to a client to obtain a second audit data set after manual review; The first language model is trained according to the second audit data set to update the first language model.
6. The method according to claim 1, characterized in that After inputting each non-specified target question in the target question into the second language model to obtain a third answer language for each non-specified target question, the method further includes: Determine a non-specified target question list in which the third reply is empty, and send the non-specified target question list to the client.
7. The method according to claim 6, characterized in that After determining that the third reply is an empty non-specified target question list, and sending the non-specified target question list to the client, the method further includes: The non-specified target question table manually written on the client is sent to the first language model, and the first language model is trained according to the non-specified target question table to update the first language model.
8. The method according to claim 1, characterized in that Before generating a first answer language of a first language model for the first question type according to a preset first prompt language of the first question type, the method further includes: An operation log of the first language model is collected, and the operation log is extracted to obtain the first question type.
9. A preference data set production device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for producing a preference data set as described in any one of claims 1 to 8.
10. A non-volatile computer storage medium for preparing a preference data set, storing computer executable instructions, characterized in that: The computer executable instructions can execute a method for preparing a preference data set as described in any one of claims 1 to 8.