User preference oriented instruction tuning data selection method
Through user preference-oriented instruction tuning data selection method, and using DPO and BiPS strategies to optimize training data, the problem of lack of user preference considerations in existing methods is solved, and the model's performance on target tasks is improved when the data volume is reduced.
Patent Information
- Application Number
- CN202510679407.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing instruction tuning data selection methods lack explicit consideration of user preference signals, resulting in poor performance of the model on target tasks, especially when the amount of training data is limited.
The user preference-oriented instruction tuning data selection method is adopted, and the preference relationship between different responses is constructed through direct preference optimization (DPO) and bidirectional preference synthesis strategy (BiPS), and the training data is optimized to better match the user preferences of the target task.
In the case of significantly reducing the amount of training data, maintain or even surpass the effect of fine-tuning of the full data, improve the generalization ability and accuracy of the model on the target tasks, and significantly improve the matching degree between the data and the target tasks.
Smart Images

Figure CN120197712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and particularly to a method for selecting instruction tuning data oriented to user preference guidance. Background Art
[0002] Instruction Tuning is an important means to improve the practical application ability of large language models (LLMs). Models represented by GPT-3 (an artificial intelligence language model) and GPT-4 perform excellently in natural language understanding and generation. The core of Instruction Tuning is to enable the model to generate high-quality answers that meet expectations according to specific instructions. In this process, the selection of training data plays a crucial role in the final effect of the model. Currently, the training of mainstream language models often relies on large-scale instruction-response datasets. However, existing research has shown that LIMA can also achieve performance comparable to GPT-4 by only using a small number of carefully constructed instruction-response pairs. Such a discovery has attracted people's attention to high-quality data selection strategies. How to improve the model performance by reasonably selecting training samples under limited training scale has become an urgent problem to be solved currently.
[0003] During the instruction tuning process, how to screen out a high-quality subset from a large amount of instruction-response data so that the fine-tuned model can achieve the best performance on specific target tasks is a technical problem that needs to be solved urgently. Most existing instruction tuning data selection methods measure the quality and diversity of samples based on the instruction-response generation process of the model. They usually focus on features such as the difficulty or diversity of the generated responses. Whether it is a general selection that does not consider specific target tasks or a targeted selection that utilizes certain target task information, existing methods lack explicit consideration of user preference signals. This means that the model may miss those training instances that truly meet the requirements of the target task and user preferences, resulting in limitations in the performance of the fine-tuned model in the target scenario. Therefore, a new data selection technical solution is needed to more effectively align the training data with the user preferences of the target task, thereby improving the model performance on the target task while significantly reducing the amount of training data. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention provides a method for selecting instruction tuning data oriented to user preferences, so as to clearly screen out training samples that meet user preferences and improve the fine-tuning effect of the LLM. By Direct Preference Optimization (DPO), the preference relationship between different responses is constructed, so that the gradient information of DPO training can be used as the preference representation of the target task. The Bidirectional Preference Synthesis (BiPS) strategy scores according to the correlation between the training samples and the positive and negative preferences by utilizing the preference features bidirectionally. The present invention can not only automatically screen high-quality data, but also accurately identify data highly relevant to the target task, improve the generalization ability of the instruction fine-tuning model on the target task, and has important research value and application significance. It can be widely applied to tasks such as the training optimization of large language models, intelligent dialogue systems, open-domain question answering, and automatic writing.
[0005] The present invention provides a method for selecting instruction tuning data oriented to user preferences, including: S1: Randomly select a training subset from the training dataset, and perform supervised fine-tuning on the pre-trained large language model on the training subset to obtain a supervised fine-tuned large language model; S2: Construct a preheating preference dataset according to the instruction, input, preference response, and non-preference response; S3: Optimize the supervised fine-tuned large language model through the preheating preference dataset according to the direct preference optimization strategy to obtain a direct preference large language model; S4: Extract a set of verification instructions from the target task to form a verification set, generate basic candidate responses to the verification instructions through a basic candidate model, and generate preference candidate responses to the verification instructions through a preference candidate model; S5: Evaluate the basic candidate response and the preference candidate response according to the evaluation model, and construct a preference pair set according to the evaluation results; S6: Calculate the preference gradient according to the preference pair set by using the preference loss function to obtain a bidirectional user preference gradient; S7: Score the training data according to the bidirectional user preference gradient, select the training dataset according to the score to obtain a high-score sample set, fine-tune the direct preference large language model through the high-score sample set to obtain an optimized large language model, and select instruction tuning data according to the optimized large language model.
[0006] Further, the supervised fine-tuned large language model has basic instruction-following capabilities, and the basic instruction-following capabilities include basic instruction understanding capabilities and basic instruction response capabilities.
[0007] Furthermore, the preheating preference dataset includes multiple triples, and each triple includes an instruction and an input, a preference response, and a non-preference response.
[0008] Furthermore, step S3 includes: S31: Initialize the policy model and the reference model by supervised fine-tuning of the large language model; S32: Optimize the policy model through the preheating preference dataset and the direct preference loss function, and constrain the optimization direction of the policy model through the reference model to obtain the direct preference large language model.
[0009] Furthermore, the evaluation model includes GPT-4.
[0010] Furthermore, the preference pair set includes a positive preference pair set and a negative preference pair set; When the preference candidate response is better than the base candidate response, the verification instruction and the preference candidate response are used as the preference output, with the base candidate response as the control, and recorded in the positive preference pair set; When the preference candidate response is worse than the base candidate response, the instruction and the preference candidate response are used as the preference output, with the base candidate response as the control, and recorded in the negative preference pair set; When the preference candidate response is equal to the base candidate response, it is not included in the preference pair set.
[0011] Furthermore, in step S6, the calculation expression of the bidirectional user preference gradient is: where, is the positive preference gradient vector set, is the Rademacher random variable, is the gradient, is the direct preference loss function, is the preference pair sample, is the policy model parameter, is the reference model parameter, is the positive preference pair set, is the number of positive preference pairs, is the real number vector set, is the dimension of the vector, is the negative preference gradient vector set, is the negative preference pair set, is the number of negative preference pairs, is the bidirectional user preference gradient, is the number of samples in the validation set.
[0012] Furthermore, step S7 includes: S71: Calculate the gradient of the fine-tuning loss function of each sample in the training dataset with respect to the sample to obtain the sample gradient; S72: Calculate the cosine similarity between the sample gradient and the set of positive preference gradients to obtain the positive preference score; S73: Calculate the cosine similarity between the sample gradient and the set of negative preference gradients to obtain the negative preference score; S74: Combine the positive preference score and the negative preference score through linear weighting to obtain the preference score; S75: Select the top high-score samples from the training set according to the preference score to obtain the high-score sample set; S76: Fine-tune the direct preference large language model through the high-score sample set to obtain the optimized large language model, and select the instruction tuning data according to the optimized large language model.
[0013] Furthermore, in step S74, the calculation expression of the preference score is: where, is the preference score, is the weight vector, is the positive preference score, is the negative preference score, is the number of samples in the training dataset, is the sample.
[0014] Furthermore, the weight vector is automatically optimized and searched through the simulated annealing algorithm.
[0015] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: The present invention can maintain or even exceed the effect of full-data fine-tuning while significantly reducing the amount of training data, has higher data efficiency, and the selected data can make the fine-tuned LLM achieve higher accuracy in tasks such as open-domain question answering, dialogue, and professional reasoning. The present invention can not only automatically screen high-quality data, but also accurately identify data highly relevant to the target task, can significantly improve the matching degree of the selected data with the target task, and improve the generalization ability of the instruction tuning model in the target task.
[0016] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of a method for selecting instruction tuning data oriented to user preference guidance provided by the present invention.
[0019] Figure 2 It is a schematic framework diagram of a method for selecting instruction tuning data oriented to user preference guidance provided by the present invention.
[0020] Figure 3 It is a schematic diagram of constructing the preference of the validation set of the present invention.
[0021] Figure 4 It is a comparison diagram of the effects of the present invention on different data sets under different training data scales.
[0022] Figure 5 It is a comparison diagram of the effects of the present invention and the task-independent data selection method under different training data scales. Detailed implementation manners
[0023] To make the purpose, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0024] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0025] The following will be combined with Figures 1 to 5Describe a method for selecting instruction tuning data oriented to user preferences in the present invention.
[0026] As Figure 1 shown, a method for selecting instruction tuning data oriented to user preferences includes: S1: Randomly select a training subset from the training dataset and perform supervised fine-tuning on the pre-trained large language model on the training subset to obtain a supervised fine-tuned large language model; The model framework is as Figure 2 shown. In some specific embodiments of the present invention, the pre-trained large language model is , and the parameters are denoted as . Randomly select a training subset from the training dataset. Through the training subset , perform fine-tuning on to obtain a supervised fine-tuned large language model . The parameters of are . The supervised fine-tuned large language model has basic instruction following capabilities, and the basic instruction following capabilities include basic instruction understanding capabilities and basic instruction response capabilities.
[0027] Use cross-entropy loss to calculate the loss function. The calculation expression of the loss function is: where is the fine-tuning loss, is the response generated by the th sample, is the true reference answer of the th sample, is the number of samples in the supervised fine-tuning stage, is the cross-entropy loss function.
[0028] S2: Construct a warm-up preference dataset according to the instruction, input, preference response, and non-preference response; The warm-up preference dataset includes multiple triples, and the triples include the instruction and input, preference response and non-preference response .
[0029] The calculation expression of the warm-up preference dataset is: where is the warm-up preference dataset, is the instruction and input of the th sample, is the preference response of the th sample, is the non-preferred response for the th sample, and both come from the training dataset, is the generated response, and is the number of samples in the policy optimization phase.
[0030] S3: Optimize the supervised fine-tuning of the large language model through the warm-up preference dataset according to the direct preference optimization strategy to obtain the direct preference large language model; S31: Initialize the policy model and the reference model through the supervised fine-tuning of the large language model; S32: Optimize the policy model through the warm-up preference dataset and the direct preference loss function, and constrain the optimization direction of the policy model through the reference model to obtain the direct preference large language model; Among them, is the policy optimization loss, is the Sigmoid function, is the temperature parameter used to scale the implicit reward difference, is the policy model, is the reference model, and is the number of samples in the policy optimization phase.
[0031] In some specific embodiments of the present invention, during the pre-training process of the direct preference large language model (DPO), since the reference model is fixed, the parameters of the policy model will ultimately be the same.
[0032] The purpose of optimization is to increase the prediction probability of the model for the preferred response while reducing the bias towards the non-preferred response, thereby guiding the model to evolve in a direction that better fits the user's preferences.
[0033] The direct preference large language model not only has the ability to follow instructions but also can align with user preferences.
[0034] The present invention uses gradients to estimate the influence of training samples and user preferences, thereby associating the training data with the target task. For a triple in the training dataset, the training gradient can be expressed as: Among them, is the gradient, is the training sample, is the dimension of the gradient vector, is the set of real number vectors, is the fine-tuning loss; The calculation expression of the influence of the training sample is: where, is the influence of the training sample, is a Rademacher random variable, , is the training data set, is the number of samples in the training data set, is the dimension of the vector.
[0035] In some specific embodiments of the present invention, .
[0036] S4: Extract a set of verification instructions from the target task to form a verification set, generate a basic candidate response for the verification instructions through the basic candidate model, and generate a preferred candidate response for the verification instructions through the preferred candidate model; Extract a set of verification instructions from the target task to form a verification set , , where, is the instruction and input of the th sample, is the number of verification set samples.
[0037] The basic candidate model is obtained by performing supervised (SFT) fine-tuning on a small amount of random data , and the preferred candidate model is obtained by performing supervised (SFT) fine-tuning on more high-quality data .
[0038] Use the basic candidate model and the preferred candidate model to generate responses of different qualities for the same instruction, and respectively obtain the basic candidate response and the preferred candidate response .
[0039] S5: Evaluate the basic candidate response and the preferred candidate response according to the evaluation model, and construct a set of preference pairs according to the evaluation results; Based on the intuition that "high-quality data should be closer to the correct preference and farther from the incorrect preference", the present invention divides the user preference into two directions, Positive preference: The response changes from unsatisfactory to ideal; Negative preference: The response degenerates from ideal to unsatisfactory; The set of preference pairs includes a set of positive preference pairs and a set of negative preference pairs; When the preferred candidate response is better than the base candidate response, the verification instruction and the preferred candidate response are used as the preferred output, and the base candidate response is used as the control and recorded in the positive preference pair set; When the preferred candidate response is worse than the base candidate response, the instruction and the preferred candidate response are used as the preferred output, and the base candidate response is used as the control and recorded in the negative preference pair set; When the preferred candidate response is equal to the base candidate response, it is not included in the preference pair set.
[0040] The calculation formula is: Among them, is the positive preference pair set, is the preferred candidate response at time is the base candidate response at time is the number of positive preference pair sets, is the number of negative preference pair sets, is the score of the evaluation model for ; is the score of the evaluation model for ; is the negative preference pair set, .
[0041] The construction process of the preference pair is as Figure 3 shown.
[0042] S6: According to the preference pair set, use the preference loss function to calculate the preference gradient and obtain the bidirectional user preference gradient; The calculation formula of the bidirectional user preference gradient is: Among them, is the positive preference gradient vector set, is the Rademacher random variable, is the gradient, is the direct preference loss function, is the preference pair sample, is the policy model parameter, is the reference model parameter, is the positive preference pair set, is the number of positive preference pair sets, is the real number vector set, is the dimension of the vector, is a set of negative preference gradient vectors, is a set of negative preference pairs, is the number of the set of negative preference pairs, is a two-way user preference gradient, is the number of samples in the validation set.
[0043] The positive preference gradient vector indicates the direction in which the model evolves towards the ideal response, and the negative preference gradient vector indicates the direction in which the model deviates from the ideal response.
[0044] S7: Score the training data according to the two-way user preference gradient, select the training data set according to the score, fine-tune the direct preference large language model through the training data set to obtain an optimized large language model, and select instruction tuning data according to the optimized large language model.
[0045] S71: Calculate the gradient of the fine-tuning loss function of each sample in the training data set with respect to the sample to obtain the sample gradient; S72: Calculate the cosine similarity between the sample gradient and the set of positive preference gradients to obtain the positive preference score, where, is the correlation between the th sample gradient and the th positive preference gradient, is the normalization factor of the positive preference gradient cosine similarity, is the th sample gradient, is the th positive preference gradient, is the transpose of the matrix.
[0046] Perform a weighted sum of the correlations between each sample in the training data set and the positive preference, where the weights are determined by the norm of the positive preference direction, to obtain the positive preference score .
[0047] S73: Calculate the cosine similarity between the sample gradient and the set of negative preference gradients to obtain the negative preference score, where, is the correlation between the th sample gradient and the th negative preference gradient, is the normalization factor of the negative preference gradient cosine similarity, is the th negative preference gradient, Perform a weighted sum of the correlation between each sample in the training dataset and the negative preference, where the weight is determined by the norm of the negative preference direction, to obtain the negative preference score 。
[0048] S74: Combine the positive preference score and the negative preference score through linear weighting to obtain the preference score; The scoring calculation expression for the training sample is: where, is the score of the training sample, is the weight vector, is the positive preference score, is the negative preference score.
[0049] Adopt the Simulated Annealing Algorithm to automatically optimize , and the calculation expression for the preference score is: where, is the preference score.
[0050] The optimization objective of the simulated annealing algorithm is to maximize the similarity between the training sample and the positive preference, while minimizing its similarity with the negative preference, so as to ensure that the finally selected data better meets the preference requirements of the target task.
[0051] S75: Select the top-ranked high-score samples from the training set according to the preference score to obtain the high-score sample set , S76: Fine-tune the direct preference large language model through the high-score sample set to obtain the optimized large language model, and select the instruction tuning data according to the optimized large language model.
[0052] The instruction tuning data includes intelligent customer service data, voice assistant data, legal / medical texts, etc.
[0053] Compared with the existing non-target task method IFD, the present invention (ProDS) can significantly improve the matching degree between the selected data and the target task. For example, the distribution deviation of the top 5% instructions selected by IFD from the actual target test set is relatively large, and the average Euclidean distance after t-SNE dimensionality reduction is about 85.85. While through the preference signal optimization screening of the target task in the present invention, the selected data set is closer to the target test set, and the average Euclidean distance after t-SNE dimensionality reduction is reduced to 67.13. The present invention is significantly superior to IFD.
[0054] The present invention has higher data efficiency and can maintain or even exceed the effect of full-data fine-tuning while significantly reducing the amount of training data. For example, using a 5%-10% training data subset selected by IFD, the performance of the fine-tuned model exceeds that of the model fine-tuned using 100% training data. This not only reduces the computational cost and storage requirements but also improves the training efficiency, making the present invention applicable to scenarios with limited computing resources.
[0055] The present invention has achieved significant performance improvements on multiple task benchmarks. Compared with traditional data selection methods such as BM25 (based on classical information retrieval algorithms), DSIR (Data Selection via Influence Ranking, using gradient influence ranking to select samples with the greatest impact on the target task), and RDS (using preference gradient direction for sample selection), the data selected by the present invention can enable the fine-tuned LLM to achieve higher accuracy in tasks such as open-domain question answering, dialogue, and professional reasoning. For example, on multiple complex task benchmarks, the present invention can exceed the effect of any existing data selection method of the same scale using only 5% of the data.
[0056] The present invention can not only automatically screen high-quality data but also accurately identify data highly relevant to the target task, improving the generalization ability of the instruction fine-tuning model on the target task.
[0057] The method of the present invention was verified using two types of instruction datasets, target-agnostic and target-relevant. In the setting of target-relevant data selection, FLAN V2 (Instruction Fine-Tuning Language Network Dataset V2), COT (Chain of Thought Dataset), DOLLY (Databricks Open Large Language Model Library), and OpenAssistant (Open Assistant Dialogue Dataset) were used as the training set, totaling approximately 270,000 diverse instruction samples, covering various data formats and complex reasoning tasks. The test sets selected three commonly used benchmark tasks: MMLU (Massive Multitask Language Understanding Evaluation Set), TYDIQA (Type-Diverse Question Answering Evaluation Set), and BBH (BIG-Bench Hard Subset Evaluation Set). The target-relevant training data does not contain samples directly related to these test tasks to ensure the reliability of the evaluation results. In the SFT warm-up stage, we used 5% of the data in the training set for fine-tuning. The preference training in the DPO warm-up stage was also constructed from 5% of the samples in the full training set. To evaluate the similarity between the training samples and the target task, we constructed validation sets from each test set. For the dataset oriented to the target task, we adopted the same validation set construction method as the LESS method.
[0058] Tables 1 and 2 show the performance comparison of different methods on three benchmark datasets under target-related settings. Using a larger-scale model (Llama2-7B) for data selection and fine-tuning can achieve better results than a smaller model (Llama32-1B). However, in the TYDIQA task, due to the rich context information provided by this dataset, the generation ability of the small model is sufficient, so Llama32-1B achieved a higher score. For MMLU, the multiple-choice format results in relatively small differences in user preferences between different responses, so the improvement of ProDS compared to the full data is limited. In contrast, BBH is an open-ended question-answering task with chain-of-thought reasoning, and the preference differences between different answers are more semantically profound, so ProDS achieved a more significant performance improvement on BBH. It is worth noting that ProDS can achieve even better or exceed the performance of the model fine-tuned with 100% of the full data by only using 5% of the subset of the entire training data, demonstrating high data utilization efficiency. In Table 2, 1B represents Llama32-1B and 7B represents Llama2-7B, represents the improvement value of the best data of other methods to the best 7B of the present invention, 1B 7B means first using the Llama32-1B model with a smaller number of parameters for sample selection (data scoring), and then using the selected data to train a larger Llama2-7B model.
[0059] Table 1 Table 2 The Alpaca dataset (containing 52,002 instruction-response samples) was used as the training set. The test sets selected five datasets, Vicuna, Koala, WizardLM, Self-Instruct, and LIMA, each containing approximately 1K artificially created open-domain or closed-domain instructions, to comprehensively evaluate the performance of the model in different scenarios. For the Alpaca-related test sets, if there is no sub-task, 10% of the data was randomly selected as the validation set; if there are sub-tasks (such as Vicuna and WizardLM), 2 instructions were randomly sampled from each sub-task as the validation set.
[0060] In a target-agnostic scenario, the subset of the Alpaca training set screened by ProDS was compared with the data trained using the full Alpaca training set. On five open-domain test sets (Vicuna, Koala, WizardLM, Self-Instruct, Lima), the comparison results evaluated by GPT-4 showed that the model fine-tuned with only 10% of the Alpaca data had comparable performance to the model fine-tuned with 100% of the data. The number of wins, ties, and losses is as Figure 4 shown, and in most cases, the two are tied or the model fine-tuned with 10% of the Alpaca data wins.
[0061] As Figure 5 shown, the effect of ProDS under different training data scales. As the proportion of the data subset increases from 5% to 20%, the average winning rate of the model on the five test sets always remains above 1. Therefore, the model fine-tuned with the data selected by ProDS stably outperforms the model fine-tuned with the full data. In addition, by comparing ProDS with the existing task-agnostic data selection method IFD, it can be found that at an extremely small data subset (5%), IFD slightly exceeds ProDS. However, as the training data scale expands, the performance of ProDS continues to improve and surpasses IFD. Therefore, although IFD can select high-quality data according to the instruction difficulty, it has deficiencies in identifying the target-related preferences of medium and low-quality samples. On the contrary, by adding preference learning on the basis of SFT, ProDS can more effectively identify the training samples beneficial to the target task, and thus shows stronger advantages at a larger data scale.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for selecting instruction tuning data oriented to user preference guidance, characterized in that, Including: S1: Obtain a supervised fine-tuned large language model by randomly selecting a training subset from the training dataset and performing supervised fine-tuning on the pre-trained large language model on the training subset; S2: Construct a warm-up preference dataset according to instructions, inputs, preferred responses, and non-preferred responses; S3: Optimize the supervised fine-tuned large language model through the warm-up preference dataset according to the direct preference optimization strategy to obtain a direct preference large language model; S4: Extract a set of validation instructions from the target task to form a validation set, generate basic candidate responses for the validation instructions through the basic candidate model, and generate preferred candidate responses for the validation instructions through the preference candidate model; S5: Evaluate the basic candidate responses and preferred candidate responses according to the evaluation model, and construct a preference pair set according to the evaluation results; S6: Calculate the preference gradient using the preference loss function according to the preference pair set to obtain the bidirectional user preference gradient; S7: Score the training data according to the bidirectional user preference gradient, select the training dataset according to the score to obtain a high-score sample set, fine-tune the direct preference large language model through the high-score sample set to obtain an optimized large language model, and select instruction tuning data according to the optimized large language model.
2. The method for selecting instruction tuning data oriented to user preference guidance according to claim 1, wherein The supervised fine-tuned large language model has basic instruction-following capabilities, and the basic instruction-following capabilities include basic instruction understanding capabilities and basic instruction response capabilities.
3. A method for selecting instruction tuning data oriented to user preference guidance according to claim 1, characterized in that, The warm-up preference dataset includes multiple triples, and the triples include instructions and inputs, preferred responses, and non-preferred responses.
4. A method for selecting instruction tuning data oriented to user preference guidance according to claim 1, characterized in that Step S3 includes: S31: Initialize the policy model and the reference model through the supervised fine-tuned large language model; S32: Optimize the policy model through the warm-up preference dataset and the direct preference loss function, and constrain the optimization direction of the policy model through the reference model to obtain a direct preference large language model.
5. A method for selecting instruction tuning data oriented to user preference guidance according to claim 1, characterized in that, The evaluation model includes GPT-4.
6. The method for selecting instruction tuning data oriented to user preference guidance according to claim 1, wherein The preference pair set includes a positive preference pair set and a negative preference pair set; When the preferred candidate response is better than the basic candidate response, use the validation instruction and the preferred candidate response as the preference output, and the basic candidate response as the control, and record it in the positive preference pair set; When the preferred candidate response is worse than the basic candidate response, use the instruction and the preferred candidate response as the preference output, and the basic candidate response as the control, and record it in the negative preference pair set; When the preferred candidate response is equal to the basic candidate response, it is not included in the preference pair set.
7. A method for selecting instruction tuning data oriented to user preference guidance according to claim 1, characterized in that In step S6, the calculation expression of the bidirectional user preference gradient is: Among them, is a set of positive preference gradient vectors, is a Rademacher random variable, is a gradient, is a direct preference loss function, is a preference pair sample, is a policy model parameter, is a reference model parameter, is a set of positive preference pairs, is the number of positive preference pair sets, is a set of real vectors, is the dimension of the vector, is a set of negative preference gradient vectors, is a set of negative preference pairs, is the number of negative preference pair sets, is a two-way user preference gradient, is the number of validation set samples.
8. A method for selecting instruction tuning data oriented to user preference guidance according to claim 1, characterized in that Step S7 includes: S71: Calculate the gradient of the fine-tuning loss function of each sample in the training dataset with respect to the sample to obtain the sample gradient; S72: Calculate the cosine similarity between the sample gradient and the positive preference gradient set to obtain the positive preference score; S73: Calculate the cosine similarity between the sample gradient and the negative preference gradient set to obtain the negative preference score; S74: Combine the positive preference score and the negative preference score through linear weighting to obtain the preference score; S75: Select the top-ranked high-score samples from the training set according to the preference score to obtain a high-score sample set; S76: Fine-tune the direct preference large language model with the high-score sample set to obtain an optimized large language model, and select instruction tuning data according to the optimized large language model.
9. A method for selecting instruction tuning data oriented to user preference guidance according to claim 8, characterized in that, In step S74, the calculation expression of the preference score is: wherein, is the preference score, is the weight vector, is the positive preference score, is the negative preference score, is the number of samples in the training data set, is the sample.
10. A method for selecting instruction tuning data oriented to user preference guidance according to claim 9, characterized in that, Automatically optimize and search the weight vector through the simulated annealing algorithm.
Citation Information
Patent Citations
Reinforcement learning alignment model training method and system based on AI feedback
CN118735002A
Chinese grammar error correction method and system based on large model fine tuning
CN119849482A
Cited By
Screening method and equipment of instruction data, medium and product
CN120653995A
Data screening method, device and equipment and computer storage medium
CN121524635A
Data screening method, device and equipment and computer storage medium
CN121524635B