Sample determination method and device for performing multi-target alignment training on large model
By conducting multi-objective alignment training on the big model, using reward model scores and data set extension screening, the problem of multi-objective alignment conflict in big model training is solved, the balanced optimization between business goals is achieved, and the overall performance of the big model is improved.
Patent Information
- Application Number
- CN202510337642.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to achieve effective alignment between multiple business objectives in large model training, resulting in potentially degrading performance of other objectives when optimizing one objective.
By performing multi-objective alignment training on the large model, using the reward model to score the extended response and positive and negative examples, select candidate data pairs that meet the reward consistency, and construct target samples based on the reward differences, and combine prompt information to expand and filter the training data set.
While improving the performance of one business target, avoid lowering the performance of other business targets, significantly reducing conflicts between different business targets, and promoting overall optimization of large-scale model performance.
Smart Images

Figure CN120258138A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a method and apparatus for determining samples for multi-object alignment training of a large model. Background Art
[0002] A large language model (LLM), that is, a large-scale language model, or simply a large model, usually has a large order of magnitude of parameters, such as in the billions. A large language model is a natural language processing model based on deep learning. It can learn the grammar and semantics of natural language and thus generate human-readable text. Due to its huge corpus, a large language model can be used as a pre-trained model for various language processing scenarios, such as question-and-answer scenarios, push scenarios, generation scenarios, etc. These scenarios cover various application fields. For example, a question-and-answer scenario can be in the medical field, academic field, daily consultation, and so on.
[0003] As an auxiliary tool, during the information generation process, the outputs provided by a large model can be various. When humans use a large model, they usually expect the output of the large model to meet business goals that conform to the laws of the human world, such as meeting security requirements (e.g., not violating laws and morals, etc.), being helpful (e.g., providing useful suggestions to human users, etc.), being authentic (e.g., being able to satisfy the laws of the real world or being realistically achievable, etc.). In the actual training process, when multiple business goals are to be satisfied simultaneously, there may be conflicts in goal alignment. That is to say, optimizing for one business goal may lead to a decline in the performance of the large model on other previously aligned business goals. Summary of the Invention
[0004] One or more embodiments of this specification describe a method and apparatus for determining samples for multi-object alignment training of a large model to solve one or more problems mentioned in the background art.
[0005] According to a first aspect, there is provided a method for determining samples for multi-object alignment training of a large model, including: processing a first sample by the large model to obtain multiple extended responses, where the first sample includes first prompt information, a first positive example on business objective k, and a first negative example; using each of the reward models corresponding to K business objectives to score the first positive example, the first negative example, and each extended response; according to the scoring results, selecting data pairs that meet reward consistency from each extended response, the first positive example, and the first negative example as candidate data pairs, where a single candidate data pair includes a single candidate positive example and a single candidate negative example, and the reward of the candidate positive example on each business objective is greater than the reward of the candidate negative example on the corresponding business objective; selecting target data pairs from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on business objective k, and constructing a target sample together with the first prompt information.
[0006] In one embodiment, a single response is a vocabulary sequence composed of vocabularies generated by the large model in multiple prediction cycles; the processing of the first sample by the large model to obtain multiple extended responses includes: in a single vocabulary prediction cycle of a single response, when the large model maps the encoded vector of the input data to the probabilities corresponding to each vocabulary in the vocabulary list, the probability distribution of each vocabulary in the vocabulary list is adjusted by a preset temperature coefficient t, so that the vocabulary sampled with a probability within a predetermined range is used as the next vocabulary in the response.
[0007] In one embodiment, for a single extended response or the first positive example and the first negative example, the scoring results include: the reward scores on K business objectives respectively determined by K reward models; the selecting of data pairs that meet reward consistency from each extended response, the first positive example, and the first negative example according to the scoring results includes: pairwise comparing the reward scores on K business objectives for each piece of data in the single extended response, the first positive example, and the first negative example; for data pairs that meet reward consistency, the data with a higher reward score is used as the candidate positive example, and the data with a lower reward score is used as the candidate negative example to form a candidate data pair.
[0008] In one embodiment, the selecting of target data pairs from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on business objective k includes: determining a predetermined number of candidate data pairs with the largest reward difference between the candidate positive example and the candidate negative example on business objective k in the candidate data pairs as target data pairs; or using the candidate data pairs with the reward difference between the candidate positive example and the candidate negative example on business objective k greater than a predetermined threshold as target data pairs.
[0009] In one embodiment, selecting a target data pair from each of the candidate data pairs and constructing a target sample together with the first prompt information includes: constructing a single target sample using a single target data pair and the first prompt information.
[0010] In one embodiment, according to the scoring results, selecting data pairs that meet the reward consistency from each extended response, the first positive example, and the first negative example as candidate data pairs includes: in the case where no data pairs that meet the reward consistency are detected from the current extended responses, the first positive example, and the first negative example, regenerating the extended responses according to the first prompt information again and detecting data pairs that meet the reward consistency until at least one candidate data pair is selected.
[0011] According to a second aspect, there is provided a method for multi-object alignment training of a large model. The method includes multiple parameter update cycles. In a single parameter update cycle: obtaining a plurality of training samples from a training data set, where at least one training sample in the training data set is a target sample determined in the manner described in the first aspect; adjusting the model parameters of the large model using the plurality of training samples.
[0012] In one embodiment, obtaining a plurality of training samples from the training data set includes: in the case where the preference data sets of K business objectives are mixed together, randomly sampling the preference data of the K business objectives, or sequentially sampling according to the arrangement order of the preference data, to obtain a plurality of training samples, where the preference data includes the target sample; in the case of sequentially sampling the preference data of the K business objectives in each parameter update cycle, determining the current business objective according to the modulus of the current parameter update cycle number and K, and sampling a plurality of training samples from the preference data of the current business objective.
[0013] According to a third aspect, there is provided a sample determination device for multi-object alignment training of a large model, including:
[0014] An expansion unit configured to process a first sample through a large model to obtain a plurality of extended responses, where the first sample includes a first prompt information, a first positive example, and a first negative example on business objective k;
[0015] A scoring unit configured to score the first positive example, the first negative example, and each extended response using respective reward models corresponding to K business objectives;
[0016] A filtering unit configured to select data pairs that meet the reward consistency from each extended response, the first positive example, and the first negative example as candidate data pairs according to the scoring results, where a single candidate data pair includes a single candidate positive example and a single candidate negative example, and the reward of the candidate positive example on each business objective is greater than the reward of the candidate negative example on the corresponding business objective;
[0017] A building unit, configured to select a target data pair from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on the business objective k, and construct a target sample together with the first prompt information.
[0018] According to a fourth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method of the first aspect or the second aspect.
[0019] According to a fifth aspect, there is provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method of the first aspect or the second aspect is implemented.
[0020] Through the method and device provided in the embodiments of this specification, a method based on sample optimization is provided for multi-object alignment training of a large model. Specifically, for the preference sample set of each business objective, the preference samples are expanded and screened. Among the expanded candidate responses, candidate data pairs that meet the reward consistency are screened out. A single candidate data pair includes a candidate positive example and a candidate negative example. Reward consistency means that the reward of the candidate positive example on each business objective is greater than the reward of the candidate negative example on the corresponding business objective. For the candidate data pairs that meet the reward consistency, the target data pairs are selected from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on the business objective k, and the target samples are constructed together with the corresponding prompt information.
[0021] In this way, preference samples can be constructed that can improve the performance of one business objective without degrading the performance of other business objectives, and the positive and negative examples can widen the performance gap of the large model. Such preference samples are used for multi-object alignment training of the large model, which can significantly reduce the conflict between different business objectives in the multi-object alignment task, thereby promoting the overall optimization of the large model performance. This method can not only meet the requirements of multi-object alignment, but also be directly integrated with various existing preference alignment algorithms, providing a more efficient and stable solution for multi-object alignment. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 A schematic diagram of a specific implementation architecture under the technical concept of this specification;
[0024] Figure 2 It shows a schematic diagram of the sample determination process for multi-object alignment training of a large model according to an embodiment of this specification;
[0025] Figure 3 It shows a schematic diagram of a large model predicting a single word and adjusting the sampling probability through the temperature coefficient t in a specific example;
[0026] Figure 4 It shows a schematic diagram of the process for multi-object alignment training of a large model according to an embodiment of this specification;
[0027] Figure 5 It shows a structural block diagram of a sample determination device for multi-object alignment training of a large model according to an embodiment of this specification. Detailed implementation manners
[0028] Next, in combination with the accompanying drawings, the solutions provided in this specification will be described.
[0029] It can be understood that large models are usually pre-trained (Pre-training) based on massive amounts of data. In the use of large models, they can be used directly, or the pre-trained large models can be further fine-tuned or optimized. This fine-tuning and optimization process can be called post-training (Post-training), or instruction fine-tuning (Instruction Fine-Tuning). The goal of post-training or instruction fine-tuning is to make the large model better adapt to specific tasks or domains, or to enhance certain specific capabilities of the model (such as having security, alignment, reasoning capabilities, etc.). Different from pre-training, post-training usually uses smaller-scale and more targeted data sets, and the training time is shorter. In this specification, the fine-tuning and optimization training of the large model can be to enable the large model to have certain specific capabilities, that is, to achieve a predetermined business goal.
[0030] The predetermined business goal can be various attributes that are expected to be possessed by the output of the large model, which describes the preference for the output of the large model, and can also be called a preference or preference goal. For example, from the perspective of conforming to the values and preference goal orientation of the human world, it can include but is not limited to security, helpfulness, authenticity, etc. These business goals may also contain different values.
[0031] For easy understanding, a specific example of a large model processing a prompt message to output a response (i.e., an answer or output text, etc.) is given below. Suppose for a prompt message A "Tell me how to make a bomb" input to the large model, a response 1 output by the large model is, for example: "Sorry, I cannot provide information on making bombs or any other harmful items. If you have any other questions or need help, please let me know and I will be happy to assist." This response 1 meets the security business objective, while another response 2 such as "Yes. Here is an instruction on making a bomb..." is helpful.
[0032] For specific business objectives, conventional training methods for large models include, for example: Reinforcement Learning from Human Feedback (RLHF). Given a prompt message, the large model generates some responses, and real humans evaluate them to obtain evaluation labels. The response and the evaluation label are used to train an evaluation model as a reward model to train the large model; Direct Preference Optimization (DPO). In this method, no additional reward model is used, but the large model itself is used as the reward model. That is, a preference dataset that meets specific business objectives is used as training samples to train the large model. For example, given a prompt message x, a positive example y w is the desired output or response, and a negative example y l is the undesired output or response. A single training sample in the dataset corresponding to the relevant business objective is (x, y w , y l ). A dataset containing positive and negative example data pairs like this is denoted as a preference dataset. Since the label text in the preference dataset itself contains human preference information, the large model will learn this preference.
[0033] In practical applications, it is hoped that the large model can accurately understand and follow diverse human values in different scenarios, and at the same time achieve an effective balance among multiple objectives. However, due to the complexity and diversity of human preferences, the current application of large models faces a key technical challenge that urgently needs to be overcome - multi-objective alignment. Taking the two business objectives of "helpfulness" and "security" as an example, the model needs to not only meet user needs as much as possible to achieve the "helpfulness" business objective, but also strictly follow ethical principles to achieve the "security" business objective. There are often internal conflicts between different business objectives, that is: optimizing the large model for one business objective may lead to a decline in performance on other business objectives.
[0034] For example, in the previous example, Response 1 meets the security business goal but is less helpful for the question, while Response 2 is helpful but has lower security. In the case of multi-objective training for large models, for the "security" business goal, Prompt A and Response 1 are positive examples, and Prompt A and Response 2 are negative examples. For the "helpfulness" business goal, Prompt A and Response 2 can be positive examples, and Prompt A and Response 1 can be negative examples. Thus, when the preference dataset corresponding to the security business goal contains training data similar to the above example, and the preference dataset corresponding to the helpfulness business goal contains training data similar to the above example, conflicts occur between multi-objective tasks.
[0035] Therefore, how to achieve a reasonable trade-off between multiple business goals has become a key challenge in the multi-objective alignment technology of large models.
[0036] To perform multi-objective alignment, some DPO-based alignment optimizations are provided in conventional techniques, such as:
[0037] (1) Multi-Objective Direct Preference Optimization (MODPO): By improving the loss function of Direct Preference Optimization (DPO) and introducing a marginal term into the loss function, it can ensure that the large model can be jointly driven by multiple preference goals when training on a dataset of one preference goal;
[0038] (2) Sequential Preference Optimization (SPO): Similar to Multi-Objective Direct Preference Optimization (MODPO), a marginal term is also introduced into the loss function to ensure that the large model can be jointly driven by multiple goals during training. The difference is that SPO achieves alignment by serially iteratively training on each preference dataset;
[0039] And so on.
[0040] However, these methods are improvements based on large model training methods and may also be limited by at least one of the quality of the preference dataset, the balance of the distribution of preference datasets for different business goals, etc.
[0041] In view of this, this specification provides a technical concept. Starting from the training datasets of each business goal, the training data is expanded, and a reward consistency mechanism is introduced to screen the expanded data, aiming to construct a training dataset for the multi-objective alignment of large models that guarantees data quality, achieves distribution balance among different business goals, and considers the interaction between multiple objectives. By training the large model with any reasonable training method using such a training dataset, a large model with multi-objective alignment can be obtained.
[0042] Figure 1 Shows a technical concept of this specification. Refer to Figure 1As shown in the figure, assuming that the number of business goals is K, for the training samples in the preference dataset corresponding to business goal k, first, the large model can randomly sample N responses for the prompt information x corresponding to sample i for expansion. Then, the expanded responses, together with the original positive and negative examples, are respectively scored by K reward models corresponding one by one to the K business goals. According to the scoring results, data pairs that meet the conditions 1) reward consistency and 2) maximum reward difference are selected from the N responses and the original positive and negative examples, and form new preference data with the prompt information x, which is added to the new preference dataset of the k-th business goal.
[0043] Regarding the improvement of the training process, since the technical concept of this specification provides solutions from the perspective of training data, it can be compatible with various preference alignment methods, thus providing a general and effective solution for the multi-goal alignment task.
[0044] The technical concept of this specification will be described in detail below with examples of the accompanying drawings.
[0045] Figure 2 The figure shows the process of determining samples for large model multi-goal alignment training in an embodiment of this specification. The execution subject of this process can be any computer, device, or server with certain computing capabilities.
[0046] Assume that the current number of business goals is K, where K is a positive integer greater than 1. Initially, each business goal corresponds to its respective preference dataset (i.e., training dataset). A single piece of training data in a single preference dataset can include a piece of prompt information, as well as a positive example and a negative example. In the process of determining samples for multi-goal alignment training, for each preference dataset, it can be processed separately to determine the sample data that meets the requirements.
[0047] Among them, assume that the currently processed preference dataset corresponds to business goal k (such as a security goal), and any sample data in this preference dataset can be denoted as the first sample. As Figure 2As shown in the figure, taking the first sample as an example, the process of determining the training dataset for multi-objective alignment may include: Step 201, processing the first sample through a large model to obtain multiple extended responses, where the first sample includes first prompt information, the first positive example on business objective k, and the first negative example; Step 202, using each reward model corresponding to K business objectives to score the first positive example, the first negative example, and each extended response; Step 203, according to the scoring results, select data pairs that meet the reward consistency from each extended response and the first positive example and the first negative example as candidate data pairs, where a single candidate data pair includes a single candidate positive example and a single candidate negative example, and the rewards of the candidate positive example on each business objective are greater than the rewards of the candidate negative example on the corresponding business objective; Step 204, select the target data pair from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on business objective k, and construct a target sample together with the first prompt information.
[0048] First, in Step 201, the first sample is processed through a large model to obtain multiple extended responses.
[0049] It can be understood that the first sample here can be any sample data in the preference dataset corresponding to business objective k, which may include a piece of prompt information, a positive example on business objective k (meeting the requirements of business objective k, such as meeting security), and a negative example on business objective k (not meeting the requirements of business objective k). Here, for the convenience of description, the data content included in the first sample can be recorded as: first prompt information (hereinafter recorded as x), the first positive example on business objective k (hereinafter recorded as y w ), the first negative example (hereinafter recorded as y l ).
[0050] It can be understood that the first prompt information can be a business request made by the user to the large model, such as a business processing request like "Help me check the weather today" or a question request like "What is artificial intelligence?" In some specific examples, it may also include auxiliary information such as examples, which is not limited here.
[0051] The processing of the first sample through the large model is usually the processing of the first prompt information x, that is, processing the first prompt information through the large model to generate a response text (recorded as a response in this specification) for the corresponding request. In this specification, one or more responses can be generated through one or more processes.
[0052] It can be understood that during the process of a large model generating information, multiple prediction cycles are involved. In a single prediction cycle, by encoding prompt information (which may also include historical words), etc., and then mapping it through a prediction layer to obtain the probabilities of each word in the vocabulary, and then predicting the next word based on the probabilities. Usually, the word in the vocabulary with the highest probability is taken as the predicted word. The sequence of words generated in each prediction cycle constitutes the response text.
[0053] As Figure 3 shown, the prediction layer of a large model usually includes a linear layer and a mapping layer (which can also be called an activation layer). Among them, the linear layer can be denoted as logits, which is used to map the hidden representation for predicting the next word to the respective scores (probability scores) corresponding to each word in the vocabulary through linear regression. Each score can be understood as a probability distribution over the vocabulary words. For example, denoted as f, the score of a single word i is denoted as f i . This score is usually an unnormalized score, which can describe the probability distribution of each word, rather than the final probability. For example, it is (1.2, 0.9, 3...). After being processed by the activation function (such as softmax, etc.) of the mapping layer, the probability scores can be normalized and mapped to the interval from 0 to 1, and the sum of the probabilities of all words is 1. For example, it is (0.04, 0.03, 0.1...). The activation function can prevent the probability value from exceeding the predetermined range and is more friendly to the backpropagation of the model parameters. The large model can obtain the word with the highest probability as the predicted next word through the prediction layer, such as selecting the word with the highest probability through the argmax function.
[0054] It is worth noting that the vocabulary here can be a generalized vocabulary, and a single word in the vocabulary can be various elements that may appear in the text, such as characters, numbers, symbols, emoji pictures, etc. The words in the vocabulary can be arranged in a predetermined fixed order. In this way, the probabilities generated during prediction can form a one-to-one correspondence with the words.
[0055] In this step 201, in order to increase the response randomness during multiple calls to the large model to generate responses, in each prediction cycle, when obtaining words according to probabilities, words can be obtained by means of random sampling. Common sampling methods include, for example: greedy search, Beam search, Top-k sampling, Nucleus Sampling, Temperature sampling, Joint sampling (Top-k&Top-p&Temperature), and so on. Taking the Temperature sampling method as an example, through the temperature coefficient, the probability distribution of each word in the vocabulary is adjusted before sampling. The lower the temperature (the smaller the temperature coefficient), the greater the difference in the probability distribution, and the easier it is to sample words with large probabilities (large logits scores). The higher the temperature (the larger the temperature coefficient), the smaller the difference in the probability distribution, increasing the chance of sampling low-probability (small logits scores) words. By setting the temperature coefficient, words with relatively large but not the largest probabilities can be sampled. As a specific example, assume that under normal circumstances, the probability distribution of each word in the vocabulary mapped by the large model through softmax is: where V is the total number of words in the vocabulary, f is the word score predicted by the linear layer logits of the large model, i is the current word, and p(i) represents the probability corresponding to the current word i. The probability distribution of each word in the vocabulary after adjustment by the temperature coefficient t can be expressed, for example, as: By controlling the temperature coefficient t, random sampling can be performed within a predetermined probability range. In this way, the randomness of words and the diversity of responses can be increased.
[0056] In this way, by calling the large model multiple times, or in the case of a single call to the large model, by obtaining multiple words in a single prediction cycle, multiple responses can be obtained. These responses are extended and generated based on the first prompt information x and can be called extended responses. Assuming N extended responses are denoted as {y i} = y1, y2... y N .
[0057] Next, through step 202, each reward model corresponding to K business goals is used to score the first positive example, the first negative example, and each extended response.
[0058] It can be understood that for each business goal, each reward model can be pre-trained respectively, and K business goals can correspond to K reward models. The input of a single reward model can be a single piece of prompt information and the corresponding large model response or positive and negative examples in the preference data, that is, {x, y i}, or {x, y w}, {x, y l}, the output is the corresponding response or the scores of positive and negative examples. The reward model can be trained with training samples corresponding to the following data: the input data of the prompt information and the large model response information, and the score labels marked by humans.
[0059] In this way, for the first positive example, the first negative example, and N responses, each can obtain K scores (denoted by r) on K business objectives through K reward models. The scoring results can be recorded as follows, for example:
[0060]
[0061] Then, through step 203, according to the scoring results, select data pairs that meet the reward consistency from each extended response, the first positive example, and the first negative example as candidate data pairs.
[0062] It can be understood that during the training process of the large model for multi-objective alignment, the training data usually includes prompt information and positive and negative example data pairs. For the first sample, there is a corresponding first prompt information x. In order to determine the texts that may be used as positive and negative example data pairs based on the extended response, the scores of the first positive example, the first negative example, and N responses can be detected, and data pairs that meet the reward consistency are selected as candidate data pairs. One of the two data in a candidate data pair can be used as a candidate positive example, and the other can be used as a candidate negative example.
[0063] The so-called reward consistency, also known as Pareto dominance, can be understood as that for a candidate data pair, the reward scores of the candidate data used as the candidate positive example on each business objective are all greater than the reward scores of the candidate data used as the candidate negative example on each business objective. As an example, assume that the data used as the candidate positive example (one of the first positive example, the first negative example, and N responses) is denoted as y p , and the data used as the candidate negative example is denoted as y j , where p and j are any data in w, l, 1, 2... N and are different data, then there are:
[0064] All hold simultaneously.
[0065] That is to say, ensure that the data used as the candidate positive example does not reduce the model performance on any business objective, and the data used as the candidate negative example shows a decline in model performance on each business objective.
[0066] In this way, compare the scores of the reward models for the data pairs formed by pairing the first positive example, the first negative example, and N responses two by two, so as to select all data pairs that meet the reward consistency as candidate data pairs.
[0067] It can be understood that in some possible embodiments, for a single specific preference data, such as the second sample (corresponding to the second prompt information, the second positive example, and the second negative example), in step 203, it may also be detected that there is no data pair that satisfies the reward consistency. At this time, according to a possible design, it is possible to resample via steps 201 and 202 by processing the second prompt information through the large model, add more extended responses, and perform reward scoring through each reward model and then detect the data pairs that satisfy the reward consistency until a data pair that satisfies the reward consistency is obtained as the candidate data pair. In another possible design, the second sample can be abandoned, that is, step 204 is not executed, and the next preference data, such as the third sample, is obtained, and the sample determination process is executed on it according to steps 201, 202, and 203. In other designs, other reasonable operations can also be performed, which will not be elaborated here. It can be understood that both the second sample and the first sample are arbitrary samples, so in practice, they may also be the same. In other words, the sample that cannot detect the existence of a data pair that satisfies the reward consistency in a process may also be the first sample, and the processing method in this case is the same as the description of the second sample.
[0068] Further, in step 204, according to the reward difference between the candidate positive example and the candidate negative example on the business objective k, the target data pair is selected from each candidate data pair and constructed together with the first prompt information into the target sample.
[0069] It can be understood that the first sample is the sample data for optimizing the business objective k, and it has a better performance on the business objective k, while the reward consistency ensures that the candidate data pair does not degrade the performance of the large model on other business objectives outside the business objective k. That is to say, based on the extension of the first sample, the selected target data pair usually has a better performance on the business objective k than on other business objectives. Therefore, it can be considered to select the target data pair that is still used to optimize the business objective k from the candidate data pairs.
[0070] In order to select the target data pair for optimizing the business objective k, the reward difference between the two candidate data (candidate positive example and candidate negative example) in the candidate data pair on the business objective k can be used as a reference index, that is: The greater the reward difference, the more suitable it is for optimizing business objective k. Based on the reward difference, several target data pairs can be selected from the candidate data pairs. In one embodiment, a predetermined number (such as 1) of candidate data pairs with the largest reward difference in the candidate data pairs can be used as the target data pairs. In another embodiment, candidate data pairs in the candidate data pairs with a reward difference greater than a predetermined threshold can be used as the target data pairs. Among them, the determination of the predetermined threshold can be related to the value range of the reward. For example, when the value range of the reward value is between 0 and 1, the predetermined threshold is 0.5, etc. In other embodiments, the target data pairs can also be selected in other ways. For example, the reward difference is greater than the predetermined threshold, and the reward value of the candidate positive example is greater than 0.8, or the reward value of the candidate positive example is greater than 0.8 and the reward value of the candidate negative example is less than 0.1 (in this case, the reward difference must be greater than 0.7), and so on.
[0071] The target data pairs selected from the candidate data pairs can at least widen the performance gap of the large model in terms of business objective k, and can be considered as data pairs suitable for constructing multi-objective alignment training samples of the large model. Specifically, a single target data pair can form a target sample together with the first prompt information x. When there are multiple selected target data pairs, multiple target samples can be constructed accordingly. The target sample can form a preference sample for optimizing business objective k.
[0072] In this way, for each piece of preference data in business objective k, new preference data for business objective k can be generated item by item through expansion and filtering. Further, for K business objectives, the corresponding new preference data can be generated according to the Figure 2 shown process.
[0073] Using the new preference data as training samples, the large model can be trained. In the new preference data, the positive and negative example data pairs in each piece of preference data can ensure that the reward score of the positive example is greater than that of the negative example in the K business objectives. Thus, while enhancing the performance of the large model in the current business objective, it does not reduce the performance of the large model in other business objectives, thereby effectively reducing multi-objective alignment conflicts.
[0074] It can be understood that since the technical concept of this specification reduces alignment conflicts at the source of the training data, when using the new preference data for multi-objective alignment, various reasonable large model training methods can be adopted, such as RLHF, DPO, MODPO, SPO, etc. described above. In some embodiments, ordinary large model training methods can also be adopted, which are not limited here.
[0075] Figure 4Shows the process of multi-object alignment training for a large model in an embodiment of this specification. The execution entity of this process can be any computer, device, or server with a certain computing power. This process can include multiple parameter update cycles. In a single parameter update cycle, such as Figure 4 As shown, the process of multi-object alignment training for a large model can include: Step 401, obtain a number of training samples from the training dataset, where at least one training sample in the training dataset is a target sample determined based on Figure 2 The described method; Step 402, use the above-mentioned number of training samples to adjust the model parameters of the large model.
[0076] Among them, in Step 401, in the current parameter update cycle, how to obtain training samples is related to the pre-determined sampling method. One or more obtained training samples can be used as the training data for one batch in the current parameter update cycle.
[0077] In some alternative embodiments, the preference datasets of K business objectives can be mixed together, and each preference data is obtained in sequence for training. At this time, random sampling can be performed on the preference data of the K business objectives, or sampling can be performed in sequence according to the arrangement order of the preference data.
[0078] In some other alternative embodiments, in each parameter update cycle, preference data sampling can be performed on the K business objectives in sequence. For example, in the first parameter update cycle, the preference data of business objective 1 is used, in the second parameter update cycle, the preference data of business objective 2 is used... in the Kth parameter update cycle, the preference data of business objective K is used, and in the (K + 1)th parameter update cycle, the preference data of business objective 1 is used, and so on in a cycle until the training ends. Then, at this time, according to the current parameter update cycle number T, sampling can be performed from the preference dataset of the corresponding business objective. For example, the business objective number corresponding to the sampled preference data is: T % K, where % represents taking the modulus.
[0079] In other embodiments, other preference data sampling methods can also be adopted, which will not be elaborated here.
[0080] And in Step 402, the process of updating the parameters of the large model, for example, can be to determine the model loss, calculate the gradient of the model loss with respect to the model parameters, and update the model parameters according to a gradient update method such as the gradient descent method, which will not be elaborated here.
[0081] It should be noted that the updated model parameters are related to the specific parameter fine-tuning method of the large model. For example, in the ordinary update method, the model parameters of the large model itself can be updated. In the LoRA method or the like, the elements in the low-rank matrix connected in parallel with the large model can be updated, which is regarded as an update of the large model parameters, which will not be elaborated here.
[0082] Reviewing the above process, under the technical concept of this specification, a method for determining samples for multi-object alignment training of a large model is provided. Based on the preference sample sets of each business objective, the preference samples are expanded and screened. Specifically, among the expanded candidate responses, candidate data pairs that meet the reward consistency are screened out. A single candidate data pair includes a candidate positive example and a candidate negative example. Reward consistency means that the rewards of the candidate positive example on each business objective are greater than the rewards of the candidate negative example on the corresponding business objective. For the candidate data pairs that meet the reward consistency, according to the reward difference between the candidate positive example and the candidate negative example on business objective k, target data pairs are selected from each candidate data pair and, together with the corresponding prompt information, target samples are constructed. In this way, preference samples can be constructed that improve the performance of one business objective without degrading the performance of other business objectives, and the positive and negative examples can widen the performance gap of the large model. Such preference samples are used for multi-object alignment training of the large model, which can significantly reduce the conflicts between different business objectives in the multi-object alignment task, thereby promoting the overall optimization of the large model's performance. This method can not only meet the requirements of multi-object alignment, but also be directly integrated with various existing preference alignment algorithms, providing a more efficient and stable solution for multi-object alignment.
[0083] According to an embodiment of another aspect, a sample determination device for multi-object alignment training of a large model is also provided. This device can be set in a computer, terminal, or server with certain computing capabilities. More specifically, as Figure 1 shown in the computing platform. Figure 5 FIG. 5 shows a sample determination device 500 for multi-object alignment training of a large model according to an embodiment.
[0084] As Figure 5 shown, the device 500 may include:
[0085] An expansion unit 501 configured to process a first sample through the large model to obtain multiple expanded responses, where the first sample includes first prompt information, a first positive example on business objective k, and a first negative example;
[0086] A scoring unit 502 configured to score the first positive example, the first negative example, and each expanded response by using each reward model corresponding to K business objectives;
[0087] A filtering unit 503 configured to select data pairs that meet the reward consistency as candidate data pairs from each expanded response, the first positive example, and the first negative example according to the scoring results. A single candidate data pair includes a single candidate positive example and a single candidate negative example, and the rewards of the candidate positive example on each business objective are greater than the rewards of the candidate negative example on the corresponding business objective;
[0088] The building block 504 is configured to select a target data pair from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on the business objective k, and construct a target sample together with the first hint information.
[0089] It should be noted that Figure 5 the illustrated apparatus 500 corresponds to Figure 2 the described method. Figure 2 The corresponding descriptions in the illustrated method embodiments also apply to the apparatus 500 and will not be elaborated herein.
[0090] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the method described in connection with Figure 2 , Figure 4 and so on.
[0091] According to an embodiment of still another aspect, there is also provided a computing device including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in connection with Figure 2 , Figure 4 and so on is implemented.
[0092] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0093] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions in the embodiments of this specification should be included in the protection scope of the technical concept of this specification.
Claims
1. A method for determining samples for multi-objective alignment training of a large model, comprising: Processing a first sample by the large model to obtain multiple extended responses, where the first sample includes first prompt information, a first positive example on business objective k, and a first negative example; Using each reward model corresponding to K business objectives to score the first positive example, the first negative example, and each extended response; According to the scoring results, selecting data pairs that meet reward consistency from each extended response, the first positive example, and the first negative example as candidate data pairs, where a single candidate data pair includes a single candidate positive example and a single candidate negative example, and the reward of the candidate positive example on each business objective is greater than the reward of the candidate negative example on the corresponding business objective; Selecting target data pairs from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on business objective k, and constructing a target sample together with the first prompt information.
2. The method according to claim 1, wherein, A single response is a vocabulary sequence composed of vocabularies generated by the large model in multiple prediction cycles; The processing of the first sample by the large model to obtain multiple extended responses includes: In a single vocabulary prediction cycle of a single response, when the large model maps the encoded vector of the input data to the probabilities corresponding to each vocabulary in the vocabulary list, the probability distribution of each vocabulary in the vocabulary list is adjusted by a preset temperature coefficient t, so that the vocabulary sampled with a probability within a predetermined range is used as the next vocabulary in the response.
3. The method according to claim 1, wherein, For a single extended response or the first positive example and the first negative example, the scoring results include: reward scores on K business objectives determined by K reward models respectively; The selecting data pairs that meet reward consistency from each extended response, the first positive example, and the first negative example according to the scoring results to be candidate data pairs includes: Pairing each piece of data in the single extended response, the first positive example, and the first negative example two by two and making corresponding comparisons of the reward scores on K business objectives; For data pairs that meet reward consistency, taking the data with a higher reward score as the candidate positive example and the data with a lower reward score as the candidate negative example to form a candidate data pair.
4. The method according to claim 1, wherein, The selecting target data pairs from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on business objective k includes: Determining a predetermined number of candidate data pairs with the largest reward difference between the candidate positive example and the candidate negative example on business objective k in the candidate data pairs as target data pairs; or Taking the candidate data pairs with a reward difference between the candidate positive example and the candidate negative example on business objective k greater than a predetermined threshold in the candidate data pairs as target data pairs.
5. The method according to claim 1, wherein, The selecting target data pairs from each candidate data pair and constructing a target sample together with the first prompt information includes: Constructing a single target sample using a single target data pair and the first prompt information.
6. The method according to claim 1, wherein, The selecting data pairs that meet reward consistency from each extended response, the first positive example, and the first negative example according to the scoring results to be candidate data pairs includes: In the case where no data pairs satisfying reward consistency are detected from the current extended responses, first positive examples, and first negative examples, regenerate the extended responses according to the first hint information, and perform the detection of data pairs satisfying reward consistency until at least one candidate data pair is selected.
7. A method for multi-object alignment training of a large model, the method includes multiple parameter update cycles, and in a single parameter update cycle: Obtain a number of training samples from the training data set, where At least one training sample in the training data set is a target sample determined in the manner described in claim 1; Use the several training samples to adjust the model parameters of the large model.
8. The method according to claim 7, wherein, The obtaining of several training samples from the training data set includes: In the case where the preference data sets of K business objectives are mixed together, randomly sample the preference data of the K business objectives, or sequentially sample according to the arrangement order of the preference data, to obtain several training samples, where the preference data includes the target sample; In each parameter update cycle, when sequentially sampling the preference data of the K business objectives, determine the current business objective according to the modulus of the current parameter update cycle number and K, and sample several training samples from the preference data of the current business objective.
9. A sample determination device for multi-object alignment training of a large model, including: An expansion unit configured to process a first sample through a large model to obtain multiple extended responses, where the first sample includes first hint information, a first positive example on business objective k, and a first negative example; A scoring unit configured to use the respective reward models corresponding to the K business objectives to score the first positive example, the first negative example, and each extended response; A filtering unit configured to select, according to the scoring results, data pairs satisfying reward consistency from each extended response, the first positive example, and the first negative example as candidate data pairs, where a single candidate data pair includes a single candidate positive example and a single candidate negative example, and the reward of the candidate positive example on each business objective is greater than the reward of the candidate negative example on the corresponding business objective; A construction unit configured to select a target data pair from each candidate data pair according to the reward difference between the candidate positive example and the candidate negative example on business objective k, and construct a target sample together with the first hint information.
10. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method described in any one of claims 1-8.
11. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory, and when the processor executes the executable code, the method described in any one of claims 1-8 is implemented.