Large model post-training method and device, computer program and storage medium
By replacing the low-quality candidate answers generated by the large model with the standard answers, and combining reward model evaluation and single-stage training, the problems of cumbersome and unstable large model training processes are solved, the accuracy and security of the generated content are improved, and training efficiency is increased.
Patent Information
- Application Number
- CN202510975987.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-28
AI Technical Summary
Existing large models suffer from problems such as cumbersome training processes, low efficiency, and instability in the early stages of training. In particular, the quality of candidate answers generated by the model varies greatly during the reinforcement learning stage, and it is easy to get stuck in local optima.
By acquiring multiple candidate answers generated by a large model, replacing low-quality or unsafe candidate answers with the standard answer, evaluating and dynamically adjusting the intervention intensity using a reward model, and combining this with a single-stage training process, the model generation strategy is optimized.
It improves the accuracy and security of content generated by large models, reduces computational resource consumption and training time, balances the stability and diversity of generated content, and achieves efficient model optimization.
Smart Images

Figure CN120851117A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, computer program, and storage medium for training large models. Background Technology
[0002] With the rapid development of artificial intelligence technology, large models have demonstrated powerful performance and broad application prospects in many fields. However, in practical applications, large models often face problems such as unstable quality of generated content, insufficient security, and inadequate alignment. Post-training methods have become a key means to improve model performance.
[0003] In related technologies, a typical post-training method employs a multi-stage training process. First, the model is trained using labeled data. This first stage optimizes model parameters to ensure the model's output closely approximates the expected results in the labeled data, thereby equipping the model with basic reasoning ability and alignment. Next, the model enters the second stage of reinforcement learning training. In this stage, the model generates multiple candidate answers, evaluates these answers using reward signals, and updates the model parameters based on the reward signals to optimize the model's generation strategy and improve the quality and diversity of the generated content.
[0004] While this multi-stage post-training method can improve model performance to some extent, it suffers from drawbacks such as cumbersome training processes, low efficiency, and instability in the initial training phase. For example, the training process requires separate training stages, resulting in long training cycles and high resource consumption. In the reinforcement learning stage, the inconsistent quality of candidate answers generated by the model can lead to training instability and a tendency to get stuck in local optima. Therefore, there is an urgent need for an efficient and stable post-training method that can solve the above problems. Summary of the Invention
[0005] In view of this, this disclosure proposes a large model post-training technique.
[0006] According to one aspect of this disclosure, a method for post-training a large model is provided, comprising:
[0007] Obtain the output set containing multiple candidate answers generated by a large model for the same request question;
[0008] Replace at least one of the multiple candidate answers with the standard answer to obtain the updated output group;
[0009] The large model is trained based on the updated output set.
[0010] In one possible implementation, replacing at least one of the plurality of candidate answers with the standard answer includes:
[0011] The reward value is calculated for each of the multiple candidate answers using a reward model.
[0012] Determine at least one candidate answer with the lowest reward value among the plurality of candidate answers;
[0013] Replace at least one candidate answer with the lowest reward value with the standard answer.
[0014] In one possible implementation, the intervention intensity of the standard answer on multiple candidate answers gradually decreases as the training phase of the large model evolves, and replacing at least one of the multiple candidate answers with the standard answer includes:
[0015] The intervention intensity of the standard answer is dynamically adjusted according to the training progress, wherein the intervention intensity is negatively correlated with the training progress.
[0016] In one possible implementation, the intervention intensity of dynamically adjusting the standard answer according to the training process includes:
[0017] In the first phase of training the large model, at least one of the multiple candidate answers is replaced with the standard answer; and / or,
[0018] In the second stage of training the large model, at least one of the multiple candidate answers is replaced with a standard answer that has been subjected to random perturbation; and / or,
[0019] In the third stage of training the large model, training is performed based on the candidate answers originally generated by the large model, wherein the first stage precedes the second stage, and the second stage precedes the third stage.
[0020] In one possible implementation, the method further includes:
[0021] Calculate the initial standard deviation of the reward value, std0;
[0022] If the standard deviation of the reward value of the current batch of data is in the interval [1 / a×std0, std0], then the current training is determined to be in the first stage; where 1 / a < 1.
[0023] If the standard deviation of the reward value of the current batch of data is in the interval [1 / b×std0, 1 / a×std0], then the current training is determined to be in the second stage; where 1 / b < 1 / a.
[0024] If the standard deviation of the reward value of the current batch of data is less than or equal to 1 / b×std0, the current training is determined to be in the third stage.
[0025] In one possible implementation, the method further includes:
[0026] Calculate the error rate of a large model;
[0027] If the error rate of the answer is higher than the preset error rate threshold, the current training is determined to be in the first stage.
[0028] If the error rate is within the first preset range, the current training is determined to be in the second stage.
[0029] If the error rate is lower than the target threshold, the current training is determined to be in the third stage, wherein the preset error rate threshold is higher than the target threshold, and the first preset interval is between the preset error rate threshold and the target threshold.
[0030] In one possible implementation, training the large model based on the updated output set includes:
[0031] Calculate the mean and variance of the reward values for candidate answers in the updated output group;
[0032] A dominance function is generated based on the mean and variance.
[0033] The large model is trained based on the aforementioned advantage function.
[0034] In one possible implementation, training the large model based on the advantage function includes:
[0035] Construct multiple input pairs based on system prompts and multiple candidate answers in the updated output group;
[0036] Each input pair is input into the large model to obtain the original predicted values corresponding to multiple candidate answers in the updated output group;
[0037] A policy optimization loss function is constructed based on the advantage function and the original predicted values, and the parameters of the large model are updated using the policy optimization loss function.
[0038] In one possible implementation, the step of constructing a policy optimization loss function based on the advantage function and the original predicted values, and using the policy optimization loss function to update the parameters of the large model, includes:
[0039] The advantage function is used as a weighting factor, and the probability logarithm corresponding to the original predicted value is weighted and summed to construct the policy optimization loss function.
[0040] The gradient of the policy optimization loss function with respect to the model parameters is calculated using the gradient backpropagation algorithm.
[0041] The parameters of the large model are updated based on the gradient.
[0042] In one possible implementation, replacing at least one of the plurality of candidate answers with the standard answer includes:
[0043] Using manually annotated standard answers, randomly replace one or more of the multiple candidate answers.
[0044] In one possible implementation,
[0045] According to another aspect of this disclosure, a large model post-training apparatus is provided, comprising:
[0046] The acquisition module is used to acquire the output set containing multiple candidate answers generated by the large model for the same request question;
[0047] The replacement module is used to replace at least one of the multiple candidate answers with the standard answer to obtain an updated output group;
[0048] The training module is used to train the large model based on the updated output set.
[0049] In one possible implementation, the replacement module is used to:
[0050] The reward value is calculated for each of the multiple candidate answers using a reward model.
[0051] Determine at least one candidate answer with the lowest reward value among the plurality of candidate answers;
[0052] Replace at least one candidate answer with the lowest reward value with the standard answer.
[0053] In one possible implementation, the intervention strength of the standard answer on multiple candidate answers gradually decreases as the training phase of the large model evolves, and the replacement module is used to:
[0054] The intervention intensity of the standard answer is dynamically adjusted according to the training progress, wherein the intervention intensity is negatively correlated with the training progress.
[0055] In one possible implementation, the replacement module is used to:
[0056] In the first phase of training the large model, at least one of the multiple candidate answers is replaced with the standard answer; and / or,
[0057] In the second stage of training the large model, at least one of the multiple candidate answers is replaced with a standard answer that has been subjected to random perturbation; and / or,
[0058] In the third stage of training the large model, training is performed based on the candidate answers originally generated by the large model, wherein the first stage precedes the second stage, and the second stage precedes the third stage.
[0059] In one possible implementation, the apparatus further includes a first-stage partitioning module, used for:
[0060] Calculate the initial standard deviation of the reward value, std0;
[0061] If the standard deviation of the reward value of the current batch of data is in the interval [1 / a×std0, std0], then the current training is determined to be in the first stage; where 1 / a < 1.
[0062] If the standard deviation of the reward value of the current batch of data is in the interval [1 / b×std0, 1 / a×std0], then the current training is determined to be in the second stage; where 1 / b < 1 / a.
[0063] If the standard deviation of the reward value of the current batch of data is less than or equal to 1 / b×std0, the current training is determined to be in the third stage.
[0064] In one possible implementation, the device further includes a second-stage partitioning module, used for:
[0065] Calculate the error rate of a large model;
[0066] If the error rate of the answer is higher than the preset error rate threshold, the current training is determined to be in the first stage.
[0067] If the error rate is within the first preset range, the current training is determined to be in the second stage.
[0068] If the error rate is lower than the target threshold, the current training is determined to be in the third stage, wherein the preset error rate threshold is higher than the target threshold, and the first preset interval is between the preset error rate threshold and the target threshold.
[0069] In one possible implementation, the training module is used for:
[0070] Calculate the mean and variance of the reward values for candidate answers in the updated output group;
[0071] A dominance function is generated based on the mean and variance.
[0072] The large model is trained based on the aforementioned advantage function.
[0073] In one possible implementation, the training module is used for:
[0074] Construct multiple input pairs based on system prompts and multiple candidate answers in the updated output group;
[0075] Each input pair is input into the large model to obtain the original predicted values corresponding to multiple candidate answers in the updated output group;
[0076] A policy optimization loss function is constructed based on the advantage function and the original predicted values, and the parameters of the large model are updated using the policy optimization loss function.
[0077] In one possible implementation, the training module is used for:
[0078] The advantage function is used as a weighting factor, and the probability logarithm corresponding to the original predicted value is weighted and summed to construct the policy optimization loss function.
[0079] The gradient of the policy optimization loss function with respect to the model parameters is calculated using the gradient backpropagation algorithm.
[0080] The parameters of the large model are updated based on the gradient.
[0081] In one possible implementation, the replacement module is used to:
[0082] Using manually annotated standard answers, randomly replace one or more of the multiple candidate answers.
[0083] According to another aspect of this disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described method.
[0084] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.
[0085] According to another aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0086] In this embodiment, after obtaining the output set containing multiple candidate answers generated by the large model for the same request question, an updated output set is obtained by replacing at least one of the multiple candidate answers with the standard answer; then, the large model is trained based on the updated output set. Thus, by replacing low-quality or insecure candidate answers, the large model can directly learn the standardized output labeled by humans during training, thereby improving the accuracy and security of the generated content. Simultaneously, the single-stage training process avoids the overhead of repeatedly switching training targets in traditional multi-stage methods, reducing computational resource consumption and training time. Furthermore, retaining unreplaced candidate answers provides the model with autonomous exploration space, balancing the stability and diversity of the generated content. This combination of dynamic replacement and single-stage training allows the model to efficiently optimize the generation strategy without relying on complex external modules or human intervention.
[0087] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0088] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0089] Figure 1 A flowchart illustrating a large model post-training method according to an embodiment of the present disclosure is shown.
[0090] Figure 2 An example diagram of a large model post-training method according to an embodiment of this disclosure is shown.
[0091] Figure 3 An example diagram of a large model post-training method according to an embodiment of this disclosure is shown.
[0092] Figure 4 A block diagram of a large model post-training apparatus according to an embodiment of the present disclosure is shown.
[0093] Figure 5 A block diagram of an apparatus for post-training a large model according to an embodiment of the present disclosure is shown. Detailed Implementation
[0094] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0095] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.
[0096] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.
[0097] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0098] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0099] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0100] In today's era of rapid development in artificial intelligence, large models have demonstrated enormous potential and broad application prospects in numerous fields. From natural language processing to multimodal tasks such as text generation, machine translation, intelligent customer service, and content recommendation, large models are playing a vital role.
[0101] However, current large-scale models suffer from drawbacks such as cumbersome training processes, low efficiency, and instability in the initial training phase. For example, training requires separate stages, resulting in long training cycles and high resource consumption. In the reinforcement learning stage, the inconsistent quality of candidate answers generated by the model can lead to training instability and a tendency to get stuck in local optima. Therefore, there is an urgent need for an efficient and stable post-training method that can solve these problems.
[0102] To meet these needs, this disclosure proposes a post-training method for large models. By optimizing the training process, this method can effectively improve the training efficiency of large models and enhance their performance, making them more accurate, reliable, and efficient in practical applications.
[0103] Figure 1 A flowchart illustrating a large model post-training method according to an embodiment of the present disclosure is shown. Figure 1 As shown, the method includes:
[0104] In step S11, the output group containing multiple candidate answers generated by the large model for the same request question is obtained;
[0105] Candidate answers are multiple potential outputs generated by a large model for the same input question. These large models are typically based on deep learning architectures (such as Transformers) and acquire broad language understanding and generation capabilities through pre-training. When a user requests a question, the large model uses its parameterized knowledge base to generate multiple candidate answers.
[0106] The generation process can be implemented through different decoding strategies, such as beam search, random sampling, and temperature scaling. The choice of these strategies will affect the quality and diversity of candidate answers. This disclosure does not limit the specific decoding strategy.
[0107] For example, when a user asks "How to make a cake", the model generates 10 candidate answers, covering different combinations of ingredients, order of steps, or decoration suggestions.
[0108] In step S12, at least one of the multiple candidate answers is replaced with the standard answer to obtain the updated output group;
[0109] Based on preset rules or dynamic evaluation, some candidate answers are replaced with manually labeled standard answers. Standard answers can be ideally labeled or authoritatively defined outputs, used to guide the accuracy and safety of the content generated by the model. In supervised training, standard answers are typically labeled by domain experts to ensure compliance with specific norms (such as factual accuracy and ethical guidelines).
[0110] For example, in medical question-answering tasks, the standard answer may be a medically validated treatment plan to avoid the model generating misleading suggestions.
[0111] There can be multiple specific replacement strategies. For example, strategies such as combining reward values from reward models, semantic analysis, or manual intervention can be used to select low-quality or high-risk answers for replacement. Specific implementation methods are available in this publication and will not be elaborated upon here.
[0112] In step S13, the large model is trained based on the updated output group.
[0113] After updating the output set, the parameters of the large model can be optimized, improving the quality and security of the content generated by the large model. The training process may combine supervised learning and reinforcement learning objectives, using a loss function to train the large model, making its output biased towards high-quality answers.
[0114] During training, the large model calculates gradients based on the replaced answer sets and adjusts parameters to reduce the probability of generating low-quality content and improve semantic alignment with the standard answer. Specific implementation details can be found in the possible implementations provided in this publication and will not be elaborated upon here.
[0115] In this embodiment, after obtaining the output set containing multiple candidate answers generated by the large model for the same request question, an updated output set is obtained by replacing at least one of the multiple candidate answers with the standard answer; then, the large model is trained based on the updated output set. Thus, by replacing low-quality or insecure candidate answers, the large model can directly learn the standardized output labeled by humans during training, thereby improving the accuracy and security of the generated content. Simultaneously, the single-stage training process avoids the overhead of repeatedly switching training targets in traditional multi-stage methods, reducing computational resource consumption and training time. Furthermore, retaining unreplaced candidate answers provides the model with autonomous exploration space, balancing the stability and diversity of the generated content. This combination of dynamic replacement and single-stage training allows the model to efficiently optimize the generation strategy without relying on complex external modules or human intervention.
[0116] In one possible implementation, replacing at least one of the multiple candidate answers with the standard answer includes: calculating reward values for each of the multiple candidate answers using a reward model; determining at least one candidate answer with the lowest reward value among the multiple candidate answers; and replacing the at least one candidate answer with the lowest reward value with the standard answer.
[0117] In this implementation, a reward model is used to filter and replace low-quality candidate answers. That is, by using quantitative evaluation of reward values, answers that need to be optimized are dynamically identified and replaced with standard answers to improve the effectiveness of model training.
[0118] Specifically, the first step is to score multiple candidate answers using a reward model to generate corresponding reward values. The reward model is a pre-trained evaluation module used to quantify and score specific attributes of the generated content, such as security, accuracy, or user satisfaction.
[0119] For example, in a customer service conversation scenario, the reward model might calculate a comprehensive score based on dimensions such as compliance of the answer, completeness of information, and friendliness of tone. The reward value awarded to each candidate answer reflects its performance in these dimensions.
[0120] Next, all candidate answers can be sorted according to their reward values to identify one or more answers with the lowest scores. Here, "lowest" can mean an absolute score below a preset threshold or the answer that ranks last within the group.
[0121] Then, low-reward answers are replaced with standard answers, ensuring that the training data includes high-quality reference outputs. For example, if a large model generates five answers about "data privacy protection," and one answer receives the lowest score from the reward model due to missing a key legal clause, it is replaced with a complete legal statement reviewed by experts.
[0122] The standard answer here can be manually labeled or generated using a large, highly reliable model; this publication does not impose any restrictions on it.
[0123] In this embodiment, a reward model is used to calculate reward values for each of the multiple candidate answers; at least one candidate answer with the lowest reward value is determined; and this candidate answer with the lowest reward value is replaced with a manually labeled standard answer. Thus, the reward model quantitatively evaluates answer quality, providing an objective basis for the replacement operation and reducing the subjectivity and inefficiency of manual screening. Secondly, replacing low-reward-value answers can precisely eliminate weaknesses in the model's generation process, such as correcting inaccurate diagnostic suggestions in medical Q&A, directly improving the security of the generated content. Furthermore, retaining unreplaced answers provides the model with room for autonomous exploration; for example, in creative writing tasks, it allows the model to generate diverse results while maintaining overall quality by replacing outcomes that deviate significantly from logic.
[0124] In one possible implementation, replacing at least one of the plurality of candidate answers with a standard answer includes: randomly replacing one or more of the plurality of candidate answers using a manually annotated standard answer.
[0125] In this embodiment, by randomly replacing one or more candidate answers with manually labeled standard answers, correct answer information can be directly introduced into the training data. This replacement operation effectively corrects erroneous answers generated by the model, making the training data more accurate and reliable. Simultaneously, the random replacement strategy increases the diversity of the training data, preventing the model from overfitting to specific error patterns and improving its generalization ability and adaptability. In this way, the model can learn a wider range of correct answer patterns during training, thereby answering questions or completing tasks more accurately in practical applications, improving the overall performance and stability of the model.
[0126] In one possible implementation, training the large model based on the updated output set includes: calculating the mean and variance of the reward values of candidate answers in the updated output set; generating an advantage function based on the mean and variance; and training the large model based on the advantage function.
[0127] When training the large model based on the updated output group, an advantage function can be generated using the reward value distribution of the candidate answer group, thereby guiding the update of model parameters. By quantifying the overall performance and individual differences of answers within the group, more refined optimization signals are provided for the training process, which retains the policy optimization ideas of traditional reinforcement learning while reducing the dependence on external value evaluation modules through intra-group statistics.
[0128] In practical implementation, the mean (μ) and variance (σ) of the reward value can be calculated for the updated candidate answer group (including the replaced standard answer and the original answer that was not replaced). 2 The mean reflects the average quality level of the answers within a group. For example, in a medical question-and-answer task, if the mean of the replaced answer group is 0.8 (out of 1), it indicates that the overall answers are close to professional standards. The variance measures the dispersion of answer quality. For example, a variance of 0.05 indicates that the quality of answers within a group is relatively concentrated, while a variance of 0.2 indicates that there are significant differences.
[0129] Subsequently, an advantage function is generated based on the mean and variance. For example, the advantage function can be represented as {A}. h A2, A3, ... A G}, where A h This is the dominance function corresponding to the replaced standard answer. The specific method for generating the dominance function based on the mean and variance will not be elaborated here.
[0130] For example, the advantage function A i It can be determined based on the following formula:
[0131] A i =r i -μ+α×σ
[0132] Where, r i Let be the reward value for the i-th candidate answer, and α be an adjustment coefficient. For example, in advertising copy generation, if the reward value r of a certain answer is... i =0.9, group mean μ=0.7, standard deviation σ=0.1, α=0.5, then A i =0.9 - 0.7 + 0.5 × 0.1 = 0.25. The value of this dominance function (dominance value) reflects the quality of an individual's answer relative to the average level.
[0133] For example, suppose a legal consultation model generates 5 candidate answers with reward values of [0.6, 0.7, 0.5, 0.8, 0.9], where the lowest-scoring answer (0.5) is replaced with the standard answer (reward value 1.0), and the updated group is [0.6, 0.7, 1.0, 0.8, 0.9]. The calculated mean μ = 0.8, variance σ² = 0.02, and standard deviation σ ≈ 0.14. Advantage function A i =r i Given -0.8 + 0.1 × 0.14, the advantage values for each answer are [-0.186, -0.086, 0.214, 0.014, 0.114]. The model maximizes the probability of answers with high advantage values (e.g., answer A corresponding to a reward value of 1.0). i =0.214), while suppressing low-dominance answers (such as A corresponding to the original answer of 0.6). i =-0.186), train the large model and gradually optimize the generation strategy.
[0134] In this embodiment, since the mean and variance of the reward system are calculated based on multiple candidate answers including the standard answer, the accuracy of the value estimation baseline (mean of reward value) can be improved. By replacing the independent value function network in traditional reinforcement learning with within-group statistics, the number of model parameters and training time are reduced. Because the within-group mean and variance dynamically reflect the overall level and volatility of the model's current generation ability, the dominance function better reflects the actual training state. Especially in the early stages of training when the variance of reward values is large, the standard deviation term in the dominance function can appropriately amplify the optimization signal of high-reward answers, accelerating model convergence.
[0135] In one possible implementation, training the large model based on the advantage function includes: constructing multiple input pairs based on system prompts and multiple candidate answers in the updated output group; inputting each input pair into the large model to obtain original predicted values corresponding to the multiple candidate answers in the updated output group; constructing a policy optimization loss function based on the advantage function and the original predicted values, and updating the parameters of the large model using the policy optimization loss function.
[0136] When training the large model based on the aforementioned advantage function, the advantage function can be combined with the model's original predictions to construct a policy optimization loss function to achieve policy optimization. This guides the model to learn the correlation between generating policies and reward feedback during the training process, thereby improving the accuracy of parameter updates.
[0137] Specifically, multiple input pairs can be constructed based on a system prompt and an updated set of candidate answers. The system prompt is a predefined instruction or contextual information used to guide the model in understanding the task objective. For example, in a medical question-answering task, the system prompt might be "Please generate diagnostic suggestions based on the latest clinical guidelines," and each input pair would consist of this prompt combined with a candidate answer.
[0138] After constructing the input pairs, each pair can be fed into the large model to obtain the corresponding raw predicted values (Logits). The raw predicted values are the model's original output values before probability normalization (such as Softmax), representing the model's confidence in different generated options. For example, when generating the answer "Daily intake of Vitamin C 100mg", the model might output higher Logits values for options like "100mg" and "200mg". Finally, the advantage function (quantifying the relative value of candidate answers) and Logits are used to construct a strategy to optimize the loss function, such as calculating the gradient through weighted summation, driving the large model parameters to update towards higher-advantage answers.
[0139] In this embodiment, by constructing input pairs, system prompts are bound to candidate answers, enabling the model to more clearly understand the optimization direction. For example, in an advertising copywriting task, the combination of the prompt "highlight the product's environmental characteristics" and the corresponding answer can help the model focus on optimizing the expression of specific selling points. Furthermore, the original predicted values directly reflect the model's preference for generated options; combined with the dominance function, the loss function can more accurately identify generation patterns that need to be strengthened or suppressed. For example, when generating legal clauses, the logits gradient corresponding to high-dominance answers is amplified, accelerating the model's alignment with professional terminology.
[0140] In one possible implementation, the step of constructing a policy optimization loss function based on the advantage function and the original predicted values, and updating the parameters of the large model using the policy optimization loss function, includes: using the advantage function as a weighting factor and performing a weighted summation with the probability logarithms corresponding to the original predicted values to construct the policy optimization loss function; calculating the gradient of the policy optimization loss function with respect to the model parameters using a gradient backpropagation algorithm; and updating the parameters of the large model according to the gradient.
[0141] In this implementation, when constructing the loss function by fusing the advantage function with the original predicted values, the advantage function can be used as a dynamic weighting factor, and its value can characterize the relative value of the candidate answers. When constructing the loss function, the advantage function value of each candidate answer can be multiplied by the logarithm of its original predicted value after Softmax transformation, and then a weighted sum can be performed.
[0142] For example, the loss function can be expressed by the following formula:
[0143] L=-∑(A i ×logP i )
[0144] Where P i The probabilities of Logits are normalized using the Softmax function. Through backpropagation, the model will increase the probability of generating high-dominance answers (such as the standard answer) and suppress low-dominance answers.
[0145] High-dominance answers are represented by large negative values in the loss function, which generate positive gradients during backpropagation, pushing the model to increase its generation probability; conversely, low-dominance answers have negative gradients that inhibit the model from producing similar results again.
[0146] During the gradient calculation phase, the derivative of the policy optimization loss function can be calculated using the gradient backpropagation algorithm. Specifically, the partial derivative of the loss function with respect to the original predicted values can be calculated, with the advantage function used as a weight in the calculation. Answers with high advantage have larger gradient components, while answers with low advantage have relatively smaller or even negative gradient components. For example, when the advantage function value of an answer is 0.5, the gradient of its probability logarithmic term will be amplified by 50%; if the advantage function value is -0.2, the gradient direction is reversed.
[0147] After obtaining the gradient, the parameters of the large model can be updated based on the preset learning rate and possible hyperparameters such as momentum and weight decay.
[0148] In this embodiment, the policy optimization loss function is constructed by weighting the advantage function as a weighting factor and summing it with the logarithms of the probabilities corresponding to the original predicted values. The gradient of the policy optimization loss function with respect to the model parameters is calculated using the gradient backpropagation algorithm. The parameters of the large model are then updated based on the gradient. Therefore, weighting by the advantage function highlights key samples, avoids the resource waste of uniformly optimizing all samples, and significantly improves training efficiency.
[0149] Figure 2 An example diagram of the large model post-training method according to an embodiment of this disclosure is shown, such as... Figure 2 As shown, the process first takes the request question q as input and scales it to the largest model, i.e., the policy model, to generate multiple candidate answers O1, O3, ..., O G The output group is then used. Then, the reward value r1, r3, ..., r for each candidate answer is calculated using the reward model. G Meanwhile, the reference model supervises the output of the policy model to ensure that the output is aligned with the output of the reference model.
[0150] Next, the manually annotated standard answer Oh A pseudo-random replacement of a candidate answer in the output group yields an updated output group. For the replaced output group, the reward value r corresponding to the standard answer is calculated using the reward model. h Then, based on all updated reward values r1,r h ,r3,…,r G The advantage function A1,A1 for each candidate answer is obtained through group computation. h A3,…,A G .
[0151] Simultaneously, new input pairs SP+O1 and SP+O are constructed using the system hints (SP) and each candidate answer in the updated output set. h ,SP+O3,…,SP+O G This is then input into the larger policy model to generate the corresponding predicted values: logits1, logits h ,logits3,…,logits G Finally, the obtained advantage function is combined with the corresponding logits, and the policy model is trained and updated using the Group Relative Policy Optimization (GRPO) loss function, thereby optimizing the model.
[0152] In one possible implementation, the intervention intensity of the standard answer on multiple candidate answers decreases progressively with the evolution of the training phase of the large model training, and replacing at least one of the multiple candidate answers with the standard answer includes: dynamically adjusting the intervention intensity of the standard answer according to the training process, wherein the intervention intensity is negatively correlated with the training process.
[0153] Intervention intensity is the degree to which manually labeled standard answers correct the candidate answer set originally generated by the model. It represents the degree to which external prior knowledge regulates the autonomous output of a large model and is a weight parameter that balances the supervisory signal and the model's freedom of exploration.
[0154] Specifically, the intensity of intervention can be reflected in the level of invasiveness when replacing candidate answers with standard answers, including: whether to replace, the number of replacements, and the degree of modification of the replaced content. For example, high intensity means: completely replacing low-reward-value candidate answers with the original standard answer (e.g., replacing an incorrect medical diagnosis with expert-reviewed text); medium intensity means: replacing with a standard answer that incorporates synonym substitutions or sentence structure adjustments (e.g., changing "once daily" to "once a day"); low intensity means: only partially replacing words in high-risk answers (e.g., correcting "absolutely effective" to "clinically shown to be effective"); zero intensity means: completely retaining the original candidate answers. It should be noted that the specific intensity levels described here are merely examples, and this disclosure does not limit the specific intensity levels.
[0155] The training process characterizes the maturity state of a large model during the training cycle, and is used to dynamically evaluate the model's evolution from underfitting to overfitting. There are various criteria for dividing the training process. In one example, it can be divided based on the stability metrics of the large model, such as the standard deviation of the reward value (see the possible implementations provided in this disclosure). In another example, it can be divided based on performance metrics, such as the mean reward or the convergence of the loss function. Furthermore, it can also be divided based on time-series metrics, such as the number of training iterations or the percentage of training cycles.
[0156] There is a negative correlation between intervention intensity and training progress. The earlier the training process, the higher the intervention intensity, and the more likely the replacement operation is to use the unmodified standard answer. Conversely, the later the training process, the lower the intervention intensity, and the more likely the replacement operation is to retain the original candidate answer or introduce a controllable perturbation.
[0157] Therefore, as the training process increases step by step, the intervention intensity will decrease step by step. The decrease in intervention intensity can be a discrete step decrease, such as a three-stage switching decrease, for example, see the possible implementation methods provided in this disclosure; or it can be a continuous gradual decrease, such as the exponential decay of the replacement probability of the standard answer.
[0158] In this embodiment, the intervention intensity of the standard answer on multiple candidate answers gradually decreases as the training phase of the large model progresses. By dynamically adjusting the intervention intensity of the standard answer according to the training process, the intervention intensity is negatively correlated with the training process. This forces high-intensity intervention in the early stages to quickly correct systematic errors, reduce the model's exploration iterations in the wrong direction, and shorten the training time to reach the basic performance threshold. Meanwhile, the intervention intensity decreases in the later stages, which can retain candidate answers generated by the model itself, encourage the exploration of unseen patterns, avoid excessive constraints on the output style by the manual standard answer, and improve the diversity of generated content.
[0159] In one possible implementation, the step of dynamically adjusting the intervention intensity of the standard answer according to the training process includes: in the first stage of training the large model, replacing at least one of the plurality of candidate answers with the standard answer; and / or, in the second stage of training the large model, replacing at least one of the plurality of candidate answers with the standard answer after adding random perturbation; and / or, in the third stage of training the large model, training is performed directly based on the candidate answers originally generated by the large model, wherein the first stage precedes the second stage, and the second stage precedes the third stage.
[0160] In this implementation, the stability of model training and the diversity of generated models are balanced through phased replacement rules. Overall, the intensity of human intervention can be gradually reduced according to the different maturity levels of model training (early, middle, and late stages), achieving a natural transition from strong guidance to autonomous exploration.
[0161] Specifically, in the first stage of training (when the model's capabilities are weak), low-quality candidate answers are directly replaced with standard answers, providing the model with a clear direction for optimization. For example, when the model generates multiple answers about "financial risk control," terminology errors or logical loopholes may frequently occur in the early stages. In this case, replacing low-scoring answers with standard answers reviewed by experts helps the model quickly establish basic generation capabilities.
[0162] In the second training phase, random perturbations are added to the standard answer. For example, in text processing tasks, synonyms in the standard answer are replaced, sentence structure is adjusted, or in image object recognition tasks, the coordinates or color parameters of the standard answer are fine-tuned. This perturbation preserves the core semantics of the standard answer while introducing controllable diversity, preventing the model from over-relying on fixed patterns. For example, in advertising copy generation, changing the standard answer "limited-time discount 50%" to "today's special offer 50% off" or "half-price carnival for one day only" maintains the accuracy of the promotional information while expanding the model's learning of diverse expressions. For coordinate detection and object detection tasks in images, manually labeled bounding boxes or coordinates can be jittered, randomly adjusted to coordinates with an Intersection over Union (IoU) of 0.5 with the correct manually labeled boxes.
[0163] In the third stage (when the model tends to stabilize), the original generated candidate answers are completely retained. For example, in creative writing tasks, the model can generate reasonable plot developments. At this time, the replacement is stopped to encourage independent innovation and avoid excessive restrictions on the generation style by manual annotation.
[0164] In this embodiment, the high-intensity replacement in the early stages of training provides the model with a stable learning baseline, especially in scenarios with scarce data or high task complexity (such as legal document generation), enabling rapid correction of systematic errors. Secondly, the perturbation design in the middle stages of training, by introducing appropriate noise, enhances the model's understanding of semantic invariance. For example, in translation tasks, the perturbed standard answer helps the model learn the equivalence of different linguistic expressions, improving generalization ability. Finally, the autonomous exploration phase in the later stages of training allows the model to explore potential optimization space on the established stability foundation. For example, in dialogue systems, the model may generate answers that better meet the user's personalized needs, rather than mechanically copying the standard answer. This progressive strategy not only reduces the cost of full-process manual intervention but also balances training efficiency and generation quality through stage-adaptive replacement rules.
[0165] In one possible implementation, the method further includes: calculating the initial standard deviation of the reward value std0; if the standard deviation of the reward value of the current batch data is in the interval [1 / a×std0, std0], determining that the current training is in the early stage of training; where 1 / a < 1; if the standard deviation of the reward value of the current batch data is in the interval [1 / b×std0, 1 / a×std0], determining that the current training is in the middle stage of training; where 1 / b < 1 / a; if the standard deviation of the reward value of the current batch data is less than or equal to 1 / b×std0, determining that the current training is in the late stage of training.
[0166] When dividing the training phases, statistical indicators (standard deviation) can be used to objectively measure the training progress, replacing the traditional subjective judgment that relies on fixed iterations or human experience, thereby achieving more intelligent training phase management.
[0167] The standard deviation (std0) of the reward values for the initial batch of data can be calculated at the beginning of training and used as a benchmark for subsequent stage divisions. For example, when the model starts training, the first 100 batches of candidate answers are selected, the score for each answer is calculated using the reward model, and the standard deviation of these scores is then used as std0. This value reflects the degree of fluctuation in the model's initial generation ability.
[0168] Then, during the subsequent training process, the standard deviation (std) of the reward value of the current batch of data can be calculated in real time and compared with std0: if std is in the interval [1 / a×std0,std0], it is determined to be the early stage of training (e.g., if a=2, the interval is [0.5std0,std0]).
[0169] If std falls within the interval [1 / b×std0, 1 / a×std0], it is considered to be in the middle of training. For example, if b = 3, the interval is [0.33std0, 0.5std0]. If std ≤ 1 / b×std0, it enters the later stage of training, such as ≤ 0.33std0. Parameters a and b must satisfy 1 / b < 1 / a (e.g., a = 2, b = 3) to ensure the progressive nature of the interval division.
[0170] For example, suppose the initial standard deviation std0 of a legal consultation model is 1.0 (the quality of the initially generated answers varies greatly):
[0171] Initially: When a batch has std=0.8 (in the range of [0.5, 1.0]), some answers generated by the model may omit key legal clauses and need to be replaced with fully annotated answers;
[0172] Mid-term: When std = 0.4 (in the range of [0.33, 0.5]), the quality of answers tends to be stable, but there is still a problem of monotonous sentence structure. At this time, the phrase "the party concerned should provide a written application" in the standard answer is perturbed to "written materials need to be submitted" to increase diversity;
[0173] Later stage: When std = 0.2 (≤ 0.33), the model can generate compliant and diverse answers, and the original answers are directly retained for training.
[0174] In this embodiment, the standard deviation dynamically reflects the model's convergence state, avoiding undertraining or overfitting caused by fixed-period division. Traditional methods rely on manual observation of the loss curve or sample generation for judgment, while the standard deviation provides a quantifiable objective indicator, particularly suitable for automated training systems. Furthermore, parameters a and b can be adjusted according to task requirements—for high-security tasks (such as medical question answering), a smaller 1 / b value (e.g., 0.2) can be set to prolong the initial intervention; for creative tasks (such as poetry generation), the threshold can be relaxed to encourage early diversity. Moreover, low-quality answers are replaced during the high-fluctuation phase in the early training stage to reduce ineffective training iterations; in the later low-fluctuation phase, computational overhead is reduced, allowing focus on model fine-tuning.
[0175] Figure 3 An example diagram of the large model post-training method according to an embodiment of this disclosure is shown, such as... Figure 3 As shown in part (a), after inputting the request question, the policy model generates multiple candidate answers O1, O2, ..., O G The reward model calculates the reward value r1, r2, ..., r for each candidate answer. G The mean and standard deviation of the reward values are calculated. The reward values are sorted from smallest to largest to determine the positions of the candidate answers to be replaced. Then, a human-made standard answer O is used. hReplace one of the original candidate answers to obtain the updated output group. The corresponding reward value is r. i The reward value r for being replaced by the standard answer h dominance function A i It was also replaced with A h .
[0176] like Figure 3 Section (b) illustrates the replacement strategies for different training phases:
[0177] Initial training phase: Replace with standard answer O h To ensure stable training.
[0178] Mid-training: Replace with the standard answer Oh-jitter, which incorporates jitter, to enhance the model's exploratory capabilities.
[0179] Later in the training phase: No more replacements are made; the model is allowed to explore autonomously while maintaining the original output answer.
[0180] The entire process integrates supervised learning and reinforcement learning ideas within the same training phase. By dynamically adjusting the replacement strategy, the model is strongly guided by human annotations in the early stage of training, balances stability and exploration in the middle stage, and fully leverages the model's own exploration capabilities in the later stage, thereby improving the model's performance and training efficiency.
[0181] In one possible implementation, the method further includes: calculating the error rate of a large model; if the error rate is higher than a preset error rate threshold, determining that the current training is in a first stage; if the error rate is in a first preset interval, determining that the current training is in a second stage; if the error rate is lower than a target threshold, determining that the current training is in a third stage, wherein the preset error rate threshold is higher than the target threshold, and the first preset interval is between the preset error rate threshold and the target threshold.
[0182] To more accurately determine the current training stage of the model, the error rate of the large model can be calculated in real time, and the training process can be dynamically divided accordingly.
[0183] Specifically, the error rate is the proportion of incorrect answers generated by the model in the current batch of data, reflecting the model's immediate performance on a specific task. Two key thresholds can be pre-designed: a higher "preset error rate threshold" and a lower "target threshold".
[0184] When the model's error rate exceeds the preset error rate threshold, it indicates that the model is still in the early stages of training. At this point, its generation ability is weak, and the output quality fluctuates greatly, classifying it as the first stage. In this stage, a high-intensity intervention training strategy can be adopted, directly replacing low-quality candidate answers with standard answers to quickly correct the model's systematic biases and prevent it from continuously exploring in the wrong direction.
[0185] As training progresses, the performance of the large model gradually improves, and the error rate begins to decrease. When the error rate falls into the "first preset interval"—that is, between the preset error rate threshold and the target threshold—training can be considered to have entered the second stage. At this stage, the model has acquired certain basic capabilities, but still suffers from local instability or limited representation. Therefore, the training strategy appropriately reduces the intensity of intervention; for example, by adding slight perturbations to the standard answer before replacement. This preserves the core semantics of the reference answer while introducing controllable diversity, helping the large model improve its generalization ability while learning stably.
[0186] When the error rate of the large model further decreases below the target threshold, training can be considered to have entered the third stage. At this point, the model has matured and possesses high accuracy and stability. The training strategy can be adjusted to minimal or zero intervention, meaning that the model's original candidate answers are completely retained during training. The focus of this stage is to unleash the large model's autonomous exploration capabilities, encouraging it to generate more creative and diverse content based on its established knowledge, and avoiding excessive restrictions on the model's output style imposed by human-provided standard answers.
[0187] In this embodiment, a dynamic stage division mechanism based on error rate allows for judging the training progress based on the model's actual performance rather than a fixed number of iterations or human experience, thereby more accurately matching the training strategy with the model state. This not only improves training efficiency and reduces unnecessary computational overhead, but also quickly corrects model biases in the early stages of training, enhances its generalization ability in the middle stages, and fully releases its generative potential in the later stages, ultimately achieving an optimal balance between model performance and training resources.
[0188] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0189] In addition, this disclosure also provides a large model post-training device, electronic device, computer-readable storage medium, and program, all of which can be used to implement any of the large model post-training methods provided in this disclosure. For the corresponding technical solutions and descriptions, please refer to the relevant records in the method section, which will not be repeated here.
[0190] Figure 4 A block diagram of a large model post-training apparatus according to an embodiment of the present disclosure is shown, such as Figure 4 As shown, the large model post-training device 20 includes:
[0191] Module 21 is used to obtain the output group containing multiple candidate answers generated by the large model for the same request question;
[0192] Replacement module 22 is used to replace at least one of the multiple candidate answers with the standard answer to obtain an updated output group;
[0193] Training module 23 is used to train the large model based on the updated output group.
[0194] In one possible implementation, the replacement module is used to:
[0195] The reward value is calculated for each of the multiple candidate answers using a reward model.
[0196] Determine at least one candidate answer with the lowest reward value among the plurality of candidate answers;
[0197] Replace at least one candidate answer with the lowest reward value with the standard answer.
[0198] In one possible implementation, the intervention strength of the standard answer on multiple candidate answers gradually decreases as the training phase of the large model evolves, and the replacement module is used to:
[0199] The intervention intensity of the standard answer is dynamically adjusted according to the training progress, wherein the intervention intensity is negatively correlated with the training progress.
[0200] In one possible implementation, the replacement module is used to:
[0201] In the first phase of training the large model, at least one of the multiple candidate answers is replaced with the standard answer; and / or,
[0202] In the second stage of training the large model, at least one of the multiple candidate answers is replaced with a standard answer that has been subjected to random perturbation; and / or,
[0203] In the third stage of training the large model, training is performed based on the candidate answers originally generated by the large model, wherein the first stage precedes the second stage, and the second stage precedes the third stage.
[0204] In one possible implementation, the apparatus further includes a first-stage partitioning module, used for:
[0205] Calculate the initial standard deviation of the reward value, std0;
[0206] If the standard deviation of the reward value of the current batch of data is in the interval [1 / a×std0, std0], then the current training is determined to be in the first stage; where 1 / a < 1.
[0207] If the standard deviation of the reward value of the current batch of data is in the interval [1 / b×std0, 1 / a×std0], then the current training is determined to be in the second stage; where 1 / b < 1 / a.
[0208] If the standard deviation of the reward value of the current batch of data is less than or equal to 1 / b×std0, the current training is determined to be in the third stage.
[0209] In one possible implementation, the device further includes a second-stage partitioning module, used for:
[0210] Calculate the error rate of a large model;
[0211] If the error rate of the answer is higher than the preset error rate threshold, the current training is determined to be in the first stage.
[0212] If the error rate is within the first preset range, the current training is determined to be in the second stage.
[0213] If the error rate is lower than the target threshold, the current training is determined to be in the third stage, wherein the preset error rate threshold is higher than the target threshold, and the first preset interval is between the preset error rate threshold and the target threshold.
[0214] In one possible implementation, the training module is used for:
[0215] Calculate the mean and variance of the reward values for candidate answers in the updated output group;
[0216] A dominance function is generated based on the mean and variance.
[0217] The large model is trained based on the aforementioned advantage function.
[0218] In one possible implementation, the training module is used for:
[0219] Construct multiple input pairs based on system prompts and multiple candidate answers in the updated output group;
[0220] Each input pair is input into the large model to obtain the original predicted values corresponding to multiple candidate answers in the updated output group;
[0221] A policy optimization loss function is constructed based on the advantage function and the original predicted values, and the parameters of the large model are updated using the policy optimization loss function.
[0222] In one possible implementation, the training module is used for:
[0223] The advantage function is used as a weighting factor, and the probability logarithm corresponding to the original predicted value is weighted and summed to construct the policy optimization loss function.
[0224] The gradient of the policy optimization loss function with respect to the model parameters is calculated using the gradient backpropagation algorithm.
[0225] The parameters of the large model are updated based on the gradient.
[0226] In one possible implementation, the replacement module is used to:
[0227] Using manually annotated standard answers, randomly replace one or more of the multiple candidate answers.
[0228] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0229] This disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0230] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0231] This disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above method.
[0232] Figure 5 A block diagram of an apparatus for post-training a large model according to an embodiment of the present disclosure is shown. For example, apparatus 1900 may be provided as a server or terminal device. (Refer to...) Figure 5 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0233] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0234] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.
[0235] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0236] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.
[0237] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information of computer-readable program instructions. These electronic circuits can execute computer-readable program instructions to implement various aspects of this disclosure.
[0238] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0239] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0240] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0241] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0242] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for post-training a large model, characterized in that, include: Obtain the output set containing multiple candidate answers generated by a large model for the same request question; Replace at least one of the multiple candidate answers with the standard answer to obtain the updated output group; The large model is trained based on the updated output set.
2. The method according to claim 1, characterized in that, The step of replacing at least one of the plurality of candidate answers with the standard answer includes: The reward value is calculated for each of the multiple candidate answers using a reward model. Determine at least one candidate answer with the lowest reward value among the plurality of candidate answers; Replace at least one candidate answer with the lowest reward value with the standard answer.
3. The method according to claim 1, characterized in that, The intervention strength of the standard answer on multiple candidate answers gradually decreases as the training phase of the large model evolves. Replacing at least one of the multiple candidate answers with the standard answer includes: The intervention intensity of the standard answer is dynamically adjusted according to the training progress, wherein the intervention intensity is negatively correlated with the training progress.
4. The method according to claim 3, characterized in that, The intervention intensity of dynamically adjusting the standard answer according to the training process includes: In the first phase of training the large model, at least one of the multiple candidate answers is replaced with the standard answer; and / or, In the second stage of training the large model, at least one of the multiple candidate answers is replaced with a standard answer that has been subjected to random perturbation; and / or, In the third stage of training the large model, training is performed based on the candidate answers originally generated by the large model, wherein the first stage precedes the second stage, and the second stage precedes the third stage.
5. The method according to claim 4, characterized in that, The method further includes: Calculate the initial standard deviation of the reward value, std0; If the standard deviation of the reward value of the current batch of data is in the interval [1 / a×std0, std0], then the current training is determined to be in the first stage; where 1 / a < 1. If the standard deviation of the reward value of the current batch of data is in the interval [1 / b×std0, 1 / a×std0], then the current training is determined to be in the second stage; where 1 / b < 1 / a. If the standard deviation of the reward value of the current batch of data is less than or equal to 1 / b×std0, the current training is determined to be in the third stage.
6. The method according to claim 4, characterized in that, The method further includes: Calculate the error rate of a large model; If the error rate of the answer is higher than the preset error rate threshold, the current training is determined to be in the first stage. If the error rate is within the first preset range, the current training is determined to be in the second stage. If the error rate is lower than the target threshold, the current training is determined to be in the third stage, wherein the preset error rate threshold is higher than the target threshold, and the first preset interval is between the preset error rate threshold and the target threshold.
7. The method according to claim 1, characterized in that, The training of the large model based on the updated output group includes: Calculate the mean and variance of the reward values for candidate answers in the updated output group; A dominance function is generated based on the mean and variance. The large model is trained based on the aforementioned advantage function.
8. The method according to claim 7, characterized in that, The training of the large model based on the advantage function includes: Construct multiple input pairs based on system prompts and multiple candidate answers in the updated output group; Each input pair is input into the large model to obtain the original predicted values corresponding to multiple candidate answers in the updated output group; A policy optimization loss function is constructed based on the advantage function and the original predicted values, and the parameters of the large model are updated using the policy optimization loss function.
9. The method according to claim 8, characterized in that, The step of constructing a policy optimization loss function based on the advantage function and the original predicted values, and using the policy optimization loss function to update the parameters of the large model, includes: The advantage function is used as a weighting factor, and the probability logarithm corresponding to the original predicted value is weighted and summed to construct the policy optimization loss function. The gradient of the policy optimization loss function with respect to the model parameters is calculated using the gradient backpropagation algorithm. The parameters of the large model are updated based on the gradient.
10. The method according to claim 1, characterized in that, The step of replacing at least one of the plurality of candidate answers with the standard answer includes: Using manually annotated standard answers, randomly replace one or more of the multiple candidate answers.
11. A large-model post-training device, characterized in that, include: The acquisition module is used to acquire the output set containing multiple candidate answers generated by the large model for the same request question; The replacement module is used to replace at least one of the multiple candidate answers with the standard answer to obtain an updated output group; The training module is used to train the large model based on the updated output set.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.
13. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method and device for model training, equipment and storage medium
CN117689003A
Answer feedback method and device applied to large language model
CN117708292A
Question and answer model training method and device, electronic equipment, storage medium and product
CN118153659A
Large model scene question and answer optimization method and system based on preference learning
CN119336960A
Answer feedback method and apparatus applied to large language model
US20250005053A1
Cited By
Decision model training method, live broadcast decision determination method and device
CN121585837A
Computing device, method, computer readable storage medium and computer program product for optimizing choice questions
CN121745318A