Reinforcement learning training method and device for large model
Patent Information
- Application Number
- CN202610693492.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-11
AI Technical Summary
然而,并行测试阶段扩展在实际使用中,训练流程复杂、计算资源消耗大,难以兼顾生成质量与验证可靠性
Smart Images

Figure CN122735818A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to one or more embodiments in the field of artificial intelligence, and more particularly to reinforcement learning training methods and apparatus for large models. Background Technology
[0002] With the continuous development of artificial intelligence technology, large models have been widely applied in various scenarios such as natural language processing, image understanding, video analysis, visual reasoning, human-computer interaction, and intelligent control. In practical applications, it is not only necessary for large models to generate accurate reasoning results or decision outputs, but also to verify and evaluate the reliability of the output results to avoid the direct use of erroneous output results in downstream tasks, thereby improving the overall security and robustness of the tasks.
[0003] In existing technologies, increasing computational load during the model inference phase—a technique known as test-time scaling—can improve model performance and inference reliability to some extent. One type of test-time scaling is parallel test-time scaling, where the model generates multiple candidate results in parallel during the inference phase and selects the best result through sorting, filtering, or validation mechanisms. However, in practical applications, parallel test-time scaling suffers from complex training processes, high computational resource consumption, and difficulty in balancing generation quality and validation reliability.
[0004] Therefore, an improved approach is needed that simplifies the training process, reduces computational resource consumption, and balances generation quality and verification reliability through reinforcement learning training of large models. Summary of the Invention
[0005] This specification describes one or more embodiments of a reinforcement learning training method and apparatus for large models, which can simplify the training process, reduce the consumption of computing resources, and at the same time take into account the quality of generation and the reliability of verification.
[0006] Firstly, a reinforcement learning training method for large models is provided, including:
[0007] Receive query information, and use the large model to generate a first number of candidate answers and a first number of corresponding verification scores for evaluating the correctness of the candidate answers; the first number is an integer greater than 1;
[0008] Based on whether each candidate answer matches the answer tag, determine the accuracy of each candidate answer and the corresponding generated reward.
[0009] The verification reward is determined based on whether each verification score is consistent with the accuracy of the candidate answer.
[0010] Based on the generation reward corresponding to each candidate answer, determine the first advantage of the large model in answer generation; based on the verification reward corresponding to each verification score, determine the second advantage of the large model in verification ability.
[0011] The parameters of the large model are adjusted with the goal of maximizing the sum of the first and second advantages.
[0012] In one possible implementation, the answer labels include discrete labels, and the step of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer label includes:
[0013] When any candidate answer is the same as a discrete label, the accuracy of the candidate answer is determined to be correct, and the corresponding generated reward is determined to be the first reward score;
[0014] When any candidate answer differs from the discrete label, the accuracy of that candidate answer is determined to be incorrect, and the corresponding generated reward is determined to be the second reward score.
[0015] Furthermore, the discrete label is the instruction name of the first operation instruction.
[0016] In one possible implementation, the answer tags include continuous tags, and the step of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer tag includes:
[0017] Calculate the similarity index between any candidate answer and continuous labels;
[0018] When the similarity index falls within the first preset range, the accuracy of the candidate answer is determined to be correct, and the corresponding generated reward is determined to be the third reward score.
[0019] When the similarity index does not belong to the first preset interval, the accuracy of the candidate answer is determined to be incorrect, and the corresponding generated reward is determined to be the fourth reward score.
[0020] In one possible implementation, the answer tags include continuous tags, and the step of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer tag includes:
[0021] Calculate the similarity index between any candidate answer and the continuous label, and use the similarity index as the accuracy of the candidate answer and the reward score for generating the reward.
[0022] Furthermore, the continuous label is the coordinate of the graphic bounding box, and the similarity index is an indicator used to measure the degree of overlap between two bounding boxes.
[0023] In one possible implementation, determining the verification reward corresponding to each verification score based on whether each verification score is consistent with the accuracy of the candidate answer includes:
[0024] Based on the accuracy of each candidate answer, the first number of candidate answers are divided into a correct answer group and an incorrect answer group;
[0025] Calculate the first average of the verification scores for each candidate answer in the correct answer group, and the second average of the verification scores for each candidate answer in the incorrect answer group;
[0026] For any candidate answer in the correct answer group, the verification reward corresponding to the verification score of the candidate answer is determined based on whether the verification score of the candidate answer is greater than the second average.
[0027] For any candidate answer in the incorrect answer group, the verification reward corresponding to the verification score of the candidate answer is determined based on whether the verification score of the candidate answer is less than the first average value.
[0028] In one possible implementation, determining the verification reward corresponding to each verification score based on whether each verification score is consistent with the accuracy of the candidate answer includes:
[0029] For any candidate answer, select other candidate answers as the comparison set for the candidate answer based on the difference between the accuracy of the candidate answer and the accuracy of other candidate answers.
[0030] Based on the ranking relationship between the verification score of the candidate answer and the verification scores of candidate answers in the comparison set, and whether it conforms to the expected ranking relationship corresponding to the difference, the verification reward corresponding to the candidate answer is determined.
[0031] In one possible implementation, determining the first advantage of the large model regarding answer generation based on the generation rewards corresponding to each candidate answer includes:
[0032] The arithmetic mean of the generation rewards for multiple candidate answers is calculated as the generation baseline;
[0033] Subtract the generation baseline from the generation reward for each candidate answer to obtain the generation advantage value corresponding to that candidate answer;
[0034] The first advantage is calculated based on the generation advantage value corresponding to each candidate answer.
[0035] In one possible implementation, determining the second advantage of the large model regarding validation capability based on the validation rewards corresponding to each validation score includes:
[0036] Calculate the arithmetic mean of the validation rewards for multiple validation scores as the validation baseline;
[0037] Subtract the verification baseline from the verification reward for each verification score to obtain the verification advantage value corresponding to that verification score;
[0038] The second advantage is calculated based on the validation advantage value corresponding to each validation score.
[0039] Further, based on the generation advantage value corresponding to each candidate answer, the first advantage is calculated, including:
[0040] By using a first mask to block out the word positions used to output the validation score, and only retaining the word positions used to generate candidate answers for calculation, the first advantage is obtained.
[0041] Further, based on the validation advantage value corresponding to each validation score, the second advantage is calculated, including:
[0042] By using a second mask to block out the word positions used to generate candidate answers, and only retaining the word positions used to output the validation score for calculation, a second advantage is obtained.
[0043] In one possible implementation, the large model is a visual-language model, and the query information includes image data or video stream data.
[0044] Secondly, a reinforcement learning training device for large models is provided, including:
[0045] The generation unit is used to receive query information and generate a first number of candidate answers and corresponding first number of verification scores for evaluating the correctness of the candidate answers using the large model; the first number is an integer greater than 1.
[0046] The first reward unit is used to determine the accuracy of each candidate answer and the corresponding generation reward based on whether each candidate answer obtained by the generation unit is consistent with the answer label.
[0047] The second reward unit is used to determine the verification reward corresponding to each verification score based on whether the accuracy of each verification score obtained by the generation unit is consistent with that of the candidate answer obtained by the first reward unit.
[0048] The advantage determination unit is used to determine the first advantage of the large model in answer generation based on the generation rewards corresponding to each candidate answer obtained by the first reward unit; and to determine the second advantage of the large model in verification capability based on the verification rewards corresponding to each verification score obtained by the second reward unit.
[0049] An adjustment unit is used to adjust the parameters of the large model with the goal of maximizing the sum of the first advantage and the second advantage obtained by the advantage determination unit.
[0050] Thirdly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of the first aspect.
[0051] Fourthly, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the method of the first aspect.
[0052] The methods and apparatus provided in the embodiments of this specification effectively solve the bottleneck problems of high memory consumption and large inference latency caused by maintaining independent value networks or verification modules in traditional dual-model schemes by constructing dual-path advantage signals for generating rewards and verifying rewards, and realizing independent calculation and collaborative optimization of the two under a single-model architecture. Simultaneously, the advantage decoupling mechanism avoids gradient conflicts in multi-task training, preventing the model from sacrificing the performance of another dimension while pursuing a single objective. This allows the model to simultaneously improve the accuracy of answer generation and the calibration of self-evaluation without the burden of additional parameters. In summary, this scheme simplifies the training process, reduces computational resource consumption, and balances generation quality and verification reliability. Attached Figure Description
[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a schematic diagram illustrating an implementation scenario of one embodiment disclosed in this specification;
[0055] Figure 2 A flowchart illustrating a reinforcement learning training method for a large model according to one embodiment is shown.
[0056] Figure 3 A schematic diagram illustrating the grouping of candidate answers according to one embodiment is shown;
[0057] Figure 4 A schematic diagram illustrating a method for determining verification rewards according to one embodiment is shown;
[0058] Figure 5 A schematic diagram of a masking mechanism according to one embodiment is shown;
[0059] Figure 6 A schematic diagram of joint optimization according to one embodiment is shown;
[0060] Figure 7 A schematic block diagram of a reinforcement learning training apparatus for a large model according to one embodiment is shown. Detailed Implementation
[0061] The solution provided in this specification will now be described with reference to the accompanying drawings.
[0062] Figure 1 This is a schematic diagram illustrating an implementation scenario of one embodiment disclosed in this specification. This implementation scenario involves reinforcement learning training of a large model. It can be understood that reinforcement learning is a machine learning paradigm where a large model learns a better policy by interacting with the environment to obtain reward signals. The large model can be referred to as the policy model in reinforcement learning. (Refer to...) Figure 1 After receiving the query information, the large model not only needs to generate the answer but also output a self-validation score. Unlike typical processing methods, it doesn't require two separate models for answer generation and validation score output; instead, a single model handles both. The self-validation score is output by the large model itself. After generating the answer, the large model automatically evaluates the accuracy of its generated answer and outputs the validation score.
[0063] The embodiments in this specification are based on the testing phase expansion to improve the performance of large models. Specifically, G candidate answers {O1, O2, ..., O...} are generated in parallel. G} and the corresponding G validation scores {S1,S2,...,S} G}, where G>1, S i The answer is O. i The validation score is given by the model, where i ranges from 1 to G. The large model is configured to have dual output capabilities. When it receives a user's query, instead of simply outputting a single final answer, it executes the following parallel processing flow: First, based on the current input context, the model generates G distinct candidate answers {O1, O2, ..., O...} using a sampling strategy. G These answers represent different reasoning paths or expressions of the model for the same question. Following each generated candidate answer O... i The model outputs a corresponding self-validation score S. iThe score can fall within a preset range, such as [0,1] or [0,100]. This score is used to quantify the model's own assessment of the candidate answer O. i Confidence level of accuracy or quality.
[0064] For example, a user poses a task, and the agent completes it by calling a tool. The agent first uses a large model to reason and derive the answer to the task, then provides the user with the tool to use based on that answer. The large model not only generates the answer but also assigns a validation score to each answer, evaluating its correctness and applicability to the task. If the query for the task is "Open Settings," the large model might generate the following answers: "Click the settings icon," with a validation score of 0.9; "Enter the system menu," with a validation score of 0.8; and "Restart your phone," with a validation score of 0.2. The validation score indicates the accuracy of the answer, with the first two answers being correct and the last one incorrect.
[0065] This approach utilizes a large model to generate multiple candidate outputs through batch sampling or batch decoding. Each candidate output includes a candidate answer and its corresponding validation score. All candidate answers and their corresponding validation scores are generated by the same neural network with shared weights, eliminating the need for dual-model collaboration, thus reducing computational resource consumption and improving deployment efficiency.
[0066] To improve the performance of large models, training is required. Although training techniques for large models are relatively mature, significant technical bottlenecks still exist when handling the dual tasks of "generation + verification". This specification provides a novel reinforcement learning training scheme for a single model architecture, aiming to simplify the training process, reduce computational resource consumption, and simultaneously ensure both generation quality and verification reliability.
[0067] Figure 2 This diagram illustrates a reinforcement learning training method for a large model according to one embodiment, which can be based on... Figure 1 The implementation scenario is shown. For example... Figure 2As shown, the reinforcement learning training method for the large model in this embodiment includes the following steps: Step 21, receiving query information and generating a first number of candidate answers and corresponding first number of verification scores for evaluating the correctness of the candidate answers using the large model; the first number is an integer greater than 1; Step 22, determining the accuracy and generation reward corresponding to each candidate answer based on whether each candidate answer is consistent with the answer label; Step 23, determining the verification reward corresponding to each verification score based on whether each verification score is consistent with the accuracy of the candidate answer; Step 24, determining the first advantage of the large model in answer generation based on the generation reward corresponding to each candidate answer; determining the second advantage of the large model in verification ability based on the verification reward corresponding to each verification score; Step 25, adjusting the parameters of the large model with the goal of maximizing the sum of the first and second advantages. The specific execution method of each of the above steps is described below.
[0068] First, in step 21, the query information is received, and a first number of candidate answers and corresponding first number of verification scores for evaluating the correctness of the candidate answers are generated using a large model; the first number is an integer greater than 1. It can be understood that the large model can be a language model, which generates tokens one by one. Each candidate answer may include several tokens, and each verification score may also include several tokens.
[0069] In the embodiments described in this specification, the query information may include text-based information, and may further include information from other modalities such as images. The large model may be a simple language model or a multimodal model. The model evaluates the accuracy of candidate answers by outputting a validation score. Accuracy reflects the correctness of the answer; correct answers receive higher scores, and incorrect answers receive lower scores.
[0070] In one example, the large model is a visual-language model, and the query information includes image data or video stream data.
[0071] In this example, the vision-language model is a multimodal model capable of simultaneously understanding images (i.e., visual information) and text (i.e., linguistic information), aiming to establish a correspondence between visual content and linguistic semantics. By combining visual perception with text generation, the large model no longer simply performs image recognition or text generation, but possesses deep visual understanding, precise action planning, and self-reflection capabilities, providing a feasible technical path for low-cost, high-efficiency intelligent applications.
[0072] The visual-language model described in this specification has a wide range of applications, such as in human-computer interaction scenarios of multimodal intelligent assistants and intelligent terminals; in automated operation and intelligent control scenarios, including but not limited to mobile device control, graphical user interface operation and robot decision-making; in intelligent decision support systems that require complex visual understanding and reasoning; and in multimodal artificial intelligence application systems that have high requirements for the reliability, stability and robustness of reasoning results.
[0073] For example, the aforementioned visual-language model can be used in a mobile automated operations and maintenance assistant to perform corresponding user interface (UI) operations based on user instructions. The user says, "Turn up the phone volume," and a screenshot of the phone screen is used as input. During the generation phase, the model generates multiple candidate answers, each corresponding to a candidate operation, including the operation's location coordinates and name, such as [{x:80, y:50}, "Click the volume button"] or [{x:10, y:10},"Click the power button"]. During the validation phase, the model outputs a validation score for each operation. For a correct operation, the operation name is "Click the volume button," and the model should output a high score, such as 0.9. For an incorrect operation, the operation name is "Click the power button," and the model should output a low score, such as 0.1.
[0074] For example, the aforementioned vision-language model can also be used in industrial quality inspection robots for defect detection and localization. Inputting product images from a factory assembly line conveyor belt, the robot is required to identify and mark the location of defects. During the generation phase, the model generates multiple candidate answers, each corresponding to a defect, including possible bounding boxes and type descriptions (such as "scratches" or "cracks"). During the validation phase, the model assigns a confidence score to each defect.
[0075] For example, the aforementioned vision-language model can also be used in in-vehicle autonomous driving assistance systems for video stream understanding. While the vehicle is in motion, cameras capture real-time video streams of the road ahead, requiring a decision on the next driving strategy, such as whether emergency braking is necessary. During the generation phase, the model generates multiple candidate answers based on the current frame, representing a series of driving strategies, such as [maintain speed], [decelerate], and [change lanes]. During the validation phase, the model predicts the safety score for each driving strategy within the next few seconds.
[0076] Understandably, there are many possible application scenarios for visual-language models, which will not be listed here.
[0077] Then, in step 22, based on whether each candidate answer matches the answer label, the accuracy of each candidate answer and the corresponding reward are determined. It is understandable that, due to differences in task types, the logic for determining the consistency of answer labels and their matching in the embodiments of this specification varies, specifically in the difference between discrete and continuous tasks. This difference is mainly reflected in the fact that discrete tasks require defining a "control sample," while continuous tasks require defining an "expected direction." Discrete tasks have discrete labels, and continuous tasks have continuous labels.
[0078] In one example, the answer labels include discrete labels. The process of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer label includes:
[0079] When any candidate answer is the same as a discrete label, the accuracy of the candidate answer is determined to be correct, and the corresponding generated reward is determined to be the first reward score;
[0080] When any candidate answer differs from the discrete label, the accuracy of that candidate answer is determined to be incorrect, and the corresponding generated reward is determined to be the second reward score.
[0081] In this example, the first reward score is greater than the second reward score, which can be 0. By introducing a strict matching mechanism for discrete labels, complex task processing is transformed into a clear binary decision ("right" or "wrong"), thus providing reinforcement learning with a clear, unambiguous, and deterministic reward signal. This rule-based matching judgment method not only significantly reduces training noise caused by fuzzy evaluations, ensuring that the model can accurately converge to the correct answer in logically demanding tasks such as mathematical reasoning and intelligent navigation, but also constructs a strong positive and negative feedback loop through differentiated first and second reward scores. This effectively guides the model to quickly distinguish and optimize the correct generation path and discard incorrect strategies, thereby significantly improving training efficiency and the accuracy of the final answer in discrete task scenarios.
[0082] Furthermore, the discrete label is the instruction name of the first operation instruction.
[0083] In this example, by concretizing discrete labels as "the name of the first operation instruction," the abstract reinforcement learning reward signal is anchored in a specific human-computer interaction semantic space, enabling the model to directly learn standardized action execution logic. This design not only eliminates the ambiguity caused by the diversity of natural language expressions (e.g., mapping "open settings" and "enter settings interface" to the same instruction name), ensuring a high degree of consistency between the generated content and the real task objective, but also significantly improves the agent's instruction compliance ability and generalization performance in scenarios such as automated operation and maintenance and user interface navigation. This allows the model to accurately output standardized operation instructions that conform to the specifications when faced with user queries with different expressions but the same intent, thereby ensuring the reliability and robustness of complex task execution.
[0084] In one example, the answer labels include continuous labels. The process of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer label includes:
[0085] Calculate the similarity index between any candidate answer and continuous labels;
[0086] When the similarity index falls within the first preset range, the accuracy of the candidate answer is determined to be correct, and the corresponding generated reward is determined to be the third reward score.
[0087] When the similarity index does not belong to the first preset interval, the accuracy of the candidate answer is determined to be incorrect, and the corresponding generated reward is determined to be the fourth reward score.
[0088] In this example, by introducing a similarity index and a first preset interval judgment mechanism, the limitations of traditional binary judgment in continuous tasks are overcome, transforming accuracy into a relative evaluation based on a quality threshold. This design effectively adapts to scenarios with continuous accuracy requirements, such as visual localization (Bounding Box), allowing the model to still receive positive feedback when there is a certain deviation between the predicted result and the true label, thus avoiding training rigidity caused by excessive pursuit of precise matching. At the same time, by setting clear threshold boundaries, it ensures that only candidate answers that meet practical quality standards are considered "correct," preserving the model's ability to explore diverse solution spaces while preventing low-quality outputs from interfering with strategy optimization. This significantly improves the robustness, fault tolerance, and practical value of the final output generated by the model in continuous tasks.
[0089] In one example, the answer labels include continuous labels. The process of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer label includes:
[0090] Calculate the similarity index between any candidate answer and the continuous label, and use the similarity index as the accuracy of the candidate answer and the reward score for generating the reward.
[0091] In this example, by introducing a direct mapping mechanism for similarity metrics, the technical limitation of traditional schemes that require forced binarization of continuous tasks ("right / wrong" judgment) is overcome, effectively preserving fine-grained quality information in quantitative metrics such as Intersection over Union (IoU) and semantic similarity. This design enables the model to perceive subtle differences in the "closeness" between candidate answers and standard answers (e.g., distinguishing the quality of IoU values of 0.45 and 0.55), thereby constructing a smooth and continuous reward signal. This guides the model not only to "reach the threshold" during optimization but also to continuously approach the standard answer, significantly improving the convergence accuracy and robustness of the generated results of large models when handling continuous multimodal tasks such as visual localization and coordinate prediction.
[0092] Furthermore, the continuous label is the coordinate of the graphic bounding box, and the similarity index is an indicator used to measure the degree of overlap between two bounding boxes.
[0093] In this example, by concretizing continuous labels as "coordinates of graphical bounding boxes" and defining similarity metrics as measures of overlap (such as Intersection over Union (IoU), the abstract visual localization task is transformed into a quantifiable geometric spatial matching problem. This design not only provides the model with accurate spatial perception feedback, enabling it to learn pixel-level or region-level precise localization capabilities, rather than relying solely on vague semantic descriptions; but also, the overlap-based quantization evaluation mechanism effectively solves the training instability problem caused by annotation noise or small biases in traditional methods. This ensures that the model can generate operation instructions with high geometric accuracy and strong spatial consistency in visual-language interaction scenarios such as object detection and user interface element clicks, thereby significantly improving the execution reliability of multimodal agents in complex visual environments.
[0094] It should be noted that the generation phase of large models may involve both discrete and continuous tasks. To better understand discrete and continuous tasks, these two task types will be explained separately below.
[0095] For discrete tasks (such as mathematical reasoning and intelligent navigation), the answer typically has only two explicit states: "correct" or "incorrect," with no intermediate states. The judgment criterion is usually deterministic rule matching. Discrete labels are used as answer labels. When any candidate answer's content is exactly the same as the discrete label, the candidate answer is judged as "correct" and assigned a higher first reward score (e.g., +1); conversely, when the candidate answer is inconsistent with the discrete label, it is judged as "incorrect" and assigned a lower second reward score (e.g., 0 or -1). For example, in an intelligent navigation task, if the correct answer label is "open the settings app," then any candidate answer generated by the model that contains this instruction name is considered correct.
[0096] For continuous tasks (such as visual localization and coordinate prediction), the answers do not have a strict binary opposition (black and white), but rather exhibit a continuously changing quality distribution. For example, in bounding box prediction tasks, as long as the overlap between the predicted box and the ground truth box reaches a certain threshold, it is considered valid, even if they are not completely identical. Continuous labels are used as answer labels. First, a similarity index (e.g., the Intersection over Union (IoU) used to measure the degree of overlap between two bounding boxes) is calculated between any candidate answer and the continuous label. Then, the accuracy is determined based on the numerical range of this similarity index: when the similarity index falls within a first preset interval (e.g., IoU > 0.5), the accuracy of the candidate answer is determined to be "correct," and it is assigned a higher third reward score; when the similarity index does not fall within the first preset interval, the accuracy of the candidate answer is determined to be "incorrect," and it is assigned a lower fourth reward score. For example, in visual localization tasks, if the continuous label is the ground truth coordinate of a graphic bounding box, its quality is quantified by calculating the IoU value between the predicted box and the ground truth box, thereby determining the allocation of reward scores.
[0097] Through the above mechanism, the embodiments of this specification can flexibly adapt to different types of task scenarios, ensuring that whether it is a discrete task with precise matching or a continuous task with fuzzy matching, reasonable and quantifiable generation rewards can be obtained, thereby providing an accurate input basis for subsequent calculation of advantageous signals.
[0098] Next, in step 23, the verification reward corresponding to each verification score is determined based on whether each verification score is consistent with the accuracy of the candidate answer. It can be understood that the accuracy of the verification score can be measured by whether the following conditions are met: "The verification score of the correct answer should be higher than the verification score of the incorrect answer" or "The verification score of the answer with higher accuracy should be higher than the verification score of the answer with lower accuracy".
[0099] In one example, determining the verification reward corresponding to each verification score based on whether each verification score is consistent with the accuracy of the candidate answer includes:
[0100] Based on the accuracy of each candidate answer, the first number of candidate answers are divided into a correct answer group and an incorrect answer group;
[0101] Calculate the first average of the verification scores for each candidate answer in the correct answer group, and the second average of the verification scores for each candidate answer in the incorrect answer group;
[0102] For any candidate answer in the correct answer group, the verification reward corresponding to the verification score of the candidate answer is determined based on whether the verification score of the candidate answer is greater than the second average.
[0103] For any candidate answer in the incorrect answer group, the verification reward corresponding to the verification score of the candidate answer is determined based on whether the verification score of the candidate answer is less than the first average value.
[0104] In this example, by introducing an inter-group cross-comparison mechanism, the calculation of validation rewards is upgraded to a dynamic relative ranking optimization. Its core effect is to force the model to establish strict confidence thresholds. Specifically, by requiring the score of the correct answer group to be higher than the average score (second average) of the incorrect answer group, and simultaneously requiring the score of the incorrect answer group to be lower than the average score (first average) of the correct answer group, the risk of the model "misjudging" incorrect answers as having high confidence or "underestimating" the value of correct answers is effectively eliminated. This design not only significantly improves the discrimination and calibration capabilities of large models in self-evaluation tasks, ensuring that their output validation scores truly reflect the distribution of answer quality, but also guides the model to actively learn the inherent logic that "good answers should receive high scores and bad answers should receive low scores" by constructing a clear positive and negative feedback loop. This achieves autonomous evolution and robustness improvement of validation capabilities without the need for external supervision signals.
[0105] Figure 3 This diagram illustrates the grouping of candidate answers according to one embodiment. (See also...) Figure 3The first number of candidate answers are denoted as S1, S2, S3, S4, S5, S6, S7, and S8. According to preset rules (such as instruction name matching), S1, S2, S3, and S4 are judged as correct, while S5, S6, S7, and S8 are judged as incorrect. Therefore, S1, S2, S3, and S4 are assigned to the correct answer group, and S5, S6, S7, and S8 are assigned to the incorrect answer group. The statistical average of the verification scores within each of the two groups is calculated as the benchmark for the dynamic threshold. The first average (μ+): The average verification score of each candidate answer in the correct answer group is calculated, i.e., (0.9+0.8+0.3+0.6) / 4=0.65. This is the average level of the verification scores of the correct answers and will be used as the penalty threshold for incorrect answers. The second average (μ-): Calculate the average verification score of each candidate answer in the incorrect answer group, i.e., (0.2+0.7+0.4+0.1) / 4=0.35. This is the average verification score of the incorrect answers and will be used as the incentive threshold for the correct answer. In the embodiments of this specification, different thresholds are used to determine the verification reward for candidate answers in different groups.
[0106] It should be noted that for continuous tasks, when determining whether a candidate answer is correct or incorrect based on a similarity metric, a minimum threshold for the difference between the two metrics can be set to ensure the effectiveness of the comparison.
[0107] Figure 4 The diagram illustrates a method for determining verification rewards according to one embodiment, which is based on Figure 3 The grouping is performed according to the illustrated embodiment. Refer to... Figure 4 This embodiment employs a dual-threshold judgment logic. The two thresholds are the aforementioned first average (μ+=0.65) and second average (μ-=0.35). For each candidate answer, its verification score is compared with the average of the verification scores of another group to determine the verification reward corresponding to that candidate answer. Specifically, for each candidate answer in the correct answer group, the reward condition is verification score > μ-. The verification score of each candidate answer is checked to see if it is greater than μ-, and the verification reward for each candidate answer is obtained based on the check result. For each candidate answer in the incorrect answer group, the reward condition is verification score < μ+. The verification score of each candidate answer is checked to see if it is less than μ+, and the verification reward for each candidate answer is obtained based on the check result. Table 1 shows the correspondence between a set of candidate answers and verification rewards.
[0108] Table 1: Correspondence between candidate answers and verification rewards
[0109]
[0110] Referring to Table 1, for the correct answer group {S1, S2, S3, S4}, we check whether the validation score of each candidate answer is greater than μ-. Candidate answer S1 has a validation score of 0.9 > 0.35, meeting the reward condition, and receives a validation reward of 1; candidate answer S2 has a validation score of 0.8 > 0.35, meeting the reward condition, and receives a validation reward of 1; candidate answer S3 has a validation score of 0.3 < 0.35, not meeting the reward condition, and receives a validation reward of 0; candidate answer S4 has a validation score of 0.6 > 0.35, meeting the reward condition, and receives a validation reward of 1. The principle for setting validation rewards for the correct answer group is to encourage the model to give high validation scores for correct answers and to penalize the model for giving low validation scores for correct answers.
[0111] For the incorrect answer group {S5, S6, S7, S8}, check whether the validation score of each candidate answer is less than μ+. The validation score of candidate answer S5 is 0.2 < 0.65, which meets the reward condition, and the validation reward is 1; the validation score of candidate answer S6 is 0.7 > 0.65, which does not meet the reward condition, and the validation reward is 0; the validation score of candidate answer S7 is 0.4 < 0.65, which meets the reward condition, and the validation reward is 1; the validation score of candidate answer S8 is 0.1 < 0.65, which meets the reward condition, and the validation reward is 1. The principle for setting validation rewards for incorrect answer groups is to encourage the model to give low validation scores for incorrect answers and to penalize the model for giving high validation scores for incorrect answers.
[0112] This specification's embodiments construct a cross-constraint mechanism: a positive constraint, requiring correct answers to have a higher average validation score than incorrect answers; and a negative constraint, requiring incorrect answers to have a lower average validation score than correct answers. This mechanism ensures that the model's output validation score is a quality metric with relative discriminative power, thus providing high-quality feedback signals for subsequent advantage calculations. This cross-constraint mechanism can also be called a preference learning mechanism based on inter-group comparisons, designed to train large models to distinguish between correct and incorrect answers, thereby improving the accuracy of their self-assessment.
[0113] It should be noted that when all of the aforementioned first number of candidate answers are either correct or incorrect, the same verification reward can be directly assigned to each candidate answer. For example, the verification reward for each candidate answer is 0.
[0114] In one example, determining the verification reward corresponding to each verification score based on whether each verification score is consistent with the accuracy of the candidate answer includes:
[0115] For any candidate answer, select other candidate answers as the comparison set for the candidate answer based on the difference between the accuracy of the candidate answer and the accuracy of other candidate answers.
[0116] Based on the ranking relationship between the verification score of the candidate answer and the verification scores of candidate answers in the comparison set, and whether it conforms to the expected ranking relationship corresponding to the difference, the verification reward corresponding to the candidate answer is determined.
[0117] In this example, by introducing a relative ranking mechanism based on within-group differences, the calculation of validation rewards is upgraded from relying on fixed thresholds (such as the mean) to a dynamic relative consistency assessment, effectively solving the evaluation bias problem caused by fluctuations in the difficulty of different batches of tasks. Its core effect lies in forcing the model to establish an internal logical alignment: "high-quality answers should receive high confidence, and low-quality answers should receive low confidence." Positive rewards are only awarded when the model's output validation score ranking strictly matches the actual accuracy ranking of the candidate answers. This design not only eliminates the dependence on specific statistics (such as the mean and median), enhancing the algorithm's versatility in both discrete and continuous tasks, but also significantly improves the calibration accuracy of large models' self-evaluation capabilities, ensuring that the output validation scores truly reflect the relative quality of the answers, thereby avoiding evaluation illusions caused by overconfidence or underconfidence.
[0118] In step 24, based on the generation reward corresponding to each candidate answer, the first advantage of the large model regarding answer generation is determined; based on the verification reward corresponding to each verification score, the second advantage of the large model regarding verification capability is determined. It can be understood that calculating the advantages based on generation and verification rewards separately decouples the different optimization objectives.
[0119] In one example, determining the first advantage of the large model in answer generation based on the generation rewards corresponding to each candidate answer includes:
[0120] The arithmetic mean of the generation rewards for multiple candidate answers is calculated as the generation baseline;
[0121] Subtract the generation baseline from the generation reward for each candidate answer to obtain the generation advantage value corresponding to that candidate answer;
[0122] The first advantage is calculated based on the generation advantage value corresponding to each candidate answer.
[0123] In this example, by introducing a dynamic baseline mechanism, the reward is transformed into a relative advantage signal, effectively eliminating the fluctuations in the reward scale caused by differences in task difficulty or sample distribution between different batches. This ensures that the model is always optimized based on performance that is "better than the current average level." This debiasing process not only prevents the model from blindly pursuing high scores in simple tasks or giving up too early in difficult tasks, but also significantly enhances the stability and convergence speed of policy updates, enabling large models to more accurately identify and strengthen those truly discriminative high-quality answer generation paths.
[0124] Further, based on the generation advantage value corresponding to each candidate answer, the first advantage is calculated, including:
[0125] By using a first mask to block out the word positions used to output the validation score, and only retaining the word positions used to generate candidate answers for calculation, the first advantage is obtained.
[0126] In this example, a masking mechanism is used to physically isolate the gradients of the generation and validation tasks. This ensures that when calculating the first advantage, model parameter updates are driven solely by the feedback from the generated answer sequence, completely shielding the interference noise that may be introduced by validation score terms. This design effectively prevents the validation signal from influencing the generation strategy, avoiding the model sacrificing the fluency or accuracy of answer generation to cater to validation rewards. In other words, it resolves the gradient conflict problem in multi-task training, thus ensuring that the large model can independently and purely optimize its language generation capabilities, significantly improving the quality and stability of the generated content.
[0127] In one example, determining the second advantage of the large model regarding validation capability based on the validation rewards corresponding to each validation score includes:
[0128] Calculate the arithmetic mean of the validation rewards for multiple validation scores as the validation baseline;
[0129] Subtract the verification baseline from the verification reward for each verification score to obtain the verification advantage value corresponding to that verification score;
[0130] The second advantage is calculated based on the validation advantage value corresponding to each validation score.
[0131] In this example, by introducing a dynamic baseline mechanism, validation feedback is transformed into a relative advantage signal, effectively eliminating evaluation bias caused by uneven distribution of sample quality within a batch (such as overall scores being too high or too low). This approach ensures that when optimizing validation capabilities, the model no longer simply pursues high scores, but rather focuses on improving the discrimination and calibration accuracy of answer quality. This guides large models to establish more sensitive self-evaluation criteria, significantly improving their robustness and accuracy in identifying true / false answers and quantifying confidence in complex tasks.
[0132] Further, based on the validation advantage value corresponding to each validation score, the second advantage is calculated, including:
[0133] By using a second mask to block out the word positions used to generate candidate answers, and only retaining the word positions used to output the validation score for calculation, a second advantage is obtained.
[0134] In this example, a masking mechanism is used to physically isolate the gradients of the validation and generation tasks. This ensures that when calculating the second advantage, model parameter updates are driven solely by feedback from the validation score sequence, completely shielding the interference noise that may be introduced by the answer text terms. This design effectively prevents the generation signal from influencing the evaluation strategy and avoids the model sacrificing the objectivity of its self-evaluation to cater to generation rewards (such as pursuing high fluency). This resolves the gradient conflict problem in multi-task training, ensuring that the large model can independently and purely optimize its self-calibration and quality judgment capabilities, significantly improving the reliability and discriminative power of the validation scores.
[0135] Figure 5 A schematic diagram of a masking mechanism according to one embodiment is shown. (Refer to...) Figure 5 In the output token sequence of a large model, a complete reasoning process involves tokens with multiple roles: answer-generating tokens responsible for logical reasoning and answer generation (such as...). <think> I need to calculate 2×3< / think> <answer> 6< / answer> ), and the verification score morphemes responsible for self-assessment (such as <score> 0.9< / score> ),in, <think> and< / think> It is a special word element, and the word element between the two corresponds to logical reasoning; <answer> and< / answer> This is another special lexical unit, and the lexical unit between the two corresponds to the answer; <score> and< / score> As another special type of lexical unit, the lexical unit between the two corresponds to the validation score. Although these lexical units belong to the same sequence, they correspond to different optimization objectives in terms of logical function: the former needs to maximize the generation reward, i.e., the correctness of the answer, while the latter needs to maximize the validation reward, i.e., the reasonableness of the score. If all lexical units are included in a unified loss calculation without distinction, the model may exhibit speculative behavior. For example, the model may deliberately generate an incorrect candidate answer, but then give an extremely low validation score to cater to the validation reward, resulting in poor quality of the final output answer. To address the above problem, the embodiments in this specification construct two complementary and mutually exclusive masks: a first mask and a second mask. Both masks can be binary vectors of the same length as the output lexical unit sequence, where a vector element of 1 represents that the position participates in the gradient calculation of the current round, and a vector element of 0 represents that the position is masked.
[0136] The first mask, used in the calculation of the first advantage, only retains the positions of the terms used to generate candidate answers (i.e., the aforementioned answer generation terms), while masking the positions of the terms used to output the validation score. When calculating the first advantage, the first mask is used to filter the gradient. At this point, the generated reward signal only affects the parameter updates of the answer generation part, while the parameters of the validation score part are frozen and unaffected by the generated reward signal. This forces the model to focus on improving the accuracy and logic of answer generation, preventing the model from sacrificing answer quality for high validation rewards.
[0137] The second mask, used in the calculation of the second advantage, only retains the lexical positions used to output the validation score; that is, the aforementioned validation score lexical positions are masked from the lexical positions used to generate candidate answers. When calculating the second advantage, the gradient is filtered using the second mask. At this point, the validation reward signal only affects the parameter updates of the validation score part, while the parameters of the answer generation part are frozen and unaffected by the validation reward signal. This forces the model to focus on improving its self-evaluation calibration capabilities, preventing the model from assigning excessively high confidence scores to incorrect answers in order to cater to the generation reward.
[0138] In this embodiment, through the collaborative operation of the two mutually exclusive masks, the generation pipeline only updates parameters related to "answering questions," eliminating interference from the "scoring" task and focusing solely on making the answer correct. The verification pipeline only updates parameters related to "scoring," eliminating interference from the "answering questions" task and focusing solely on making the judgment accurate. This design ensures that two distinct optimization objectives evolve in parallel within the same model architecture, avoiding training oscillations caused by gradient cancellation and eliminating the possibility of the model obtaining false high scores, thereby significantly improving the overall performance and robustness of large models under complex tasks.
[0139] Understandably, depending on the application scenario, the meaning and form of the answer in the output word sequence of a large model will also differ. Figure 5 In mathematical calculation scenarios, the answer is a number, but the answer is not limited to this. For example, in the scenario of a mobile intelligent assistant, the output word sequence is " <think> The item I want to select is... I need to click...< / think> <answer> Click(box=(220,350,480,510))< / answer> <score> 0.8< / score> In this scenario, the answer is the name of the operation instruction and the coordinates of the graphic bounding box.
[0140] Finally, in step 25, the parameters of the large model are adjusted with the goal of maximizing the sum of the first advantage and the second advantage. It is understood that the sum of the first advantage and the second advantage covers the entire output word sequence of the large model, allowing for joint optimization of the large model.
[0141] Figure 6 A schematic diagram illustrating joint optimization according to one embodiment is shown. (Refer to...) Figure 6By processing candidate answers and validation scores in parallel, a unified advantage signal is generated to guide the iterative update of the policy model. Different reward functions are used to process the data from the two paths: first, reward calculation is performed, and then the candidate answers are processed into an answer sequence. , … Using a verifiable reward function, the generation reward for each answer is calculated, denoted as . The score sequence S1, S2…S is formed by verifying the scores. G Using the preference reward function, the validation reward corresponding to each validation score is calculated, denoted as . Then, advantage calculation is performed based on the generated reward. Calculate the first advantage, denoted as ; Based on verification rewards The second advantage is calculated and denoted as Finally, the first and second advantages mentioned above are combined by weighted summation or other aggregation operations to generate a unified advantage. This unified advantage integrates the objective generation effect and subjective quality evaluation of the answer and will be fed back to the large model to guide its parameter updates, thereby achieving end-to-end optimization of the large model's behavioral strategy.
[0142] In this context, the large model serves as the policy model in reinforcement learning. The entire output sequence of the large model (including candidate answers and validation scores) is treated as a series of discrete actions. The generation of each word is a decision made by the model based on the current state (input and historical outputs). The goal of the training process is to maximize the expected reward (i.e., the sum of the first advantage and the second advantage). Through backpropagation, the weight parameters of the large model are continuously adjusted to change its output probability distribution. For the generation part, parameter adjustment makes the model more inclined to output word sequences consistent with the standard answer. For the validation part, parameter adjustment makes the model more inclined to output score characters that accurately distinguish between right and wrong. As a policy model, the large model not only passively responds to queries but also actively optimizes its internal reasoning logic and evaluation criteria through reinforcement learning feedback loops, achieving a transformation from static knowledge storage to dynamic capability evolution.
[0143] In the embodiments of this specification, the verifiable reward function is used to calculate the generation reward. It quantifies the correctness of the generated answer by calculating the consistency between the candidate answer and the answer label (such as the instruction name or bounding box coordinates), thereby giving the corresponding reward score. The preference reward function is used to calculate the verification reward. It determines the corresponding reward score by comparing the quality of different candidate answers, thereby guiding the model to allocate higher verification scores to better answers, that is, giving higher verification scores to correct answers and lower verification scores to incorrect answers.
[0144] The embodiments in this specification can be based on the following overall optimization objective function: Reinforcement learning is used to train a large model, among which... Represents the first mask. Represents the second mask. Representing the primary advantage, This represents the second advantage. It decouples the traditional single-policy optimization process into two independent and parallel optimization subtasks through a masking mechanism: answer generation optimization and verification capability optimization.
[0145] Alternatively, the overall optimization objective function can also be: ,in, Representing the primary advantage, Representing the second advantage, Represents the KL divergence. To verify the target weight, is the KL regularization coefficient.
[0146] The method provided in the embodiments of this specification effectively solves the bottleneck problems of high memory consumption and large inference latency caused by maintaining independent value networks or verification modules in traditional dual-model schemes by constructing dual-path advantage signals for generating rewards and verifying rewards, and realizing independent calculation and collaborative optimization of the two under a single-model architecture. Simultaneously, the advantage decoupling mechanism avoids gradient conflicts in multi-task training, preventing the model from sacrificing the performance of another dimension while pursuing a single objective. This allows the model to simultaneously improve the accuracy of answer generation and the calibration of self-evaluation without the burden of additional parameters. In summary, this scheme simplifies the training process, reduces computational resource consumption, and balances generation quality and verification reliability.
[0147] According to another embodiment, a reinforcement learning training apparatus for large models is also provided, which is used to perform the methods provided in the embodiments of this specification. Figure 7 A schematic block diagram of a reinforcement learning training apparatus for a large model according to one embodiment is shown. Figure 7 As shown, the device 700 includes:
[0148] The generation unit 71 is used to receive query information and generate a first number of candidate answers and corresponding first number of verification scores for evaluating the correctness of the candidate answers using the large model; the first number is an integer greater than 1.
[0149] The first reward unit 72 is used to determine the accuracy of each candidate answer and the corresponding generation reward based on whether each candidate answer obtained by the generation unit 71 is consistent with the answer label.
[0150] The second reward unit 73 is used to determine the verification reward corresponding to each verification score based on whether the accuracy of each verification score obtained by the generation unit 71 is consistent with that of the candidate answer obtained by the first reward unit 72.
[0151] The advantage determination unit 74 is used to determine the first advantage of the large model in answer generation based on the generation rewards corresponding to each candidate answer obtained by the first reward unit 72; and to determine the second advantage of the large model in verification ability based on the verification rewards corresponding to each verification score obtained by the second reward unit 73.
[0152] The adjustment unit 75 is used to adjust the parameters of the large model with the goal of maximizing the sum of the first advantage and the second advantage obtained by the advantage determination unit 74.
[0153] Optionally, as an embodiment, the answer label includes a discrete label. The first reward unit 72 is specifically used to determine the accuracy of any candidate answer as correct and to determine the corresponding generated reward as a first reward score when any candidate answer is the same as the discrete label; and to determine the accuracy of any candidate answer as incorrect and to determine the corresponding generated reward as a second reward score when any candidate answer is different from the discrete label.
[0154] Furthermore, the discrete label is the instruction name of the first operation instruction.
[0155] Optionally, as an embodiment, the answer label includes a continuous label, and the first reward unit 72 is specifically used to calculate the similarity index between any candidate answer and the continuous label; when the similarity index belongs to a first preset interval, the accuracy of the candidate answer is determined to be correct, and the corresponding generated reward is determined to be a third reward score; when the similarity index does not belong to the first preset interval, the accuracy of the candidate answer is determined to be incorrect, and the corresponding generated reward is determined to be a fourth reward score.
[0156] Optionally, as an embodiment, the answer label includes a continuous label, and the first reward unit 72 is specifically used to calculate the similarity index between any candidate answer and the continuous label, and use the similarity index as the accuracy corresponding to the candidate answer and the reward score for generating the reward.
[0157] Furthermore, the continuous label is the coordinate of the graphic bounding box, and the similarity index is an indicator used to measure the degree of overlap between two bounding boxes.
[0158] Optionally, as an embodiment, the second reward unit 73 includes:
[0159] Grouping subunits are used to divide the first number of candidate answers into correct answer groups and incorrect answer groups based on the accuracy of each candidate answer;
[0160] The average value calculation subunit is used to calculate the first average value of the verification scores of each candidate answer in the correct answer group after the grouping subunit has divided the data, and the second average value of the verification scores of each candidate answer in the incorrect answer group.
[0161] The first reward subunit is used to determine the verification reward corresponding to the verification score of any candidate answer in the correct answer group by calculating the second average value obtained by the subunit based on whether the verification score of the candidate answer is greater than the average value.
[0162] The second reward subunit is used to determine the verification reward corresponding to the verification score of any candidate answer in the incorrect answer group by calculating the first average value obtained by the subunit based on whether the verification score of the candidate answer is less than the average value.
[0163] Optionally, as an embodiment, the second reward unit 73 includes:
[0164] The set selection sub-unit is used to select other candidate answers as the comparison set for any candidate answer based on the difference between the accuracy of the candidate answer and the accuracy of other candidate answers.
[0165] The reward determination subunit is used to determine the verification reward corresponding to the candidate answer based on whether the ranking relationship between the verification score of the candidate answer and the verification scores of the candidate answers in the comparison set obtained by the set selection subunit conforms to the expected ranking relationship corresponding to the difference.
[0166] Optionally, as an embodiment, the advantage determining unit 74 includes:
[0167] The first baseline calculation subunit is used to calculate the arithmetic mean of the generation rewards of multiple candidate answers as the generation baseline;
[0168] The first advantage value calculation subunit is used to subtract the generation baseline obtained by the first baseline calculation subunit from the generation reward of each candidate answer to obtain the generation advantage value corresponding to the candidate answer;
[0169] The first advantage calculation subunit is used to calculate the generation advantage value corresponding to each candidate answer obtained by the first advantage value calculation subunit, and to calculate the first advantage.
[0170] Optionally, as an embodiment, the advantage determining unit 74 includes:
[0171] The second baseline calculation subunit is used to calculate the arithmetic mean of the verification rewards of multiple verification scores as the verification baseline;
[0172] The second advantage value calculation subunit is used to subtract the verification baseline obtained by the second baseline calculation subunit from the verification reward of each verification score to obtain the verification advantage value corresponding to the verification score.
[0173] The second advantage calculation subunit is used to calculate the second advantage based on the verification advantage value corresponding to each verification score obtained by the second advantage value calculation subunit.
[0174] Furthermore, the first advantage calculation subunit is specifically used to use a first mask to block out the word positions used to output the verification score, and only retain the word positions used to generate candidate answers to participate in the calculation, thereby obtaining the first advantage.
[0175] Furthermore, the second advantage calculation subunit is specifically used to use a second mask to block out the word positions used to generate candidate answers, and only retain the word positions used to output the verification score to participate in the calculation, thus obtaining the second advantage.
[0176] Optionally, as an example, the large model is a visual-language model, and the query information includes image data or video stream data.
[0177] The apparatus provided in the embodiments of this specification effectively solves the bottleneck problems of high memory consumption and large inference latency caused by maintaining independent value networks or verification modules in traditional dual-model schemes by constructing dual-path advantage signals for generating rewards and verifying rewards, and realizing independent calculation and collaborative optimization of the two under a single-model architecture. Simultaneously, the advantage decoupling mechanism avoids gradient conflicts in multi-task training, preventing the model from sacrificing the performance of another dimension while pursuing a single objective. This allows the model to simultaneously improve the accuracy of answer generation and the calibration of self-evaluation without the burden of additional parameters. In summary, this scheme simplifies the training process, reduces computational resource consumption, and balances generation quality and verification reliability.
[0178] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform a combination Figure 2 The method described.
[0179] According to another embodiment, a computing device is also provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements a combination... Figure 2 The method described.
[0180] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.
[0181] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.
Claims
1. A reinforcement learning training method for large models, comprising: Receive query information, and use the large model to generate a first number of candidate answers and a corresponding first number of verification scores for evaluating the correctness of the candidate answers; The first number is an integer greater than 1; Based on whether each candidate answer matches the answer tag, determine the accuracy of each candidate answer and the corresponding generated reward. The verification reward is determined based on whether each verification score is consistent with the accuracy of the candidate answer. Based on the generation reward corresponding to each candidate answer, determine the first advantage of the large model in answer generation; based on the verification reward corresponding to each verification score, determine the second advantage of the large model in verification capability. The parameters of the large model are adjusted with the goal of maximizing the sum of the first and second advantages.
2. The method of claim 1, wherein, The answer labels include discrete labels. The process of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer label includes: When any candidate answer is the same as a discrete label, the accuracy of the candidate answer is determined to be correct, and the corresponding generated reward is determined to be the first reward score; When any candidate answer differs from the discrete label, the accuracy of that candidate answer is determined to be incorrect, and the corresponding generated reward is determined to be the second reward score.
3. The method of claim 2, wherein, The discrete label is the instruction name of the first operation instruction.
4. The method of claim 1, wherein, The answer tags include continuous tags. The process of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer tag includes: Calculate the similarity index between any candidate answer and continuous labels; When the similarity index falls within the first preset range, the accuracy of the candidate answer is determined to be correct, and the corresponding generated reward is determined to be the third reward score. When the similarity index does not belong to the first preset interval, the accuracy of the candidate answer is determined to be incorrect, and the corresponding generated reward is determined to be the fourth reward score.
5. The method of claim 1, wherein, The answer tags include continuous tags. The process of determining the accuracy and corresponding reward for each candidate answer based on whether each candidate answer matches the answer tag includes: Calculate the similarity index between any candidate answer and the continuous label, and use the similarity index as the accuracy of the candidate answer and the reward score for generating the reward.
6. The method of claim 4 or 5, wherein, The continuous label is the coordinate of the graphic bounding box, and the similarity index is an indicator used to measure the degree of overlap between two bounding boxes.
7. The method of claim 1, wherein, The determination of the verification reward corresponding to each verification score based on whether each verification score is consistent with the accuracy of the candidate answer includes: Based on the accuracy of each candidate answer, the first number of candidate answers are divided into a correct answer group and an incorrect answer group; Calculate the first average of the verification scores for each candidate answer in the correct answer group, and the second average of the verification scores for each candidate answer in the incorrect answer group; For any candidate answer in the correct answer group, the verification reward corresponding to the verification score of the candidate answer is determined based on whether the verification score of the candidate answer is greater than the second average. For any candidate answer in the incorrect answer group, the verification reward corresponding to the verification score of the candidate answer is determined based on whether the verification score of the candidate answer is less than the first average value.
8. The method of claim 1, wherein, The determination of the verification reward corresponding to each verification score based on whether each verification score is consistent with the accuracy of the candidate answer includes: For any candidate answer, select other candidate answers as the comparison set for the candidate answer based on the difference between the accuracy of the candidate answer and the accuracy of other candidate answers. Based on the ranking relationship between the verification score of the candidate answer and the verification scores of candidate answers in the comparison set, and whether it conforms to the expected ranking relationship corresponding to the difference, the verification reward corresponding to the candidate answer is determined.
9. The method of claim 1, wherein, The step of determining the first advantage of the large model in answer generation based on the generated rewards corresponding to each candidate answer includes: The arithmetic mean of the generation rewards for multiple candidate answers is calculated as the generation baseline; Subtract the generation baseline from the generation reward for each candidate answer to obtain the generation advantage value corresponding to that candidate answer; The first advantage is calculated based on the generation advantage value corresponding to each candidate answer.
10. The method of claim 1, wherein, The determination of the second advantage of the large model regarding validation capability based on the validation rewards corresponding to each validation score includes: Calculate the arithmetic mean of the validation rewards for multiple validation scores as the validation baseline; Subtract the verification baseline from the verification reward for each verification score to obtain the verification advantage value corresponding to that verification score; The second advantage is calculated based on the validation advantage value corresponding to each validation score.
11. The method of claim 9, wherein, The first advantage is calculated based on the generated advantage value corresponding to each candidate answer, including: By using a first mask to block out the word positions used to output the validation score, and only retaining the word positions used to generate candidate answers for calculation, the first advantage is obtained.
12. The method of claim 10, wherein, The second advantage is calculated based on the validation advantage value corresponding to each validation score, including: By using a second mask to block out the word positions used to generate candidate answers, and only retaining the word positions used to output the validation score for calculation, a second advantage is obtained.
13. The method of claim 1, wherein, The large model is a vision-language model, and the query information includes image data or video stream data.
14. A reinforcement learning training device for large models, comprising: The generation unit is used to receive query information and generate a first number of candidate answers and corresponding first number of verification scores for evaluating the correctness of the candidate answers using the large model. The first number is an integer greater than 1; The first reward unit is used to determine the accuracy of each candidate answer and the corresponding generation reward based on whether each candidate answer obtained by the generation unit is consistent with the answer label. The second reward unit is used to determine the verification reward corresponding to each verification score based on whether the accuracy of each verification score obtained by the generation unit is consistent with that of the candidate answer obtained by the first reward unit. The advantage determination unit is used to determine the first advantage of the large model in answer generation based on the generation rewards corresponding to each candidate answer obtained by the first reward unit. Based on the verification rewards corresponding to each verification score obtained from the second reward unit, the second advantage of the large model in terms of verification capability is determined. An adjustment unit is used to adjust the parameters of the large model with the goal of maximizing the sum of the first advantage and the second advantage obtained by the advantage determination unit.
15. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-13.