A critic-driven unified reinforcement learning method for test-time self-improvement

CN122840158APending Publication Date: 2026-09-29RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610896006.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0008]本公开的实施例的目的在于提供一种用于可验证自然语言复杂推理任务的面向测试时自改进的批评驱动统一强化学习方法,适用于数学推理、代码生成、自动化答题和程序修复等文本任务,旨在解决现有技术无法在同一策略模型内联合优化求解、批评与批评引导重探三种能力,同时在测试时完全脱离外部监督、实现自主持续改进的问题

Benefits of technology

针对自然语言数学推理和代码生成等可验证文本任务,将求解、验证、批评和重探能力联合训练到同一策略模型中,使模型在测试时无需外部教师模型或真实答案信号即可执行自主迭代改进;通过结果验证器提供确定性奖励,通过结构化批评和延迟奖励筛选真正有助于重探的战略性提示;通过显式丢弃错误初始解并清空或屏蔽其上下文,减少模型复制错误推理路径的锚定偏差;通过经验回放和统一GRPO目标,使零通过困难样本获得有效梯度信号,从而提高自然语言复杂推理任务的求解准确率和测试时扩展能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840158A_ABST
    Figure CN122840158A_ABST
Patent Text Reader

Abstract

The disclosure provides a test-oriented self-improving criticism-driven unified reinforcement learning method. It includes steps 1 to 4: Step 1: initial solution generation, obtaining the solution and verification of the initial average correct rate; Step 2: criticism generation using selective sampling and mixed advantage; Step 3: criticism-guided re-exploration, explicitly discarding the initial incorrect solution in the context of guided re-exploration to eliminate anchoring bias; Step 4: build a unified training target to realize capability fusion, so as to internalize the solution, criticism and criticism-guided re-exploration into the same model, make the model realize continuous self-improvement, and finally obtain and output the model that can be used for verifiable natural language complex reasoning tasks. In this way, the problem that the prior art cannot jointly optimize the three capabilities of solution, criticism and criticism-guided re-exploration in the same strategy model, while completely separating from external supervision during testing and realizing autonomous continuous improvement is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer science, and more specifically, to a test-time self-improving criticism-driven unified reinforcement learning method. Background Technology

[0002] The performance enhancement paradigm for Large Language Models (LLMs) is shifting from "expanding training computing power" to "expanding inference-time computing power." Inference-time computing power scaling primarily manifests in two ways: (1) Parallel scaling—improving performance by simultaneously generating and evaluating multiple candidate answers, such as Best-of-N and Majority Voting; and (2) Sequential scaling—guiding the model to deeper levels of reasoning using longer thought chains and Reinforcement Learning with Verifiable Rewards (RLVR). Representative achievements of the latter, such as DeepSeek-R1 and OpenAI o1, have demonstrated outstanding reasoning capabilities in fields such as mathematics and coding.

[0003] However, existing technologies have the following significant shortcomings in supporting iterative self-improvement: (1) Parallel expansion methods (such as majority voting and Best-of-N) only sample independent solutions in parallel and cannot iteratively improve the same answer. Essentially, they are breadth expansion rather than depth improvement.

[0004] (2) Current inference models based on sequential expansion do not explicitly optimize self-verification and self-correction capabilities simultaneously during training, resulting in insufficient self-verification reliability during testing and difficulty in forming an effective self-improvement loop.

[0005] (3) Existing critique-guided exploration methods rely on external signals to trigger the critique process, including using stronger teacher models to provide external feedback or relying on ground-truth to judge the correctness of the solution. These external dependencies are not available during testing, making it impossible for the model to autonomously run the "solution-verification-critique-re-exploration" closed loop during the reasoning stage.

[0006] (4) Existing self-correction methods (such as SCoRe, STILL-3, etc.) attach the initial erroneous solution to the context when regenerating, which causes the model to produce obvious anchoring bias, tending to repeat previous errors, limiting the diversity of exploration and the final error correction effect.

[0007] In summary, none of the existing technologies can jointly optimize the solution, critique and critique-guided re-exploration capabilities within the same strategy model, while also being able to completely detach from external supervision during testing and achieve autonomous continuous improvement. Summary of the Invention

[0008] The purpose of this disclosure is to provide a test-time self-improving criticism-driven unified reinforcement learning method for verifiable natural language complex reasoning tasks. It is applicable to text tasks such as mathematical reasoning, code generation, automated question answering, and program repair. It aims to solve the problem that existing technologies cannot jointly optimize the solution, criticism, and criticism-guided re-exploration capabilities within the same policy model, while achieving autonomous continuous improvement completely independent of external supervision during testing.

[0009] In one general aspect, a test-time self-improving criticism-driven unified reinforcement learning method for verifiable natural language complex reasoning tasks is provided, comprising steps one through four: Step 1: Generate an initial solution for the training problem, which includes the problem text, constraints, and verifiable objectives, and obtain the initial average accuracy through the result validator to achieve automatic verification of the problem solution results; Step 2: Use selective sampling, structured criticism generation, and hybrid dominance functions to train criticism, so that the same policy model outputs validation judgments and strategic hints; Step 3: Conduct critical guided re-exploration. Explicitly discard the initial solution that was judged to be wrong at the context level, and construct re-exploration prompts based only on the original question and strategic hints to eliminate anchoring bias caused by erroneous reasoning paths; Step 4: Construct a unified training objective to achieve capability fusion, thereby internalizing the solution, critique, and critique-guided re-exploration into the same policy model, enabling the model to achieve continuous self-improvement, and finally obtaining and outputting a test-time self-improving language model that can be used for verifiable complex natural language reasoning tasks.

[0010] Step 1: The specific implementation of the initial solution generation step is as follows: For each verifiable natural language complex inference training problem x, start from the current policy model π θ Sample K initial solutions to form an initial solution set S. init ={y1,…,y K}; where training problem x includes mathematical word problems and / or code generation problems, and contains problem text, constraints, target answer or test cases; the result validator is used to verify each initial solution y i Assign binary reward r i ∈{0,1}, and calculate the initial average accuracy R. Acc (x)=(1 / K)·Σr iThe result verifier is pre-configured as either a mathematical reasoning result verifier or a code generation result verifier according to the training task branch. Specifically, when the current training batch belongs to the mathematical reasoning task branch, the mathematical reasoning result verifier is used to verify the initial solution. The mathematical reasoning result verifier can be implemented based on math_verify, and its verification process includes: first, extracting candidate final answers from the model output text according to preset answer format and priority rules; then, standardizing the format of the candidate final answers, including removing or correcting format differences in common LaTeX expressions, boxed environments, unit text, percentages, sets, intervals, matrices, or equations; subsequently, converting the standardized candidate answers into a unified symbolic representation; finally, comparing the candidate answers with the standard answers, including string normalization comparison, numerical precision tolerance comparison, symbolic expression equivalence comparison, set or interval equivalence comparison, matrix expression equivalence comparison, and equivalence comparison after relational direction reversal. If the candidate answer and the standard answer are determined to be equivalent through any of the above applicable comparison methods, a reward r is output. i =1, otherwise output reward r i =0. When the current training batch belongs to the code generation task branch, the initial solution is verified using a code generation result validator. The code generation result validator can be implemented based on EvalPlus, and its verification process includes: first, extracting the program code to be executed from the model output text, and performing necessary post-processing on the code according to the task entry function, function signature, or problem hints; then loading the program code in a controlled execution environment, and executing the target entry function based on the original test input and the enhanced test input respectively; wherein, the enhanced test input is used to cover boundary cases and robustness cases not covered by the original test cases; then, comparing the program execution result with the expected result generated by the preset assertion or benchmark implementation. If the program completes the run within the time and memory limits by being compiled or interpreted, and passes all the original tests and enhanced tests, then the reward r is output. i =1; If a syntax error, compilation error, runtime error, timeout, memory limit violation, or test output mismatch occurs, the reward r will be output. i =0.

[0011] Step 2: The specific implementation of selective sampling and mixed-advantage criticism generation is as follows: based on R... Acc (x) Divide the training problem into a zero-pass group, a partial pass group, and a full pass group, where the zero-pass group satisfies R. Acc (x)=0, partially passing through the group satisfies 0. <R Acc (x)<1, the complete set satisfies R Acc(x)=1; According to the preset sampling budget, negative samples are preferentially drawn from the zero-pass group, and positive and negative samples with the same problem context are drawn from the partial pass group, and then in the criticism training batch B. crit To maintain a 1:1 balance between positive samples with r=1 and negative samples with r=0, in order to avoid the verification judgment degenerating into a single-class prediction; For each sampled (problem, solution) pair (x, y) i The current strategy model generates structured criticisms in the order of "verification and judgment generation—error type localization—hint abstraction—hint filtering". i =(v i ;h i First, based on the validator reward r i Generate verification judgment v i ∈{Correct,Incorrect}, where Incorrect corresponds to the self-verification field Error_Found=True, and Correct corresponds to Error_Found=False; subsequently in r i When the value is 0, the error category is determined, which includes at least answer extraction errors, constraint omissions, incomplete reasoning paths, symbolic or numerical calculation errors, and code compilation or testing failures. Then, the error category is converted into a strategic hint h that does not reveal the final answer or reiterate the initial incorrect solution. i h i Used to indicate constraints, inference directions, or program boundary conditions that need to be re-examined; finally, h is detected through prompt filtering rules. i With y i The degree of overlap between segments and whether they contain direct answers are considered. If the degree of overlap exceeds the threshold or a direct answer is contained, a formatting penalty is applied to ensure that higher-level concepts are used for the re-exploration phase. The training objective employs a hybrid advantage function oriented towards criticism: for criticism c i The verification segment tokens are set with a verification mask m. v (t)=1 and assign a fixed positive odds A v =1.0; Set m for prompt segment words v (t)=0, and the average correctness of the re-exploration solution set is used. Format constraints and rewards R format (c i ), and the gain item. and the penalty for leaking P overlap (h i ,y i ) constitutes delayed reward , where R format (c i =-λ1·I[Missing required field]-λ2·max(0,L) min-L(h i )) / L min -λ3·I [Hint contains direct answer]; Hint segment advantage A hint =R(c i )-b(x), b(x)=R Acc (x); Ultimate critical lexical advantage A crit (t)=m v (t)·A v +(1-m v (t))·A hint This ensures that only strategic hints that improve re-exploration accuracy and satisfy structural constraints receive positive gradients.

[0012] Step 3: The specific implementation method for conducting criticism-guided re-exploration to eliminate anchoring bias is as follows: First, based on the result validator's reward r... i Determine the initial set of erroneous solutions E={y i |r i =0}, during testing, the initial solution to be re-explored is determined based on the model's self-verification output Error_Found=True (corresponding to the verification judgment Incorrect); for each incorrect initial solution y i It is only used as an object of evaluation in the criticism generation stage and is not included in the re-exploration prompts; the strategic prompts h i Concatenate with the original question x to construct the guiding prompt x' i =Template(x,h i ), where Template(x,h i This includes the original question stem, constraints, hints, and output format requirements; The explicit discarding includes: removing the initial erroneous solution y from the lexical sequence of the re-exploration prompt word. i Its reasoning steps, final answer, code snippets, and related information. i The corresponding historical dialogue identifier; cleared by y before invoking the strategy model for re-exploration decoding. i The generated decoding buffer and position state; perform a fragment matching check on the retest prompt to confirm that there is no fragment from y. i After processing the continuous word segments, input them into the policy model so that the model is based solely on the original question x and the higher-level cue h. i Reconstruct the reasoning path from the new state to avoid continuing to pay attention to or replicate erroneous contexts; The model is based on the guiding prompt word x' checked through fragment matching. i Generate M re-exploration solutions S guided ={y' i,1 ,…,y' i,M}, and the result validator obtains the reward r' respectively. i,j ; Calculate the average accuracy of the re-exploration solution set ,Will Feedback to critic c i As a major component of delayed reward, it forms a closed training loop of "criticism - explicit discard - re-exploration - verification - reward".

[0013] Step 4: The specific method for constructing a unified training objective is as follows: construct a joint optimization objective J based on GRPO. total =η s ·J solve +η c ·J critique +η g ·J guided η s η c and η g The weighting coefficients are set to 1:1:1 by default; where, for any subtask data D and advantage A, J is used. GRPO (D,A)=E[(1 / |o|)·Σ t min(ρ t A,clip(ρ t ,1-ε,1+ε)A)-βD KL (π θ ||π ref )],ρ t ε is the ratio of the probability of the current strategy to that of the old strategy at the t-th word, β is the pruning coefficient, and β is the KL constraint coefficient. (1) Solve for the objective J solve : Select r' from the retest solutions of the zero-pass group i,j =1 successfully guided solution S success , will S success The context of the prompt h i After masking, it is replayed to the initial ungroup S. init Constructing augmented set S aug =S init ∪MaskHint(S success ), and in S aug The relative advantage of the group is recalculated internally, so that the initial solutions that were originally all wrong have a negative advantage, and the solutions are successfully guided to have a positive advantage, thereby injecting an effective learning signal into the difficult samples; (2) Criticism Target J critique : Structured criticism i For the output, use the verification mask m v (t) Differentiate between the verification segment and the prompt segment. For the verification segment, use fixed-dominance supervised learning to learn Error_Found judgments consistent with Correct / Incorrect. For the prompt segment, use A... hint =R(c i )-RAcc (x) leverages the centralized advantage of delayed rewards and enhances accuracy, usability, and non-disclosure through joint constraints of format penalties, length penalties, and penalties for overlapping with incorrect solutions; (3) Guiding the solution of objective J guided : Containing only the original question x and strategic hint h i The guiding prompt word x' i As input, optimize policy π θ (y'|x,h) generates a re-exploration solution, and in each S guided Internally, the relative advantage of the group is calculated according to the validator reward, so that the policy learning follows its own prompts to generate new reasoning paths; The three sub-objectives independently calculate the mean, standard deviation, and word-level loss within their respective sub-task groups, and then sum them with equal weights after normalization according to the number of effective words. When any sub-task batch is empty, the corresponding target weights are reset to zero and the remaining weights are renormalized to maintain the consistency of optimization dimensions and gradient stability during training.

[0014] The technical effects to be achieved by the embodiments of the present invention are as follows: For verifiable text tasks such as natural language mathematical reasoning and code generation, the solution, verification, critique, and re-exploration capabilities are jointly trained into the same policy model, enabling the model to perform autonomous iterative improvement during testing without the need for an external teacher model or real answer signals. A deterministic reward is provided through a result validator, and strategic cues that truly contribute to re-exploration are filtered through structured critique and delayed rewards. Anchoring bias in the model's replication of erroneous reasoning paths is reduced by explicitly discarding incorrect initial solutions and clearing or masking their context. Experience replay and a unified GRPO objective enable zero to obtain effective gradient signals through difficult samples, thereby improving the solution accuracy and test-time scalability for complex natural language reasoning tasks. Attached Figure Description

[0015] The above and other objects and features of this disclosure will become clearer from the following description taken in conjunction with the accompanying drawings.

[0016] Figure 1 This is an architectural diagram illustrating a test-time self-improving criticism-driven unified reinforcement learning method according to an embodiment of the present disclosure; Detailed Implementation

[0017] The following detailed embodiments are provided to aid the reader in gaining a comprehensive understanding of the methods, apparatus, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but may be changed as will become clear upon understanding this disclosure, except for operations that must occur in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.

[0018] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein are provided only to illustrate some of the many feasible ways of implementing the methods, apparatus, and / or systems described herein, which will become clear upon understanding the disclosure of this application.

[0019] As used herein, the term “and / or” includes any one of the associated listed items and any combination of any two or more.

[0020] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, assemblies, regions, layers, or parts, these components, assemblies, regions, layers, or parts should not be limited by these terms. Rather, these terms are used only to distinguish one component, assembly, region, layer, or part from another. Thus, without departing from the teaching of the examples described herein, the first component, first assembly, first region, first layer, or first part referred to as the first component, first assembly, first region, first layer, or first part may also be referred to as the second component, second assembly, second region, second layer, or second part.

[0021] In the specification, when an element (such as a layer, region, or substrate) is described as being "on" another element, "connected to," or "bonded to" another element, the element may be directly "on" another element, directly "connected to," or "bonded to" the other element, or one or more other elements may be present in between. Conversely, when an element is described as being "directly on" another element, "directly connected to," or "directly bonded to" another element, no other elements may be present in between.

[0022] The terminology used herein is for the purpose of describing various examples only and is not intended to limit disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. The terms “comprising,” “including,” and “having” indicate the presence of the described features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0023] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains upon understanding this disclosure. Unless expressly defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and in this disclosure, and shall not be interpreted in an idealized or overly formalistic manner.

[0024] Furthermore, in the description of the examples, detailed descriptions of well-known related structures or functions will be omitted when it is believed that such detailed descriptions would lead to a vague interpretation of this disclosure.

[0025] Figure 1 This is a schematic diagram illustrating a test-time self-improving criticism-driven unified reinforcement learning method according to an embodiment of the present disclosure.

[0026] To achieve the aforementioned objectives, the present invention employs the following technical framework: Figure 1 As shown. Figure 1 The left side compares the key differences between the CURE re-exploration strategy (explicitly discarding erroneous solutions and re-exploring only based on high-level cues) and existing methods (attaching context to erroneous solutions); the right side demonstrates the specific implementation of the joint training framework's ability to simultaneously optimize the solution, critique, and guide the solution within a single policy model.

[0027] This invention proposes a test-time self-improving critique-driven Unified Reinforcement Learning (CURE) framework for verifiable complex natural language reasoning tasks. The task can be a mathematical word problem, symbolic computation problem, code generation problem, or program repair problem. The input is a natural language question stem and constraints, and the output is the answer, reasoning process, or executable code. CURE integrates solving, critique, and critique-guided re-exploration capabilities into a single policy model, enabling the model to autonomously execute an iterative closed loop of "solving-verifying-critique-re-exploration" without any external supervision during testing, achieving continuous self-improvement. The implementation flowchart is shown below. Figure 1 As shown, the specific steps are as follows: Step 1: The specific implementation of the initial solution generation step is as follows: For each verifiable natural language complex inference training problem x, start from the current policy model π θ Sample K initial solutions to form an initial solution set S. init ={y1,…,y K}; where training problem x includes mathematical word problems and / or code generation problems, and contains problem text, constraints, target answer or test cases; the result validator is used to verify each initial solution y i Assign binary reward r i ∈{0,1}, and calculate the initial average accuracy R. Acc (x)=(1 / K)·Σr i The result verifier is pre-configured as either a mathematical reasoning result verifier or a code generation result verifier according to the training task branch. Specifically, when the current training batch belongs to the mathematical reasoning task branch, the mathematical reasoning result verifier is used to verify the initial solution. The mathematical reasoning result verifier can be implemented based on math_verify, and its verification process includes: first, extracting candidate final answers from the model output text according to preset answer format and priority rules; then, standardizing the format of the candidate final answers, including removing or correcting format differences in common LaTeX expressions, boxed environments, unit text, percentages, sets, intervals, matrices, or equations; subsequently, converting the standardized candidate answers into a unified symbolic representation; finally, comparing the candidate answers with the standard answers, including string normalization comparison, numerical precision tolerance comparison, symbolic expression equivalence comparison, set or interval equivalence comparison, matrix expression equivalence comparison, and equivalence comparison after relational direction reversal. If the candidate answer and the standard answer are determined to be equivalent through any of the above applicable comparison methods, a reward r is output. i =1, otherwise output reward r i =0. When the current training batch belongs to the code generation task branch, the initial solution is verified using a code generation result validator. The code generation result validator can be implemented based on EvalPlus, and its verification process includes: first, extracting the program code to be executed from the model output text, and performing necessary post-processing on the code according to the task entry function, function signature, or problem hints; then loading the program code in a controlled execution environment, and executing the target entry function based on the original test input and the enhanced test input respectively; wherein, the enhanced test input is used to cover boundary cases and robustness cases not covered by the original test cases; then, comparing the program execution result with the expected result generated by the preset assertion or benchmark implementation. If the program completes the run within the time and memory limits by being compiled or interpreted, and passes all the original tests and enhanced tests, then the reward r is output. i =1; If a syntax error, compilation error, runtime error, timeout, memory limit violation, or test output mismatch occurs, the reward r will be output. i =0.

[0028] Step Two: Critique Generation – Selective Sampling and Hybrid Dominance. A selective sampling strategy is used to construct the criticism training batch Bcrit: based on RAcc(x), the training problem is divided into a zero-pass group, a partial-pass group, and a full-pass group. Negative samples are preferentially extracted from the zero-pass group to drive the 0-to-1 breakthrough, while positive and negative control samples under the same problem are extracted from the partial-pass group to learn the validation boundary, and a 1:1 balanced ratio of positive to negative samples is enforced. For each sampled (problem, solution) pair (x, yi), the model generates criticism ci=(vi;hi) according to the process of "validation judgment generation—error type localization—hint abstraction—hint filtering". Here, vi is Correct or Incorrect, and hi is a strategic hint that does not reveal the final answer and does not reiterate the initial incorrect solution, used to indicate constraints, reasoning directions, or program boundary conditions that need to be reconsidered. The training objective uses a hybrid dominance function: the validation segment uses a fixed positive dominance of 1.0; the hint segment uses a hybrid dominance function. Calculate the delayed reward and obtain the centralization advantage with b(x)=RAcc(x) as the baseline, thereby encouraging only hints that can improve re-exploration accuracy, satisfy format constraints and do not replicate incorrect solutions.

[0029] Step 3: Critique-Guided Re-Generation – Eliminating Anchoring Bias. The initial set of erroneous solutions E={yi|ri=0} is determined based on ri=0. During testing, the solution to be re-generated is determined based on the model's self-validation output Error_Found=True (corresponding to the validation judgment Incorrect). For each erroneous initial solution yi, it is only used in the criticism generation stage and not written into the re-generation prompt. The strategic hint hi is concatenated with the original question x to construct the guiding prompt x'i=Template(x,hi). Before the guided re-generation, yi and its reasoning steps, final answer, code snippets, and historical dialogue identifiers are removed from the lexical sequence, and the decoding cache generated by yi is cleared. Subsequently, a fragment matching check is performed to confirm that there are no continuous fragments from yi in the re-generation prompt. The model generates M re-generation solutions Sguided={y'i,1,…,y'i,M} based only on the original question x and the high-level hint hi. r'i,j is obtained from the result validator and calculated. This allows for delayed rewards for criticism (ci), forming a closed loop of "criticism—explicit discard—re-exploration—verification—reward".

[0030] Step 4: Unified GRPO Objective. Based on GRPO, a joint optimization objective Jtotal = ηs·Jsolve + ηc·Jcritique + ηg·Jguided is constructed, with the default ηs:ηc:ηg = 1:1:1. The solve objective Jsolve selects successful guided solutions Ssuccess from the re-exploration solutions of the zero-pass group, masks their context, and reverts them to the initial solution group. Saug = Sinit∪MaskHint(Ssuccess) is constructed, and the group relative advantage is recalculated to inject gradients into hard samples. The critique objective Jcritique distinguishes between validation and cue segments using a validation mask, employing fixed advantage and delayed reward-centered advantage respectively, combined with format penalties, length penalties, and penalties for overlapping with incorrect solutions. The guided solve objective Jguided optimizes πθ(y'|x,h), enabling policy learning to generate new inference paths within a context containing only the original problem and strategic cues. The three sub-objectives are independently normalized within their respective sub-task groups and averaged according to the number of effective lexical units, thereby maintaining consistent optimization dimensions and gradient stability.

[0031] Test-Time Self-Improvement Loop. After training, the model performs up to T rounds of iterative self-improvement for natural language math problems or code generation problems during inference: ① Generate candidate solutions to the problem and output the self-verification field Error_Found; ② If Error_Found=False (corresponding to the verification judgment Correct), directly output the answer or code; ③ If Error_Found=True (corresponding to the verification judgment Incorrect), generate strategic hints, delete the incorrect solutions and their historical context, and re-explore to generate new solutions based only on the original problem and hints; ④ Repeat the verification, criticism, and re-exploration process until the verification is correct or the maximum number of rounds is reached. Since the model has internalized the self-verification and hint generation capabilities, the entire self-improvement loop does not require an external teacher model or real answer signals during testing.

[0032] While some embodiments of this disclosure have been shown and described, those skilled in the art will understand that modifications may be made to these embodiments without departing from the principles and spirit of this disclosure, which are defined by the claims and their equivalents.

Claims

1. A unified reinforcement learning method driven by criticism for test-time self-improvement, characterized in that, For the natural language text reasoning task, steps one through four are included: Step 1: Initial solution generation. Initial solutions are generated for the training problem, which includes the problem text, constraints, and verifiable objectives. The initial average accuracy is obtained through the result validator, thereby achieving automatic verification of the problem solution results. Step 2: Utilize selective sampling and hybridization advantages to generate critiques; Step 3: Conduct a critical guided re-exploration, explicitly discarding the initial erroneous solution in the context of the guided re-exploration to eliminate anchoring bias; Step 4: Construct a unified training objective to achieve capability fusion, thereby internalizing the solution, critique, and critique-guided re-exploration into the same model, enabling the model to continuously improve itself, and finally obtaining and outputting a test-time self-improving language model that can be used for verifiable natural language text complex reasoning tasks.

2. The test-time self-improvement criticism-driven unified reinforcement learning method as described in claim 1, characterized in that, The specific implementation of the initial solution generation step is as follows: For each verifiable natural language text complex reasoning training problem x, from the current policy model π... θ Sample K initial solutions to form an initial solution set S. init ={y1,…,y K }; where training problem x includes mathematical word problems and / or code generation problems, and contains problem text, constraints, target answer or test cases; the result validator is used to verify each initial solution y i Assign binary reward r i ∈{0,1}, and calculate the initial average accuracy R. Acc (x)=(1 / K)·Σr i The result validator is pre-configured as a mathematical reasoning result validator or a code generation result validator according to the training task branch, without setting an online task type identification unit; when the training task branch is a mathematical reasoning task, the mathematical reasoning result validator performs answer extraction, format normalization, unified symbol expression conversion, and equivalence comparison with the standard answer on the model output; when the training task branch is a code generation task, the code generation result validator performs code extraction, controlled execution, original test input and enhanced test input verification on the model output, and outputs a binary reward based on the execution result.

3. The test-time self-improvement criticism-driven unified reinforcement learning method as described in claim 2, characterized in that, The specific method for generating criticisms using selective sampling and hybrid advantages is as follows: Based on R... Acc (x) Divide the training problem into a zero-pass group, a partial pass group, and a full pass group, where the zero-pass group satisfies R. Acc (x)=0, partially passing through the group satisfies 0. <R Acc (x)<1, the complete set satisfies R Acc (x)=1; According to the preset sampling budget, negative samples are preferentially drawn from the zero-pass group, and positive and negative samples with the same problem context are drawn from the partial pass group, and then in the criticism training batch B. crit To maintain a 1:1 balance between positive samples with r=1 and negative samples with r=0, in order to avoid the verification judgment degenerating into a single-class prediction; For each sampled (problem, solution) pair (x, y) i The current strategy model generates structured criticisms in the order of "verification and judgment generation—error type localization—hint abstraction—hint filtering". i =(v i ;h i First, based on the validator reward r i Generate verification judgment v i ∈{Correct,Incorrect}, where Incorrect corresponds to the self-verification field Error_Found=True, and Correct corresponds to Error_Found=False; subsequently in r i When =0, the error category is located. The category includes at least answer extraction error, constraint omission, incomplete reasoning path, symbol or numerical calculation error, and code compilation or test failure. Then, the error category is converted into a strategic hint h that does not reveal the final answer and does not restate the initial incorrect solution. i h i Used to indicate constraints, inference directions, or program boundary conditions that need to be re-examined; finally, h is detected through prompt filtering rules. i With y i The degree of overlap between segments and whether they contain direct answers are considered. If the degree of overlap exceeds the threshold or a direct answer is contained, a formatting penalty is applied to ensure that higher-level concepts are used for the re-exploration phase. The training objective employs a hybrid advantage function oriented towards criticism: for criticism c i The verification segment tokens are set with a verification mask m. v (t)=1 and assign a fixed positive odds A v =1.0; Set m for prompt segment words v (t)=0, and the average correctness of the re-exploration solution set is used. Format constraints and rewards R format (c i ), and the gain item. and the penalty for leaking P overlap (h i ,y i ) constitutes delayed reward , where R format (c i =-λ1·I[Missing required field]-λ2·max(0,L) min -L(h i )) / L min -λ3·I [Hint contains direct answer]; Hint segment advantage A hint =R(c i )-b(x), b(x)=R Acc (x); Ultimate critical lexical advantage A crit (t)=m v (t)·A v +(1-m v (t))·A hint This ensures that only strategic hints that improve re-exploration accuracy and satisfy structural constraints receive positive gradients.

4. The test-time self-improvement criticism-driven unified reinforcement learning method as described in claim 3, characterized in that, The specific method for conducting critical guided re-exploration is as follows: first, based on the reward r of the result validator... i Determine the initial set of erroneous solutions E={y i |r i =0}, during testing, the model self-verifies and outputs Error_Found=True, corresponding to the verification judgment Incorrect, and determines that the initial solution should be re-explored; For each incorrect initial solution y i It is only used as an object of evaluation in the criticism generation stage and is not included in the re-exploration prompts; the strategic prompts h i Concatenate with the original question x to construct the guiding prompt x' i =Template(x,h i ), where Template(x,h i This includes the original question stem, constraints, hints, and output format requirements; The explicit discarding includes: removing the initial erroneous solution y from the lexical sequence of the re-exploration prompt word. i Its reasoning steps, final answer, code snippets, and related information. i The corresponding historical dialogue identifier; cleared by y before invoking the strategy model for re-exploration decoding. i The generated decoding buffer and position state; perform a fragment matching check on the retest prompt to confirm that there is no fragment from y. i After processing the continuous word segments, input them into the policy model so that the model is based solely on the original question x and the higher-level cue h. i Reconstruct the reasoning path from the new state to avoid continuing to pay attention to or replicate erroneous contexts; The model is based on the guiding prompt word x' checked through fragment matching. i Generate M re-exploration solutions S guided ={y' i,1 ,…,y' i,M }, and the result validator obtains the reward r' respectively. i,j ; Calculate the average accuracy of the re-exploration solution set ,Will Feedback to critic c i As a major component of delayed reward, it forms a closed training loop of "criticism - explicit discard - re-exploration - verification - reward".

5. The test-time self-improvement criticism-driven unified reinforcement learning method as described in claim 4, characterized in that, The specific method for constructing a unified training objective is as follows: constructing a joint optimization objective J based on GRPO. total =η s ·J solve +η c ·J critique +η g ·J guided η s η c and η g The weighting coefficients are set to 1:1:1 by default; where, for any subtask data D and advantage A, J is used. GRPO (D,A)=E[(1 / |o|)·Σ t min(ρ t A,clip(ρ t ,1-ε,1+ε)A)-βD KL (π θ ||π ref )],ρ t ε is the ratio of the probability of the current strategy to that of the old strategy at the t-th word, β is the pruning coefficient, and β is the KL constraint coefficient. (1) Solve for the objective J solve : Select r' from the retest solutions of the zero-pass group i,j =1 successfully guided solution S success , will S success The context of the prompt h i After masking, it is replayed to the initial ungroup S. init Construct augmented set S aug =S init ∪MaskHint(S success ), and in S aug The relative advantage of the group is recalculated, so that the initial solution that was originally all wrong gains a negative advantage, and the solution is successfully guided to gain a positive advantage, thereby injecting an effective learning signal into the difficult sample; (2) Criticism Target J critique : Structured criticism i For the output, use the verification mask m v (t) Differentiate between the verification segment and the prompt segment. For the verification segment, use fixed-dominance supervised learning to learn Error_Found judgments consistent with Correct / Incorrect. For the prompt segment, use A... hint =R(c i )-R Acc (x) leverages the centralized advantage of delayed rewards and enhances accuracy, usability, and non-disclosure through joint constraints of format penalties, length penalties, and penalties for overlapping with incorrect solutions; (3) Guiding the solution of objective J guided : Containing only the original question x and strategic hint h i The guiding prompt word x' i As input, optimize policy π θ (y'|x,h) generates a re-exploration solution, and in each S guided Internally, the relative advantage of the group is calculated according to the validator reward, so that the policy learning follows its own prompts to generate new reasoning paths; The three sub-objectives independently calculate the mean, standard deviation, and word-level loss within their respective sub-task groups, and then sum them with equal weights after normalization according to the number of effective words. When any sub-task batch is empty, the corresponding target weights are reset to zero and the remaining weights are renormalized to maintain the consistency of optimization dimensions and gradient stability during training.