Optimization method, device, medium and product of model

By identifying the first correct answer and the boundaries of subsequent reflection steps in the large reasoning model, and combining this with a reward mechanism to optimize training, the problem of overthinking in the large reasoning model is solved, thus improving the accuracy and efficiency of reasoning.

CN122264129APending Publication Date: 2026-06-23INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-05-26
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Large reasoning models exhibit overthinking during the reasoning process, leading to wasted computing resources, increased reasoning latency, and reduced accuracy of answers. Existing technologies cannot effectively suppress overthinking and improve accuracy.

Method used

By acquiring the inference dataset of the model to be processed, identifying the location of the first correct answer and the boundaries of subsequent reflection steps, and combining the design of a reward mechanism, generating redundant step correction rewards, training and optimizing the model, and suppressing the phenomenon of excessive reflection.

Benefits of technology

Effectively identify and suppress redundant verification steps in the model, improve inference accuracy and computational efficiency, and achieve refined management and control of the inference trajectory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122264129A_ABST
    Figure CN122264129A_ABST
Patent Text Reader

Abstract

This application discloses a model optimization method, device, medium, and product, relating to the field of artificial intelligence technology. The method includes: pre-acquiring a reasoning dataset for the model, each dataset comprising an original question, a standard answer to the original question, and the reasoning trajectory text of the model to be processed for the original question; identifying a first node for identifying the location of the first correct answer and a second node for identifying the boundaries of subsequent reflection steps based on the standard answer and the reasoning trajectory text; during training and optimization of the model to be processed, fusing the first and second nodes to generate a redundant step correction reward to suppress overthinking; and by explicitly locating the first correct answer and the boundaries of subsequent reflection steps, fusing the trajectories of the answer discovery stage and unnecessary post-answer verification stages into the redundant step correction reward, effectively identifying redundant verification steps and improving the model's reasoning accuracy while suppressing overthinking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to methods, devices, media and products for optimizing models. Background Technology

[0002] As large reasoning models demonstrate superior performance in complex reasoning tasks such as mathematics and science, extending the thought chain to enhance reasoning ability has become a key technological approach in this field. However, large reasoning models commonly suffer from "overthinking," where the model continues to generate numerous unnecessary, repetitive, or irrelevant verification reasoning steps even after arriving at the correct answer. Related technologies typically address overthinking by directly controlling the length of the thought chain, such as by designing specific prompts to guide the model to generate a more concise thought chain, or by intervening during the model's reasoning process to terminate it early or dynamically control its length. However, these technologies cannot accurately control the reasoning trajectory, affecting the accuracy of large reasoning models. Therefore, how to improve reasoning accuracy while suppressing overthinking in large reasoning models is a pressing technical problem that needs to be solved. Summary of the Invention

[0003] This application provides methods, devices, media, and products for optimizing models, in order to at least address the problem in related technologies of how to improve reasoning accuracy while suppressing overthinking in large reasoning models.

[0004] This application provides a method for optimizing a model, including:

[0005] Obtain the inference dataset of the model to be processed; wherein, the inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question;

[0006] Based on the standard answer, identify the first node in the reasoning trajectory text that marks the location of the first correct answer and the second node that marks the boundary of subsequent reflection steps;

[0007] Based on the first and second nodes, determine the redundant steps in the reasoning trajectory text and correct the reward.

[0008] The reward is adjusted based on the redundant steps, and the model to be processed is trained and optimized to obtain the target model.

[0009] This application provides a model optimization method. For the model to be processed, an inference dataset of the model is pre-acquired. Each inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question. Based on the standard answer and the inference trajectory text, a first node is identified to mark the location of the first correct answer and a second node is identified to mark the boundary of subsequent reflection steps. When the model to be processed is trained and optimized, the first and second nodes are fused to generate a redundant step correction reward to suppress overthinking. This application combines the structured analysis of the model's inference trajectory with the design of a reward mechanism. By explicitly locating the first correct answer and the boundary of subsequent reflection steps, the trajectory of the answer discovery stage and the unnecessary post-answer verification stage are fused into the redundant step correction reward. Compared with related technologies that only focus on the length of the thought chain compression, this method can effectively identify redundant verification steps, thereby suppressing model overthinking and improving the model's inference accuracy.

[0010] This application also provides a reasoning method for the model, including:

[0011] Obtain the problem to be reasoned;

[0012] The question to be reasoned is input into the target model to trigger the target model to reason about the answer to the question; wherein, the target model is obtained based on the optimization method of any of the models mentioned above;

[0013] Determine the target answer based on the output of the target model.

[0014] This application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the optimization method of any of the above models.

[0015] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the optimization method of any of the above models.

[0016] This application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the optimization method for any of the above models.

[0017] This application combines the structured analysis of the model's reasoning trajectory with the design of a reward mechanism. By explicitly locating the boundaries of the first correct answer and subsequent reflection steps, the trajectory of the answer discovery stage and the unnecessary post-answer verification stage is integrated into the redundant step correction reward. This effectively identifies redundant verification steps, suppressing overthinking in the model while retaining effective verification steps. Therefore, it solves the technical problem in related technologies of how to improve reasoning accuracy while suppressing overthinking in large reasoning models, thus improving the model's reasoning accuracy. Attached Figure Description

[0018] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic diagram of an optimized system architecture for a model provided in an embodiment of this application;

[0020] Figure 2 A flowchart illustrating the optimization method for the model provided in this application embodiment. Figure One ;

[0021] Figure 3 A flowchart illustrating the optimization method for the model provided in this application embodiment. Figure Two ;

[0022] Figure 4 A flowchart illustrating the optimization method for the model provided in this application embodiment. Figure Three ;

[0023] Figure 5 A flowchart illustrating a model reasoning method provided in this application embodiment. Figure One ;

[0024] Figure 6 A schematic diagram of the structure of the optimization device for the model provided in the embodiments of this application;

[0025] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0027] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0028] First, let's explain the terms that appear in the embodiments of this application:

[0029] Large Language Model: A natural language processing model based on deep learning that can generate and understand human language.

[0030] Large reasoning models: Large reasoning models are a class of large language models specifically designed and optimized for complex reasoning tasks. The core features of large reasoning models lie in their significantly extended thought processes and enhanced step-by-step reasoning capabilities.

[0031] Thought Chain: A technique that enhances a model's reasoning ability by generating multi-step reasoning paths.

[0032] Reasoning trajectory: The reasoning trajectory refers to the complete dynamic process record of a large reasoning model when dealing with a complex problem, from the start of thinking to the final output of the answer.

[0033] Reward hacking: Reward hacking refers to a model discovering a flaw or incompleteness in the reward function and taking a series of strategies or behaviors that maximize the reward score but violate the designer's original intention and expected goals.

[0034] Cue words: Cue words are input instructions or text cues provided by the user to the model to guide the model to generate the expected output.

[0035] Overthinking is a common problem in large inference models, leading not only to a huge waste of computational resources and a significant increase in inference latency and cost, but more seriously, redundant reflection steps may introduce errors or interference, ultimately reducing the accuracy of the model's final answer. From the perspective of reinforcement learning, the root cause of overthinking lies in the mismatch of reward signals and the imperfection of credit allocation. Solutions to the overthinking problem in large inference models can be mainly divided into three categories: The first category is cue-based optimization methods. These methods guide the model to generate more concise thought chains by designing specific cue words. For example, by constraining the length of the model's output or emphasizing the conciseness of the reasoning. While simple to implement, these methods are highly dependent on the quality of the cue word design, and forced length constraints often negatively impact the reasoning depth of complex problems, leading to decreased accuracy and a lack of universality and stability. The second category is inference-time control methods. These methods intervene during the model's inference process to achieve early termination or dynamic control of the inference length. They typically rely on heuristic rules, such as monitoring the convergence of the inference state or the appearance of specific markers. However, these methods have significant limitations: heuristic rules are difficult to generalize and have poor adaptability to different tasks and model architectures; intervention only at the end of inference fails to fundamentally address the cause of overthinking, merely treating the symptoms, not the root cause; and there is a risk of errors interrupting effective thought chains, potentially harming model performance. The third category is post-training optimization methods. These methods can be further subdivided into supervised fine-tuning methods and reinforcement learning-based methods. Supervised fine-tuning utilizes datasets containing short, high-quality thought chains to instill efficient inference patterns, but its performance heavily depends on the size and quality of the collected dataset and is prone to overfitting to specific tasks or data distributions, exhibiting poor cross-task generalization ability. Reinforcement learning-based methods are more direct, typically employing the following strategies: first, introducing an explicit length penalty term to directly negatively incentivize long outputs in the reward function; second, designing a length-aware reward function that uses output length as a dimension for reward calculation; and third, constructing a complex step-by-step verification mechanism to allocate rewards to intermediate steps in the thought chain. However, these strategies share a common problem: simple length penalties or rewards may oversimplify the optimization objective and fail to distinguish between necessary deep reasoning and redundant reflection; while complex verification mechanisms are usually accompanied by high computational costs and low training efficiency, making them difficult to apply on a large scale.

[0036] Therefore, related technologies have significant shortcomings in controlling the inference trajectory. Most solutions only focus on compressing the output length, failing to distinguish between the solution discovery phase and the unnecessary post-answer verification phase at the trajectory level. Standard training objectives primarily reward the correctness of the final answer, lacking effective supervision on when the inference process should terminate. This results in subsequent inference steps generated after the model produces the first correct answer not being reasonably penalized, and may even accumulate positive rewards through spurious associations, forming reward hacking behavior, manifested as persistent overthinking. Therefore, how to improve inference accuracy while suppressing overthinking in large inference models is a pressing technical problem that needs to be solved.

[0037] To address the aforementioned technical issues, embodiments of this application provide a model optimization method, device, medium, and product. Based on the standard answer and the inference trajectory text output by the model to be processed, a first node for identifying the location of the first correct answer and a second node for identifying the boundaries of subsequent reflection steps are identified. When training and optimizing the model to be processed, the first and second nodes are fused to generate a redundant step correction reward for suppressing over-reflection, thereby achieving optimized training of the model to be processed.

[0038] Specifically, this application requires a mechanism that can explicitly model the internal structure of the reasoning trajectory, accurately identify the location of the first correct answer and the boundaries of subsequent reflection steps, and provide fine-grained reward signals accordingly, so as to fundamentally suppress overthinking and achieve an optimized balance between accuracy and computational efficiency.

[0039] It is understood that the embodiments of this application are mainly used to suppress overthinking in models based on thought chain reasoning. The solutions of the embodiments of this application have broad applicability, not limited to large inference models, but also extend to large-scale models that include multimodal information processing such as vision and hearing, including multimodal large models of large inference models, such as those for image reasoning, speech understanding, and video content analysis. For example, in image reasoning tasks, computation can be simplified by structurally recognizing the trajectories of key feature extraction; in speech understanding tasks, the semantic parsing node sequences can be optimized. By adapting the structured representation of inference trajectories under different modalities, the embodiments of this application can effectively solve the redundant computation problem commonly found in multimodal large models and improve their inference efficiency. Essentially, any model that adopts a thought chain or similar chain-based reasoning mechanism can apply the technical solutions of the embodiments of this application.

[0040] In some possible implementations, the core technical problem to be solved by the embodiments of this application can be systematically described in the following three aspects:

[0041] The core issue is the loss of control over the inference trajectory caused by reward signal mismatch: Reinforcement learning training paradigms primarily rely on the correctness of the final answer as the reward signal, a design fundamentally flawed. It cannot provide fine-grained supervision of the internal structure of the inference trajectory, causing the model to be unable to perceive when to stop inference. Specifically, after generating the first correct answer, subsequent redundant reflection steps cannot be effectively identified and penalized, and may even receive positive rewards through accidental correct associations, thus exacerbating reward hacking behavior. This application aims to solve this reward mismatch problem by designing a mechanism that can accurately identify key nodes in the inference process and provide differentiated reward signals accordingly. This fundamentally guides the model to form a moderate inference habit, achieving refined management and control of the inference trajectory.

[0042] The second core challenge is quantifying reflection steps due to the ambiguity of the reasoning process structure: Overthinking mainly manifests as a large amount of unnecessary verification reasoning after the answer appears. However, related technologies lack the ability to explicitly model the internal structure of the reasoning trajectory, making it impossible to clearly define the boundary between key reasoning and redundant reflection. This results in the inability to effectively quantify and penalize reflection steps. The technical challenge that this application aims to overcome is how to automatically and accurately locate the moment when the first correct answer appears in a lengthy thought chain without relying on heuristic rules and while ensuring high reliability, and to delineate the subsequent pure reflection interval accordingly. Only by achieving accurate location and counting of reflection steps can a reasonable inhibitory reward signal be designed for them.

[0043] The third core aspect addresses the balance between suppressing redundant reflection and maintaining necessary reasoning depth: Simply punishing long outputs or encouraging short outputs is a crude control strategy that may inadvertently impair the model's necessary and in-depth thinking process for complex problems. This can lead to the model sacrificing reasoning depth in pursuit of brevity, ultimately reducing its ability to solve difficult problems. Therefore, this application's embodiments need to address a key balance issue: how to design an intelligent reward mechanism that not only effectively suppresses redundant reflection after the answer but also fully encourages and retains the crucial reasoning steps necessary for problem-solving "before the answer." This requires the reward function to have context-aware capabilities, dynamically adjusting its tolerance for reflection based on the actual difficulty of the problem, ensuring that efficiency is improved without compromising the model's core ability to solve complex problems.

[0044] In summary, the embodiments of this application directly address the core pain points of related technologies in the optimization of large inference models, and are committed to building a new reward mechanism that can delve into the inference trajectory, distinguish between useful and redundant inference, and dynamically balance efficiency and performance, in order to fundamentally suppress overthinking and promote the development of large inference models towards a more efficient and reliable direction.

[0045] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] The specific application environment architecture or hardware architecture upon which the optimization method for the model depends is described here. (References) Figure 1 , Figure 1 This is a schematic diagram of an optimization system architecture for a model provided in an embodiment of this application. The optimization system for this model is a computer device. Figure 1 As shown, the above architecture includes at least one of a data acquisition device 101, a processing device 102, and a display device 103.

[0047] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the optimized system architecture of the model. In other feasible embodiments of this application, the above architecture may include more or fewer components than illustrated, or combine some components, or split some components, or arrange different components, which can be determined according to the actual application scenario and is not limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of both.

[0048] In the specific implementation process, the data acquisition device 101 may include an input / output interface or a communication interface, and the data acquisition device 101 can be connected to the processing device through the input / output interface or the communication interface.

[0049] The processing device 102 can identify a first node for marking the location of the first correct answer and a second node for marking the boundary of subsequent reflection steps based on the standard answer and the reasoning trajectory text output by the model to be processed. When training and optimizing the model to be processed, the first and second nodes are fused to generate a redundant step correction reward to suppress the phenomenon of over-reflection, thereby achieving optimized training of the model to be processed.

[0050] The display device 103 can also be a touch screen or the screen of a terminal device, used to receive user commands while displaying the above-mentioned content, so as to realize interaction with the user.

[0051] It should be understood that the aforementioned processing device can be implemented by a processor reading instructions from memory and executing those instructions, or it can be implemented by a chip circuit.

[0052] Furthermore, the network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0053] Figure 2 A flowchart illustrating the optimization method for the model provided in this application embodiment. Figure One ,like Figure 2 As shown, embodiments of this application provide a method for optimizing a model, which is described in detail below:

[0054] S201: Obtain the inference dataset for the model to be processed.

[0055] The inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question.

[0056] In one possible implementation, the model to be processed can be the large inference model described in the above embodiments or a multimodal large model including the large inference model. In specific applications, it can be any model that adopts a thought chain or similar chain-like inference mechanism.

[0057] In one possible implementation, the inference dataset is obtained by: collecting publicly available benchmark datasets of the MindChain framework, generating it using manual annotation, or generating it automatically without intervention from the model to be processed and then filtering it.

[0058] In one possible implementation, the specific method for obtaining the inference dataset of the model to be processed is as follows: First, obtain multiple original questions and their corresponding standard answers. The standard answers can be manually annotated or collected from the dataset. After obtaining the original questions, input them into the model to be processed to obtain the corresponding inference trajectory text.

[0059] In some possible implementations, to improve the accuracy and generalization ability of the inference dataset, one original question in the inference dataset of this application embodiment can correspond to multiple inference trajectory texts. That is, the inference dataset can be generated by sampling the original question multiple times for the model to be processed. Specifically, for the same original question, the model to be processed is controlled to perform multiple independent inferences under different random sampling parameter settings (such as higher temperature parameters and / or kernel sampling parameters) to obtain multiple different inference trajectory texts corresponding to the original question.

[0060] For example, in each training iteration of reinforcement learning, for each original problem, the sampling count can be set to 8 times, the temperature coefficient to 1.0, and the kernel sampling coefficient to 1.0, so that the model to be processed generates 8 inference trajectory texts. These 8 trajectories together constitute the inference dataset for this problem in this training round, which is used for subsequent redundant steps to correct reward calculations and policy updates.

[0061] It should be noted that the above-described sampling frequency and parameter settings are merely an exemplary implementation. Those skilled in the art can reasonably adjust the specific values ​​of the sampling frequency and sampling parameters based on factors such as the size of the model to be processed, the complexity of the task, and available computing resources. This application embodiment does not impose specific limitations in this regard.

[0062] Optionally, during partial sampling, different forms or levels of guidance cues can be provided to the model to be processed. For example, different styles of inference framework instructions or examples of some intermediate steps can be embedded in the cues to guide the model to generate different inference trajectories.

[0063] Optionally, noise can be selectively injected into the model input or during the generation process when generating partial inference trajectories. By adjusting the noise levels, the generalization ability of the inference dataset and the efficiency of model optimization can be improved.

[0064] In one possible implementation, the inference dataset can cover problems of different domains and difficulties to ensure the generalization ability of the trained reward mechanism.

[0065] In some embodiments, the original question may be at least one of text, image, or audio format, and correspondingly, the standard answer is not limited to text, image, or audio formats.

[0066] In some embodiments, the model to be processed can be applied to intelligent code generation. Accordingly, the original question and the corresponding standard answer are in text format. The optimization method of the model in this embodiment can suppress unnecessary redundant comments and repetitive validation logic after code generation, thereby improving development efficiency.

[0067] In some embodiments, the model to be processed can be applied to assist in medical diagnosis. Accordingly, the original question and the corresponding standard answer are in text and / or image format. Through the model optimization method in the embodiments of this application, repeated verification after the answer is determined is avoided in case analysis and image diagnosis reasoning, while ensuring the integrity of key diagnostic logic.

[0068] In some embodiments, the model to be processed can be applied to autonomous driving decision-making. Accordingly, the original question and the corresponding standard answer are at least one of text, image, or audio formats. The optimization method of the model in this application embodiment optimizes redundant decision verification steps in driving scenario reasoning, improves real-time response speed, and adapts to scenarios with limited edge computing resources.

[0069] S202: Based on the standard answer, identify the first node in the reasoning trajectory text that marks the location of the first correct answer and the second node that marks the boundary of subsequent reflection steps.

[0070] In one possible implementation, once the first node and the second node are determined, labels are added to the first node and the second node so that subsequent redundant steps can correct the reward calculation.

[0071] In one possible implementation, the label corresponding to the first node is <1st_answer>, which indicates the location where the first correct answer appears, and the label corresponding to the second node is <1st_answer>, which represents the boundary of subsequent reflection steps.

[0072] S203: Based on the first and second nodes, determine the redundant steps in the reasoning trajectory text and correct the reward.

[0073] S204: Adjust the reward based on the redundant steps, train and optimize the model to be processed to obtain the target model.

[0074] In one possible implementation, the step of correcting the reward based on redundant steps and training and optimizing the model to be processed to obtain the target model includes: determining a target policy optimization function for reinforcement learning based on the corrected reward based on redundant steps and a preset policy optimization algorithm; and training the model to be processed using reinforcement learning based on the target policy optimization function to obtain the target model.

[0075] The preset policy optimization algorithm can be the Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) algorithm. This algorithm is particularly suitable for incorporating a single but clear efficiency preference signal, such as a reward for redundancy step correction, into the model policy optimization, which can promote the model to stably learn efficient inference behavior. The preset policy optimization algorithm can also be any policy optimization algorithm, such as proximal policy optimization, and this application embodiment does not impose specific limitations.

[0076] This application combines redundant step correction rewards with a pre-defined policy optimization algorithm to construct a target policy optimization function for reinforcement learning. This function is then used to train the model, enabling it to directly and stably learn efficient reasoning strategies within the reinforcement learning framework. By explicitly guiding the model to optimize its internal reasoning path generation process through reward signals, redundant thinking is effectively suppressed. Furthermore, the model adaptively balances reasoning depth and efficiency during training, automatically evolving a simpler, more focused reasoning pattern without relying on manually designed complex heuristics. Ultimately, this results in a target model that improves both computational efficiency and answer accuracy.

[0077] In one possible implementation, the embodiments of this application employ a policy optimization algorithm based on reinforcement learning, specifically as follows:

[0078] Training setup: During training, each instance (each instance corresponding to the same original problem) is sampled multiple times (the number of samples is not specifically limited), and a high temperature (e.g., 1.0) and kernel sampling (top-p) coefficient (e.g., 1.0) are set to explore diverse inference paths. During optimization, traditional length penalties are disabled, and the inference efficiency is guided entirely by a redundant step correction reward mechanism.

[0079] Optimization Objective: Taking the DAPO algorithm framework as an example, the objective function for optimizing the reward substitution strategy involves correcting redundant steps. parameter, This refers to the reward for correcting redundant steps corresponding to the i-th reasoning trajectory.

[0080] After training, the model significantly reduced the proportion of reflective steps in its reasoning trajectory when solving complex reasoning problems. This achieved the goal of greatly improving reasoning efficiency while maintaining or enhancing accuracy.

[0081] This application provides a model optimization method. For the model to be processed, an inference dataset of the model is pre-acquired. Each inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question. Based on the standard answer and the inference trajectory text, a first node is identified to mark the location of the first correct answer and a second node is identified to mark the boundary of subsequent reflection steps. When the model to be processed is trained and optimized, the first and second nodes are fused to generate a redundant step correction reward to suppress overthinking. This application combines the structured analysis of the model's inference trajectory with the design of a reward mechanism. By explicitly locating the first correct answer and the boundary of subsequent reflection steps, the trajectory of the answer discovery stage and the unnecessary post-answer verification stage are fused into the redundant step correction reward. Compared with related technologies that only focus on the length of the thought chain compression, this method can effectively identify redundant verification steps, thereby suppressing model overthinking and improving the model's inference accuracy.

[0082] In one possible implementation, embodiments of this application achieve model optimization through a highly accurate node identification method; correspondingly, Figure 3 A flowchart illustrating the optimization method for the model provided in this application embodiment. Figure Two ,like Figure 3 As shown, the method includes:

[0083] S301: Obtain the inference dataset of the model to be processed.

[0084] The inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question.

[0085] S302: Segment the inference trajectory text to obtain multiple text fragments.

[0086] In one possible implementation, the inference trajectory text is segmented to obtain multiple text fragments, including: segmenting the inference trajectory text according to a preset character length to obtain multiple text fragments of a preset character length.

[0087] It is understood that the preset character length can be determined according to the actual situation, and this application embodiment does not impose specific restrictions on it.

[0088] This application's embodiments utilize a preset character length to perform regular segmentation of the inference trajectory text, efficiently and uniformly transforming a continuous, unstructured inference trajectory text stream into a series of independently processable text fragment units. This fixed-length segmentation strategy is simple to implement, computationally inexpensive, and avoids the errors and delays that may be introduced by relying on complex semantic understanding for sentence segmentation. It also provides a standardized input format for subsequent node recognition processing, ensuring the stability and repeatability of the processing flow.

[0089] In one possible implementation, embodiments of this application employ a fixed window for text segmentation. Since the original thought chain may have an irregular structure and be quite long, the reasoning trajectory text is first segmented into continuous, non-overlapping segments of a fixed length (e.g., 1000 characters). This design aims to maintain the reasoning order while achieving localized analysis, avoiding the accumulation of errors caused by long contexts.

[0090] S303: Based on the standard answer, perform recognition processing on the text fragment to determine whether the first node and the second node exist in the text fragment.

[0091] In one possible implementation, the text fragment is subjected to recognition processing based on the standard answer to determine whether a first node and a second node exist in the text fragment, including: performing a first recognition processing on the text fragment based on the standard answer to identify multiple nodes to be filtered in the text fragment; and performing a second recognition processing on the multiple nodes to be filtered to obtain the first node and the second node.

[0092] This application proposes a two-stage identification mechanism. By first broadly screening potential correct nodes, and then precisely identifying the first correct answer node and subsequent redundant boundaries, the identification process becomes hierarchical and refined. This method effectively reduces the difficulty of direct, one-step identification, improves the fault tolerance and robustness for analyzing complex and lengthy reasoning trajectories, ensures the accuracy of key node location, and provides a more reliable foundation for subsequent reward calculation.

[0093] In one possible implementation, the correct answer can be an answer that is exactly the same as the standard answer, or an answer whose semantic similarity to the standard answer is greater than a preset similarity threshold, or an answer that shares the same core keywords as the standard answer. Semantic similarity can be determined by calculating the cosine similarity of the text embedding vectors or using a natural language inference model; core keywords can be determined by extracting named entities, technical terms, or through algorithms from the standard answer. It is understood that the preset similarity threshold and the selection criteria for core keywords can be adjusted according to the specific task domain and accuracy requirements, and this application embodiment does not impose specific limitations on them.

[0094] In one possible implementation, a first identification process is performed on the text fragment based on the standard answer to determine multiple nodes to be filtered in the text fragment, including: performing a matching process between the text fragment and the standard answer to determine whether there is correct text in the text fragment corresponding to the standard answer; if there is correct text in the text fragment corresponding to the standard answer, then the end node of the correct text is determined as the node to be filtered.

[0095] This application's embodiments achieve precise identification of the positions of all correct answers appearing in the reasoning trajectory by using the standard answer as a benchmark and performing refined matching processing on the segmented text fragments. First, using the standard answer as the standard ensures the objectivity and accuracy of the identification results, avoiding semantic omissions caused by rule-based or keyword matching. Second, by locating the end node of the correct text and marking it as the second node to be filtered, the originally continuous and unstructured reasoning trajectory text is transformed into a discrete, indexable node sequence. No matter how many times the model repeatedly outputs the correct answer in the same reasoning trajectory, it can be completely captured, laying a data foundation for the accurate measurement of reflection reward.

[0096] In one possible implementation, the matching method can be based on information such as similarity and keywords, or it can use the model to be processed in the embodiments of this application for matching and identification. The embodiments of this application can also use external models for matching and identification; specifically, it can be any pre-trained artificial intelligence model, or an external model or annotation tool with strong reasoning capabilities.

[0097] In some embodiments, this application uses external large-scale models such as annotation tools to assist in annotation, which can explore the model's self-annotation mode or train a dedicated annotation model through transfer learning, thereby reducing annotation costs and latency and adapting to large-scale dataset scenarios.

[0098] In one possible implementation, a matching process is performed between the text fragment and the standard answer to determine whether there is correct text in the text fragment corresponding to the standard answer. This includes: obtaining a preset recognition model; inputting the original question, the standard answer, and the text fragment into the preset recognition model; and determining whether there is correct text in the text fragment corresponding to the standard answer based on the output of the preset recognition model.

[0099] This application's embodiments introduce an external large language model with strong semantic understanding and reasoning capabilities as a discriminator to undertake the core task of answer matching. It leverages the mature capabilities of readily available open-source large models, eliminating the need to retrain a dedicated model for answer matching, thus reducing computational costs. The model enables efficient and accurate recognition. It can be smoothly extended to more vertical fields such as code generation, scientific reasoning, and medical diagnosis without modifying the core algorithm logic.

[0100] In one possible implementation, the preset recognition model is configured with multiple prompt word templates, each corresponding to a task type. Accordingly, the original question, standard answer, and text fragment are input into the preset recognition model, including: determining the target prompt word template based on the task type of the original question; and inputting the original question, standard answer, and text fragment into the preset recognition model that calls the target prompt word template.

[0101] This application's embodiments differentiate prompt word templates based on the characteristics of reasoning trajectories for different task types and dynamically call them, enabling flexible and accurate recognition schemes for different task types, further improving recognition accuracy.

[0102] In one possible implementation, a second identification process is performed on multiple nodes to be screened to obtain a first node and a second node, including: determining the first-time appearance of the node to be screened and the non-first-time appearance of the node to be screened based on the positions of the multiple nodes to be screened in the inference trajectory text; determining the first-time appearance of the node to be screened as the first node; and determining the non-first-time appearance of the node to be screened as the second node.

[0103] By sorting the positions of all correct answers in chronological order, the system completes the crucial transformation from the first correct answer to the subsequent reflection nodes. No additional training or heuristic threshold settings are required, ensuring absolute accuracy in the transformation process and simplifying the operation.

[0104] In one possible implementation, this application embodiment performs structured analysis on the lengthy inference trajectory generated by the model, accurately locating two key nodes: the position where the first correct answer appears (marked as <1st_answer>) and the boundary of subsequent reflection steps (marked as). After completing the text segmentation, the specific identification method is as follows:

[0105] First, external model-assisted annotation is used: each text fragment, along with the original question and its standard answer, is input into an external model (which can be an annotation tool model) with strong reasoning capabilities. Using carefully designed prompt word templates, the external model is guided to determine whether the current fragment contains a correct answer statement. If a fragment contains n correct answer statements, the model outputs n corresponding statements, and appends a label to each statement containing a correct answer, where n is any positive integer.

[0106] The prompt word templates here can be flexibly set according to different task types. In this embodiment, the prompt word templates can be called from a pre-stored database or the prompt word templates can be received from user input. In one possible implementation, the prompt word templates are divided into "mathematical task templates" (as shown in Table 1) and "scientific task templates" (as shown in Table 2), which can be selected according to the task type information of the training data.

[0107] In one possible implementation, Table 1 is an illustrative table of the first type of prompt word templates provided in the embodiments of this application, and Table 2 is an illustrative table of the second type of prompt word templates provided in the embodiments of this application. Tables 1 and 2 are merely illustrative and do not affect the protection scope of the embodiments of this application.

[0108] Table 1. Schematic diagram of the first type of prompt word template

[0109]

[0110] Table 2. Schematic diagram of the second type of prompt word template.

[0111]

[0112] After label recognition is completed, a key label transformation is performed: once all segments are labeled, the earliest label appearing in the entire reasoning sequence is replaced with <1st_answer>, thus marking the first appearance of the correct solution during the reasoning process. This step clearly divides the reasoning trajectory into the necessary reasoning stage before answer discovery and the reflection and verification stage after answer discovery.

[0113] By labeling <1st_answer> and key nodes, the previously ambiguous reasoning trajectory is clearly divided into necessary reasoning stages and redundant reflection stages, transforming the model's reasoning process from a black box into a traceable and analyzable one. This structured reasoning trajectory not only helps developers locate problems in the model's reasoning (such as premature termination or excessive reflection), but also provides a clear basis for subsequent iterative optimization of the model (such as adjusting the reasoning strategy and optimizing reward parameters), further enhancing the engineering application value of large-scale reasoning models.

[0114] S304: Based on the first and second nodes, determine the redundant steps in the reasoning trajectory text to correct the reward.

[0115] S305: Adjust the reward based on the redundant steps, train and optimize the model to be processed to obtain the target model.

[0116] The implementation methods of steps S301 and S201 are the same, and the implementation methods of steps S304-S305 are the same as those of steps S203-S204, which will not be described in detail here.

[0117] This application embodiment segments the inference trajectory text, transforming it into multiple structured text fragments. These fragments are then identified based on the standard answer, enabling automated and accurate location of key positions in the inference process. Dividing the inference trajectory into independent segments of fixed length effectively avoids interference from information decay and error accumulation in long contexts, laying the foundation for large-scale parallel processing. While maintaining the inference order, it achieves localized analysis, avoiding error accumulation caused by long contexts, and further improving the inference accuracy of the optimized model.

[0118] In one possible implementation, embodiments of this application optimize the model by modifying the reward through multi-dimensional redundant steps; correspondingly, Figure 4 A flowchart illustrating the optimization method for the model provided in this application embodiment. Figure Three ,like Figure 4 As shown, the method includes:

[0119] S401: Obtain the inference dataset for the model to be processed.

[0120] The inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question.

[0121] S402: Based on the standard answer, identify the first node in the reasoning trajectory text that marks the location of the first correct answer and the second node that marks the boundary of subsequent reflection steps.

[0122] The implementation of steps S401-S402 can be found in steps S301-S303 of the above embodiments, and will not be repeated here.

[0123] S403: Based on the first node, determine the answer to the reasoning trajectory text and discover the reward.

[0124] In one possible implementation, determining the answer discovery reward based on the first node includes: if the reasoning trajectory text includes the first node, then determining the answer discovery reward as a first preset reward value; if the reasoning trajectory text does not include the first node, then determining the answer discovery reward as a second preset reward value.

[0125] The first preset reward value is a positive number; the second preset reward value is a negative number.

[0126] Answer discovery reward This indicates that the reward model successfully finds the first correct answer in this dimension. In one possible implementation, a penalty is applied if the <1st_answer> label is not present in the reasoning trajectory. If the marker is successfully marked, a reward will be given. ).

[0127] The method for determining the answer discovery reward in this embodiment establishes a clear and reinforcing learning signal mechanism by setting explicit positive and negative reward values. When the reasoning trajectory contains the first node, indicating a successful correct answer, a positive reward is given, directly motivating the model to learn and reproduce the key behavior of successfully finding the answer. Conversely, when the trajectory does not contain the first node, a negative reward is applied, explicitly penalizing the reasoning path that fails to reach a valid conclusion. This effectively incentivizes answer discovery behavior.

[0128] S404: Based on the second node, determine the reflection reward for the reasoning trajectory text.

[0129] In one possible implementation, the reflection reward for the reasoning trajectory text is determined based on the second node, including: determining the tolerance range of reflection steps for the original question corresponding to the reasoning trajectory text; determining the number of second nodes in the reasoning trajectory text as the number of reflection steps for the reasoning trajectory text; and determining the reflection reward based on the tolerance range of reflection steps and the number of reflection steps.

[0130] This application's embodiments introduce a dynamic standard—the tolerance range for reflective steps—combined with statistics on the actual number of reflective steps in the inference trajectory, to achieve adaptive and quantitative punishment for overthinking behavior in the model. This not only accurately measures the severity of redundant reflection but also allows for dynamic adjustment of the punishment scale based on the actual difficulty of the problem or the model's current capability level, making the reward mechanism more flexible and reasonable, and further improving the accuracy of model optimization.

[0131] In one possible implementation, determining the tolerance range of the reflection steps for the original question corresponding to the reasoning trajectory text includes: obtaining the empirical pass rate of the original question; wherein the empirical pass rate is calculated from the results of multiple reasonings of the model to be processed for the original question; if the empirical pass rate is not less than a preset pass rate threshold, then the tolerance range of the reflection steps is determined as the preset tolerance range of the reflection steps; if the empirical pass rate is less than the preset pass rate threshold, then the tolerance range of the reflection steps is determined based on at least one of the number of standard answers, the number of reasoning trajectory texts and the number of second nodes, and the preset tolerance range of the reflection steps.

[0132] This method of dynamically determining the tolerance range, by introducing empirical pass rates, achieves personalized and context-aware reward strategies. For problems that the model can stably solve, a pre-set strict tolerance range is used to suppress overthinking; for problems that the model is not yet proficient at, the tolerance range is appropriately relaxed or dynamically adjusted to avoid imposing inappropriate penalties when the model still needs to explore and experiment, thus hindering its learning of necessary deep reasoning. This enhances the intelligence and stability of the training process, ensuring that the reward mechanism promotes efficiency without compromising the model's ability to solve difficult problems.

[0133] In one possible implementation, the tolerance range for reflection steps includes a minimum tolerance threshold and a maximum tolerance threshold. Accordingly, the reflection count reward is determined based on the tolerance range and the number of reflection steps, including: if the number of reflection steps is not greater than the minimum tolerance threshold, the reflection count reward is determined to be a third preset reward value; wherein the third preset reward value is a positive number; if the number of reflection steps is greater than the minimum tolerance threshold but not greater than the maximum tolerance threshold, the reflection count reward is calculated based on the minimum tolerance threshold, the maximum tolerance threshold, and the number of reflection steps; if the number of reflection steps is greater than the maximum tolerance threshold, the reflection count reward is determined to be 0.

[0134] A three-stage reward calculation rule based on the tolerance range is constructed, creating a clear, continuous, and well-guided penalty gradient. Positive rewards are given when the number of reflection steps is minimal, encouraging concise reasoning; as the number of reflection steps remains within an acceptable range, the reward decreases with increasing steps, providing a smooth optimization signal; and when the number of reflection steps exceeds the maximum tolerance threshold, the reward is reset to zero. This structured reward function precisely guides the model to control the number of reflection steps within an ideal range, avoiding excessive penalties for minor reflections and suppressing serious redundant behavior.

[0135] It is understood that the first preset reward value, the second preset reward value, the preset pass rate threshold, and the third preset reward value can all be determined according to the actual situation, and the embodiments of this application do not impose specific restrictions.

[0136] S405: Rewards for discovering answers and reflection times, and rewards for correcting redundant steps in the reasoning trajectory text.

[0137] In one possible implementation, the redundant step correction reward for the reasoning trajectory text is determined based on the answer discovery reward and the reflection count reward, including: obtaining the reasoning result corresponding to the reasoning trajectory text; determining the correctness reward for the reasoning trajectory text based on the reasoning result and the standard answer; and determining the redundant step correction reward as the sum of the answer discovery reward, the reflection count reward, and the correctness reward.

[0138] The embodiments of this application enable the reward signal to comprehensively and evenly reflect the performance of the reasoning process in three key dimensions: result correctness, thinking agility, and path simplicity. This guides the model to collaboratively optimize multiple objectives during reinforcement learning training, ultimately producing reasoning behavior that is both accurate and efficient, effectively avoiding the problem of sacrificing accuracy in pursuit of saving computing resources.

[0139] In one possible implementation, the number of reflections is rewarded. This indicates that this dimension is central to curbing overthinking. This represents the number of labels in the trajectory (i.e., the number of reflection steps).

[0140] First, a reasonable boundary for the number of reflection steps is dynamically set based on the empirical pass rate p of the problem. For relatively easy problems ( Set a fixed tolerance range (minimum tolerance threshold). Maximum tolerance threshold For difficult problems () Then, based on the distribution of reflection steps observed in the correct samples, the number of reflection steps is dynamically determined. and This allows for more verification space.

[0141] Then, based on Relationship with boundary values, calculation :

[0142] when hour, (Reflection should be moderate, and no punishment will be imposed).

[0143] when hour, (Linearly decreasing penalty to suppress excessive reflection).

[0144] when hour, (Severely punish excessive reflection).

[0145] In special cases, if the model terminates immediately after providing the first answer (unlabeled), set... , This indicates that correct behavior should be rewarded, and that negative impacts on efficient behavior should be avoided.

[0146] On the other hand, if dynamically determined and This is also considered a situation of appropriate reflection, and is set up .

[0147] S406: Adjust the reward based on the redundant steps, train and optimize the model to be processed to obtain the target model.

[0148] This application's embodiments achieve a refined and structured evaluation of reasoning efficiency by integrating two independently measurable dimensions—answer discovery reward and reflection count reward—into the redundant step correction reward. The answer discovery reward directly incentivizes the model to quickly locate the correct answer, while the reflection count reward explicitly penalizes unnecessary repetition or divergent thinking after arriving at the answer. This makes the reward signal more interpretable and guiding, enabling more precise guidance for the model to optimize its reasoning strategy. On the one hand, it encourages efficient and direct attainment of the goal; on the other hand, it suppresses ineffective mental loops. Thus, during training, the model spontaneously seeks the optimal balance between reasoning speed and logical rigor, achieving precise suppression of excessive reflection.

[0149] Through the refined design of the three-dimensional reward function, a synergistic improvement in computing power and accuracy was achieved:

[0150] On the one hand, redundant reflection is suppressed and efficiency is improved by rewarding the number of reflections. On the other hand, effective reasoning is encouraged by rewarding the discovery of answers, and the reliability of the final correctness is ensured by rewarding the correctness of the result. At the same time, the reflection boundary is dynamically adjusted to adapt to the difficulty of the problem and avoids accidentally damaging the necessary reasoning process for complex problems. In the end, both efficiency and accuracy are improved, with the model's reasoning efficiency increasing while the accuracy is relatively improved.

[0151] In one possible implementation, the calculation process for the redundant step correction reward is as follows:

[0152] If the <1st_answer> marker does not appear in the inference trajectory, it indicates that the model's initial inference direction is incorrect, it failed to generate a valid conclusion, or the inference process is contradictory. In this case, [-1.0, -1.0, -1.0] is returned directly to mark the sample as invalid inference.

[0153] In one possible implementation, embodiments of this application also identify tags. Tags are used to identify structured delimiters indicating "end of reasoning phase, start of answer output".

[0154] If the reasoning trajectory ends with <1st_answer>, and the trajectory contains but does not contain, it indicates that the model did not perform subsequent reflection after initially obtaining the correct answer. In this case, the reasoning trajectory... The score, which is the reward for the number of reflections and the reward for the final correctness, is the return. If the inference path contains <1st_answer> but does not appear, it indicates that the output is incomplete and may be truncated before the final conclusion is generated, returning [1.0, -1.0, -1.0].

[0155] If the reasoning trajectory contains both <1st_answer> and , then the calculation proceeds normally. , , The score.

[0156] The reward for correcting redundant steps is obtained by summing the reward for finding the answer, the reward for reflecting on the frequency of reflection, and the reward for final accuracy.

[0157]

[0158] This application's embodiment protects the structured recognition technology of reasoning trajectory: through a three-level linkage mechanism of "fixed window text segmentation + external model-assisted annotation + key label conversion", it accurately locates the position of the first correct answer (<1st_answer>) and the boundary of the reflection step (), and for the first time realizes the explicit division of "necessary reasoning stage" and "redundant reflection stage" in the reasoning trajectory, solving the pain point of existing technologies being unable to quantify the reflection steps.

[0159] This application's embodiment protects the design of a three-dimensional dynamic reward function: innovatively constructing a joint reward system of "answer discovery reward + reflection count reward + final correctness reward", wherein the reflection count reward achieves precise suppression of redundant reflection through the "dynamic boundary delineation of problem pass rate + linear decreasing penalty" mechanism, while avoiding "accidental injury" to the necessary reasoning depth.

[0160] This application's embodiments protect the optimization and integration of reinforcement learning strategies: the redundant step correction reward mechanism is deeply embedded in the reinforcement learning framework, and through training settings such as multi-round sampling and high temperature coefficient exploration, the traditional length penalty is disabled, guiding the model to autonomously learn the reasoning mode of "efficient reasoning + knowing when to stop", forming a complete technical closed loop.

[0161] Furthermore, the embodiments of this application also have the following technical effects: low implementation cost: no need to rely on expensive dedicated hardware or complex annotation tools, the external model-assisted annotation process is simple and reusable, reducing training complexity; low integration difficulty: the technical closed loop is clear, and it can be directly embedded into the reinforcement learning training process of existing large models without major modifications to the model structure; large scalability potential: it avoids the problems of high computational cost and low training efficiency of complex verification mechanisms, and can support the training of large-scale datasets, meeting the mass production needs of industrial applications.

[0162] This application also provides a reasoning method for a model. Figure 5 A flowchart illustrating a model reasoning method provided in this application embodiment. Figure One ,like Figure 5 As shown, the method includes:

[0163] S601: Obtain the problem to be reasoned.

[0164] S602: Input the question to be reasoned into the target model to trigger the target model to reason about the answer to the question to be reasoned.

[0165] The target model is obtained based on the optimization method of the model in the above embodiments. The optimization methods of the model in the above embodiments are all applicable to the target model in the embodiments of this application.

[0166] S603: Determine the target answer based on the output of the target model.

[0167] In one possible implementation, after the target model outputs the target answer, a reasoning dataset is generated based on the question to be reasoned and the target answer, which is used for dynamic optimization of the target model, further improving the reasoning efficiency of the target model.

[0168] This application's embodiments achieve efficient, accurate, and cost-effective automated reasoning by deploying an optimized target model in a real-world question-answering scenario. Since the target model has been trained using the aforementioned optimization methods, it can initiate a concise and focused reasoning mode when faced with a new question. This not only inherits the accuracy guarantee of the target model but also improves the response speed of a single reasoning iteration and reduces computational resource consumption, enabling the large-scale application of high-quality, low-latency intelligent reasoning services in real-world products.

[0169] Figure 6 This is a schematic diagram of the structure of the optimization device for the model provided in an embodiment of this application. Figure 6 As shown, embodiments of this application also provide a model optimization device, which includes: a first acquisition module 701, a first determination module 702, a second determination module 703, and an optimization module 704.

[0170] The first acquisition module 701 is used to acquire the inference dataset of the model to be processed; wherein the inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question;

[0171] The first determining module 702 is used to determine, based on the standard answer, a first node in the reasoning trajectory text to identify the location where the first correct answer appears and a second node to identify the boundary of subsequent reflection steps;

[0172] The second determining module 703 is used to determine the redundant steps correction reward of the reasoning trajectory text based on the first node and the second node;

[0173] The optimization module 704 is used to correct the reward based on the redundant steps and to train and optimize the model to be processed in order to obtain the target model.

[0174] In one possible implementation, the first determining module 702 is specifically used for:

[0175] The inference trajectory text is segmented to obtain multiple text fragments;

[0176] Based on the standard answer, the text fragment is processed to identify whether a first node and a second node exist in the text fragment.

[0177] In one possible implementation, the first determining module 702 is specifically used for:

[0178] Based on the standard answer, the text fragment undergoes a first identification process to identify multiple nodes to be filtered within the text fragment;

[0179] A second identification process is performed on multiple nodes to be screened to obtain the first node and the second node.

[0180] In one possible implementation, the first determining module 702 is specifically used for:

[0181] The text fragments are matched with the standard answers to determine whether there is correct text in the text fragments that corresponds to the standard answers.

[0182] If a text fragment contains correct text corresponding to the standard answer, then the end node of the correct text is determined as the node to be filtered.

[0183] In one possible implementation, the first determining module 702 is specifically used for:

[0184] Obtain the preset recognition model;

[0185] Input the original question, standard answer, and text fragment into the preset recognition model;

[0186] Based on the output of the preset recognition model, determine whether there is correct text in the text fragment that corresponds to the standard answer.

[0187] In one possible implementation, the preset recognition model is configured with multiple prompt word templates, and each prompt word template corresponds to a task type.

[0188] Accordingly, the first determining module 702 is specifically used for:

[0189] Determine the target prompt word template based on the task type of the original question;

[0190] Input the original question, standard answer, and text fragment into the preset recognition model that calls the target prompt word template.

[0191] In one possible implementation, the first determining module 702 is specifically used for:

[0192] Based on the positions of multiple nodes to be filtered in the inference trajectory text, determine the first occurrence of the node to be filtered and the non-first occurrence of the node to be filtered;

[0193] The first node to be filtered is designated as the first node;

[0194] Nodes that are not appearing for the first time are identified as the second node.

[0195] In one possible implementation, the first determining module 702 is specifically used for:

[0196] Based on a preset character length, the inference trajectory text is segmented to obtain multiple text fragments of the preset character length.

[0197] In one possible implementation, the second determining module 703 is specifically used for:

[0198] Based on the first node, determine the answer to the reasoning trajectory text and discover the reward;

[0199] Based on the second node, determine the reward for the number of reflections on the reasoning trajectory text;

[0200] Rewards are given based on the number of answers found and the number of reflections, as well as rewards for correcting redundant steps in the reasoning trajectory text.

[0201] In one possible implementation, the second determining module 703 is specifically used for:

[0202] If the reasoning trajectory text includes the first node, then the reward for finding the answer is determined to be the first preset reward value; where the first preset reward value is a positive number.

[0203] If the reasoning trajectory text does not include the first node, then the reward for finding the answer is determined to be the second preset reward value; where the second preset reward value is a negative number.

[0204] In one possible implementation, the second determining module 703 is specifically used for:

[0205] Determine the tolerance range of the reflection steps for the original question corresponding to the reasoning trajectory text;

[0206] The number of second nodes in the reasoning trajectory text is determined as the number of reflection steps in the reasoning trajectory text;

[0207] The reward for the number of reflections is determined based on the tolerance range and the number of reflection steps.

[0208] In one possible implementation, the second determining module 703 is specifically used for:

[0209] Obtain the empirical pass rate of the original problem; whereby the empirical pass rate is calculated from the results of multiple inferences by the model to be processed for the original problem;

[0210] If the empirical pass rate is not less than the preset pass rate threshold, then the tolerance range of the reflection step is determined to be the preset tolerance range of the reflection step.

[0211] If the empirical pass rate is less than the preset pass rate threshold, the tolerance range of the reflection steps is determined based on at least one of the number of standard answers, reasoning trajectory texts, and second nodes, as well as the preset tolerance range of the reflection steps.

[0212] In one possible implementation, the tolerance range of the reflection step includes a minimum tolerance threshold and a maximum tolerance threshold; correspondingly, the second determining module 703 is specifically used for:

[0213] If the number of reflection steps is not greater than the minimum tolerance threshold, then the reflection count reward is determined to be the third preset reward value; where the third preset reward value is a positive number.

[0214] If the number of reflection steps is greater than the minimum tolerance threshold but not greater than the maximum tolerance threshold, then the reflection count bonus is calculated based on the minimum tolerance threshold, the maximum tolerance threshold, and the number of reflection steps.

[0215] If the number of reflection steps exceeds the maximum tolerance threshold, then the reflection reward is set to 0.

[0216] In one possible implementation, the second determining module 703 is further specifically used for:

[0217] Determine whether the reasoning trajectory text contains either a first-type or a second-type reflection phenomenon; wherein, the judgment condition for the first-type reflection phenomenon is that the number of reflection steps is 0, and the judgment condition for the second-type reflection phenomenon is that the highest tolerance threshold is not greater than the lowest tolerance threshold in the preset tolerance range of reflection steps and the number of reflection steps is not greater than the highest tolerance threshold.

[0218] If the reasoning trajectory text exhibits a first-type reflection phenomenon, then the reflection count reward is determined to be the correctness reward of the reasoning trajectory text;

[0219] If the reasoning trajectory text exhibits a second type of reflection phenomenon, then the reward for the number of reflections is determined to be the third preset reward value.

[0220] In one possible implementation, the second determining module 703 is further specifically used for:

[0221] Obtain the reasoning result corresponding to the reasoning trajectory text;

[0222] Based on the reasoning results and the standard answer, a reward is given for the correctness of the reasoning trajectory text.

[0223] The sum of the reward for finding the answer, the reward for reflecting on the number of times, and the reward for correctness is determined as the reward for correcting redundant steps.

[0224] In one possible implementation, optimization module 704 is specifically used for:

[0225] Based on the redundant steps, the reward and the preset policy optimization algorithm are corrected to determine the target policy optimization function for reinforcement learning;

[0226] Based on the objective policy optimization function, the model to be processed is trained by reinforcement learning to obtain the objective model.

[0227] This application also provides a model inference device, which includes: a second acquisition module, an input module and a third determination module.

[0228] The second acquisition module is used to acquire the problem to be reasoned.

[0229] The input module is used to input the question to be reasoned into the target model, so as to trigger the target model to reason about the answer to the question; wherein, the target model is obtained based on the optimization method of the model in the above embodiment;

[0230] The third determining module is used to determine the target answer based on the output of the target model.

[0231] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Figure 7 As shown, the electronic device 80 provided in this embodiment includes at least one processor 801 and a memory 802. In one possible implementation, the electronic device 80 further includes a communication component 803. The processor 801, memory 802, and communication component 803 are connected via a bus.

[0232] In a specific implementation, at least one processor 801 executes computer execution instructions stored in memory 802, causing at least one processor 801 to execute the above-described model optimization or model inference method embodiments.

[0233] The specific implementation process of processor 801 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0234] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0235] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0236] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0237] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in the embodiments of the optimization or inference methods of any of the above-described models when it is run.

[0238] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0239] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the above-described optimization or inference method embodiments of any of the models.

[0240] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the above-described optimization or inference method embodiments of any of the models.

[0241] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.

[0242] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0243] The above provides a detailed description of the optimization or inference of a model provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for optimizing a model, characterized in that, include: Obtain the inference dataset of the model to be processed; wherein, the inference dataset includes the original question, the standard answer to the original question, and the inference trajectory text of the model to be processed for the original question; Based on the standard answer, determine a first node in the reasoning trajectory text to identify the location where the first correct answer appears and a second node to identify the boundary of subsequent reflection steps; Based on the first node and the second node, determine the redundant step correction reward of the reasoning trajectory text; The reward is corrected based on the redundant steps, and the model to be processed is trained and optimized to obtain the target model.

2. The method according to claim 1, characterized in that, The step of determining, based on the standard answer, a first node in the reasoning trajectory text to identify the location of the first correct answer and a second node to identify the boundary of subsequent reflection steps include: The inference trajectory text is segmented to obtain multiple text fragments; Based on the standard answer, the text fragment is processed for identification to determine whether the first node and the second node exist in the text fragment.

3. The method according to claim 2, characterized in that, The step of performing recognition processing on the text fragment based on the standard answer to determine whether the first node and the second node exist in the text fragment includes: Based on the standard answer, the text fragment undergoes a first recognition process to identify multiple nodes to be filtered within the text fragment; A second identification process is performed on multiple nodes to be screened to obtain the first node and the second node.

4. The method according to claim 3, characterized in that, The first identification process, based on the standard answer, is performed on the text fragment to identify multiple nodes to be filtered within the text fragment, including: The text fragment is matched with the standard answer to determine whether there is correct text in the text fragment that corresponds to the standard answer. If the text fragment contains correct text corresponding to the standard answer, then the end node of the correct text is determined as the node to be filtered.

5. The method according to claim 4, characterized in that, The step of matching the text fragment with the standard answer to determine whether there is correct text in the text fragment corresponding to the standard answer includes: Obtain the preset recognition model; The original question, the standard answer, and the text fragment are input into the preset recognition model; Based on the output of the preset recognition model, determine whether there is correct text in the text fragment that corresponds to the standard answer.

6. The method according to claim 5, characterized in that, The preset recognition model is configured with multiple prompt word templates, and each prompt word template corresponds to a task type. Accordingly, inputting the original question, the standard answer, and the text fragment into the preset recognition model includes: Based on the task type of the original question, determine the target prompt word template; The original question, the standard answer, and the text fragment are input into the preset recognition model that invokes the target prompt word template.

7. The method according to claim 3, characterized in that, The second identification process for multiple nodes to be screened, to obtain the first node and the second node, includes: Based on the positions of the multiple nodes to be filtered in the inference trajectory text, the first-time appearance of the nodes to be filtered and the non-first-time appearance of the nodes to be filtered are determined; The first node to be filtered is identified as the first node; The non-first-time occurrence of the node to be screened is determined as the second node.

8. The method according to claim 2, characterized in that, The inference trajectory text is segmented to obtain multiple text fragments, including: The inference trajectory text is segmented according to a preset character length to obtain multiple text fragments of the preset character length.

9. The method according to any one of claims 1 to 8, characterized in that, The step of determining the redundant step correction reward of the inference trajectory text based on the first node and the second node includes: Based on the first node, the answer discovery reward is determined from the reasoning trajectory text; Based on the second node, determine the reflection count reward for the reasoning trajectory text; Based on the reward for finding the answer and the reward for the number of reflections, the redundant step correction reward for the reasoning trajectory text is determined.

10. The method according to claim 9, characterized in that, The step of determining the answer discovery reward for the reasoning trajectory text based on the first node includes: If the reasoning trajectory text includes the first node, then the answer discovery reward is determined to be a first preset reward value; wherein, the first preset reward value is a positive number; If the reasoning trajectory text does not include the first node, then the answer discovery reward is determined to be a second preset reward value; wherein, the second preset reward value is a negative number.

11. The method according to claim 9, characterized in that, The step of determining the reflection count reward for the reasoning trajectory text based on the second node includes: Determine the tolerance range of the reflection steps for the original question corresponding to the reasoning trajectory text; The number of the second node in the reasoning trajectory text is determined as the number of reflection steps in the reasoning trajectory text; The reward for the number of reflections is determined based on the tolerance range of the reflection steps and the number of reflection steps.

12. The method according to claim 11, characterized in that, The tolerance range of the reflection steps for determining the original question corresponding to the reasoning trajectory text includes: Obtain the empirical pass rate of the original problem; wherein, the empirical pass rate is calculated from the results of multiple inferences by the model to be processed for the original problem; If the empirical pass rate is not less than a preset pass rate threshold, then the tolerance range of the reflection step is determined to be the preset tolerance range of the reflection step. If the empirical pass rate is less than the preset pass rate threshold, then the reflection step tolerance range is determined based on at least one of the standard answer, the reasoning trajectory text, and the number of second nodes, as well as the preset reflection step tolerance range.

13. The method according to claim 12, characterized in that, The tolerance range of the reflection steps includes a minimum tolerance threshold and a maximum tolerance threshold; Accordingly, determining the reflection count reward based on the tolerance range of the reflection steps and the number of reflection steps includes: If the number of reflection steps is not greater than the minimum tolerance threshold, then the reflection count reward is determined to be a third preset reward value; wherein, the third preset reward value is a positive number. If the number of reflection steps is greater than the minimum tolerance threshold but not greater than the maximum tolerance threshold, then the reflection count reward is calculated based on the minimum tolerance threshold, the maximum tolerance threshold, and the number of reflection steps. If the number of reflection steps is greater than the maximum tolerance threshold, then the reflection count reward is determined to be 0.

14. The method according to claim 13, characterized in that, The step of determining the reflection count reward based on the tolerance range of the reflection steps and the number of reflection steps further includes: Determine whether the reasoning trajectory text contains a first type of reflection phenomenon or a second type of reflection phenomenon; wherein, the judgment condition for the first type of reflection phenomenon is that the number of reflection steps is 0, and the judgment condition for the second type of reflection phenomenon is that the highest tolerance threshold is not greater than the lowest tolerance threshold in the preset reflection step tolerance range and the number of reflection steps is not greater than the highest tolerance threshold. If the reasoning trajectory text exhibits the first type of reflection phenomenon, then the reflection count reward is determined to be the correctness reward of the reasoning trajectory text; If the reasoning trajectory text exhibits the second type of reflection phenomenon, then the reflection count reward is determined to be the third preset reward value.

15. The method according to claim 9, characterized in that, The step of determining the redundant step correction reward of the reasoning trajectory text based on the reward for discovering the answer and the reward for the number of reflections includes: Obtain the reasoning result corresponding to the reasoning trajectory text; Based on the reasoning results and the standard answer, determine the correctness reward of the reasoning trajectory text; The sum of the answer discovery reward, the reflection count reward, and the correctness reward is determined as the redundant step correction reward.

16. The method according to any one of claims 1 to 8, characterized in that, The step of correcting the reward based on the redundant steps and training and optimizing the model to be processed to obtain the target model includes: Based on the redundant steps, the reward is corrected and the preset policy optimization algorithm is used to determine the target policy optimization function for reinforcement learning; The model to be processed is trained by reinforcement learning according to the target policy optimization function to obtain the target model.

17. A reasoning method for a model, characterized in that, include: Obtain the problem to be reasoned; The question to be reasoned is input into the target model to trigger the target model to reason about the answer to the question to be reasoned; wherein the target model is obtained based on the optimization method of the model as described in any one of claims 1 to 16; The target answer is determined based on the output of the target model.

18. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the optimization method for the model as claimed in any one of claims 1 to 16.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the optimization method for the model as described in any one of claims 1 to 16.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the optimization method for the model as described in any one of claims 1 to 16.