Large model reasoning capability optimization method and system and storage medium

By introducing a self-game mechanism into the big model and using supervised fine-tuning data sets for iterative optimization, the problem of low efficiency of existing large model inference ability optimization methods is solved, and more efficient inference ability improvement and generalization ability enhancement is achieved.

CN119940485AInactive Publication Date: 2025-05-06XIAMEN YUANTING INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510413458.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing large-model inference capability optimization methods are inefficient and rely on a large amount of manual annotation data and complex human feedback.

Method used

The self-game mechanism is introduced to allow the big model to conduct confrontation training with itself, and by initializing the main model and the opponent model, iterative optimization is performed using supervised fine-tuning datasets until the main model converges.

Benefits of technology

It effectively reduces data dependence and computing resource consumption, improves the inference ability and generalization ability of large models, and improves the efficiency of inference ability optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940485A_ABST
    Figure CN119940485A_ABST
Patent Text Reader

Abstract

The invention discloses a large model reasoning capability optimization method and system and a storage medium, and the method comprises the steps: S10, initializing a to-be-fine-tuned large language model, and taking the to-be-fine-tuned large language model as an initial version of a main model and an opponent model; s20, obtaining prompt information and question content from the supervision fine tuning data set, inputting the prompt information and the question content into the opponent model, and generating a corresponding opponent model response; s30, optimizing weight parameters in the master model by minimizing a logic loss function through a first preset formula by utilizing a real response in the supervised fine tuning data set, and training the master model to distinguish an opponent model response from the real response; s40, maximizing an evaluation value of the main model to the generation response through a second preset formula so as to update a weight parameter in the opponent model; a regularization item is introduced into the second preset formula; s50, taking the trained main model as a new opponent model to replace the current opponent model; and S60, repeating the steps S20 to S50 until the main model converges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large model reasoning capability optimization method, system and storage medium. Background Art

[0002] Large models can demonstrate excellent performance in a variety of tasks through large-scale pre-training and task fine-tuning. At present, the fine-tuning technology of large models mainly relies on two methods: supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF). Among them, the typical implementation of RLHF is proximal policy optimization (PPO). The core goal of these technologies is to improve model performance by expanding the data scale.

[0003] Specifically, SFT relies on a large amount of manually annotated high-quality data to improve its reasoning ability by fine-tuning model parameters; while RLHF requires two key components: the Actor model and the Critic model. Taking the natural language generation task as an example, the Actor is a text generation model responsible for generating the corresponding response based on the input prompt information; the Critic is a neural network used to evaluate the quality of the response generated by the Actor and output a state value (usually a floating point number). Through the continuous interaction between the Actor and the Critic, the model's capabilities are improved, but this interaction process often leads to a decrease in fine-tuning efficiency.

[0004] That is, existing methods for optimizing the reasoning capabilities of large models are inefficient. Summary of the invention

[0005] To solve the above problems, the present invention provides a method for optimizing the reasoning capability of a large model, which comprises the following steps: S10, initializing the large language model to be fine-tuned and using it as the initial version of the main model and the opponent model; S20, obtaining prompt information and question content from the supervised fine-tuning dataset, and inputting them into the opponent model to generate a corresponding opponent model response; S30, using the real responses in the supervised fine-tuning dataset, optimizing the weight parameters in the main model by minimizing the logistic loss function through a first preset formula, and training the main model to distinguish between the opponent model response and the real response; S40, maximizing the evaluation value of the main model for the generated response through a second preset formula to update the weight parameter in the opponent model; a regularization term is introduced into the second preset formula, and the evaluation value is specifically the probability or confidence that the main model determines that a certain response is a true response; S50, using the trained main model as a new opponent model to replace the current opponent model; S60, repeating steps S20 to S50 until the main model converges.

[0006] Optional, a large language model to be fine-tuned, specifically loading the BPE tokenizer based on the Transformer architecture and initializing the Zephyr-7B-SFT-Full model weights.

[0007] Optionally, the hyperparameters of the Zephyr-7B-SFT-Full model include at least the learning rate, optimizer, number of attention heads, number of hidden layers, vocabulary size, and data precision.

[0008] Optionally, in S20, the prompt information is concatenated with the question content and then input into the opponent model in json format.

[0009] Optionally, the first preset formula is as follows: The first preset formula is as follows: ; in, is the main model, f is a highly expressive function, is a series of highly expressive function classes, is the number of iterations, E is the evaluation value, For prompt information, For a true response, is the opponent model response, is the logistic loss function, specifically a monotonically decreasing and convex loss function, and .

[0010] Optionally, the regularization term is specifically a KL regularization term.

[0011] Optionally, the second preset formula is as follows: ; in, The probability of generating a response for the adversary model that is indistinguishable from the primary model, Represents the evaluation value gap between the response generated by the opponent model and the true response. The evaluation value is the distribution Calculated, represents the distribution of the adversary model input, Indicates that the true response follows the true response distribution, is the true response distribution, Indicates that the opponent model response follows the opponent model response distribution, is the opponent model response distribution, Indicates that the true response obeys the initial data distribution to be fine-tuned, is the evaluation value of the main model, λ is the regularization parameter, and >0, is the expectation under the distribution of the adversary model input, is the KL regularization term, is the initial data distribution to be fine-tuned, is the opponent model response distribution at the tth iteration.

[0012] Optionally, in S60, the main model is saved during each training iteration; The large model reasoning capability optimization method further includes S70, after the main model converges, using the SFT data set to evaluate the saved main model, and selecting the optimized main model according to the evaluation result.

[0013] Corresponding to the large model reasoning capability optimization method, the present invention provides a large model reasoning capability optimization system, which includes: The initialization module is used to initialize the large language model to be fine-tuned and use it as the initial version of the main model and the opponent model; The opponent model response generation module is used to perform the opponent model response generation step: obtain the prompt information and question content from the supervised fine-tuning dataset, input them into the opponent model, and generate the corresponding opponent model response; A main model training module, configured to perform a main model training step: using the real responses in the supervised fine-tuning dataset, optimizing the weight parameters in the main model by minimizing the logistic loss function by a first preset formula, and training the main model to distinguish between the opponent model responses and the real responses; The opponent model updating module is used to execute the opponent model updating step: maximizing the evaluation value of the main model for the generated response through the second preset formula to update the weight parameters in the opponent model; the regularization term is introduced into the second preset formula, and the evaluation value is specifically the probability or confidence that the main model judges that a certain response is a true response; The iterative module is used to execute the replacement step: using the trained main model as the new opponent model to replace the current opponent model; and iteratively executing the opponent model response generation step, the main model training step, the opponent model update step and the replacement step in sequence until the main model converges.

[0014] In addition, to achieve the above-mentioned objectives, the present invention also provides a computer-readable storage medium, on which a large model reasoning capability optimization program is stored. When the large model reasoning capability optimization program is executed by a processor, the steps of the large model reasoning capability optimization method described above are implemented.

[0015] The present invention introduces a self-game mechanism to allow the large model to conduct adversarial training with itself, without relying on a large amount of manually labeled data and complex human feedback, effectively reducing data dependence and computing resource consumption. At the same time, by continuously iteratively optimizing the main model and the opponent model, the reasoning ability and generalization ability of the large model are improved, and the optimization efficiency of the large model's reasoning ability is improved.

[0016] The present invention loads the BPE word segmenter based on the Transformer architecture and initializes the weights of the Zephyr-7B-SFT-Full model, which provides a good initial performance and a stable training foundation for the model, and helps to improve the efficiency and effect of subsequent training.

[0017] The present invention splices prompt information with question content and then inputs the result into the opponent model in JSON format, which makes it easier for the model to understand and process input data and improves the efficiency and accuracy of data input.

[0018] The present invention optimizes the weight parameters in the main model by minimizing the logical loss function through the first preset formula, so that the main model can more accurately distinguish between generated responses and real responses, thereby improving the recognition ability and reasoning performance of the model.

[0019] The present invention introduces the KL regularization term as a regularization term to prevent the opponent model parameters from excessive deviation, maintain the stability and consistency of the model, and avoid instability that may occur during the training process.

[0020] The present invention maximizes the evaluation value of the main model for the generated response through the second preset formula, further optimizes the parameters of the opponent model, makes the generated response closer to the real response, and improves the generation ability of the model and the adversarial training effect.

[0021] The present invention saves the main model during each iterative training and uses the SFT data set for evaluation after the main model converges, thereby ensuring the stability and reliability of the model, being able to select the optimal model based on the evaluation results, and improving the practical application effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1A simplified flow chart of an embodiment of a large model reasoning capability optimization method of the present invention; Figure 2 It is a schematic diagram of the large model reasoning ability optimization process of the present invention; Figure 3 A framework diagram of an embodiment of a large model reasoning capability optimization system of the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0024] like Figure 1 As shown, a large model reasoning capability optimization method of the present invention comprises the following steps: S10, initializing the large language model to be fine-tuned and using it as the initial version of the main model and the opponent model; S20, obtaining prompt information and question content from the supervised fine-tuning dataset, and inputting them into the opponent model to generate a corresponding opponent model response; S30, using the real responses in the supervised fine-tuning dataset, optimizing the weight parameters in the main model by minimizing the logistic loss function through a first preset formula, and training the main model to distinguish between the opponent model response and the real response; S40, maximizing the evaluation value of the main model for the generated response through a second preset formula to update the weight parameter in the opponent model; a regularization term is introduced into the second preset formula, and the evaluation value is specifically the probability or confidence that the main model determines that a certain response is a true response; S50, using the trained main model as a new opponent model to replace the current opponent model; S60, repeating steps S20 to S50 until the main model converges.

[0025] The present invention introduces a self-game mechanism to allow the large model to conduct adversarial training with itself, without relying on a large amount of manually labeled data and complex human feedback, effectively reducing data dependence and computing resource consumption. At the same time, by continuously iteratively optimizing the main model and the opponent model, the reasoning ability and generalization ability of the large model are improved, and the optimization efficiency of the large model's reasoning ability is improved.

[0026] In this embodiment, the large language model to be fine-tuned specifically loads the BPE word segmenter based on the Transformer architecture and initializes the Zephyr-7B-SFT-Full model weights.

[0027] The present invention loads the BPE word segmenter based on the Transformer architecture and initializes the weights of the Zephyr-7B-SFT-Full model, which provides a good initial performance and a stable training foundation for the model, and helps to improve the efficiency and effect of subsequent training.

[0028] In this embodiment, the hyperparameters of the Zephyr-7B-SFT-Full model include at least learning rate, optimizer, number of attention heads, number of hidden layers, vocabulary size, and data accuracy.

[0029] Preferably, learning rate: 2e-5, using Adam optimizer and cosine learning rate scheduling, number of attention heads: 32, number of hidden layers: 32, vocabulary size: 32000, data precision: bfloat16.

[0030] In this embodiment, in S20, the prompt information and the question content are concatenated and input into the opponent model in json format.

[0031] The present invention splices prompt information with question content and then inputs the result into the opponent model in JSON format, which makes it easier for the model to understand and process input data and improves the efficiency and accuracy of data input.

[0032] In this embodiment, when S30 trains the main model, it is to maximize the true response distribution of the main model. and the opponent model response distribution The difference in evaluation value between: (Formula 1); in, For the main model, is the number of iterations, f is a highly expressive function, and E is the evaluation value, which is the distribution Calculated, represents the distribution of the adversary model input, Indicates that the true response follows the true response distribution, is the true response distribution, Indicates that the opponent model response follows the opponent model response distribution, is a series of highly expressive function classes, which will be determined in subsequent derivations. Since the function class depends on the opponent model response distribution , so in The subscript t is added. For prompt information, For a true response, Respond to the opponent model. The value of (i.e., the evaluation value) reflects the main model's response source is the true response distribution Rather than the opponent model response distribution Ideally, when When , the main model should give a high evaluation value, and when , a low evaluation value should be given. Furthermore, the weight parameters in the main model should be optimized by minimizing the logistic loss function through the first preset formula below.

[0033] In this embodiment, the first preset formula is as follows: (Formula 2) in, is the logistic loss function, specifically a monotonically decreasing and convex loss function, and .

[0034] The present invention optimizes the weight parameters of the main model by minimizing the logical loss function through the first preset formula, so that the main model can more accurately distinguish between the opponent model response and the real response, thereby improving the recognition ability and reasoning performance of the model.

[0035] In this embodiment, the regularization term is specifically a KL regularization term.

[0036] The present invention introduces the KL regularization term as a regularization term to prevent the opponent model parameters from excessive deviation, maintain the stability and consistency of the model, and avoid instability that may occur during the training process.

[0037] Updating the opponent model is to obtain The weight parameters of the opponent model for the iteration , it can be understood that the weight parameters can also be referred to as parameters for short. When faced with two responses (the opponent model response and the true response) to the same prompt information x, the main model evaluates the values ​​of the opponent model response and the true response. Then, it is inferred that the response with a higher evaluation value comes from the true response distribution, while the response with a lower evaluation value is attributed to the opponent model. Subsequently, the opponent model aims to find a better large language model so that the main model cannot distinguish between the two responses, which is specifically achieved by maximizing the evaluation value of the main model for the generated response through the second preset formula, and updating the parameters of the opponent model. In order to prevent the opponent model of the t+1 iteration from deviating too much from the opponent model of the t iteration and to stabilize the adversarial process, the present invention adds a KL regularization term. Therefore, the second preset formula is as follows: (Formula 3); in, The probability of generating a response for the adversary model that is indistinguishable to the primary model (i.e., the primary model cannot distinguish the adversary model response from the adversary model response distribution), It represents the evaluation value gap between the response generated by the opponent model and the real response. It is hoped that the gap between the two will be minimized, so as to better train the main model to improve its discrimination ability, so that the main model can achieve better results. The evaluation value is a distribution of Calculated, represents the distribution of the adversary model input, Indicates that the true response follows the true response distribution, is the true response distribution, Indicates that the opponent model response follows the opponent model response distribution, is the opponent model response distribution, Indicates that the true response obeys the initial data distribution to be fine-tuned, is the evaluation value of the main model, λ is the regularization parameter, and >0, is the expectation under the distribution of the adversary model input, is the KL regularization term, is the initial data distribution to be fine-tuned, is the opponent model response distribution at the tth iteration.

[0038] The above equation 3 has a closed solution : (Formula 4); in, It refers to a closed solution. is the only closed solution, The closed solution for the opponent model.

[0039] It should be noted that Not necessarily in the large language model parameter space ; is the probability space of parameter θ. Since we hope that the closed solution in the probability space This can be achieved by a large language model with parameters θ, i.e. , solve ∝ It turns out that: ,in, is the probability distribution of the large language model, is the opponent model distribution for the tth iteration.

[0040] so, Select the function class for: (Formula 5); in is the large language model parameter space considered, given the above formula Select. Optimize formula 2 to get Depend on Parameter changes are in the following form: (Formula 6); Substituting equation 6 into equation 4 yields: , is a closed solution with parameter θ, learned from Equation 2 This is the optimization target of the opponent model (optimal large language model parameters), that is, at this time the opponent model weight parameter update is completed.

[0041] The present invention maximizes the evaluation value of the main model's response to the opponent model through the second preset formula, further optimizes the weight parameters of the opponent model, makes the opponent model response closer to the real response, and improves the model's generation ability and adversarial training effect.

[0042] In this embodiment, if the logistic loss function of the main model no longer decreases significantly and basically tends to a stable state in S50, it can be considered that the main model has converged.

[0043] In this embodiment, the results of S30 and S40 can be integrated into an end-to-end training objective. Specifically, Formula 4 is substituted into Formula 2 to obtain the following update rule for the opponent model: ; in, represents the weight parameter of the t+1th iteration, represents the weight parameter of the tth iteration, The specific training objectives are defined as follows: (Formula 7); in, is the conditional probability in the large language model parameter space, is the conditional probability in the parameter space for the tth iteration.

[0044] Specifically, the schematic diagram of the large model reasoning capability optimization process of the present invention can be referred to Figure 2 , select the opponent model from t iterations to train the main model, so as to obtain The adversary model parameters for the iteration Then, the parameters are directly copied to form a new adversary model, which is then used to train the main model in the t+2th iteration.

[0045] In this embodiment, in S60, the main model is saved during each iterative training; In this embodiment, the large model reasoning capability optimization method further includes S70, after the main model converges, using the SFT data set to evaluate the saved main model, and selecting the optimized main model according to the evaluation result.

[0046] Preferably, after the iteration exceeds the preset number of times, if the logistic loss function of the main model no longer decreases significantly and basically tends to a stable state, the saved main model is evaluated using the SFT dataset ultraChat200k, 50k data are randomly extracted from it, and the basic response is generated using Zephyr-7B-SFT-Ful. The Huggingface Open LLM ranking is used as the evaluation benchmark to select the best model as the optimized main model.

[0047] The present invention saves the main model during each iterative training and uses the SFT data set for evaluation after the main model converges, thereby ensuring the stability and reliability of the model, being able to select the optimal model based on the evaluation results, and improving the practical application effect of the model.

[0048] like Figure 3 As shown, the present invention also provides a large model reasoning capability optimization system, which includes: An initialization module 10 is used to initialize the large language model to be fine-tuned and use it as the initial version of the main model and the opponent model; The response result generation module 20 is used to execute the response result generation step: obtain the prompt information and question content from the supervised fine-tuning data set, and input them into the opponent model to generate the corresponding response result; The main model training module 30 is used to perform the main model training step: using the real response in the supervised fine-tuning data set, optimizing the main model parameters by minimizing the logistic loss function by a first preset formula, and training the main model to distinguish the response result from the real response; The opponent model updating module 40 is used to perform the opponent model updating step: maximizing the evaluation value of the main model for the generated response through the second preset formula to update the parameters of the opponent model; the regularization term is introduced in the second preset formula, and the evaluation value is specifically the probability or confidence that the main model determines that a certain response is a true response; The iteration module 50 is used to execute the replacement step: using the trained main model as the new opponent model to replace the current opponent model; and iteratively executing the response result generation step, the main model training step, the opponent model updating step and the replacement step in sequence until the main model converges.

[0049] The embodiment of the present invention further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the memory in the above embodiment; or a computer-readable storage medium that exists independently and is not installed in a device. The computer-readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement Figure 1 The large model reasoning capability optimization method shown. The computer readable storage medium can be a read-only memory, a disk or an optical disk, etc.

[0050] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system embodiment and the storage medium embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0051] Furthermore, in this document, the terms "comprises," "comprising," or any other variation thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.

[0052] The above description shows and describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the invention, through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not depart from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A method for optimizing the reasoning capability of a large model, characterized in that: The following steps are involved: S10, initializing the large language model to be fine-tuned and using it as the initial version of the main model and the opponent model; S20, obtaining prompt information and question content from the supervised fine-tuning dataset, and inputting them into the opponent model to generate a corresponding opponent model response; S30, using the real responses in the supervised fine-tuning dataset, optimizing the weight parameters in the main model by minimizing the logistic loss function through a first preset formula, and training the main model to distinguish between the opponent model response and the real response; S40, maximizing the evaluation value of the main model for the generated response through a second preset formula to update the weight parameters in the opponent model; A regularization term is introduced into the second preset formula, and the evaluation value is specifically the probability or confidence that the main model determines that a certain response is a true response; S50, using the trained main model as a new opponent model to replace the current opponent model; S60, repeating steps S20 to S50 until the main model converges.

2. The large model reasoning capability optimization method according to claim 1 is characterized in that: The large language model to be fine-tuned, specifically loading the BPE tokenizer based on the Transformer architecture and initializing the Zephyr-7B-SFT-Full model weights.

3. The large model reasoning capability optimization method according to claim 2 is characterized in that: The hyperparameters of the Zephyr-7B-SFT-Full model include at least the learning rate, optimizer, number of attention heads, number of hidden layers, vocabulary size, and data accuracy.

4. The large model reasoning capability optimization method according to claim 1 is characterized in that: In S20, the prompt information is concatenated with the question content and then input into the opponent model in json format.

5. The large model reasoning capability optimization method according to claim 1 is characterized in that: The first preset formula is as follows: ; in, is the main model, f is a highly expressive function, is a series of highly expressive function classes, is the number of iterations, E is the evaluation value, For prompt information, For a true response, is the opponent model response, is the logistic loss function, specifically a monotonically decreasing and convex loss function, and .

6. The large model reasoning capability optimization method according to claim 5 is characterized in that: The regularization term is specifically a KL regularization term.

7. The large model reasoning capability optimization method according to claim 6 is characterized in that: The second preset formula is as follows: ; in, The probability of generating a response for the adversary model that is indistinguishable from the primary model, Represents the evaluation value gap between the response generated by the opponent model and the true response. The evaluation value is the distribution Calculated, represents the distribution of the adversary model input, Indicates that the true response follows the true response distribution, is the true response distribution, Indicates that the opponent model response follows the opponent model response distribution, is the opponent model response distribution, Indicates that the true response obeys the initial data distribution to be fine-tuned, is the evaluation value of the main model, λ is the regularization parameter, and >0, is the expectation under the distribution of the inputs to the adversary model, is the KL regularization term, is the initial data distribution to be fine-tuned, is the opponent model response distribution at the tth iteration.

8. The large model reasoning capability optimization method according to claim 1, characterized in that: In S60, the main model is saved during each training iteration; The process also includes S70, after the main model converges, using the SFT data set to evaluate the saved main model, and selecting an optimized main model according to the evaluation result.

9. A large model reasoning capability optimization system, characterized in that: include: The initialization module is used to initialize the large language model to be fine-tuned and use it as the initial version of the main model and the opponent model; The opponent model response generation module is used to perform the opponent model response generation step: obtain the prompt information and question content from the supervised fine-tuning dataset, input them into the opponent model, and generate the corresponding opponent model response; A main model training module, configured to perform a main model training step: using the real responses in the supervised fine-tuning dataset, optimizing the weight parameters in the main model by minimizing the logistic loss function by a first preset formula, and training the main model to distinguish between the opponent model responses and the real responses; The opponent model updating module is used to perform the opponent model updating step: maximizing the evaluation value of the main model to the generated response through a second preset formula to update the weight parameters in the opponent model; A regularization term is introduced into the second preset formula, and the evaluation value is specifically the probability or confidence that the main model determines that a certain response is a true response; The iterative module is used to execute the replacement step: using the trained main model as the new opponent model to replace the current opponent model; and iteratively executing the opponent model response generation step, the main model training step, the opponent model update step and the replacement step in sequence until the main model converges.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a large model reasoning capability optimization program, and when the large model reasoning capability optimization program is executed by the processor, the steps of the large model reasoning capability optimization method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Code processing model training method and device, electronic equipment and storage medium

    CN116820429A

  • Method and device for reinforcement learning of large language model

    CN117808120A

  • Urban function composition ratio interactive generation method based on special large language model

    CN119721045A

  • Image recognition method, device and equipment and computer readable storage medium

    CN119741509A

  • Method, apparatus and system for consistency enhanced large language models

    US20250013873A1

Cited By

  • Network security task processing method and device, computer equipment and readable storage medium

    CN120880774A

  • Cybersecurity task processing methods, devices, computer equipment, and readable storage media

    CN120880774B

  • Optimization method and system for brain-like cross-modal continuous learning neural network

    CN121212268A