Language model reinforcement learning training extension method and device based on reflection
By generating target reflection thinking chains in large language models and performing reinforcement learning training expansion, the problem of insufficient expansion potential of reinforcement learning from human feedback is solved, and the ability of large language models to solve complex problems and the understanding of reinforcement learning training is improved.
Patent Information
- Application Number
- CN202510144398.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-13
AI Technical Summary
In related technologies, less attention is paid to the expansion potential and attributes of reinforcement learning from human feedback, which leads to limitations in practical understanding of large-scale reinforcement learning training and reduces the ability of large-language models to solve complex problems.
By inputting the target natural language inference problem into the pre-constructed large language model, the target reflection thinking chain of the large language model is generated, and based on this, reinforcement learning training extension is carried out, the target inference model is obtained, and its relationship between the target performance and the generation length is evaluated to determine the target validity of the inference extension.
It effectively improves the ability of large language models to solve complex problems, expands the understanding and application of reinforcement learning training, and improves the performance of the model in complex tasks.
Smart Images

Figure CN120146142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of reinforcement learning, and particularly relates to a method and device for expanding the reinforcement learning training of a language model based on reflection. Background Art
[0002] LLMs (Large Language Models) have revolutionized the field of natural language processing by learning extensive language patterns from large datasets. A key step in these models is RLHF (Reinforcement Learning from Human Feedback), which helps align the behavior of the model with human intentions and improve performance in various tasks such as text generation, code, and mathematical reasoning. This approach has been successfully applied to leading models such as ChatGPT, Llama, etc., and has brought significant improvements. RLHF optimizes the behavior of the model by integrating external feedback (e.g., human preferences and answer supervision), thereby enhancing its generation ability. The process starts with training a reward model using human-annotated preference data or inference data with correctness labels, and the reward model serves as an agent for human supervision. Then, reinforcement learning is used to optimize the policy model through iterative feedback from the reward model.
[0003] The Scaling Law of large language models is one of the key factors in the success of current language models. As computing resources (e.g., more data and larger models) increase, their performance will be significantly improved. However, although the pre-training stage and supervised fine-tuning have been widely studied, the scalability of reinforcement learning, including training expansion and inference expansion, has rarely been studied in depth. In related technologies, the emergence of OpenAI - o1 has demonstrated the potential of scaled reinforcement learning in inference tasks, and some studies have attempted to explore the greater potential of reinforcement learning. For example, the effects of PPO (Proximal Policy Optimization) and DPO (Divergence-constrained Policy Optimization) in different task expansions have been compared.
[0004] However, in related technologies, less attention has been paid to the expansion potential and properties of reinforcement learning from human feedback, resulting in limited practical understanding of large-scale reinforcement learning training and reducing the ability of large language models to solve complex problems, which urgently needs to be solved. Summary of the Invention
[0005] This application provides a method and device for expanding the reinforcement learning training of a language model based on reflection, so as to solve the problem that in the related art, less attention is paid to the expansion potential and attributes of reinforcement learning from human feedback, resulting in limitations in the practical understanding of large-scale reinforcement learning training and reducing the ability of large language models to solve complex problems.
[0006] The first aspect embodiment of this application provides a method for expanding the reinforcement learning training of a language model based on reflection, including the following steps: inputting a target natural language reasoning problem into a pre-constructed large language model to generate a target reflection thought chain of the large language model; based on the target reflection thought chain, performing reinforcement learning training expansion on the large language model to obtain a target reasoning model; evaluating the relationship between the target performance and the generation length of the target reasoning model to obtain an evaluation result, and using the evaluation result to determine the target effectiveness of the reasoning expansion of the target reasoning model.
[0007] Optionally, in an embodiment of this application, the step of inputting a target natural language reasoning problem into a pre-constructed large language model to generate a target reflection thought chain of the large language model includes: inputting the target natural language reasoning problem into the pre-constructed large language model to generate multiple answers, and performing evaluation processing on each answer to obtain a processing result, and based on the processing result, determining an initial thought chain; determining a target reasoning pattern in the initial thought chain, and using the target reasoning pattern and the large language model to optimize the initial thought chain to generate a target reflection thought chain of the large language model.
[0008] Optionally, in an embodiment of this application, the step of performing reinforcement learning training expansion on the large language model based on the target reflection thought chain to obtain a target reasoning model includes: incorporating the information entropy of each answer as a reward into the loss function to encourage the large language model to perform target exploration in the reinforcement learning training to generate an exploration result; punishing the target answer in the exploration result to construct the target reasoning model.
[0009] Optionally, in an embodiment of this application, the step of evaluating the relationship between the target performance and the generation length of the target reasoning model to obtain an evaluation result includes: inputting the natural language reasoning problem into the target reasoning model to generate a long thinking text; using a preset length to segment the long thinking text to obtain multiple segmented thinking texts; inputting the multiple segmented thinking texts into a target summary model to generate a final answer, and determining the evaluation result according to the final answer.
[0010] Optionally, in an embodiment of this application, the target reflection thought chain is expressed as:
[0011]
[0012] Among them, p is the extracted reasoning pattern, q is the natural language reasoning problem, and a ′ is the chain of reflective thinking.
[0013] In the second aspect of the embodiments of the present application, a device for expanding the reinforcement learning training of a language model based on reflection is provided, including: an input module, configured to input a target natural language reasoning problem into a pre-constructed large language model to generate a target chain of reflective thinking of the large language model; an acquisition module, configured to perform reinforcement learning training expansion on the large language model based on the target chain of reflective thinking to obtain a target reasoning model; an evaluation module, configured to evaluate the relationship between the target reasoning model in terms of target performance and generation length to obtain an evaluation result, and use the evaluation result to determine the target effectiveness of the reasoning expansion of the target reasoning model.
[0014] Optionally, in an embodiment of the present application, the input module includes: an input unit, configured to input the target natural language reasoning problem into the pre-constructed large language model to generate multiple answers, and perform evaluation processing on each answer to obtain a processing result, and based on the processing result, determine an initial chain of thinking; a determination unit, configured to determine a target reasoning pattern in the initial chain of thinking, and use the target reasoning pattern and the large language model to optimize the initial chain of thinking to generate a target chain of reflective thinking of the large language model.
[0015] Optionally, in an embodiment of the present application, the acquisition module includes: a first generation unit, configured to incorporate the information entropy of each answer as a reward into a loss function to encourage the large language model to perform target exploration in reinforcement learning training to generate an exploration result; a construction unit, configured to punish the target answer in the exploration result to construct the target reasoning model.
[0016] Optionally, in an embodiment of the present application, the evaluation module includes: an input unit, configured to input the natural language reasoning problem into the target reasoning model to generate a long thinking text; an acquisition unit, configured to use a preset length to segment the long thinking text to obtain multiple segmented thinking texts; a second generation unit, configured to input the multiple segmented thinking texts into a target summary model to generate a final answer, and determine the evaluation result according to the final answer.
[0017] Optionally, in an embodiment of the present application, the target chain of reflective thinking is expressed as:
[0018]
[0019] Among them, p is the extracted reasoning pattern, q is the natural language reasoning problem, and a ′ is the chain of reflective thinking.
[0020] The third aspect of the embodiments of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the method for expanding the reinforcement learning training of the language model based on reflection as described in the above embodiments.
[0021] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the method for expanding the reinforcement learning training of the language model based on reflection as described above.
[0022] The fifth aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed, it is used to implement the method for expanding the reinforcement learning training of the language model based on reflection as described above.
[0023] The embodiments of the present application can input a target natural language reasoning problem into a pre-constructed large language model to generate a target chain of reflective thinking of the large language model, so as to expand the reinforcement learning training of the large language model to obtain a target reasoning model, and then evaluate the relationship between the target performance and the generated length of the target reasoning model, so as to use the evaluation result to determine the target effectiveness of the reasoning expansion of the target reasoning model, effectively improving the ability of the large language model to solve complex problems. Thus, it solves the problem in the related technology that less attention is paid to the expansion potential and attributes of reinforcement learning from human feedback, resulting in limitations in the practical understanding of large-scale reinforcement learning training and reducing the ability of the large language model to solve complex problems.
[0024] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings
[0025] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0026] Figure 1 is a flowchart of a method for expanding the reinforcement learning training of a language model based on reflection according to an embodiment of the present application;
[0027] Figure 2 is a flowchart of a method for expanding the reinforcement learning training of a language model based on reflection in a specific embodiment of the present application;
[0028] Figure 3Schematic diagram for measuring the effectiveness of the inference expansion ability of a specific embodiment of the present application;
[0029] Figure 4 Structural schematic diagram of a reinforcement learning training expansion device for a language model based on reflection provided according to an embodiment of the present application;
[0030] Figure 5 Structural schematic diagram of an electronic device provided according to an embodiment of the present application. Detailed implementation manners
[0031] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.
[0032] The method and device for expanding reinforcement learning training of a language model based on reflection according to an embodiment of the present application will be described below with reference to the accompanying drawings. Regarding the problem in the related art mentioned in the above background technology that less attention is paid to the expansion potential and properties of reinforcement learning from human feedback, resulting in limited practical understanding of large-scale reinforcement learning training and reducing the ability of large language models to solve complex problems, the present application provides a method for expanding reinforcement learning training of a language model based on reflection. In this method, a target natural language inference problem can be input into a pre-constructed large language model to generate a target reflection thought chain of the large language model, so as to expand the reinforcement learning training of the large language model to obtain a target inference model, and then evaluate the relationship between the target performance and the generation length of the target inference model, so as to determine the target effectiveness of the inference expansion of the target inference model using the evaluation result, effectively improving the ability of the large language model to solve complex problems. Thus, the problem in the related art that there is a limitation in the practical understanding of large-scale reinforcement learning training and the ability of large language models to solve complex problems is solved.
[0033] Specifically, Figure 1 Flow schematic diagram of a method for expanding reinforcement learning training of a language model based on reflection provided according to an embodiment of the present application.
[0034] As Figure 1 shown, the method for expanding reinforcement learning training of a language model based on reflection includes the following steps:
[0035] In step S101, a target natural language inference problem is input into a pre-constructed large language model to generate a target reflection thought chain of the large language model.
[0036] In an embodiment of the present application, the target reflection thought chain is a thought chain with a reflection process.
[0037] It can be understood that in the embodiments of the present application, the target natural language inference problem can be input into a pre-constructed large language model. For example, a natural language inference problem q can be input into the large language model to generate multiple answers in the following steps. Then, the large language model is used to evaluate the answers according to the reference answers and summarize them into an inference thought chain, so that a thought chain with a reflection process can be determined based on the inference thought chain, and further an effective starting point can be provided for the reinforcement learning model.
[0038] Among them, in an embodiment of the present application, inputting the target natural language inference problem into a pre-constructed large language model to generate the target reflection thought chain of the large language model includes: inputting the target natural language inference problem into a pre-constructed large language model to generate multiple answers, and performing evaluation processing on each answer to obtain a processing result, and determining an initial thought chain based on the processing result; determining the target inference mode in the initial thought chain, and using the target inference mode and the large language model to optimize the initial thought chain to generate the target reflection thought chain of the large language model.
[0039] In the actual execution process, first, the embodiments of the present application can construct thought chain data with a reflection process. For example, in order to generate thought chain data with rich inference modes (including reflection and self-verification), the present application proposes a method of integrating multiple thought processes of each prompt into a thought chain with reflection and search, which combines incorrect attempts, correction processes, and verified successful methods. Specifically, given a natural language inference problem q, first generate multiple answers (a 1 , a 2 , …, a n )(n is generally 3-5) for q, and evaluate the accuracy of each answer according to the reference answer.
[0040] Next, guide the large language model to perform self-evaluation on each answer: for incorrect or partially correct answers, the large language model tries to identify the position where the error occurs and point out the reason for the error, and at the same time tries to correct it with reference to the correct answer; for correct responses, prompt the large language model to perform mutual verification among multiple correct answers. As self-verification, (q, a i , c i ) can be obtained, where c i is the self-evaluation corresponding to each answer.
[0041] Finally, these evaluated answers including error correction and mutual verification are integrated into a unified linear inference chain, that is, the initial thought chain, which is expressed as:
[0042]
[0043] Among them, q is a natural language inference question, and c i is the self-assessment corresponding to each answer, and a n is the generated answer, is the initial thought chain.
[0044] The resulting initial thought chain not only shows a more complex reasoning pattern but also includes a trial-and-error and reflection process from the initial attempt to the final solution. However, the initial thought chain constructed by the above process may contain repeated reasoning processes. For example, simply listing different methods and finally giving the correct answer.
[0045] Furthermore, the embodiments of the present application can use a large language model to abstract the core reasoning pattern from the initial thought chain Among them, p is the extracted reasoning pattern, that is, reflection, verification, error correction, etc.; then, based on the extracted reasoning pattern, use the large language model to process the initial thought chain for optimization, so as to obtain a thought chain containing a reflective reasoning process, that is, a reflective thought chain a′, which is expressed as:
[0046]
[0047] Among them, p is the extracted reasoning pattern, q is the natural language inference question, and a ′ is the reflective thought chain.
[0048] In step S102, based on the target reflective thought chain, the large language model is extended by reinforcement learning training to obtain a target reasoning model.
[0049] In the embodiments of the present application, the target reasoning model is a reasoning model with long thinking.
[0050] It can be understood that the embodiments of the present application can extend the large language model by reinforcement learning training based on the target reflective thought chain. For example, through over-sampling, entropy loss, etc. in the following steps, diverse exploration in reinforcement learning training can be encouraged, and ngram and rules, etc. can be used to impose penalties on unreasonable answers, such as punishing garbage texts that appear in exploration, so as to obtain a reasoning model with long thinking, and further achieve effective and stable extended training of reinforcement learning.
[0051] Among them, in one embodiment of the present application, extending the large language model by reinforcement learning training based on the target reflective thought chain to obtain a target reasoning model includes: incorporating the information entropy of each answer as a reward into the loss function to encourage the large language model to perform target exploration in reinforcement learning training to generate exploration results; punishing the target answers in the exploration results to construct a target reasoning model.
[0052] As a possible implementation, the embodiments of the present application can, on the basis of maximizing the reward function of the existing reinforcement learning RLOO algorithm, propose to adopt various methods during the reinforcement learning stage to encourage exploration and punish unreasonable behaviors, so as to achieve effective expansion of reinforcement learning.
[0053] For example, to encourage the exploration of large language models, the present application expands sampling, including generating multiple answers for each question during reinforcement learning training, that is:
[0054] {(q 1 ,a 1 ),…,(q k ,a k )},k≥32
[0055] Thereby, a wider range of reasoning paths can be captured. At the same time, the information entropy of each answer is incorporated as a reward into the loss function, that is:
[0056]
[0057] where V is the language model vocabulary set, L RL is the loss function of reinforcement learning, β is the weight of the entropy reward, H is the calculation of entropy, w is the word in the vocabulary, and q is the natural language inference question.
[0058] By incorporating the information entropy of each answer as a reward into the loss function, the large language model is motivated to explore higher information gain.
[0059] Also, for example, to improve the training stability, the embodiments of the present application need to punish bad responses while encouraging exploration. Specifically, an adjusted reward function is set: if there is repeated and overly long text (repeated n-grams or exceeding the predefined length threshold) or garbage text (mixed languages or garbled characters) in the response, the reward is -1. This can guide the large language model to avoid generating repeated, lengthy, or garbled text, and build a reasoning model with long thinking, thereby improving the generation quality and training stability.
[0060] In step S103, the relationship between the target reasoning model's target performance and generation length is evaluated to obtain an evaluation result, so as to use the evaluation result to determine the target effectiveness of the target reasoning model's reasoning expansion.
[0061] It can be understood that the embodiments of the present application can evaluate the relationship between the target inference model in terms of target performance and generation length. For example, the relationship between the target inference model in terms of target performance and generation length can be evaluated based on the "thinking - summarizing" method, and the evaluation result is used to determine the effectiveness of the expansion of the target inference model in the inference stage, thereby improving the inference efficiency and accuracy of the model and enhancing the generalization ability of the model.
[0062] Among them, in one embodiment of the present application, to evaluate the relationship between the target inference model in terms of target performance and generation length to obtain an evaluation result, it includes: inputting a natural language inference problem into the target inference model to generate a long thinking text; using a preset length to segment the long thinking text to obtain multiple segmented thinking texts; inputting the multiple segmented thinking texts into the target summary model to generate a final answer, and determining the evaluation result according to the final answer.
[0063] In some embodiments, the embodiments of the present application can evaluate the inference expansion ability of the model after reinforcement learning, that is, the target inference model. For example, based on the method of thinking summary, the relationship between the performance of the inference model and the generation length is evaluated. Specifically, an additional summary model π is trained in the present application. φ , summary model π φ can generate a concise final answer (q,a)→s according to the given chain - of - thought reasoning (referred to as "thinking"). Given the problem and the long thinking answer of the model, in order to evaluate how the model performance changes with the thinking length, this "thinking" text is truncated to a predetermined length, for example, in intervals of 1024, and the truncated thinking text is input into the summary model, so that the summary model generates a final answer based on this thinking process, and finally obtains {(q,a 1024 ,s 1 ),(q,a 2048 ,s 2 ),…,(q,a 1024*k ,s k )}. By evaluating the accuracy of the summary answers {s 1 ,…,s k}, the relationship between the inference thinking length and performance of the model can be effectively evaluated, so that the effectiveness of the inference expansion of the model can be verified and the generalization ability of the model can be enhanced.
[0064] For example, as Figure 2 shown, the working principle of the embodiments of the present application is elaborated in detail below with a specific embodiment.
[0065] Step S201: The large - language model generates multiple answers, that is, a natural language inference problem is input into the large - language model to generate multiple answers.
[0066] Step S202: Evaluate whether each solution is correct. If not, execute Step S203; otherwise, execute Step S204.
[0067] Step S203: Error analysis and correction. That is, for incorrect or partially correct solutions, the large language model attempts to identify the location where the error occurs, point out the reason for the error, and at the same time try to correct it with reference to the correct solution.
[0068] Step S204: Self-verification. That is, for correct responses, prompt the large language model to mutually verify among multiple correct answers as self-verification.
[0069] Step S205: Fusion, optimization, and rewriting of multiple solutions.
[0070] Step S206: Chain of thought with reflection. That is, based on the fusion of multiple processes, construct a chain of thought with reflection.
[0071] Step S207: Reinforcement learning training: Encourage exploration and punish errors. That is, encourage diverse exploration in reinforcement learning training through methods such as oversampling and entropy loss, and impose penalties on unreasonable solutions using methods such as ngram and rules.
[0072] Step S208: Inference model, that is, an inference model with long thinking.
[0073] Step S209: Expansion law based on thought summary. That is, based on the way of thought summary, determine the effectiveness of the expansion of the inference model in the inference stage, thereby enhancing the generalization ability of the model.
[0074] For example, the embodiments of the present application can verify the effectiveness of the reinforcement learning training expansion and inference expansion methods of the large language model based on mathematical and scientific tasks.
[0075] First, the present application evaluated the proposed method on mathematical and physical chemistry tasks. MATH500 is data of high school-level difficult math problems, Omni-math-500 covers high school high-difficulty and college math problems, and AIME mainly includes high-difficulty math competition problems. GPQA is mainly a collection of college physical chemistry and biology problems. Among them, Table 1 is the evaluation result table on mathematical and physical chemistry tasks, and the specific Table 1 is as follows:
[0076] Table 1
[0077]
[0078] As can be seen from Table 1, in the experiments based on the open-source model Qwen, the method proposed by the present application achieved remarkable results on these datasets.
[0079] Secondly, asFigure 3 As shown, it demonstrates the effectiveness of using the method proposed in this application to measure the inference expansion ability of the model. From Figure 3 it can be seen that by using the method proposed in this application, as the amount of reinforcement learning training increases, the performance of the model shows a near logarithmic-linear inference expansion trend as the thinking length grows. This result indicates that the inference expansion ability of the model continuously enhances after reinforcement learning, and at the same time, this inference expansion ability can be effectively measured.
[0080] According to the method for expanding reinforcement learning training of a reflection-based language model proposed in an embodiment of this application, a target natural language inference problem can be input into a pre-constructed large language model to generate a target reflection thought chain of the large language model, thereby expanding the reinforcement learning training of the large language model to obtain a target inference model, and then evaluating the relationship between the target performance and the generation length of the target inference model to determine the target effectiveness of the inference expansion of the target inference model using the evaluation result, effectively improving the ability of the large language model to solve complex problems. Thus, it solves the problem that in the related art, the practical understanding of large-scale reinforcement learning training has limitations, reducing the ability of the large language model to solve complex problems.
[0081] Next, a reflection-based language model reinforcement learning training expansion device proposed in an embodiment of this application will be described with reference to the accompanying drawings.
[0082] Figure 4 is a block diagram of a reflection-based language model reinforcement learning training expansion device according to an embodiment of this application.
[0083] As Figure 4 shown, the reflection-based language model reinforcement learning training expansion device 10 includes: an input module 100, an acquisition module 200, and an evaluation module 300.
[0084] Specifically, the input module 100 is configured to input a target natural language inference problem into a pre-constructed large language model to generate a target reflection thought chain of the large language model.
[0085] The acquisition module 200 is configured to perform reinforcement learning training expansion on the large language model based on the target reflection thought chain to obtain a target inference model.
[0086] The evaluation module 300 is configured to evaluate the relationship between the target performance and the generation length of the target inference model to obtain an evaluation result, and use the evaluation result to determine the target effectiveness of the inference expansion of the target inference model.
[0087] Optionally, in an embodiment of this application, the input module 100 includes: an input unit and a determination unit.
[0088] Among them, an input unit is configured to input a target natural language inference problem into a pre-constructed large language model to generate multiple answers, evaluate each answer to obtain a processing result, and determine an initial thought chain based on the processing result.
[0089] A determination unit is configured to determine a target inference pattern in the initial thought chain, and use the target inference pattern and the large language model to optimize the initial thought chain to generate a target reflective thought chain of the large language model.
[0090] Optionally, in an embodiment of the present application, the acquisition module 200 includes: a first generation unit and a construction unit.
[0091] Among them, the first generation unit is configured to incorporate the information entropy of each answer as a reward into the loss function to encourage the large language model to perform target exploration in the reinforcement learning training to generate an exploration result.
[0092] The construction unit is configured to punish the target answer in the exploration result to construct a target inference model.
[0093] Optionally, in an embodiment of the present application, the evaluation module 300 includes: an input unit, an acquisition unit, and a second generation unit.
[0094] Among them, the input unit is configured to input a natural language inference problem into the target inference model to generate a long thinking text.
[0095] The acquisition unit is configured to use a preset length to segment the long thinking text to obtain multiple segmented thinking texts.
[0096] The second generation unit is configured to input the multiple segmented thinking texts into the target summary model to generate a final answer, and determine an evaluation result according to the final answer.
[0097] Optionally, in an embodiment of the present application, the target reflective thought chain is expressed as:
[0098]
[0099] Among them, p is the extracted inference pattern, q is the natural language inference problem, and a ′ is the reflective thought chain.
[0100] It should be noted that the foregoing explanation of the embodiment of the reinforcement learning training extension method for a reflection-based language model also applies to the reflection-based language model reinforcement learning training extension device of this embodiment, and will not be elaborated here.
[0101] The reinforcement learning training expansion device for a language model based on reflection proposed according to an embodiment of the present application can input a target natural language inference problem into a pre-constructed large language model to generate a target reflection thought chain of the large language model, thereby performing reinforcement learning training expansion on the large language model to obtain a target inference model, and then evaluating the relationship between the target performance and the generation length of the target inference model to determine the target effectiveness of the inference expansion of the target inference model using the evaluation results, effectively improving the ability of the large language model to solve complex problems. Thus, the problem in the related art that the practical understanding of large-scale reinforcement learning training has limitations and reduces the ability of the large language model to solve complex problems is solved.
[0102] Figure 5 The structure diagram of the electronic device provided by the embodiment of the present application. The electronic device may include:
[0103] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0104] When the processor 502 executes the program, it implements the method for expanding the reinforcement learning training of the language model based on reflection provided in the above embodiment.
[0105] Furthermore, the electronic device further includes:
[0106] A communication interface 503 for communication between the memory 501 and the processor 502.
[0107] The memory 501 is used to store a computer program executable on the processor 502.
[0108] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0109] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5It is represented by only one thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0110] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.
[0111] The processor 502 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.
[0112] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the above-mentioned method for extending reinforcement learning training of a reflection-based language model.
[0113] This embodiment also provides a computer program product, including a computer program. When the computer program is executed, it is used to implement the above-mentioned method for extending reinforcement learning training of a reflection-based language model.
[0114] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0115] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0116] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application pertain.
[0117] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0118] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0119] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above-described embodiment methods can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0120] In addition, in each of the embodiments of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0121] The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for extending language model reinforcement learning training based on reflection, characterized in that: The following steps are involved: Inputting a target natural language reasoning problem into a pre-built large language model to generate a target reflective thinking chain of the large language model; Based on the target reflective thinking chain, the large language model is expanded through reinforcement learning training to obtain a target reasoning model; The target reasoning model is evaluated in terms of the relationship between the target performance and the generation length to obtain an evaluation result, so as to use the evaluation result to determine the target effectiveness of the reasoning extension of the target reasoning model.
2. The method according to claim 1, characterized in that The step of inputting the target natural language reasoning problem into the pre-built large language model to generate the target reflective thinking chain of the large language model includes: Inputting the target natural language reasoning problem into the pre-built large language model to generate multiple answers, and evaluating each answer to obtain a processing result, and determining an initial thinking chain based on the processing result; Determine a target reasoning pattern in the initial thinking chain, and optimize the initial thinking chain using the target reasoning pattern and the large language model to generate a target reflective thinking chain of the large language model.
3. The method according to claim 2, characterized in that The step of performing reinforcement learning training and expansion on the large language model based on the target reflective thinking chain to obtain a target reasoning model includes: Incorporating the information entropy of each solution as a reward into the loss function to encourage the large language model to perform target exploration in reinforcement learning training to generate exploration results; Penalizing the target solution in the exploration result to build the target reasoning model.
4. The method according to claim 1, characterized in that: The evaluating the relationship between the target performance and the generation length of the target reasoning model to obtain an evaluation result includes: Inputting the natural language reasoning question into the target reasoning model to generate a long thinking text; Segmenting the long thinking text according to a preset length to obtain a plurality of segmented thinking texts; The plurality of segmented thought texts are input into a target summary model to generate a final answer, and the evaluation result is determined according to the final answer.
5. The method according to claim 1, characterized in that The target reflection thinking chain is expressed as: Among them, p is the extracted reasoning pattern, q is the natural language reasoning problem, and a ′ For reflective thinking chain.
6. A language model reinforcement learning training extension device based on reflection, characterized in that: include: An input module, used for inputting a target natural language reasoning problem into a pre-built large language model to generate a target reflective thinking chain of the large language model; An acquisition module, used to perform reinforcement learning training and expansion on the large language model based on the target reflective thinking chain to obtain a target reasoning model; An evaluation module is used to evaluate the relationship between the target performance and the generation length of the target reasoning model to obtain an evaluation result, so as to use the evaluation result to determine the target effectiveness of the reasoning extension of the target reasoning model.
7. The device according to claim 6, characterized in that The input module comprises: An input unit, used for inputting the target natural language reasoning problem into the pre-built large language model to generate multiple answers, and performing evaluation processing on each answer to obtain a processing result, and determining an initial thinking chain based on the processing result; A determination unit is used to determine a target reasoning pattern in the initial thinking chain, and optimize the initial thinking chain using the target reasoning pattern and the large language model to generate a target reflective thinking chain of the large language model.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the reflection-based language model reinforcement learning training expansion method as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the reflection-based language model reinforcement learning training expansion method as described in any one of claims 1-5.
10. A computer program product, comprising a computer program, characterized in that The computer program is executed by a processor to implement the reflection-based language model reinforcement learning training expansion method as described in any one of claims 1 to 5.
Citation Information
Cited By
Model training method and device, equipment and storage medium
CN120541529A
Data set generation and answer verification method and device, and medium
CN121072731A
A dataset generation and answer verification method, device, and medium
CN121072731B
Model training method and platform and text reasoning method and platform
CN121119025A