Question solving method and system based on chain reasoning and program reasoning fusion
By combining chain reasoning with procedural reasoning and using self-distillation technology to train the model, the balance problem between inference logic and computational accuracy of computers when analyzing mathematical problems is solved, achieving more efficient and answering complex mathematical problems.
Patent Information
- Application Number
- CN202510150138.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
When computers analyze mathematical problems, it is difficult for computers to balance the reasoning logic and computational accuracy, resulting in insufficient accuracy when solving complex mathematical problems.
Using a fusion method based on chain reasoning and procedural reasoning, the mathematical problem is split into multiple subtasks through a large language model, and combining the results of chain reasoning and procedural reasoning, the model is trained using self-distillation technology to generate comprehensive answers.
It significantly improves the model's inference and calculation ability, and can solve complex mathematical problems more accurately, which is both logically coherent and computationally accurate.
Smart Images

Figure CN120146180A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a problem-solving method and system based on the fusion of chain reasoning and program reasoning. Background Art
[0002] In recent years, large language models (such as GPT-3, GPT-4, etc.) have made remarkable progress in natural language processing (NLP) and automated reasoning. These models rely on a large number of parameters and training data and can demonstrate powerful capabilities in understanding, generating, and reasoning natural language texts. However, when solving mathematical problems, large models are often prone to errors, especially in scenarios that require precise numerical calculations.
[0003] Chain of Thought (CoT) is a reasoning method that allows large models to show their thinking steps step by step, suitable for solving mathematically complex problems. The advantage of CoT is that it enables the model to decompose complex tasks through a series of logical derivations. However, the details of this reasoning are usually not precise enough, and errors often occur when it comes to specific numerical calculations. Therefore, although CoT has advantages in reasoning logic, its accuracy may be insufficient when solving complex mathematical problems.
[0004] Program of Thought (PoT) is another problem-solving approach that relies on the model to generate code (such as Python) to perform specific calculations. This method is good at numerical calculations and task execution, especially suitable for tasks that require multiple-step calculations or precise answers. For example, the model will generate a piece of code and obtain the accurate result by executing this code. However, the limitation of PoT is that although it can complete complex numerical calculations, it is not ideal in scenarios that require comprehensive reasoning such as multiple-choice questions and logical derivations.
[0005] Patent document CN117668212A discloses a problem-solving method and device. The method includes determining a text to be answered; inputting the text to be answered into a problem-solving model, and in the problem-solving model, determining problem-solving information corresponding to the text to be answered, where the problem-solving information includes a problem-solving expression and a problem-solving service interface identifier; determining a problem-solving service interface according to the problem-solving service interface identifier, so as to call a problem-solving service through the problem-solving service interface to perform problem-solving processing on the problem-solving expression and obtain a problem-solving result returned by the problem-solving service interface; and determining an answer text corresponding to the text to be answered according to the problem-solving result and the problem-solving expression.
[0006] Patent document CN117251533A discloses a method for generating math problems and their solution processes, including: 1) quantitatively analyzing the characteristics of math problems to obtain a dataset containing the chapters to which the math problems belong and the problem characteristics; 2) constructing a framework and a generalized math model for generating math problems and their solution processes based on elementary matrix transformations; 3) designing an algorithm for generating math problems and their solution processes; 4) inputting the measurement content, examined knowledge, and relevant characteristic parameters of math problems, and outputting math problems and their detailed solution processes; 5) quantitatively analyzing the difficulty of the output math problems and relevant measurement indexes to evaluate the quality of the generated math problems. Summary of the Invention
[0007] The purpose of the present invention is to provide a problem-solving method and system based on the integration of chain reasoning and program reasoning, which can effectively solve the balance problem between reasoning logic and calculation accuracy in the process of a computer analyzing math problems.
[0008] To achieve the first object of the present invention, the following technical solution is provided: A problem-solving method based on the integration of chain reasoning and program reasoning, including the following steps: Input a math problem, and parse it through chain reasoning and program reasoning to obtain a chain reasoning process and a program reasoning result; Input the math problem into a pre-constructed large language model to output prompt words related to problem-solving; Form a training set with the prompt words, the chain reasoning process, and the program reasoning result; Construct an initial model based on the large language model framework, and the initial model includes a prompt word extraction module, a chain reasoning module, a program reasoning module, and an answer generation module; The prompt word extraction module is used to obtain the prompt words related to problem-solving in the input math problem; The chain reasoning module reasons based on the input prompt words to output a chain reasoning process; The program reasoning module reasons based on the input prompt words to output a program reasoning result; The answer generation module combines the input chain reasoning process and the program reasoning result in an attention mechanism manner to output a comprehensive answer; Use the training set to perform self-distillation training on the initial model to obtain a problem-solving model for comprehensively answering math problems; Input the math model to be solved into the problem-solving model to output a comprehensive answer including the chain reasoning process and the program reasoning result.
[0009] By combining chain reasoning with process reasoning and using self-distillation technology to train the model, the present invention enables the model to significantly improve its reasoning and computing capabilities and perform better in solving complex mathematical problems.
[0010] Specifically, the chain reasoning method refers to splitting the input mathematical problem into multiple subtasks through a large language model, analyzing each subtask to output the corresponding problem-solving process, and concatenating all the problem-solving processes as the chain reasoning process for output.
[0011] Specifically, the process of splitting into multiple subtasks is as follows: Input the mathematical problem into the large language model to extract the prompt words regarding mathematical knowledge; Match the extracted prompt words with the mathematical prior knowledge for similarity, and output the mathematical prior knowledge with the maximum similarity as the theme of the corresponding subtask.
[0012] Specifically, the program reasoning method refers to generating the corresponding calculation code for the input mathematical problem through a large language model and running the calculation code to output the final numerical value as the program reasoning result.
[0013] Specifically, the answer generation module uses the chain reasoning process as the head and the program reasoning result as the tail to construct a logically coherent comprehensive answer.
[0014] Specifically, the attention mechanism refers to introducing the weights of each subtask in the chain reasoning process into the program reasoning module during training to generate a program reasoning result that conforms to the reasoning logic in the chain reasoning process.
[0015] Specifically, the self-distillation training means that after inputting the data set into the initial model to obtain the initial comprehensive answer, the initial comprehensive answer and the data set are combined to form a new data set and the initial model is fine-tuned for the second time.
[0016] To achieve the second object of the present invention, the following technical solution is provided: A problem-solving system implemented by the above-mentioned problem-solving method based on the integration of chain reasoning and program reasoning, which includes an interaction unit, an analysis unit, and an output unit; The interaction unit is used to input the text content of the mathematical problem; The analysis unit generates the comprehensive answer corresponding to the mathematical problem according to the input text content; The output unit is used to visually output the comprehensive answer.
[0017] Compared with the prior art, the beneficial effects of the present invention: By combining Chain of Thought (CoT) and Process of Thought (PoT), and using self-distillation technology to optimize the mathematical problem-solving ability of large models, the model can significantly improve its reasoning and calculation ability and perform better in solving complex mathematical problems by merging the answers of two different reasoning methods and using self-distillation for training. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Schematic diagram of the problem-solving method based on the integration of chain reasoning and program reasoning provided in this embodiment; Figure 2 Generation process of the comprehensive answer provided in this embodiment; Figure 3 Flow schematic diagram of the self-distillation training provided in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0020] As Figure 1 shown, the problem-solving method based on the integration of chain reasoning and program reasoning provided in this embodiment includes the following steps: Input a math problem and parse it through chain reasoning and program reasoning to obtain the chain reasoning process and the program reasoning result; Input the math problem into a pre-constructed large language model to output prompt words related to problem-solving; Form a training set with the prompt words, the chain reasoning process, and the program reasoning result; Build an initial model based on the large language model framework, where the initial model includes a prompt word extraction module, a chain reasoning module, a program reasoning module, and an answer generation module; The prompt word extraction module is used to obtain the prompt words related to problem-solving in the input math problem; The chain reasoning module reasons based on the input prompt words to output the chain reasoning process; The program reasoning module reasons based on the input prompt words to output the program reasoning result; The answer generation module combines the input chain-of-thought reasoning process and the program reasoning result in an attention mechanism manner to output a comprehensive answer. The initial model is trained by self-distillation using a training set to obtain a problem-solving model for comprehensively solving math problems. The math model to be solved is input into the problem-solving model to output a comprehensive answer containing the chain-of-thought reasoning process and the program reasoning result.
[0021] More specifically, in this embodiment, the chain-of-thought reasoning process (CoT) is that the model provides a logical reasoning solution process through step-by-step derivation, but there may be deficiencies in precise numerical calculations.
[0022] The program reasoning result (PoT) is that the model directly solves numerical calculation problems by generating calculation codes (such as Python), ensuring the accuracy of calculations, but the coherence in logical reasoning is poor.
[0023] Such as Figure 2 As shown, it is the generation process of the comprehensive answer provided by this embodiment, which includes selection and merging. Selection means that for the generated CoT and PoT answers, a mechanism is designed to judge whether the solutions of these two methods are correct for specific problems, and merging means combining the logical reasoning process of CoT and the calculation results of PoT to form a comprehensive answer with both logical coherence and calculation accuracy.
[0024] More specifically, in the process of merging CoT and PoT answers, the method of answering CoT first and then PoT is adopted for merging. The data generated in this way enables the trained model to calculate the weight of the logical reasoning part generated by the CoT part when generating the PoT answer. Therefore, the generated PoT is more logical than the original.
[0025] Such as Figure 3 As shown, it is the process of self-distillation training provided by this embodiment. The self-distillation training includes: generating new training data: using the merged comprehensive answer above as a new training sample to further train the model. Through self-distillation technology, the model optimizes itself using the comprehensive answer it generates, reducing the dependence on external large-scale data sets. As the training process progresses, the model can learn better reasoning and calculation abilities from the merged CoT and PoT answers.
[0026] The specific process is as follows: S1. Generation of chain-of-thought reasoning (CoT) and process reasoning (PoT): First, for each math problem, use a large model to generate corresponding CoT answers and PoT answers respectively. Among them, CoT answers: The model provides a logical reasoning solution process through step-by-step derivation, but there may be deficiencies in precise numerical calculations. PoT answers: The model directly solves numerical calculation problems by generating calculation code (such as Python) to ensure the accuracy of calculations, but the coherence in logical reasoning is not good.
[0027] S2. Answer selection and merging: S21. Select the correct solution method: For the generated CoT and PoT answers, design a mechanism to judge whether the solutions of these two methods are correct for specific problems. S22. Merge to generate a new answer: Combine the two correct answers. The specific method can be to combine the logical reasoning process of CoT and the calculation results of PoT to form a comprehensive answer that has both logical coherence and calculation accuracy. S23. Attention mechanism optimization: In the process of merging CoT and PoT answers, adopt the method of merging CoT answers first and then PoT answers. The data generated in this way can enable the trained model to calculate the weight of the logical reasoning part generated by the CoT part when generating PoT answers. Therefore, the generated PoT is more logical than the original one.
[0028] S3. Self-distillation training: S31. Generate new training data: Use the above-mentioned merged comprehensive answer as a new training sample to further train the model.
[0029] S32. Self-distillation optimization: Through self-distillation technology, the model uses the comprehensive answers it generates to optimize itself, reducing the dependence on external large-scale data sets. As the training process progresses, the model can learn better reasoning and calculation abilities from the merged CoT and PoT answers.
[0030] This embodiment also provides a problem-solving system, which is implemented by the above-mentioned problem-solving method based on the fusion of chain reasoning and program reasoning. It includes an interaction unit, an analysis unit, and an output unit. The interaction unit is used to input the text content of math problems. The analysis unit generates a comprehensive answer corresponding to the math problem according to the input text content. The output unit is used to visually output the comprehensive answer.
[0031] In addition, the terms "upper", "lower", "inner", "outer", "front", and "rear" are for descriptive purposes only and should not be construed as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0032] Of course, the above are only specific embodiments of the present invention and do not limit the scope of implementation of the present invention. Any equivalent changes or modifications made according to the structure, features, and principles described in the scope of the patent application of the present invention should be included in the scope of the patent application of the present invention.
[0033] Finally, it should be noted that the above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or can easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A problem-solving method based on the fusion of chain reasoning and procedural reasoning, characterized in that: The following steps are involved: Input math problems and analyze them through chain reasoning and procedural reasoning to obtain chain reasoning process and procedural reasoning results; Input math problems into a pre-built large language model to output prompt words related to solving the problem; The prompt words, chain reasoning process and program reasoning results form a training set; Building an initial model based on a large language model framework, the initial model includes a prompt word extraction module, a chain reasoning module, a program reasoning module and an answer generation module; The prompt word extraction module is used to obtain prompt words related to solving the problem in the input math problem; The chain reasoning module performs reasoning according to the input prompt words to output the chain reasoning process; The program reasoning module performs reasoning according to the input prompt words to output the program reasoning result; The answer generation module uses an attention mechanism to merge the input chain reasoning process with the program reasoning result to output a comprehensive answer; The initial model is trained by self-distillation using the training set to obtain a problem-solving model for comprehensively solving math problems; The mathematical model of the substitute solution is input into the problem-solving model to output a comprehensive answer including the chain reasoning process and the procedural reasoning results.
2. The problem-solving method based on the fusion of chain reasoning and procedural reasoning according to claim 1 is characterized in that: The chain reasoning method refers to splitting the input math problem into multiple subtasks through a large language model, analyzing each subtask to output the corresponding problem-solving process, and connecting all the problem-solving processes in series as a chain reasoning process for output.
3. The problem-solving method based on the fusion of chain reasoning and procedural reasoning according to claim 2 is characterized in that: The process of splitting into multiple subtasks is as follows: Input math problems into a large language model to extract clue words about math knowledge; The extracted prompt words are matched with mathematical prior knowledge in similarity, and the mathematical prior knowledge with the greatest similarity is output as the topic of the corresponding subtask.
4. The problem-solving method based on the fusion of chain reasoning and procedural reasoning according to claim 1 is characterized in that: The program reasoning method refers to generating corresponding calculation codes for input math problems through a large language module, and running the calculation codes to output final values as program reasoning results.
5. The problem-solving method based on the fusion of chain reasoning and procedural reasoning according to claim 1 is characterized in that: The answer generation module uses the chain reasoning process as the head and the program reasoning result as the tail to construct a logically coherent comprehensive answer.
6. The problem-solving method based on the fusion of chain reasoning and procedural reasoning according to claim 1 is characterized in that: The attention mechanism refers to introducing the weights of each subtask in the chain reasoning process into the program reasoning module during the training process to generate a program reasoning result that conforms to the reasoning logic in the chain reasoning process.
7. The problem-solving method based on the fusion of chain reasoning and procedural reasoning according to claim 1 is characterized in that: The self-distillation training refers to inputting the data set into the initial model to obtain an initial comprehensive answer, combining the initial comprehensive answer with the data set to form a new data set and performing secondary fine-tuning on the initial model.
8. A problem-solving system, characterized in that: The problem-solving method based on the fusion of chain reasoning and program reasoning as described in any one of claims 1 to 7 is implemented, and includes an interaction unit, an analysis unit and an output unit; The interactive unit is used to input the text content of the math problem; The analysis unit generates a comprehensive answer corresponding to the math problem based on the input text content; The output unit is used to visually output the comprehensive answer.
Citation Information
Patent Citations
Mathematical question and answering process generation method thereof
CN117251533A
Question solving method and device
CN117668212A
Mathematical problem answering model training method and device
CN116595159A
Text processing model training method, text processing method and question and answer processing method and device
CN118627543A
Multi-instance controllable image generation method based on cross attention redistribution
CN118628611A
Cited By
Large model problem solving method and system integrating understanding and reasoning
CN120598748A