Inference Task Execution Method, Electronic Device, Readable Storage Medium, and Program Product
Through the method of bilingual model collaboration, the first language model is used to generate initial ideas and optimize the second language model, combining the step generation process of real-time scoring and correction, the problem of logical errors in complex inference tasks is solved, and efficient and low-cost high-quality inference data generation is achieved.
Patent Information
- Application Number
- CN202510352937.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-25
AI Technical Summary
When performing complex inference tasks in the prior art, the generated inference chain may have logical errors and cannot meet the user's accuracy needs.
Using the method of bilingual model collaboration, firstly use the first language model to generate preliminary ideas for task execution, and then correct and optimize the initial ideas through the second language model to ensure the accuracy of task execution ideas. Then, each step of the task is gradually generated using the first language model, and after each step is generated, the second language model is used to score and correct it in real time, and the generation process is dynamically adjusted.
Through the dual-model collaborative reflection method, the accuracy of task execution ideas and the authenticity of the reasoning chain are improved, the quality of the generated data is ensured, while reducing the computing cost and improving efficiency.
Smart Images

Figure CN119886360B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a method for executing an inference task, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] NLP (Natural Language Processing) can help a computer understand and process human language, enabling the computer to read and understand natural language. With the rapid development of artificial intelligence technology, the accuracy requirement of language models in executing complex inference tasks is also getting higher and higher.
[0003] Related technologies improve the accuracy of language models in executing complex inference tasks by generating multi-step inference paths or combining search algorithms. However, the inference chains generated by this method may have logical errors and cannot meet the accuracy requirements of users. Summary of the Invention
[0004] The present invention provides a method for executing an inference task, an electronic device, a computer-readable storage medium, and a computer program product, which can generate high-quality inference data efficiently and at low cost, and effectively improve the execution accuracy and execution efficiency of complex inference tasks.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] On the one hand, the present invention provides a method for executing an inference task, including:
[0007] Inputting a task to be processed and a prompt into a pre-trained task processing model; the task processing model includes a first language model and a second language model; according to the prompt, using the first language model to generate an initial task execution idea corresponding to the task to be processed, using the second language model and the task to be processed to analyze the initial task execution idea, and using the first language model to determine a task execution idea according to the initial task execution idea and the result of the idea analysis; according to the prompt, based on the single-step analysis result output by the second language model, using the first language model to gradually generate each task execution step according to the task execution idea, and the single-step analysis result is obtained by analyzing a single initial task execution step output by the first language model.
[0008] On the other hand, the present invention provides a method for executing an inference task, including:
[0009] Input the training sample set of the inference task and the prompt words into a pre-built task processing model. The task processing model includes a first language model and a second language model that have completed pre-training. The labels of each inference task data sample in the training sample set of the inference task include at least the question, the task execution idea, the idea adjustment process, the task execution steps, the step reflection process, and the task execution result. According to the prompt words, use the first language model to generate the initial idea prediction data corresponding to the question sample of the current inference task data sample, and use the second language model and the question sample to analyze the initial idea prediction data. And use the first language model to determine the idea prediction data based on the initial idea prediction data and the idea analysis prediction result. According to the prompt words, based on the single-step analysis prediction result obtained by the second language model analyzing a single initial task execution prediction step output by the first language model, use the first language model to gradually generate the task execution prediction steps according to the idea prediction data, and use the question sample, the initial idea prediction data, the idea analysis prediction result, each initial task execution prediction step, the single-step analysis prediction result, and each task execution prediction step as the prediction data. Based on the labels of each inference task data sample and their corresponding prediction data, continuously train the task processing model until the preset model training stop condition is met.
[0010] The present invention also provides an electronic device, including a memory and a processor. When the processor executes the computer program stored in the memory, it implements the steps of any one of the inference task execution methods in the above method embodiments.
[0011] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the inference task execution method in the above method embodiments.
[0012] Finally, the present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, it implements the steps of any one of the above inference task execution methods.
[0013] The advantages of the technical solution provided by the present invention are as follows: The first language model is used to generate a preliminary idea for task execution. Considering that it may introduce errors, the second language model is used for correction and optimization to improve the accuracy of the task execution idea, thereby facilitating the improvement of task execution accuracy. After determining the task execution idea, the first language model is used to gradually generate each step of task execution. Whenever a task execution step is generated, the second language model is used to perform real-time scoring and correction on the task execution step generated by the first language model, thereby dynamically adjusting the generation process of each task execution step to ensure its accuracy and logical coherence. Based on this dual-model collaborative reflection method, the second language model is used to perform real-time scoring and correction on the inference path generated by the first language model, and the incorrect parts are adjusted in a timely manner to ensure the authenticity of the inference chain, providing a fine-grained quality guarantee for the generated data. In addition, since the task processing model has direct inference ability through generating long-thinking training data and does not need to be called multiple times, it can not only improve the task execution efficiency but also reduce a large amount of computing costs, thereby achieving efficient and low-cost generation of high-quality inference data and effectively improving the execution accuracy and execution efficiency of complex inference tasks.
[0014] In addition, the present invention also provides corresponding implementation electronic devices, computer-readable storage media, and computer program products for the inference task execution method, further making the method more practical, and the electronic devices, computer-readable storage media, and computer program products have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic diagram of the hardware composition framework applicable to the inference task execution method provided by the present invention;
[0017] Figure 2 It is a schematic flowchart of an inference task execution method provided by the present invention;
[0018] Figure 3 It is a schematic flowchart of the dual-model collaborative generation of the task execution idea in an exemplary example provided by the present invention;
[0019] Figure 4 It is a schematic flowchart of the dual-model collaborative generation of the task execution steps in an exemplary example provided by the present invention;
[0020] Figure 5 Schematic flowchart of another method for performing an inference task provided by the present invention;
[0021] Figure 6 Schematic diagram of the generation process of an inference task training sample set in an exemplary example provided by the present invention;
[0022] Figure 7 Structural framework diagram under an exemplary embodiment of an inference task execution device provided by the present invention;
[0023] Figure 8 Structural framework diagram under another exemplary embodiment of an inference task execution device provided by the present invention;
[0024] Figure 9 Structural diagram of an exemplary embodiment of an electronic device provided by the present invention. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Among them, the terms "first", "second", "third", etc. in the specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment or illustration". Any embodiment described herein as "exemplary" should not be construed as being superior to or better than other embodiments.
[0026] With the development of natural language processing technology, the understanding ability and generation ability of language models have been greatly improved, such as in language tasks such as ordinary conversations and text generation. However, for complex inference tasks such as mathematical problem solving, code generation, and logical reasoning, these tasks require multi-step reasoning, logical deduction, and the recognition and correction of error patterns.
[0027] Related technologies improve the accuracy of inference trajectories and enhance mathematical logic reasoning capabilities by generating multi-step reasoning paths or combining search algorithms. For example, methods such as Chain of Thought, self-consistency, Tree of Thought, and reinforcement learning training are used. However, these methods lack systematic generation of high-quality reasoning data, resulting in the model training relying on limited and low-quality labeled data, restricting the generalization ability. Moreover, there are also problems of high computational cost and low efficiency. Taking the Chain of Thought method as an example, this method guides the language model to generate multi-step answers by providing examples of step-by-step reasoning, significantly improving the interpretability and answer accuracy of the language model. However, it requires a large number of high-quality reasoning chain examples, which usually need to be manually labeled, with a high cost. For models with a small number of parameters, it is difficult for this method to spontaneously generate correct reasoning chains, and the accuracy is relatively low. In addition, the reasoning chains generated by this method may have logical errors or redundant steps, and generating multi-step reasoning paths requires multiple calls to the model, resulting in high computational costs. It can be seen that related technologies cannot efficiently, low-costly, and highly accurately execute complex reasoning tasks.
[0028] In view of this, the present invention uses a first language model to generate a preliminary idea for task execution, and corrects and optimizes this idea through a second language model. After determining a task execution idea with high accuracy, the first language model is used to gradually generate each execution step for processing the task to be processed. Whenever an execution step of the task is generated, the second language model is used to perform real-time scoring and correction on the execution step of the task generated by the first language model, dynamically adjusting the generation process of each execution step of the task to ensure its accuracy and logical coherence. Based on this dual-model collaborative reflection method, complex reasoning tasks can be executed efficiently, with high precision, and at low cost.
[0029] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the reasoning task execution method depends, the specific application environment architecture or specific hardware architecture is described herein. The following is combined with Figure 1 Some possible application scenarios related to the technical solution of the present invention are introduced by way of example, which may include the following content:
[0030] Such as Figure 1As shown in the figure, the hardware composition framework may include a first electronic device 11 and a second electronic device 12, which are connected by a network 13. The first electronic device 11 deploys a processor for the inference task execution method described in the embodiments of the present invention. The second electronic device 12 may be a user terminal deployed to provide a human-computer interaction interface, such as a smart phone, a tablet, and a personal computer. Among them, the human-computer interaction interface may be an interface of a corresponding application software, or an interface opened in a browser through a specified URL (Uniform Resource Locator). One of the application scenarios of the embodiments of the present invention can be realized through the interaction between the second electronic device 12 and the user. In this application scenario, the user can input a task to be processed through the human-computer interaction interface, and the human-computer interaction interface sends the task to be processed to the first electronic device 11 through the network 13. The first electronic device 11 may be a server, for example. The first electronic device 11 processes the task to be processed by performing all or part of the steps in the inference task execution described in the method embodiments of the present invention, and combines the initial task execution idea, the idea analysis result, the task execution idea, each initial task execution step, the single-step analysis result, and each task execution step into a task processing result, and sends it to the second electronic device 12 through the network 13. The task processing result can be displayed at the corresponding position of the human-computer interaction interface, or can be displayed to the user in other ways such as voice and text form.
[0031] It should be noted that the above application scenario is only shown for the convenience of understanding the idea and principle of the present invention, and the embodiments of the present invention are not limited in this regard. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, the various non-limiting embodiments of the present invention will be described in detail below with reference to the drawings and specific embodiments.
[0032] First, please refer to Figure 2 , Figure 2 which is a schematic flowchart of an inference task execution method provided in this embodiment. This embodiment may include the following content:
[0033] S201: Input the task to be processed and the prompt word into a pre-trained task processing model.
[0034] In this step, the tasks to be processed are tasks that require the execution of complex reasoning tasks, including but not limited to mathematical and logical reasoning related question-and-answer tasks, code generation tasks, and legal reasoning tasks. Mathematical and logical reasoning related question-and-answer tasks such as math problem solving tasks or word problem solving tasks. The tasks to be processed can be input in any form, such as text, tables, graphics, videos, and voices, and at least include questions, which are, for example, math problems described in visual form and / or language form. The questions in this embodiment support any form of expression, such as text form, image form, or voice form, and the stem and the picture may appear at the same time. The prompt word is a prompt word that tells the task processing model how to perform the task, and can be flexibly set according to the actual application scenario.
[0035] Among them, the task processing model is a pre-trained language model. The network model structure of the task processing model of the present invention includes at least a first language model and a second language model. On this basis, when the task to be processed is allowed to be input in multiple ways, a corresponding encoding module can be set in the input layer. If voice input is supported, a network model for voice-to-text or voice-to-image conversion needs to be added to the input layer of the task processing model. If visual or text information input is supported, an image encoder and a text encoder need to be added to the input layer of the task processing model. The first language model and the second language model are language models that use different network model structures and can process natural language tasks. The present invention does not limit the model types of the first language model and the second language model. The scale or parameter amount of the first language model and the second language model may be the same or different, which does not affect the implementation of the present invention. In order to reduce the resources required for the task processing model, the performance of the first language model and the second language model of the present invention are different. The first language model can select a language model with a parameter amount of less than 7b, such as LLAMA2-7B (name of the artificial intelligence network model) or GPT-3.5 (name of the artificial intelligence network model), and the second language model can select a language model with a parameter amount of more than 7b, for example, GPT-4 (name of the artificial intelligence network model) or Claude-3 (name of the artificial intelligence network model). The task processing model is built using the first language model and the second language model, and the task processing model is trained using any model training method using a training sample data set that matches the task to be processed, and the trained task processing model is used to execute the task to be processed. Of course, if the accuracy requirements of the task to be processed are high and the available computing resources and storage resources are relatively abundant, the task processing model may also include a larger-scale third language model or more language models, and the third language model is used to reflect and correct the output of the second language model, and the quality and diversity of the reasoning chain are further improved through the collaboration of three models or more models.
[0036] S202: Generate an initial task execution idea corresponding to the task to be processed using the first language model according to the prompt word, analyze the initial task execution idea using the second language model and the task to be processed, and determine the task execution idea based on the initial task execution idea and the idea analysis result using the first language model.
[0037] Among them, the initial task execution idea refers to the first language model's idea of how to execute the task to be processed. Since the task execution idea output by the first language model may introduce errors, in order to avoid description, it can be defined as the initial task execution idea. In order to determine whether there are errors in the initial task execution ideas, based on the reflection mechanism in the human thinking process, the second language model can be used to analyze the initial task execution ideas, thereby simulating the human thinking process from intuitive judgment to in-depth deliberation. For the convenience of description, the analysis result of the initial task execution ideas by the second language model is defined as the idea analysis result. The idea analysis result at least includes relevant information on whether the initial task execution ideas are wrong and how to improve the initial task execution ideas. The task execution ideas used to assist in generating the execution task steps are finally determined by the initial task execution ideas and the idea analysis results. If the idea analysis results determine that the initial task execution ideas are highly accurate and there is no need for improvement, the initial task execution ideas are the task execution ideas. If the idea analysis results determine that the initial task execution ideas are not accurate and / or there are areas that need improvement, the initial task execution ideas are adjusted according to the idea analysis results, that is, the first language model regenerates the initial task execution ideas based on the idea analysis results and the tasks to be processed. The adjusted initial task execution ideas or the newly generated initial task execution ideas are the final task execution ideas, thereby improving the logical rigor and practicality of the generated data.
[0038] S203: According to the prompt word, the single-step analysis result obtained by analyzing the single initial task execution step output by the first language model based on the second language model is analyzed, and each task execution step is gradually generated according to the task execution idea using the first language model.
[0039] In this embodiment, the first language model is used to obtain the task execution idea according to the above steps, and each task execution step for executing the to-be-processed task is generated step by step. For the convenience of description, it is defined as the initial task execution step. By gradually generating the reasoning chain, the logical rigor and integrity of such complex reasoning tasks can be ensured. To improve the task execution accuracy, whenever the first language model generates an initial task execution step, the second language model is used to analyze the initial task execution step, and the analysis result is defined as the single-step analysis result. The single-step analysis result at least includes whether there is an error in the initial task execution step and relevant information on how to improve the initial task execution step. The final output task execution steps are jointly determined based on the initial task execution steps output by the first language model and the single-step analysis result. If the single-step analysis result determines that the accuracy of the initial task execution step is high and there is no need for improvement, then the initial task execution step is the final task execution step. If the single-step analysis result determines that the accuracy of the initial task execution step is not high and / or there is a need for improvement, then the initial task execution step is adjusted according to the single-step analysis result, that is, the first language model regenerates an initial task execution step based on the single-step analysis result and the to-be-processed task. The adjusted initial task execution idea, that is, the newly generated initial task execution step, is the final task execution step for this step, simulating the human's step-by-step derivation thinking mode when solving complex problems, thereby improving the logical rigor and practicality of the generated data. According to actual needs, the task processing model can output the final answer obtained by integrating each task execution step as the final task processing result of the to-be-processed task. Of course, the initial task execution idea, the idea analysis result, the task execution idea, each initial task execution step, the single-step analysis result, each task execution step, and the final answer can also be combined into a reasoning data chain, and the reasoning data chain is output as the final task processing result of the to-be-processed task. Of course, the data content included in the reasoning data chain can be flexibly adjusted according to needs, which does not affect the implementation of the present invention.
[0040] In the technical solution provided in this embodiment, a preliminary idea for task execution is generated by using a first language model. Considering that it may introduce errors, a second language model is used for correction and optimization to improve the accuracy of the task execution idea, thereby facilitating the improvement of the task execution accuracy. After determining the task execution idea, each step of task execution is gradually generated by using the first language model. Whenever a task execution step is generated, the second language model is used to perform real-time scoring and correction on the task execution step generated by the first language model, thereby dynamically adjusting the generation process of each task execution step to ensure its accuracy and logical coherence. Based on this dual-model collaborative reflection method, the second language model is used to perform real-time scoring and correction on the inference path generated by the first language model, and the incorrect parts are adjusted in a timely manner to ensure the authenticity of the inference chain, providing a fine-grained quality guarantee for the generated data. In addition, since the task processing model is given direct inference ability by generating long-thought training data and does not need to be called multiple times, it can not only improve the task execution efficiency but also reduce a large amount of computing costs, thereby achieving efficient and low-cost generation of high-quality inference data and effectively improving the execution accuracy and execution efficiency of complex inference tasks.
[0041] In the above embodiment, there is no limitation on how to analyze the initial task execution idea. Based on the above embodiment, the present invention also provides an exemplary implementation manner of using a second language model and a task to be processed to analyze the initial task execution idea and generate an idea analysis result, which may include the following contents:
[0042] Use the second language model and the idea analysis prompt words to score the initial task execution idea, and output task execution idea modification opinions according to the task to be processed and the initial task execution idea.
[0043] In this embodiment, whether there are errors in the initial task execution idea can be reflected in a quantitative form, that is, by scoring. The prompt words in the above embodiment at least include idea analysis prompt words, and the idea analysis prompt words are used to prompt the second language model to score the initial task execution idea and give idea modification opinions. The number of parameters of the first language model and the second language model is different, and the number of parameters of the second language model is greater than that of the first language model. Correspondingly, the model performance of the second language model is better than that of the first language model, and the output of the first language model can be effectively corrected by the second language model. Using a weak model to execute tasks and a strong model to assist can also reduce the resources required by the task processing model and improve the task execution efficiency.
[0044] When the second language model outputs suggestions for modifying the task execution idea, if the score of the initial task execution idea is lower than the preset qualified score threshold for the idea, the first language model regenerates the task execution idea according to the suggestions for modifying the task execution idea; if the score of the initial task execution idea is not lower than the preset qualified score threshold for the idea, the initial task execution idea is used as the task execution idea.
[0045] Among them, the preset qualified score threshold for the idea can be flexibly set according to actual needs. For example, it can be 60 points. For example, the task to be processed is a math problem-solving task. Correspondingly, the task to be processed includes at least the problem. The first language model is a weak model, and the second language model is a strong model. Figure 3For example, the thought prompt for the weak model is "Analyze the above problem and generate a solution idea. Do not answer the problem and do not generate specific solution steps". Based on the input problem, the weak model generates a preliminary solution idea according to this prompt. For example, if the math problem is "Solve the equation x^2 - 5x + 6 = 0", the preliminary solution idea generated by the weak model is "Use the factorization method to factor the equation into two linear factors and then solve for the roots". The thought prompt for the strong model is "Score the above solution idea according to its accuracy, with the score range from 0 to 10. The more accurate the idea, the higher the score, and 6 is the passing score. Output the scoring result in the format of 'What is the score for the above analysis' and give modification suggestions. Do not generate the modified solution idea. Do not answer the problem and do not generate specific solution steps". Based on this thought prompt, the strong model scores and reflects on the preliminary solution idea generated by the weak model, and the output task execution thought modification suggestion is: "The score for the above analysis is 8 points. Modification suggestion: The factorization method is applicable to this equation, but it is necessary to clearly point out that the equation can be factored into (x - 2)(x - 3) = 0 and explain how to obtain the roots through factorization". If the task execution thought modification suggestion is: "The score for the above analysis is 5 points. Modification suggestion: The factorization method is applicable to this equation, but it is necessary to clearly point out that the equation can be factored into (x - 2)(x - 3) = 0 and explain how to obtain the roots through factorization", then the weak model will regenerate the solution idea according to the task execution thought modification suggestion of the strong model. When regenerating the task execution thought again, the thought prompt for the weak model can be "Generate a new solution idea according to the modification suggestions for the solution idea in the above text. Do not answer the problem and do not generate specific solution steps". Correspondingly, the corrected solution idea generated by the weak model can be: "Use the factorization method to factor the equation x^2 - 5x + 6 = 0 into (x - 2)(x - 3) = 0, and then obtain the roots by solving (x - 2) = 0 and (x - 3) = 0". Finally, the "Problem - Solution Idea - Thought Reflection - Thought Correction" can be saved as structured data. For example, for the above math problem, the saved data format is: Problem: Solve the equation x^2 - 5x + 6 = 0. Solution Idea: Use the factorization method to factor the equation into two linear factors and then solve for the roots. Thought Reflection: The score for the above analysis is 6 points. Modification suggestion: The factorization method is applicable to this equation, but it is necessary to clearly point out that the equation can be factored into (x - 2)(x - 3) = 0 and explain how to obtain the roots through factorization. Thought Correction: Use the factorization method to factor the equation x^2 - 5x + 6 = 0 into (x - 2)(x - 3) = 0, and then obtain the roots by solving (x - 2) = 0 and (x - 3) = 0.
[0046] As can be seen from the above, in this embodiment, the weak model generates a preliminary idea, and then the strong model corrects and optimizes it through scoring and modification opinions, guiding the weak model to gradually correct errors, and finally generating inference data with rigorous logic and clear structure. By collaborating with the strong and weak models to simulate the human thinking process from intuitive judgment to in-depth deliberation, it can not only generate high-quality task execution idea data, effectively improve the logical rigor and practicality of the generated data, but also provide rich reflective learning samples for subsequent model training, significantly enhancing the model's inference ability and dynamic correction ability.
[0047] In the above embodiment, there is no limitation on how to analyze the initial task execution steps. Based on the above embodiment, the present invention also provides an exemplary implementation method for step reflection and correction, which may include the following:
[0048] According to the task to be processed and the task execution idea, use the first language model to generate the current task execution steps for executing the task to be processed according to the generated content and output them in the generated format; use the task to be processed, the task execution idea, and the current task execution steps as step reflection data and input them into the second language model; use the second language model to score the current task execution steps according to the step reflection prompt words and give step modification opinions; if the score of the current task execution steps is lower than the preset qualified score threshold for steps, the first language model regenerates the current task execution steps based on the step regeneration prompt words according to the step modification opinions.
[0049] In this embodiment, the current task execution steps may refer to any task execution step. The prompt words at least include the generated content, the generated format, and the step reflection prompt words. The generated content is used to limit the content output by the first language model and the second language model to prevent the output of irrelevant content or redundant information. The generated format is used to limit the format of the content output by the first language model and the second language model. In order to better process the data, the first language model will not generate multiple task execution steps or directly output the final answer at one time, but gradually generate strictly according to the requirements of the prompt words. The step reflection prompt words are used to prompt the second language model to score the last task execution step output by the first language model and give step modification opinions. For example, in the code generation task, the second language model may point out that the variable naming of a certain step is not standard or the logic is not clear, and the first language model can regenerate a better code segment according to the feedback.
[0050] In this embodiment, the process of using the task processing model to execute the task to be processed may be as follows: According to the task to be processed and the task execution idea, use the first language model to generate the first task execution step for executing the task to be processed according to the generated content, and output it according to the generated format; Use the task to be processed, the task execution idea, and the first task execution step as step reflection data and input them into the second language model; Use the second language model to score the first task execution step according to the step reflection prompt, and give step modification opinions; If the score of the first task execution step is lower than the preset step passing score threshold, the first language model regenerates the first task execution step according to the step modification opinions; If the score of the first task execution step is greater than or equal to the preset step passing score threshold, the first language model does not need to regenerate the first task execution step according to the step modification opinions. Then, according to the task to be processed and the task execution idea, use the first language model to generate the second task execution step for executing the task to be processed according to the generated content, and output it according to the generated format; Use the task to be processed, the task execution idea, the newly generated first task execution step, or the first task execution step and the second task execution step as step reflection data and input them into the second language model; Use the second language model to score the second task execution step according to the step reflection prompt, and give step modification opinions; If the score of the second task execution step is lower than the preset step passing score threshold, the first language model regenerates the second task execution step according to the step modification opinions; If the score of the second task execution step is greater than or equal to the preset step passing score threshold, the first language model does not need to regenerate the second task execution step according to the step modification opinions. Repeat the above process until the number of currently generated steps is equal to the total number of step generations or the currently generated step is the last step, and then output all the currently generated task execution steps as the task processing result.
[0051] To make those skilled in the art more clearly understand the implementation process of the above embodiment, taking the task to be processed as a math problem-solving task as an example, correspondingly, the task to be processed includes at least a problem, the first language model is a weak model, and the second language model is a strong model, such as Figure 4As shown, according to the input question corresponding to the task to be processed and the task execution idea, following the prompt "According to the above solution idea, generate the first step of the problem-solving steps, and start with'step 1,'. Only generate one step of the problem-solving steps", call the weak model to generate the first step of the problem-solving steps for the math problem-solving task. For example, the math problem-solving task is "Solve the equation \(x^{2}-5x + 6 = 0\)", and the first problem-solving step generated by the weak model is "Step 1, factor the equation \(x^{2}-5x + 6 = 0\) into \((x - 2)(x - 3)=0\)". Combine the input question, the task execution idea, and the existing step-by-step solution, that is, the first problem-solving step, and follow the prompt "Do not score the solution idea. Score the last problem-solving step according to the solution idea, in the range of 0 to 10 points. The more accurate the step, the higher the score, and 6 points is the passing line. Output the scoring result in the format 'How many points is the score of the problem-solving step' and give the modification opinion. Do not generate new problem-solving steps", call the strong model to score and reflect on the last problem-solving step. For example, the math problem-solving task is "Solve the equation \(x^{2}-5x + 6 = 0\)". If the first problem-solving step generated is "Step 1, factor the equation \(x^{2}-5x + 6 = 0\) into \((x - 2)(x - 3)=0\)". If the strong model scores this step and outputs the step modification opinion: "The score of the problem-solving step is 9 points. Modification opinion: The factorization step is correct, but it can be supplemented with an explanation of how to obtain \((x - 2)(x - 3)=0\) through factorization", then there is no need to regenerate. If the strong model scores this step and outputs the step modification opinion as: "The score of the problem-solving step is 5 points. Modification opinion: The factorization step is incomplete, and it is necessary to clearly explain how to obtain \((x - 2)(x - 3)=0\) through factorization", then follow the prompt "According to the modification opinion of the problem-solving step in the above text, regenerate this problem-solving step. Do not generate subsequent steps", call the weak model to regenerate this problem-solving step according to the step modification opinion. The first problem-solving step regenerated by the weak model is "Step 1, factor the equation \(x^{2}-5x + 6 = 0\) into \((x - 2)(x - 3)=0\), and the specific method is to find two numbers whose product is 6 and sum is -5, getting -2 and -3". When the first problem-solving step is completed, according to the existing problem-solving steps, such as the first \(n - 1\) steps, follow the prompt "According to the above solution idea and the existing problem-solving (first \(n - 1\) steps) steps, generate the next (the \(n\)th) problem-solving step, and start with'step n,'. If all the problem-solving steps already exist, return the final answer in the form of a box. Only generate one step of the problem-solving step", call the weak model to generate the next problem-solving step. For example, based on the existing first problem-solving step, the second problem-solving step generated by the weak model can be "Step 2, solve the equation \((x - 2)=0\) to get \(x = 2\)". If the problem-solving steps are completed at this time, the weak model will generate the final answer and return it in the form of a box. For example, the output can be: "Final answer: \(\boxed{2},\boxed{3}\)".
[0052] As can be seen from the above, in each step of the generation process of this embodiment, by restricting the generated content and format, it is ensured that only one clear task execution step is output in each step, avoiding the generation of redundant or irrelevant information. This not only improves the controllability of the generated data but also ensures the logical coherence of the reasoning chain. Further, through step-level reflection and correction, the generation process of each step of the problem-solving steps can be dynamically adjusted. By simulating the self-checking and correction of each step of derivation when humans solve problems, its accuracy and logical coherence are ensured, thus generating high-quality reasoning chain data. The scoring and feedback of the second language model not only help the weak model correct errors but also provide fine-grained quality assurance for the generated data. That is to say, this embodiment uses a relatively weak large model and a relatively strong large model to collaborate to generate a task execution idea and task execution steps that include the reflection and correction process, which is more in line with the thinking process of humans when solving complex problems. This process ensures the correctness and integrity of the reasoning process by cyclically calling the model to control the generation of step-level answers and finally outputting the final answer.
[0053] To further improve the execution efficiency of the reasoning task, based on the above embodiment, the present invention also limits the total number of steps generated in the prompt. Under the limitation of the total number of steps generated, each task execution step is gradually generated, which may include the following content:
[0054] If the current number of task execution steps is greater than the total number of steps generated, then all the currently generated task execution steps are output as the task execution result; if the current number of task execution steps is less than or equal to the total number of steps generated, according to the task to be processed and the task execution idea, the first language model is used to generate the current task execution step for executing the task to be processed according to the generated content in the prompt and output it according to the generated format in the prompt; the task to be processed, the task execution idea, and the current task execution step are used as step reflection data and input into the second language model. When the score of the current task execution step is less than the preset step passing score threshold, the first language model is used to regenerate the current task execution step; if the current task execution step is the last step, then all the currently generated task execution steps are output as the task execution result; if the current task execution step is not the last step, then according to the task to be processed and the task execution idea, the first language model is used to generate the next task execution step of the current task execution step according to the generated content.
[0055] In this embodiment, the step of generating the total number is used to limit the maximum number of generated steps. If the number of steps executed in the current task reaches the total number of generated steps, the generation process will be forcibly terminated and the current result will be returned. For example, if the total number of generated steps is 10 steps and the number of steps executed in the current task is 10 steps, the generation process will be forcibly terminated and the current result will be returned. The process of executing the task to be processed in combination with the total number of generated steps in the prompt can be as follows: According to the task to be processed and the task execution idea, use the first language model to generate the first task execution step for executing the task to be processed according to the generated content, and output it in the generated format; Use the task to be processed, the task execution idea, and the first task execution step as step reflection data and input it into the second language model; Use the second language model to score the first task execution step according to the step reflection prompt and give step modification opinions; If the number of steps executed in the current task is not greater than the total number of generated steps, according to the task to be processed and the task execution idea, use the first language model to generate the second task execution step for executing the task to be processed according to the generated content, and output it in the generated format; Use the task to be processed, the task execution idea, the newly generated or initial first task execution step, and the second task execution step as step reflection data and input it into the second language model; Use the second language model to score the second task execution step according to the step reflection prompt and give step modification opinions; If the number of steps executed in the current task is equal to the total number of generated steps, the newly generated first task execution step and the second task execution step will be output as the task execution result.
[0056] As can be seen from the above, in this embodiment, by dynamically controlling the number of generated steps, it is possible to prevent the occurrence of infinite generation of steps or redundant steps, while ensuring the generation efficiency, avoiding step omission or logical errors caused by insufficient generation ability of the task processing model. In addition, by dynamically adjusting the total number of generated steps, it is possible to adapt to different task complexity requirements and has strong scalability.
[0057] In order to further improve the execution efficiency of the inference task, based on the above embodiment, the present invention also limits the number of reflections in the prompt, and gradually generates each task execution step under the limitation of the number of reflections, which may include the following content:
[0058] If the number of times the current task execution step generated by the first language model is less than the number of reflections + 1, then the task to be processed, the task execution idea, and the current task execution step are used as step reflection data and input into the second language model; the second language model is used to score the current task execution step according to the step reflection prompt words and give step modification opinions; if the score of the current task execution step is lower than the preset step passing score threshold, the first language model regenerates the current task execution step according to the step modification opinions; if the number of times the current task execution step generated by the first language model is equal to the number of reflections + 1 and the current task execution step is not the last step, then according to the task to be processed and the task execution idea, the first language model is used to generate the next task execution step of the current task execution step according to the generated content in the prompt words.
[0059] In this embodiment, by restricting the number of reflections, such as up to 3 times, it is possible to effectively prevent the task processing model from falling into an infinite loop of "reflection - modification" during the execution of the task to be processed. If the preset step passing score threshold cannot be reached after multiple corrections, the current step can be recorded and the next step can be continued to generate to ensure the integrity of the overall inference chain. The process of executing the task to be processed by combining the total number of steps generated in the prompt and the number of reflections can be as follows: According to the task to be processed and the task execution idea, use the first language model to generate the first task execution step for executing the task to be processed according to the generated content and output it in the generated format; Use the task to be processed, the task execution idea, and the first task execution step as step reflection data and input them into the second language model; Use the second language model to score the first task execution step according to the step reflection prompt and give step modification opinions; Use the first language model to generate a new first task execution step again based on the step modification opinions. When the number of generations of the first task execution step is less than the number of reflections + 1, use the task to be processed, the task execution idea, and the current task execution step as step reflection data and input them into the second language model again; If the number of generations of the first task execution step is equal to the number of reflections + 1 and the current task execution step is not the last step, then according to the task to be processed and the task execution idea, use the first language model to generate the next task execution step of the current task execution step according to the generated content in the prompt. If the number of current task execution steps is not greater than the total number of steps generated, according to the task to be processed and the task execution idea, use the first language model to generate the second task execution step for executing the task to be processed according to the generated content and output it in the generated format; Use the task to be processed, the task execution idea, the newly generated or initial first task execution step, and the second task execution step as step reflection data and input them into the second language model; Use the second language model to score the second task execution step according to the step reflection prompt and give step modification opinions; If the number of current task execution steps is equal to the total number of steps generated, then output the newly generated first task execution step and the second task execution step as the task execution result.
[0060] As can be seen from the above, in this embodiment, by restricting the number of reflections, while ensuring the generation efficiency, it is possible to avoid step omissions or logical errors caused by insufficient generation capabilities of the task processing model. In addition, by dynamically adjusting the number of generation reflections, it is possible to adapt to different task complexity requirements and has strong scalability.
[0061] The present invention also provides another method for executing an inference task, such as Figure 5 shown, which may include the following contents:
[0062] S501: Input the inference task training sample set and the prompt into a pre - built task processing model.
[0063] In this embodiment, the task processing model may include a first language model and a second language model that have completed pre-training. Pre-training is a strategy for training deep learning models, which uses large-scale datasets to preliminarily train the models so that the first language model and the second language model learn general feature representations. This process is similar to the basic learning stage of humans before learning new knowledge, where they accumulate experience through extensive reading and observation. The first language model and the second language model that have completed pre-training refer to designing language model training tasks based on large-scale corpora (including language training materials such as sentences and paragraphs), training large-scale neural network algorithm structures to learn and implement, and the final large-scale neural network algorithm structures and parameters are pre-trained language models. Subsequently, for other tasks, feature extraction or task fine-tuning can be performed on the basis of this model to achieve specific task purposes. The idea of pre-training is to first train a task to obtain a set of model parameters, then use this set of model parameters to initialize the network model parameters, and then use the initialized network model to train other tasks to obtain models adapted to other tasks. By pre-training on a large-scale corpus, the neural language representation model can learn powerful language representation capabilities and extract rich syntactic and semantic information from the text. The pre-trained language model can provide word elements (tokens) containing rich semantic information and sentence-level features for downstream tasks, or directly perform fine-tuning for downstream tasks on the pre-trained model to conveniently and quickly obtain downstream-specific models. The neural network algorithm structures of the first language model and the second language model can be CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), etc., or models constructed with attention networks, such as transformer (transformer network model), bert (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), Clip (Contrastive Language–Image Pre-training), etc. This application does not make any limitations here. An attention network refers to a network model that uses the attention mechanism for training. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence and enabling the model to finally obtain a more accurate output.
[0064] Among them, the training sample set for the reasoning task is a training sample set matching the task to be processed, which may include a large number of reasoning task data samples to provide high-quality and diverse data support for subsequent model training. Each reasoning task data sample has a label, and the labels of each reasoning task data sample at least include the question, the task execution idea, the idea adjustment process, the task execution steps, the step reflection process, and the task execution result. The reasoning task data samples can be obtained, for example, from a mathematics competition data set, an open-source data set, code generation task data, and a high-quality mathematics corpus. For example, the mathematics competition data can be crawled from international mathematics competition websites, covering multiple fields such as algebra, geometry, and combinatorics; for the open-source data set, public resources such as GSM8K (data set name) and MATH (data set name) can be used, covering a wide range of mathematics problems and their solutions; the code generation task data is collected from platforms such as LeetCode (code library name), ATCoder (code library name), and Codeforces (code library name), covering topics such as algorithms and data structures; in addition, mathematics corpora such as MathPile (corpus name) and OpenWebMath (corpus name) can also be used to extract high-quality mathematics texts such as textbooks and papers to ensure the diversity and professionalism of the data.
[0065] S502: According to the prompt words, use the first language model to generate the initial idea prediction data corresponding to the question sample of the current reasoning task data sample, and use the second language model and the question sample to analyze the initial idea prediction data; and use the first language model to determine the idea prediction data based on the initial idea prediction data and the idea analysis prediction result.
[0066] In this embodiment, for the sake of distinction, during the training process of the task processing model, the task execution idea first output by the first language model is defined as the initial idea prediction data, the idea analysis result generated by the second language model analyzing the initial idea prediction data is defined as the idea analysis prediction result, and the finally determined task execution idea is defined as the idea prediction data.
[0067] S503: According to the prompt words, based on the single-step analysis prediction result obtained by the second language model analyzing the single initial task execution prediction step output by the first language model, use the first language model to gradually generate the task execution prediction steps according to the idea prediction data, and use the question sample, the initial idea prediction data, the idea analysis prediction result, each initial task execution prediction step, the single-step analysis prediction result, and each task execution prediction step as the prediction data.
[0068] In this embodiment, for the convenience of distinction, during the training process of the task processing model, the task execution steps output by the first language model for the first time are defined as the initial task execution prediction steps, the step analysis results generated by the second language model analyzing the initial task execution prediction steps are defined as the single-step analysis prediction results, and the finally determined task execution steps are defined as the task execution prediction steps. For example, taking the math problem-solving task as an example, the prediction data may include the problem, the problem-solving idea, the reflection on the idea, the problem-solving steps, the reflection on the steps, and the final answer. Such long-thinking data can effectively enhance the reasoning ability of the task processing model.
[0069] S504: Continuously train the task processing model based on the labels of each inference task data sample and its corresponding prediction data until the preset model training stop condition is met.
[0070] Among them, the first language model and the second language model can be directly fine-tuned through the labels of each inference task data sample and its corresponding prediction data. Fine-tuning means that on the basis of using the first language model and the second language model, for the downstream task of complex reasoning, further small-scale training is carried out using the downstream data of each inference task data sample to adjust the model parameters so that it can better adapt to complex reasoning tasks, and finally a model adapted to the complex reasoning task and the inference task data sample set is obtained. During the fine-tuning process, most layers of the pre-trained model, that is, the first language model and the second language model, are usually frozen, and only the newly added layers are trained or a small number of key layers are adjusted. This can not only retain the features learned by the pre-trained model but also quickly adapt to the specific requirements of the new task. In addition, choosing the appropriate learning rate and number of training epochs is also the key to the success of fine-tuning. The preset training stop condition can be, for example, that the number of iterations reaches the preset iteration threshold, or the model has converged, or the prediction accuracy reaches the preset accuracy threshold, which does not affect the implementation of the present invention.
[0071] As can be seen from the above, in this embodiment, the task processing model is trained by generating long-thinking data, endowing the task processing model with direct reasoning ability. Compared with other reasoning enhancement methods that call the model multiple times, it can reduce a large amount of computational cost. Further, through the reflection method of dual-model collaboration, logical errors can be generated and corrected, generating a clearer reasoning chain that is more in line with the human thinking process, and improving the performance of the task processing model in executing complex reasoning tasks. Further, through training with long-thinking data covering multiple fields such as mathematics and code, the task processing task performs better in complex tasks and has enhanced generalization ability.
[0072] It should be noted that there is no strict order of execution between the steps in the present invention. As long as it conforms to the logical order, these steps can be executed simultaneously or in a certain preset order. Figure 2 and Figure 5This is just a schematic way and does not mean that the execution order can only be like this.
[0073] To further improve the performance of the task processing model, based on the above embodiments, the present invention can also preprocess the inference task training sample set, and the preprocessing process may include:
[0074] Whenever an original inference task data sample is obtained, check whether there is a sample data in the current inference task training sample set that is the same as the original inference task data sample; if there is no sample data in the current inference task training sample set that is the same as the original inference task data sample, and if the original inference task data sample has missing values, fill the original inference task data sample according to the data type of the original inference task data sample; if the original inference task data sample has outliers, adjust the original inference task data sample according to the outlier type to obtain an inference task data sample; convert the inference task data sample into a target format and store it in the current inference task training sample set.
[0075] In this embodiment, data cleaning is performed on the obtained original inference task data sample, including but not limited to missing value processing, outlier detection and processing, and duplicate data deletion. For example, the Pandas library (database name) can be used to detect and process missing values, fill numerical data with the mean value, and fill categorical data with the mode; the IQR (Interquartile Range) method is used to detect outliers, and data outside the range is corrected or deleted; the drop_duplicates() method (function name) of Pandas is used to delete duplicate records and only keep the first occurrence to ensure data uniqueness. For the sake of management, the inference task data samples obtained after data cleaning can be unified in format, such as converting fields such as dates and numerical values into a unified data type and unifying the formats of fields such as dates and currencies to ensure data readability and consistency, as Figure 6 shown.
[0076] As can be seen from the above, in this embodiment, by performing data cleaning and format unification on the inference task data samples, the readability and consistency of the data are ensured, which is beneficial to obtaining high-quality and diverse inference task data samples and improving the performance and training efficiency of the task processing model.
[0077] Considering that a large amount of storage resources are required during the model training process, to avoid low or even failed model training efficiency due to insufficient storage resources, based on the above embodiments, the inference task data samples can also be stored and managed in the following ways, which may include the following contents:
[0078] Obtain data reading parameters; according to the data reading parameters, read the corresponding number of data samples from the inference task training sample set and load them into the memory as a data sample block.
[0079] Among them, the data reading parameters include the number of rows read each time. For example, the chunksize parameter (parameter name) can be set, and the chunksize parameter is used to specify the number of rows read each time. When chunksize is set, the read_csv() method (function name) will return an iterable object, and the data can be processed block by block by iterating over this object. The chunked data of the inference task data samples can be uniformly stored in CSV (format name), JSON (format name) or database format for subsequent use.
[0080] As can be seen from the above, in this embodiment, reading data in chunks can avoid memory overflow and ensure the training effect of the task processing model.
[0081] In order to further improve the performance of the task processing model, based on the above embodiment, the present invention also provides a sample screening embodiment, which may include the following content:
[0082] If the current inference task data sample belongs to the first task type, then when the task execution result of the current inference task data sample is not completely consistent with each task execution prediction step, mark the current inference task data sample as an error sample and delete the current inference task data sample from the inference task training sample set; if the current inference task data sample belongs to the second task type, then when the task execution result of the current inference task data sample does not match each task execution prediction step, and / or each task execution prediction step fails to pass the corresponding correctness verification condition, mark the current inference task data sample as an error sample and delete the current inference task data sample from the inference task training sample set.
[0083] In this embodiment, the prediction data can be saved in a structured format, which may include six fields: "question - problem-solving idea - idea reflection - problem-solving steps - step reflection - final answer". For example, if the task to be processed is "Solve the equation x^2 - 5x + 6 = 0", the data format of the prediction data is as follows: Question: Solve the equation x^2 - 5x + 6 = 0. Initial idea prediction data: Use the factorization method to factor the equation into two linear factors and then solve for the roots. Idea analysis prediction result: The above analysis score is 8 points. The factorization method is applicable to this equation, but it is necessary to clearly indicate that the equation can be factored into (x - 2)(x - 3) = 0 and explain how to obtain the roots through factorization. Task execution steps: Step 1, factor the equation x^2 - 5x + 6 = 0 into (x - 2)(x - 3) = 0; Step 2, solve the equation (x - 2) = 0 to get x = 2; Step 3, solve the equation (x - 3) = 0 to get x = 3. Step reflection: The score for Step 1 is 9 points, the score for Step 2 is 10 points, and the score for Step 3 is 10 points. Final answer: \(\boxed{2},\boxed{3}\).
[0084] Among them, the first type of task is a task that does not require verifying the effect, but only needs to compare whether the correct answer in the label is consistent with the predicted answer, such as a math problem-solving task. The second type of task is a task that not only compares whether the correct answer in the label is consistent with the predicted answer, but also needs to verify the effect of the generated content, such as a code generation task. For example, if the task to be processed is a math problem-solving task and the inference task training sample is math-type data, only the predicted answer generated needs to be compared with the original answer in the label. For example, the content inside the box in the finally generated predicted answer (such as \(\boxed{2},\boxed{3}\)) can be extracted and compared with the original answer in the label. If the two are exactly the same, this piece of data is retained; otherwise, it is marked as incorrect and deleted from the inference task training sample set. For example, if the original answer in the label is \(\boxed{2},\boxed{3}\) and the generated predicted answer is \(\boxed{2},\boxed{4}\), then this piece of data will be filtered out. If the task to be processed is a code generation task, it is also necessary to verify the correctness of the generated code through the test cases in the label. That is to say, the generated code will be compared with the code in the label, or matched with the test cases, and the code will be run to check whether it passes all the test cases. If the generated code can pass all the test cases, this piece of data is retained; if the generated code cannot pass all the test cases, it is marked as incorrect and excluded. For example, if the label contains the test case "Input: [1, 2, 3] Output: 6" and the generated code outputs 5 when the input is [1, 2, 3], then this piece of data will be deleted from the inference task training sample set.
[0085] For example, through the above method, the prediction data is saved in a structured format of JSON:
[0086] ```JSON
[0087] {
[0088] "problem (i.e., the problem)": "Solve the equation \(x^2 - 5x + 6 = 0\)",
[0089] "solution_idea (i.e., the initial idea prediction data)": "Use the factorization method to factor the equation into two linear factors and then solve for the roots.",
[0090] "idea_feedback" (i.e., the idea analysis prediction result): "The above analysis score is 8 points. Modification suggestion: The factorization method is applicable to this equation, but it is necessary to clearly point out that the equation can be factored into \((x - 2)(x - 3)=0\) and explain how to obtain the roots through factorization.",
[0091] "steps (i.e., the initial task execution prediction steps)":
[0092] {"step": "Step 1 Factor the equation \(x^2 - 5x + 6 = 0\) into \((x - 2)(x - 3)=0\)", "step_feedback" (i.e., the single-step analysis prediction result): "The score for Step 1 is 9 points."},
[0093] {"step": "Step 2 Solve the equation \((x - 2)=0\) to get \(x = 2\)", "step_feedback": "The score for Step 2 is 10 points."},
[0094] {"step": "Step 3 Solve the equation \((x - 3)=0\) to get \(x = 3\)", "step_feedback": "The score for Step 3 is 10 points."}
[0095] ,
[0096] "final_answer" (i.e., the task processing result): "\(\boxed{2},\boxed{3}\)"
[0097] }
[0098] ```。
[0099] As can be seen from the above, based on using the generated long-thinking data to improve the reasoning ability of the task processing model in complex tasks such as mathematics and code, the correct long-thinking data is filtered out by comparing the generated final answer with the answer in the label to train the task processing model. The code type data is filtered through test case verification to ensure the functional correctness. Ultimately, it can ensure the accuracy, practicality, and logical rigor of the generated long-thinking data, improve the quality of the generated data, provide high-quality learning samples for model training, and thus enhance the reasoning ability of the task processing model.
[0100] The above embodiments do not make any limitations on how to train the task processing model. Based on the above embodiments, the present invention also provides an exemplary training method for the task processing model, which may include the following content:
[0101] When minimizing the difference between the label of each inference task data sample and its corresponding predicted data and meeting the preset training effect condition, an initial task processing model is obtained; the reward score of the predicted data of each inference task data sample is determined according to the preset reward rule. Based on the label, corresponding predicted data, and reward score of each inference task data sample, the initial task processing model is continuously adjusted until the preset model training stop condition is met, and a trained task processing model is obtained.
[0102] In this embodiment, the inference task data sample is long-thinking data, that is, long-thinking data such as containing questions, problem-solving ideas, idea reflections, problem-solving steps, step reflections, and final answers. It is unified into a structured format, and it is ensured that each inference task data sample contains a complete inference chain and fine-grained supervision signals. The task processing model is built based on the first language model and the second language model, and the pre-trained weights are loaded. Based on the pre-configured training parameters, such as the learning rate (such as 2e-5), batch size (such as 32), and training cycle (such as 3 epochs), each inference task data sample is used to perform supervised fine-tuning on the task processing model. The optimization goal is to minimize the difference between the generated inference chain prediction data and the corresponding label, ensuring that the task processing model can fully learn the logical structure and details of the inference chain. During the fine-tuning process, the model performance of the task processing model is evaluated on the validation set regularly, and metrics such as accuracy and / or inference chain integrity are used to monitor the training effect. If overfitting or performance degradation occurs during the training of the task processing model, the learning rate can be adjusted or the data diversity can be increased. When the model is initially optimized using long-thinking data, on the basis of supervised fine-tuning, reinforcement learning is carried out to guide the task processing model to generate higher-quality inference chains through a dynamic reward method, further optimizing the inference ability of the task processing model.
[0103] As an exemplary reward method, the first reward score of each inference task data sample can be determined according to the difference between the task execution prediction step and the task execution result of each inference task data sample; the second reward score of each inference task data sample can be determined according to the difference between the format of the prediction data of each inference task data sample and the preset output format; the reward score of each inference task data sample can be determined according to the first reward score and the second reward score of each inference task data sample; for the first type of inference task data samples with a reward score greater than or equal to the preset reward threshold, based on the labels of each first type of inference task data sample and their corresponding prediction data, a positive gradient is used to train the task processing model; for the second type of inference task data samples with a reward score less than the preset reward threshold, based on the labels of each second type of inference task data sample and their corresponding prediction data, a negative gradient is used to train the task processing model.
[0104] In this embodiment, the reward score can include an accuracy reward score and a format reward score. The accuracy reward score is used to evaluate whether the final answer generated by the model is correct, such as whether the answer to a math problem is within the box; the format reward score is used to ensure that the reasoning process meets the structured requirements, such as clear steps and logical coherence. The preset reward threshold can be flexibly determined according to the actual situation. A positive gradient is given for a high reward score, and a negative gradient is given for a low reward score. For example, when the task to be processed is a math problem-solving task, the accuracy reward score can be calculated by comparing the generated final answer with the standard answer in the label, and the format reward score can be determined according to the integrity and logical coherence of the reasoning chain. After determining the reward score, the model parameters and experience priorities will be dynamically updated during training to maximize the reward signal. For example, PPO (Proximal Policy Optimization, a policy gradient-based reinforcement learning algorithm) combined with GAE (Generalized Advantage Estimation, a method for estimating the advantage function in reinforcement learning) can be used to calculate the advantage function to ensure the steady improvement of the performance of the task processing model in the inference task.
[0105] As can be seen from the above, in this embodiment, the task processing model is trained through two training stages of supervised fine-tuning and reinforcement learning, combining the supervision signal of high-quality data and the dynamically optimized reward mechanism, and making full use of long-thought data to significantly improve the accuracy of the task processing model in processing complex inference tasks and ensure that the performance of the task processing model in the inference task reaches the optimal.
[0106] To further improve the answer generation accuracy of the task processing model for mathematical logic tasks, based on the inference task training sample set of the above embodiment, a numerical operation data set can be further included. An exemplary generation process of the numerical operation data set:
[0107] By changing the values and operators in the target numerical operation relation expression, multiple original problem data are generated; for the mixed operations in each original problem data, calculate one operator with the highest current priority from left to right each time, and gradually obtain the calculation results of each original problem data; use the original problem data as the problem, the calculation result of the original problem data as the correct answer, and the calculation process of the original problem data as the answer analysis process to form a set of numerical operation data; based on each set of numerical operation data, construct a numerical operation data set.
[0108] In this embodiment, the target numerical operation expression is any obtained or existing numerical operation relation expression, which can be a unary operation relation expression, a binary operation relation expression or a mixed operation relation expression. The value is any number in the target numerical relation expression. The change operations of the operator and the value include modification, addition or deletion. There are many target numerical operation relation expressions, and the original problem data are the operation relation expressions obtained by changing the values and / or operators of the target data operation expression. Among them, as an efficient and simple way to generate original problem data: multiple binary operation relation expressions can be obtained, such as: A + B, A - B, A + B - C, A * B + C, A * (B + C). Replace the letters in at least one binary operation relation expression with random values. The random values can be random integers, decimals, or fractions, which do not affect the implementation of the present invention, and generate binary operation data; randomly set unary operators on at least one value of each binary operation data, that is, add unary operators in front of the integers, decimals, and fractions in the binary operation data to obtain multiple original problem data. For example, the binary operation data is A + B - C, and a unary operator sin is added before B, and the generated original problem data is: "A + sin(B) - C". The binary operation data is A * B + C, and a unary operator is added after C, and the generated original problem data is: "A * B + C!". Among them, "!" is a unary operator representing a logical NOT operation. More conveniently, a binary operation template can be pre-constructed, and binary operation relation expressions can be directly obtained from the binary operation template. To further improve the construction efficiency of the numerical operation data set, those skilled in the art can also perform amplification processing on the current open-source mathematical logic databases, such as: bella (database name), gsm8k (database name), ape (database name), blossom (database name) according to the above method.
[0109] After generating the original problem data, according to the operation laws, rules, and calculation priorities of the operators, detailed solution steps need to be listed. The generation of detailed solution steps can be achieved by combining the conversion of infix-postfix expressions. Each calculation only calculates the operation relationship with the highest priority at the current step and progresses step by step until the final result is obtained. For unary operations, the result can be directly calculated. For binary operations, the calculation result of each original problem data is also obtained step by step in the way of calculating one operator with the highest priority from left to right each time. For example, the calculation process of the calculation result: 5×7+3+9 / 3+2.5**(3-1)=35+3+9 / 3+2.5**(3-1)=38+9 / 3+2.5**(3-1)=38+3+2.5**(3-1)=41+2.5**(3-1)=41+2.5**2=41+6.25=47.25.
[0110] As can be seen from the above, in this embodiment, by constructing the original problem data in the inference task training sample set and realizing the detailed solution steps by combining the conversion of infix-postfix expressions for the original problem data, a numerical operation data set is obtained, enabling the task processing model to learn the mapping relationship from the input problem to the output answer, thus possessing the ability of numerical calculation tasks, improving the correctness of the task processing model for basic addition, subtraction, multiplication, and division operations and complex mixed operations, and being beneficial to improving the accuracy of executing more complex logical reasoning tasks.
[0111] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. The present invention also provides a corresponding device for the inference task execution method, further making the method more practical. Among them, the device can be described from the perspective of functional modules and the perspective of hardware respectively. The inference task execution device provided by the present invention is introduced below. The device is used to implement the inference task execution method provided by the present invention. In this embodiment, the inference task execution device may include or be divided into one or more program modules. The one or more program modules are stored in a storage medium and executed by one or more processors to complete the inference task execution method disclosed in Embodiment 1. The program modules referred to in this embodiment refer to a series of computer program instruction segments that can complete specific functions and are more suitable for describing the execution process of the inference task execution device in the storage medium. The following description will specifically introduce the functions of each program module in this embodiment. The inference task execution device described below can be correspondingly referred to the inference task execution method described above.
[0112] From the perspective of functional modules, please refer to Figure 7 ,Figure 7 The following is a schematic structural framework diagram of the inference task execution device provided in this embodiment. The device may include:
[0113] A task input module 701, configured to input a task to be processed and a prompt into a pre-trained task processing model; the task processing model includes a first language model and a second language model.
[0114] An idea determination module 702, configured to generate an initial task execution idea corresponding to the task to be processed by using the first language model according to the prompt, analyze the initial task execution idea by using the second language model and the task to be processed, and determine the task execution idea by using the first language model according to the initial task execution idea and the idea analysis result.
[0115] A task execution module 703, configured to analyze the single-step analysis result obtained by the second language model for a single initial task execution step output by the first language model according to the prompt, and gradually generate each task execution step by using the first language model according to the task execution idea.
[0116] Exemplarily, in some implementation manners of this embodiment, the above-mentioned idea determination module 702 may further be configured to: score the initial task execution idea by using the second language model and the idea analysis prompt, and output a task execution idea modification opinion according to the task to be processed and the initial task execution idea; wherein, the idea analysis prompt is used to prompt the second language model to score the initial task execution idea and give an idea modification opinion.
[0117] As an exemplary implementation manner of the above embodiment, the above-mentioned idea determination module 702 may further be configured to: if the score of the initial task execution idea is lower than a preset idea qualification score threshold, the first language model regenerates the task execution idea according to the task execution idea modification opinion; if the score of the initial task execution idea is not lower than the preset idea qualification score threshold, use the initial task execution idea as the task execution idea.
[0118] Exemplarily, in some other embodiments of this embodiment, the above task execution module 703 may also be used to: according to the task to be processed and the task execution idea, use the first language model to generate the current task execution steps for executing the task to be processed according to the generated content, and output them in the generated format; use the task to be processed, the task execution idea, and the current task execution steps as step reflection data, and input them into the second language model; use the second language model to score the current task execution steps according to the step reflection prompt words, and give step modification opinions; if the score of the current task execution steps is lower than the preset step passing score threshold, the first language model regenerates the current task execution steps based on the step modification opinions and the step regeneration prompt words; wherein, the step reflection prompt words are used to prompt the second language model to score the last task execution step output by the first language model and give step modification opinions.
[0119] Exemplarily, in some other embodiments of this embodiment, the above task execution module 703 may also be used to: if the number of current task execution steps is greater than the total number of steps to be generated, output all the currently generated task execution steps as the task execution result; if the number of current task execution steps is less than or equal to the total number of steps to be generated, according to the task to be processed and the task execution idea, use the first language model to generate the current task execution steps for executing the task to be processed according to the generated content in the prompt words, and output them in the generated format in the prompt words; use the task to be processed, the task execution idea, and the current task execution steps as step reflection data, and input them into the second language model, and if the score of the current task execution steps is less than the preset step passing score threshold, use the first language model to regenerate the current task execution steps; if the current task execution step is the last step, output all the currently generated task execution steps as the task execution result; if the current task execution step is not the last step, according to the task to be processed and the task execution idea, use the first language model to generate the next task execution step of the current task execution step according to the generated content.
[0120] Exemplarily, in some other embodiments of this embodiment, the above task execution module 703 can also be used to: if the number of times of the current task execution step generated by the first language model is less than the number of reflection times + 1, then use the task to be processed, the task execution idea, and the current task execution step as step reflection data and input them into the second language model; use the second language model to score the current task execution step according to the step reflection prompt words and give step modification opinions; if the score of the current task execution step is lower than the preset step passing score threshold, the first language model regenerates the current task execution step according to the step modification opinions; if the number of times of the current task execution step generated by the first language model is equal to the number of reflection times + 1 and the current task execution step is not the last step, then according to the task to be processed and the task execution idea, use the first language model to generate the next task execution step of the current task execution step according to the generated content in the prompt words.
[0121] From the perspective of functional modules, please refer to Figure 8 , Figure 8 which is a schematic structural framework diagram of the inference task execution device provided in this embodiment in another embodiment. The device may include:
[0122] A sample input module 801, configured to input an inference task training sample set and prompt words into a pre-built task processing model. The task processing model includes a first language model and a second language model that have completed pre-training. The labels of each inference task data sample in the inference task training sample set at least include questions, task execution ideas, idea adjustment processes, task execution steps, step reflection processes, and task execution results.
[0123] A model training module 802, configured to use the first language model to generate initial idea prediction data corresponding to the question sample of the current inference task data sample according to the prompt words, and use the second language model and the question sample to analyze the initial idea prediction data; and use the first language model to determine the idea prediction data according to the initial idea prediction data and the idea analysis prediction result; according to the prompt words, based on the single-step analysis prediction result obtained by the second language model analyzing a single initial task execution prediction step output by the first language model, use the first language model to gradually generate task execution prediction steps according to the idea prediction data, and use the question sample, the initial idea prediction data, the idea analysis prediction result, each initial task execution prediction step, the single-step analysis prediction result, and each task execution prediction step as prediction data; continuously train the task processing model based on the labels of each inference task data sample and their corresponding prediction data until the preset model training stop condition is met.
[0124] Exemplarily, in some embodiments of the present embodiment, the above device may further include a data loading module, which is further configured to: obtain data reading parameters; the data reading parameters include the number of rows read each time; according to the data reading parameters, read corresponding rows of data samples from the inference task training sample set and load them into the memory as data sample blocks.
[0125] Exemplarily, in some other embodiments of the present embodiment, the above device may further include a sample collection module, which is further configured to: whenever an original inference task data sample is obtained, check whether there is a sample data in the current inference task training sample set that is the same as the original inference task data sample; if there is no sample data in the current inference task training sample set that is the same as the original inference task data sample, and if the original inference task data sample has missing values, fill the original inference task data sample according to the data type of the original inference task data sample; if the original inference task data sample has outliers, adjust the original inference task data sample according to the outlier type to obtain an inference task data sample; convert the inference task data sample into a target format and store it in the current inference task training sample set.
[0126] Exemplarily, in some other embodiments of the present embodiment, the above device may further include a sample collection module, which is further configured to: if the current inference task data sample belongs to the first task type, when the task execution result of the current inference task data sample is not completely consistent with each task execution prediction step, mark the current inference task data sample as an error sample and delete the current inference task data sample from the inference task training sample set; if the current inference task data sample belongs to the second task type, when the task execution result of the current inference task data sample does not match each task execution prediction step, and / or, each task execution prediction step fails to pass the corresponding correctness verification condition, mark the current inference task data sample as an error sample and delete the current inference task data sample from the inference task training sample set.
[0127] Exemplarily, in some other embodiments of the present embodiment, the above model training module 802 may further be configured to: when minimizing the difference between the labels of each inference task data sample and their corresponding predicted data and meeting the preset training effect condition, obtain an initial task processing model; determine the reward scores of the predicted data of each inference task data sample according to the preset reward rule, and continuously adjust the initial task processing model based on the labels, corresponding predicted data and reward scores of each inference task data sample until the preset model training stop condition is met, and obtain a trained task processing model.
[0128] As an exemplary implementation of the above embodiments, the above model training module 802 can also be used to: determine the first reward score of each inference task data sample according to the difference between the task execution prediction step and the task execution result of each inference task data sample; determine the second reward score of each inference task data sample according to the difference between the format of the prediction data of each inference task data sample and the preset output format; determine the reward score of each inference task data sample according to the first reward score and the second reward score of each inference task data sample; for the first type of inference task data samples with the reward score greater than or equal to the preset reward threshold, based on the labels of each first type of inference task data samples and their corresponding prediction data, use the positive gradient to train the task processing model; for the second type of inference task data samples with the reward score less than the preset reward threshold, based on the labels of each second type of inference task data samples and their corresponding prediction data, use the negative gradient to train the task processing model.
[0129] The inference task execution device mentioned above is described from the perspective of functional modules. Further, the present invention also provides an electronic device, which is described from the perspective of hardware. Figure 9 FIG. is a schematic structural diagram of the electronic device provided by an embodiment of the present invention. The electronic device includes a memory 901 and a processor 902. A computer program is stored in the memory 901, and the processor 902 is configured to run the computer program to execute the steps in any of the above embodiments of the inference task execution method.
[0130] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any of the above embodiments of the inference task execution method when running.
[0131] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0132] An embodiment of the present application also provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the inference task execution method are implemented.
[0133] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the inference task execution method.
[0134] The above has introduced in detail an inference task execution method, an electronic device, a computer-readable storage medium, and a computer program product provided by the present invention. Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. Whether the units and algorithm steps of each example described in the disclosed embodiments are executed in the form of electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A method for executing an inference task, characterized in that: include: Input the task to be processed and the prompt word into the pre-trained task processing model; The task processing model includes a first language model and a second language model, and the parameter amount of the second language model is greater than the parameter amount of the first language model; the prompt words at least include a thought analysis prompt word; According to the prompt word, the first language model is used to generate an initial task execution idea corresponding to the task to be processed, the second language model and the idea analysis prompt word are used to score the initial task execution idea, and according to the task to be processed and the initial task execution idea, a task execution idea modification suggestion is output, and the first language model is used to determine the task execution idea according to the initial task execution idea and the idea analysis result; wherein the idea analysis prompt word is used to prompt the second language model to score the initial task execution idea and give an idea modification suggestion; According to the prompt word, based on the single-step analysis result output by the second language model, the first language model is used to gradually generate each task execution step according to the task execution idea; the single-step analysis result is obtained by analyzing the single initial task execution step output by the first language model.
2. The method for executing an inference task according to claim 1, characterized in that: Determining a task execution idea according to the initial task execution idea and the idea analysis result by using the first language model includes: If the score of the initial task execution idea is lower than the preset idea passing score threshold, the first language model regenerates the task execution idea according to the task execution idea modification suggestion; If the score of the initial task execution idea is not lower than the preset idea qualification score threshold, the initial task execution idea is used as the task execution idea.
3. The method for executing an inference task according to claim 1, characterized in that: The prompt words at least include generated content, generated format and step reflection prompt words. According to the prompt words, based on the single-step analysis result output by the second language model, the first language model is used to gradually generate each task execution step according to the task execution idea, including: According to the task to be processed and the task execution idea, using the first language model according to the generated content, generating current task execution steps for executing the task to be processed, and outputting them according to the generated format; Input the pending task, the task execution idea and the current task execution steps as step reflection data into the second language model; Using the second language model to score the current task execution step according to the step reflection prompt words, and giving step modification suggestions; if the score of the current task execution step is lower than the preset step passing score threshold, the first language model regenerates the current task execution step based on the step regeneration prompt words according to the step modification suggestions; The step reflection prompt word is used to prompt the second language model to score the last task execution step output by the first language model and provide step modification suggestions.
4. The method for executing an inference task according to any one of claims 1 to 3, characterized in that: The prompt word at least includes the total number of steps generated. According to the prompt word, based on the single-step analysis result output by the second language model, and using the first language model according to the task execution idea, each task execution step is gradually generated, including: If the number of current task execution steps is greater than the total number of steps generated, all currently generated task execution steps are output as task execution results; If the number of execution steps of the current task is less than or equal to the total number of steps generated, according to the task to be processed and the task execution idea, using the first language model and the generated content in the prompt word, generate the current task execution steps for executing the task to be processed, and output them according to the generated format in the prompt word; The task to be processed, the task execution idea and the current task execution step are input into the second language model as step reflection data, and when the score of the current task execution step is less than a preset step passing score threshold, the current task execution step is regenerated using the first language model; If the current task execution step is the last step, all currently generated task execution steps are output as the task execution result; if the current task execution step is not the last step, based on the task to be processed and the task execution idea, the first language model is used to generate the next task execution step of the current task execution step according to the generated content.
5. The method for executing an inference task according to any one of claims 1 to 3, characterized in that: The prompt word at least includes the number of reflections. According to the prompt word, based on the single-step analysis result output by the second language model, and using the first language model according to the task execution idea, each task execution step is gradually generated, including: If the number of execution steps of the current task generated by the first language model is less than the number of reflections+1, the task to be processed, the task execution idea and the current task execution steps are input into the second language model as step reflection data; Using the second language model to score the current task execution step according to the step reflection prompt words, and giving step modification suggestions; if the score of the current task execution step is lower than the preset step passing score threshold, the first language model regenerates the current task execution step according to the step modification suggestions; If the number of current task execution steps generated by the first language model is equal to the number of reflections + 1, and the current task execution step is not the last step, then according to the task to be processed and the task execution idea, the first language model is used to generate the next task execution step of the current task execution step according to the generated content in the prompt word.
6. A method for executing an inference task, characterized in that: include: Inputting the reasoning task training sample set and the prompt words into a pre-built task processing model, wherein the task processing model includes a first language model and a second language model that have been pre-trained, and the labels of each reasoning task data sample in the reasoning task training sample set at least include questions, task execution ideas, idea adjustment process, task execution steps, step reflection process and task execution results; the parameter amount of the second language model is greater than the parameter amount of the first language model; and the prompt words at least include idea analysis prompt words; According to the prompt words, the first language model is used to generate initial idea prediction data corresponding to the problem sample of the current reasoning task data sample, the initial idea prediction data is scored using the second language model, the idea analysis prompt words and the problem sample, and the corresponding idea modification opinions are output according to the task to be processed and the initial idea prediction data; and the first language model is used to determine the idea prediction data according to the initial idea prediction data and the idea analysis prediction result; wherein the idea analysis prompt words are used to prompt the second language model to score the initial idea prediction data and give idea modification opinions; According to the prompt word, a single initial task execution prediction step output by the first language model is analyzed based on the second language model, and the single-step analysis prediction result is obtained, and the first language model is used to predict the data according to the idea, and the task execution prediction steps are gradually generated, and the problem sample, the initial idea prediction data, the idea analysis prediction result, each initial task execution prediction step, the single-step analysis prediction result and each task execution prediction step are used as prediction data; Based on the labels of each inference task data sample and its corresponding prediction data, the task processing model is continuously trained until a preset model training stop condition is met.
7. The method for executing an inference task according to claim 6, characterized in that: Before inputting the inference task training sample set and prompt words into the pre-built task processing model, it also includes: Obtaining data reading parameters; the data reading parameters include the number of rows read each time; According to the data reading parameters, data samples of a corresponding number of rows are read from the inference task training sample set, and loaded into the memory as data sample blocks.
8. The method for executing an inference task according to claim 6, characterized in that: Before inputting the inference task training sample set and prompt words into the pre-built task processing model, it also includes: Whenever an original reasoning task data sample is obtained, a search is performed in the current reasoning task training sample set to determine whether there is sample data identical to the original reasoning task data sample; When there is no sample data identical to the original reasoning task data sample in the current reasoning task training sample set, if there is a missing value in the original reasoning task data sample, then the original reasoning task data sample is filled with data according to the data type of the original reasoning task data sample; if there is an abnormal value in the original reasoning task data sample, then the original reasoning task data sample is adjusted according to the abnormal value type to obtain the reasoning task data sample; The reasoning task data sample is converted into a target format and stored in the current reasoning task training sample set.
9. The method for executing an inference task according to claim 6, characterized in that: After the step-by-step generation of task execution prediction steps, it also includes: If the current reasoning task data sample belongs to the first task type, when the task execution result of the current reasoning task data sample is not completely consistent with each task execution prediction step, mark the current reasoning task data sample as an error sample, and delete the current reasoning task data sample from the reasoning task training sample set; If the current reasoning task data sample belongs to the second task type, when the task execution result of the current reasoning task data sample does not match the task execution prediction steps, and / or the task execution prediction steps cannot pass the corresponding correctness verification conditions, the current reasoning task data sample is marked as an error sample and the current reasoning task data sample is deleted from the reasoning task training sample set.
10. The method for executing an inference task according to any one of claims 6 to 9, characterized in that: Based on the labels of each inference task data sample and its corresponding prediction data, the task processing model is continuously trained until a preset model training stop condition is met, including: When the difference between the label of each inference task data sample and its corresponding predicted data is minimized and the preset training effect conditions are met, the initial task processing model is obtained; The reward score of the predicted data of each reasoning task data sample is determined according to the preset reward rules. Based on the label of each reasoning task data sample, the corresponding predicted data and its reward score, the initial task processing model is continuously adjusted until the preset model training stop condition is met to obtain a trained task processing model.
11. The method for executing an inference task according to claim 10, characterized in that: Based on the labels of each reasoning task data sample, the corresponding prediction data and its reward score, the initial task processing model is continuously adjusted, including: Determine a first reward score for each reasoning task data sample according to the difference between the task execution prediction step and the task execution result of each reasoning task data sample; Determining a second reward score for each reasoning task data sample according to a difference between a format of the predicted data of each reasoning task data sample and a preset output format; Determine a reward score for each reasoning task data sample according to the first reward score and the second reward score of each reasoning task data sample; For the first-category reasoning task data samples whose reward scores are greater than or equal to the preset reward threshold, the task processing model is trained using forward gradient based on the labels of each first-category reasoning task data sample and its corresponding prediction data; For the second type of reasoning task data samples whose reward scores are less than the preset reward threshold, the task processing model is trained using negative gradient based on the labels of each second type of reasoning task data sample and its corresponding prediction data.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for executing an inference task as described in any one of claims 1 to 11 when executing the computer program.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the reasoning task execution method according to any one of claims 1 to 11 are implemented.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the reasoning task execution method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Multi-agent thinking chain negotiation enhancement generation method
CN118821834A