Method and system for enhancing coprocessing capability of large model and Agent agent based on thinking chain

By generating logically coherent thought chains and accurate prompts, and coordinating large language models with agent intelligence, the problems of inaccurate generation and high annotation costs in complex reasoning tasks are solved, thereby improving the reasoning ability and efficiency of agent intelligence.

CN121052279APending Publication Date: 2025-12-02JIANGSU HAIRUO INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511159930.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing thought chain hinting methods generate inaccurate thought chains in complex reasoning tasks, have high annotation costs, and rely heavily on large-scale models, which limits the performance and application of agent intelligence in complex reasoning tasks.

Method used

By collecting and cleaning complex reasoning task datasets, recording each step of the reasoning process in detail, generating logically coherent thought chains using pre-trained large language models, optimizing thought chains by combining reinforcement learning and few-shot methods, generating accurate prompts, driving the agent to reason, and optimizing thought chains and prompts through evaluation and feedback.

Benefits of technology

It improves the performance of agent intelligence in complex reasoning tasks, reduces the annotation cost of thought chains, enhances the accuracy of thought chains, reduces the dependence on large language model scale, and realizes efficient and accurate reasoning of agent intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052279A_ABST
    Figure CN121052279A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for enhancing coprocessing capability of a large model and an Agent agent based on a thinking chain, and relates to the technical field of artificial intelligence, and the method comprises the steps: collecting a data set containing different types of complex reasoning tasks, carrying out data cleaning and labeling, labeling questions and answers, and carrying out data cleaning and labeling; the reasoning process of each step is recorded in detail to form a thinking chain; learning the marked data containing the thinking chain by using a pre-trained large language model to enable the marked data containing the thinking chain to master inference step disassembly and logic generation capabilities so as to generate the thinking chain; the large language model generates prompt information including problems, reasoning logic and expected steps based on the generated thinking chain in combination with specific problems, and provides an accurate reasoning guide framework for the Agent agent; the prompt information is input into the Agent agent to execute the reasoning task; the reasoning result of the Agent agent is evaluated, and the thinking chain and prompt information are optimized according to the evaluation result. According to the method, the performance of the Agent on a complex reasoning task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically a method and system for enhancing the collaborative processing capabilities of a large model and an agent based on thought chain enhancement. Background Technology

[0002] With the continuous development of artificial intelligence technology, agent-based systems using large language models have demonstrated powerful capabilities in natural language processing tasks. However, when dealing with tasks requiring complex reasoning, such as arithmetic reasoning, common sense reasoning, and symbolic reasoning, the performance of agent-based systems still has certain limitations. Traditional standard cue learning methods are ineffective on these complex tasks and cannot fully tap the reasoning potential of agent-based systems.

[0003] In recent years, Chain of Thought (CoT) hints have emerged as a novel approach. They encourage agents to explain their reasoning processes by providing a series of intermediate reasoning steps as "thought chains" during few-shot learning, significantly improving agent performance on arithmetic, common sense, and symbolic reasoning tasks. However, current CoT hints still have some limitations, such as the generated thought chains not necessarily being factually accurate, high annotation costs, and large model requirements. For example, in some complex arithmetic reasoning tasks, existing CoT hints may generate inaccurate thought chains, leading to incorrect results for the agent. Furthermore, the substantial manual annotation work keeps the generation cost of thought chains high, limiting the widespread application of this method. In addition, the dependence on large-scale models makes this method difficult to implement in resource-constrained scenarios. Summary of the Invention

[0004] This invention addresses the needs and shortcomings of current technological development by providing a method and system for enhancing the collaborative processing capabilities of large models and agents based on thought chains. This improves the performance of agents on complex reasoning tasks, while reducing the annotation cost of thought chains, increasing the factual accuracy of thought chains, and reducing dependence on model size.

[0005] Firstly, the present invention provides a method for enhancing the collaborative processing capabilities of a large-scale model based on a thought chain and an agent-based intelligent agent. The technical solution adopted to solve the above-mentioned technical problems is as follows:

[0006] A method for enhancing the collaborative processing capabilities of a large-scale model and an agent based on thought chains includes the following steps:

[0007] S1. Collect datasets covering different types of complex reasoning tasks. After cleaning and removing noise, duplicates and invalid data, label the data. Not only label the questions and answers, but also record each step of the reasoning process in detail to form a thought chain, providing an accurate basis for the subsequent large language model to learn reasoning logic and assist the agent in reasoning.

[0008] S2. Utilize a pre-trained large language model to learn from labeled data containing thought chains, enabling it to master the ability to decompose reasoning steps and generate logic, thereby generating task-adaptive thought chains that can be directly used to guide the Agent's reasoning.

[0009] S3, the large language model, is based on the generated thought chain and combines specific questions to generate prompt information containing questions, reasoning logic and expected steps, providing the agent with a precise reasoning guidance framework;

[0010] S4. Input the prompts generated by the large language model into the agent, and drive the agent to carry out reasoning tasks according to the thought chain and step guidance in the prompts, so as to realize the transformation of the reasoning logic of the large language model into the reasoning behavior of the agent.

[0011] S5. Evaluate the reasoning results of the Agent and feed the evaluation results back to the large language model. The large language model optimizes the thought process and prompts based on the evaluation results.

[0012] Optionally, in step S2, during the generation of the thought chain, reinforcement learning, zero-shot thought chain prompting, and few-shot thought chain prompting are used to collaboratively optimize the large language model, enabling it to generate logically coherent and clearly defined thought chains, providing interpretable path support for the reasoning process of complex tasks.

[0013] Optionally, step S4 is performed to monitor and record the entire reasoning process of the Agent in detail, including the reasoning path, key decision nodes, and result output, so as to provide data support for subsequent analysis of reasoning logic vulnerabilities and optimization of Agent performance.

[0014] Optionally, step S5 is performed to evaluate the reasoning results of the Agent using three evaluation metrics: accuracy, recall, and F1 score. Subsequently, the thought process and prompts are optimized based on the evaluation results.

[0015] Secondly, the present invention provides a system based on the collaborative processing capabilities of a large-scale thinking chain-enhanced model and an agent-based intelligent agent. The technical solution adopted to solve the above-mentioned technical problems is as follows:

[0016] A system based on the collaborative processing capabilities of a large-scale model enhanced by thought chain and an agent intelligent agent, comprising:

[0017] The collection and preprocessing module is used to collect datasets covering different types of complex reasoning tasks. After cleaning and removing noise, duplicates and invalid data, the data is labeled. Not only are the questions and answers labeled, but each step of the reasoning process is recorded in detail to form a thought chain, providing an accurate basis for the subsequent large language model to learn reasoning logic and assist the agent in reasoning.

[0018] The thought chain generation module is used to learn from labeled data containing thought chains by using a pre-trained large language model, enabling it to master the ability to decompose reasoning steps and generate logic, and then generate task-adaptive thought chains that can be directly used to guide the reasoning of the agent.

[0019] The prompt information generation module is used to assist the large language model in generating prompt information containing the question, reasoning logic and expected steps based on the generated thought chain and specific questions, providing the agent intelligent agent with an accurate reasoning guidance framework.

[0020] The reasoning task execution module is used to input the prompt information generated by the large language model into the agent, and drive the agent to carry out reasoning tasks according to the thought chain and step guidance in the prompt, so as to realize the transformation of the reasoning logic of the large language model into the reasoning behavior of the agent.

[0021] The results evaluation and optimization module is used to evaluate the reasoning results of the agent and feed the evaluation results back to the large language model. The large language model optimizes the thought process and prompts based on the evaluation results.

[0022] Optionally, in the process of generating thought chains using a pre-trained large language model, the thought chain generation module employs reinforcement learning, zero-shot thought chain prompting, and few-shot thought chain prompting methods to collaboratively optimize the large language model, enabling it to generate logically coherent and clearly defined thought chains, thus providing interpretable path support for the reasoning process of complex tasks.

[0023] Optionally, the inference task execution module supports real-time monitoring and recording functions to monitor and record the inference process of the Agent in detail, including inference path, key decision nodes and result output, providing data support for subsequent analysis of inference logic vulnerabilities and optimization of Agent performance.

[0024] Optionally, the result evaluation and optimization module uses three evaluation metrics—accuracy, recall, and F1 score—to evaluate the reasoning results of the agent. Subsequently, based on the evaluation results, the thought process and prompts are optimized.

[0025] Thirdly, the present invention also provides an electronic device comprising: a memory and at least one processor;

[0026] The memory contains computer programs;

[0027] The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the method for collaborative processing capabilities between the large model and the agent intelligence based on the thought chain as described in the first aspect.

[0028] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program that can be executed by a processor to implement the method for collaborative processing capabilities between a large model based on thought chain enhancement and an agent as described in the first aspect.

[0029] The method and system of the present invention, which are based on the collaborative processing capabilities of a large-scale model enhanced by thought chain and an agent, have the following advantages compared with the prior art:

[0030] This invention provides detailed thought chains and prompts through a large language model, guiding the agent to make accurate inferences and thus improving the agent's performance on complex reasoning tasks. Simultaneously, it reduces the annotation cost of thought chains, improves the factual accuracy of thought chains, and reduces dependence on the scale of the large language model. The evaluation and optimization mechanism for the agent's reasoning results can promptly identify errors and deficiencies in the thought chains and prompts, allowing for corresponding adjustments and improvements. This enables the agent to operate more accurately and efficiently in complex reasoning tasks, providing stronger support for the application of artificial intelligence in various fields. Attached Figure Description

[0031] Appendix Figure 1 This is a flowchart of the method according to Embodiment 1 of the present invention;

[0032] Appendix Figure 2 This is a module connection block diagram of Embodiment 2 of the present invention. Detailed Implementation

[0033] To make the technical solution, the technical problem solved, and the technical effect of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments.

[0034] Example 1:

[0035] Combined with appendix Figure 1 This embodiment proposes a method based on the collaborative processing capability of a large-scale model enhanced by thought chain and an agent, which includes the following steps:

[0036] S1. Collect datasets covering different types of complex reasoning tasks. After cleaning and removing noise, duplicates, and invalid data, label the data. Not only are the questions and answers labeled, but each step of the reasoning process is also recorded in detail to form a thought chain, providing an accurate basis for the subsequent large language model to learn reasoning logic and assist the agent in reasoning.

[0037] For example, for arithmetic reasoning tasks, the dataset should include mathematical problems of different difficulty levels; for common sense reasoning tasks, the dataset should contain common sense knowledge from various fields.

[0038] When cleaning data, noisy, duplicate, and invalid data should be removed to improve data quality.

[0039] When annotating the reasoning process, it should be as detailed as possible, clearly recording each reasoning step to provide accurate evidence for the subsequent generation of thought chains.

[0040] S2. Using a pre-trained large language model, learn from labeled data containing thought chains, enabling it to master the ability to decompose reasoning steps and generate logic, thereby generating task-adaptive thought chains that can be directly used to guide the agent's reasoning.

[0041] In step S2, during the generation of the thought chain, reinforcement learning, zero-shot thought chain prompting, and few-shot thought chain prompting are used to collaboratively optimize the large language model, enabling it to generate logically coherent and clearly defined thought chains, providing interpretable path support for the reasoning process of complex tasks.

[0042] Reinforcement learning methods can encourage large language models to generate thought processes that more closely resemble correct reasoning by setting reward mechanisms. For example, a higher reward is given when the thought process generated by the large language model accurately guides the agent to the correct answer; conversely, a lower reward is given if it fails to do so. Through continuous learning and adjustment, large language models can gradually generate more accurate and reasonable thought processes.

[0043] Meanwhile, we can learn from the Zero-shot-CoT prompting method by adding "Let's think step by step" to the end of the prompt words to stimulate the thinking chain of the large language model; or we can use the Few-shot-CoT prompting method to write thinking chain samples as prompt words so that the large language model can learn how to deduce the thinking chain.

[0044] S3, the large language model, is based on the generated thought chain and combines specific questions to generate prompts containing questions, reasoning logic, and expected steps, providing the agent with a precise reasoning guidance framework.

[0045] When generating prompts, ensure that the problem is clearly stated, the thought process is logically coherent, and the expected reasoning steps are reasonable and feasible.

[0046] S4. Input the prompts generated by the large language model into the agent, and drive the agent to carry out reasoning tasks according to the thought chain and step guidance in the prompts, so as to realize the transformation of the reasoning logic of the large language model into the reasoning behavior of the agent.

[0047] Step S4 involves monitoring and recording the entire reasoning process of the Agent, including the reasoning path, key decision nodes, and output results, providing data support for subsequent analysis of reasoning logic vulnerabilities and optimization of Agent performance.

[0048] S5. Evaluate the reasoning results of the Agent and feed the evaluation results back to the large language model. The large language model optimizes the thought process and prompts based on the evaluation results.

[0049] Step S5 is executed, specifically using three evaluation metrics—accuracy, recall, and F1 score—to evaluate the reasoning results of the agent. Subsequently, based on the evaluation results, the thought process and prompts are optimized.

[0050] For example, if the agent's reasoning on a certain problem is inaccurate, the reasons can be analyzed, the thought process or prompts adjusted, and then retraining and reasoning can be performed until the agent's reasoning performance reaches a satisfactory level. During the evaluation process, different types of reasoning tasks should be evaluated separately to more accurately understand the agent's performance in various aspects. Simultaneously, evaluation results and the optimization process should be recorded promptly to provide a reference for subsequent research and improvement.

[0051] Example 2:

[0052] Combined with appendix Figure 2 This embodiment proposes a system based on the collaborative processing capabilities of a thought chain-enhanced large model and an agent, which includes:

[0053] The collection and preprocessing module is used to collect datasets covering different types of complex reasoning tasks. After cleaning and removing noise, duplicates and invalid data, the data is labeled. Not only are the questions and answers labeled, but each step of the reasoning process is recorded in detail to form a thought chain, providing an accurate basis for the subsequent large language model to learn reasoning logic and assist the agent in reasoning.

[0054] The thought chain generation module is used to learn from labeled data containing thought chains by using a pre-trained large language model, enabling it to master the ability to decompose reasoning steps and generate logic, and then generate task-adaptive thought chains that can be directly used to guide the reasoning of the agent.

[0055] The prompt information generation module is used to assist the large language model in generating prompt information containing the question, reasoning logic and expected steps based on the generated thought chain and specific questions, providing the agent intelligent agent with an accurate reasoning guidance framework.

[0056] The reasoning task execution module is used to input the prompt information generated by the large language model into the agent, and drive the agent to carry out reasoning tasks according to the thought chain and step guidance in the prompt, so as to realize the transformation of the reasoning logic of the large language model into the reasoning behavior of the agent.

[0057] The results evaluation and optimization module is used to evaluate the reasoning results of the agent and feed the evaluation results back to the large language model. The large language model optimizes the thought process and prompts based on the evaluation results.

[0058] In this embodiment, the thought chain generation module uses a pre-trained large language model to generate thought chains. During this process, reinforcement learning, zero-shot thought chain hints, and few-shot thought chain hints are used to collaboratively optimize the large language model, enabling it to generate logically coherent and clearly defined thought chains, thus providing interpretable path support for the reasoning process of complex tasks.

[0059] In this embodiment, the inference task execution module supports real-time monitoring and recording functions, which are used to monitor and record the inference process of the Agent in detail, including the inference path, key decision nodes and result output, so as to provide data support for subsequent analysis of inference logic vulnerabilities and optimization of Agent performance.

[0060] In this embodiment, the result evaluation and optimization module uses three evaluation metrics—accuracy, recall, and F1 score—to evaluate the reasoning results of the agent. Subsequently, based on the evaluation results, the thought chain and prompt information are optimized.

[0061] Example 3:

[0062] This embodiment also provides an electronic device, including: a memory and a processor;

[0063] The memory stores the instructions executed by the computer.

[0064] The processor executes computer execution instructions stored in the memory, causing the processor to execute the method of the first embodiment based on the enhanced large model of thought chain and the collaborative processing capability of agent intelligent agent.

[0065] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.

[0066] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory devices, or other volatile solid-state storage devices.

[0067] Example 4:

[0068] This embodiment also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the method of the enhanced large model based on thought chain and the collaborative processing capability of the agent intelligent agent in Embodiment 1. Specifically, a system or device equipped with a storage medium can be provided, on which software program code implementing the method described in Embodiment 1 is stored, and the computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.

[0069] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.

[0070] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0071] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0072] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for enhancing the collaborative processing capabilities of a large-scale model based on thought chain enhancement and an agent intelligent agent, characterized in that, Includes the following steps: S1. Collect datasets covering different types of complex reasoning tasks. After cleaning and removing noise, duplicates and invalid data, label the data. Not only label the questions and answers, but also record each step of the reasoning process in detail to form a thought chain, providing an accurate basis for the subsequent large language model to learn reasoning logic and assist the agent in reasoning. S2. Utilize a pre-trained large language model to learn from labeled data containing thought chains, enabling it to master the ability to decompose reasoning steps and generate logic, thereby generating task-adaptive thought chains that can be directly used to guide the Agent's reasoning. S3, the large language model, is based on the generated thought chain and combines specific questions to generate prompt information containing questions, reasoning logic and expected steps, providing the agent with a precise reasoning guidance framework; S4. Input the prompts generated by the large language model into the agent, and drive the agent to carry out reasoning tasks according to the thought chain and step guidance in the prompts, so as to realize the transformation of the reasoning logic of the large language model into the reasoning behavior of the agent. S5. Evaluate the reasoning results of the Agent and feed the evaluation results back to the large language model. The large language model optimizes the thought process and prompts based on the evaluation results.

2. The method for collaborative processing capabilities between a large-scale model based on thought chain enhancement and an agent intelligent agent as described in claim 1, characterized in that, In step S2, during the generation of the thought chain, reinforcement learning, zero-shot thought chain prompting, and few-shot thought chain prompting are used to collaboratively optimize the large language model, enabling it to generate logically coherent and clearly defined thought chains, providing interpretable path support for the reasoning process of complex tasks.

3. The method for enhancing the collaborative processing capabilities of a large-scale model and an agent based on a thought chain, as described in claim 2, is characterized in that... Step S4 involves monitoring and recording the entire reasoning process of the Agent, including the reasoning path, key decision nodes, and output results, providing data support for subsequent analysis of reasoning logic vulnerabilities and optimization of Agent performance.

4. The method for enhancing the collaborative processing capabilities of a large-scale model and an agent based on a thought chain, as described in claim 3, is characterized in that... Step S5 is executed, using three evaluation metrics—accuracy, recall, and F1 score—to evaluate the reasoning results of the agent. Subsequently, based on the evaluation results, the thought process and prompts are optimized.

5. A system based on the collaborative processing capabilities of a large-scale model enhanced by thought chain and an agent, characterized in that, It includes: The collection and preprocessing module is used to collect datasets covering different types of complex reasoning tasks. After cleaning and removing noise, duplicates and invalid data, the data is labeled. Not only are the questions and answers labeled, but each step of the reasoning process is recorded in detail to form a thought chain, providing an accurate basis for the subsequent large language model to learn reasoning logic and assist the agent in reasoning. The thought chain generation module is used to learn from labeled data containing thought chains by using a pre-trained large language model, enabling it to master the ability to decompose reasoning steps and generate logic, and then generate task-adaptive thought chains that can be directly used to guide the reasoning of the agent. The prompt information generation module is used to assist the large language model in generating prompt information containing the question, reasoning logic and expected steps based on the generated thought chain and specific questions, providing the agent intelligent agent with an accurate reasoning guidance framework. The reasoning task execution module is used to input the prompt information generated by the large language model into the agent, and drive the agent to carry out reasoning tasks according to the thought chain and step guidance in the prompt, so as to realize the transformation of the reasoning logic of the large language model into the reasoning behavior of the agent. The results evaluation and optimization module is used to evaluate the reasoning results of the agent and feed the evaluation results back to the large language model. The large language model optimizes the thought process and prompts based on the evaluation results.

6. The system based on the collaborative processing capability of the enhanced thinking chain model and the agent intelligent agent as described in claim 5, characterized in that, The thought chain generation module utilizes a pre-trained large language model to generate thought chains. During this process, reinforcement learning, zero-shot thought chain hints, and few-shot thought chain hints are used to collaboratively optimize the large language model, enabling it to generate logically coherent and clearly structured thought chains. This provides interpretable path support for the reasoning process of complex tasks.

7. The system based on the collaborative processing capability of the enhanced thinking chain model and the agent intelligent agent as described in claim 6, characterized in that, The inference task execution module supports real-time monitoring and recording functions, which are used to monitor and record the inference process of the Agent in detail, including the inference path, key decision nodes and result output, so as to provide data support for subsequent analysis of inference logic vulnerabilities and optimization of Agent performance.

8. The system based on the collaborative processing capability of the enhanced large model of thought chain and the agent intelligent agent according to claim 7, characterized in that, The result evaluation and optimization module uses three evaluation metrics—accuracy, recall, and F1 score—to evaluate the reasoning results of the agent. Subsequently, based on the evaluation results, the thought process and prompts are optimized.

9. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the method as described in any one of claims 1 to 4, which is based on the collaborative processing capability of a large model and an agent intelligence based on the thought chain enhancement model.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method for collaborative processing capabilities between a large model based on thought chain enhancement and an agent as described in any one of claims 1 to 4.