A Method for Constructing a Multi-Agent Reasoning Framework Based on the Idea of Metacognition

By constructing metacognitive agents and central agents in large language models and simulating human thinking patterns, the problem of large language models lacking serial computing mode and "thinking" capabilities in complex problems is solved, and its reasoning ability is significantly improved.

CN119476492BActive Publication Date: 2025-06-10EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411613108.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-06-10
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

When solving complex logic problems such as mathematics, existing large language models lack serial computing mode and "thinking" capabilities, which limits their reasoning capabilities.

Method used

The multi-agent reasoning framework is constructed using metacognitive ideas, and by constructing metacognitive agents and central agents, it simulates the thinking mode of humans when solving complex problems, and introduces serial computing mode and "thinking" capabilities.

Benefits of technology

It effectively alleviates the problem of insufficient serial computing mode and "thinking" ability in the reasoning ability of large language models, and comprehensively improves its reasoning ability, especially when dealing with complex problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476492B_ABST
    Figure CN119476492B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a multi-agent reasoning framework based on metacognitive thinking. The feature is a method of using a multi-agent reasoning architecture to simulate the human brain's thinking mode. This method first constructs a metacognitive agent training data set and trains an open-source large language model based on the thinking injection training paradigm to make it a metacognitive agent with thinking ability. Then, it constructs a central agent training data set and trains the open-source large language model based on the supervised training paradigm to enable it to accurately evaluate the difficulty of tasks. Finally, the open-source large language model without additional training is used as a cognitive agent, and combined with the metacognitive agent and the central agent, a multi-agent reasoning framework that can effectively simulate the human metacognitive thinking mode is constructed. Compared with the prior art, the present invention has the advantages of improving the model's ability to solve complex problems and effectively alleviating the two major limitations that hinder the lack of serial computing mode and the lack of "thinking" ability in the reasoning ability of large language models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of metacognition, and specifically to a method for constructing a multi-agent reasoning framework based on metacognitive thinking. Background Art

[0002] Large Language Models (LLMs) are natural language processing (NLP) models based on deep learning technology. They usually use billions to hundreds of billions of parameters and are trained with large-scale text data. They can generate, understand, and process human language, and have extremely strong language generation and semantic understanding capabilities.

[0003] Currently, the reasoning capabilities of existing large language models (such as ChatGPT and Llama, etc.) are insufficient, specifically manifested as difficulty in solving problems with relatively complex logic such as mathematics. There are mainly two reasons for this: 1) Lack of serial computing mode. Large language models are all based on the Transformer architecture, and they always follow the parallel computing mode during the reasoning process. On the contrary, the human brain usually activates the parallel conduction mode of neurons, that is, fast thinking, when solving simple problems, and activates the serial conduction mode of neurons, that is, slow thinking, when solving complex problems. Therefore, the lack of serial computing mode in large language models limits their ability to handle complex problems; 2) Lack of "thinking" ability. Large language models are trained with large-scale human corpora. Although these corpora contain rich knowledge, humans usually do not write down their thinking processes. Therefore, large language models trained with these corpora usually lack the ability to deeply "think".

[0004] Existing solutions for enhancing the reasoning capabilities of large language models are mainly divided into the following two types: 1) Prompt-based methods, such as Chain of Thought, Tree of Thought, etc. These methods guide large language models to output intermediate reasoning steps through prompt engineering. Although these methods simulate the thinking process of the human brain when solving complex problems, they do not change the parallel computing mode of large language models, limiting the upper limit of their reasoning capabilities; 2) Multi-step reasoning frameworks based on multi-agents, such as the CAMEL framework. These methods improve the reasoning capabilities through the collaborative calculation of multiple models. This type of method introduces a serial computing mode through the collaborative manner of multiple models, but a single model lacks the "thinking" ability, resulting in difficulty in fully leveraging the advantages of the serial computing mode. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for constructing a multi-agent reasoning framework based on the metacognitive thought in view of the insufficient reasoning ability of existing large language models. The multi-agent reasoning framework constructed by adopting the metacognitive thought realizes the serial computing mode of the large language model, effectively simulates the thinking mode used by humans when solving complex problems. In this thinking mode, humans will introspect, evaluate and self-regulate, so as to improve their performance. Combining the metacognitive thought to construct a multi-agent reasoning framework can effectively alleviate two major limitations that hinder the reasoning ability of large language models: 1) lack of serial computing mode; 2) lack of "thinking" ability, thus comprehensively improving the reasoning ability of large language models. This method collects the training data for constructing metacognitive agents, trains the large language model through the thinking injection training paradigm to obtain metacognitive agents, then collects the training data for constructing central agents, and trains the large language model through the supervised fine-tuning training paradigm to obtain central agents. This agent is used to control when to activate metacognitive agents. Finally, a multi-agent reasoning framework is constructed in combination with the metacognitive thought. The method is simple and has good use effects, greatly improving the ability of the model to solve complex problems, and has a wide range of application scenarios and application values.

[0006] The specific technical solution for realizing the purpose of the present invention is: a method for constructing a multi-agent reasoning framework combined with the metacognitive thought, which is characterized by adopting a multi-agent reasoning architecture to enhance the reasoning ability of the large language model to simulate the thinking mode of the human brain. The method includes the following specific steps:

[0007] Step 1: Construct a training data set for metacognitive agents

[0008] First, design a thinking injection prompt (Prompt) P t , which is used to stimulate the thinking process of the large language model. Then, splice P t with the user instruction prompt to obtain an input prompt and input it into the large language model to obtain an output result, which is specifically expressed by the following formula (a):

[0009] LLM 0 (x i ) = {z i , y i} (a).

[0010] Among them, LLM 0 represents a naive large language model without additional training. z i and y i in the output result respectively represent the thinking process and the answer. Subsequently, a judge model is used to judge the quality of y i in the output, which is specifically expressed by the following formula (b):

[0011] Judge(y i ) = label, label ∈ {pos, neg} (b).

[0012] Among them, Judge represents the judge model, which is used to judge the quality of y i and generate a label label for it. label = pos / neg represents that the current y i is a positive sample / negative sample.

[0013] Finally, the training dataset D meta of the metacognitive agent is represented by the following formula (c):

[0014]

[0015] Among them, and represent the positive and negative samples in the answer respectively, and D meta represents the training dataset of the metacognitive agent.

[0016] Step 2. Obtain the metacognitive agent through the training paradigm of thought injection

[0017] Use the Direct Preference Optimization algorithm to train an LLM meta on D obtained in Step 1 0 to obtain the metacognitive agent Agent meta . The loss function for training is represented by the following formula (d):

[0018]

[0019] Among them, π = LLM 0 represents the large language model to be trained.

[0020] By minimizing the above loss function, the "slow thinking" process Agent shown in the following formula (e) can be obtained meta :

[0021]

[0022] Among them, Agent meta simulates the "slow thinking" process of the human brain and is used to process complex problems. It implicitly generates a thinking process while generating an answer, which enables Agent meta to allocate more computational resources to complex problems, effectively alleviating the limitation of the lack of "thinking" ability mentioned in the background technology. Specifically, the "complex thinking process" Agent meta (x) shown in the following formula (f):

[0023] Agent meta (x) = {z latent , y} (f).

[0024] Among them, x, z latent and y represent the input, implicit thinking process, and answer output respectively.

[0025] Step 3: Construct the central agent training dataset

[0026] To ensure the generalization of the central agent, first collect data samples from multiple fields (mathematics, logical reasoning, general knowledge, medicine, law) and process them into the form of question-and-answer pairs as shown in the following formula (g):

[0027] Multi_data = {Q i , A i} (g).

[0028] Then, divide the data sample pairs into two categories, simple and difficult, according to the difficulty level of the questions and label them. Specifically, it is expressed as the following formula (h):

[0029] D hub = {Q i , A i , R i} (h).

[0030] Among them, D hub represents the central agent training dataset, and R i ∈{easy, hard} represents the difficulty level label of the question.

[0031] Step 4: Obtain the central agent based on the supervised fine-tuning training paradigm

[0032] Use the supervised fine-tuning training paradigm for classification tasks to train an LLM hub on D obtained in Step 3 0 , and obtain the central agent Agent hub . The loss function for training is expressed as the following formula (i):

[0033]

[0034] Among them, π = LLM 0 represents the large language model to be trained, and R and represent the true difficulty label and the predicted difficulty label respectively.

[0035] By minimizing the above loss function, the central agent Agent hub can be obtained. Specifically, it is expressed as the following formula (j):

[0036]

[0037] Among them, represents finding π that minimizes , and the Agent hub is used to judge the difficulty level of the input question to determine whether to enable the Agent obtained in the second step meta , specifically expressed by the following formula (k):

[0038]

[0039] Step 5: Construct a multi-agent reasoning framework based on the metacognitive thought

[0040] Inspired by the human metacognitive thought, this multi-agent reasoning framework consists of a cognitive agent Agent cog , the metacognitive agent Agent obtained in the second step meta and the central agent Agent obtained in the fourth step hub , and is specifically expressed by the following formula (l):

[0041] MetalInf = (Agent cog , Agent meta , Agent hub ) (l).

[0042] Among them, MetaInf represents the multi-agent reasoning framework constructed based on the metacognitive thought, and Agent cog is served by an ordinary large language model LLM without additional training 0 .

[0043] In MetaInf, Agent cog is used to handle simple tasks, and Agent meta is used to handle complex tasks. In addition, in order to balance performance and resource consumption, this multi-agent reasoning framework uses Agent hub to analyze the difficulty level of the task, so as to activate Agent cog when receiving a simple task, and activate Agent meta with stronger reasoning ability but greater resource consumption when receiving a difficult task. Therefore, MetaInf has two working modes: 1) simple task processing; 2) complex task processing.

[0044] When Agent hub judges that the current task is simple, MetaInf will activate the simple task processing mode, that is, only use Agent cog to solve the task, specifically expressed by the following formulas (m) - (n):

[0045] Agent hub (Q) = easy (m);

[0046] O = Agent cog (Q) (n).

[0047] Among them, Q represents the query of the current task, and O represents the output result of MetaInf.

[0048] When Agent hub judges that the current task is difficult, MetaInf will activate the difficult task processing mode, that is, first use Agent cog to generate an intermediate result, and then use Agent meta , combined with the optimization instruction to optimize the intermediate result to obtain the final output, specifically expressed as the following formulas (o) to (q):

[0049] Agent hub (Q) = hard (o);

[0050] S = Agent cog (Q) (p);

[0051] O = Agent meta (P refine , S) (q).

[0052] Among them, S represents the intermediate result, and P refine represents the optimization instruction. This working mode introduces a serial computing mode, effectively alleviating the limitation of the lack of serial computing mode mentioned in the background technology.

[0053] The present invention has the following beneficial technical effects and remarkable technical progress compared with the prior art:

[0054] 1) The present invention can effectively alleviate the two major limitations of the lack of serial computing mode and the lack of "thinking" ability that hinder the inference ability of large language models.

[0055] 2) The present invention enables large language models to have thinking ability through a novel injection thinking training paradigm, effectively alleviating the limitation of the lack of "thinking" ability.

[0056] 3) Based on large language models with thinking ability, combined with the human metacognition idea, a multi-agent reasoning framework is constructed, introducing a serial computing mode, effectively alleviating the limitation of the lack of serial computing mode.

[0057] 4) The present invention combines the metacognition idea in cognitive neuroscience to simultaneously solve the above two main limitations.

[0058] Comprehensively improve the reasoning ability of large language models. Description of the Drawings

[0059] Figure 1 It is a flowchart of a multi-agent reasoning framework based on metacognitive thinking proposed by the present invention. Detailed Implementation Modes

[0060] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and general knowledge in the art, and the present invention has no particularly restricted content.

[0061] Embodiment 1

[0062] Refer to Figure 1 , the construction of a multi-agent reasoning framework based on metacognitive thinking specifically includes the following steps:

[0063] Step 1: Construct a metacognitive agent training dataset

[0064] The thinking injection prompt words and user instruction prompt words are both written by human experts. The data used can be obtained from the public datasets TriviaQA and AlpacaEval. LLM 0 and Judge can be any open-source large language model. First, design a thinking injection prompt word (Prompt) P t , which is used to stimulate the thinking process of the large language model. Then, splice P t with the user instruction prompt word to obtain an input prompt and input it into the large language model to obtain an output result, which is specifically expressed by the following formula (a):

[0065] LLM 0 (x i ) = {z i , y i} (a).

[0066] Among them, LLM 0 represents an ordinary large language model without additional training. In the output result, z i and y i represent the thinking process and the answer respectively. Subsequently, a judge model is used to judge the quality of y i in the output, which is specifically expressed by the following formula (b):

[0067] Judge(y i ) = label, label ∈ {pos, neg} (b).

[0068] Among them, Judge represents the judge model, which is used to judge the quality of y i and generate a label label for it. label = pos / neg represents that the current y i is a positive sample / negative sample.

[0069] Finally, the training dataset D of the metacognitive agent meta is represented by the following formula (c) as:

[0070]

[0071] Among them, and respectively represent the positive and negative samples D in the answer meta represents the training dataset of the metacognitive agent.

[0072] Step 2: Obtain the metacognitive agent through the thought injection training paradigm

[0073] The LLM 0 adopted can be any open-source large language model. Using the Direct Preference Optimization algorithm, train an LLM meta on the D obtained in Step 1 0 to obtain the metacognitive agent Agent meta . The loss function for training is represented by the following formula (d) as:

[0074]

[0075] Among them, π = LLM 0 represents the large language model to be trained. By minimizing the above loss function, Agent meta can be obtained, which is specifically represented by the following formula (e) as:

[0076]

[0077] Agent meta simulates the "slow thinking" process of the human brain and is used to process complex problems. It implicitly generates a thinking process while generating an answer, which enables Agent meta to allocate more computational resources to complex problems, effectively alleviating the limitation of the lack of "thinking" ability mentioned in the background technology. It is specifically represented by the following formula (f) as:

[0078] Agent meta (x) = {z latent , y} (f).

[0079] Among them, x, zlatent x and y represent the input, implicit thinking process, and answer output respectively.

[0080] Step 3: Construct the training dataset for the central agent

[0081] The data used can be obtained from the public datasets TriviaQA and AlpacaEval, and the LLM 0 can be any open-source large language model. To ensure the generalization of the central agent, first collect data samples from multiple fields (mathematics, logical reasoning, general knowledge, medicine, law) and process them into the form of question-and-answer pairs, which is specifically represented by the following formula (g):

[0082] Multi_data = {Q i , A i} (g).

[0083] Then, divide the data sample pairs into two categories, simple and difficult, according to the difficulty level of the questions and label them, which is specifically represented by the following formula (h):

[0084] D hub = {Q i , A i , R i} (h).

[0085] Among them, D hub represents the training dataset for the central agent, and R i ∈ {easy, hard} represents the difficulty level label of the question.

[0086] Step 4: Obtain the central agent based on the supervised fine-tuning training paradigm

[0087] The LLM 0 used can be any open-source large language model. Use the supervised fine-tuning training paradigm for the classification task to train an LLM hub on the D 0 obtained in Step 3 to get the central agent Agent hub , and the loss function for training is represented by the following formula (i):

[0088]

[0089] Among them, π = LLM 0 represents the large language model to be trained, and R and represent the true difficulty label and the predicted difficulty label respectively.

[0090] By minimizing the above loss function, Agent hub can be obtained, which is specifically represented by the following formula (j):

[0091]

[0092] Agent hub Used to judge the difficulty level of the input problem to determine whether to enable the Agent obtained in Step 2 meta , specifically expressed as the following formula (k):

[0093]

[0094] Step Five: Construct a multi-agent reasoning framework based on the metacognitive thought

[0095] Use P refine Written by human experts and inspired by the human metacognitive thought, this multi-agent reasoning framework consists of a cognitive agent Agent cog , the metacognitive agent Agent obtained in Step 2 meta and the central agent Agent obtained in Step 4 hub , and is specifically expressed as the following formula (l):

[0096] MetaInf = (Agent cog , Agent meta , Agent hub ) (l).

[0097] Among them, MetaInf represents the multi-agent reasoning framework constructed based on the metacognitive thought, and Agent cog is served by an ordinary large language model LLM without additional training 0 .

[0098] In MetaInf, Agent cog is used to handle simple tasks, and Agent meta is used to handle complex tasks. In addition, in order to balance performance and resource consumption, this multi-agent reasoning framework uses Agent hub to analyze the difficulty level of the task, so as to activate Agent cog when receiving a simple task, and activate Agent with stronger reasoning ability but greater resource consumption meta when receiving a difficult task. Therefore, MetaInf has two working modes: 1) simple task processing; 2) complex task processing

[0099] When Agent hub judges that the current task is simple, MetaInf will activate the simple task processing mode, that is, only use Agent cog to solve the task, specifically expressed as the following formulas (m) to (n):

[0100] Agent hub (Q) = easy (m);

[0101] O=Agent cog (Q) (n).

[0102] Among them, Q represents the query of the current task, and O represents the output result of MetaInf.

[0103] When Agent hub When the current task is judged to be difficult, MetaInf will activate the difficult task processing mode, that is, first use Agent cog Generate intermediate results and then use Agent meta Combine the optimization instructions to optimize the intermediate results to get the final output, as shown in the following formulas (o) to (q):

[0104] Agent hub (Q) = hard (o);

[0105] S=Agent cog (Q) (p);

[0106] O=Agent meta (P refine ,S) (q).

[0107] Among them, S represents the intermediate result, P refine This working mode introduces a serial computing mode, which effectively alleviates the limitation of the lack of a serial computing mode mentioned in the background technology.

[0108] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A method for constructing a multi-agent reasoning framework based on metacognitive thinking, characterized in that: The multi-agent reasoning architecture is used to enhance the reasoning ability of the large language model to simulate the thinking mode of the human brain. The method specifically includes the following steps: Step 1: Build a metacognitive agent training dataset 1-1: Injecting thoughts into prompt words P and user command prompt words Perform splicing and get input prompt Input it into the large language model and obtain the output result shown in the following formula (a): LLM0(x i )={z i ,y i } (a); Among them, LLM0 is a common large language model without additional training; z i and i The thinking process and the answer respectively; 1-2: Use the judge model shown in the following formula (b) to judge the output result y i The pros and cons: Judge(y i )=label,label∈{pos,neg} (b); Among them, Judge is the representative judge model that judges whether y is good or bad, and generates a label label for it. Label = pos / neg represents the current y i is a positive sample / negative sample; 1-3: Metacognitive Agent Training Dataset D meta It is expressed by the following formula (c): in, and are the positive and negative samples in the answer respectively; Step 2: Acquire metacognitive agents through thought-injection training paradigm 2-1: Using the direct preference optimization algorithm, in the metacognitive agent training dataset D mea The metacognitive agent Agent is trained as shown in the following formula (d) meta The loss function L is: Among them, π=LLM0 is the large language model to be trained; 2-2: Minimize the above loss function L and obtain the "slow thinking" agent shown in the following formula (e) meta : in, It means to find the π that makes L the smallest; The "slow thinking" agent meta While generating the answer, it implicitly generates the following thinking process Agent represented by formula (f) meta (x): Agent meta (x)={z latent ,y} (f); Among them, x, z latent and y represent the input, implicit thinking process, and answer output respectively; Step 3: Build a central agent training dataset 3-1: Collect data samples from multiple fields such as mathematics, logical reasoning, general knowledge, medicine and law, and process them into question-answer pairs Multi_data using the following formula (g): Multi_data={Q i ,A i } (g); Among them, Q i , and A i are questions and answers respectively; 3-2: Divide the data samples into two categories: simple and difficult according to the difficulty of the problem, and label them with R using the following formula (h): i : D hub ={Q i ,A i ,R i } (h); Among them, D hub is the central agent training dataset; R i ∈{easy,hard} represents the difficulty label of the problem; Step 4: Obtain the central agent through supervised fine-tuning training paradigm 4-1: Using the supervised fine-tuning training paradigm for classification tasks on the central agent training dataset D hub Train LLM0 to get the central agent Agent hub , the training uses the loss function L represented by the following formula (i) π : Among them, π = LLM0 is the large language model to be trained; R and They are the real difficulty label and the predicted difficulty label respectively; 4-2: By minimizing the loss function L π , we get the central agent Agent represented by the following formula (j): hub : 4-3: Using the Central Agent hub The agent that determines the difficulty of the input question and is represented by the following formula (k) hub (Q) Determine whether to enable "Slow Thinking": Step 5: Construct a multi-agent reasoning framework based on metacognitive thinking The multi-agent reasoning framework MetaInf is represented by the cognitive agent Agent as follows (l): cog , Metacognitive Agent meta and the central agent hub Built with: MetaInf=(Agent cog ,Agent meta ,Agent hub ) (l); The multi-agent reasoning framework MetaInf has two working modes for processing simple tasks and complex tasks. The cognitive agent cog Acts as a cognitive agent for simple tasks for the untrained large language model LLM0; The metacognitive agent meta Intelligent agents for handling complex tasks; The central agent hub To introduce the serial computing mode, the Agent is activated when a simple task is received. cog , activate Agent when receiving difficult tasks meta , the specific processing is as follows: 5-1: Agent hub When the current task is judged to be simple, MetaInf activates the simple task processing mode, that is, only using Agent cog Solve the task, which is specifically expressed by the following (m) to (n) formulas: Agent hub (Q)=easy (m); O=Agent cog (Q) (n); Among them, Q is the query of the current task; O is the output result of MetaInf; 5-2: Being an Agent hub When the current task is judged to be difficult, MetaInf activates the difficult task processing mode, that is, first use Agent cog Generate intermediate results and then use Agent meta The intermediate results are optimized by combining the optimization instructions to obtain the final output, which is specifically expressed by the following formulas (o) to (q): Agent hub (Q)=hard (o); S=Agent cog (Q) (p); Q=Agent meta (P refine ,S) (q); Among them, S is the intermediate result; P refine ,S is the optimization instruction.

Citation Information

Patent Citations

  • Method and device for saving reasoning computing power of AI large model

    CN118211660A

  • Complex problem reasoning method based on dynamic collaboration of large language model and domain knowledge base

    CN118798368A