Sustainable learning multi-agent reasoning method and system based on thinking map
By constructing a thinking diagram and a hierarchical knowledge base, multi-agent systems optimize task decomposition and collaborative reasoning, solving the shortcomings of existing systems in task decomposition and resource consumption, and achieving efficient and flexible complex task processing.
Patent Information
- Application Number
- CN202510362038.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
AI Technical Summary
The existing multi-agent system is not flexible enough in task decomposition and coordination mechanisms, and lacks effective utilization of historical empirical data, resulting in excessive consumption of computing resources and inefficient efficiency, making it difficult to adapt to the needs of complex tasks.
By constructing a thinking map, introducing a historical experience library and a hierarchical knowledge base, using multiple agents to work collaboratively, including task planning agents, step execution agents, task solution agents and reflection agents, and combining external calling mechanisms to optimize task decomposition and continuous learning.
It improves the efficiency and flexibility of multi-agent systems in complex tasks, reduces computing resource consumption, enhances the system's adaptability and knowledge retrieval capabilities, and realizes continuous learning and optimization.
Smart Images

Figure CN120258141A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to multi-agent collaboration technology, and particularly to sustainable learning multi-agent collaborative reasoning technology based on mind maps. Background Art
[0002] An agent is a specific application form of artificial intelligence AI, emphasizing autonomous behavior and interaction capabilities in a specific environment. It refers to an entity that can perceive, make decisions, and execute actions in an environment, which can be software, hardware, or a combination of both. Agents can be AI-driven or rule-based non-learning systems; they are usually based on rule systems, planning algorithms, reinforcement learning, or multi-agent collaboration technology; and are applicable to scenarios that require interaction with the environment and autonomous decision-making, such as autonomous driving and robotics.
[0003] A multi-agent system is a collection of multiple agents that jointly complete complex tasks through interaction and collaboration. Communication, collaboration, and competition are required among the multiple agents in the system. Multi-agent systems are of great significance in fields such as autonomous driving, collaborative robotics, and intelligent decision support. With the rapid development of technologies such as artificial intelligence, the Internet of Things, and big data, it has become possible and is gradually widely applied to decompose and process complex tasks through multi-agent collaboration.
[0004] In terms of intelligent decision support, through the collaborative reasoning of a multi-agent system, the efficiency and accuracy of task execution can be effectively improved, thereby providing support for decision-making in complex scenarios. However, with the increase in task complexity, the existing multi-agent systems mainly rely on the emergent capabilities of large language models themselves to achieve task decomposition. Although this method can complete tasks to a certain extent, it often has problems such as unreasonable results and low efficiency when facing complex tasks. In addition, existing systems usually lack effective utilization of historical data and cannot optimize the reasoning process through continuous learning. At the same time, many systems rely on a single large model for task processing, resulting in excessive consumption of computing resources and high deployment costs.
[0005] Most existing multi-agent systems rely on a large model for task decomposition and then task processing. Its disadvantages are: (1) The task decomposition and coordination mechanism is not flexible enough to meet the requirements of complex tasks; (2) There is a lack of effective utilization of historical experience data, and the reasoning process cannot be optimized through continuous learning; (3) After task decomposition, the large model still needs to rely on itself to execute specific tasks, resulting in excessive consumption of computing resources. Although the large model has rich general knowledge and reflective reasoning capabilities, its efficiency is low when executing specific tasks and it is difficult to meet the needs of practical applications; (4) A single knowledge base is difficult to meet the requirements of complex tasks. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a collaborative technology that decomposes complex tasks more efficiently and has sustainable learning by constructing a mind map and coordinating among multiple sustainable learning agents.
[0007] The technical solution adopted by the present invention to solve the above technical problem is a sustainable learning multi-agent reasoning method based on a mind map, including the following steps:
[0008] The historical experience library receives the text data input by the user and retrieves the historical experience data through vector similarity calculation; if there is an entry whose similarity with the input text exceeds the preset threshold, the entry with the highest matching degree is directly output as the answer to this query; otherwise, the text data is forwarded to the task planning agent;
[0009] The task planning agent constructs a mind map based on the text data, then performs task planning based on the mind map, decomposes the task into several subtasks, and generates the corresponding sub-steps for solving each subtask and transmits them to the step execution agent;
[0010] The step execution agent executes the sub-steps according to the description and parameters of the sub-steps, stores each sub-step and its corresponding execution result in a dictionary, and passes the dictionary to the task solving agent; in order to enhance the flexibility of the system, the present invention introduces an external call mechanism, and calls external interfaces, knowledge bases, and function mini-models according to the description of the sub-steps during execution; allows the system to call mini-models, external tools, or knowledge bases according to task requirements, or dynamically adjust the task execution strategy by receiving real-time data from external systems, so as to ensure that the system can effectively handle complex tasks in a changing environment;
[0011] The task solving agent analyzes the dictionary received from the step execution agent according to the text data input by the user, uses a large language model to organize the analysis result and outputs the answer to this query; when receiving the task execution optimization suggestion from the reflection agent, re-outputs the answer to this query according to the task optimization suggestion;
[0012] The reflection agent outputs task execution optimization suggestions to the task planning agent or / and the task solving agent according to the historical conversation data; when the task planning agent receives the task execution optimization suggestion from the reflection agent, re-forms a mind map according to the text data and the task optimization prompt; when the task solving agent receives the task execution optimization suggestion from the reflection agent, re-outputs the answer to this query according to the task optimization suggestion.
[0013] Correspondingly, the present invention further provides a multi-agent system for using the above method, including a task planning agent, a step execution agent, a task solving agent, a reflection agent, a historical experience library, and a knowledge base. Each agent collaborates to complete the reasoning by implementing the above steps.
[0014] Further, after outputting the task execution optimization suggestions, the reflection agent also collects user feedback, and then evaluates the execution results of the task solving agent in combination with the user feedback and historical conversation data. The execution results with qualified evaluation results are sent to the historical experience library as historical experience data, and the execution results with unqualified evaluation results are discarded.
[0015] The historical experience library receives and stores the historical experience data from the reflection agent.
[0016] Further, for the execution results with unqualified evaluation results, the reflection agent will request the user to provide the correct solution and add the information provided by the user to the historical experience library for subsequent learning and improvement.
[0017] The historical experience library receives and stores the information provided by the user for the unqualified execution results from the reflection agent.
[0018] Preferably, the knowledge base includes a general knowledge base and a dedicated knowledge base. The hierarchical knowledge base system of the present invention includes three parts: a dedicated knowledge base, a general knowledge base, and a historical experience library. (1) Dedicated knowledge base: Designed for problems and tasks in specific fields, it contains specific information, rules, etc. required to handle problems in these fields; (2) General knowledge base: Covers a wider range of information and is suitable for solving basic cross-domain problems; (3) Historical experience library: Used to store the experience data accumulated in previous interaction processes, and continuously improves the system's problem-solving ability through continuous learning and optimization.
[0019] Further, the system stores historical experience in the knowledge base as the basis for continuous learning, enabling the agent to continuously optimize the reasoning process from actual experience.
[0020] To enhance the system's complex problem-solving ability, the task planning agent of the present invention uses a large model as the core "brain", responsible for task decomposition and coordination, while specific subtasks are processed by functional small models or external tools, thereby reducing the consumption of computing resources while ensuring performance. Specifically, the task planning agent includes a thinking generator, a thinking aggregator, a thinking enhancer, and a thinking solver.
[0021] Through the query matching algorithm and continuous learning mechanism of the historical experience library, the inference efficiency and adaptability of the system are improved. The collaborative inference mechanism combining the mind map and the hierarchical knowledge base enhances the inference ability of the system in different tasks and scenarios. At the same time, the intelligent agent system equipped with a dedicated knowledge base works in coordination with the external call mechanism, further improving the knowledge retrieval and generation ability of the system, enabling it to better adapt to complex tasks and changing environments. Through the evaluation of the reflection agent and the user feedback mechanism, the system can continuously optimize the task execution results and store valuable execution results in the historical experience library to achieve continuous learning and performance improvement.
[0022] The beneficial effects of the present invention are as follows: by introducing a mind map to optimize the task decomposition process, using historical experience data to construct a multi-level knowledge base for continuous learning, and reducing the deployment cost by calling dedicated small models or external tools, the efficiency and flexibility of the multi-agent system are significantly improved. Brief Description of the Drawings
[0023] Figure 1 It is a schematic diagram of the multi-agent collaborative inference workflow.
[0024] Figure 2 It is a task decomposition mind map.
[0025] Figure 3 It is a hierarchical knowledge base structure Detailed Embodiments
[0026] The framework of the sustainable learning multi-agent collaborative inference system based on the mind map is as Figure 1 shown, including a task planning agent, a step execution agent, a task solving agent, a reflection agent, a historical experience library, a general knowledge base, and a dedicated knowledge base;
[0027] The historical experience library receives the text data input by the user and retrieves the historical experience data through vector similarity calculation. If there is an entry whose similarity to the input text exceeds the preset threshold, the entry with the highest matching degree is directly output as the answer to this query; otherwise, the text data is forwarded to the task planning agent; it receives and stores the historical experience data from the reflection agent and the information provided by the user for the unqualified execution results;
[0028] The task planning agent is used to construct a mind map based on the text data, then perform task planning based on the mind map, decompose the task into several subtasks, and generate the corresponding sub-steps for solving each subtask and transfer them to the step execution agent; when receiving the task execution optimization suggestion from the reflection agent, it reformulates the mind map according to the text data and the task optimization prompt;
[0029] The step execution agent is used to execute sub-steps according to the descriptions and parameters of sub-steps, store each sub-step and its corresponding execution result in a dictionary, and pass the dictionary to the task-solving agent; when executing, it calls external interfaces, general knowledge bases, specialized knowledge bases, and functional mini-models according to the descriptions of sub-steps;
[0030] The task-solving agent is used to analyze the dictionary received from the step execution agent according to the text data input by the user, use a large language model to organize the analysis results and output the answer to this query; when receiving task execution optimization suggestions from the reflection agent, re-output the answer to this query according to the task optimization suggestions;
[0031] The reflection agent is used to output task execution optimization suggestions to the task planning agent or / and the task-solving agent according to historical conversation data; collect user feedback, and combine historical conversation data to evaluate the execution results of the task-solving agent, evaluate the execution results of the task-solving agent, and send the qualified execution results to the historical experience database as historical experience data, and discard the unqualified execution results; for the unqualified execution results, request the user to provide the correct solution, and add the information provided by the user to the historical experience database for subsequent learning and improvement.
[0032] The task planning agent includes a thinking generator, a thinking aggregator, a thinking enhancer, and a thinking solver;
[0033] The thinking generator is used to generate multiple thinking nodes by expanding from a single thinking; when starting, use the text data input by the user as the initial thinking node;
[0034] The thinking aggregator is used to select several thinking nodes for combination, form new thinking nodes and their dependencies, and add them to the thinking map;
[0035] The thinking enhancer is used to re-adjust and optimize the formed thinking nodes, and adjust the thinking map according to the optimization results;
[0036] The thinking solver generates a path according to the optimized thinking map and determines the sub-steps corresponding to the sub-tasks.
[0037] Each thinking node contains two attributes: a question and a score. The question of the initial thinking node is the text data input by the user, and the score is obtained through a predefined scoring rule;
[0038] The predefined scoring rules of the thinking generator include five dimensions: feasibility, contribution, complexity, risk, and resource consumption. Points will be deducted if the requirements are not met in a certain dimension; nodes with a score greater than or equal to the preset score are considered qualified, otherwise they are unqualified;
[0039] The predefined scoring rules of the thought aggregator include five dimensions: relevance, importance, repetition, effectiveness, and simplicity. Points will be deducted if the requirements are not met in any dimension. Nodes with scores greater than or equal to the preset score are considered qualified, otherwise they are unqualified.
[0040] The thinking enhancer readjusts and optimizes unqualified thinking nodes. If the score of the readjusted thinking node is greater than or equal to the preset score, the thinking node will be retained; otherwise, it will not be adopted.
[0041] The thinking solver uses a depth-first search algorithm to traverse the thinking map, obtain all possible paths, calculate the total scores of all thinking nodes on each path, and sort them by score. It selects the three paths with the highest scores as alternatives, and determines one by one whether the paths in the alternatives can solve the problem represented by the text data input by the user. If so, the task planning is terminated and the path is output as the sub-step corresponding to solving the subtask.
[0042] The system combines the collaborative work of multiple agents, mind maps, hierarchical knowledge bases and historical experience bases to form an efficient task decomposition and collaborative reasoning solution. Through the collaboration of multiple agents, the system solves the problems input by users step by step. The mind map is introduced as a task planning tool to decompose complex tasks into several subtasks, and execute them in order according to the dependencies of the subtasks. In addition, the system designs a three-layer knowledge base system, including a dedicated knowledge base, a general knowledge base and a historical experience base, in order to provide multi-level knowledge support and continuous learning. In order to enhance the flexibility of the system, the system obtains resources from external tools, knowledge bases or small models as needed when executing tasks, and can adjust the task execution strategy according to the real-time data provided by the external system to ensure efficient operation in a dynamic environment.
[0043] The multi-agent system of the present invention includes four core agents, each of which is responsible for different functions and cooperates with each other to complete tasks, including: (1) Task planning agent: responsible for receiving questions input by users, planning tasks through mind maps, decomposing tasks into several subtasks, and generating corresponding execution substeps. (2) Step execution agent: according to the description and parameters of the subtask, it performs corresponding actions, and may call external interfaces to obtain real-time data or query the knowledge base. (3) Task solving agent: integrates the execution results of subtasks, generates the final answer, and answers the user's initial question. (4) Reflection agent: analyzes the entire task execution process, combines user feedback and self-criticism mechanism, evaluates the task execution effect, and generates improvement suggestions to promote system optimization.
[0044] Design of a three - level knowledge base, namely a dedicated knowledge base, a general knowledge base, and a historical experience base: (1) The dedicated knowledge base contains professional information and rules for specific fields, used to handle field - specific tasks, ensuring that the system can accurately solve problems in specific industries or disciplines; (2) The general knowledge base contains basic knowledge across fields, used to solve some general and non - professional problems, providing extensive support to ensure that the system can handle diverse task requirements; (3) The historical experience base stores the experience data accumulated by the system in past tasks, continuously optimizing the system's execution strategies and models. Through the design of the three - level knowledge base, the system can not only handle complex field tasks but also adapt to cross - field requirements and achieve continuous learning and optimization during long - term operation.
[0045] Correspondingly, a sustainable learning multi - agent collaborative reasoning method based on a mind map for the above - mentioned system includes the following steps:
[0046] 1) According to the problem input by the user, use a query - matching algorithm based on the historical experience base to find corresponding entries. If matching entries that reach the preset similarity threshold exist, directly output the answer to this query. Otherwise, enter step 2), and pass the user input to the task - planning agent for processing;
[0047] 2) The task - planning agent receives the problem input by the user, conducts task planning through a mind map, and decomposes the task into several sub - steps;
[0048] 2 - 1) Starting from the problem input by the user, construct an initial mind map and generate new mind nodes.
[0049] Evaluate these new nodes to determine whether they meet the requirements. Qualified nodes will be incorporated into the mind map;
[0050] 2 - 2) Select several nodes from the already qualified nodes for combination to form new aggregated nodes. Evaluate the aggregated nodes, and qualified nodes will be added to the mind map;
[0051] 2 - 3) Adjust and optimize the unqualified nodes and re - evaluate. If the adjusted nodes meet the requirements, retain them and make corresponding adjustments to the mind map;
[0052] 2 - 4) By traversing the mind map, screen out the three paths with the highest scores and determine whether these paths can solve the problem. If so, output the paths as the task - planning result; otherwise, repeat the relevant operations until the end condition is met;
[0053] 3) The step - execution agent performs corresponding actions according to the descriptions and parameters of the sub - steps, specifically including:
[0054] 3-1) If a sub-step involves external information, the step execution agent calls the corresponding external interface according to the description of the sub-step to obtain the corresponding external real-time data; if no external information is required, directly query the dedicated knowledge base or complete the current instruction according to the general knowledge base;
[0055] 3-2) Based on the external information or general knowledge base information, combined with the knowledge in the dedicated knowledge base, generate the input of the functional small model for solving the sub-step generated by the task planning agent. If a functional small model in a specific domain is not required, directly generate the result corresponding to the current sub-step and enter step 3-4);
[0056] 3-3) The functional small model executes according to the input and generates the result corresponding to the sub-step;
[0057] 3-4) Execute each sub-step in sequence until all sub-steps are completed;
[0058] 3-5) After the step execution agent finishes execution, the system stores the sub-step and its corresponding execution result in a dictionary Res_dict, and the stored information is {step 1: result 1,..., step n: result n}, and then passes this dictionary to the task solving agent;
[0059] 4) The task solving agent receives the dictionary Res_dict containing all sub-steps and their execution results. Based on the question initially proposed by the user, this agent integrates and summarizes the information in Res_dict to generate a precise answer to the initial question. Finally, the task solving agent passes the obtained answer result to the reflection agent for further analysis.
[0060] 5) The reflection agent analyzes the entire dialogue process through self-criticism and requests feedback from humans. Then, it combines the user feedback information with its self-criticism mechanism based on historical conversations for comprehensive analysis. The analysis results include the evaluation and reflection on the current task execution process and results, so as to diagnose whether there are deficiencies in the task results. If deficiencies are found, the agent will generate task suggestions for the next round and return to step 1). Otherwise, enter step 6).
[0061] 6) In order to continuously improve the performance of the agent, the controller requests the user to evaluate the task completion situation and decides how to process the execution data according to the scoring result. The specific steps include:
[0062] 6-1) After the reflection agent analyzes the task process, the controller requests the user to rate the task completion situation (1-10 points).
[0063] 6-2) Determine whether to save the result according to the score: if the score ≥ the preset threshold, save it to the historical experience database for future reference; if the score < the preset threshold, discard the data this time.
[0064] 6-3) If the task score is lower than the preset threshold, the system will ask the user to provide the correct solution and add this information to the historical experience database for subsequent learning and improvement.
[0065] 6-4) The system regularly extracts data from the historical experience database. Both successful examples and solutions corrected by users are used for model fine-tuning and updating. In this way, the agent can continuously learn the latest effective strategies and gradually improve its robustness and adaptability.
[0066] The sustainable learning multi-agent collaborative reasoning method based on the mind map realizes the efficient decomposition and collaborative reasoning of tasks. To enhance the system's continuous learning ability and retrieval and generation ability, the process of continuous learning is achieved by using the data in the historical experience database to fine-tune and update the model. At the same time, with the support of the general knowledge base and the dedicated knowledge base, the system's adaptability and learning ability are further improved.
[0067] Embodiment
[0068] The specific implementation steps of each component of the multi-agent system during collaborative reasoning are as follows:
[0069] Step 1: Query and match the problem. The historical database is designed as Figure 3 shown below. The specific query steps are as follows:
[0070] Step 1-1: Design the historical experience database, create a two-dimensional data table in the relational database to store historical experience data, including fields: primary key Id, query or problem Query, answer or result Answer; Query and Answer are values Value; Id is the key Key;
[0071] Step 1-2: Design a vector database table corresponding to the historical experience database table, create a collection in different vector databases such as Milvus and Faiss to specifically store vector data, including field values: primary key vector Vector_id, vector Vector representing the value Value. Associate the Id field of the historical experience database with the Vector_id field of the vector database table through foreign key constraints to ensure that each vector data corresponds to a historical experience entry, facilitating the search for relevant entries by calculating the similarity of query vectors;
[0072] Step 1-3: Use the pre-trained DistilBERT model to convert the preprocessed query into a vector representation;
[0073] Step 1-4: For a new query vector, calculate its cosine similarity with each vector in the vector database using the following formula:
[0074]
[0075] Step 1-5: If there are entries in the vector database greater than the preset threshold, filter out the Id of the matching vector with the highest similarity from the entries greater than the threshold, and retrieve the corresponding answer from the historical experience database according to the Id and return it directly to the user. Otherwise, go to Step 2;
[0076] Step 2: The task planning agent realizes task planning and decomposition with the help of a thinking graph, including a thinking generator, a thinking aggregator, a thinking enhancer, and a thinking solver. The thinking graph consists of thinking nodes and directed edges between the nodes. The thinking nodes store task sub-steps, and the directed edges store the dependency relationships between the nodes. The thinking generator expands from a single thinking node to generate multiple thinking nodes; the thinking aggregator can select several thinking nodes for combination to form multiple new thinking nodes; the thinking enhancer re-adjusts and optimizes the existing thinking nodes; the thinking solver generates paths based on the optimized thinking nodes and determines the sub-steps corresponding to the sub-tasks. The task decomposition thinking graph is as Figure 2 shown, and the specific steps are as follows:
[0077] Step 2-1: Each thinking node contains two attributes, "question" and "score". When the task starts, use the question input by the user as the initial thinking node to construct the initial thinking graph. The thinking generator generates several new thinking nodes based on the initial thinking node and takes them as the leaf nodes of the initial thinking node. At the same time, the thinking generator evaluates the new thinking nodes according to the predefined scoring rules. This scoring rule discriminates the thinking from five dimensions: feasibility, contribution, complexity, risk, and resource consumption. If it does not meet the requirements in a certain dimension, 2 points will be deducted, and the full score is 10 points. Subsequently, the new thinking nodes and their dependency relationships are added to the thinking graph. Nodes with a new thinking node score greater than or equal to 6 points are considered qualified, otherwise they are unqualified.
[0078] Step 2-2: The thinking aggregator selects several nodes from the qualified thinking nodes to aggregate and generate new aggregated thinking nodes, supporting the generation of multiple different aggregated thinking nodes at one time, and the dependency nodes of each aggregated thinking node can be different. After aggregation, the aggregator scores the aggregated thinking nodes from five dimensions: relevance, importance, repeatability, effectiveness, and simplicity. If it does not meet the requirements in a certain dimension, 2 points will be deducted, and the full score is 10 points. Subsequently, the new thinking nodes and their dependency relationships are added to the thinking graph. Nodes with a new thinking node score greater than or equal to 6 points are considered qualified, otherwise they are unqualified.
[0079] Step 2-3: The thought enhancer pushes all unqualified thought nodes onto the stack, combines them with the existing qualified thought nodes, re-adjusts and optimizes the thought nodes on the stack in sequence, and scores them according to the scoring rules in Step 2-1. If the score of a re-adjusted thought node is greater than or equal to 6 points, the node is retained; otherwise, it is not adopted and backtracked. If the dependency relationship between the qualified thought nodes after adjustment changes, the thought map needs to be adjusted.
[0080] Step 2-4: The thought solver traverses the thought map using the depth-first search algorithm to obtain all possible paths. During this process, all paths containing unqualified thought nodes will be discarded. For the remaining paths, calculate the sum of the scores of all nodes in each path and sort them by score. Select the 3 paths with the highest scores. Selecting 3 paths can provide multiple alternative solutions while avoiding a significant increase in computational complexity due to too many paths. Then, the solver determines whether the selected paths can solve the initial problem. If there is at least one path that can solve the problem, the task planning is terminated, and these paths are output as the sub-steps corresponding to solving the subtask. Otherwise, repeat the thought generation, aggregation, and enhancement operations (Steps 2-2, 2-3, 2-4) until one of the following conditions is met: (1) The thought solver actively terminates the task planning; (2) The number of thought aggregations reaches 3 times. Limiting the number of thought aggregations to 3 times is based on the balance of the efficiency and resource consumption of the actual task planning. After multiple aggregations, the optimization space of the path may gradually decrease, while the computational cost will continue to increase.
[0081] Step 3: The step execution agent executes the corresponding tasks according to the task list, including the following steps:
[0082] Step 3-1: The step execution agent receives the task list from the task planning agent and performs the corresponding actions according to the description and parameters of each subtask. Each subtask may contain multiple sub-steps, and the step execution agent determines and executes the corresponding sub-steps according to the specific content of the subtask;
[0083] Step 3-2: If the subtask involves external information, the step execution agent parses the subtask description, determines the required external interfaces, and accesses the corresponding interfaces to obtain real-time data. These external interfaces may include sensor data, API calls, or other system interfaces; if the subtask does not require external information, the step execution agent queries the dedicated knowledge base or the general knowledge base to complete the instructions of the subtask;
[0084] Step 3-3: Based on external information or information from the general knowledge base, combined with the knowledge in the dedicated knowledge base, the step execution agent generates answers for the sub-steps corresponding to the subtasks. If the subtask is relatively complex or requires specific domain support, it generates the input for specific functional small models, which are designed specifically for specific tasks or subtasks and have high pertinence and efficiency. If functional small models are not required, the step execution agent directly generates the results of the current subtask;
[0085] Step 3-4: The specific domain functional small models perform calculations or inferences based on the input generated in Step 3-3 and generate answers corresponding to the sub-steps. The calculation or inference results of each sub-step provide necessary information for the final result of the subtask. These functional small models can be machine learning models, rule engines, or other forms of inference systems;
[0086] Step 3-5: The step execution agent sequentially executes each sub-step in the subtask until all sub-steps are completed;
[0087] Step 3-6: When all sub-steps are executed, the step execution agent stores each sub-step and its corresponding execution result in a dictionary Res_dict in the format {Step 1: Result 1,..., Step n: Result n}.
[0088] Step 3-7: The step execution agent passes this Res_dict dictionary to the task-solving agent for further processing and generating the final answer;
[0089] Step 4: The task-solving agent generates the final answer to the user's initial query based on the execution results provided by the step execution agent. The specific steps are as follows:
[0090] Step 4-1: The task-solving agent receives the Res_dict dictionary from the step execution agent, which contains each step and its corresponding result, in the format {Step 1: Result 1,..., Step n: Result n};
[0091] Step 4-2: Based on the question initially proposed by the user, the task-solving agent analyzes and summarizes the entire conversation history and the initial question based on predefined prompt words, conducts a detailed analysis of the results in the Res_dict dictionary, and identifies key information and data points;
[0092] Step 4-3: The task-solving agent uses the general capabilities of the large language model to understand and summarize the currently obtained analysis results and finally outputs the answer to the question;
[0093] Step 5: The reflection agent accepts the task solution. The final answer of the solution agent analyzes the entire conversation process through the self-criticism mechanism and conducts reasonable reflection in combination with the user feedback to ensure that the answer is reasonable, correct, and well-organized. The specific steps are as follows:
[0094] Step 5-1: The reflection agent combines the historical conversation data to conduct self-criticism and comprehensively evaluate the execution result of the task-solving agent. The self-criticism mechanism analyzes the potential deficiencies in the current task execution process based on the experience and rules in the historical conversation and proposes possible improvement suggestions.
[0095] Step 5-2: If the reflection agent gives the next suggestion, the multi-agent system will re-enter Step 2 or Step 4 to incorporate these suggestions and optimize the task execution. If no further suggestions are put forward, the system enters the user feedback stage and waits for the user's feedback on the final answer. The reflection agent also requires manual input from humans to provide feedback to verify the accuracy of its self-criticism. The effectiveness of the current task execution is comprehensively analyzed and evaluated by combining the user feedback and the self-criticism of the reflection agent itself. If deficiencies are found in the feedback, the reflection agent will generate new task suggestions and re-enter Step 2 or Step 4 to optimize the task execution.
[0096] Step 5-3: In the user feedback stage, if the user finds that the answer does not consider some key factors, they can enter specific prompt words to point out these omissions. After receiving the prompt, the multi-agent system will return to Step 2 to re-plan the task and ensure that the omitted factors are addressed in the revised plan. If the user does not provide feedback input or believes that the requirements have been met and no input is needed, the reflection agent determines that the task has been successfully executed and no next task is required.
[0097] Step 6: To continuously improve the performance of the agent, the controller requests the user to rate the task completion based on the evaluation result of the reflection agent and decides how to process the execution data according to the rating result. The specific steps are as follows:
[0098] Step 6-1: After the reflection agent completes the task evaluation, the controller requests the user to rate the task completion, and the rating range is from 1 to 10 points.
[0099] Step 6-2: According to the user rating, the controller decides whether to save the task execution result to the historical experience library. If the user rating ≥ the preset threshold, the relevant data of the task execution (including the query, answer, and corresponding vector representation) will be saved to the historical experience library and the corresponding vector database. If the rating < the preset threshold, the task execution result of this time will be discarded to avoid having a negative impact on subsequent tasks.
[0100] Step 6-3: If the user's score is lower than the preset threshold, the system asks the user to provide solutions and stores the information provided by these users in the historical experience database;
[0101] Step 6-4: The system regularly extracts data from the historical experience database. The extracted content includes successful examples and solutions provided by users. This data will be used for the fine-tuning and updating of the agent model so that the system can continuously learn and optimize strategies, gradually enhancing the robustness and adaptability of the agent.
Claims
1. A sustainable learning multi-agent reasoning method based on mind maps, characterized in that It includes the following steps: The historical experience database receives the text data input by the user and retrieves the historical experience data through vector similarity calculation. If there is an entry whose similarity with the input text exceeds the preset threshold, the entry with the highest matching degree is directly output as the answer to this query. Otherwise, the text data is forwarded to the task planning agent; The task planning agent forms a mind map based on the text data, then conducts task planning based on the mind map, decomposes the task into several subtasks, and generates the corresponding sub-steps for solving each subtask and passes them to the step execution agent; The step execution agent executes the sub-steps according to the description and parameters of the sub-steps, stores each sub-step and its corresponding execution result in a dictionary, and passes the dictionary to the task solving agent; when executing, it calls external interfaces, knowledge bases, and functional mini-models according to the description of the sub-steps; The task solving agent analyzes the dictionary received from the step execution agent according to the text data input by the user, and uses a large language model to organize the analysis result and output the answer to this query; When receiving the task execution optimization suggestion from the reflection agent, re-output the answer to this query according to the task optimization suggestion; The reflection agent outputs the task execution optimization suggestion to the task planning agent and / or the task solving agent according to the historical conversation data; When the task planning agent receives the task execution optimization suggestion from the reflection agent, it re-forms a mind map according to the text data and the task optimization; when the task solving agent receives the task execution optimization suggestion from the reflection agent, it re-outputs the answer to this query according to the task optimization suggestion.
2. The method according to claim 1, wherein After outputting the task execution optimization suggestion, the reflection agent also conducts user feedback collection, and then combines the user feedback and historical conversation data to evaluate the execution result of the task solving agent. The execution results with qualified evaluation results are sent to the historical experience database as historical experience data, and the execution results with unqualified evaluation results are discarded; The historical experience database receives and stores the historical experience data from the reflection agent.
3. The method according to claim 1, wherein The knowledge base includes a general knowledge base and a dedicated knowledge base; the general knowledge base is applicable to solving cross-domain basic tasks; The dedicated knowledge base is applicable to solving tasks in specific fields.
4. The method according to claim 1, wherein The historical experience database stores historical experience data in a two-dimensional data table in key-value form.
5. The method according to claim 1, wherein The mind map is composed of directed edges between mind nodes. The mind nodes store task sub-steps, and the directed edges store the dependency relationships between the nodes.
6. The method according to claim 5, wherein The task planning agent includes a mind generator, a mind aggregator, a mind enhancer, and a mind solver; The mind generator is used to expand from a single mind to generate multiple mind nodes; when starting, the text data input by the user is used as the initial mind node; The mind aggregator is used to select several mind nodes for combination, form new mind nodes and their dependency relationships and add them to the mind map; The mind enhancer is used to re-adjust and optimize the formed mind nodes, and adjust the mind map according to the optimization result; The mind solver generates a path according to the optimized mind map and determines the sub-steps corresponding to the subtasks.
7. The method according to claim 6, wherein Each thinking node contains two attributes: question and score. The question of the initial thinking node is the text data input by the user, and the score is obtained according to the predefined scoring rules. The predefined scoring rules of the thought generator include five dimensions: feasibility, contribution, complexity, risk, and resource consumption. Points will be deducted if any dimension fails to meet the requirements. Nodes with scores greater than or equal to the preset score are considered qualified, otherwise they are unqualified. The predefined scoring rules of the thought aggregator include five dimensions: relevance, importance, repetition, effectiveness, and simplicity. Points will be deducted if the requirements are not met in any dimension. Nodes with scores greater than or equal to the preset score are considered qualified, otherwise they are unqualified. The thinking enhancer readjusts and optimizes unqualified thinking nodes. If the score of the readjusted thinking node is greater than or equal to the preset score, the thinking node will be retained; otherwise, it will not be adopted. The thinking solver uses a depth-first search algorithm to traverse the thinking map, obtain all possible paths, calculate the total scores of all thinking nodes on each path, and sort them by score. It selects the three paths with the highest scores as alternatives, and determines one by one whether the paths in the alternatives can solve the problem represented by the text data input by the user. If so, the task planning is terminated and the path is output as the sub-step corresponding to solving the subtask.
8. The method according to claim 1, wherein Functional models are machine learning models, rule engines, or other forms of reasoning systems for specific tasks.
9. The method according to claim 2, wherein For execution results that are not qualified, the reflective agent will ask the user to provide the correct solution and add the user-provided information to the historical experience library for subsequent learning and improvement; The historical experience library receives and stores user-provided information about unqualified execution results from the reflective agent.
10. A sustainable learning multi-agent system based on mind maps, characterized in that, It includes task planning agent, step execution agent, task solving agent, reflection agent, historical experience base, knowledge base and small functional model; The historical experience database is used to receive the text data input by the user to query whether there is a matching entry in the historical experience data. If so, the matching entry is directly output as the answer to this query; if not, the text data is forwarded to the task planning agent; The task planning agent is used to form a mind map according to the text data, and then perform task planning based on the mind map, decompose the task into several subtasks, and generate substeps corresponding to solving each subtask and pass them to the step execution agent; after receiving the task execution optimization suggestions from the reflective agent, the mind map is re-formed according to the text data and task optimization; The step execution agent is used to execute the sub-steps according to the description and parameters of the sub-steps, store each sub-step and its corresponding execution result in a dictionary, and pass the dictionary to the task solving agent; during execution, the external interface, knowledge base and functional mini-model are called according to the description of the sub-steps; Knowledge bases include general knowledge bases and special knowledge bases. General knowledge bases are suitable for solving cross-domain basic tasks. Specialized knowledge bases are suitable for solving tasks in specific fields; Functional models are machine learning models, rule engines, or other forms of reasoning systems for specific tasks; The task-solving agent is used to analyze the dictionary received from the step-executing agent according to the text data input by the user, and use the large language model to organize the analysis results and output the answer to this query; when receiving the task execution optimization suggestions from the reflection agent, re-output the answer to this query according to the task optimization suggestions; The reflection agent is used to output task execution optimization suggestions to the task planning agent and / or the task-solving agent according to the historical conversation data.
Citation Information
Cited By
Multi-agent driven intelligent experiment design system, method, equipment and medium
CN120430422A
Intelligent agent self-adaptive evolution method and device of double-cooperation mechanism and electronic equipment
CN121835739A
An agent self-adaptive evolution method and device based on a double synergy mechanism and an electronic device
CN121835739B