Large model agent knowledge enhancement method and device based on self-awareness

By building knowledge bases and scenario determination criteria, the self-awareness ability of large-scale model agents is trained to enable them to selectively call knowledge, which solves the problem of poor execution of large-scale language models in complex decision-making and strategy planning tasks, and reduces the cost and complexity of knowledge injection, and improves decision stability and efficiency.

CN120218172APending Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510374384.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing large language models deal with complex decision-making and strategic planning tasks, they are difficult to capture the deep structure and logic of the task, resulting in poor execution results and lack of understanding and application of structured knowledge.

Method used

By constructing a knowledge base and knowledge selection module for agent planning, heuristic agent situation determination criteria are designed, including three scenarios: fast thinking, slow thinking and knowledge-based thinking. A two-stage training method combining supervised fine-tuning and reinforcement learning is adopted to train the knowledge-based self-awareness ability of large model agents so that they can selectively call knowledge.

Benefits of technology

It realizes that large-model agents can perform fast thinking, slow thinking and knowledge-based thinking when making motion decisions. They only introduce external knowledge during knowledge-based thinking, which reduces the cost and complexity of knowledge injection and improves the decision-making stability and efficiency of the agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218172A_ABST
    Figure CN120218172A_ABST
Patent Text Reader

Abstract

The invention discloses a large model agent knowledge enhancement method and device based on self-awareness, and the method comprises the steps: constructing an agent planning-oriented knowledge base, and providing a knowledge selection module for the knowledge base to form a knowledge system; designing a heuristic agent scene judgment criterion, wherein the heuristic agent scene judgment criterion comprises three scenes including fast thinking, slow thinking and knowledge-based thinking; according to a scene judgment criterion, constructing a self-consciousness training data set for the intelligent agent by utilizing the constructed knowledge system; on the constructed self-consciousness training data set, a supervised fine tuning and reinforcement learning combined two-stage training method is adopted to train the knowledge-based self-consciousness ability of the large model agent; and the trained large model agent performs reasoning and decision making, and the large model agent can determine different situations and make different decisions by itself in a form of outputting special marks, so that the cost and complexity of knowledge injection can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a method for enhancing the knowledge of a large model intelligent agent based on self-awareness. Background Art

[0002] The emergence of large language models marks a key transformation in the development of intelligent agents. The various GPT series large models announced have pushed the era of large language models to a new height. The introduction of these models has not only accelerated the application of AI intelligent agents in the field of natural language processing but also provided a new development opportunity for AI intelligent agents. Different from traditional AI intelligent agents that mainly demonstrate specific capabilities in specific tasks, large model intelligent agents mainly show broader applicability and depth in the ability to understand and generate language through their excellent language processing capabilities. This improvement in ability has not only changed the application fields of AI intelligent agents but also opened up new paths for understanding and applying artificial intelligence.

[0003] In the field of AI, building intelligent agents that can adapt to complex environments and efficiently complete tasks has been a long-term pursuit of people. With the progress of technology, especially in the application of artificial intelligence, knowledge-enhanced AI intelligent agents have shown their indispensable value. Such intelligent agents can not only improve their problem-solving ability by integrating a wide range of knowledge systems but also make more reasonable and accurate decisions in a changing environment.

[0004] First of all, most existing AI models, especially text-based large language models, mainly rely on text data during training. Although this method has achieved remarkable results in language understanding and generation, it often lacks the understanding and application of planning knowledge. For example, in tasks that require complex decision-making and strategy planning, models trained only with text often have difficulty capturing the deep structure and logic of the tasks, resulting in poor execution effects. Therefore, the training of intelligent agents not only needs to process language data but also integrate structured knowledge to make up for the lack of this planning knowledge. In specific applications, such as professional fields like medicine, law, or finance, intelligent agents not only need to process a large amount of information but also ensure that their decisions and actions are consistent with known facts and logical principles in the real world. The core of this ability lies in that intelligent agents can access and utilize rich structured knowledge resources, which provide in-depth insights and verified data about the real world, making the decisions of intelligent agents more reliable and effective. In addition, for intelligent agents operating in the physical world, such as autonomous vehicles or service robots, understanding and applying physical laws become particularly important. These intelligent agents need to predict and plan their actions based on physical laws, such as calculating the movement trajectory of an object or operating a robotic arm to reach an accurate position. Only by fully integrating these world models and physical knowledge can intelligent agents operate safely and effectively in the real world.

[0005] However, pre-trained LLMs often perform poorly in specific domains, such as web navigation or emerging technology applications, due to the lack of sufficient domain-specific knowledge. To address this issue, the AI research community has begun to explore how to more effectively integrate offline experiences and domain knowledge into LLMs to enhance their ability to make decisions and perform tasks.

[0006] The potential of knowledge-enhanced AI agents in various application domains will be increasingly recognized and utilized. According to the form of knowledge storage, existing knowledge enhancement methods include symbolic knowledge enhancement and parametric knowledge enhancement. Projects such as KnowAgent, AutoGuide, and WKM have successfully demonstrated that through knowledge-enhanced prompting capabilities, the behavior and decision-making processes of agents can be greatly optimized, thus showing higher efficiency and adaptability when facing complex challenges.

[0007] When large model agents process human needs, hallucinatory (unreasonable) actions often occur. This is the essential idea of KnowAgent. KnowAgent models all the actions that may be required to complete a task in the form of a graph, where the nodes of the graph are actions and the edges are the constraint relationships between two connected actions. Feeding this graph to the large model agent in natural language form can, to some extent, alleviate the hallucinatory actions. However, there are the following drawbacks: Drawback 1: Although GPT-4 is introduced, the action knowledge of KnowAgent still mainly relies on human experts to construct, which has a large labor cost and extremely poor transferability between domains, affecting the learning efficiency of the agent; Drawback 2: The knowledge injection of KnowAgent is carried out mindlessly, that is, regardless of the real needs of the agent. When the agent does not need knowledge, mindless knowledge injection may instead lengthen the text, increase the text understanding difficulty of the large model, and reduce the stability of the agent's decision-making. At the same time, it increases the cost of model training and inference.

[0008] WKM can be considered as proposed based on KnowAgent, aiming to solve the problems of high labor cost and poor transferability of symbolic knowledge. WKM enables the agent to self-synthesize knowledge, avoiding the cost of manually writing knowledge, and at the same time proposes a parametric knowledge enhancement mechanism, that is, training knowledge into a parametric model and using the generalization of the model to enhance the generalization of knowledge. However, WKM still has the following drawbacks: Drawback 1: The knowledge injection of WKM is also carried out mindlessly, that is, regardless of the real needs of the agent. When the agent does not need knowledge, mindless knowledge injection may instead lengthen the text, increase the text understanding difficulty of the large model, and reduce the stability of the agent's decision-making. At the same time, it increases the cost of model training and inference. Summary of the Invention

[0009] In view of the above, the object of the present invention is to provide a method for enhancing the knowledge of large model agents based on self-awareness, enabling the agents to start from their true needs during the process of natural language processing in human-computer interaction to generate decisions. The agents with self-awareness can selectively call knowledge when needed, and make decisions by directly generating actions or changing actions after reflection when knowledge is not needed, reducing the cost and complexity of knowledge injection.

[0010] To achieve the above object of the invention, an embodiment provides a method for enhancing the knowledge of large model agents based on self-awareness, including the following steps:

[0011] Construct a knowledge base for agent planning, and equip the knowledge base with a knowledge selection module to form a knowledge system;

[0012] Design heuristic agent scenario determination criteria, including three scenarios: fast thinking, slow thinking, and knowledge-based thinking;

[0013] According to the scenario judgment criteria, use the constructed knowledge system to construct a self-awareness training dataset for the agent;

[0014] On the constructed self-awareness training dataset, adopt a two-stage training method combining supervised fine-tuning and reinforcement learning to train the knowledge-based self-awareness ability of the large model agent;

[0015] Enable the trained large model agent to make reasoning decisions, and the large model agent can self-determine different scenarios and make different decisions by outputting special marks.

[0016] Preferably, constructing a knowledge base for agent planning includes:

[0017] Use a lightweight large model as the first agent to explore in the environmental task space, and collect three different task trajectories: the successful trajectory where the agent successfully completes the task, the indirectly successful trajectory where the agent fails in the first task but succeeds after reflection, and the failed trajectory where the agent cannot complete the task;

[0018] Design prompt words to let the first agent summarize the successful, reflective, and failed experiences from the three trajectories to form knowledge, limit the number of knowledge to N, and when the number of knowledge exceeds N, the first agent merges the existing knowledge to construct a knowledge base.

[0019] Preferably, the lightweight large model includes GPT-4o, and the knowledge selection module includes using DeepSeek-v3 as the knowledge selection model.

[0020] Preferably, designing heuristic agent scenario determination criteria, including three scenarios: fast thinking, slow thinking, and knowledge-based thinking, includes:

[0021] The situation where the agent is at time t is represented by the historical trajectory h t The correct action required for the next step is a t+1 while the predicted action made by the agent is When the agent makes a wrong prediction, it is allowed to conduct a self-reflection once and generate a predicted action again Then the following judgment is made:

[0022] 1) When , it means that the agent can directly make the correct decision without reflection and without knowledge, and this situation is called fast thinking;

[0023] 2) When and , it means that the agent needs to reflect to make the correct decision, but the modification process still relies on the agent's own ability to make the correct decision without external knowledge, and this situation is called slow thinking;

[0024] 3) When and , it means that the agent cannot rely on its own ability to make the correct decision and needs to introduce knowledge as a guide, and this situation is called knowledge-based thinking.

[0025] Preferably, according to the situation judgment criterion, a self-awareness training data is constructed for the agent by using the constructed knowledge system, including:

[0026] Given the historical action pair (h t , a t+1 ), let the agent make the following decision according to the historical trajectory:

[0027] 1) When the agent makes a correct decision according to the historical trajectory, it conforms to the fast thinking situation. At this time, the output action y = a t+1 is directly the training data;

[0028] 2) When the agent makes a wrong decision according to the historical trajectory and makes a correct decision after rethinking, it conforms to the slow thinking situation. The thinking process is set as ret. At this time, the reflection mark [reflection] is introduced to record the slow thinking situation, and the output action y with the following structure is constructed as the training data:

[0029]

[0030] where <r>and< / r> is a special mark around the reflection thought chain ret;

[0031] 3) When the agent makes a decision error based on the historical trajectory, the historical trajectory is input into the knowledge selection module of the knowledge system. The knowledge selection module selects the most appropriate piece of knowledge know in the knowledge base and returns it to the agent. The agent makes a decision based on the introduced knowledge know, introduces an external knowledge marker [knowledge] to mark the knowledge-based thinking scenario, and constructs the output action y with the following structure as training data:

[0032] y = [[knowledge] <k>know< / k> ,a t+1

[0033] where, <k>and< / k> is a special marker around the knowledge know;

[0034] Construct a self-awareness training data set D based on the output actions in the above three cases self .

[0035] Preferably, on the constructed self-awareness training data set, a two-stage training method combining supervised fine-tuning and reinforcement learning is used to train the knowledge-based self-awareness ability of the large model agent, including:

[0036] The first-stage training: Use a regression loss function to train the large model agent to obtain a reference proxy agent. The reference proxy agent explores in the self-awareness training data set and collects error action outputs. The error action outputs and the correct action outputs in the self-awareness training data set form positive and negative data pairs and form a data set D pair ;

[0037] The second-stage training: Based on the data set D pair Perform reinforcement learning on the large model agent to improve the knowledge-based self-awareness ability. The loss functions used include DPO loss, PPO loss, or GRPO loss;

[0038] Preferably, when performing reinforcement learning on the large model agent based on the data set D pair to improve the knowledge-based self-awareness ability, an SFT loss is also introduced to normalize the output length of the large model agent to stabilize the second-stage training.

[0039] Preferably, enable the trained large model agent to perform inference and decision-making. The large model agent can self-determine different scenarios and make different decisions by outputting special markers, including:

[0040] During the inference and decision-making process of the trained large model agent, if it directly outputs a decision action, it means that the current is a fast-thinking scenario, and the directly output decision action is directly put into the historical trajectory for the next round of decision-making;

[0041] ​If the reflection tag [reflection] is output, it indicates a slow thinking scenario at this time. Let the large model agent further conduct self-reflection, and put the decision-making actions after reflection into the historical trajectory for the next round of decision-making.

[0042] If the knowledge tag [knowledge] is output, it means that knowledge is needed at this time. Call the knowledge selection module to select the most appropriate knowledge from the knowledge base and put it into the context, so that the agent can make decision-making actions based on the knowledge.

[0043] To achieve the above-mentioned invention purpose, the embodiment also provides a knowledge enhancement device for a large model agent based on self-awareness, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned knowledge enhancement method for a large model agent based on self-awareness.

[0044] To achieve the above-mentioned invention purpose, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned knowledge enhancement method for a large model agent based on self-awareness.

[0045] Compared with the prior art, the beneficial effects of the present invention at least include:

[0046] The present invention constructs a scenario determination criterion and constructs a self-training data set based on the scenario determination criterion, so that when the large model agent makes action decisions, it can perform three scenarios: fast thinking, slow thinking, and knowledge-based thinking, and only introduce external knowledge during knowledge-based thinking, reducing the cost and complexity of knowledge injection. Brief Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 is a flowchart of the knowledge enhancement method for a large model agent based on self-awareness provided by the embodiment;

[0049] Figure 2 is a flowchart of self-awareness learning provided by the embodiment. Detailed Embodiments

[0050] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain the present invention and do not limit the protection scope of the present invention.

[0051] The technical concept of the present invention is as follows: A method for knowledge enhancement of a large model intelligent agent based on self-awareness is provided, enabling the intelligent agent based on the large language model to autonomously adjust knowledge utilization. Specifically, the method of the present invention is data-centered, endowing the intelligent agent with knowledge-based self-awareness like humans. Specifically, a heuristic situation judgment criterion is designed to mark special marks on the autonomous exploration trajectory of the intelligent agent to collect training data, and then through a two-stage training process, the intelligent agent can switch between different situations by generating specific special marks, thereby achieving the best planning effect at the lowest cost.

[0052] As Figure 1 shown, a method for knowledge enhancement of a large model intelligent agent based on self-awareness provided by the embodiment includes the following steps:

[0053] S1, construct a knowledge base for intelligent agent planning, and equip the knowledge base with a knowledge selection module to form a knowledge system.

[0054] In the embodiment, an automated knowledge system construction method is adopted. Different from the automatic differentiation knowledge base construction method of WKM, the present invention does not require expert trajectories. Specifically, a lightweight large model such as the GPT-4o model is selected as the first intelligent agent to explore in the environmental task space, and three different task completion trajectories are collected: a successful trajectory (i.e., the intelligent agent successfully completes the task), an indirectly successful trajectory (the intelligent agent originally fails, but solves the failure problem through self-reflection and achieves success), and a failed trajectory (the task trajectory that the intelligent agent cannot complete). Then, prompt words are designed to let the first intelligent agent summarize the successful, reflective and failed experiences from the three trajectories to form knowledge. To reduce the redundancy of the knowledge base, the number of knowledge is limited to N. When the number of knowledge exceeds N, let the first intelligent agent merge the existing knowledge, thereby constructing the final knowledge base. In addition, a knowledge selection module is also equipped for the knowledge base, and DeepSeek-v3 is used as the knowledge selection model to select the knowledge required by the intelligent agent in the knowledge base according to the environment where the intelligent agent is located.

[0055] S2, design a heuristic intelligent agent situation determination criterion, including three situations: fast thinking, slow thinking and knowledge-based thinking.

[0056] In the embodiment, in order to classify the situation where the intelligent agent is located to judge whether the intelligent agent needs knowledge in different situations, a heuristic situation judgment criterion is designed, specifically including:

[0057] The situation where the agent is at time t is represented by the historical trajectory h t The correct action required for the next step is a t+1 , and the predicted action made by the agent is When the agent makes a wrong prediction, it is allowed to conduct a self-reflection once and generate a predicted action again Then the following judgment is made:

[0058] 1) When , it indicates that the agent can directly make a correct decision without reflection and without knowledge. This situation is called fast thinking;

[0059] 2) When and , it indicates that the agent needs to make a correct decision through reflection, but the modification process still relies on the agent's own ability to make a correct decision without external knowledge. This situation is called slow thinking;

[0060] 3) When and , it indicates that the agent cannot make a correct decision relying on its own ability and needs to introduce knowledge as a guide. This situation is called knowledge-based thinking.

[0061] This heuristic agent situation determination criterion will guide the subsequent construction of training data, enabling the agent to self-select whether to introduce knowledge according to its own situation.

[0062] S3. According to the situation judgment criterion, use the constructed knowledge system to construct a self-awareness training data set for the agent.

[0063] In the embodiment, a data-driven method is designed to enable the agent to have self-awareness, so as to selectively introduce knowledge. Specifically, given the historical action pair (h t , a t+1 ), let the agent make the following decision according to the historical trajectory:

[0064] 1) When the agent makes a correct decision according to the historical trajectory, it obviously conforms to the fast thinking situation. At this time, the output action y = a t+1 is directly used as the training data;

[0065] 2) When the agent makes a wrong decision according to the historical trajectory and makes a correct decision after rethinking, it conforms to the slow thinking situation. Set the thinking process as ret. At this time, introduce the reflection mark [reflection] to record the slow thinking situation, and construct the output action y with the following structure as the training data:

[0066]

[0067] where <r>and< / r>is a special marker around the reflection thinking chain ret;

[0068] 3) When the agent makes a decision error based on the historical trajectory, the historical trajectory is input into the knowledge selection module of the knowledge system. The knowledge selection module selects the most appropriate piece of knowledge know from the knowledge base and returns it to the agent. The agent makes a decision based on the introduced knowledge know, introduces an external knowledge marker [knowledge] to mark the knowledge-based thinking scenario, and constructs an output action y with the following structure as training data:

[0069] y = [[knowledge] <k>know< / k> ,a t+1

[0070] where, <k>and< / k> is a special marker around the knowledge know;

[0071] Construct a self-awareness training dataset D based on the output actions in the above three cases self , and this training dataset D self customizes different outputs according to the agent's own capabilities, so that the agent can learn to recognize different scenarios according to its own capabilities.

[0072] S4. On the constructed self-awareness training dataset, use a two-stage training method combining supervised fine-tuning and reinforcement learning to train the knowledge-based self-awareness ability of the large model agent.

[0073] In the embodiment, a two-stage training method is adopted to enable the agent with the decision-making strategy of π θ to have knowledge-based self-awareness. First, in the first-stage training, an autoregressive loss function L SFT is used to train a reference agent, whose decision-making strategy is π ref . Let the reference agent explore in the self-awareness training dataset and collect the error action output y p . The error action output and the correct action output in the self-awareness training dataset form a positive and negative data pair and form a dataset D pair ;

[0074]

[0075] where, π θ (y|h t ) represents that based on the historical trajectory h t the agent outputs the decision action y according to its own decision-making strategy π θ ;

[0076] Then, in the second-stage training, based on the dataset D pair ​Reinforcement learning is performed on the large model agent to improve the knowledge-based self-awareness ability. Among them, the loss functions adopted include DPO loss, PPO loss, or GRPO loss; among them, the DPO loss L DPO is expressed as:

[0077]

[0078] where σ() is the sigmoid function, and β is a hyperparameter that controls the deviation between the control and the decision-making policy π ref , and π θ (y|h t ) represents the probability distribution of the agent outputting the decision-making action y based on the historical trajectory h t according to its own decision-making policy π θ , and π ref (y|h t ) represents the probability distribution of the agent outputting the decision-making action y based on the historical trajectory h t according to the decision-making policy π ref , and π θ (y p |h t ) represents the probability distribution of the agent outputting the decision-making action y t according to its own decision-making policy π θ based on the historical trajectory h p , and π ref (y p |h t ) represents the probability distribution of the agent outputting the decision-making action y t according to its own decision-making policy π ref based on the historical trajectory h p .

[0079] In the second-stage training, since the space of correct actions is very narrow relative to the natural language space, according to the existing work, the normalized SFT loss (i.e., L NLL ) is introduced into this stage again and normalized by the output length to stabilize the training in the second stage:

[0080]

[0081] Then the total training loss in the second stage is:

[0082] L RPO = L DPO + αL NLL

[0083] where α is a hyperparameter used to control the ratio between the two losses.

[0084] S5. Enable the trained large model agent to make inference decisions. The large model agent can self-determine different situations and make different decisions by outputting special tags.

[0085] In the embodiment, during the inference decision-making process of the trained large model agent, if it directly outputs a decision action, it indicates that the current is a fast-thinking scenario, and the directly output decision action is put into the historical trajectory for the next round of decision-making; if it outputs a reflection tag [reflection], it means that this is a slow-thinking scenario, and the large model agent is allowed to further self-reflect, and the decision action after reflection is put into the historical trajectory for the next round of decision-making; if it outputs a knowledge tag [knowledge], it means that knowledge is needed at this time, and the knowledge selection module is called to select the most appropriate knowledge from the knowledge base and put it into the context, so that the agent makes a decision action based on the knowledge.

[0086] Through the above method, the large model agent can have self-awareness, so as to selectively call knowledge, and therefore no special requirements are made for the source of knowledge. In addition to the lightweight knowledge system construction scheme proposed in this solution, feasible solutions also include manual writing, large model self-synthesis, symbolic knowledge, parametric knowledge, etc., as long as it is reasonable, and no other special requirements are made.

[0087] The embodiment also tests the solution of the present invention on two datasets, ALFWorld and WebShop, and two models, Llama-8B and Gemma-2B. That is, the large model agent of the present invention uses two models, Llama-8B and Gemma-2B, to test the knowledge reasoning application of the present invention in the process of human-computer interaction dialogue and the shopping field. And it is compared with many methods including KnowAgent and WKM. The experimental results are shown in Tables 1 and 2 below:

[0088] Table 1 Experimental Results on ALFWorld Dataset

[0089]

[0090] Table 2 Experimental Results on Web Shop Dataset

[0091]

[0092] In the two tables, KnowSelf is the method of the present invention, using the average reward as the evaluation index. The best results are in bold, and the symbol represents the prompt-based baseline, and the symbol represents the baseline based on fine-tuning training. The knowledge enhancement ratio represents the ratio of actions enhanced by knowledge.

[0093] It can be found from the experimental results that the proposed solution of the present invention exceeds all knowledge enhancement (Know% = 100%) and non-knowledge enhancement (Know% = 0%) methods at an extremely low knowledge introduction rate (Know%), achieving SOTA (state-of-the-art, that is, the best effect relative to all baselines) in terms of performance. Moreover, the performance of the small Llama-8B model even approaches that of the powerful GPT-4 on Reflexion (a powerful baseline). The above experimental results indicate that the method of the present invention can greatly reduce the proportion of knowledge introduction while improving performance, selectively introducing knowledge using the self-awareness of the agent, and fully balancing the three scenarios of fast thinking, slow thinking, and knowledge thinking according to its own capabilities. The experimental phenomena also fully demonstrate that mindlessly introducing knowledge in existing knowledge enhancement methods (such as KnowAgent, WKM, etc.) is not the optimal approach.

[0094] Although the method of the present invention focuses on the dynamic knowledge selection of large model agents in the field of agent planning, it can also be extended and applied to the following fields:

[0095] 1. Fast and slow thinking fields

[0096] With the emergence of inference models such as OpenAI O1 and DeepSeek-R1, large models have entered the era of slow thinking. However, this has also brought some problems. Excessively slow thinking will lead to a rapid increase in the inference cost of large models, which is obviously unnecessary in some scenarios. Therefore, enabling large models to have the self-awareness to decide when to perform fast thinking and when to perform slow thinking is an issue that needs to be solved. The present invention provides an idea for this problem. By removing the knowledge module, the present invention can be simplified into an architecture for fast and slow thinking selection, and an agent that can make a choice between fast and slow thinking according to its own capabilities can be trained.

[0097] 2. Knowledge boundaries of large models

[0098] Determining the knowledge boundaries of large models is an important issue for alleviating large model hallucinations and factual errors. The ability of large models to selectively output answers or refuse to answer according to their own capabilities is the key to realizing a stable and reliable large model. The solution of the present invention to judge and train the knowledge-based self-awareness of large model agents is extremely similar to judging and training the factual knowledge self-awareness of large models. The method adaptation can be completed by facing factual Q&A data during the data construction stage.

[0099] Based on the same inventive concept, the embodiment also provides a large model agent knowledge enhancement device based on self-awareness, including a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, it is used to implement the above-mentioned large model agent knowledge enhancement method based on self-awareness, specifically including the following steps:

[0100] S1. Construct a knowledge base for agent planning and equip the knowledge base with a knowledge selection module to form a knowledge system;

[0101] S2. Design heuristic agent scenario determination criteria, including three scenarios: fast thinking, slow thinking, and knowledge-based thinking;

[0102] S3. According to the scenario judgment criteria, use the constructed knowledge system to build a self-awareness training dataset for the agent;

[0103] S4. On the constructed self-awareness training dataset, adopt a two-stage training method combining supervised fine-tuning and reinforcement learning to train the knowledge-based self-awareness ability of the large model agent;

[0104] S5. Make the trained large model agent perform reasoning and decision-making. The large model agent can self-determine different scenarios and make different decisions by outputting special marks.

[0105] The computing device provided by the embodiment, at the hardware level, in addition to including a processor and a memory, also includes other necessary hardware for other services such as an internal bus, a network interface, and a memory. The memory is a non-volatile memory. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the self-awareness-based large model agent knowledge enhancement method described in S1-S5 above. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.

[0106] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the self-awareness-based large model agent knowledge enhancement method described above, specifically including the following steps:

[0107] S1. Construct a knowledge base for agent planning and equip the knowledge base with a knowledge selection module to form a knowledge system;

[0108] S2. Design heuristic agent scenario determination criteria, including three scenarios: fast thinking, slow thinking, and knowledge-based thinking;

[0109] S3. According to the scenario judgment criteria, use the constructed knowledge system to build a self-awareness training dataset for the agent;

[0110] S4. On the constructed self-awareness training dataset, adopt a two-stage training method combining supervised fine-tuning and reinforcement learning to train the knowledge-based self-awareness ability of the large model agent;

[0111] S5 enables the trained large model agent to make inference decisions. The large model agent can self-determine different situations and make different decisions by outputting special markers.

[0112] In the embodiments, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.

[0113] The specific embodiments described above have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for enhancing knowledge of a large model agent based on self-awareness, characterized in that: The following steps are involved: Build a knowledge base for agent planning and equip the knowledge base with a knowledge selection module to form a knowledge system; Design heuristic agent scenario judgment criteria, including three scenarios: fast thinking, slow thinking, and knowledge-based thinking; According to the situational judgment criteria, the constructed knowledge system is used to construct a self-awareness training dataset for the intelligent agent; On the constructed self-awareness training dataset, a two-stage training method combining supervised fine-tuning and reinforcement learning is used to train the knowledge-based self-awareness ability of the large model agent; The trained large model agent is enabled to make reasoning decisions. The large model agent can self-judge different situations and make different decisions by outputting special tags.

2. The method for enhancing knowledge of a large model agent based on self-awareness according to claim 1 is characterized in that: Build a knowledge base for agent planning, including: Use the lightweight large model as the first agent to explore the environment task space and collect three different task trajectories: a successful trajectory where the agent successfully completes the task, an indirect successful trajectory where the agent fails the task for the first time but succeeds after reflection, and a failed trajectory where the agent cannot complete the task; The prompt words are designed to let the first agent summarize the experiences of success, reflection and failure from the three trajectories to form knowledge, and the amount of knowledge is limited to N. When the amount of knowledge exceeds N, the first agent merges the existing knowledge to construct a knowledge base.

3. The method for enhancing knowledge of a large model agent based on self-awareness according to claim 2 is characterized in that: The lightweight large model includes GPT-4o, and the knowledge selection module includes DeepSeek-v3 as the knowledge selection model.

4. The method for enhancing knowledge of a large model agent based on self-awareness according to claim 1 is characterized in that: Design heuristic agent scenario judgment criteria, including three scenarios: fast thinking, slow thinking and knowledge-based thinking, including: Set the situation of the agent at time t to use the historical trajectory h t Indicates that the correct action required for the next step is a t+1 , and the predicted action made by the agent is When the agent makes a wrong prediction, it is allowed to self-reflect and generate a predicted action again Then make the following judgment: 1) When When , it means that the agent can make the correct decision directly without reflection or knowledge. This situation is called fast thinking. 2) When and When , it means that the agent needs to reflect to make the correct decision, but the modification process still relies on the agent's own ability to make the correct decision and does not require external knowledge. This situation is called slow thinking; 3) When and When , it means that the intelligent agent cannot make correct decisions based on its own ability and needs to introduce knowledge as a guide. This situation is called knowledge-based thinking.

5. The method for enhancing knowledge of a large model agent based on self-awareness according to claim 1 is characterized in that: According to the situational judgment criteria, the constructed knowledge system is used to construct self-awareness training data for the agent, including: Given a historical action pair (h t ,a t+1 ), let the agent make the following decisions based on the historical trajectory: 1) When the agent makes the correct decision based on the historical trajectory, it meets the fast thinking scenario, and the output action y = a t+1 Directly for training data; 2) When the agent makes a wrong decision based on the historical trajectory and makes a correct decision after rethinking, it meets the slow thinking scenario and is set to ret in the reflective thinking chain. At this time, the reflection tag [reflection] is introduced to record the slow thinking scenario and construct the output action y of the following structure as training data: in <r> and< / r> It is a special mark around the reflective thinking chain ret; 3) When the agent makes a wrong decision based on the historical trajectory, the historical trajectory is input into the knowledge selection module of the knowledge system. The knowledge selection module selects the most appropriate knowledge know in the knowledge base and returns it to the agent. The agent makes a decision based on the introduced knowledge know, introduces an external knowledge tag [knowledge] to mark the knowledge-based thinking scenario, and constructs the output action y of the following structure as training data: y=[[knowledge] <k> know< / k> ,a t+1 ] in, <k> and< / k> It is a special mark around the knowledge know; Based on the output actions of the above three situations, a self-awareness training dataset D is constructed. self .

6. The method for enhancing knowledge of a large model agent based on self-awareness according to claim 1 is characterized in that: On the constructed self-awareness training dataset, a two-stage training method combining supervised fine-tuning and reinforcement learning is used to train the knowledge-based self-awareness ability of the large model agent, including: One-stage training: The large model agent is trained using a regression loss function to obtain a reference agent. The reference agent explores the self-awareness training dataset and collects the wrong action outputs. The wrong action outputs and the correct action outputs in the self-awareness training dataset form positive and negative data pairs and form a dataset D pair ; Two-stage training: based on data set D pair Reinforcement learning is performed on large model agents to improve their knowledge-based self-awareness capabilities. The loss functions used include DPO loss, PPO loss, or GRPO loss;.

7. The method for enhancing knowledge of a large model agent based on self-awareness according to claim 6 is characterized in that: Based on the dataset D pair When performing reinforcement learning on large model agents to improve their knowledge-based self-awareness capabilities, SFT loss is also introduced to normalize the output length of the large model agent to stabilize the two-stage training.

8. The method for enhancing knowledge of a large model agent based on self-awareness according to claim 1 is characterized in that: The trained big model agent can make inference decisions. The big model agent can self-judge different situations and make different decisions by outputting special tags, including: If the trained large model agent directly outputs the decision action during the reasoning and decision-making process, it means that the current situation is fast thinking, and the output decision action is directly put into the historical track to make the next round of decision; If the reflection mark [reflection] is output, it means that this is a slow thinking scenario, allowing the large model agent to further reflect on itself and put the decision-making actions after reflection into the historical track for the next round of decision-making; If the knowledge tag [knowledge] is output, it means that knowledge is needed at this time. The knowledge selection module is called to select the most appropriate knowledge from the knowledge base and put it into the context, so that the intelligent agent can make decision actions based on the knowledge.

9. A large model intelligent agent knowledge enhancement device based on self-awareness, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the self-awareness-based large model agent knowledge enhancement method described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the method for enhancing knowledge of a large model intelligent agent based on self-awareness as described in any one of claims 1-8 is implemented.

Citation Information

Cited By

  • Intelligent agent optimization method and device, equipment, storage medium and product

    CN120996185A

  • Construction method of breast cancer patient nutrition tutoring large model

    CN121306427A

  • Memory management method and device of intelligent agent

    CN122242560A

  • Memory management method and apparatus for an agent

    CN122242560B