Adaptive learning agent based on large language model
Through adaptive learning agents, the situational memory knowledge base is constructed, which solves the shortcomings of large-model agents in task and environment adaptation, and realizes efficient and low-cost application in the production environment.
Patent Information
- Application Number
- CN202510448933.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
AI Technical Summary
Existing large-model agents have insufficient performance in task adaptation and environmental adaptation, mainly due to the lack of episodic memory modeling and the reliance on expensive manual annotation trajectory data, which cannot be efficiently generalized to various tasks.
Adaptive methods are used to build an episodic memory knowledge base. Through the agent's execution of commands multiple times, the highest performance trajectory is sampled, the episodic memory knowledge base is generated, and decisions are made based on the current situation.
Improves the performance capabilities of agents in different tasks and environments, reduces costs and improves efficiency, and is suitable for production environments.
Smart Images

Figure CN120354940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an adaptive learning agent based on a large language model. Background Art
[0002] In order to achieve continuous learning and rapid task adaptation, researchers have explored methods to promote the learning of agents by using memory mechanisms. However, the existing work on large model agents mainly focuses on the modeling of semantic memory, while paying little attention to the modeling of episodic memory, so that past experiences and events are not utilized, which weakens the performance of large model agents in specific tasks. In addition, the existing work mainly collects trajectory data based on rules or manual annotation, which is very expensive and cannot be generalized to every task, and does not utilize the information of the model and the task itself.
[0003] Therefore, it is necessary to provide an adaptive learning agent based on a large language model, so that the agent can quickly adapt to different tasks and environments, and can be applied to the production environment at lower cost and higher efficiency. Summary of the Invention
[0004] The purpose of the present invention is to provide an adaptive learning agent based on a large language model, so that the agent can quickly adapt to different tasks and environments, and can be applied to the production environment at lower cost and higher efficiency.
[0005] In order to solve the problems existing in the prior art, the present invention provides an adaptive learning agent based on a large language model, including:
[0006] An episodic memory generation unit configured to construct an episodic memory knowledge base by an adaptive method, and the episodic memory knowledge base is a set of states combined with behaviors;
[0007] An agent reasoning unit configured to, when the agent faces a new situation, dynamically retrieve relevant experiences and decision-making trajectories from the episodic memory knowledge base according to the specific state of the current situation; make a reasonable decision based on the historical experiences and decision-making trajectories and in combination with the current situation information.
[0008] Optionally, in the adaptive learning agent based on the large language model,
[0009] The episodic memory knowledge base is obtained based on the trajectories of the large model agent.
[0010] Optionally, in the adaptive learning agent based on the large language model, the construction method of the episodic memory knowledge base is as follows:
[0011] The agent executes each command multiple times;
[0012] Perform result sampling, select the trajectory with the highest performance score from the results as the elements for constructing episodic memory;
[0013] Construct an episodic memory knowledge base based on the elements for constructing episodic memory.
[0014] Compared with the prior art, the present invention has the following advantages:
[0015] (1) The present invention enables the agent to quickly adapt to different tasks and environments, enabling it to be applied to the production environment at lower cost and higher efficiency.
[0016] (2) The present invention can enhance the performance of the large model agent in specific tasks and requires less manual input. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flowchart of the reasoning of the agent provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] The following will describe the specific embodiments of the present invention in more detail with reference to the schematic diagrams. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0019] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application.
[0020] To achieve continuous learning and rapid task adaptation, researchers have explored methods of using memory mechanisms to promote agent learning. However, the existing work of large model agents mainly focuses on the modeling of semantic memory, while paying little attention to the modeling of episodic memory, so that past experiences and events are not utilized, which weakens the performance of large model agents in specific tasks. In addition, the existing work mainly collects trajectory data based on rules or manual annotation, which is very expensive and cannot be generalized to every task, and does not utilize the information of the model and the task itself.
[0021] To solve the problems existing in the prior art, the present invention provides an adaptive learning agent based on a large language model, including:
[0022] The episodic memory generation unit is configured to construct an episodic memory knowledge base using an adaptive method. The episodic memory knowledge base is a collection of states combined with behaviors. Further, the episodic memory knowledge base is obtained based on the trajectories of the large model agent.
[0023] Specifically, the construction method of the episodic memory knowledge base is as follows: The agent executes each command multiple times; performs result sampling, selects the trajectory with the highest performance score from the results as the element for constructing the episodic memory; constructs the episodic memory knowledge base based on the elements for constructing the episodic memory.
[0024] The adaptive method provided by the present invention is more convenient compared to the existing trajectory generation methods based on rules or manual annotation, and can be generalized to each task. After multiple samplings and the completion of training, a high-quality episodic memory knowledge base is obtained, which can provide high-quality reference when the model makes inferences and decisions, helping the large model agent make more reasonable and efficient decisions.
[0025] The agent reasoning unit is configured to, when the agent faces a new situation, dynamically retrieve the relevant experiences and decision-making trajectories from the episodic memory knowledge base according to the specific state of the current situation; make reasonable and efficient decisions based on the historical experiences and decision-making trajectories in combination with the current situation information. Since the episodic memory knowledge base is constructed based on the trajectories of the agent's own best performance, it can provide more accurate and personalized decision support for the agent compared to the traditional methods based on fixed rules or external knowledge bases.
[0026] In one embodiment, the reasoning process of the large model agent is as Figure 1 described. The large model agent completes specific tasks in the episodic memory generation unit. The agent executes each task multiple times; performs result sampling, selects the trajectory with the highest performance score from the results as the element for constructing the episodic memory; constructs the episodic memory knowledge base based on the elements for constructing the episodic memory. During reasoning, the large model agent refers to the episodic memory knowledge base and makes reasonable and efficient decisions in combination with the current situation information.
[0027] The present invention uses an adaptive method to enable the large model agent to learn from the trajectory information it generates, so as to enhance the performance ability of the large model agent in different environments.
[0028] In summary, compared with the prior art, the present invention has the following advantages:
[0029] (1) The present invention enables the agent to quickly adapt to different tasks and environments, enabling it to be applied to the production environment at lower cost and higher efficiency.
[0030] (2) The present invention can enhance the performance of large model agents in specific tasks and requires less manual input.
[0031] The above are only the preferred embodiments of the present invention and do not impose any restrictive effects on the present invention. Any person skilled in the art, within the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification and other changes to the disclosed technical solution and technical content of the present invention, which are all within the content of the technical solution of the present invention and still fall within the protection scope of the present invention.
Claims
1. An adaptive learning intelligent agent based on a large language model, characterized in that, Including: An episodic memory generation unit configured to construct an episodic memory knowledge base using an adaptive method, where the episodic memory knowledge base is a collection of states combined with behaviors; An agent reasoning unit configured to, when the agent faces a new situation, dynamically retrieve relevant experiences and decision-making trajectories from the episodic memory knowledge base according to the specific state of the current situation; and make a reasonable decision based on the historical experiences and decision-making trajectories in combination with the current situation information.
2. The adaptive learning agent based on a large language model according to claim 1, wherein The episodic memory knowledge base is obtained based on the trajectories of the large model agent.
3. The adaptive learning agent based on a large language model according to claim 2, characterized in that, The construction method of the episodic memory knowledge base is as follows: The agent executes each command multiple times; Perform result sampling, select the trajectory with the highest performance score from the results as the element for constructing the episodic memory; Construct an episodic memory knowledge base based on the elements for constructing the episodic memory.