Memory management agent based on large language model
Through the combination of memory bank and reinforcement learning training unit, the agent continues to learn in interaction with the environment, overcomes the limitations of context length, and realizes efficient memory management, solves the problems of high computing complexity and relying on manual rules in the existing technology, and realizes independent learning and low-cost applications.
Patent Information
- Application Number
- CN202510448875.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
AI Technical Summary
Existing language agents are limited by fixed context lengths when processing long context tasks, and cannot effectively process ultra-long sequences or dynamically changing environmental information, and the computational complexity increases linearly with the growth of context, and rely on external storage and manual rules for memory management.
Using memory bank and reinforcement learning training units, including memory writing, retrieval units and update units, automatically determines storage information through large language models, and optimizes memory management with forgetting mechanisms and reward signals. The agent continues to learn in interaction with the environment and overcomes context length limitations.
It realizes efficient application of agents in different environments and tasks, reduces computational complexity, and optimizes memory management through independent learning and reduces labor costs.
Smart Images

Figure CN120354885A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a memory management agent based on a large language model. Background Art
[0002] When existing language agents process long-context tasks, they are usually limited by a fixed context length, resulting in an inability to effectively process ultra-long sequences or dynamically changing environmental information, and the computational complexity increases linearly with the growth of the context. In addition, existing methods usually rely on external storage and manually set rules for memory management, and the agent cannot autonomously learn how to optimize the access and use of memory.
[0003] Therefore, it is necessary to provide a memory management agent based on a large language model, enabling the agent to continuously learn during the interaction with the environment, overcome the context length limitation through a loop mechanism, and achieve efficient application with a low computational complexity. Summary of the Invention
[0004] The purpose of the present invention is to provide a memory management agent based on a large language model, enabling the agent to continuously learn during the interaction with the environment, overcome the context length limitation through a loop mechanism, and achieve efficient application with a low computational complexity.
[0005] To solve the problems existing in the prior art, the present invention provides a memory management agent based on a large language model, including:
[0006] A memory bank, which includes a memory writing unit, a memory retrieval unit, and a memory update unit;
[0007] A reinforcement learning training unit configured to train the agent using reinforcement learning methods so that it learns to manage the memory bank and use the memory to complete tasks.
[0008] Optionally, in the memory management agent based on a large language model,
[0009] The memory writing unit is used to, during the process of the agent executing tasks, the large language model automatically determines the information to be stored in the memory bank according to the current environment, interaction content, and reasoning results;
[0010] The memory retrieval unit is used to, when the agent faces a new task, the agent retrieves relevant information from the memory bank according to the current environment to assist in decision-making;
[0011] The memory update unit is used to continuously modify outdated or incorrect memories, and also delete unimportant memories through a forgetting mechanism. Unimportant memories are memories that have not been used for more than a preset time.
[0012] Optionally, in the large language model-based memory management agent, during the training of the agent, the agent obtains a reward signal through interaction with the environment, and the reward signal is based on the task completion quality and the efficiency of memory management.
[0013] Compared with the prior art, the present invention has the following advantages:
[0014] (1) By designing an adaptive memory management mechanism, the agent can continuously learn during the interaction with the environment, and overcome the context length limitation through a loop mechanism, achieving efficient application with a low computational complexity.
[0015] (2) Through reinforcement learning training, the agent can obtain efficient memory management capabilities, achieve continuous learning, and quickly adapt to different environments and tasks.
[0016] (3) The present invention does not rely on a large amount of manually labeled data, and only needs to deploy an interaction environment and design rewards, so the labor cost is low. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the architecture of the agent provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The following will describe the specific embodiments of the present invention in more detail with reference to the schematic diagrams. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0019] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application.
[0020] Existing language agents are usually limited by a fixed context length when dealing with long-context tasks, resulting in the inability to effectively process ultra-long sequences or dynamically changing environmental information, and the computational complexity increases linearly with the growth of the context. In addition, existing methods usually rely on external storage and manually set rules for memory management, and the agent cannot autonomously learn how to optimize the access and use of memory.
[0021] To solve the problems existing in the prior art, the present invention provides a memory management intelligent agent based on a large language model, including:
[0022] A memory bank, which includes a memory writing unit, a memory retrieval unit, and a memory update unit;
[0023] The memory writing unit is used for the large language model to automatically determine the information to be stored in the memory bank according to the current environment, interaction content, and reasoning result during the process of the intelligent agent executing tasks;
[0024] The memory retrieval unit is used for the intelligent agent to retrieve relevant information from the memory bank according to the current environment to assist in decision-making when the intelligent agent faces a new task;
[0025] The memory update unit is used to continuously modify outdated or incorrect memories, and also delete unimportant memories through a forgetting mechanism. Unimportant memories are memories that have not been used for more than a preset time. For example, memories that have not been run for more than 6 months are unimportant memories.
[0026] A reinforcement learning training unit, configured to train the intelligent agent using reinforcement learning methods so that it learns to manage the memory bank and use memories to complete tasks.
[0027] Further, during the process of training the intelligent agent, the intelligent agent obtains a reward signal through interaction with the environment. The reward signal is based on the task completion quality and the efficiency of memory management (such as retrieval accuracy and storage space utilization rate). Through reinforcement learning, the intelligent agent gradually optimizes its memory management strategy, such as which information needs to be stored, when to update the memory bank, and how to extract the most relevant content from the memory bank. After training, the intelligent agent can flexibly call memory information in different task scenarios, significantly improving the decision-making efficiency and task performance.
[0028] In one embodiment, the architecture of the intelligent agent is as Figure 1 shown, including working memory and long-term memory. The read and write functions are completed between the working memory and the long-term memory, and the intelligent agent continuously cycles and optimizes, so as to combine external observations to make accurate output actions.
[0029] The memory management intelligent agent provided by the present invention uses a cyclic mechanism to manage the context, making it not restricted by the traditional context length. At the same time, by introducing a dynamic read and write memory bank, the intelligent agent can continuously learn during the interaction with the environment. In addition, the memory management ability of the intelligent agent is trained through reinforcement learning, enabling it to reasonably update the memory bank to achieve the purpose of efficient and continuous learning.
[0030] In summary, compared with the prior art, the present invention has the following advantages:
[0031] (1) The present invention designs an adaptive memory management mechanism, enabling the agent to continuously learn during the interaction with the environment and overcome the context length limitation through a loop mechanism, achieving efficient application with relatively low computational complexity.
[0032] (2) Through reinforcement learning training, the agent can acquire efficient memory management capabilities, achieve continuous learning, and quickly adapt to different environments and tasks.
[0033] (3) The present invention does not rely on a large amount of manually labeled data, only requires the deployment of an interactive environment and the design of rewards, so the labor cost is relatively low.
[0034] The above are only the preferred embodiments of the present invention and do not impose any limitation on the present invention. Any person skilled in the art within the technical field, without departing from the technical solution of the present invention, making any form of equivalent replacement or modification and other changes to the technical solution and technical content disclosed by the present invention, all belong to the content that has not departed from the technical solution of the present invention and still falls within the protection scope of the present invention.
Claims
1. A memory management intelligent agent based on large language models, characterized in that, Comprising: A memory bank, which includes a memory writing unit, a memory retrieval unit, and a memory update unit; A reinforcement learning training unit configured to train an agent using reinforcement learning methods so that it learns to manage the memory bank and complete tasks using the memory.
2. The memory management agent based on a large language model according to claim 1, wherein The memory writing unit is used for, during the process of the agent executing a task, the large language model automatically determines the information to be stored in the memory bank according to the current environment, interaction content, and reasoning results; The memory retrieval unit is used for when the agent faces a new task, the agent retrieves relevant information from the memory bank according to the current environment to assist in decision-making; The memory update unit is used for continuously modifying outdated or incorrect memories, and also deleting unimportant memories through a forgetting mechanism, where unimportant memories are memories that have not been used for more than a preset time.
3. The memory management intelligent agent based on a large language model according to claim 1, characterized in that, During the process of training the agent, the agent obtains a reward signal through interaction with the environment, and the reward signal is based on the task completion quality and the efficiency of memory management.
Citation Information
Cited By
Large model memory management method and device and medium
CN121050663A