Large language model long-term memory method based on human cognitive inspiration
Through the long-term memory method of large language model inspired by human cognition, using sliding window and event cutting theory, the problem of large language model lacking dynamic memory and contextualization is solved, dynamic memory and contextualization characteristics are realized, and knowledge management and content generation efficiency in interaction is improved.
Patent Information
- Application Number
- CN202510684502.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-15
AI Technical Summary
The existing large language models lack dynamic memory and contextualization characteristics, making it difficult to continuously accumulate and update knowledge in multiple interactions.
The long-term memory method of large language model inspired by human cognition is adopted, and the semantic similarity distribution and local information entropy of historical interactive text vectors are calculated through sliding windows, event boundaries are determined, memory partitions are managed as real-time, event and long-term memory, and related historical events are retrieved using contextual memory in the interaction.
The dynamic memory and contextualization characteristics of the large language model are realized, the knowledge accumulation and update efficiency in multiple interactions is improved, and coherent and accurate content is generated.
Smart Images

Figure CN120493992A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of large language models, and specifically relates to a long-term memory method for large language models based on human cognitive inspiration. Background Art
[0002] With the widespread adoption of the Transformer architecture, large language models have rapidly become a core research direction in the field of artificial intelligence. Since the explosive growth of technological breakthroughs and applications based on large language models, a variety of advanced language models have emerged, significantly driving the leap forward in natural language processing technology. These models, pre-trained on massive amounts of text data, have demonstrated exceptional language understanding and generation capabilities, enabling them to handle increasingly complex and diverse tasks. Despite their success in many fields, existing large language models lack long-term memory and are unable to continuously accumulate, update, and apply knowledge across multiple interactions.
[0003] Research on long-term memory mechanisms in large language models is still in its early stages, and a unified theoretical framework has yet to emerge. Mainstream implementations fall into two categories: parametric memory and storage memory. Parametric memory, through fine-tuning or knowledge editing, directly embeds historical interaction records or domain expertise into model parameters. This approach internalizes memory, eliminating the need for the model to rely on external databases when accessing memory. However, due to the high computational cost associated with adjusting model parameters, memory writes are inefficient. In contrast, storage memory relies on external storage tools, storing interaction records in a repository and often combining retrieval-augmented generation techniques to enable dynamic access. Storage memory can be further categorized into text-based and token-based based on the type of stored data. Text-based memory typically stores plain text vectors or triple structures similar to knowledge graphs and is suitable for storing semantically oriented information. Token-based memory stores key-value vectors from specific Transformer layers and can capture fine-grained representations within the model.
[0004] Parametric memory and storage memory each have their own advantages and disadvantages. Parametric memory does not require an external database when reading memory, but the computational cost is very high when writing memory, making it difficult to adapt to the needs of dynamic or frequent memory updates. Storage memory is convenient for writing memory, but performs poorly in memory reading efficiency, especially in scenarios where frequent and rapid access to historical information is required. Storage memory requires the design of efficient storage modes and retrieval algorithms to ensure the accuracy and speed of memory calls. In addition, these existing methods lack human-likeness and find it difficult to effectively simulate the dynamic and contextual characteristics of human memory. In other words, existing long-term memory methods for large language models lack dynamic memory and contextualization.
[0005] Therefore, how to implement a long-term memory method for large language models with dynamic memory and contextual characteristics is an urgent problem to be solved in this field. Summary of the Invention
[0006] The present invention addresses the shortcomings of the prior art by providing a long-term memory method for large language models based on human cognition. This method incorporates the event segmentation theory, hierarchical memory architecture, and natural forgetting curves from human cognitive psychology into the long-term memory mechanism of large language models. This method, based on human cognition, calculates the semantic similarity distribution of historical interaction text vectors using a sliding window and calculates the local information entropy of the window. Event boundaries are determined based on the difference in local information entropy values between adjacent windows. Each discrete event is used as the model's memory, and memory partitions are managed. When the event memory partition is full, memory events with low retention rates are transferred to the long-term memory area. When the content of the interaction involves historical information, a contextualized memory retrieval method is used to retrieve relevant historical events from the event memory partition. The large language model generates content for the user based on the current context and the retrieved relevant historical events. This method utilizes memory encoding and storage based on event segmentation theory, multi-level dynamic memory management, and event-level contextualized memory retrieval, thereby addressing the lack of dynamic memory and contextualization in existing long-term memory methods for large language models.
[0007] In order to achieve the above objectives, the present invention adopts the following technical solutions:
[0008] The present invention proposes a long-term memory method for a large language model based on human cognition inspiration, which is characterized by comprising the steps of:
[0009] Obtain historical interaction texts between users and a large language model from a dataset, and convert the historical interaction texts into text vectors;
[0010] The semantic similarity distribution of text vectors within the window is calculated based on the sliding window mechanism, and the local information entropy of the window is calculated. For adjacent sliding windows, the difference in their local information entropy values is calculated. If the entropy difference between adjacent windows exceeds a predefined threshold, it is marked as an event boundary.
[0011] Based on all event boundaries, discrete events are formed, and a vector database is used to store independent events as the memory of the large language model;
[0012] Partitioning the memory of the large language model into immediate memory, event memory, and long-term memory;
[0013] When the event memory partition is full, the memory events with low retention rate are compressed and transferred to the long-term memory area;
[0014] When the content of the user's interaction with the large language model involves previous historical information, the contextual memory retrieval method is used to retrieve related historical events in the event memory partition. When the event memory partition is not hit, the long-term memory partition is used to retrieve related historical events based on semantic similarity.
[0015] The large language model combines the current context and retrieved relevant historical events to generate coherent and accurate content for users;
[0016] A natural forgetting mechanism is introduced in the long-term memory partition. Within the set time interval, the system scans the long-term memory partition and decides whether to retain the memory unit based on the retention rate of each memory unit.
[0017] Furthermore, the historical interaction texts between the user and the large language model are recorded as a sequence {T1, T2, ..., T n}, where T i Represents the text content of the i-th conversation record; each interaction record T is embedded using a pre-trained large language model or a specific text embedding model i Convert to a fixed-dimensional text vector v i , and obtain the vector sequence {v1,v2,…,v n}.
[0018] Furthermore, we define the sliding window w and window size m, and each window w t Contains m consecutive text vectors {v t ,v t+1 ,…,v t+m-1} and slide backward with a step size s; in each sliding window w t , calculate the entropy value H(w t ), the calculation formula is as follows:
[0019]
[0020] Among them, P ij is the vector v in the window i and vector v j The normalized cosine similarity between the two interaction records reflects the weight of the correlation between the two interaction records in the entire window. The calculation formula is as follows:
[0021]
[0022] For adjacent sliding windows w t and w t+1 , calculate the local information entropy difference ΔH(w t ,w t+1 ):
[0023] ΔH(w t,w t+1 )=|H(w t+1 )-H(w t )|
[0024] If ΔH(w t ,w t+1 ) exceeds the predefined entropy difference threshold δ, the dividing point between the sliding windows is marked as a potential event boundary.
[0025] Furthermore, after detecting event boundaries, the continuous interaction sequence is segmented into discrete events, and each event is regarded as an independent memory unit; each event contains an event identifier, timestamp, and a text vector of the event content, and is stored using a vector database.
[0026] Furthermore, inspired by the hierarchical structure of human memory, the memory of the large language model is divided into immediate memory, event memory and long-term memory; among them, immediate memory corresponds to short-term information within the context window of the large language model, which is used to quickly store and process the content of the current interaction, event memory stores medium-term memory units generated by the event cutting algorithm, and long-term memory is used to store low-frequency information that has not been accessed for a long time.
[0027] Furthermore, when the event memory partition is full, the system compresses the memory events with low retention rates and transfers them to the long-term memory area. The retention rate of events is calculated as follows:
[0028]
[0029] Among them, Freq(M i ) represents event M i Recency (M i ) represents event M i The recency of the event is the time interval between the current time and the last recall time of the event, and λ is the time decay factor.
[0030] Furthermore, when the content of the interaction between the user and the large language model involves previous historical information, the contextual memory retrieval method is used to retrieve relevant historical events in the event memory partition. Specifically, when the content of the interaction between the user and the large language model involves previous historical information, the recall probability score of each memory event is calculated to comprehensively evaluate its relevance and importance in the current context. The recall probability calculation formula is as follows:
[0031]
[0032] Where S(M i ) represents the memory event M i The recall probability score, cos(Q,M i) represents the current query vector Q and the memory event vector M i The semantic cosine similarity of i Represents memory event M i The number of historical recalls, Δt i Indicates the current task Q and memory event M i The time interval is λ, and λ is the time decay factor;
[0033] When the event memory partition is not hit, relevant historical events are retrieved in the long-term memory partition according to semantic similarity. Specifically, when the number of recalls of a long-term memory exceeds the set threshold, the memory event is considered to have continuous value to the current task or situation, and it is dynamically recalled to the event memory partition.
[0034] Furthermore, a random number r is generated for each memory unit in the range of [0,1] and compared with the retention rate of the unit; if the random number r is greater than the retention rate, the memory unit is eliminated, thereby simulating the natural forgetting process in human memory.
[0035] The present invention further provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0036] Compared with the existing technology, it has the following beneficial effects:
[0037] The present invention's large language model long-term memory method based on human cognitive inspiration obtains historical interaction texts between users and large language models from a data set and converts the historical interaction texts into text vectors; calculates the semantic similarity distribution of the historical interaction text vectors based on a sliding window, and calculates the local information entropy of the window; determines the event boundary based on the difference in local information entropy values of adjacent windows; uses each discrete event as the memory of the model to manage the memory partitions; when the event memory partition is full, transfers memory events with low retention rates to the long-term memory area; when the content of the interaction involves historical information, uses a contextualized memory retrieval method to retrieve relevant historical events in the event memory partition; the large language model generates content for the user in combination with the current context and the retrieved relevant historical events. The present invention adopts memory encoding and storage based on event cutting theory, multi-level dynamic memory management, and event-level contextualized memory retrieval methods, so that the large language model long-term memory method has dynamic memory and contextualized characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 Schematic diagram of a long-term memory method for a large language model based on human cognition inspiration provided by an embodiment of the present invention.
[0040] Figure 2 This is a diagram of the overall architecture of the long-term memory method for a large language model based on human cognition inspiration provided by an embodiment of the present invention.
[0041] Figure 3 This is an example diagram of a multi-round conversation between a user and a large language model provided by an embodiment of the present invention.
[0042] Figure 4 A schematic diagram of system prompt words for a large language model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0044] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0045] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0046] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0047] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0048] The present invention provides a long-term memory method for a large language model based on human cognitive inspiration. Figure 1 As shown in the figure, the long-term memory method of a large language model based on human cognition inspiration includes the following steps. The overall architecture diagram of the long-term memory method of a large language model based on human cognition inspiration is as follows: Figure 2 shown.
[0049] A long-term memory method for a large language model based on human cognition inspiration, characterized by comprising the steps of:
[0050] The historical interaction texts between the user and the large language model are obtained from the dataset and converted into text vectors.
[0051] Specifically, the historical interaction texts between the user and the large language model are recorded as a sequence {T1, T2, ..., T n}, where T i Represents the text content of the i-th conversation record; each interaction record T is embedded using a pre-trained large language model or a specific text embedding model i Convert to a fixed-dimensional text vector v i , and obtain the vector sequence {v1,v2,…,v n}.
[0052] The semantic similarity distribution of text vectors within the window is calculated according to the sliding window mechanism, and the local information entropy of the window is calculated; for adjacent sliding windows, the difference in their local information entropy values is calculated. If the entropy value difference between adjacent windows exceeds a predefined threshold, it is marked as an event boundary.
[0053] Specifically, define the sliding window w and window size m, each window w t Contains m consecutive text vectors {v t ,v t+1 ,…,v t+m-1} and slide backward with a step size s; in each sliding window w t , calculate the entropy value H(w t ), the calculation formula is as follows:
[0054]
[0055] Among them, P ij is the vector v in the window i and vector v j The normalized cosine similarity between the two interaction records reflects the weight of the correlation between the two interaction records in the entire window. The calculation formula is as follows:
[0056]
[0057] For adjacent sliding windows w t and w t+1 , calculate the local information entropy difference ΔH(w t ,w t+1 ):
[0058] ΔH(w t ,w t+1 )=|H(w t+1 )-H(w t )|
[0059] If ΔH(w t ,w t+1 ) exceeds the predefined entropy difference threshold δ, the dividing point between the sliding windows is marked as a potential event boundary.
[0060] Based on all event boundaries, discrete events are formed, and a vector database is used to store independent events as the memory of the large language model.
[0061] After detecting event boundaries, the continuous interaction sequence is divided into discrete events, and each event is regarded as an independent memory unit; each event contains an event identifier, timestamp, and a text vector of the event content, and is stored using a vector database.
[0062] Specifically, the database may be Faiss, Pinecone, etc.
[0063] The memory of the large language model is partitioned and managed, including immediate memory, event memory and long-term memory.
[0064] Inspired by the hierarchical structure of human memory, the memory of the large language model is divided into immediate memory, event memory, and long-term memory; immediate memory corresponds to short-term information within the context window of the large language model, which is used to quickly store and process the content of the current interaction; event memory stores medium-term memory units generated by the event cutting algorithm; long-term memory is used to store low-frequency information that has not been accessed for a long time.
[0065] Specifically, immediate memory corresponds to short-term information within the context window of a large language model. It is used to quickly store and process the content of the current interaction. It has the smallest capacity but the fastest retrieval speed. Event memory stores medium-term memory units generated by the event segmentation algorithm and has a larger storage capacity and faster retrieval speed. Long-term memory is used to store low-frequency information that has not been accessed for a long time. It has the largest storage capacity and slower retrieval speed.
[0066] When the event memory partition is full, the memory events with low retention rates are compressed and transferred to the long-term memory area.
[0067] When the event memory partition is full, the system compresses the memory events with low retention rates and transfers them to the long-term memory area. The event retention rate is calculated as follows:
[0068]
[0069] Among them, Freq(M i ) represents event M i Recency (M i ) represents event M i The recency of the event is the time interval between the current time and the last recall time of the event, and λ is the time decay factor.
[0070] When the content of the interaction between the user and the large language model involves previous historical information, the contextual memory retrieval method is used to retrieve related historical events in the event memory partition; when the event memory partition is not hit, the relevant historical events are retrieved in the long-term memory partition according to semantic similarity.
[0071] Specifically, when the content of the interaction between the user and the large language model involves previous historical information, the contextual memory retrieval method is used to retrieve relevant historical events in the event memory partition. Specifically, when the content of the interaction between the user and the large language model involves previous historical information, the recall probability score of each memory event is calculated to comprehensively evaluate its relevance and importance in the current context. The recall probability calculation formula is as follows:
[0072]
[0073] Where S(M i ) represents the memory event M i The recall probability score, cos(Q,M i ) represents the current query vector Q and the memory event vector M i The semantic cosine similarity of i Represents memory event M i The number of historical recalls, Δt i Indicates the current task Q and memory event M i The time interval is λ, and λ is the time decay factor;
[0074] When the event memory partition is not hit, relevant historical events are retrieved in the long-term memory partition according to semantic similarity. Specifically, when the number of recalls of a long-term memory exceeds the set threshold, the memory event is considered to have continuous value to the current task or situation, and it is dynamically recalled to the event memory partition.
[0075] The large language model combines the current context and retrieved relevant historical events to generate coherent and accurate content for users.
[0076] Specifically, the large language model combines the current context and retrieved relevant historical events to generate more coherent and accurate content for users based on system prompt words.
[0077] A natural forgetting mechanism is introduced in the long-term memory partition. Within the set time interval, the system scans the long-term memory partition and decides whether to retain the memory unit based on the retention rate of each memory unit.
[0078] A probability-based random elimination strategy is introduced in the long-term memory partition. After a set time interval (such as weekly or monthly), the system will scan the long-term memory partition and decide whether to retain each memory unit based on its retention rate.
[0079] A random number r is generated for each memory unit in the range of [0,1] and compared with the retention rate of the unit; if the random number r is greater than the retention rate, the memory unit is eliminated, thus simulating the natural forgetting process in human memory.
[0080] In a specific embodiment, Figure 3 As shown in the figure, users engage in multiple rounds of conversations with the large language model, focusing on topics such as science fiction preferences, recommendations, and sharing of life events. Throughout this interaction, the large language model's long-term memory method, inspired by human cognition, is implemented in the following steps:
[0081] Data acquisition and vector conversion: We obtain interaction records between users and large language models from open source datasets. These records are text sequences arranged in chronological order, such as T1, "I like to read science fiction novels. Can you recommend some works to me?", and T2, "The Three-Body Problem series, the Foundation series, The Time Machine, and The Mars Trilogy are all good science fiction novels." We use a pre-trained general text vector model (such as Qwen's text-embedding-v2) to transform each interaction record T into a vector. i Convert to a vector v of fixed dimension (such as 1536 dimensions) i , forming a vector sequence {v1,v2,…,v n}.
[0082] Calculate local information entropy: Define a sliding window w, set the maximum window size m = 4, start the window size from 2, and the step size s = 1. Take the first window w1 as an example, it contains vectors v1, v2, and the second window w2 contains vectors v1, v2, v3. Taking the calculation of the local information entropy of window w2 as an example, first calculate the cosine similarity between the vectors in the window:
[0083]
[0084] Normalize the cosine similarity between vectors to get the probability distribution:
[0085]
[0086] Calculate the entropy value H(w1) of the text vector distribution within the window:
[0087]
[0088] Step (3) Detect event boundaries: For adjacent sliding windows w t and w t+1 , calculate the local information entropy difference ΔH(w t ,w t+1 ):
[0089] ΔH(w t ,w t+1 )=|H(w t+1 )-H(w t )|
[0090] Preset the entropy difference threshold δ (for example, 0.6), if ΔH(w t ,w t+1 ) exceeds δ, the dividing point between the two windows is marked as a potential event boundary. For example, if H(w1) = 0.83, H(w2) = 1.09, and H(w3) = 1.76, then ΔH(w2, w3) = 1.76 - 1.09 = 0.67 > 0.6, then it is considered that there is a significant semantic change between the two windows. The system will mark this position as a potential event boundary, and the new sliding window will start from this event boundary.
[0091] Event storage: Based on the detected event boundaries, continuous interaction sequences are divided into discrete events. For example, event M1: {User: I like reading science fiction novels. Can you recommend some works to me? Intelligent assistant: The Three-Body series, the Foundation series, The Time Machine, and the Mars trilogy are all very good science fiction novels. User: I also like watching science fiction movies, and my favorite is Interstellar. Intelligent assistant: Interstellar is indeed a classic, and the Dune series is also a good science fiction movie. I recommend you to watch it.}, Event M2: {User: I participated in a table tennis competition last week and won the runner-up. Intelligent assistant: Wow, congratulations! It’s amazing to win the runner-up in a table tennis competition. You must have very good skills.} Each event contains a unique identifier, a timestamp, and a text vector of the event content. A vector database (such as Faiss) is used to store each independent event.
[0092] Memory partition management: Based on the hierarchical caching concept, the large language model memory is divided into immediate memory, event memory, and long-term memory. Immediate memory corresponds to short-term information within the context window of the large language model, such as the latest interaction between the user and the large language model. It has the smallest capacity and the fastest retrieval speed, and is stored directly within the window of the large language model. Event memory stores medium-term memory units generated by the event segmentation algorithm, such as events M1 and M2 segmented in step 4. It has a large storage capacity and fast retrieval speed and is stored in the vector database. Long-term memory is used to store low-frequency information that has not been accessed for a long time. It has the largest storage capacity and a slower retrieval speed and is stored in the vector database.
[0093] Memory event scheduling: Assume that the capacity limit of the event memory partition is set to store 100 events. When the event memory partition reaches the preset capacity limit, the system will process the memory events with low retention rates, generate event summaries to compress the event content, and then transfer these compressed events to the long-term memory partition. During this process, the system calculates the retention rate of each event. Taking event M1 as an example, the retention rate calculation formula is:
[0094]
[0095] Assume that λ = 0.5, the recall frequency of event M1 Freq(M1) = 5 (i.e., it has been retrieved by the system 5 times), and the recentness Recency(M1) = 30 (assuming that it is in days, indicating that it was last recalled 30 days ago from the current time). Then Assuming that the calculated retention rate of event M1 is the lowest, event M1 is compressed and transferred to the long-term memory area for storage.
[0096] Contextual memory retrieval: The user asks "Do you remember what my favorite science fiction movie is?" The interaction involves previous historical information. The system calculates the probability score of memory event recall:
[0097]
[0098] Taking event M1 as an example, the semantic cosine similarity between the current query vector Q and the text vector of event M1 is cos(Q,M1)=0.95, and the historical recall times of event M1 is c i =1, the time interval between the current task and the event is Δt1 = 3 days, λ = 0.5, then The events in the event memory partition are sorted according to the scores to retrieve the most relevant historical events.
[0099] Long-term memory partition retrieval: If the event memory partition is not hit, the system searches the long-term memory partition based on semantic similarity. Assuming the recall threshold is set to 4 times, if a long-term memory item is recalled 5 times, the system deems the memory event potentially valuable for the current task or context and dynamically transfers it from the long-term memory partition to the event memory partition for subsequent rapid retrieval.
[0100] Large language model generates responses: The large language model combines the current context and the retrieved relevant historical events M1, based on predefined system prompt words (such as Figure 4 ), generating the response "Your favorite science fiction movie is Interstellar".
[0101] Long-term memory forgetting: Schedule weekly scans of the long-term memory partition. Generate a random number r (in the range [0, 1], for example, r = 0.46) for each memory cell and compare it to the cell's retention rate (e.g., a cell with a retention rate of 0.38). If r is greater than the retention rate, the cell is eliminated, simulating the natural forgetting process in human memory and freeing up storage space in the long-term memory partition.
[0102] The computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0103] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions that, when executed, enable a processor to perform a large language model long-term memory method based on human cognition inspiration.
[0104] The processor is used to provide computing and control capabilities to support the operation of the entire computer equipment.
[0105] The internal memory provides an environment for the operation of computer programs in non-volatile storage media. When the computer program is executed by the processor, the processor can execute a long-term memory method of a large language model based on human cognition inspiration.
[0106] The network interface is used to communicate with other devices over the network. Those skilled in the art will appreciate that the above-described computer device structure is merely a partial structure related to the present invention and does not limit the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0107] The processor is used to run a computer program stored in the memory, which implements the long-term memory method of a large language model based on human cognition inspiration as described in Example 1.
[0108] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0109] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0110] The present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the long-term memory method for a large language model based on human cognition inspiration as described in Example 1.
[0111] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0112] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0113] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0114] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.
[0115] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.
Claims
1. A long-term memory method for large language models based on human cognition inspiration, characterized by: Including steps: Obtain historical interaction texts between users and a large language model from a dataset, and convert the historical interaction texts into text vectors; The semantic similarity distribution of text vectors within the window is calculated based on the sliding window mechanism, and the local information entropy of the window is calculated. For adjacent sliding windows, the difference in their local information entropy values is calculated. If the entropy difference between adjacent windows exceeds a predefined threshold, it is marked as an event boundary. Based on all event boundaries, discrete events are formed, and a vector database is used to store independent events as the memory of the large language model; Partitioning the memory of the large language model into immediate memory, event memory, and long-term memory; When the event memory partition is full, the memory events with low retention rate are compressed and transferred to the long-term memory area; When the content of the user's interaction with the large language model involves previous historical information, the contextual memory retrieval method is used to retrieve related historical events in the event memory partition. When the event memory partition is not hit, the long-term memory partition is used to retrieve related historical events based on semantic similarity. The large language model combines the current context and retrieved relevant historical events to generate coherent and accurate content for users; A natural forgetting mechanism is introduced in the long-term memory partition. Within the set time interval, the system scans the long-term memory partition and decides whether to retain the memory unit based on the retention rate of each memory unit.
2. The method according to claim 1, characterized in that The historical interaction texts between the user and the large language model are recorded as a sequence {T1, T2, ..., T n }, where T i Represents the text content of the i-th conversation record; each interaction record T is embedded using a pre-trained large language model or a specific text embedding model i Convert to a fixed-dimensional text vector v i , and obtain the vector sequence {v1,v2,…,v n }.
3. The method according to claim 1, characterized in that Define sliding window w and window size m, each window w t Contains m consecutive text vectors {v t ,v t+1 ,…,v t+m-1 } and slide backward with a step size s; in each sliding window w t , calculate the entropy value H(w t ), the calculation formula is as follows: Among them, P ij is the vector v in the window i and vector v j The normalized cosine similarity between the two interaction records reflects the weight of the correlation between the two interaction records in the entire window. The calculation formula is as follows: For adjacent sliding windows w t and w t+1 , calculate the local information entropy difference ΔH(w t ,w t+1 ): ΔH(w t ,w t+1 )=|H(w t+1 )-H(w t )| If ΔH(w t ,w t+1 ) exceeds the predefined entropy difference threshold δ, the dividing point between the sliding windows is marked as the event boundary.
4. The method according to claim 1, wherein After detecting event boundaries, the continuous interaction sequence is divided into discrete events, and each event is regarded as an independent memory unit; each event contains an event identifier, timestamp, and a text vector of the event content, and is stored using a vector database.
5. The method according to claim 1, characterized in that Inspired by the hierarchical structure of human memory, the memory of the large language model is divided into immediate memory, event memory, and long-term memory; Among them, immediate memory corresponds to short-term information within the context window of the large language model, which is used to quickly store and process the content of the current interaction. Event memory stores medium-term memory units generated by the event cutting algorithm, and long-term memory is used to store low-frequency information that has not been accessed for a long time.
6. The method according to claim 1, characterized in that When the event memory partition is full, the system compresses the memory events with low retention rates and transfers them to the long-term memory area. The event retention rate is calculated as follows: Among them, Freq(M i ) represents event M i Recency (M i ) represents event M i The recency of the event is the time interval between the current time and the last recall time of the event, and λ is the time decay factor.
7. The method according to claim 1, characterized in that When the content of the interaction between the user and the large language model involves previous historical information, the contextual memory retrieval method is used to retrieve related historical events in the event memory partition. Specifically, when the content of the interaction between the user and the large language model involves previous historical information, the recall probability score of each memory event is calculated to comprehensively evaluate its relevance and importance in the current context. The recall probability calculation formula is as follows: Where S(M i ) represents the memory event M i The recall probability score, cos(Q,M i ) represents the current query vector Q and the memory event vector M i The semantic cosine similarity of i Represents memory event M i The number of historical recalls, Δt i Indicates the current task Q and memory event M i The time interval is λ, and λ is the time decay factor; When the event memory partition is not hit, relevant historical events are retrieved in the long-term memory partition according to semantic similarity. Specifically, when the number of recalls of a long-term memory exceeds the set threshold, the memory event is considered to have continuous value to the current task or situation, and it is dynamically recalled to the event memory partition.
8. The method according to claim 1, characterized in that A random number r is generated for each memory unit in the range of [0,1] and compared with the retention rate of the unit; if the random number r is greater than the retention rate, the memory unit is eliminated, thus simulating the natural forgetting process in human memory.
9. A computer device, characterized in that: The device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Cited By
Hierarchical urban operation agent memory management method and system
CN120910313A
A hierarchical city operation intelligent agent memory management method and system
CN120910313B
Memory optimization system for large language model
CN121328718A
Request processing method and device of intelligent assistant based on large model and related equipment
CN121388190A
Dynamic compression and efficient recall collaborative end-side robot memory system management method
CN121552393A