Proxy memory dynamic retrieval scheduling method and system based on exponential decay model
Patent Information
- Application Number
- CN202611013231.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]传统代理记忆检索调度方式大多采用固定排序与静态权重机制,不会随时间推移和访问行为变化调整记忆的检索优先级,无法模拟自然状态下的记忆遗忘规律
一、本发明通过引入指数衰减模型模拟记忆遗忘规律,结合多维度属性构建动态优先级计算体系,能够贴合实际使用场景量化记忆可检索价值。依托分层阈值划分不同调度状态,搭配差异化延迟计算逻辑,可对检索任务进行精细化分流管控,优先保障高价值记忆快速响应检索需求,同时合理调控低优先级记忆的检索响应节奏。结合访问场景与记忆状态动态调整核心属性参数,实现记忆状态的自适应强化,让记忆活跃度随访问行为自主更新。整套机制摒弃固定检索排序模式,依据历史访问行为与时间变化实时调整检索顺序,大幅提升检索匹配的精准度与整体调度效率,适配高频、多类型记忆检索的运行需求。
Smart Images

Figure CN122838468A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent agent data management technology, specifically to an agent memory dynamic retrieval and scheduling method and system based on an exponential decay model. Background Technology
[0002] With the rapid development of AI agent technology, agent systems need to continuously retain massive amounts of interactive memory data to support core functions such as dialogue interaction, task execution, and business response. The effective retrieval and scheduling of memory data directly determines the service quality and operational efficiency of intelligent agents. Currently, various intelligent agent platforms accumulate a large amount of historical interaction content, business tags, access records, and other memory information, and the data scale continues to expand with usage time. The industry generally relies on retrieval and scheduling mechanisms to manage memory data, aiming to quickly match retrieval needs and retrieve effective memory content. Optimizing retrieval logic based on dimensions such as memory activity and access frequency has also become a mainstream development direction in this field. How to simulate the forgetting and reinforcement characteristics of human memory and build a dynamic retrieval system based on data access patterns has become a key research topic in the field of intelligent agent memory management, and related technologies are gradually evolving towards dynamic, refined, and layered approaches.
[0003] Traditional proxy memory retrieval scheduling methods mostly employ fixed sorting and static weighting mechanisms, failing to adjust memory retrieval priorities over time and with changes in access behavior, thus unable to simulate the natural forgetting curve of memories. Long-unaccessed memories continue to occupy active retrieval queue resources, not only lengthening retrieval computation time but also reducing the hit rate of effective memories. Most existing solutions lack robust state partitioning and latency control logic, failing to differentiate response rhythms for memories with varying activity levels, resulting in inefficient allocation of retrieval resources. Furthermore, traditional technologies lack a systematic hierarchical storage design, making it difficult to automatically archive inactive memories. Massive amounts of data remain concentrated in active areas, continuously increasing system storage and computational load. For re-accessed archived memories, there are no corresponding transition control strategies, easily disrupting the original retrieval order, and memory attributes cannot be automatically updated based on access scenarios, resulting in significant shortcomings in overall adaptability and operational stability. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies by providing a dynamic retrieval and scheduling method and system for agent memories based on an exponential decay model, and to build a complete technical solution for the full lifecycle management of intelligent agent memories. The solution achieves unified collection, standardization, and partitioned storage of memory data, access records, and configuration parameters. It calculates memory recall based on the exponential decay law, generates retrieval priorities by combining static weights, divides different scheduling states according to priorities, sets differentiated retrieval delays, and outputs retrieval results in an orderly manner. The system dynamically updates memory stability parameters and applies boundary constraints based on access scenarios, while automatically archiving memories through periodic scanning and setting up observation and control mechanisms for backtracking memories. The overall technology abandons the static retrieval mode, realizing automated operation of memory retrieval, state updates, and hierarchical storage, effectively improving retrieval efficiency and resource utilization, and can adapt to the massive memory management needs of various intelligent agent scenarios.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On the one hand, a proxy memory dynamic retrieval and scheduling method based on an exponential decay model, the specific steps of which are as follows:
[0006] S1 collects the memorized text content, category labels, creation time, initial stability parameters and initial static weights, collects user active access and system callback access records and timestamps, loads weight grading standards, stability parameter boundaries, delay parameters, priority thresholds, archiving thresholds, learning rates, depth factors, fixed depth factors and scanning cycles corresponding to gradient degradation states, and stores them in partitions after standardization. S2, based on the data stored in S1, calculates the time interval between the current time and the last access timestamp for the memories in the active retrieval queue, uses the exponential decay forgetting model, calculates the real-time recall based on the time interval and the current stability parameter, and then combines the real-time recall with the static weight to calculate the comprehensive retrieval priority value. S3: Based on the comprehensive retrieval priority calculated in S2, the comprehensive retrieval priority is compared with the first threshold and the second threshold to determine the scheduling state. The retrieval delay is calculated based on the scheduling state and the delay parameter. Memory not lower than the second threshold is formed into an execution queue in descending order of priority, and the retrieval results are returned in sequence according to the delay. S4. For the memory of successful access in S3, determine the depth factor based on the access scenario, calculate the updated stability parameter according to the depth factor, learning rate and current stability parameter, perform truncation processing when the updated stability parameter exceeds the boundary, and update the last access timestamp. S5, traverse the active retrieval queue according to the scanning cycle, and obtain the comprehensive retrieval priority of each memory using the calculation method of S2. Memory whose priority is lower than the archiving threshold for multiple consecutive rounds is moved to the offline storage repository; when a user is detected to actively access an archived memory, the archived memory is moved back to the active retrieval queue and an observation period is started, during which scheduling is restricted. When S5 detects that a user has initiated an access action to a memory in the offline repository, it automatically copies the memory from the offline repository back to the active retrieval queue and adds an observation period status mark to the memory. Memory with the observation period mark does not execute regular scheduling rules, but only independent exclusive scheduling rules. After the observation period ends, the observation period mark is cleared, and the corresponding regular scheduling status is matched according to the access records during the observation period.
[0007] Furthermore, the memory text content collected in S1 includes the complete original interactive text at the time of memory generation and the associated context identifier; the category label includes the memory business category identifier and the access permission identifier; the creation time is the system standard time at the moment of memory generation; the initial stability parameter and the initial static weight are the initial attribute values assigned when the category label is matched at the moment of memory generation; the collected user-initiated access records include the access initiator user identifier and the access-triggered interaction event identifier; the system callback access records include the access-triggered system task identifier and the calling module identifier; the timestamps corresponding to all access records are the system standard timestamps at the moment of access, and the timestamps are bound and stored one by one with the corresponding access records.
[0008] Furthermore, the weight grading standard loaded in S1 is a pre-configured mapping rule between memory classification labels and static weight levels; the stability parameter boundary is a pre-configured upper and lower limit rule for the legal values of the stability parameter; the latency parameter is a pre-configured baseline value rule for calculating retrieval latency; the priority threshold is a pre-configured two-level judgment rule for memory scheduling state division; the learning rate is a pre-configured amplitude rule for the iterative update of the stability parameter; the depth factor is a pre-configured amplitude rule for memory reinforcement corresponding to different access scenarios; the scanning cycle is a pre-configured interval rule for global scanning of the full memory; all configuration parameters are fixed rules that take effect globally and are uniformly stored in the global configuration partition after loading.
[0009] Furthermore, in S2, the system first verifies the validity of the received retrieval request. After the verification is successful, it retrieves a list of all memory entries currently in the active retrieval queue and performs calculations on each memory entry in the order they were entered into the database. First, it calculates the time interval by the difference between the last access timestamp of a single memory and the current system standard time. Then, it calculates the real-time recall based on the time interval and the current stability parameter of the corresponding memory. Finally, it calculates the comprehensive retrieval priority by combining the real-time recall with the initial static weight of the corresponding memory. The calculation result of a single memory is bound to the unique identifier of the corresponding memory and temporarily stored in a temporary cache. The data in the temporary cache is only valid within the processing cycle of this retrieval request and is automatically cleared after the current retrieval is completed.
[0010] Furthermore, in S3, the system first reads the first threshold and the second threshold from the global configuration partition and completes the parameter validity verification. After the verification is passed, it reads the comprehensive retrieval priority stored in the temporary cache one by one. The comprehensive retrieval priority is first compared with the first threshold. If the priority is not lower than the first threshold, it is marked as a normal retrieval state. If the priority is lower than the first threshold, it is compared with the second threshold. If the priority is not lower than the second threshold, it is marked as a gradient degradation state. If the priority is lower than the second threshold, it is marked as a pending archiving state. Each state is marked as an identifier field with a fixed format, which is bound to the corresponding unique identifier stored.
[0011] Furthermore, in S3, the system matches the corresponding delay calculation rule according to the state flag bound to each memory, and performs queue admission screening after completing the delay value calculation of all memories; only memory entries with normal retrieval state flags and gradient degradation state flags are retained, and memory entries with pending archiving state flags are removed; after screening, the memory entries are sorted according to a fixed rule from high to low comprehensive retrieval priority values, and a retrieval execution queue is generated and the queue order is locked. The queue order will not be adjusted during this retrieval processing cycle.
[0012] Furthermore, in S4, after the system completes the output of the retrieval result for a single memory and confirms the successful access, it retrieves the initiator identifier of this access behavior. If the initiator identifier is a user identifier, it is determined to be a user-initiated access scenario, and the depth factor corresponding to the user access scenario is matched. If the initiator identifier is a system server identifier, it is determined to be a system callback access scenario, and the depth factor corresponding to the system access scenario is matched. If the memory being accessed has a gradient degradation state marker, the fixed depth factor corresponding to the gradient degradation state is matched first, and the scenario is no longer determined based on the initiator identifier.
[0013] Furthermore, after the system completes the iterative calculation of the stability parameters in S4, it immediately triggers the parameter boundary verification process; it reads the upper and lower threshold values of the stability parameters from the global configuration partition, and compares the new stability parameters obtained by the iterative calculation with the upper and lower threshold values respectively; if the new stability parameter value is greater than the upper threshold, the upper threshold is taken as the final update value; if the new stability parameter value is less than the lower threshold, the lower threshold is taken as the final update value; if the new stability parameter value is within the upper and lower limit range, the iteratively calculated value is retained; the final update value synchronously overwrites the corresponding memory's original stored stability parameter field.
[0014] Furthermore, in S5, when the system reaches the trigger node of the preset scanning cycle, it automatically starts a global scan, and the scanning traversal range covers all memory entries in the currently active retrieval queue; it pulls the current comprehensive retrieval priority of the corresponding memory and the priority history of previous global scans one by one; if the priority results of a single memory in N consecutive rounds of scanning are all lower than the archiving threshold, then the memory entry is removed from the storage partition of the active retrieval queue, and all attribute fields, historical access records and all status flags of the memory are completely copied and written to the offline repository. The writing process does not modify any original data content; where N is a preset positive integer and N≥3.
[0015] On the other hand, the proxy memory dynamic retrieval and scheduling system based on the exponential decay model includes a data acquisition and storage module, a priority calculation module, a retrieval and scheduling module, a parameter update module, and a memory archiving and migration module. The data acquisition and storage module is used to collect and remember text content, category tags, creation time, initial stability parameters, initial static weights, user active access records, system callback access records and corresponding timestamps, load various global configuration parameters, standardize all data and store it in partitions, and maintain global configuration partitions, active retrieval queue storage partitions and offline storage partitions. The priority calculation module is used to receive retrieval requests and complete legality verification. Based on the last access time interval of the memory and the current stability parameters, it calculates the real-time recall rate using the exponential decay forgetting model, calculates the comprehensive retrieval priority by combining the initial static weights, and stores the calculation results in a temporary cache after binding them with the unique identifier of the memory. The retrieval scheduling module is used to read two-level priority thresholds, divide the normal retrieval state, gradient degradation state, and pending archiving state according to the comprehensive retrieval priority, calculate the retrieval delay based on the scheduling state and delay parameters, filter valid memory entries and generate a retrieval execution queue in descending order of priority, and output the retrieval results in sequence according to the retrieval delay. The parameter update module is used to match the corresponding depth factor according to the access initiator or memory state after successful memory access, iteratively calculate the new stability parameter in combination with the learning rate, and perform truncation processing according to the upper and lower limits of the stability parameter, and synchronously update the memory's stability parameter and the last access timestamp. The memory archiving and migration module is used to traverse the active retrieval queue according to a preset scanning cycle, determine whether the memory meets the archiving conditions, and perform offline migration; when it is detected that a memory in the offline repository is accessed, the memory is migrated back to the active retrieval queue and an observation period mark is added, a dedicated scheduling rule is executed on the memory during the observation period, and the mark is cleared and matched with the normal scheduling status after the observation period ends; wherein, the data acquisition and storage module corresponds to step S1 of claim 1; the priority calculation module corresponds to step S2; the retrieval scheduling module corresponds to step S3; the parameter update module corresponds to step S4; and the memory archiving and migration module corresponds to step S5. Beneficial effects
[0016] Compared with existing technologies, this agent memory dynamic retrieval scheduling method and system based on the exponential decay model has the following advantages: I. This invention simulates the forgetting curve of memory by introducing an exponential decay model and constructs a dynamic priority calculation system based on multi-dimensional attributes, enabling the quantification of the retrieval value of memories in practical use cases. By dividing different scheduling states into hierarchical thresholds and combining them with differentiated delay calculation logic, it allows for refined traffic management of retrieval tasks, prioritizing high-value memories for rapid response to retrieval needs while rationally controlling the retrieval response pace of low-priority memories. By dynamically adjusting core attribute parameters based on access scenarios and memory states, it achieves adaptive reinforcement of memory states, allowing memory activity to update autonomously with access behavior. The entire mechanism abandons fixed retrieval sorting patterns, adjusting the retrieval order in real time based on historical access behavior and time changes, significantly improving the accuracy of retrieval matching and overall scheduling efficiency, adapting to the operational needs of high-frequency, multi-type memory retrieval.
[0017] II. This invention, through a periodic global scanning mechanism and a hierarchical storage architecture, automatically identifies and archives long-term inactive memories offline, effectively reducing the data volume of the active queue and lowering the resource consumption of daily retrieval operations. Dedicated migration and observation control logic is designed for archived memories, ensuring normal retrieval of offline data while preventing migration from interfering with the regular retrieval order in the short term, thus balancing data integrity and system stability. Multiple data storage partitions and independent cache areas are defined to standardize the entire data flow and temporary data management process, ensuring that all types of original data are not tampered with. Differentiated processing rules are designed based on user access and system callback behaviors to adapt to the usage needs of diverse users. The overall architecture possesses good versatility and robustness, and can stably support the storage, retrieval, and full lifecycle management of large-scale proxy memories for a long period.
[0018] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0020] Figure 1 The flowchart shows the proxy memory dynamic retrieval scheduling method based on the exponential decay model. Figure 2 This is a framework diagram of a proxy memory dynamic retrieval scheduling system based on an exponential decay model. Detailed Implementation
[0021] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0022] Reference Figure 1One embodiment of the present invention proposes a dynamic retrieval and scheduling method for proxy memories based on an exponential decay model. The method employs an exponential decay forgetting model to quantify the recallability of memories, divides the scheduling state into hierarchical levels with dual thresholds, and completes the memory archiving and relocation mechanism by combining partitioned storage with periodic global scanning. It can dynamically change the retrieval priority and retrieval response latency according to memory access behavior and access duration, and simultaneously complete the iterative optimization of memory attribute parameters and automatic hot and cold data diversion and control, so as to realize the dynamic scheduling, orderly response and full life cycle management of proxy memories in retrieval scenarios.
[0023] The method described in this embodiment specifically includes: collecting the text content, category tags, time information, and initial attribute parameters corresponding to the proxy memory; simultaneously collecting two types of behavior records—user-initiated access and system callback access—and their corresponding timestamp data; loading various global configuration rules preset by the system; standardizing all collected and loaded data; dividing the data into different storage partitions for independent storage based on its purpose; calculating the time span of a single memory from its last access to the current moment based on the memory data already stored in the partition and the access timestamp; using the exponential decay forgetting model to solve for the real-time recallability of the memory; and combining the initial static weights inherent in the memory to calculate the comprehensive retrieval priority corresponding to each memory; comparing the calculated comprehensive retrieval priority with the two-level thresholds preset by the system in sequence; dividing different memory scheduling states according to the comparison results; calculating the retrieval delay time of a single memory based on the corresponding scheduling state and delay-related rules; filtering out valid memory entries that can participate in the retrieval response; and generating a retrieval execution list after sorting by priority. The system sequentially outputs search results to the queue according to the search delay corresponding to each entry. When a memory completes a search and confirms a successful access, it matches the depth factor based on the initiating entity of the access and the current scheduling state of the memory. It then iteratively calculates the memory's original stability parameters based on the preset learning rate. Simultaneously, it truncates the calculation results according to the legal value range of the stability parameters and updates the last access timestamp of the memory to provide a time basis for subsequent calculations. According to the fixed scanning interval set by the system, it periodically performs a full-entry traversal check on the active search queue. Combining the priority data obtained from multiple rounds of scanning of a single memory, it determines whether the entry meets the archiving criteria and migrates the qualified memories from the active queue to the offline repository. When a memory in the offline repository is triggered for access, it is automatically migrated back to the active search queue and an observation period status flag is added. Memoryes in the observation period execute independent scheduling and control logic. After the observation period ends, the status flag is cleared, and the regular scheduling state is rematched based on the access records generated during the observation period.
[0024] Specifically, the overall idea behind the implementation of this invention is to build a closed-loop agent memory management system that covers data access, computation and analysis, retrieval and scheduling, parameter updates, and data flow. The system first completes the unified collection, format standardization, and partitioned storage of various types of raw data, behavioral data, and system configuration rules, providing standardized and complete data support for all subsequent operations and scheduling. It introduces an exponential decay model to simulate the objective law of memory gradually weakening over time, and combines this with the inherent weight of the memory itself to quantify priority, using this as the basis for retrieval resource allocation and response order arrangement. It classifies scheduling states through two-level thresholds, setting differentiated retrieval delay calculation logic for different states, allowing memories with different activity levels to have appropriate response rhythms. After a memory is successfully accessed, the system adjusts the enhancement magnitude based on the access scenario and state information, iteratively updates stability parameters, and ensures that all attribute parameters remain within the compliant range by relying on parameter boundary constraints. It identifies long-term low-activity memories through periodic global scanning and performs offline archiving, reducing the data volume and computational pressure of the active queue. For scenarios where archived memories are accessed again, a migration process and observation mechanism are set up to avoid scheduling anomalies in short-term migration of memories. Ultimately, it achieves dynamic management of the entire process of proxy memory from generation and storage, real-time retrieval, parameter iteration to cold and hot memory circulation.
[0025] Optionally, the collection of memory data, access records, and global configuration parameters, and the completion of standardized partitioned storage, specifically includes: collecting the text content, category tag, creation time, initial stability parameters, and initial static weight corresponding to each memory record; simultaneously summarizing user-initiated access records and system callback access records generated during system operation; and extracting the system timestamp corresponding to each access behavior; batch reading the entire set of configuration rules pre-configured by the system, including weight grading standards, stability parameter boundaries, latency parameters, two-level priority thresholds, archiving thresholds, learning rates, depth factors corresponding to multiple scenarios, fixed depth factors specific to gradient degradation states, and global scanning cycles; uniformly performing standardized processing such as format verification, field completion, and format conversion on all the above data and rules; removing invalid data with missing fields or incorrect formats; and then dividing different storage partitions according to data functions and usage scenarios to complete the final storage.
[0026] The collected memory text content includes not only the complete original interactive text retained during memory creation, but also identification information for associating with contextual content. Category tags are divided into two types, used to mark the business category to which the memory belongs and the access permission range corresponding to the memory. The creation time accurately records the system standard time at the moment the memory is generated. Initial stability parameters and initial static weights are not uniformly assigned values, but are initialized during the memory generation stage based on the corresponding level matched to the bound category tags, becoming inherent attributes of the memory. User-initiated access records record the identity of the user initiating the access operation and the identifier of the interaction event triggered by this access. System callback access records record the identifier of the background system task that triggered the access behavior and the identifier of the functional module responsible for calling the memory. Each access record is bound to the system standard timestamp of the time the behavior occurred, ensuring a one-to-one correspondence between the access behavior and the time information, preventing time sequence errors.
[0027] The system's weighting grading standard is a pre-established mapping relationship between memory classification labels and static weight levels; the stability parameter boundary clearly defines the maximum and minimum allowed values of the stability parameter, delineating the legal range of the parameter; the latency parameter serves as the basic benchmark rule used in the retrieval latency operation; two-level priority thresholds are used to distinguish different scheduling states and are the core basis for determining the retrieval state; the archiving threshold serves as the criterion for determining whether active memories need to be transferred to offline storage; the learning rate is used to control the magnitude of change in each iteration update of the stability parameter; the depth factor distinguishes multiple types and is used to regulate the enhancement magnitude of memory attributes under different access scenarios; the scanning cycle specifies the time interval for the system to perform a full traversal detection of the active retrieval queue. All the above configuration rules are uniformly stored in an independent partition after loading and take effect throughout the entire system. All subsequent operations and scheduling processes are executed according to this set of rules.
[0028] For example, during the data collection and partitioned storage execution phase, the system connects to the proxy memory generation port and the access behavior log port to continuously collect new memory data and historical access behavior data. For each new memory, the system fully retains the original interaction text and context association identifier, and automatically matches the corresponding initial parameters based on the business segment to which the memory belongs and the open access permissions, ensuring that memories with the same category of tags have a unified initial attribute standard. For all access behaviors, the system distinguishes between two scenarios: human-initiated operations and automatic calls from background programs, fully recording information such as the behavior initiator, triggering event, and calling module, and using the system's unified clock to mark timestamps, ensuring that each behavior log has complete time sequence information. When loading global configuration rules, the system reads all preset control rules at once, performs rule format verification, and stores them centrally. After all data is integrated, the system divides the data according to partitioning logic, storing configuration rules separately, and assigning basic memory data and behavior logs to the corresponding partitions of the active queue. Different partitions are isolated from each other and can be read and written independently, which can avoid data interference and improve the efficiency of subsequent data retrieval. The entire process will be executed automatically during the system startup phase, and it also supports incremental collection of new data during operation, ensuring the real-time performance and integrity of the data source.
[0029] Optionally, the calculation of real-time recall based on access time interval and exponential decay model, and the solution of comprehensive retrieval priority by combining static weights, specifically includes: when the system receives an externally issued retrieval request, it first verifies the legitimacy of the request's initiation channel, message format, and access permissions. If the verification result fails, the current retrieval process is terminated directly and a prompt message is provided; if the verification passes, the system reads the list of all valid memory entries in the currently active retrieval queue. The system processes each memory sequentially according to the order in which the memory entries were stored in the queue. For a single memory, the timestamp bound to the last access of the memory is extracted, and the time difference between the timestamp and the current system standard time is calculated. This time difference is the last access time interval of the memory. The calculated time interval and the current effective stability parameter of the memory are substituted into the exponential decay forgetting model formula to calculate the current real-time recall of the memory. Then, the real-time recall is multiplied by the initial static weight assigned during the memory initialization phase to obtain the comprehensive retrieval priority corresponding to the memory. The calculation result of a single memory is bound to the memory's unique identifier and temporarily stored in a cache area specifically allocated for this retrieval process. The data in this cache area is only valid within the current retrieval processing cycle. Once the entire retrieval task is completed, the system automatically clears all temporary data in the cache to avoid redundant data accumulation that would consume system resources.
[0030] The mathematical expression for calculating real-time recall is:
[0031] In the formula, For the real-time recall of memories, To remember the time interval between the last visit and the current time, To remember the current stability parameters, It is a natural constant.
[0032] The formula for calculating the attenuation coefficient is: In the formula Represents the exponential decay coefficient. This is a parameter for memory stability, used to determine how quickly memory recall decays over time. The mathematical expression for calculating the overall retrieval priority value is:
[0033] In the formula, To optimize the search priority values, This is the real-time recall rate calculated in the previous step. The initial static weights for memory.
[0034] For example, in the priority calculation stage, after a retrieval request enters the system, it first enters the verification stage. The verification mechanism intercepts illegal requests, requests with abnormal formats, and requests without access permissions, reducing invalid calculations at the source. For valid retrieval requests that pass verification, the system traverses all memory entries in the active queue, performing calculations one by one according to the order in which the entries were added to the database. The application of the exponential decay model aligns with the gradual weakening of memories over time; the longer the access interval, the lower the recallability of the corresponding memory. The stability parameter changes the rate of weakening; the larger the parameter value, the smoother the rate of memory weakening. After the recallability calculation is completed, a second calculation is performed based on the memory's fixed static weight. The static weight is determined by the memory's business importance and permission level; the higher the importance of the memory, the higher the static weight, and the higher the final overall priority. All calculation results are bound to the memory's unique number and stored in a temporary cache. The method of calculating and caching one by one can adapt to different sizes of memory entries. After this round of retrieval, the temporary cache is automatically cleared, ensuring the rational use of system memory resources and preventing historical cache residues from multiple retrieval tasks.
[0035] Optionally, the step of dividing the scheduling state based on the comprehensive retrieval priority, calculating the retrieval delay, and filtering entries to generate a retrieval execution queue specifically includes: the system reads a pre-set first priority threshold and a second priority threshold from the global configuration partition, and simultaneously performs compliance checks on the two threshold parameters to confirm that the parameters are within the system's allowed value range. After completing the parameter verification, the system reads the memory comprehensive retrieval priority data stored in the temporary cache area one by one, and sequentially performs the scheduling state determination work. When the comprehensive retrieval priority of a single memory is greater than or equal to the first threshold, the memory is marked as being in a normal retrieval state; when the comprehensive retrieval priority is less than the first threshold but greater than or equal to the second threshold, the memory is marked as being in a gradient degradation state; when the comprehensive retrieval priority is less than the second threshold, the memory is marked as being in a pending archiving state. Each scheduling state corresponds to a set of fixed-format identifier fields, and the state identifier and the unique identifier of the memory are bound together and stored uniformly.
[0036] After completing the status marking of all memories, the system matches the corresponding retrieval delay calculation logic based on the status identifier of each memory, and calls the dedicated calculation formula to solve the retrieval delay duration of each memory one by one. After all delay calculations are completed, the system performs a retrieval item filtering process, directly removing all memory items marked as pending archiving, and retaining only two types of items: those in normal retrieval status and those in gradient degradation status, for subsequent retrieval responses. The valid memory items after filtering are strictly sorted in descending order of comprehensive retrieval priority. After sorting, the overall queue order is locked, and the queue order will not be adjusted throughout the entire processing cycle of this retrieval task, ultimately forming the formal retrieval execution queue. The system outputs the corresponding memory retrieval content sequentially according to the retrieval delay duration of each memory item in the queue, completing the entire retrieval response.
[0037] The mathematical expression for calculating the retrieval delay is as follows:
[0038] In the formula, For the retrieval delay of a single memory, The system's preset base delay parameters, This is the overall retrieval priority value for this memory. The minimum delay parameter preset by the system. It is a natural constant.
[0039] The formula for calculating the delay attenuation coefficient is: In the formula Represents the delay attenuation coefficient. This is a numerical value representing the overall retrieval priority, used to control the degree to which retrieval latency decreases with changing priority. This formula ensures that retrieval latency is maintained regardless of priority. Take any positive real value, retrieval delay Always fall Within the range, and with a higher priority value, the retrieval latency decreases strictly and monotonically, enabling faster response from high-priority memory.
[0040] For example, in the execution stages of scheduling status determination, delay calculation, and queue generation, two-level thresholds serve as the dividing criteria for state division. These thresholds remain fixed during system operation and only change when configuration rules are manually modified, ensuring a unified and stable scheduling standard. Three scheduling states correspond to different activity levels of the memory: the normal retrieval state represents memory with high recent access frequency and strong activity, making it the primary target for retrieval responses; the gradient degradation state represents memory activity levels that have decreased but still have retrieval value, requiring an appropriate increase in response latency; and the pending archiving state represents memory activity levels that are extremely low, temporarily ceasing participation in the current retrieval response. The calculation logic for retrieval latency is inversely related to priority; higher-priority memories have shorter retrieval waiting times and can output results first, aligning with the scheduling logic of prioritizing high-value, high-activity memories. The item filtering stage directly removes pending archiving items, reducing invalid result output and lowering system computation and data transmission pressure. After queue sorting, a locking operation is performed to prevent order chaos during retrieval, ensuring a smooth and orderly retrieval output rhythm. This entire logic can adapt to various business scenarios such as concurrent retrieval and batch retrieval.
[0041] Optionally, the steps of matching depth factors, updating stability parameters, and synchronizing timestamps after a successful access specifically include: after the retrieval content of a single memory is output and the system verifies that the access behavior was successfully executed, firstly, extracting the initiating entity identifier of this access behavior, and combining the entity identifier with the currently marked scheduling state of the memory to match the depth factor required for this operation. If the memory currently has a gradient degradation state identifier, then the type of the initiating entity is no longer distinguished, and the fixed depth factor corresponding to the gradient degradation state is directly selected for operation; if the memory does not have a gradient degradation state identifier, then the access scenario is distinguished according to the initiating entity: when the initiating entity identifier is a user-side identifier, it is determined to be a user-initiated access scenario, and the depth factor specific to this scenario is matched; when the initiating entity identifier is a system server-side identifier, it is determined to be a system callback access scenario, and the depth factor specific to this scenario is matched.
[0042] After determining the depth factor, the system combines the stability parameters currently in use before the memory update with the system's preset learning rate, and substitutes these values into the iterative formula to calculate the original values of the updated stability parameters. After the iterative calculation is complete, the system immediately initiates a parameter boundary verification process. It reads the upper and lower threshold values of the stability parameters from the global configuration partition and compares the newly generated stability parameter values with these two thresholds one by one. If the iterated value is greater than the upper threshold, the upper threshold is used as the final updated value for the memory's stability parameter; if the iterated value is less than the lower threshold, the lower threshold is used as the final updated value; if the iterated value is between the upper and lower thresholds, the original value obtained from the iterative calculation is directly retained. The final determined parameter value overwrites the original stability parameters of the memory, and the last access timestamp of the memory is updated to the system timestamp of the current access.
[0043] The mathematical expression for calculating the updated stability parameter is as follows:
[0044] In the formula, These are the stability parameters that were not truncated after iterative calculation. To remember the current stability parameters before the update, The system's preset learning rate, This represents the depth factor obtained from the current match. The formula for calculating the gain adjustment term in this formula is: In the formula Representative parameter: gain adjustment term The system's preset learning rate, The depth factor obtained from the current match is used to control the gain of a single access behavior on the memory stability parameter.
[0045] For example, in the execution phase of stability parameter update and timestamp synchronization, successful access is used as the initiation condition for parameter update. Subsequent calculations are only triggered after the retrieval results are output normally and the access behavior is fully executed, thus avoiding invalid accesses from interfering with memory attribute parameters. The depth factor plays a role in regulating the magnitude of memory reinforcement. Different access scenarios correspond to different reinforcement strengths. The reinforcement effect is stronger for manual active access, while the reinforcement effect is relatively gentle for regular calls from the backend system. For memories in a gradient degradation state, a fixed depth factor is used uniformly to ensure that the parameter change rhythm of such memories remains consistent. The learning rate is a globally unified rule that controls the magnitude of parameter changes brought about by each access, preventing drastic parameter fluctuations caused by a single access. The parameter boundary truncation mechanism ensures compliant parameter operation, limiting the stability parameter to a fixed range. This prevents the parameter from increasing infinitely and causing the memory weakening rate to completely stagnate, and also prevents the parameter from decreasing infinitely and causing the memory to fail rapidly. The synchronous update of the last access timestamp accurately records the time node of each valid access, providing accurate time basis for the next round of recall and priority calculation, forming a linkage mechanism between parameter iteration and time recording.
[0046] Optionally, the process of performing memory archiving, offline memory migration, and observation period control according to the scanning cycle specifically includes: when the system runtime reaches a preset scanning cycle node, an active retrieval queue global scan task is automatically triggered, covering all stored memory entries within the active queue, performing a check on each entry without omission. The system reads the comprehensive retrieval priority calculated for the current round of a single memory entry, and simultaneously retrieves the priority historical data retained during the previous global scans of that memory, comprehensively determining whether the priority value of that memory has been lower than the system's preset archiving threshold for multiple consecutive rounds. If the condition of failing to meet the threshold for multiple consecutive rounds is met, the memory is determined to be long-term low-activity data, and an archiving operation is performed: the memory entry is removed from the storage partition corresponding to the active retrieval queue, and all data content of the memory, including attribute fields, historical access records, and status identifiers, is completely copied and written to the offline repository as is. The entire data migration process does not modify or delete any original data of the memory.
[0047] When the system detects that a memory in the offline repository has been accessed, it automatically initiates a migration process, copying and migrating all data of that memory from the offline repository back to the active retrieval queue. Simultaneously, an observation period status flag is added to that memory. Memoryes with an observation period flag no longer follow the system's general scheduling rules; instead, they are subject to independently configured scheduling logic, which restricts their participation in priority calculation, queue sorting, and retrieval response. When the preset observation period ends, the system automatically clears the observation period status flag, retrieves all access records generated by that memory during the entire observation period, recalculates the comprehensive retrieval priority based on these records, and reclassifies the scheduling status using a two-level priority threshold. This allows the memory to officially join the system's regular retrieval scheduling system and participate in all subsequent retrieval processes according to general rules.
[0048] For example, in the execution phases of memory archiving, migration, and observation period management, periodic global scanning is the core step in achieving hot and cold data separation. A fixed scanning interval ensures the system regularly reviews the active queue and promptly identifies inactive memories that have been inaccessible for extended periods. Using "continuous rounds of priority below the archiving threshold" as the archiving criterion avoids erroneous archiving caused by instantaneous priority fluctuations, improving the accuracy of archiving operations. When migrating memories to the offline repository, a complete copy method is used to ensure that the original data is not lost or tampered with, and all information in the memory can be completely restored during subsequent migration. An automatic migration mechanism when offline memories are accessed ensures that historical archived data can be retrieved and used normally, without becoming inaccessible due to being moved to offline storage. The observation period setting is mainly used to manage newly migrated memories. These memories have been offline for a long time, and their data activity and attribute parameters differ from regular active memories. Using dedicated scheduling rules to restrict them in the short term prevents them from suddenly joining the regular queue and disrupting the overall retrieval scheduling order. After the observation period ends, the system re-determines the status based on the access behavior during the period, allowing the migration memory to gradually adapt to the normal operating logic and achieve a smooth transition of bidirectional flow of hot and cold data.
[0049] A dynamic retrieval and scheduling system for proxy memory based on an exponential decay model.
[0050] Reference Figure 2 The present invention also provides a proxy memory dynamic retrieval scheduling system based on an exponential decay model. The system is used to execute the proxy memory dynamic retrieval scheduling method based on an exponential decay model described above. The system includes: a data acquisition and storage module, a priority calculation module, a retrieval scheduling module, a parameter update module, and a memory archiving and migration module.
[0051] The data acquisition and storage module is used to collect and store text content, category tags, creation time, initial stability parameters, and initial static weights. It also collects user-initiated access records, system callback access records, and timestamps corresponding to various access records. It loads all global configuration rules, including weight grading standards, stability parameter boundaries, latency parameters, priority thresholds, archiving thresholds, learning rates, depth factors, and scanning cycles. It performs standardized processing and legality verification on all collected data and configuration rules. It builds and maintains three types of storage areas according to functional partitions: a global configuration partition, an active retrieval queue storage partition, and an offline repository, completing the partitioned persistent storage of all data and rules.
[0052] The priority calculation module receives externally initiated retrieval requests and performs request validity checks. After the checks pass, it reads all memory entries in the active retrieval queue. It calculates the time interval between the last access to each memory, solves the real-time recall rate based on the exponential decay forgetting model, and calculates the comprehensive retrieval priority by combining the initial static weights. It binds the calculation result of each memory with the memory's unique identifier and stores it in a temporary cache area dedicated to this round of retrieval. After the entire retrieval process is completed, all data in the temporary cache is automatically cleared.
[0053] The retrieval scheduling module reads two-level priority thresholds from the global configuration partition and performs parameter verification. Based on the numerical value of the comprehensive retrieval priority, it divides memories into three categories: normal retrieval state, gradient degradation state, and pending archiving state, and binds a corresponding status identifier to each memory. Combining the memory status identifier with the delay calculation rules, it calculates the retrieval delay time for each memory. It removes memory entries marked as pending archiving, arranges the remaining valid entries in descending order according to the comprehensive retrieval priority, generates a retrieval execution queue, and locks the queue order. According to the retrieval delay corresponding to each entry, it outputs the memory retrieval results sequentially.
[0054] The parameter update module is used to identify the initiator of the current access and the current scheduling status of the memory after a successful memory retrieval access, and match the corresponding type of depth factor; combine the original stability parameters of the memory, the system learning rate, and the matched depth factor to iteratively calculate new stability parameters; retrieve the upper and lower limits of the stability parameters, perform truncation processing on the iteration results to ensure that the parameters fall within the legal value range; update the original stability parameters of the memory with the final parameters, and replace the timestamp of the last access of the memory with the system timestamp of the current access.
[0055] The memory archiving and migration module is used to periodically perform full-entry traversal checks on the active retrieval queue according to the system's preset scanning cycle. Combining the priority historical data generated by multiple rounds of scanning of a single memory, it determines whether the entry meets the archiving conditions and migrates qualified memories to the offline repository for archiving. It also monitors the access behavior of the offline repository in real time. When an offline memory is accessed, it completes the data migration and adds an observation period mark to the memory, enabling dedicated scheduling rules for memories within the observation period. After the observation period ends, the status mark is automatically cleared, and the scheduling status is re-determined based on the access records during the observation period, allowing the memory to be integrated into the system's regular retrieval scheduling process.
[0056] It should be noted that the interconnections between the various modules described above do not necessarily represent direct or indirect connections. Any indirect connection method is applicable to the embodiments of this invention as long as it achieves the purpose of this invention. The above descriptions are merely exemplary embodiments of this invention and should not be construed as limiting the scope of the invention.
[0057] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A proxy memory dynamic retrieval and scheduling method based on an exponential decay model, characterized in that, The specific steps of this method are as follows: S1 collects the memorized text content, category labels, creation time, initial stability parameters and initial static weights, collects user active access and system callback access records and timestamps, loads weight grading standards, stability parameter boundaries, delay parameters, priority thresholds, archiving thresholds, learning rates, depth factors, fixed depth factors and scanning cycles corresponding to gradient degradation states, and stores them in partitions after standardization. S2, based on the data stored in S1, calculates the time interval between the current time and the last access timestamp for the memories in the active retrieval queue, uses the exponential decay forgetting model, calculates the real-time recall based on the time interval and the current stability parameter, and then combines the real-time recall with the static weight to calculate the comprehensive retrieval priority value. S3: Based on the comprehensive retrieval priority calculated in S2, the comprehensive retrieval priority is compared with the first threshold and the second threshold to determine the scheduling state. The retrieval delay is calculated based on the scheduling state and the delay parameter. Memory not lower than the second threshold is formed into an execution queue in descending order of priority, and the retrieval results are returned in sequence according to the delay. S4. For the memory of successful access in S3, determine the depth factor based on the access scenario, calculate the updated stability parameter according to the depth factor, learning rate and current stability parameter, perform truncation processing when the updated stability parameter exceeds the boundary, and update the last access timestamp. S5, traverse the active retrieval queue according to the scanning cycle, and use the calculation method of S2 to obtain the comprehensive retrieval priority of each memory. Memory whose priority is lower than the archiving threshold for multiple consecutive rounds is moved to the offline storage repository. When it is detected that a user actively accesses the archived memory, the archived memory is moved back to the active retrieval queue and an observation period is started. During the observation period, scheduling is restricted.
2. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, The memory text content collected in S1 includes the complete original interactive text at the time of memory generation and the associated context identifier; the category label includes the memory business category identifier and the access permission identifier; the creation time is the system standard time at the moment of memory generation; the initial stability parameter and the initial static weight are the initial attribute values assigned after matching the category label at the moment of memory generation; the collected user-initiated access records include the access initiator user identifier and the access-triggered interaction event identifier; the system callback access records include the access-triggered system task identifier and the calling module identifier; the timestamps corresponding to all access records are the system standard timestamps at the moment of access, and the timestamps are bound and stored one by one with the corresponding access records.
3. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, The weight grading standard loaded in S1 is a pre-configured mapping rule between memory classification labels and static weight levels; the stability parameter boundary is a pre-configured upper and lower limit rule for the legal values of the stability parameter; the delay parameter is a pre-configured baseline value rule for calculating retrieval delay; the priority threshold is a pre-configured two-level judgment rule for memory scheduling state division; the learning rate is a pre-configured amplitude rule for the iterative update of the stability parameter; the depth factor is a pre-configured amplitude rule for memory reinforcement corresponding to different access scenarios; the scanning cycle is a pre-configured interval rule for full memory global scanning; all configuration parameters are fixed rules that are globally effective and are uniformly stored in the global configuration partition after loading.
4. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, In S2, the system first verifies the validity of the received retrieval request. After the verification is successful, it retrieves a list of all memory entries currently in the active retrieval queue and performs calculations on each memory entry in the order they were entered into the database. First, it calculates the time interval by the difference between the last access timestamp of a single memory and the current system standard time. Then, it calculates the real-time recall based on the time interval and the current stability parameters of the corresponding memory. Finally, it calculates the comprehensive retrieval priority by combining the real-time recall with the initial static weight of the corresponding memory. The calculation result of a single memory is bound to the unique identifier of the corresponding memory and temporarily stored in a temporary cache. The data in the temporary cache is only valid within the processing cycle of this retrieval request and is automatically cleared after the current retrieval is completed.
5. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, In S3, the system first reads the first threshold and the second threshold from the global configuration partition and performs parameter validity verification. After the verification is passed, it reads the comprehensive retrieval priority stored in the temporary cache one by one. The comprehensive retrieval priority is first compared with the first threshold. If the priority is not lower than the first threshold, it is marked as a normal retrieval state. If the priority is lower than the first threshold, it is compared with the second threshold. If the priority is not lower than the second threshold, it is marked as a gradient degradation state. If the priority is lower than the second threshold, it is marked as a pending archiving state. Each state is marked as an identifier field with a fixed format, which is bound to the corresponding unique identifier stored.
6. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, In S3, the system matches the corresponding delay calculation rule according to the state flag bound to each memory, and performs queue admission screening after completing the delay value calculation of all memories; only memory entries with normal retrieval state flags and gradient degradation state flags are retained, and memory entries with pending archiving state flags are removed; after screening, the memory entries are sorted according to a fixed rule from high to low comprehensive retrieval priority value, and a retrieval execution queue is generated and the queue order is locked. The queue order will not be adjusted during this retrieval processing cycle.
7. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, In S4, after the system completes the output of the retrieval result for a single memory and confirms the successful access, it retrieves the initiator identifier of this access behavior. If the initiator identifier is a user identifier, it is determined to be a user-initiated access scenario, and the depth factor corresponding to the user access scenario is matched. If the initiator identifier is a system server identifier, it is determined to be a system callback access scenario, and the depth factor corresponding to the system access scenario is matched. If the memory being accessed has a gradient degradation state marker, the fixed depth factor corresponding to the gradient degradation state is matched first, and the scenario is no longer determined based on the initiator identifier.
8. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, After the system completes the iterative calculation of stability parameters in S4, the parameter boundary verification process is immediately triggered. The upper and lower threshold values of the stability parameter are read from the global configuration partition. The new stability parameter obtained by iterative calculation is compared with the upper and lower threshold values respectively. If the new stability parameter value is greater than the upper threshold, the upper threshold is taken as the final update value. If the new stability parameter value is less than the lower threshold, the lower threshold is taken as the final update value. If the new stability parameter value is within the upper and lower limit range, the iteratively calculated value is retained. The final update value synchronously overwrites the corresponding original stored stability parameter field.
9. The proxy memory dynamic retrieval and scheduling method based on the exponential decay model according to claim 1, characterized in that, In step S5, when the system reaches the trigger node of the preset scanning cycle, it automatically starts a global scan, and the scanning traversal range covers all memory entries in the currently active retrieval queue; it pulls the current comprehensive retrieval priority of the corresponding memory and the priority history of previous global scans one by one; if the priority results of a single memory in N consecutive rounds of scanning are all lower than the archiving threshold, the memory entry is removed from the storage partition of the active retrieval queue, and all attribute fields, historical access records and all status flags of the memory are completely copied and written to the offline repository. The writing process does not modify any original data content; where N is a preset positive integer and N≥3.
10. A proxy memory dynamic retrieval and scheduling system based on an exponential decay model, the system being applicable to the proxy memory dynamic retrieval and scheduling method based on an exponential decay model as described in any one of claims 1-9, characterized in that, It includes a data acquisition and storage module, a priority calculation module, a retrieval and scheduling module, a parameter update module, and a memory, archiving, and rollback module; The data acquisition and storage module is used to collect and remember text content, category tags, creation time, initial stability parameters, initial static weights, user active access records, system callback access records and corresponding timestamps, load various global configuration parameters, standardize all data and store it in partitions, and maintain global configuration partitions, active retrieval queue storage partitions and offline storage partitions. The priority calculation module is used to receive retrieval requests and complete legality verification. Based on the last access time interval of the memory and the current stability parameters, it calculates the real-time recall rate using the exponential decay forgetting model, calculates the comprehensive retrieval priority by combining the initial static weights, and stores the calculation results in a temporary cache after binding them with the unique identifier of the memory. The retrieval scheduling module is used to read two-level priority thresholds, divide the normal retrieval state, gradient degradation state, and pending archiving state according to the comprehensive retrieval priority, calculate the retrieval delay based on the scheduling state and delay parameters, filter valid memory entries and generate a retrieval execution queue in descending order of priority, and output the retrieval results in sequence according to the retrieval delay. The parameter update module is used to match the corresponding depth factor according to the access initiator or memory state after successful memory access, iteratively calculate the new stability parameter in combination with the learning rate, and perform truncation processing according to the upper and lower limits of the stability parameter, and synchronously update the memory's stability parameter and the last access timestamp. The memory archiving and migration module is used to traverse the active retrieval queue according to a preset scanning cycle, determine whether the memory meets the archiving conditions, and perform offline migration. When access to a memory in the offline repository is detected, the memory is migrated back to the active retrieval queue and marked with an observation period. A dedicated scheduling rule is executed on the memory during the observation period. After the observation period ends, the mark is cleared and the memory is matched with the regular scheduling status. The data acquisition and storage module corresponds to step S1 of claim 1; the priority calculation module corresponds to step S2; the retrieval and scheduling module corresponds to step S3; the parameter update module corresponds to step S4; and the memory archiving and migration module corresponds to step S5.