Hierarchical indexing and parallel retrieval method for agent dialogue memory
Patent Information
- Application Number
- CN202611318390.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的一个目的在于提出一种智能体对话记忆的分层索引及并行检索方法,针对现有技术中不同时间尺度和访问频率的记忆共用存储与索引策略、迁移并发时可见性不一致以及跨层追加检索缺少定向边界的问题,提出以统一记忆单元为基础,根据驻留效用进行热层、温层和冷层动态分层并配置异构索引,通过路标、落位凭证和索引世代形成迁移可见状态,以单次世代快照执行预算化跨层并行检索,并对异构候选进行分位校准、证据融合、上界终止和迁移重复项折叠的技术方案,使检索访问集中于与查询和迁移路径相关的有限索引区域,并保持迁移并发下的候选一致性和可追溯性
[0044]1、通过依据语义重要性、时间和访问统计形成驻留效用,并为热层、温层和冷层配置不同的索引结构,使高频精确访问、语义相似访问和长期归档分别使用相适配的检索路径,减少对全部历史记忆的遍历。
Smart Images

Figure CN122838588A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer information storage and retrieval technology, and in particular to a hierarchical indexing and parallel retrieval method for intelligent agent dialogue memory. Background Technology
[0002] As conversational agents, code-based agents, and tool-invoking agents continue to run, dialogue text, code snippets, tool results, and runtime logs accumulate. Current technologies typically store this information uniformly in a database or vector library, or partition it only by data type, employing the same or similar indexing and retrieval strategies. When the scale of historical data expands, high-frequency short-term memory, mid-term task knowledge, and low-frequency long-term records collectively consume storage and indexing resources. During retrieval, a large amount of data unrelated to the current query may still be accessed, leading to increased memory usage, storage costs, and retrieval latency.
[0003] On the other hand, semantically similar memories across different sessions or tasks are difficult to reuse effectively, and a single index cannot simultaneously adapt to exact matching, semantically similar retrieval, and low-cost archiving. If memories are dynamically migrated between different storage layers, asynchronous updates of the source and target indexes may cause missed or duplicate recalls. Without cross-layer directional association, parallel retrieval, and comparable result fusion mechanisms, append retrieval can easily degenerate into aimless, slow layer scanning, thereby increasing tail latency and affecting the stability of agent response and decision-making.
[0004] Therefore, there is a need for a hierarchical indexing and parallel retrieval method for agent dialogue memory that can overcome the shortcomings of the existing technologies. Summary of the Invention
[0005] One objective of this invention is to propose a hierarchical indexing and parallel retrieval method for intelligent agent dialogue memory. Addressing the problems of shared storage and indexing strategies for memories with different time scales and access frequencies, inconsistent visibility during migration concurrency, and lack of directional boundaries in cross-layer append retrieval in existing technologies, this invention proposes a method based on unified memory units. It dynamically hierarchizes memory into hot, warm, and cold layers according to residency utility and configures heterogeneous indexes. Migration visibility is established through roadmaps, location credentials, and index generations. Budgeted cross-layer parallel retrieval is performed using single-generation snapshots. Furthermore, the invention employs quantile calibration, evidence fusion, upper bound termination, and migration duplicate folding techniques for heterogeneous candidates. This ensures that retrieval access is concentrated in a limited index area related to the query and migration path, while maintaining candidate consistency and traceability under migration concurrency.
[0006] This invention provides a hierarchical indexing and parallel retrieval method for intelligent agent dialogue memory, including:
[0007] S1. Obtain the memory data generated by the agent and encapsulate it into a memory unit containing the original memory identifier, content version, domain identifier, time information, access statistics and semantic representation;
[0008] S2. Determine the residency utility based on the semantic importance, time information and access statistics of the memory unit, place the memory unit into the hot layer, warm layer or cold layer according to the hierarchical rules with hysteresis interval, and establish a layer-differentiated index.
[0009] S3. In response to the change of the residence layer of the memory unit, a migration record is established with the original memory identifier, a road sign item pointing to the target retrieval area is written in the source layer, the target placement certificate is associated with the index generation and then submitted to form a migration visibility state determined by the index generation.
[0010] S4. Parse the memory query to obtain the search conditions, read the committed index generation once as a query snapshot, allocate candidate budgets to each level of index and search in parallel, consume only the text items and sign items visible in the query snapshot, access the target search area along the hit sign, and obtain a heterogeneous candidate set.
[0011] S5. Convert the candidate scores of each index branch to the quantile scale of the corresponding historical score distribution and merge them. Control the branch termination and candidate budget transfer based on the upper bound of the remaining candidate fusion scores of each index branch to obtain the fusion candidate set.
[0012] S6. Based on the original memory identifier, content version, and migration duplicates in the submitted index generation folding fusion candidate set associated with the candidate, select candidates according to the candidate fusion score and output the memory context.
[0013] Optionally, S1 includes: using any one of the following as the memory data: dialogue text, code snippets, tool call results, or runtime logs;
[0014] The domain identifier includes at least one of a session identifier and a task identifier; the time information includes the generation time and the most recent access time; and the access statistics include the number of accesses and the access interval.
[0015] Semantic encoding is performed on the memory data, and entity fingerprints or code symbol fingerprints are extracted. The resulting semantic representations and symbol fingerprints are associated with the original memory identifiers and content versions and written into the memory unit.
[0016] Optionally, S2 includes: weighting the semantic importance, the time decay determined by time information, and the access activity determined by access statistics, and correcting the weighted combination result in combination with the reconstruction cost of the memory unit to obtain the residency utility, and the larger the residency utility, the higher the priority of residing in the hot layer.
[0017] Set an inflow threshold for two adjacent layers to migrate into a hotter layer and an outflow threshold for migrate out of a hotter layer, such that the inflow threshold is greater than the outflow threshold, and keep the layer where the memory cell is located unchanged when the residency utility is equal to either threshold.
[0018] In the hot layer, establish an exact inverted index or hash index and a time ring index; in the warm layer, establish an approximate vector index and a task entity adjacency index; and in the cold layer, establish a summary prototype index, a time bucket index, and an archive address index.
[0019] Optionally, S3 includes: causing the migration record to sequentially go through the states of preparation, placement, submission, and recycling;
[0020] In the ready state, a fixed-length roadmap entry is generated, which records the source layer, target layer, old index key digest, new index key digest, target bucket identifier, migration generation, and status bit.
[0021] After the target text and target index are written, the placement certificate containing the target bucket identifier and the target index verification value is generated. When the placement certificate is generated, the migration record is switched to the placement state. The road sign item, placement certificate and index generation are atomically associated. After the migration generation is set to the associated index generation value, the record is switched to the committed state.
[0022] When the target text write fails, the target index write fails, the placement credential verification fails, or the commit is interrupted before the commit state, the migration record is maintained or rolled back to the ready state and the source text is kept visible to the query snapshot. The commit is then retried after the text is rewritten and the placement credential verification is passed.
[0023] If there is no active query snapshot with a committed index generation smaller than the migration generation, the source text is reclaimed and switched to the reclaimed state. Before this, the source text is retained for the active query snapshot to read, and the signpost item is retained for a preset retention period.
[0024] Furthermore, the visibility of the query snapshot to migration records is determined by comparing the migration generation with the committed index generation in the query snapshot;
[0025] Make the target text item or the landmark item visible to the query snapshot when the migration generation is not greater than the committed index generation in the query snapshot.
[0026] When the migration generation is greater than the committed index generation in the query snapshot, or when the migration record has not yet formed a committed migration generation, the source text item becomes visible to the query snapshot and the target text item and target index are not included in the candidate list.
[0027] When both the source text and the target text corresponding to the same original memory identifier are visible to the query snapshot, both are sent as candidates for folding, and their respective source layer, content version, migration generation as the submission generation, and migration path are retained as the basis for folding.
[0028] Optionally, S4 includes: decomposing the search conditions into precise conditions, semantic conditions, time conditions and domain conditions, and selecting corresponding index branches for each, using permission conditions as new filtering conditions, and allocating the candidate budget to each index branch according to the query type, historical hit rate of each layer and access cost;
[0029] For migration records in the preparation state, consume source text items; for migration records in the settled but not committed state, only consume source text items visible in the query snapshot, and make the target text and target index only used for settlement credential verification and not enter the candidate; for migration records in the committed or recycled state, consume target text items or access the target bucket or summary prototype range defined by the target bucket identifier along the road sign items.
[0030] Furthermore, the permission conditions include the access subject identifier, the set of allowed domain identifiers, and the memory visibility level;
[0031] Before the index branch returns to the candidate, the index entries and roadmap entries are filtered according to the permission conditions, and the target text accessed along the roadmap entries continues to be subject to the same or less restrictive permission conditions as the source text as access security constraints.
[0032] Optionally, S5 includes: maintaining the historical score cumulative distribution according to the index branch and query type, and mapping the original score of each index branch to the quantile value in the historical score cumulative distribution;
[0033] The candidate fusion score is obtained by weighting and combining the quantile value, time evidence, importance evidence, access evidence, and migration path evidence indicating whether the candidate was obtained along the road sign item.
[0034] Based on the upper bound of the quantile values that have not yet returned candidates and the upper bound of the contribution that can be obtained by time evidence, importance evidence, access evidence and migration path evidence, the upper bound of the remaining candidate fusion score is determined according to the same weighted combination rule as the candidate fusion score, so that each index branch returns to the upper bound of the remaining candidate fusion score when returning candidates, and the upper bound of the remaining candidate fusion score is used as the control amount for branch termination and candidate budget transfer.
[0035] Furthermore, obtain the preset number of items parameter K, where K is a positive integer;
[0036] When the number of current valid candidates reaches K, the current valid candidates are sorted in descending order of candidate fusion score, and the index branch is terminated when the upper bound of the remaining candidate fusion score of an index branch does not reach the candidate fusion score of the Kth candidate.
[0037] Extract entity conditions from the precise conditions or semantic conditions, extract task conditions from the domain conditions, and determine that there is a gap in the corresponding evidence when no candidate satisfying the corresponding conditions is found in the current valid candidates for entity conditions, time conditions or task conditions.
[0038] The unused candidate budget of the index branch is transferred to the target bucket where the signpost has been hit and the evidence gap exists. The target bucket is a storage area defined by the target bucket identifier within the target retrieval area pointed to by the signpost. The target bucket identifier is used as a budget transfer constraint to prevent the unused candidate budget from being transferred to a target bucket not pointed to by the signpost.
[0039] Optionally, S6 includes: receiving the fusion candidate set and query snapshot, obtaining a preset number of items parameter K, where K is a positive integer, and setting the content version to a version number that increments with the content update of the same original memory identifier;
[0040] For candidates with the same original memory identifier and content version, sort them in descending order of submission generation and select the first candidate whose submission generation is not greater than the submitted index generation in the query snapshot. When multiple candidates have the same first submission generation, prioritize retaining the target text in the submitted state, and merge the hit index type and migration path of the remaining parallel candidates into the retained candidate.
[0041] For candidates with the same original memory identifier but different content versions, sort them in descending order of version number and retain the first candidate whose submitted generation is no greater than the submitted index generation in the query snapshot;
[0042] The candidates after folding are sorted according to their fusion scores, and the top K are selected. The source layer, content version, submission generation, hit index type, and migration path of the selected candidates are written into the traceable field and output along with the memory context.
[0043] The beneficial effects of this invention are:
[0044] 1. By forming residency utility based on semantic importance, time and access statistics, and configuring different index structures for hot, warm and cold layers, high-frequency precise access, semantically similar access and long-term archiving use appropriate retrieval paths respectively, reducing the traversal of the entire historical memory.
[0045] 2. By using migration records, source layer roadmaps, target location credentials, and index generations with the original memory identifier as the key, queries can determine the visibility of text items and roadmap items in a single committed generation, and fold migration double-write candidates according to the original identifier, version, and committed generation, reducing the risk of missed recall and duplicate recall during the migration process.
[0046] 3. By using cross-layer parallel candidate budgeting, heterogeneous score quantile calibration, upper bound of remaining candidate fusion scores, and transferring budgets only to target buckets that have roadmap hits and evidence gaps, the scope of slow-layer append access is limited and lossless early termination is supported, thereby reducing retrieval tail latency. Attached Figure Description
[0047] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0048] Figure 1 This is a flowchart of a hierarchical indexing and parallel retrieval method for intelligent agent dialogue memory according to the present invention.
[0049] Figure 2 This is a flowchart of the S5 heterogeneous score fusion and parallel retrieval termination method based on upper bound of the present invention. Detailed Implementation
[0050] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0051] refer to Figures 1-2 A hierarchical indexing and parallel retrieval method for agent dialogue memory, comprising:
[0052] S1. Obtain the memory data generated by the agent and encapsulate it into a memory unit containing the original memory identifier, content version, domain identifier, time information, access statistics and semantic representation;
[0053] S2. Determine the residency utility based on the semantic importance, time information and access statistics of the memory unit, place the memory unit into the hot layer, warm layer or cold layer according to the hierarchical rules with hysteresis interval, and establish a layer-differentiated index.
[0054] S3. In response to the change of the residence layer of the memory unit, a migration record is established with the original memory identifier, a road sign item pointing to the target retrieval area is written in the source layer, the target placement certificate is associated with the index generation and then submitted to form a migration visibility state determined by the index generation.
[0055] S4. Parse the memory query to obtain the search conditions, read the committed index generation once as a query snapshot, allocate candidate budgets to each level of index and search in parallel, consume only the text items and sign items visible in the query snapshot, access the target search area along the hit sign, and obtain a heterogeneous candidate set.
[0056] S5. Convert the candidate scores of each index branch to the quantile scale of the corresponding historical score distribution and merge them. Control the branch termination and candidate budget transfer based on the upper bound of the remaining candidate fusion scores of each index branch to obtain the fusion candidate set.
[0057] S6. Based on the original memory identifier, content version, and migration duplicates in the submitted index generation folding fusion candidate set associated with the candidate, select candidates according to the candidate fusion score and output the memory context.
[0058] In this specific embodiment, S1 includes:
[0059] This specific implementation uses a dialogue agent memory service running in a server cluster as the execution entity. The memory service receives dialogue text, code snippets, tool call results, and running logs from the session event bus. It writes the tenant number, session number, task number, event type, generation time, payload address, and access subject carried by each event into the access record, and verifies the event attribution using the tenant number and session number. When the task number is missing, the task number is set as a session-level reserved value and the event is not discarded. When the payload is empty or the integrity verification fails, the event is written to the isolation queue and no searchable memory unit is generated, so that subsequent processing uses memory data with clear source and complete content.
[0060] The memory service generates an original memory identifier for verified events. The original memory identifier is obtained by concatenating the tenant number, session number, event sequence number, and random salt and then performing a SHA-256 digest. When the same original memory identifier is first entered into the database, the content version is set to 1. Subsequently, any change in the text verification digest, text digest update, code snippet replacement, semantic encoder version change, or fingerprint extractor version change that causes any change in the text verification digest, text digest, semantic representation digest, or fingerprint set digest will increment the content version without changing the original memory identifier when the update transaction is successfully committed. The original payload is written to the text object database as an immutable object, so that multiple content updates of the same logical memory can be aggregated by the original memory identifier and distinguished by the content version.
[0061] The memory service establishes a memory unit record, with fields including original memory identifier, content version, session identifier, task identifier, memory type, generation time, last access time, number of accesses, access interval, text object address, text digest, semantic representation, entity fingerprint set, code symbol fingerprint set, visibility level, creation generation, and verification digest. The generation time is taken from the trusted timestamp of the event bus, the last access time is equal to the generation time when it is first written, the number of accesses is initialized to zero, and the access interval is initialized to a preset upper limit of 86,400 seconds. Each time a retrieved result is actually used, the last access time, the number of accesses, and the difference between two adjacent use times are written in an atomic update manner.
[0062] For dialogue text and tool call results, the preprocessor performs Unicode normalization, structured field expansion, sliding segmentation of more than 4096 lexical units, and 128 lexical unit overlap between adjacent segments. For code snippets, it retains language type, file path, function boundary, and line number. For runtime logs, it retains time, level, component name, trace identifier, and exception stack. The preprocessed text snippets are written to the snippet table with a composite key composed of the original memory identifier, content version, and snippet number. Code snippets lacking language type are entered into the general code lexer. Log payloads that cannot be parsed are saved as original text snippets and parsing status bits are set.
[0063] The semantic encoder employs a dual-encoding Transformer fine-tuned using samples from same-domain dialogue and code retrieval. Each text segment is encoded as a 768-dimensional vector and normalized using the L2 norm. The normalization relation is as follows: ,symbol Indicates the first The dimensionless semantic representation of a fragment, symbol Indicates the segment number, symbol The mapping function representing the semantic encoder, symbol Indicates the first A sequence of lexical segments, symbols The L2 norm of the encoding result, symbol This represents the dimensionless lower limit for preventing division by zero, and in this specific embodiment, it is set as... When the encoding result contains non-numerical components or its L2 norm is lower than the lower limit, the memory service marks the segment as to be re-encoded and retains only the exact index material;
[0064] The entity fingerprint extractor extracts normalized entity types, entity names, and aliases from dialogue text and tool results. It extracts package names, class names, function names, parameter lists, and called symbols from code snippets. It concatenates the entity type with the lowercase normalized name to calculate a 128-bit fingerprint. It concatenates the code language, qualified symbol name, and number of parameters to calculate a 128-bit code symbol fingerprint. When there is a conflict between different texts with the same fingerprint, a 16-bit check suffix is appended to the fingerprint record. The fingerprint, original surface form, fragment number, and character offset are written together into the fingerprint table, so that precise conditions can locate the specific evidence fragment inside the memory unit.
[0065] For memory data containing multiple text fragments, the memory service calculates the memory-level semantic representation using the number of effective lexical units in the fragments as the weight, and normalizes the resulting vector again according to the aforementioned L2 norm rule. At the same time, it writes the text summary, fragment number, encoder version, and fingerprint extractor version, which are no more than 512 bytes in length, into the memory unit. When the encoder version changes, a new content version is created and the old version is retained, instead of overwriting the semantic representation that has already been referenced by the index generation.
[0066] After encapsulation, the memory service writes the memory unit metadata, text object verification digest, and persistent outbound records within the same database transaction. The persistent outbound records include an idempotent key composed of the original memory identifier and content version, a random salt persisted when the original memory identifier was first generated and reused in subsequent retries, a domain identifier, generation time, last access time, number of accesses, access interval, semantic representation address, fingerprint set address, delivery status, and number of retries. After the transaction is committed, the outbound delivery device scans records awaiting delivery or those that have timed out, and sends them as encapsulation completion messages to the layered decision-making system. The system defines a queue and marks outbound records as confirmed only after receiving a persistent acknowledgment from the queue. When the process is interrupted between transaction commit and message enqueueing, the recovery scanner re-delivers the message using the same idempotent key. S2 first writes the idempotent key into the persistent receive table and rejects duplicate messages with a unique constraint, then determines the resident utility using the message and memory unit records. When a transaction fails, unreferenced temporary text objects are deleted and no deliverable outbound records are generated. When outbound records fail to retry consecutively, the memory unit is retained and the message is entered into the alarm queue, thus closing the recovery link between transaction commit, message enqueueing, and downstream idempotent consumption.
[0067] In this specific embodiment, S2 includes:
[0068] The layer manager receives the encapsulation completion message output by S1 and periodically scans memory units that need to be re-evaluated due to changes in access statistics. It normalizes semantic importance, time decay, access activity, and reconstruction cost and writes them into the residency evaluation record. This record uses the original memory identifier and content version as keys and also includes the current residency layer, last migration time, evaluation time, indicator source version, and current residency utility. The semantic importance inference features consist of the memory unit's one-hot encoding of the memory type, the logarithmic number of effective lexical terms in the text, the number of entity fingerprints, the number of code symbol fingerprints, whether it contains an exception stack, whether it contains a tool success result, the number of subsequent adoption confirmations within the task, and the text summary vector. The numerical features are scaled according to the feature scaling table released with the model. The offline labeled task completion contribution samples are used to train the gradient boosting tree. The inference output is truncated to 0 to 1, and the model identifier, model version, feature pattern version, inference time, and feature summary are written into the residency evaluation record. When the model file is missing, the feature pattern is incompatible, or the inference timeout occurs, the median importance of confirmed samples of the same task and memory type is used first. If there are no confirmed samples of the same task, the median importance of the same tenant and memory type in the last 30 days is used. If neither of these exists, the fixed prior value in the version-based residency strategy table is used. The fixed prior values for dialogue text, code snippets, tool results, and running logs are 0.50, 0.60, 0.55, and 0.40, respectively, so that the output is still determined in the cold start state.
[0069] The time decay is calculated based on the interval between the evaluation time and the most recent access time. ,symbol Indicates the first The dimensionless time decay of each memory cell, ranging from 0 to 1, with the sign... Indicates the memory cell number, symbol Indicates the number of seconds between the evaluation time and the most recent access time of that memory cell, denoted by the symbol. This represents the decay time constant configured according to memory type, and the unit is seconds. In this specific implementation, the dialogue text, code snippet, tool call result and running log are set to 604800 seconds, 2592000 seconds, 1209600 seconds and 172800 seconds respectively. When the timestamp goes backward, the time interval is set to zero, and the interval exceeding 180 days is truncated to 180 days.
[0070] The number of active visits is determined by the number of visits and the visit interval in the most recent 30 days. The hierarchical manager first performs logarithmic scaling on the number of visits, and then forms a reciprocal interval evidence based on a preset 3600-second visit interval baseline. The two items are then weighted and merged to 0 to 1. The reconstruction cost is obtained by dividing the number of bytes read from the main text object, the average time spent on semantic recoding, and the average time spent on index reconstruction by the corresponding deployment percentile upper limit and then weighting them. The deployment percentile upper limit is derived from the 95th percentile of the successful tasks in the most recent seven days and is released with the configuration version. If less than 100 samples are collected, the previous valid version is used.
[0071] Resident utility Calculation, symbol Indicates the first The dimensionless residency utility of each memory cell, with larger values indicating higher priority for residing in the thermal layer, symbol... Symbols representing semantic importance weights Indicates the first The dimensionless semantic importance of each memory unit, symbol Indicates the weight of visit activity, symbol Indicates the first The dimensionless access activity of a memory cell, symbol Represents the reconstruction cost weight, symbol Indicates the first The dimensionless reconstruction cost of a memory unit, symbolic Indicates the weight of time decay, symbol The aforementioned time decay amount is represented by the symbol. This represents a function that truncates the result to between 0 and 1. In this specific implementation, the four weights are 0.35, 0.25, 0.20, and 0.20 respectively and are stored in the residency policy table.
[0072] The residency policy table uses tenant level and memory type as lookup keys. The fields include four weights, hot / warm migration threshold, hot / warm migration threshold, warm / cold migration threshold, warm / cold migration threshold, minimum residency duration, policy version, and effective time. The table is generated by performing a grid search on samples of the hit rate, average read cost, and reconstruction time in the last 30 days. The release conditions are that the latency at the end of the verification set does not increase and the cache usage does not exceed the quota. In this specific implementation, a dialogue text record sets the hot / warm migration threshold to 0.72, the hot / warm migration threshold to 0.58, the warm / cold migration threshold to 0.46, the warm / cold migration threshold to 0.32, and the minimum residency duration to 1800 seconds.
[0073] The layer manager marks memory cells currently in the warm layer and whose dwell utility is greater than the hot-temperature migration threshold as migrating to the hot layer; it marks memory cells currently in the hot layer and whose dwell utility is less than the hot-temperature migration threshold as migrating out to the warm layer; it marks memory cells currently in the cold layer and whose dwell utility is greater than the warm-cold migration threshold as migrating to the warm layer; it marks memory cells currently in the warm layer and whose dwell utility is less than the warm-cold migration threshold as migrating out to the cold layer. When the dwell utility is equal to any threshold or the minimum dwell time since the last migration has not been reached, the original layer remains unchanged. If no policy record is found, the previous effective version is used and the status is set to pending review.
[0074] The hot layer stores memory unit metadata and high-frequency text fragments in an in-memory key-value library, builds an exact hash index based on entity fingerprint, code symbol fingerprint and domain identifier, and builds a 24-hour cyclic time ring index based on the generation time minute slot. The warm layer constructs an HNSW approximate vector index based on memory-level semantic representation and constructs a task entity adjacency index based on task identifier and entity fingerprint. The cold layer builds a summary prototype index based on the cluster center within the task, builds a time bucket index based on the generation month and domain identifier, and builds an archive address index based on the original memory identifier and content version.
[0075] Each index item at each layer uniformly includes the original memory identifier, content version, domain identifier, text object address, source layer, index key summary, index generation, visibility level, and status bit. When there is a collision in the hot layer hash bucket, it is verified by the complete fingerprint. The warm layer vector index first writes a tombstone for deleted items and rebuilds it when the proportion of tombstones reaches 10%. The cold layer summary prototype consists of up to 256 cluster centers per task and is recalculated every monthly archive. If any index construction fails, the original resident layer and its effective index remain unchanged, and the reason for the failure is written to the retry queue.
[0076] When the stratification decision result is different from the current resident stratum and all target index materials are ready, the stratification manager outputs a resident stratum change event. The event carries the original memory identifier, content version, source stratum, target stratum, old index key summary, new index key summary, suggested target bucket, and strategy version. S3 uses this to create a migration record. If the stratification decision result remains unchanged, only the resident evaluation record is updated without triggering a migration, thus forming a reproducible continuous processing chain between resident utility, hysteresis decision, and differentiated indexes.
[0077] In this specific embodiment, S3 includes:
[0078] The migration coordinator receives residency layer change events output by S2, acquires a mutex lease using the original memory identifier and content version, and creates a migration record in the migration log. The migration record fields include migration transaction identifier, original memory identifier, content version, source layer, target layer, source text address, target text address, old index key digest, new index key digest, target bucket identifier, migration generation, status bit, placement credential address, creation time, and number of retries. The initial state is set to ready. Before acquiring the record, it first cleans up duplicate placeholder records with expired leases and no status log updates according to the original memory identifier and content version, and retains the record with the highest status generation and a valid verification digest. If an incomplete record with the same key already exists, the record is reused instead of creating a new parallel migration.
[0079] In the preparation state, the migration coordinator determines the target retrieval region based on the bucketing rules of the target layer index. The hot layer uses the high-order bits of the entity fingerprint or code symbol fingerprint as the hash bucket; the warm layer uses vector entry point partitions and task entity adjacency partitions; and the cold layer uses a bucketing key composed of the task digest prototype and month. A fixed 128-byte roadmap entry is written to the source layer. This roadmap entry sequentially uses 1 byte layout version, 1 byte source layer, 1 byte target layer, 12 bytes old index key digest, 12 bytes new index key digest, 16 bytes target bucket identifier, 12 bytes original memory identifier digest, 8 bytes migration generation, and 8 bytes... The migration record contains the following data: content version, status bit, tenant and domain federated digest, visibility level, permission snapshot identifier, access control list digest, checksum, and reserved area. The digest field uses the truncated SHA-256 value of the corresponding complete field, and the migration record retains the complete value for secondary verification. The permission snapshot is stored in the protected migration permission table with the permission snapshot identifier and is retained for at least 24 hours after the road sign item expires. Access to the target text is prohibited along the road sign when the field value is inconsistent with the migration record or migration permission table, the complete digest verification fails, or the reference becomes invalid.
[0080] The migration coordinator copies or references the source text to the target text address, and then writes it to the target index entry according to the target layer rules. During the writing process, the migration transaction identifier is used as the idempotent key. After both the target text and the target index are persisted, the corresponding index entry in the target bucket is read and a checksum is calculated. A placement certificate containing the migration transaction identifier, target bucket identifier, target text checksum, target index checksum, number of index entries, and generation time is generated. Only when the two checksums in the certificate are consistent with the result of this write is the migration record switched from ready to placed.
[0081] The index generation service maintains a monotonically increasing committed generation counter. The migration coordinator requests the next generation for the settled records. In the same commit transaction, it writes the landmark item, the settled credential address, and the requested index generation into the migration record, sets the migration generation to the index generation, updates the effective generation of the target index item, and switches the status bit to committed. At the same time, it writes a persistent generation publication record containing the migration transaction identifier, the generation to be published, the publication status, and the commit log sequence number. After the transaction is committed, the generation publisher advances the global committed generations to a level not less than the generation to be published using a compare-and-swap method. After successful advancement, it marks the publication record as confirmed. If the process is interrupted after the migration transaction is committed but before the global generation is advanced, the recovery scanner replays the unconfirmed publication records according to the commit log sequence number. Repeatedly publishing the same generation does not change the result. If any field write fails within a transaction, no committed migration or generation to be published records are generated, thus allowing queries to determine whether the migration is complete with a single generation read.
[0082] When the target text write fails, the target index write fails, the placement credential verification fails, or the commit transaction is interrupted before the committed state is formed, the migration coordinator will record or roll back to the ready state, retain the visibility flags of the source text and source index, clear the target index candidate flags that have not been committed, and re-execute the target write and credential verification according to the migration transaction identifier. The retry interval adopts a backoff sequence of 1 second, 2 seconds, 4 seconds to a maximum of 60 seconds. After five consecutive failures, it will be transferred to the manual review queue but the source text will not be hidden.
[0083] After reading the globally committed generations, the query snapshot retains this value until the end of the query. For each migration record, if the migration record has not yet formed a committed migration generation or the migration generation is greater than the committed index generation in the query snapshot, the source text item is visible to the query snapshot while the target text item and target index do not enter the candidate. If the migration generation is not greater than the committed index generation in the query snapshot, the target text item and the roadmap item used for directional expansion are visible to the query snapshot. When the migration record is committed, the source index has not been reclaimed, and the double-read flag is 1, the source text item is allowed to enter the same query snapshot as a migration duplicate candidate. The double-read flag is set to 1 when the committed transaction is completed and set to 0 before the source index is reclaimed. The comparison rule is executed by the index scanner when reading the index item, rather than rereading the generation in the middle of the query.
[0084] The index scanner writes migration visibility as a Boolean criterion. and ,symbol The symbol indicates the visibility of the target text item to the query snapshot and takes the value 0 or 1. This indicates that the source text item is visible to the query snapshot and takes the value 0 or 1. Represents the migration generation of a migration record and is a non-negative integer, with sign... Indicates a committed index generation in the query snapshot and is a non-negative integer, with the sign... Indicates the migration state, symbol , , and The symbols represent the states of being ready, placed, submitted, and retrieved, respectively. The double-read flag indicates that the migration has been committed and the source index is still retained; it is set to 1 when the source index is reclaimed; and it is set to 0 before the source index begins to be reclaimed. This is an indicator function that takes 1 if the condition is true and 0 otherwise. When the migration generation has not yet been formed, it is treated as positive infinity to keep the source text visible.
[0085] Road sign visibility by Judgment, symbol This indicates that the road sign is visible and takes the value 0 or 1. Represents the migration generation of a migration record and is a non-negative integer, with sign... Indicates the committed index generation in the query snapshot and is a non-negative integer, with the sign... Indicates the migration state, symbol Indicates a committed status, symbol Indicates a recycled status, symbol Indicates the end time of the retention period for road sign items, in milliseconds, symbol Indicates the query start time in milliseconds, symbol This is an indicator function that takes 1 if the condition is true and 0 otherwise; if the query start time is not earlier than the retention period expiration time, it is prohibited to extend along the landmark and to retrieve the submitted target index by the original memory identifier; if the landmark checksum, permission snapshot identifier or access control list summary are inconsistent, the landmark is removed from the current candidate path and a reconstruction task is registered.
[0086] When both the source and target texts corresponding to the same original memory identifier are visible to the query snapshot, i.e., within the transition window where the migration record has been committed, the migration generation is no greater than the query snapshot generation, the double read flag is 1, and the source index has not yet been reclaimed, both are sent as candidates for folding. The migration coordinator marks the source text candidate as a migration duplicate candidate and the target text candidate as a canonical candidate. Both retain their respective source layer, content version, the migration generation as the committed generation, the migration transaction identifier, and the migration path as the basis for folding. The coordinator verifies the text digests of the two texts according to the original memory identifier and content version. If the digests are inconsistent, both candidates are retained and a content conflict flag is set instead of deduplicating directly at the physical address in the index layer. If the digests are consistent, they are still handed over to S6 for folding according to the canonical candidate priority rule.
[0087] The migration coordinator first deduplicates the query identifier, snapshot generation, and lease time in the active snapshot register, then deletes dirty records whose leases have expired and have no running heartbeats, and calculates the minimum active generation. ,symbol The smallest committed index generation represents the active query snapshot and is a non-negative integer. Represents the cleaned-up active query set, symbol Indicates the query sequence number, symbol Indicates the first The committed index generations of each active query are non-negative integers. When the active query set is empty, the minimum active generation is set to the current global committed generation plus 1.
[0088] The activity snapshot registration table records the query snapshot generation and lease expiration time using query identifiers. The reclaimer only allows setting the double-read flag to 0 in the same reclamation transaction, deleting the source text and source index, and switching the migration record to reclamation status when the publication record of the persistent generation in this migration has been confirmed as successfully published, the globally committed generation is not less than the generation in this migration, and the minimum active generation is not less than the generation in this migration. When the publication record is pending confirmation, publication fails, or the globally committed generation is still less than the generation in the migration, the source text, source index, and double-read flag are maintained regardless of whether the active query set is empty, and the recovery scanner continues to publish to avoid reclamation of the source end before the target is visible to any new snapshot. After reclamation, the landmark item still exists according to the 24-hour retention period set in this specific implementation. The corresponding permission snapshot in the migration permission table is retained for at least 24 hours after the landmark expires. It is only deleted when the landmark retention period expires and there is no active snapshot reference, so that late index branches can still complete targeted access along the existing target bucket identifier under permission constraints.
[0089] Each migration record is appended to a state log for every state change. The state log contains the previous state, the next state, the operator, the verification digest, the index generation, and the time. State transitions are only allowed from prepared to settled, settled to committed, committed to recycled, and, in case of failure, from settled back to prepared. After the state log is written, the verifier compares the previous state, the next state, and the generation value one by one. If an illegal transition or generation rollback is found, the record is frozen while keeping the last confirmed visible text. The record is frozen by the verifier reading the most recent valid settled credential and the source target text verification digest. If the two are inconsistent, the recycling step is frozen, the forced double-read flag is kept at 1, the source text remains visible, and the target text and target index are marked as content conflict isolation state, preventing them from being used as canonical candidate outputs. If the two are consistent, the target text remains visible but the recycling step is prohibited. After the commit is completed, a migration visibility event containing the target bucket identifier and migration generation is published to the query router for S4 to update the beacon access cache.
[0090] In this specific embodiment, S4 includes:
[0091] The query coordinator receives the memory query submitted by the agent, parses the query text, access subject identifier, session scope, task scope, time scope, number of returned items, and query type, organizes the quoted phrase, entity fingerprint, and code symbol fingerprint into precise conditions, organizes the semantic encoding vector into semantic conditions, organizes the start and end times into time conditions, organizes the session identifier and task identifier into domain conditions, and organizes the access subject identifier, the set of allowed domain identifiers, and memory visibility level into permission conditions. The permission conditions are used as new filtering conditions for each index branch. In this specific implementation, the number of returned items is 20 when there is a lack of returned items. The parsed query conditions, permission conditions, and number of returned items are used as inputs for subsequent query snapshots, branch budgets, and parallel retrieval.
[0092] Before initiating any index scan, the query coordinator reads the committed generation of the index generation service only once and forms a query snapshot. The query snapshot records the query identifier, the committed index generation, the start time, the permission condition summary, and the lease expiration time. During the entire query, all hot, warm, and cold branches carry the same snapshot value. When a branch times out and retryes, the original snapshot is still used instead of reading the new generation, thus avoiding different branches in a single query making different visibility judgments for the same migration.
[0093] The branch planner selects the hot layer hash index and inverted index based on the precision condition, the hot layer time ring index and cold layer time bucket index based on the time condition, the warm layer HNSW index and cold layer summary prototype index based on the semantic condition, and the warm layer task entity adjacency index based on the task condition. After the branch is determined, the number of enabled branches m is obtained. The query coordinator reads the oversampling coefficient and the upper limit of the candidate budget from the candidate budget strategy table with version. The oversampling coefficient is multiplied by the number of returned items K and rounded up. The result is then truncated to the upper limit of the candidate budget. If the result is lower than the number of enabled branches m, it is set to m, thus obtaining the integer candidate total budget B. Therefore, B is not less than m. If the upper limit of the candidate budget is less than m, a strategy anomaly is recorded and m is used to replace the upper limit. Subsequently, a branch task containing the branch identifier, index type, source layer, query type, initial budget, consumed budget, remaining budget, access cost estimate, and timeout period is created for each branch. In this specific implementation, the oversampling coefficients for factual precise query and semantic recall query are 4 and 8, respectively, and the upper limit of the candidate budget is 512. Independent tasks are executed in parallel by the thread pool.
[0094] Candidate total budget by Assignment, symbol Indicates the first The candidate budget obtained from each index branch is in items, and the symbol is... Indicates the branch number, symbol This indicates the total candidate budget for this query, determined by the aforementioned candidate budget strategy table and the number of branches enabled, and is not less than [amount missing]. Positive integers, sign Indicates the first The dimensionless historical hit rate of each branch after smoothing, strictly greater than 0 and not greater than 1, symbol. Indicates the first Each branch returns a normalized access cost for a valid candidate that is not less than the strict positive lower bound. ,symbol This represents the strict positive lower bound of the normalized access cost, and in this specific implementation, it is taken as 0.01. Indicates the branch number used for summation, symbol Indicates the number of index branches used this time, and is a positive integer, sign... and symbols They represent the first The smoothed historical hit rate of each branch being strictly positive and not less than The normalized access cost is such that the normalized denominator is strictly greater than 0; after rounding, the difference is made up by rounding down the decimal remainder of the ratio, and the budget and excess are calculated. Time only from Branches with a budget greater than 1 are deducted in ascending order of their ratio, until the budget for each branch is still not less than 1 and the total budget equals 1 / 2. ;
[0095] Historical hit rate and access cost are stored in a branch statistics table. This table uses query type, source layer, and index type as keys, and the fields include seven-day query count, effective hit count, average read bytes, 95th percentile latency, normalized access cost, sample version, and update time. The historical hit rate is smoothed by adding one, i.e., adding 1 to the effective hit count and dividing by adding 2 to the seven-day query count, ensuring that it remains strictly greater than 0 even in cases of no hits or no samples. The normalized denominator for read bytes and latency is adjusted when the 95th percentile of the corresponding full branch is below the strictly positive baseline. The strict positive baseline is used to fuse two normalization results with equal weight. When the fused result is less than 0.01, it is set to 0.01 as the normalized access cost, thus avoiding zero cost and zero denominator. For a task entity query record, the smoothed historical hit rate of the temperature layer adjacency index is recorded as 0.63 and the normalized access cost is recorded as 0.22. When the number of samples is less than 200, it is weighted with the strictly positive global prior according to the number of samples. When the table lookup fails, the same source layer statistics are used. If it is still missing, the hit rate of 0.5 and the access cost of 1.0 are used without canceling the branch.
[0096] When scanning index entries, each index branch first compares the effective generation of the index entry with the committed index generation in the query snapshot, and then filters according to permission conditions. The main text is only read when the domain identifier of the index entry belongs to the set of allowed domain identifiers, the visibility level is not higher than the authorized level of the accessing subject, and the accessing subject meets the tenant isolation rules. For signpost entries, the fixed layout version and checksum are first verified. Then, the immutable source permission snapshot is read from the migration permission table with the permission snapshot identifier, and the tenant and domain federated summary, visibility level, and access control list summary are checked. Subsequently, the intersection of the query permission conditions, the source permission snapshot, and the target main text permission conditions is taken. The target main text is only read when the intersection allows access. If the permission snapshot is missing, expired, or the summary is inconsistent, the signpost is rejected and the target main text cannot be accessed first. Entries that fail permission verification are neither returned as candidates nor consume the main text reading budget.
[0097] For migration records in the ready state, the branch only consumes source text items visible in the query snapshot. For migration records that have been settled but not committed, the target text and target index are only used to verify the settlement credentials and do not enter the candidate list. For migration records that have been committed or reclaimed, the branch consumes the target text items or the hit signpost items. The signpost items provide the target bucket identifier, target layer, new index key summary, tenant and domain federated summary, visibility level, permission snapshot identifier, and access control list summary. It does not trigger a full scan of the target layer. The query coordinator only limits the extension task to the target bucket or the corresponding summary prototype range after completing the aforementioned permission intersection verification.
[0098] Each branch returns a candidate record that uniformly includes the original memory identifier, content version, source layer, text object address, text verification summary, migration transaction identifier, source text verification summary, target text verification summary, index type, original score, matching fragment, domain identifier, generation time, last access time, access count, semantic importance, commit generation, migration path, target bucket identifier, content conflict flag, and permission verification summary. The original score of the precise index is the term or fingerprint matching ratio, the vector index is the cosine similarity, the time index is the time proximity, and the summary prototype index is the prototype similarity. Candidates that have not undergone migration have their migration transaction identifier and source and target text verification summaries set to empty. Candidates that have undergone migration use the migration transaction identifier as a strong consistent lookup reference to the S3 migration record and carry the source and target text verification summaries from that record. The content conflict flag is directly taken from the flag in the S3 migration record or the target isolated state index item and is prohibited from being rewritten by branches. Each original score maintains its own branch scale and is calibrated by S5.
[0099] The parallel executor sets up an independent read-only cursor and a bounded queue with a capacity twice the candidate budget for each branch. Candidates are enqueued one by one according to the internal priority of the branch. When the queue is full, the corresponding cursor is paused without discarding the verified candidate. The coordinator uses an idempotent key composed of the query identifier, branch identifier and candidate sequence number to eliminate duplicate returns caused by network retries, and aggregates the original score boundaries of candidates and remaining candidates every 20 milliseconds. Slow branches will not block other branches from submitting the candidates they have obtained.
[0100] When a branch exhausts its budget, reaches the unified query deadline, or exits abnormally, it closes the scanning cursor. When a branch receives a recoverable pause flag from S5, it does not close the cursor but sets it to a recoverable pause state, persistently retaining the next scan position, the next unreturned candidate original score boundary, the remaining budget, and the Kth bucket truncation score used during the pause. The branch, along with returned candidates, unused budget, cursor continue token, the next unreturned candidate original score boundary, and the set of hit landmark target buckets, is written into the heterogeneous candidate set. If a branch exits abnormally, the results of other branches are retained, and the unused budget of the abnormal branch is marked as transferable. The query coordinator merges the records of all normal branches, paused branches, and abnormal branches to obtain a heterogeneous candidate set containing the branch identifier, candidate record, unused budget, cursor continue token, the next unreturned candidate original score boundary, and the set of hit landmark target buckets. The heterogeneous candidate set and the same query snapshot are then given to S5 for quantification and fusion. When S5 issues a recovery flag, the query coordinator uses the same cursor continue token to recover the paused branch, without repeating consumed candidates.
[0101] In this specific embodiment, S5 includes:
[0102] The fusion coordinator receives the heterogeneous candidate set output by S4, and reads the historical score distribution records according to query type, source layer, and index type. The sampling object of the record is the candidate original score of the query completed in the last 30 days that has passed the permission verification and the text integrity verification. The sampling unit is one candidate. The record saves the original score histogram, sample number, quantile breakpoint, distribution version, and update time. The exact index, vector index, time index, task entity adjacency index, and summary prototype index maintain the distribution respectively. The original scores with different dimensions and monotonic directions are not directly added together. When the sample number is less than 500, the current histogram is smoothed by weighting the global distribution of the same type according to the sample number.
[0103] For index branches where higher scores are preferred, the fusion coordinator follows... Obtain the quantile value, sign Indicates the first The first index branch The candidate values are dimensionless quantiles in the historical distribution, ranging from 0 to 1, with the sign... Indicates the index branch number, symbol This indicates the candidate numbers returned in priority within this branch, symbol... Indicates the first Each index branch corresponds to the cumulative distribution function of the historical raw scores, with the sign... Indicates the first The first index branch For each candidate original score, the branch with the smaller distance or cost is better is subtracted from the corresponding cumulative distribution value to keep all percentile values so that the larger the value, the better the candidate.
[0104] The fusion coordinator calculates time evidence, importance evidence, access evidence, and migration path evidence from candidate records. Time evidence is obtained by dividing the interval from the query time to the generation time by the time scale corresponding to the query type and then exponentially decaying it. Importance evidence is obtained by directly using S2 normalized semantic importance. Access evidence is obtained by using the log-normalized value of the number of accesses in the last 30 days. Migration path evidence is set to 1 for direct text hits, 0.95 for candidates obtained along valid landmarks and whose credentials have been verified, and 0 for candidates that need to wait for permission review and are temporarily excluded from the valid set. All four types of evidence are truncated to 0 to 1.
[0105] Candidate fusion scores by Calculation, symbol Indicates the first There are 1 candidate dimensionless fusion scores, ranging from 0 to 1, with larger values indicating better candidates. Indicates the candidate fusion number, symbol Indicates quantile evidence weights, symbol Indicates the first 1 candidate dimensionless quantile value, symbol Indicates the weight of time evidence, symbol Indicates the first One candidate dimensionless time evidence, symbol The symbol represents the weight of material evidence. Indicates the first Evidence of the dimensionless importance of each candidate, symbol Indicates the access evidence weight, symbol Indicates the first One candidate dimensionless access evidence, symbol Indicates the evidence weight of migration path, symbol Indicates the first Evidence for one candidate dimensionless migration path;
[0106] Five fusion weights are stored in the fusion strategy table and satisfy the conditions that they are non-negative and their sum is 1. In this specific implementation, the weights are set to 0.45, 0.15, 0.15, 0.15 and 0.10 for factual accurate queries and 0.55, 0.10, 0.15, 0.10 and 0.10 for semantic recall queries. The fusion strategy table uses tenant level and query type as keys and also records time scale, sample version, effective time and rollback version. The weights are trained by pairwise sorting of the most recent 30 days of manual adoption and subsequent actual reference tags in the dialogue, and the release condition is that the cumulative gain of the first 20 items of the holdout set does not decrease. When the table lookup fails, the previous effective strategy is used.
[0107] When each index branch returns a candidate, it gives the upper bound of the quantile value of the next candidate that has not yet been returned based on the index sorting key of the unscanned region. The fusion coordinator then combines this upper bound with the upper bound of the contribution that can be obtained from time evidence, importance evidence, access evidence, and migration path evidence, according to the same fusion weight, to obtain the upper bound of the fusion score of the remaining candidates in that branch. When the target bucket time range, domain range, or landmark path type is known, the corresponding evidence upper bound is tightened. When a smaller upper bound cannot be proven, 1 is used as the evidence upper bound, thus ensuring that the upper bound does not underestimate the achievable fusion score of any candidate that has not yet been returned.
[0108] The preset number of items parameter K in the query request is a positive integer, and in this specific implementation, it is set to 20 by default. When receiving candidates, the fusion coordinator first establishes a final quota bucket according to the original memory identifier. Different content versions, different migration generations, migration double-write candidates, and cross-index candidates under the same original memory identifier each occupy only one counting position. Candidates with a content conflict flag of 1 are first excluded from the bucket. Then, the current representative candidate is selected in descending order of content version, descending order of visible commit generation, and descending order of fusion score. The bucket representative score is the fusion score of the representative candidate and all hit evidence is retained. The fusion coordinator distinguishes the branch status into normal scanning state, recoverable paused state, terminated preparation state, and confirmed closed state. It also distinguishes between budget exhaustion, scan completion, unified query deadline reached, and abnormal state. Exit is recorded as an unrecoverable hard boundary state; when the number of valid buckets with different original memory identifiers has not reached K items, no normal branch is paused due to the upper bound of the score. When the number reaches K items, the representative score of the Kth final quota bucket in the current sort is read as the truncation score. If the upper bound of the remaining candidate fusion score of a normal scan branch is less than the truncation score, a recoverable pause flag is sent to that branch and the cursor continue token is persistently retained. If the two are equal, at least one candidate with the same score is obtained and the boundary is resolved according to the lexicographical order of the original memory identifiers. The number of valid buckets and the score of the Kth bucket are recalculated after each addition of a candidate, cross-version replacement, permission elimination, or content conflict isolation. If the number of valid buckets drops below K, or the new score of the Kth bucket is lower than any recoverable... If the upper bound of the remaining candidate fusion score saved in the paused branch is not found, a recovery flag is sent to the corresponding branch and the scanning continues from its cursor with a continuation token. The fusion coordinator repeatedly executes receive, recalculate, pause, and resume. Only when there are no branches in the normal scanning state, the input queue and in-transit candidates are all empty, the number of effective buckets reaches K, and the remaining upper bound of each recoverable paused branch is strictly less than the current Kth bucket score, will each recoverable paused branch be converted to the termination preparation state and a query termination record with an integrity confirmation status of pending confirmation be generated. If all branches enter the unrecoverable hard boundary state and the number of effective buckets is still less than K, the fusion coordinator first waits for the input queue and in-transit candidates to be empty and checks the S5 budget transfer record. Only when there are no branches that meet the original query permission conditions, Exhaustion-type query termination records are generated only when there are road sign target bucket constraints, directional budget transfer tasks with remaining transferable budget greater than zero that have not yet been executed, or when all such tasks have been executed and their target branches have entered an unrecoverable hard boundary state. Their integrity confirmation status is set to pending confirmation, the hard boundary reasons for each branch and the current actual number of candidates are recorded, and the actual candidate set is used for post-integrity verification in S6. The termination preparation state retains the cursor continue token and is not considered irreversible termination. Only after S6 completes the post-integrity verification and returns a success flag for integrity confirmation will S5 switch the termination preparation state to the confirmed closed state and destroy the continue token. Only then is the upper bound termination irreversible. Branches in an unrecoverable hard boundary state are not reopened due to S6 feedback.Thus, high-content version replacement or post-completeness elimination enables the bucket representative to reopen the termination preparation branch when the score is reduced, and each original memory identifier still occupies only one return slot after S6 folding.
[0109] The fusion coordinator extracts entity conditions from precise or semantic conditions, task conditions from domain conditions, and builds an evidence coverage bitmap by combining time conditions. If there are no candidates that meet the entity conditions, time conditions, or task conditions among the current valid candidates, the corresponding bit is marked as an evidence gap. The budget transfer is only allowed when the terminated or abnormal branch has an unused budget, S4 records a hit signpost item, and the target bucket pointed to by the signpost item can provide the missing evidence type.
[0110] Budget transfer records include source branch, number of transfer items, evidence gap type, signpost item identifier, target bucket identifier, target layer, and expiration time. The query coordinator limits unused budgets to the storage area constrained by the target bucket identifier within the target retrieval area pointed to by the signpost item. The target branch scanning conditions must simultaneously meet the target bucket identifier and the original query permission conditions. Target buckets that are not hit by signposts cannot receive budgets even if their historical hit rate is greater than that of hit target buckets. The cumulative transferred budget for the same signpost target bucket shall not exceed one-quarter of the total candidate budget.
[0111] The historical score distribution builder samples from the query audit flow every hour. It first deduplicates by query identifier, branch identifier, and candidate sequence number, and removes traffic from permission failures, integrity failures, and manual testing. Then, it counts 4096 buckets in an equal-width initial interval and compresses them into 256 quantile breakpoints using an order-preserving merging method. When the original score is for similarity, the effective range is limited to -1 to 1. When the original score is for matching ratio or time proximity, the effective range is limited to 0 to 1. Samples outside the range are truncated and included in the anomaly count. When the anomaly ratio reaches 0.5%, the release of a new distribution version is paused.
[0112] Before the distribution version is released, the previous day's hold-out query is used for calibration and verification. It is required that the adjacent quantile breakpoints do not decrease, the quantile values are all between 0 and 1, and the candidate priority after mapping does not decrease when the original score of the same branch is monotonically improved. If the verification is successful, the current distribution version is replaced with an atomic pointer. If the verification fails, the previous version is used. The distribution version read by the fusion coordinator is fixed in one query and the version number is written into the fusion evidence record of each candidate.
[0113] The fusion coordinator takes the next unreturned candidate sorting key reported by the branch as the real-time measurement object. The measurement unit is the raw score of a candidate branch. The raw score is converted into the upper bound of the quantile value through the current fixed distribution version. At the same time, the target bucket time range, domain range and migration path type are read to calculate the upper bound of other evidence contributions. When the upper bound of the quantile or any evidence upper bound is less than 0 or greater than 1, it is truncated to 0 to 1 and the branch is recorded as abnormal. When the upper bound cannot be proved from the sorting key, 1 is directly used without prematurely terminating the branch.
[0114] New target bucket candidates still undergo quantile transformation and are merged with the same weights using the historical score distribution of their respective index branches. Then, the truncation score of the Kth item in the final bucket formed by the original memory identifier and the remaining upper bounds of each branch are recalculated. After all normal scan branches have entered the termination preparation state or the unrecoverable hard boundary state, and after generating normal query termination records or exhaustion query termination records, the fusion coordinator output includes the candidate fusion score, each evidence component, the hit index type, the source layer, the commit generation, the migration path, the target bucket identifier, the text verification summary, and the migration transaction identifier. The fusion candidate set consists of the source text validation digest, the target text validation digest, and the content conflict flag. The text validation digest, migration transaction flag, source text validation digest, target text validation digest, and content conflict flag all retain the S4 value for each candidate and do not participate in the score calculation. This set, the query snapshot, and the query termination record are then handed over to S6 to fold the migration duplicates. Only after receiving the integrity confirmation success flag returned by S6 will the fusion coordinator set the query termination record as confirmed and close and destroy the cursor continue tokens of each paused branch, making the upper bound termination irreversible.
[0115] In this specific embodiment, S6 includes:
[0116] The folder receives the fusion candidate set, query snapshot, and query termination records with integrity confirmation status pending from the output of S5. It reads the preset number of items parameter K and verifies that it is a positive integer. It verifies that each candidate carries the text verification summary and content conflict flag transmitted by S3 through S4 and S5 as is. For candidates that have undergone migration, it also verifies that they carry the migration transaction identifier and the text verification summary of the source and target ends. It performs a strong consistency lookup on the S3 migration record with the migration transaction identifier. If the lookup value is inconsistent or the reference is invalid, the candidate is moved into the integrity isolation set. Each candidate is organized into a folded record with the original memory identifier, content version, and commit generation as the core. The content version uses the version number incremented by the same original memory identifier each time the content is updated. The commit generation is the committed migration generation of the candidate's associated migration record. The initial text that has not undergone migration uses its creation generation as the commit generation. Candidates with a commit generation greater than the committed index generation in the query snapshot are discarded. If the content conflict flag is missing, the candidate is regarded as having incomplete integrity information and is moved into the isolation set, and is not treated as conflict-free.
[0117] Content version increment rules are written as ,symbol This indicates a new content version after a valid retrieval update for the same original memory identifier, and is a positive integer. The value represents the version of the content before the update and is a positive integer. The update of the effective retrieved content adopts the same criterion as S1, that is, when any one of the text verification summary, text summary, semantic representation summary or fingerprint set summary changes and the update transaction is successfully submitted, the increment is executed. Text error correction, summary update, code snippet replacement, semantic encoder version change and fingerprint extractor version change are all processed according to this criterion. When the same update transaction is submitted repeatedly, the existing version is maintained idempotently with the update transaction identifier. When the version number overflows the preset 64-bit upper limit, the original memory identifier is frozen and enters the manual migration process.
[0118] As a result, the folder first uses a pre-grouped key Establish candidate pregroups, symbols Indicates a folded pregroup key, symbol Symbols representing primitive memory identifiers Indicates the content version and is a positive integer; within each pregroup, after selecting the commit generation by query snapshot, the final composite key is used. Create a fold group, symbol Indicates the folded final compound bond, symbol The selected candidate is associated with a committed index generation and is a non-negative integer. Before entering the pregrouping, the candidate is checked for the following: the original memory identifier is not empty, the content version is greater than zero, the committed generation is not greater than the query snapshot generation, and the text digest is valid. If any of the checks are not met, the candidate is moved to the rejection set and does not participate in the matching.
[0119] For pregroups with the same original memory identifier and content version, the result folder first sorts them in descending order of commit generation and selects the highest commit generation whose commit generation is not greater than the commit generation of the query snapshot index. Then, it only forms folding groups with the same final composite key for candidates that are equal to the highest commit generation. When multiple candidates have the same highest commit generation, the target text with the migration status of committed is retained first, followed by the target text corresponding to the recycled record, and then the source text visible in the old snapshot. The hit index type, matching fragment, original score, quantile value, migration path and target bucket identifier of the remaining parallel candidates are incorporated into the traceable fields of the retained candidates.
[0120] For candidates with the same original memory identifier but different content versions, the result folder sorts them in descending order of content version. Within each version, the aforementioned commit generation selection is performed. Only the highest content version whose commit generation is no greater than the commit index generation of the query snapshot is retained. If the highest version text is not visible due to permission conditions, it will not automatically fall back to the old version with wider permissions. Instead, the invisible item is removed from the candidate group of the original memory identifier and the old version is retained only if an old version that still meets the same permission conditions exists.
[0121] The generation selection of candidates in the same version is written as ,symbol This indicates that the same original memory identifier and content version pregrouping are selected commit generations under the query snapshot and are non-negative integers. Indicates candidate The associated commit generation is a non-negative integer, and the sign is... Indicates the candidate sequence number within the same pregroup, symbol This indicates that the committed index generation in the query snapshot is a non-negative integer; if the set is empty, the pregroup is deleted; if the set is not empty, only candidates whose committed generation is equal to the selected committed generation are retained, and the final composite key is formed by the original memory identifier, content version and the selected committed generation before entering the state priority comparison.
[0122] Before merging the hit evidence, the result folder checks the content conflict flags written by S3. If any source target candidate is marked as having a content conflict under the same original memory identifier, content version, and commit generation, all text candidates of the final fold group are moved to the content integrity isolation set. Neither the source nor the target text is selected, and no valid logical candidate slots are occupied. The migration transaction identifier, the text verification summary at both ends, and the query snapshot generation are recorded for manual review. At the same time, the logical candidate bucket is reported as invalid to S5. Based on this, S5 keeps the query termination record as pending confirmation and reopens the termination preparation branch with a remaining upper bound not lower than the updated Kth bucket score according to the cursor continuation token stored therein, or reopens all termination preparation branches with remaining budget when there are less than K valid buckets. Stop the prepared branch, continue scanning and resubmit the fusion candidate to S6 until the valid logical candidate reaches K or the budget and unified query period have been exhausted after integrity verification; for fold groups without content conflicts, when merging hit evidence, do not recalculate the original scores of different index branches, but retain the branch percentile values and evidence components given by S5. The fusion score of the retained candidate is the maximum value of the fusion scores of each candidate under the same content version and commit generation. Hit index types are deduplicated in a fixed order of exact, semantic, time, task entity and summary prototype. Migration paths are deduplicated according to migration transaction identifiers and retain source layer, target layer, road sign item identifier and target bucket identifier, so that a single migration double write will not occupy two return slots with two memory texts;
[0123] After completing the folding within the group, the result folder sorts the candidates in descending order of fusion score. Candidates with the same fusion score are then ordered sequentially by content version (descending), submission generation (descending), semantic importance (descending), and finally, lexicographical order of the original memory identifier (ascending). The valid range for both candidate fusion score and semantic importance is 0 to 1. Candidates exceeding this range are rejected, and a fusion calibration anomaly is recorded. The top K items are then selected as candidates. If, after integrity review, content conflict isolation, or permission eviction, there are fewer than K items, the result folder does not return an integrity confirmation success flag but instead sends a logical candidate bucket failure and replenishment request to S5. S5 keeps the query termination record in a pending confirmation state and reopens the branch in the termination preparation state that still has remaining budget and valid cursor continuation tokens. If the number of valid buckets is less than K, the branch is reopened. Open all eligible termination preparation branches. When the number of valid buckets reaches K, reopen termination preparation branches whose remaining upper bound is not lower than the current Kth bucket score. Do not reopen branches that are unrecoverable hard boundary states due to budget exhaustion, scan completion, unified query deadline reaching, or abnormal exit. Supplement candidates and perform S5 fusion and S6 post-integrity verification again until the number of valid candidates reaches K and passes all verifications, or all branches are in unrecoverable hard boundary states and there are no more termination preparation branches to reopen. In the latter case, perform post-integrity verification on existing candidates for exhaustion-type query termination records. After verification, even if the actual number of valid records is less than K, return the integrity confirmation success flag and output according to the actual number of valid records, without exceeding the query snapshot or permission conditions to scan unplanned areas.
[0124] To verify the folding rules, the folder performs deterministic validation before deployment using four states: prepared for coverage, placed, committed, and recycled, as well as replay sets of old and new query snapshots. It requires that the same input outputs the same original memory identifier, content version, and commit generation sequence regardless of the candidate arrival order, and that each output original memory identifier occupies only one return slot. During runtime, a shadow fold is extracted once every thousand queries. If the shadow fold result is inconsistent with the main fold result, the main result is retained, the input snapshot is recorded, and the release of new rule versions is paused.
[0125] The result folder constructs a memory context based on the matching fragments of the selected candidates, extracts adjacent text fragments according to the time and task conditions in the query, and avoids splicing different logical memories into a fragment by using the original memory identifier as the boundary. For code memory, it retains the file path, function boundary and line number; for tool call results, it retains the tool name, call time and result status; for runtime log, it retains the trace identifier and exception stack summary. When a single context exceeds 2048 words, it prioritizes retaining the hit fragment and the 512 words before and after it.
[0126] For each output item, the traceable fields are written as the original memory identifier, content version, source layer, submission generation, set of hit index types, fusion score, evidence component, migration path set, target bucket identifier, text verification summary, and query snapshot generation. Before output, the access subject identifier, set of allowed access domain identifiers, and memory visibility level are checked again. For candidates with inconsistent text object verification summaries, the output is stopped and written to the integrity review queue.
[0127] The memory context, along with the query identifier, query snapshot generation, number of returned items, and folding statistics, is written to the response object. The folding statistics record the number of input candidates, the number of duplicate items migrated within the same version, the number of items eliminated across versions, the number of items eliminated with permissions, and the final output. After consuming the response, the agent only sends an adoption confirmation message to the memory units that are actually included in the inference hints or decision-making basis. When all candidates have completed strong consistency back lookup, content conflict isolation, permission review, and quota verification, and there is no need to request S5 to supplement them again, the result folder returns an integrity confirmation success flag to S5. S5 then confirms the query termination record and destroys the paused cursor. The access statistics processor in S1 atomically updates the most recent access time, access count, and access interval accordingly, so that the current output forms a closed loop for the resident utility calculation of the next cycle.
[0128] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0129] This invention establishes a determinable technical relationship between the physical layer where the memory resides, the visible state of the migration, and the query candidate generation process through the cooperation of hierarchical heterogeneous indexes, migration landmarks, and generational snapshots. This allows queries to avoid traversing all indexes at each layer to confirm the migration result, and enables physically dispersed memories that are associated with landmarks or task entity relationships to be directionally included in the candidate set.
[0130] This invention uses a transition state machine to constrain the visibility range of source and target candidates, uses the upper bound of the remaining candidate fusion score of the same scale to control the termination of parallel branches, and limits the unused budget to the target bucket pointed to by the signpost and with evidence gaps, thereby reducing invalid slow layer accesses while maintaining the integrity of the first K search results.
Claims
1. A hierarchical indexing and parallel retrieval method for intelligent agent dialogue memory, characterized in that, include: S1. Obtain the memory data generated by the agent and encapsulate it into a memory unit containing the original memory identifier, content version, domain identifier, time information, access statistics and semantic representation; S2. Determine the residency utility based on the semantic importance, time information and access statistics of the memory unit, place the memory unit into the hot layer, warm layer or cold layer according to the hierarchical rules with hysteresis interval, and establish a layer-differentiated index. S3. In response to the change of the residence layer of the memory unit, a migration record is established with the original memory identifier, a road sign item pointing to the target retrieval area is written in the source layer, the target placement certificate is associated with the index generation and then submitted to form a migration visibility state determined by the index generation. S4. Parse the memory query to obtain the search conditions, read the committed index generation once as a query snapshot, allocate candidate budgets to each level of index and search in parallel, consume only the text items and sign items visible in the query snapshot, access the target search area along the hit sign, and obtain a heterogeneous candidate set. S5. Convert the candidate scores of each index branch to the quantile scale of the corresponding historical score distribution and merge them. Control the branch termination and candidate budget transfer based on the upper bound of the remaining candidate fusion scores of each index branch to obtain the fusion candidate set. S6. Based on the original memory identifier, content version, and migration duplicates in the submitted index generation folding fusion candidate set associated with the candidate, select candidates according to the candidate fusion score and output the memory context.
2. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 1, characterized in that, S1 includes: The memory data is defined as any one of the following: dialogue text, code snippets, tool call results, or runtime logs; the domain identifier includes at least one of session identifier and task identifier; the time information includes generation time and recent access time; and the access statistics include access count and access interval. Semantic encoding is performed on the memory data, and entity fingerprints or code symbol fingerprints are extracted. The resulting semantic representation and symbol fingerprint are associated with the original memory identifier and content version and written into the memory unit.
3. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 1, characterized in that, S2 includes: weighting the semantic importance, the time decay determined by time information, and the access activity determined by access statistics, and correcting the weighted combination result by combining the reconstruction cost of the memory unit to obtain the residency utility, wherein the larger the residency utility, the higher the priority of residing in the hot layer; setting a migration threshold for migration into the hotter layer and a migration threshold for migration out of the hotter layer for two adjacent layers, such that the migration threshold is greater than the migration threshold, and keeping the layer where the memory unit is located unchanged when the residency utility is equal to either threshold; establishing an exact inverted index or hash index and a time ring index in the hot layer, establishing an approximate vector index and a task entity adjacency index in the warm layer, and establishing a summary prototype index, a time bucket index, and an archive address index in the cold layer.
4. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 1, characterized in that, S3 includes: sequentially navigating the migration record through the states of preparation, placement, commit, and recycling; generating a fixed-length roadmap entry in the preparation state, the roadmap entry recording the source layer, target layer, old index key digest, new index key digest, target bucket identifier, migration generation, and status bit; generating a placement credential containing the target bucket identifier and target index checksum after the target text and target index are written, and switching the migration record to the placement state when the placement credential is generated, atomically associating the roadmap entry, placement credential, and index generation, and setting the migration generation to... After the associated index generation value is reached, the system switches to the committed state. If the target text write fails, the target index write fails, the placement credential verification fails, or the commit is interrupted before the committed state, the migration record is maintained or rolled back to the ready state, and the source text remains visible to the query snapshot. The system re-submits the data and re-submits after rewriting and passing the placement credential verification. If there is no active query snapshot with a committed index generation smaller than the migration generation, the source text is reclaimed and the system switches to the reclaimed state. Before this, the source text is retained for the active query snapshot to read, and the signpost item is retained for a preset retention period.
5. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 4, characterized in that, S4 includes: decomposing the search conditions into precise conditions, semantic conditions, time conditions, and domain conditions, and selecting corresponding index branches for each; using permission conditions as new filtering conditions; and allocating the candidate budget to each index branch based on the query type, historical hit rate at each layer, and access cost; consuming source text items for migration records in the preparation state; consuming source text items visible in the query snapshot for migration records in the placed but not committed state; and ensuring that the target text and target index are used only for placement credential verification and do not enter the candidate list for migration records in the committed or reclaimed state; consuming target text items or accessing the target bucket or summary prototype range defined by the target bucket identifier for migration records in the committed or reclaimed state.
6. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 5, characterized in that, S5 includes: The historical score cumulative distribution is maintained according to the index branch and query type, and the original score of each index branch is mapped to the quantile value in the historical score cumulative distribution. The candidate fusion score is obtained by weighting and combining the quantile value, time evidence, importance evidence, access evidence, and migration path evidence indicating whether the candidate was obtained along the road sign item. Based on the upper bound of the quantile values that have not yet returned candidates, as well as the upper bound of the contribution that can be obtained from time evidence, importance evidence, access evidence, and migration path evidence, the upper bound of the remaining candidate fusion score is determined according to the same weighted combination rule as the candidate fusion score, so that each index branch returns to the upper bound of the remaining candidate fusion score when returning to the candidate, and the upper bound of the remaining candidate fusion score is used as the control quantity for branch termination and candidate budget transfer.
7. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 6, characterized in that, Get the preset number of items parameter K, where K is a positive integer; when the number of current valid candidates reaches K items, sort the current valid candidates in descending order of candidate fusion score, and terminate the index branch when the upper bound of the remaining candidate fusion score of an index branch does not reach the candidate fusion score of the sorted Kth candidate. Extract entity conditions from the precise conditions or semantic conditions, extract task conditions from the domain conditions, and determine that there is a gap in the corresponding evidence when no candidate satisfying the corresponding conditions is found in the current valid candidates for entity conditions, time conditions or task conditions. The unused candidate budget of the index branch is transferred to the target bucket where the signpost has been hit and the evidence gap exists. The target bucket is a storage area defined by the target bucket identifier within the target retrieval area pointed to by the signpost. The target bucket identifier is used as a budget transfer constraint to prevent the unused candidate budget from being transferred to a target bucket not pointed to by the signpost.
8. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 4, characterized in that, The visibility of the query snapshot to the migration record is determined by comparing the migration generation with the committed index generation in the query snapshot; when the migration generation is not greater than the committed index generation in the query snapshot, the target text item or the signpost item is made visible to the query snapshot. When the migration generation is greater than the committed index generation in the query snapshot, or when the migration record has not yet formed a committed migration generation, the source text item is made visible to the query snapshot and the target text item and target index are not included in the candidates; when both the source text and target text corresponding to the same original memory identifier are visible to the query snapshot, both are sent to the folding process as candidates, and their respective source layer, content version, the migration generation as the committed generation and migration path are retained as the basis for folding.
9. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 5, characterized in that, The permission conditions include the access subject identifier, the set of allowed domain identifiers, and the memory visibility level; before the index branch returns candidates, the index items and road sign items are filtered according to the permission conditions, and the target text accessed along the road sign items continues to use the same or less restrictive permission conditions as the source text as access security constraints.
10. The hierarchical indexing and parallel retrieval method for agent dialogue memory according to claim 8, characterized in that, S6 include: Receive the fusion candidate set and query snapshot, obtain the preset number of items parameter K, where K is a positive integer, and set the content version to the version number that increments with the content update of the same original memory identifier; for candidates with the same original memory identifier and content version, sort them in descending order of submission generation and select the first candidate whose submission generation is not greater than the submitted index generation in the query snapshot; when multiple candidates have the same first submission generation, prioritize retaining the target text in the submitted state, and merge the hit index type and migration path of the remaining parallel candidates into the retained candidate; For candidates with the same original memory identifier but different content versions, sort them in descending order of version number and retain the first candidate whose submission generation is no greater than the submitted index generation in the query snapshot; sort the folded candidates according to the candidate fusion score and select the top K items, write the source layer, content version, submission generation, hit index type and migration path of the selected candidates into the traceable field, and output them with the memory context.