Self-evolution artificial intelligent agent system based on memory enhancement and reflection mechanism

CN122221898BActive Publication Date: 2026-09-18BEIJING FUTURE INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610342128.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-19
Publication Date
2026-09-18
Estimated Expiration
2046-03-19

AI Technical Summary

Technical Problem

[0005]本发明所要解决的技术问题在于针对智能体持续学习中的知识污染与性能衰退

Benefits of technology

[0021] 1. This self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms establishes a low-level memory isolation mechanism by dividing the main working memory area and the reflection sandbox area in the memory management module. Combined with the cognitive conflict detection module, the extracted local causal subgraphs and reference local subgraphs are transformed into weighted adjacency matrices, and the square root operation of the sum of squares of the matrix difference elements is performed to output the topological difference quantity. This enables the system to perform independent and isolated topological comparisons and intercept logical deviation strategies when new reflection rules are generated. This achieves the effect of ensuring the safety and accuracy of the global causal graph evolution without interfering with the execution of regular tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122221898B_ABST
    Figure CN122221898B_ABST
Patent Text Reader

Abstract

This invention discloses a self-evolving artificial intelligence agent system based on memory enhancement and reflection mechanisms. The system includes a task execution monitoring module, a memory management module, a cognitive conflict detection module, a counterfactual scenario generation module, and an asynchronous shadow verification module. The task execution monitoring module collects scores to trigger the reflection mechanism and generate a local causal subgraph. The memory management module isolates and injects the local causal subgraph in the reflection sandbox area. The cognitive conflict detection module compares the graph of the main working memory area, calculates the topological difference, and locates conflict nodes. The counterfactual scenario generation module perturbs the preconditions to construct a counterfactual boundary test scenario. The asynchronous shadow verification module instantiates a shadow agent to perform verification, performing momentum weight merging based on the weighted comprehensive score, or generating a hash feature string to append to the taboo list for blocking. This invention achieves safe and controlled evolution of the agent's causal graph, avoiding performance degradation caused by absorbing erroneous logic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms. Background Technology

[0002] Artificial intelligence agents refer to intelligent computing entities that can perceive the external environment, make autonomous decision-making inferences, and perform physical or software-level actions. With the development of reinforcement learning and knowledge graph technologies, artificial intelligence agents are gradually evolving from being able to perform only pre-programmed fixed tasks to complex decision-making systems with continuous learning, experience accumulation, and autonomous evolution capabilities, and are being widely applied in business scenarios such as autonomous driving, robot control, and complex system scheduling.

[0003] In existing technological applications, in order to adapt to dynamically changing task requirements, AI agents typically rely on experience replay and online fine-tuning mechanisms to update their knowledge base or decision-making models. The system continuously collects environmental status feedback during task execution. When encountering task execution failures or poor performance, the AI ​​agent uses this failure data to generate new empirical rules or adjust model parameters, and writes this newly extracted knowledge directly into the underlying storage system or global knowledge base so that it can call the updated strategy to cope when facing the same or similar tasks in the future.

[0004] However, existing agent self-evolution mechanisms often lack secure isolation and systematic structural verification of newly added empirical rules. When an agent generates new causal logic based on feedback from local failed tasks and directly integrates it into the global system, if the new logic itself has inherent biases or structural conflicts with the existing knowledge system, it will directly pollute the existing main working memory. This unverified knowledge overlay will destroy the original topology of the underlying data, causing decision confusion and performance crashes when the agent handles routine tasks that it could previously solve successfully. It is impossible to guarantee the security and accuracy of the global knowledge system evolution without interfering with normal business execution. Summary of the Invention

[0005] The technical problem to be solved by this invention is the knowledge pollution and performance degradation in the continuous learning of intelligent agents.

[0006] The first aspect of this invention provides a self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms.

[0007] The system includes a task execution monitoring module, a memory management module, a cognitive conflict detection module, a counterfactual scenario generation module, and an asynchronous shadow verification module;

[0008] The task execution monitoring module collects the environmental state transition sequence and the task score fed back by the environment during the intelligent agent's task execution in real time. The task execution monitoring module is equipped with dual threshold judgment logic: calculate the cumulative reward value after the task execution trajectory ends, and when the cumulative reward value is lower than the preset task success score lower limit, issue a reflection start signal; calculate the fluctuation ratio between the current cumulative reward value and the average cumulative reward value of the historical execution cycle, and when the fluctuation ratio exceeds the preset performance jitter threshold, trigger the reflection mechanism simultaneously.

[0009] The memory management module divides the main working memory area and the reflection sandbox area at the logical storage level. The main working memory area stores the global causal graph through a graph database, while the reflection sandbox area stores the improvement reflection rules to be verified. Within the reflection sandbox area, the memory management module allocates an independent memory mirror space through the memory page table isolation mechanism at the operating system level, and injects the local causal subgraph into the memory mirror space in the form of a local weighted adjacency matrix.

[0010] The cognitive conflict detection module performs a topological comparison between the local causal subgraphs generated in the reflection sandbox area and the existing rules in the main working memory area. When performing the topological comparison, the cognitive conflict detection module extracts the central anchor point of the local causal subgraph, performs a breadth-first search of a preset depth in the global causal graph, performs data filtering through the attribute matching indicator function, and obtains reference local subgraphs from the global causal graph.

[0011] The cognitive conflict detection module transforms the local causal subgraph into a local weighted adjacency matrix and the reference local subgraph into a reference adjacency matrix based on a node mapping sequence of a unified dimension. The cognitive conflict detection module calculates the sum of squares of the differences between corresponding elements in these two matrices and then performs a square root operation on the sum of squares to obtain a value used to characterize the topological difference. When the value of the topological difference is greater than the preset topological tolerance threshold, the cognitive conflict detection module traverses the elements in the difference matrix whose absolute value exceeds the single-sided anomaly threshold, reverse maps to generate a core conflict node set, and obtains the set of precondition nodes corresponding to the in-degree edge and the set of subsequent influence nodes corresponding to the out-degree edge. The union of the above sets is used to generate a conflict boundary node set.

[0012] The counterfactual scenario generation module constructs counterfactual boundary test scenarios based on the topological difference points located and by perturbing the precondition parameters of conflict nodes. The counterfactual scenario generation module is configured with a variety of perturbation operators: the numerical boundary perturbation operator calculates negative and positive boundary values ​​for the historical numerical range of continuous environmental variables; the logical reversal perturbation operator extracts the complement category label for discrete environmental variables; and the conditional concealment perturbation operator replaces the constraint values ​​of selected environmental variables with null identifiers according to the preset concealment probability. The counterfactual scenario generation module performs a set union operation on the counterfactual boundary test scenario set and the benchmark normal scenario set extracted from the global high-order benchmark database to construct a dedicated verification dataset, and adds an independent binary traceability identifier to each scenario data in the dedicated verification dataset.

[0013] The asynchronous shadow verification module instantiates a shadow agent in an isolated computing environment and performs regression verification on the policy in the reflection sandbox area using the counterfactual boundary test scenario. The asynchronous shadow verification module splits the test results according to the source identification bit and calculates the average reward value of the counterfactual scenario and the average reward value of the benchmark scenario respectively. The asynchronous shadow verification module uses the preset verification weight coefficient to perform a weighted summation operation on the average reward value of the counterfactual scenario and the average reward value of the benchmark scenario to obtain the comprehensive verification score.

[0014] The asynchronous shadow verification module returns a merge control signal or a failure path blocking signal to the memory management module based on the comprehensive verification score. When the comprehensive verification score is greater than or equal to the system acceptance threshold, the memory management module uses a preset momentum decay rate to perform a weighted smooth update on the historical weight value and the reflection weight value, and overwrites the final weight value to the global causal graph. When the comprehensive verification score is less than the system acceptance threshold, the memory management module maps the feature string hash of the corresponding reflection directed edge to a hash feature string and appends it to the taboo list using a hash set data structure, while triggering the resource reclamation operation in the reflection sandbox area.

[0015] A second aspect of this invention provides a method for a self-evolving artificial intelligence agent based on memory enhancement and reflection mechanisms;

[0016] This method specifically elucidates the innovative principle of the present invention. By constructing a closed-loop process of experience reflection isolation, quantification testing, and controlled merging, it solves the performance crash problem in the continuous learning process of the intelligent agent. The specific innovative parts are as follows:

[0017] In terms of cognitive conflict detection and quantification, this invention transforms the comparison of graph topology into numerical operations of weighted adjacency matrices. The system extracts the node mappings of the local graph and the global reference graph and constructs a square matrix. By using the logic of calculating the sum of squares of the difference elements of the matrix and performing square root operations, the abstract structural differences are mapped to precise mathematical norms. This operation enables the system to accurately intercept logical deviation strategies using preset thresholds.

[0018] In terms of constructing counterfactual boundary scenarios, this invention adopts a scenario generation logic based on causal reverse tracing. The system locates the core conflict node through unilateral anomaly judgment and performs reverse search in the global causal graph to lock the preconditions. It performs perturbation operations such as numerical out-of-bounds or logical reversal on the preconditions and directly generates boundary environment data for the vulnerability points of the improved rules. This process avoids the inefficiency of random verification data and improves the adversarial strength of the test samples.

[0019] In terms of knowledge merging and failure prevention, this invention adopts a graph evolution logic based on weighted mean evaluation and momentum smoothing. The system verifies the reward mean of the weight coefficients to balance the counterfactual performance and the benchmark performance, ensuring that the strategy does not decay its existing capabilities when dealing with boundary anomalies. For graph branches that pass verification, a momentum decay mechanism is used to perform weight fusion to avoid drastic changes in node weights. For blocked graph branches, the source node, target node and operator of the erroneous causal relationship are concatenated and then subjected to secure hash calculation to generate a feature string and append it to a taboo list using a hash set data structure, thus intercepting the secondary generation of erroneous logic from the bottom layer.

[0020] The present invention, by adopting the above technical solution, can bring the following beneficial effects:

[0021] 1. This self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms establishes a low-level memory isolation mechanism by dividing the main working memory area and the reflection sandbox area in the memory management module. Combined with the cognitive conflict detection module, the extracted local causal subgraphs and reference local subgraphs are transformed into weighted adjacency matrices, and the square root operation of the sum of squares of the matrix difference elements is performed to output the topological difference quantity. This enables the system to perform independent and isolated topological comparisons and intercept logical deviation strategies when new reflection rules are generated. This achieves the effect of ensuring the safety and accuracy of the global causal graph evolution without interfering with the execution of regular tasks.

[0022] 2. This self-evolving AI agent system based on memory enhancement and reflection mechanisms uses a counterfactual scenario generation module to perform reverse search on core conflict nodes to lock the set of precondition nodes. It then calls perturbation operators such as numerical out-of-bounds and logic reversal to perform feature negation operations on environmental variables. The generated counterfactual test scenarios and benchmark scenarios are then combined to construct a dedicated verification dataset. This allows the system to generate verification environments containing extreme features to overcome the blindness of conventional testing. It achieves the effect of accurately exposing the vulnerabilities of newly added graph logic under boundary conditions and improving the strength of agent strategy verification adversarial.

[0023] 3. This self-evolving AI agent system based on memory enhancement and reflection mechanisms calculates a comprehensive score by weighted summation of the average reward values ​​of two test scenarios through an asynchronous shadow verification module. When the score meets the target, the historical weight values ​​of the global causal graph are smoothly updated using the momentum decay rate. When the score does not meet the target, the concatenated feature string is hashed and mapped to a hash feature string and added to the taboo list. This allows the agent to take into account the stability of its historical task performance when absorbing new strategies, thus achieving a controlled evolution of the graph that avoids drastic changes in graph weights and continuously intercepts the secondary generation of invalid logic from the bottom layer. Attached Figure Description

[0024] Figure 1 This is a logical architecture diagram of the agent self-evolution system based on causal graph sandbox and counterfactual shadow verification of the present invention;

[0025] Figure 2 This is a flowchart of the agent self-evolution method based on causal graph sandbox and counterfactual shadow verification of the present invention;

[0026] Figure 3 This is a schematic diagram of the global causal graph data structure of the present invention;

[0027] Figure 4 This invention provides flowcharts for task feedback monitoring and reflection triggering, and reflection rule parsing and sandbox injection.

[0028] Figure 5 The flowcharts for subgraph isomorphic matching and reference subgraph extraction, and the flowchart for calculating topological differences are shown below.

[0029] Figure 6 The flowcharts for locating and marking the core conflict node set and the reverse extraction flowchart of the preconditions for core conflict nodes are provided in this invention.

[0030] Figure 7 This is a flowchart illustrating the context perturbation and counterfactual scenario generation process of the present invention.

[0031] Figure 8This is a flowchart of the counterfactual scenario and benchmark sample fusion process and a flowchart of the shadow agent instantiation and multi-threaded testing process of this invention.

[0032] Figure 9 This is a flowchart of the weighted calculation process for the comprehensive verification score of this invention;

[0033] Figure 10 This is a flowchart of the spectrum merging and failure blocking process of the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] See attached document Figure 1 The present invention provides an agent self-evolution system based on causal graph sandbox and counterfactual shadow verification, which may include: a task execution monitoring module 110, configured to drive the agent to execute a target task in the main working memory environment, wherein the task execution monitoring module 110 collects the environmental state transition sequence and the task score fed back by the environment in real time during the task execution process;

[0036] The memory management module 120 is configured to divide the main working memory area 121 and the reflection sandbox area 122 at the logical storage level. The main working memory area 121 stores the global knowledge graph through a graph database, and the reflection sandbox area 122 stores the reflection logic to be verified. The main working memory area 121 and the reflection sandbox area 122 are independent of each other in physical address space and communicate across areas through a controlled access interface.

[0037] The cognitive conflict detection module 130 is configured to perform graph structure analysis on the improved rules received in the reflection sandbox area 122, and to perform topological comparison between the generated local causal subgraph and the existing rules in the main working memory area 121.

[0038] The counterfactual scenario generation module 140 is configured to construct counterfactual boundary test cases by modifying the precondition parameters of the conflict nodes based on the topological difference points located by the cognitive conflict detection module 130.

[0039] The asynchronous shadow verification module 150 is configured to instantiate shadow agents in an isolated computing environment. The asynchronous shadow verification module 150 uses test cases provided by the counterfactual scenario generation module 140 to perform regression verification on the strategies in the reflection sandbox area 122.

[0040] The task execution monitoring module 110 communicates with the memory management module 120 through a data interface. The task execution monitoring module 110 sends the collected execution trajectory with the task score lower than the preset threshold to the memory management module 120. The memory management module 120 determines the storage space allocation size of the reflection sandbox area 122 based on the task score drop value. The main working memory area 121 and the reflection sandbox area 122 maintain a read-write isolation state.

[0041] The reflection sandbox 122 receives text instructions forwarded by the task execution monitoring module 110, and then transforms the text instructions into a weighted adjacency matrix containing entity nodes and causal edges. ;

[0042] The cognitive conflict detection module 130 is connected to the main working memory area 121 and the reflection sandbox area 122 via an internal bus. The cognitive conflict detection module 130 obtains the weighted adjacency matrix. And retrieve the reference adjacency matrix of the main working memory 121 that contains the same entity nodes. The cognitive conflict detection module 130 calculates the matrix difference. The formula for locating conflict points is as follows:

[0043] (1),

[0044] in, Represents the weighted adjacency matrix The Middle Line 1 The element values ​​of the column, Represents the reference adjacency matrix The value of the element at the corresponding position in the middle. Indicates the total number of nodes involved;

[0045] The cognitive conflict detection module 130 will measure the matrix difference. Node coordinates exceeding the preset tolerance are sent to the counterfactual scenario generation module 140. The counterfactual scenario generation module 140 extracts causal constraints from the main working memory area 121 based on the node coordinates and performs a negation operation on the state features in the causal constraints to generate counterfactual boundary test cases.

[0046] The counterfactual scenario generation module 140 connects to the asynchronous shadow verification module 150 through a signaling channel, and the counterfactual scenario generation module 140 pushes the synthesized boundary test case sequence to the task queue of the asynchronous shadow verification module 150.

[0047] After executing all test cases, the asynchronous shadow verification module 150 returns the final merging control signal to the memory management module 120. When the signal is a pass instruction, the memory management module 120 releases the isolation between the main working memory area 121 and the reflection sandbox area 122. The memory management module 120 then uses a graph fusion algorithm to merge the weighted adjacency matrix... The corresponding topology is merged into the global knowledge graph of the main working memory area 121.

[0048] See attached document Figure 2 This invention provides an agent self-evolution method based on causal graph sandbox and counterfactual shadow verification, which may include:

[0049] In step S210, the task execution monitoring module 110 drives the agent to call the global causal graph in the main working memory area 121 to execute the target task, and the task execution monitoring module 110 obtains the task score fed back by the environment in real time.

[0050] In step S220, the task execution monitoring module 110 determines whether the task score meets the preset task achievement threshold. When the task score does not meet the preset task achievement threshold, the task execution monitoring module 110 triggers the agent to generate improvement and reflection rules.

[0051] In step S230, the memory management module 120 intercepts the improved reflection rules and stores them in the reflection sandbox area 122. The reflection sandbox area 122 performs semantic parsing on the improved reflection rules and constructs a local weighted adjacency matrix containing entity nodes and causal edges. ;

[0052] In step S240, the cognitive conflict detection module 130 extracts the adjacency matrix with local weights from the main working memory area 121. Reference adjacency matrix of overlapping nodes The cognitive conflict detection module 130 calculates the local weighted adjacency matrix. With reference adjacency matrix Topological differences between Topological differences The calculation logic is referenced in the aforementioned formula (1);

[0053] Step S250, the cognitive conflict detection module 130 locates the topological difference. For nodes whose median values ​​exceed a preset edge weight threshold, a core conflict node set is determined. The counterfactual scenario generation module 140 obtains the core conflict node set and backtracks the corresponding causal preconditions in the main working memory area 121. The causal preconditions include a set of environmental state variables related to the core conflict node set. ;

[0054] Step S260, the counterfactual scenario generation module 140 processes the set of environmental state variables. Perform state inversion processing to generate counterfactual boundary test scenarios. The state inversion process includes adjusting the values ​​in the set of environmental state variables to the complement range outside the preset constraint range. The counterfactual scenario generation module 140 merges the counterfactual boundary test scenario with the benchmark test scenario extracted from the historical database to construct a dedicated verification dataset.

[0055] In step S270, the asynchronous shadow verification module 150 instantiates a shadow agent in the reflection sandbox area 122. The reflection sandbox area 122 sends the policy parameters involved in the improved reflection rules to the asynchronous shadow verification module 150. The asynchronous shadow verification module 150 controls the shadow agent to load the improved reflection rules and executes the test in the dedicated verification dataset environment.

[0056] In step S280, the asynchronous shadow verification module 150 summarizes the task reward values ​​of the shadow agent in the counterfactual boundary test scenario and the benchmark test scenario, and calculates the comprehensive verification score. The calculation formula is as follows:

[0057] (2),

[0058] in, This represents the average reward value of the shadow agent in the counterfactual boundary test scenario. This represents the average reward value of the shadow agent in the benchmark test scenario. The preset verification weight coefficients, and ;

[0059] In step S290, the memory management module 120 receives the comprehensive verification score returned by the asynchronous shadow verification module 150. When the comprehensive verification score When the threshold value is greater than or equal to the preset merging threshold, the memory management module 120 removes the physical isolation between the main working memory area 121 and the reflection sandbox area 122, and merges the local weight adjacency matrix. The corresponding topology is merged into the main working memory area 121;

[0060] Step S300, when the comprehensive verification score is... When the threshold value is less than the preset merging threshold, the memory management module 120 blocks the improvement and reflection rules from entering the main working memory area 121. The memory management module 120 records the index identifier of the improvement and reflection rules as an invalid path in the reflection sandbox area 122 and clears the corresponding local weight adjacency matrix in the reflection sandbox area 122. .

[0061] See attached document Figure 3This invention provides a data structure definition and storage method for a global causal graph, which may include: the main working memory area 121 uses a graph database as the underlying storage medium for persistent storage of a set of nodes. Sum of edges The constructed global causal graph ;

[0062] Node set Each node in Represents an independent environment state, entity object, or action parameter; edge set Each edge in Represents the causal logical connection between two nodes, each node It includes a node identifier, a type attribute, and a state feature vector. The node identifier is used for unique indexing within the graph database. The type attribute is used to distinguish whether the current node belongs to an environment observation node, an agent action node, or a task target node. The state feature vector is used to store the numerical representation of the node in a multi-dimensional space. The dimension of the state feature vector is determined by the feature dimension of the environment observation space.

[0063] edge set Stored as an adjacency list, each edge Includes source node index, target node index, causal type, and weight coefficient. Causal types are used to define the logical influence of a source node on a target node. Causal association has a set of operators, where represents a positive promoting effect, represents a negative inhibiting effect, represents a boolean effect triggered by a specific precondition, and represents a weighted coefficient. The range of values ​​is set within a closed interval. Internally, it is used to quantify the confidence level of causal relationships;

[0064] The main working memory area 121 is physically stored through a distributed file system. The main working memory area 121 establishes a hash index based on a combination key of node identifier and type attribute. The storage logic layer stores the global causal graph. It is divided into multiple interconnected sub-plots to support parallel retrieval and update operations;

[0065] During the retrieval process, the task execution monitoring module 110 sends a query request to the main working memory area 121 based on the context features of the current task. The main working memory area 121 returns the local causal subgraph with the highest cosine similarity to the feature vector of the current task.

[0066] Reflecting on the local weighted adjacency matrix generated in Sandbox 122 When merging into the main working memory 121, the main working memory 121 executes a node matching algorithm. If the local weighted adjacency matrix... If a node already exists in the main working memory 121, then update the weight coefficient of the corresponding edge. :

[0067] (3),

[0068] in, The original weight values ​​in the main working memory area 121, To reflect on the feedback weight values ​​calculated in Sandbox 122, The preset forgetting factor, and The frequency is dynamically adjusted according to changes in the task environment;

[0069] If the local weight adjacency matrix Includes new nodes that do not exist in the main working memory 121, and the main working memory 121 in the global causal graph. The system dynamically creates new node items and their associated adjacency list records; the initial weight value of the newly created node is set to the preset confidence level.

[0070] Through the above data structure definition and storage method, the main working memory area 121 realizes the structured storage and dynamic evolution support of causal relationships, providing a basic data source for cognitive conflict detection.

[0071] See attached document Figure 4 This invention provides an implementation method for task feedback monitoring and reflection trigger threshold determination, which may include: a task execution monitoring module 110 establishing a real-time polling mechanism for the agent's execution trajectory; the task execution monitoring module 110 collecting action commands output by the agent and state feedback signals returned by the external environment at each time step; the state feedback signals including environmental observation vectors. And scalar values ​​that characterize the task objectives;

[0072] The task execution monitoring module 110 is equipped with a task score calculation unit. The task score calculation unit performs numerical processing on the status feedback signal according to a preset reward function to generate an instant score for the current task. Instant score The computational logic is defined as the environmental observation vector. With the target state vector The negative correlation function of the Euclidean distance between them;

[0073] Task execution monitoring module 110 will provide real-time scores. The task is stored in the cache queue. The task execution monitoring module 110 compares and determines whether the task has met the target threshold. The task execution monitoring module 110 then calculates the cumulative reward value after the task execution trajectory ends. When the cumulative reward value The score is below the preset minimum for task success. At that time, the task execution monitoring module 110 sends a reflection start signal to the memory management module 120;

[0074] The task execution monitoring module 110 compares and determines the score drop rate, and retrieves the average cumulative reward value of the previous execution cycle. Average cumulative reward value Taken from the preset sliding window Historical execution records within the organization;

[0075] Task execution monitoring module 110 calculates the current cumulative reward value. Compared with average cumulative reward value volatility ratio Volatility ratio The calculation formula is as follows:

[0076] (4),

[0077] When volatility ratio Exceeding the preset performance jitter threshold At that time, the task execution monitoring module 110 determines that the agent has an unmatched causal logic deviation in the current task and triggers the reflection mechanism simultaneously;

[0078] After receiving the reflection start signal, the memory management module 120 extracts the execution trajectory data stored in the cache queue of the task execution monitoring module 110. The execution trajectory data includes the environmental state observation vector sequence, the action probability distribution sequence, and the environmental feedback sequence. The memory management module 120 inputs the execution trajectory data into the reflection engine inside the intelligent body. The reflection engine calculates the gradient change rate of the instantaneous score sequence and locates the time step with the largest negative gradient fluctuation as the failure inflection point.

[0079] The reflection engine locates the cause of failure based on the turning point and generates targeted improvement reflection rules. The memory management module 120 allocates a unique storage handle for the improvement reflection rules in the reflection sandbox area 122. The memory management module 120 locks the write permissions of the reflection sandbox area 122, allowing only the reflection engine to write the improvement reflection rules and only allowing the cognitive conflict detection module 130 to read the improvement reflection rules in the reflection sandbox area 122.

[0080] Through the above-mentioned task feedback monitoring and threshold determination logic, this embodiment realizes the function of accurately locking the reflection time based on both task performance and performance fluctuations, providing a data input basis for subsequent isolation verification in the reflection sandbox area 122.

[0081] See attached document Figure 4This invention provides an implementation method for local causal subgraph parsing and sandbox isolation injection of reflection rules, which may include: a memory management module 120 receiving natural language improved reflection rules generated by a reflection engine; the memory management module 120 calling an internally configured entity recognition unit to extract features from the improved reflection rules; the entity recognition unit identifying environmental variable attributes, agent action attributes, and expected causal relationships involved in the improved reflection rules; the entity recognition unit calculating the semantic vector similarity between the extracted attributes and existing nodes in the main working memory area 121; and performing entity alignment operations to determine the local causal subgraph. Node identifier;

[0082] The memory management module 120 performs topology construction of the local causal subgraph and creates state nodes based on the identified environmental variable attributes. Create action nodes based on the agent's action attributes The memory management module 120 manages the state nodes based on causal relationships. With action nodes Establish directed edges between them Forming a local causal subgraph ;

[0083] Memory management module 120 pairs of local causal subgraphs Each directed edge in Weight initialization is performed, and the memory management module 120 extracts the logical confidence factor from the improvement and reflection rules. The memory management module 120 calculates the directed edges using a mapping function. initial weights The calculation formula is as follows:

[0084] (5),

[0085] in, To improve the text confidence score of the reflection rules, The preset confidence center value, This is the scaling factor;

[0086] The memory management module 120 allocates an independent memory mirror space within the reflection sandbox area 122. This memory mirror space is allocated through the operating system's underlying memory page table isolation mechanism. The memory management module 120 then uses a local causal subgraph... Using the weighted adjacency matrix Injected into the memory image space in the form of [the following].

[0087] Reflecting on Sandbox Zone 122: Establishing a Weighted Adjacency Matrix The access control list, the reflection sandbox area 122 is configured with a logical consistency verification unit, the logical consistency verification unit checks the injected weight adjacency matrix Perform self-loop detection and isolated node detection, if the weighted adjacency matrix If an action node has a logical infinite loop or no preconditions are defined, the logical consistency verification unit sends a correction instruction to the memory management module 120.

[0088] Reflecting on the Sandbox Zone 122, the virtualization layer isolates the external environment from the weighted adjacency matrix. Unauthorized tampering was addressed by the sandbox zone 122, which restricted access handles for inter-process communication (IPC) and only opened the data access interface based on the replica read protocol to the asynchronous shadow verification module 150. The asynchronous shadow verification module 150 obtained the weighted adjacency matrix through the data access interface. A read-only image;

[0089] Memory management module 120 completes the local causal graph. After the injection, the status flag of the reflection sandbox area 122 is updated to the pending verification state. The status flag is used to trigger the cognitive conflict detection module 130 to start the topological difference calculation process.

[0090] Through the aforementioned local causal subgraph parsing and sandbox isolation injection logic, this embodiment realizes the transformation of unstructured reflective text into a mathematical model with topological structure, and ensures that the reflective results are in a controlled and isolated state before verification.

[0091] See attached document Figure 5 This invention provides an implementation method for subgraph isomorphic matching and local reference subgraph extraction, which may include: a cognitive conflict detection module 130 monitoring the status flag of the reflection sandbox area 122; when the status flag shows a pending verification state, the cognitive conflict detection module 130 reading the local causal subgraph injected into the reflection sandbox area 122. Local causal subgraph Includes a set of reflection nodes With the set of directed edges of reflection ;

[0092] The cognitive conflict detection module 130 sends a homogeneous retrieval request to the main working memory area 121, and the cognitive conflict detection module 130 extracts the set of reflection nodes. The central anchor point, **the cognitive conflict detection module 130 calculates the set of reflection nodes.** The out-degree values ​​of each node are used to select the node with the largest out-degree value of the directed edge as the central anchor point. The cognitive conflict detection module 130 obtains the node identifier of the central anchor point.

[0093] The cognitive conflict detection module 130 uses the node identifier of the central anchor point as the query primary key in the global causal graph of the main working memory area 121. Perform hash matching to locate the global causal graph. The cognitive conflict detection module 130 uses the initial candidate nodes as the starting point in the global causal graph. Execute preset depth Breadth-first search, with a preset depth Values ​​are local causal subgraphs The longest path hop count;

[0094] During the breadth-first search process, the cognitive conflict detection module 130 extracts data on adjacent nodes and connected causal edges along the traversal path. The module performs dual verification of node categories and edge constraints, comparing the set of adjacent nodes with the set of reflection nodes. The corresponding node type attributes in the cognitive conflict detection module 130 compare the set of connected causal edges and the set of reflected directed edges. The causal operator type of the corresponding edge;

[0095] Cognitive conflict detection module 130 uses attribute matching indicator function Perform data filtering and attribute matching indicator functions. The calculation formula is as follows:

[0096] (6);

[0097] in, Global causal graph The adjacent nodes encountered during the traversal. For the set of reflection nodes The corresponding node in and These are the associated directed edges. This represents a function that retrieves node type attributes. This function represents the method for obtaining the type of a causal operator.

[0098] Cognitive conflict detection module 130 removes attribute matching indicator function The output is The cognitive conflict detection module 130 combines the verified adjacent nodes and connected causal edge data from the global causal graph. Local reference subgraphs are obtained by mid-segmentation. Local reference sub-image The main working memory area 121 and the local causal subgraph Historical baseline data for topology alignment;

[0099] The cognitive conflict detection module calculates a set of 130 reflection nodes. With local reference subgraph Union of reference node sets The cognitive conflict detection module has 130 pairs of unions. All nodes within the node are globally renumbered to generate a node mapping sequence with a unified dimension.

[0100] The cognitive conflict detection module 130, based on the node mapping sequence, divides the local causal subgraph... Transform into a local weight adjacency matrix in square matrix form The cognitive conflict detection module 130, based on the same node mapping sequence, will select the local reference subgraph. Transformed into a square matrix of reference adjacency matrices Local weighted adjacency matrix With reference adjacency matrix The matrix dimensions are all ,in Union The total number of nodes included;

[0101] When local reference subgraph When a specific directed edge defined by the node mapping sequence is missing, the cognitive conflict detection module 130 will refer to the adjacency matrix. The corresponding row and column element values ​​are initialized to zero. The cognitive conflict detection module 130 initializes the local weight adjacency matrix. With reference adjacency matrix Stored in the internal cache for later use in topology difference comparison calculations.

[0102] See attached document Figure 5 This invention provides an implementation of the logic for calculating the topological difference of a weighted adjacency matrix, which may include: a cognitive conflict detection module 130 reading a local weighted adjacency matrix from an internal cache. With reference adjacency matrix Local weighted adjacency matrix With reference adjacency matrix Having the same matrix dimensions Matrix dimension It depends on the number of node unions generated in the preceding isomorphic matching step;

[0103] The cognitive conflict detection module 130 performs matrix interpolation operations and traverses the local weight adjacency matrix. row index With column index The cognitive conflict detection module 130 will use the local weighted adjacency matrix The Middle Line 1 Column elements Subtract the reference adjacency matrix The element at the corresponding position in Generate the difference matrix Difference matrix elements in The calculation formula is as follows:

[0104] (7);

[0105] in, and All are positive integers, and , ;

[0106] Cognitive conflict detection module 130 pairs of difference matrices Each element in The cognitive conflict detection module 130 performs a square operation to obtain a squared difference matrix. It then sums all elements in the squared difference matrix to obtain the total squared difference. Finally, it performs a square root operation on the total squared difference and outputs a local weighted adjacency matrix. With reference adjacency matrix Topological differences between Topological differences For Frobenius norm, topological dissimilarity The calculation formula is as follows:

[0107] (8);

[0108] The cognitive conflict detection module 130 extracts the preset topological tolerance threshold. The cognitive conflict detection module 130 compares topological differences. With topology tolerance threshold ;

[0109] When topological differences Greater than the topology tolerance threshold At that time, the cognitive conflict detection module 130 determines the local causal subgraph within the reflection sandbox area 122. Global causal graph in main working memory 121 There is a logical conflict; the cognitive conflict detection module 130 extracts a preset one-sided anomaly threshold. The cognitive conflict detection module traverses the difference matrix 130 times. When the difference matrix absolute value of elements in Greater than the one-sided anomaly threshold At that time, the cognitive conflict detection module 130 extracts the corresponding row index. With column index The cognitive conflict detection module 130 will index the rows. With column index Mapping back to the global causal graph The entity node identifiers in the data are combined to generate the core conflict node set. The cognitive conflict detection module 130 sends an exception handling instruction to the counterfactual scenario generation module 140, and sets the core conflict nodes. and difference matrix Transmitted to counterfactual scenario generation module 140;

[0110] When topological differences Less than or equal to the topology tolerance threshold At that time, the cognitive conflict detection module 130 determines the local causal subgraph. Once the topology consistency condition is met, the cognitive conflict detection module 130 sends a regular test command to the asynchronous shadow verification module 150.

[0111] See attached document Figure 6 This invention provides an implementation method for accurately locating and marking the boundaries of a core conflict node set, which may include: a cognitive conflict detection module 130 receiving a difference matrix output by a computational stage. The cognitive conflict detection module 130 extracts a preset unilateral anomaly threshold. One-sided anomaly threshold Used to constrain the maximum allowable weight deviation of a single causal relationship edge;

[0112] The cognitive conflict detection module iterates through the difference matrix 130 times. all elements The cognitive conflict detection module has 130 pairs of elements. Perform absolute value calculation to obtain the absolute value of the element. The cognitive conflict detection module compares the absolute values ​​of 130 elements. With unilateral anomaly threshold ;

[0113] When the absolute value of the element Greater than the one-sided anomaly threshold At that time, the cognitive conflict detection module 130 extracts elements. Corresponding row index With column index The cognitive conflict detection module 130 will index the rows. If the index is identified as the source of the conflict, the column index will be... The index of the target node is identified as a conflict.

[0114] The cognitive conflict detection module 130 reads the node mapping sequence generated in the previous steps. Based on the node mapping sequence, the cognitive conflict detection module 130 reverse maps the conflict source node index and the conflict target node index to the entity node identifier in the main working memory area 121. The cognitive conflict detection module 130 summarizes the mapped entity node identifiers to generate the core conflict node set. Core conflict node set The formula for generating it is as follows:

[0115] (9),

[0116] in, This is the inverse mapping function for the node mapping sequence;

[0117] The global causal map of the cognitive conflict detection module 130 in the main working memory area 121 Searching for core conflict node sets The included entity nodes and the cognitive conflict detection module 130 extract the global causal graph. The middle point refers to the core conflict node set The cognitive conflict detection module 130 extracts the source nodes corresponding to the in-degree edges of all entity nodes as a set of precondition nodes. ;

[0118] The cognitive conflict detection module 130 extracts a global causal graph. The core conflict node set The cognitive conflict detection module 130 extracts the target nodes corresponding to the out-degree edges from all entity nodes in the dataset as a set of subsequent influencing nodes. ;

[0119] The cognitive conflict detection module 130 sets the core conflict nodes. Precondition Node Set and the set of subsequent affected nodes Take the union of the sets to generate a set of conflict boundary nodes. Set of conflict boundary nodes The calculation formula is as follows:

[0120] (10);

[0121] The cognitive conflict detection module 130 is a set of conflict boundary nodes in the reflection sandbox area 122. The cognitive conflict detection module 130 attaches a pending status flag to the nodes within the node set, and then sets up the conflict boundary nodes with pending status flags. The data is sent to the counterfactual scenario generation module 140 as the data source for the counterfactual scenario generation module 140 to construct test cases.

[0122] See attached document Figure 6 The present invention provides an implementation method for reverse extraction of the preconditions of core conflict nodes, which may include: a counterfactual scenario generation module 140 receiving a set of conflict boundary nodes with pending status identifiers sent by a cognitive conflict detection module 130. and core conflict node set ;

[0123] The counterfactual scenario generation module 140 establishes a data read connection with the main working memory area 121, and the counterfactual scenario generation module 140 traverses the core conflict node set. Target conflict nodes The counterfactual scenario generation module 140 targets conflict nodes. Starting with the global causal graph in the main working memory area 121 The process involves calling a depth-first search algorithm to perform a reverse path traversal.

[0124] The counterfactual scenario generation module 140 configures the termination condition for reverse path traversal. The termination condition is set to the traversal path reaching the endpoint of the environment observation node with the type attribute, or the traversal depth reaching a preset depth threshold. The counterfactual scenario generation module 140 extracts all nodes traversed during the reverse path traversal, performs deduplication, and constructs the target conflict nodes. The set of causal preceding nodes ;

[0125] Counterfactual scenario generation module 140 analyzes the set of causal preconditions. The state feature vectors of each node in the system contain continuous environmental variables and discrete environmental variables. The counterfactual scenario generation module 140 reads the historical value range corresponding to the continuous environmental variables by querying the attribute dictionary bound to the node in the main working memory area 121. The counterfactual scenario generation module 140 reads the historical category label corresponding to the discrete environmental variables by querying the attribute dictionary bound to the node in the main working memory area 121.

[0126] The counterfactual scenario generation module 140 combines historical numerical ranges with historical category labels, and targets the conflict nodes. Construct the initial condition constraint set Initial condition constraint set The expression is as follows:

[0127] (11),

[0128] in, For continuous environmental variables, A set of indices for continuous environment variables. For historical numerical ranges of continuous environmental variables; For discrete environmental variables, A set of indices for discrete environment variables. Historical category labels for discrete environmental variables;

[0129] The counterfactual scenario generation module 140 generates the initial set of condition constraints. Write to the internal register of the counterfactual scenario generation module 140, and the counterfactual scenario generation module 140 updates the target conflict node. The processing status bit indicates the parameter perturbation of the preconditions to be executed.

[0130] See attached document Figure 7 The present invention provides an implementation of generative context perturbation logic, which may include: a counterfactual scenario generation module 140 reading an initial condition constraint set from an internal register. The counterfactual scenario generation module 140 extracts the initial condition constraint set. Continuous and discrete environmental variable sequences in the data;

[0131] The counterfactual scenario generation module 140 calls the numerical out-of-bounds perturbation operator to process the continuous environment variable sequence. The counterfactual scenario generation module 140 targets continuous environment variables. Historical value range The numerical range width is calculated, and the counterfactual scenario generation module 140 generates the scenario according to a preset offset coefficient. Generate negative out-of-bounds values ​​below the lower limit. And positive out-of-bounds values ​​exceeding the numerical upper limit. Negative out-of-bounds value With positive outbound value The calculation formula is as follows:

[0132] (12);

[0133] (13);

[0134] Wherein, offset coefficient It is a real number greater than zero;

[0135] The counterfactual scenario generation module 140 will generate negative out-of-bounds values. and positive out-of-bounds value Write into the continuous perturbation candidate set;

[0136] The counterfactual scenario generation module 140 calls the logic flip perturbation operator to process the discrete environment variable sequence, and the counterfactual scenario generation module 140 extracts the discrete environment variables. History category tags The counterfactual scenario generation module 140 retrieves discrete environment variables from the global attribute dictionary in the main working memory area 121. The counterfactual scenario generation module 140 removes historical category labels from the entire category space. The complement category set is obtained, and the counterfactual scenario generation module 140 extracts the alternative category labels from the complement category set and writes the alternative category labels into the discrete perturbation candidate set.

[0137] The counterfactual scenario generation module 140 calls the conditional concealment perturbation operator and iterates through the initial condition constraint set. All environmental variables, the counterfactual scenario generation module 140, are determined based on a preset concealment probability. Replace the constraint values ​​of the selected environment variables with null identifiers. Null identifier The counterfactual scenario generation module 140 carries a null value identifier to instruct the shadow agent to perform action deductions without obtaining observation data of selected environmental variables. Selected environment variables are written into the hidden candidate set;

[0138] The counterfactual scenario generation module 140 performs orthogonal sampling calculations from the continuous disturbance candidate set, the discrete disturbance candidate set, and the hidden candidate set to generate a multidimensional disturbance parameter combination;

[0139] The counterfactual scenario generation module 140 inputs the multidimensional perturbation parameter combination into the pre-configured generative language model. The counterfactual scenario generation module 140 controls the generative language model to perform text conversion according to the preset scenario splicing instruction template. The scenario splicing instruction template includes the historical context baseline field, the multidimensional perturbation parameter slot field, and the structured output constraint identifier.

[0140] The counterfactual scenario generation module 140 acquires the structured test text output by the generative language model and encapsulates the structured test text into a set of counterfactual boundary test scenarios. The counterfactual scenario generation module 140 will generate a set of counterfactual boundary test scenarios. Stored in a dedicated test case database for subsequent shadow verification processes.

[0141] See attached document Figure 8This invention provides an implementation method for a mechanism for splicing and fusing counterfactual scenarios and high-order benchmark samples, which may include: a counterfactual scenario generation module 140 reading a set of counterfactual boundary test scenarios stored in a test case database. The counterfactual scenario generation module generates 140 counterfactual boundary test scenario sets. Total number of samples included ;

[0142] The counterfactual scenario generation module 140 sends a benchmark sample extraction request to the main working memory area 121. The main working memory area 121 is configured with a global high-order benchmark database, which stores environmental observation execution records of agents whose task scores are higher than the preset task success threshold in historical execution cycles.

[0143] Counterfactual scenario generation module 140 is based on the total number of samples and preset fusion ratio coefficient Calculate the required number of high-order benchmark samples Demand quantity The calculation formula is as follows:

[0144] (14),

[0145] in, This is a preset percentage parameter for counterfactual scenarios in the fused dataset, with a value range of [value range missing]. , This indicates the logic for rounding up;

[0146] The counterfactual scenario generation module 140 performs a random sampling operation without replacement in the global high-order benchmark database. The counterfactual scenario generation module 140 randomly selects an equal number of environmental observation execution records as required, generating a benchmark normal scenario set. Used to test the agent's ability to maintain its existing historical task performance in subsequent verification;

[0147] The counterfactual scenario generation module provides a set of 140 counterfactual boundary test scenarios. Compared with the baseline normal scenario set After performing data format normalization processing, the counterfactual scenario generation module 140 will generate a set of counterfactual boundary test scenarios. Scene text and baseline normal scene set The environmental observation execution records are uniformly mapped to standard test frames that include state observation space and action space;

[0148] The counterfactual scenario generation module 140 will convert the counterfactual boundary test scenario set into a standard test frame. Compared with the baseline normal scenario set Perform set union operations to construct a dedicated validation dataset. Dedicated validation dataset The calculation formula is as follows:

[0149] (15);

[0150] Counterfactual scenario generation module 140 is a dedicated verification dataset. Each piece of scene data is attached with an independent traceability identifier. The traceability identifier contains a binary value, which is used to distinguish the data source category of the scene data to which the subsequent modules belong, whether it is counterfactual generated data or historical benchmark data.

[0151] The counterfactual scenario generation module 140 will use a dedicated verification dataset. The counterfactual scenario generation module 140 serializes and encapsulates the data into a structured data stream according to the timestamp. The counterfactual scenario generation module 140 pushes the structured data stream to the task waiting queue of the asynchronous shadow verification module 150 through the signaling channel. The asynchronous shadow verification module 150 parses the structured data stream and configures the initial state of the virtual test environment.

[0152] See attached document Figure 8 The present invention provides an implementation method for instantiation of shadow agent sandbox and multi-threaded replay test, which may include: the asynchronous shadow verification module 150 monitors the internally configured task waiting queue, and when the task waiting queue receives the structured data stream pushed by the counterfactual scenario generation module 140, the asynchronous shadow verification module 150 sends a resource allocation instruction to the reflection sandbox area 122.

[0153] Reflecting on the sandbox area 122, parsing the structured data stream, and reading the dedicated validation dataset. The total number of samples, the reflection sandbox area 122, based on the total number of samples and the preset single-node batch processing capacity, calculates the target thread allocation number, and divides the independent computing node set and memory address pool in the virtualization layer;

[0154] The asynchronous shadow verification module 150 instantiates a shadow agent process in the set of independent computing nodes, configures dual-source data read permissions for the shadow agent process, and extracts the global causal graph from the main working memory area 121. And clone the shadow reasoning graph in the memory address pool of the reflection sandbox area 122. ;

[0155] The asynchronous shadow verification module 150 utilizes the local weight adjacency matrix within the reflection sandbox area 122. Overwrite the shadow reasoning graph The local topological weights of the corresponding node mapping sequence are used by the asynchronous shadow verification module 150 to overwrite the shadow inference graph. Load the decision preference tensor into the shadow agent process;

[0156] The asynchronous shadow verification module 150 is configured with a multi-threaded task scheduler. The multi-threaded task scheduler allocates the dedicated verification dataset according to the tracing flag carried in the standard test frame. The test is split into a counterfactual test subset and a benchmark test subset. The multi-threaded task scheduler slices the counterfactual test subset and the benchmark test subset and distributes them to multiple parallel virtual execution threads.

[0157] The shadow agent process synchronously reads standard test frames in various parallel virtual execution threads. It extracts state observation space features from these frames and uses these features as input variables in the shadow inference graph. Perform forward reasoning calculations of causal relationships within the topological path and output the test action matrix;

[0158] The asynchronous shadow verification module 150 constructs a virtual reward calculation unit in the reflection sandbox area 122. The virtual reward calculation unit receives the test action matrix and calculates the environmental feedback reward value of the current execution step based on the test action matrix and the state transition evaluation model configured in the standard test frame.

[0159] The asynchronous shadow verification module 150 summarizes the environment feedback reward values ​​output by all virtual execution threads and generates a test result set. Test result set The mathematical expression is as follows:

[0160] (16),

[0161] in, For dedicated validation datasets The first in A standard test frame, For the shadow agent process targeting the first The test action matrix output by each standard test frame. The reward evaluation function is built into the virtual reward calculation unit;

[0162] The asynchronous shadow verification module 150 establishes a task execution replay log, which includes a source identification bit, a test action matrix, and a corresponding environmental feedback reward value. The asynchronous shadow verification module 150 stores the task execution replay log in the memory address pool of the reflection sandbox area 122. The asynchronous shadow verification module 150 sends an execution completion signal to the internal scoring calculation component to trigger the comprehensive verification score calculation process.

[0163] See attached document Figure 9 This invention provides an implementation of a weighted calculation model for comprehensive verification scores, which may include: an asynchronous shadow verification module 150 reading task execution replay logs stored in the reflection sandbox area 122, the task execution replay logs containing a set of test results generated in previous steps. And the corresponding traceability identifier;

[0164] The asynchronous shadow verification module 150 parses the source identification bit in the task execution replay log, and then sets up the test results based on the source identification bit. Split into a counterfactual verification reward subset Comparison with benchmark validation subset Counterfactual verification reward subset Includes environmental feedback reward values ​​with counterfactual generation traceability identifiers, and a subset of benchmark verification rewards. Includes environmental feedback reward values ​​with historical baseline traceability identifiers;

[0165] The asynchronous shadow verification module 150 calculates the counterfactual verification reward subset. The arithmetic mean of all environmental feedback reward values ​​is used to output the average reward value for the counterfactual scenario. The asynchronous shadow verification module calculates a subset of the benchmark verification reward. The arithmetic mean of all environmental feedback reward values ​​is used to output the average reward value of the baseline scene. Average reward value in counterfactual scenarios Average reward value compared to benchmark scenario The calculation formula is as follows:

[0166] , (17),

[0167] in, For counterfactual verification reward subset The total number of elements, For counterfactual verification reward subset The single environmental feedback reward value; Validate the reward subset as a benchmark The total number of elements, Validate the reward subset as a benchmark The single environmental feedback reward value;

[0168] The asynchronous shadow verification module 150 extracts the preset verification weight coefficients. Verify the weighting coefficients Set as For real numbers within the range, the asynchronous shadow verification module 150 will average the reward value in the counterfactual scenario. Average reward value in benchmark scenarios and verification weight coefficients Substituting into the formula for calculating the comprehensive verification score, the comprehensive verification score is... The calculation formula is as follows:

[0169] (18);

[0170] The asynchronous shadow verification module 150 extracts the preset system acceptance threshold. The asynchronous shadow verification module compared 150 points to achieve a comprehensive verification score. With system acceptance threshold When the comprehensive verification score Greater than or equal to the system acceptance threshold At that time, the asynchronous shadow verification module 150 generates a graph compilation and merging instruction. The asynchronous shadow verification module 150 sends the graph compilation and merging instruction to the memory management module 120 through the internal communication bus. The graph compilation and merging instruction is used to trigger the memory management module 120 to execute the graph topology update and the writing process of the main working memory area 121.

[0171] When the comprehensive verification score Less than the system acceptance threshold At that time, the asynchronous shadow verification module 150 generates a failure path blocking instruction and sends the failure path blocking instruction to the memory management module 120. The failure path blocking instruction is used to trigger the memory management module 120 to mark the local weight adjacency matrix in the reflection sandbox area 122. This identifies the failed path and triggers resource reclamation of the memory address pool in the reflection sandbox area 122, clearing the local weight adjacency matrix. Data storage.

[0172] See attached document Figure 10 This invention provides an implementation of a graph compilation and merging instruction and a failure path blocking mechanism, which may include:

[0173] The memory management module 120 continuously monitors the internal communication bus of the system and receives the graph compilation and merging instructions or the failure path blocking instructions sent by the asynchronous shadow verification module 150.

[0174] When a graph compilation and merging instruction is received, the memory management module 120 reads the local weight adjacency matrix within the reflection sandbox area 122. With the corresponding node mapping sequence, the memory management module 120 stores the global causal graph in the main working memory area 121. Perform node matching search in the middle;

[0175] For the local weight adjacency matrix A global causal graph exists within it. The newly missed nodes in the global causal graph; memory management module 120. Assign a new node identifier and write the state feature vector of the new node into the underlying graph database of the main working memory area 121;

[0176] For the local weight adjacency matrix With global causal graph For shared overlapping directed edges, the memory management module 120 performs a smooth update calculation of the weight coefficients, and extracts the overlapping directed edges in the global causal graph. Historical weight values and in the local weight adjacency matrix Reflection weight value The memory management module 120 calculates the momentum decay rate based on the preset momentum decay rate. Calculate the final weight value after merging. The calculation formula is as follows:

[0177] (19);

[0178] Among them, momentum decay rate The range of values ​​is The memory management module 120 will assign the final weight value. Overwrite to global causal graph In the adjacency list records;

[0179] When a failed path blocking command is received, the memory management module 120 reads the local weighted adjacency matrix in the reflection sandbox area 122. The memory management module 120 extracts the local weight adjacency matrix. The memory management module 120 extracts the source node identifiers of all reflected directed edges contained therein. Target node identifier and causal type operators ;

[0180] The memory management module 120 stores the source node identifiers according to a preset delimiter. Target node identifier and causal type operators The strings are sequentially concatenated into a joint feature string. The memory management module 120 then calls a secure hash algorithm to perform a hash mapping on the joint feature string, generating a fixed-length hash feature string. Hash feature string The calculation formula is as follows:

[0181] (20);

[0182] in, This represents the string concatenation operator. For secure hash mapping functions;

[0183] The memory management module 120 has a taboo list configured in the main working memory area 121. Taboo List A hash set data structure is used for persistent storage, and the memory management module 120 stores the hash feature string. Add to forbidden list Chinese taboo list Used to intercept the same erroneous logical entity alignment during the subsequent reflection engine generation phase;

[0184] After completing a merge or block operation, the memory management module 120 sends a resource release control signal to the reflection sandbox area 122. Upon receiving the resource release control signal, the reflection sandbox area 122 releases the restrictions on the local weighted adjacency matrix. The allocated memory address pool mapping relationship is reflected. The status flag of the sandbox area 122 is reset to the initial free state. The memory management module 120 revokes the dual-source data read permission of the shadow agent process and calls the operating system interface to terminate the running thread of the shadow agent process in the set of independent computing nodes.

Claims

1. A self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms, characterized in that, include: The task execution monitoring module (110) is configured to collect the environmental state transition sequence and the task score fed back by the environment during the intelligent agent's task execution process in real time. The memory management module (120) includes a main working memory area (121) and a reflection sandbox area (122) divided at the logical storage level. The main working memory area (121) stores a global causal graph through a graph database, and the reflection sandbox area (122) stores improvement reflection rules to be verified. The cognitive conflict detection module (130) is configured to perform a topological comparison between the local causal subgraph generated in the reflection sandbox area (122) and the existing rules in the main working memory area (121); The counterfactual scenario generation module (140) is configured to construct a counterfactual boundary test scenario based on the topological difference points located by the cognitive conflict detection module (130) by perturbing the precondition parameters of the conflict nodes. The asynchronous shadow verification module (150) is configured to instantiate a shadow agent in an isolated computing environment, perform regression verification on the policy in the reflection sandbox area (122) using the counterfactual boundary test scenario, and return a merge control signal or a failure path blocking signal to the memory management module (120) based on the verification result.

2. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 1, characterized in that, The task execution monitoring module (110) is configured with dual threshold determination logic: Calculate the cumulative reward value after the task execution trajectory ends. When the cumulative reward value The score is below the preset minimum for task success. At that time, a reflection initiation signal is issued; Calculate the current cumulative reward value Average cumulative reward value over historical execution cycles volatility ratio When the volatility ratio Exceeding the preset performance jitter threshold At the same time, a reflection mechanism is triggered.

3. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 1, characterized in that, The memory management module (120) allocates an independent memory mirror space within the reflection sandbox area (122) through the memory page table isolation mechanism at the operating system level, and sets the local causal subgraph using a local weighted adjacency matrix. It is injected into the memory image space in the form of [a specific type of injection].

4. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 1, characterized in that, The cognitive conflict detection module (130), when performing topological comparison, extracts the central anchor point of the local causal subgraph and performs a depth-based comparison in the global causal graph. A breadth-first search is performed, and data filtering is carried out using an attribute matching indicator function to segment reference local subgraphs from the global causal graph. .

5. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 4, characterized in that, The cognitive conflict detection module (130) transforms the local causal subgraph into a local weighted adjacency matrix based on a node mapping sequence of a unified dimension. The reference local subgraph Transform into a reference adjacency matrix And calculate the topological difference in the form of Frobenius norm. The calculation formula is as follows: , in, The total number of nodes in the union of the nodes. and They are respectively the first in the corresponding matrix Line 1 The element values ​​of the column.

6. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 5, characterized in that, When the topological difference When the value exceeds the preset topological tolerance threshold, the cognitive conflict detection module (130) traverses the elements in the difference matrix whose absolute value exceeds the unilateral anomaly threshold and generates a core conflict node set through reverse mapping. And further obtain the set of precondition nodes corresponding to the in-degree edges. The set of subsequent affected nodes corresponding to out-degree edges The union of the nodes generates the set of conflict boundary nodes. .

7. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 1, characterized in that, The counterfactual scenario generation module (140) is configured with a variety of perturbation operators: The numerical out-of-bounds perturbation operator is used to calculate negative and positive out-of-bounds values ​​for historical numerical ranges of continuous environmental variables. Logical flip perturbation operator is used to extract complement category labels for discrete environmental variables; The conditional concealment perturbation operator is used to replace the constraint value of a selected environment variable with the null identifier NULL according to a preset concealment probability.

8. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 1, characterized in that, The counterfactual scenario generation module (140) generates the counterfactual boundary test scenario set. Compared with the set of benchmark normal scenarios extracted from the global high-order benchmark database Perform set union operation to build a dedicated validation dataset. and for the dedicated verification dataset Each piece of scene data is appended with an independent binary traceability identifier.

9. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 8, characterized in that, The asynchronous shadow verification module (150) calculates the average reward value for the counterfactual scenario based on the test results of the source tracing identifier. Average reward value compared to benchmark scenario And calculate the comprehensive verification score. The calculation formula is as follows: , in, The preset verification weight coefficients, and .

10. The self-evolving artificial intelligence system based on memory enhancement and reflection mechanisms according to claim 9, characterized in that: When the comprehensive verification score When the momentum decay rate is greater than or equal to the system acceptance threshold, the memory management module (120) determines the system's response based on the preset momentum decay rate. Update the weight coefficients of overlapping directed edges and overwrite the updated final weight values ​​into the global causal graph; when the comprehensive verification score... When the value is less than the system acceptance threshold, the memory management module (120) maps the feature string hash of the corresponding reflection directed edge to a hash feature string and appends it to the taboo list stored in the hash set data structure. middle.

Citation Information

Patent Citations

  • Data enhancement method and system based on multi-agent self-evolution and hybrid evaluation

    CN121211014A

  • Virtual energy consumption simulation system and method based on generative agent

    CN121302848A