A world behavior model self-evolution control method and system based on four-map linkage

CN122840104APending Publication Date: 2026-09-29LOVE AVATAR TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611317789.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供一种基于四图联动的世界行为模型自演化控制方法及系统,旨在解决现有智能体治理中反馈未直接驱动图结构更新、图修改依赖人工配置、跨图依赖易失效且缺少验证门控的问题

Benefits of technology

(一)实现反馈数据向结构化控制信号的有效转换,提高行为修正的精准性:本发明将用户评价、人工纠偏、业务结果和运行指标等反馈数据映射至反馈图,并通过反馈归因分析计算反馈节点与状态节点、行为节点及副作用节点之间的归因分值,使反馈能够被量化为具体的边权更新、节点新增、节点拆分、路径重排或异常分支生成等结构化修改指令,克服了现有技术中反馈仅被用作评分或样本标签而无法直接驱动控制策略调整的缺陷,实现了对智能体长期运行中积累的行为偏差的精准识别和修正。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840104A_ABST
    Figure CN122840104A_ABST
Patent Text Reader

Abstract

This invention discloses a self-evolutionary control method and system for a world behavior model based on four-graph linkage, belonging to the field of intelligent agent control technology. The method includes: collecting intelligent agent operation events and mapping them to a world behavior model composed of a state graph, behavior graph, side effect graph, and feedback graph; calculating behavioral deviations based on incremental events from the four graphs and performing attribution analysis on feedback events; generating a multi-graph change plan when evolution triggering conditions are met; performing cross-graph dependency verification and reconstruction atomically to generate candidate versions; performing structural verification, historical replay, and shadow operation verification on the candidate versions; if verification passes, releasing it as a new stable version and retaining a rollback pointer; otherwise, rolling back. This invention achieves feedback-driven continuous evolution through four-graph linkage and can be applied to scenarios of long-term intelligent agent operation governance and multimodal model fine-tuning management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent agent control technology, and in particular to a self-evolutionary control method and system based on a four-graph linkage world behavior model. Background Technology

[0002] As large language models, multimodal models, and agent systems evolve from single-turn question answering to long-term task execution, cross-tool collaboration, cross-modal perception, and continuous feedback optimization, the system's outcome is no longer determined by a single model inference. Instead, it is influenced by the task state, environmental conditions, behavioral choices, tool side effects, user feedback, and historical policy versions. Relying solely on model parameter fine-tuning, static rule bases, or single-feedback optimization is insufficient to continuously describe and correct the behavioral biases that gradually accumulate in agents over long-term operation.

[0003] In existing intelligent agent governance systems, the state management module is typically responsible for representing task state transitions, the behavior description module or workflow engine is responsible for representing action paths, audit logs are used to record tool calls and execution results, the feedback system is used to save user ratings or manual corrections, and the multimodal model fine-tuning platform converts some samples into training data. While each of these modules can perform its local function, they lack a unified joint representation and linkage update mechanism for state-behavior-side effects-feedback.

[0004] In the existing systems described above, feedback information is typically used only for scoring, sample labeling, or manual review results, and is not used to drive structured adjustments to control strategies. Updates to control strategies or graph structures are mostly done offline manually, lacking unified triggering rules for edge weight updates, node additions and deletions, path rearrangements, and abnormal branch generation. When a behavior path undergoes local adjustments, the associated side effect records and feedback attribution relationships may not be synchronously corrected, leading to inconsistencies in cross-module dependencies. Furthermore, after automatic adjustments to control strategies, existing systems lack scenario replay verification, simulation verification, risk gating, and version rollback mechanisms for candidate versions, making it difficult to guarantee the stability of automatic adjustment results in real-world operating environments. Current understandings of "continuous learning" are generally limited to the continued training of model parameters and have not yet fully addressed the engineering implementation of continuous updates to the agent's control strategy structure. Summary of the Invention

[0005] The purpose of this invention is to provide a self-evolutionary control method and system for a world behavior model based on four-graph linkage, aiming to solve the problems in existing intelligent agent governance where feedback does not directly drive graph structure updates, graph modifications rely on manual configuration, cross-graph dependencies are prone to failure, and there is a lack of verification gating.

[0006] To achieve the above objectives, this invention provides a self-evolutionary control method based on a four-graph linkage world behavior model, comprising the following steps: S1. Collect the running events of the intelligent agent, standardize the running events, and map them to the world behavior model composed of state graph, behavior graph, side effect graph and feedback graph, generate four-graph incremental events and establish cross-graph associations; S2. Based on the incremental events in the four graphs, calculate the behavioral deviation of the current running path relative to the expected path of the stable version; and based on the feedback events in the feedback graph, determine the candidate nodes and edges associated with the feedback events in the state graph, behavior graph, and side effect graph, perform attribution analysis on the candidate nodes and edges, and generate attribution results. S3. When the behavioral deviation and attribution results meet the preset evolution trigger conditions, a multi-graph change plan is generated. The multi-graph change plan includes structured modification operations on the nodes and edges of at least one graph in the state graph, behavior graph, side effect graph, and feedback graph, as well as the execution dependency order of the modification operations. S4. Based on the multi-graph change plan, perform structured modification operations in an atomic manner to perform cross-graph dependency verification and reconstruction on the state graph, behavior graph, side effect graph, and feedback graph, and generate candidate world behavior model versions; S5. Perform structural consistency verification, replay verification based on historical running data, and shadow running verification without affecting the real environment on the candidate world behavior model version, and generate verification results. S6. When the verification results meet the preset hard constraints and release conditions, release the candidate world behavior model version as a new stable version and retain the rollback pointer; when the verification results do not meet the hard constraints or release conditions, roll back to the stable version.

[0007] Preferably, the behavioral deviation in S2 is calculated according to the following formula: ; in, For path Behavioral deviation score; The deviation between the actual state and the expected state; The difference between the actual behavior sequence and the reference path; Considering the risks of side effects and the intensity of transmission; To provide feedback on the discrepancy between the results and the objectives; This is due to performance drift compared to relatively stable historical versions; , , , , These are the weighting coefficients.

[0008] Preferably, the attribution analysis described in S2 is performed in the following manner: For feedback nodes and candidate nodes Calculate the attribution score according to the following formula. : ; in, For time proximity, For dependency degree, For counterfactual discrepancy, For the credibility of the evidence, For the correlation of side effect transmission pathways, to These are the weighting coefficients; When the attribution score exceeds the attribution threshold and the number of supporting samples exceeds the minimum support threshold, the candidate node will be... Add to the evolutionary candidate set.

[0009] Preferably, the evolution triggering condition described in S3 is determined by the following combination of conditions: ; in, Scoring for behavioral deviations This is the threshold for deviation strength. This represents the current cumulative number of supporting samples. The minimum support sample threshold; For the credibility of the evidence, As a threshold for the credibility of evidence; This indicates that the cooldown period has ended. The cooldown period started counting from the last evolution trigger or version release. For logical AND operator.

[0010] Preferably, the cross-graph dependency verification and reconstruction in S4 includes edge weight updates, which are performed according to the following formula: ; in, For the first wheel edge The weights; The attenuation coefficient; The learning rate is positively rewarded. To provide feedback on credibility; To provide positive incentives; This is the deviation penalty coefficient; This is a penalty item for deviations; This is the risk penalty coefficient; Side effects risk item; This is a truncation function.

[0011] Preferably, the cross-graph dependency verification and reconstruction in S4 also includes node addition, and the addition score is calculated according to the following formula: ; in, For the new model The new score; for With existing nodes The maximum similarity; To support the sample size; For the expected impact; , , These are the weighting coefficients.

[0012] Preferably, the hard constraints described in S6 include one or more of the following, and releasing a new version requires satisfying all configured hard constraints: The failure rate of critical tasks does not exceed a preset threshold; High-risk side effects must not be increased; Permissions and data boundary checks passed; There are no dangling nodes or loop conflicts in the four-graph reference; Key scene playback passed; Paths marked as not automatically modifiable were not deleted.

[0013] Preferably, after a new stable version is released in S6, it also includes: Reward values ​​are generated based on the actual running results of the new stable version; Based on the reward value, the weight coefficients in the calculation of behavioral bias, the weight coefficients in the attribution analysis, the threshold values ​​in the evolution triggering conditions, and the learning rate or verification strategy for edge weight updates are automatically updated. If the candidate version verification fails and a rollback is performed, the current control parameters remain unchanged or the control parameters are updated based on the running data of the stable version after the rollback.

[0014] Secondly, the present invention also provides a self-evolving control system based on a four-graph linkage world behavior model, comprising: The event acquisition module is used to collect the running events of the intelligent agent and perform standardized processing. The four-graph mapping module is used to map standardized runtime events to state graphs, behavior graphs, side effect graphs, and feedback graphs, establish cross-graph relationships, and generate incremental events in the four graphs. The deviation identification module is used to receive incremental events from four graphs, compare the current running path with the expected path of the stable version, calculate the behavior deviation score, and output the deviation object. The feedback attribution module is used to receive feedback events and node and edge information of the four graphs, calculate attribution scores for feedback events in the feedback graph and candidate nodes and edges in the state graph, behavior graph, and side effect graph, and output the attribution results. The change plan generation module is used to obtain deviation objects and attribution results, and generate multi-graph change plans when the evolution triggering conditions are met. The multi-graph reconstruction module is used to perform structured modification operations atomically based on the multi-graph change plan, perform cross-graph dependency verification and reconstruction, and generate candidate world behavior model versions. The version verification module is used to perform structural consistency verification, replay verification based on historical running data, and shadow running verification without affecting the real environment on the candidate world behavior model version, and generate verification results. The rollback module is used to obtain the verification results. When the verification is successful, a new version is released and the rollback pointer is retained. When the verification fails, the rollback is executed to a stable version. The continuous learning module is used to update control parameters based on the actual running results of the new version or the running data of the stable version after rollback.

[0015] Therefore, the self-evolutionary control method and system based on a four-graph linkage world behavior model, as described above, have the following beneficial effects: (i) Effective conversion of feedback data into structured control signals to improve the accuracy of behavior correction: This invention maps feedback data such as user evaluation, manual correction, business results and operation indicators to a feedback graph, and calculates the attribution scores between feedback nodes and state nodes, behavior nodes and side effect nodes through feedback attribution analysis. This enables feedback to be quantified into specific structured modification instructions such as edge weight update, node addition, node splitting, path rearrangement or abnormal branch generation. This overcomes the defect in the prior art where feedback is only used as a score or sample label and cannot directly drive the adjustment of control strategy. It realizes the accurate identification and correction of behavior deviations accumulated in the long-term operation of the intelligent agent.

[0016] (ii) Ensuring consistency of four-graph linkage through cross-graph atomic reconstruction and reducing the risk of system conflicts caused by local updates: This invention uses a multi-graph change plan to uniformly carry out structured modification operations on the state graph, behavior graph, side effect graph, and feedback graph and their execution dependency order. Cross-graph reconstruction is performed atomically within the candidate version transaction, and cross-graph dependency verification is performed during the reconstruction process. This ensures that when adding a behavior node, the preceding state association, potential side effect association, and feedback collection association are established simultaneously. When deleting or merging nodes, the version mapping and evidence reference are retained. This overcomes the defects of the prior art in which local graph modification leads to cross-graph dependency failure, causal break, and strategy conflict, and ensures the consistency and integrity of the association relationship between the four graphs.

[0017] (III) Constructing a verifiable and rollbackable self-evolutionary engineering control link to improve the stability and security of long-term operation: This invention sequentially performs structural consistency verification, replay verification based on historical operation data, and shadow operation verification in a real environment for candidate world behavior model versions. It also sets hard constraints such as critical task failure rate, high-risk side effects, permission boundaries, dangling references, and key scenario replay. Release is only allowed when all hard constraints are met and the comprehensive verification score reaches the release threshold. When verification fails, the system quickly restores to the parent stable version through the rollback pointer. This overcomes the shortcomings of existing technologies that lack version verification and rollback mechanisms and cannot guarantee the stability and controllability of self-evolution results. At the same time, the control parameters are continuously updated based on the real operation results after release, realizing continuous optimization within the safety boundary.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a self-evolutionary control method for a world behavior model based on four-graph linkage, according to the present invention. Detailed Implementation

[0020] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Example 1 A self-evolving control method based on a four-graph linkage world behavior model is proposed. In this embodiment, the world behavior model is a versioned graph model composed of a state graph, a behavior graph, a side effect graph, and a feedback graph. It is used to describe the state conditions, optional behaviors, behavioral side effects, feedback results, and their associated rules of an agent in a specific business domain and operating environment. The world behavior model can be represented as: ; in, For the first Version status diagram For behavior graphs, This is a diagram showing the side effects. For feedback charts, For cross-graph mapping and control rule set, For version, evidence, verification, and release information, four types of graphs establish cross-graph associations (i.e., cross-graph mapping) through unified task identifiers, agent instance identifiers, time windows, event identifiers, evidence summaries, and version identifiers. The system allows different graphs to use different storage engines, but uses unified addressing and change plan objects at the control layer to ensure that modifications to the four graphs in the same evolution round can be committed, verified, and rolled back as a whole.

[0023] The intelligent agent involved in this embodiment can be an autonomous decision-making system based on a large language model or a multimodal model, or it can be a project management intelligent agent, a multimodal customer service intelligent agent, a process intelligent agent, or a control module in a model fine-tuning management platform. The intelligent agent completes its predetermined tasks by calling external tools, executing action sequences, and receiving environmental feedback and user evaluations.

[0024] like Figure 1 As shown, the present invention specifically includes the following steps: S1. Collect the running events of the intelligent agent, standardize the running events, and map them to the world behavior model composed of state graph, behavior graph, side effect graph and feedback graph, generate four-graph incremental events and establish cross-graph associations; Specifically, S1 is executed by the event acquisition module and the four-graph mapping module working together.

[0025] S11. Collect operational events of the agent. The operational event collection module receives various raw events generated during the operation of the agent in a streaming or batch manner. Operational events include at least the following categories: (1) State events: These include task phase switching, environment state changes, resource state changes, context state updates, and model state changes. For example, "the project phase switches from development to testing" in the project management agent is a state event.

[0026] (2) Behavioral events: These include decision records, action execution records, tool call records, subtask start and end records, and control node jump records. For example, when an agent calls an external API to query a database or executes a code snippet to complete data analysis, these are all behavioral events.

[0027] (3) Side effect events: These include data change records, permission impact records, resource consumption records, external system impact records, and risk event records. For example, if an agent performs a deletion operation that reduces the number of database records, or calls a high-cost computing service that consumes resource quotas, these are all side effect events.

[0028] (4) Feedback events: These include user evaluations, manual correction records, business result indicators, quality indicator scores, and failure sample markings. For example, a user clicking "unsatisfied" with the agent's answer and submitting a revised text, or a human reviewer revising the automated approval result, are all feedback events.

[0029] S12. Standardize the runtime events. Standardization refers to adding a unified timestamp, task identifier, agent instance identifier, session identifier, event type label, and version identifier to each original event, forming a standardized runtime event object. The standardized runtime events have a unified data structure, which facilitates consumption by the subsequent four-graph mapping module.

[0030] The data structure for standardized runtime events is shown in Table 1: Table 1

[0031] S13. Map standardized operational events to state diagrams, behavior diagrams, side effect diagrams, and feedback diagrams to generate four-graph incremental events. The four-graph mapping module distributes standardized events to the corresponding graph storage structures according to event type. The four graphs and their technical functions in this embodiment are shown in Table 2: Table 2

[0032] The four graphs mentioned above are not statically stored, independent data sets, but are interconnected through a unified addressing mechanism and a cross-graph mapping table. Specifically, state nodes in the state graph can be referenced by behavior nodes through precondition edges in the behavior graph; behavior nodes in the behavior graph can be associated with side effect nodes through "generation" edges in the side effect graph; and feedback nodes in the feedback graph can point to state nodes or behavior nodes through attribution edges. A four-graph incremental event refers to the smallest unit of operation that performs a single node addition, edge update, or attribute modification on any graph.

[0033] S14. Establish cross-graph associations (i.e., cross-graph mapping). The four-graph mapping module establishes cross-graph reference relationships for nodes and edges in the four graphs while generating incremental events. Cross-graph associations include at least: pre-constraint associations between state nodes and action nodes, causal associations between action nodes and side-effect nodes, evidence associations between action nodes and feedback nodes, and attribution associations between feedback nodes and state nodes. Cross-graph associations are stored in a mapping table, with each record containing the source graph type and node identifier, the target graph type and node identifier, the association type, and the association weight.

[0034] S2. Based on the incremental events in the four graphs, calculate the behavioral deviation of the current running path relative to the expected path of the stable version; and based on the feedback events in the feedback graph, determine the candidate nodes and edges associated with the feedback events in the state graph, behavior graph, and side effect graph, perform attribution analysis on the candidate nodes and edges, and generate attribution results. Specifically, S2 is executed by the deviation identification module and the feedback attribution module working together.

[0035] S21. Based on the four graph incremental events, calculate the behavioral deviation of the current running path relative to the expected path of the stable version.

[0036] The deviation identification module reads the current running path from the four-graph incremental events after each task window or time window ends. The actual record is retrieved, and the corresponding expected path reference is read from the stable version of the world behavior model. For the path Behavioral deviations are calculated using the following formula: ; in, For path Behavioral deviation score (dimensionless score, the larger the value, the more significant the deviation). The deviation between the actual state and the expected state is calculated based on the Euclidean distance or Hamming distance between the actual state vector and the expected state region. To determine the difference between the actual behavior sequence and the reference path, the degree of difference between the actual behavior sequence and the high-quality reference path is calculated based on the edit distance or the length of the longest common subsequence. To assess the risk and intensity of side effects, the weights of the risk edges associated with the current behavioral path in the side effect graph are cumulatively calculated. To determine the discrepancy between the feedback results and the target, calculations are made based on the number and severity of negative feedback associated with the current path in the feedback graph. For performance drift relative to historical stable versions, it is calculated based on the degree of deviation of the current path from the previous stable version in key performance indicators (success rate, latency, cost); , , , , These are weighting coefficients used to adjust the contribution of various deviations to the overall score. Each weighting coefficient is a non-negative real number and can be configured according to the business domain or updated by the continuous learning module based on historical verification results.

[0037] Taking a project management agent as an example, if the agent chooses to "close the task directly" when the "dependency task is not confirmed" state, but the expected path of the stable version requires "dependency check" before making a decision in the same state, then... This reflects the degree of state matching (in this example, the degree of state matching is relatively high). This reflects the sequence difference between "direct shutdown" and "decision based on post-check" (the difference is relatively large in this case). This reflects the risks of rework and delays caused by "direct closure" (the risk is higher in this case). The records of manual corrections reflecting this user behavior (in this case, there is clear negative feedback) are summarized. This will exceed the normal threshold.

[0038] S22. Based on the feedback events in the feedback graph, determine the candidate nodes and edges associated with the feedback events in the state graph, behavior graph, and side effect graph.

[0039] The feedback attribution module reads feedback nodes from the feedback graph. (For example, a user review stating "The category result is incorrect"), and based on cross-graph associations, tracing the state nodes, behavior nodes, and side effect nodes associated with this feedback node to form a candidate node set. For each candidate node... The feedback attribution module further collects the following five types of evidence: (1) the time interval between the feedback node and the candidate node; (2) whether the candidate node is located on the behavior path evaluated by the feedback node; (3) whether there is a direct cross-graph association edge between the feedback node and the candidate node; (4) the credibility of the feedback source (e.g., the authority of manual review is higher than that of automatic scoring); (5) the length of the side effect propagation path between the feedback event and the candidate node.

[0040] S23. Perform attribution analysis on candidate nodes and edges to generate attribution results.

[0041] For feedback nodes and candidate nodes Calculate the attribution score according to the following formula. : ; in, Temporal proximity represents the proximity between the occurrence time of the feedback event and the candidate node, with a value in the range [0,1]. The shorter the time interval, the higher the value. The dependency degree represents the length and strength of the dependency chain between the candidate node and the path containing the feedback event. Its value is in the range [0,1], and the higher the value, the more direct the path dependency. The counterfactual variance is the probability estimate that the feedback result will change significantly if the candidate node does not occur. It takes values ​​in the range [0,1]. A higher value indicates a more critical contribution of the candidate node to the feedback result. The credibility of the evidence represents the credibility of the feedback source, with a value in the range [0,1]. Human review has a higher credibility than automatic scoring. The side effect propagation path correlation represents the path strength of the side effects generated by the candidate node propagating to the feedback observation point through the side effect graph, and its value is in the interval [0,1]. to These are weighting coefficients, all of which are non-negative real numbers and can be updated by the continuous learning module. When the attribution score exceeds the attribution threshold and the number of supporting samples exceeds the minimum support threshold, the candidate node will be... Add to the evolutionary candidate set. An attribution score exceeding the attribution threshold indicates a statistically significant causal relationship between the candidate node and the feedback; a number of supporting samples exceeding the minimum support threshold indicates that the attribution relationship has been repeatedly verified in different time windows or different agent instances, and is not a random event.

[0042] Both the attribution threshold and the minimum support threshold are system configuration parameters. The typical value range for the attribution threshold is [0.3, 0.7], and the typical value for the minimum support threshold is 3 to 10 independent samples. When the attribution scores of multiple candidate nodes are close (i.e., the difference between the maximum attribution score and the second largest attribution score is less than the preset competition threshold), the system retains multiple candidate relationships and compares and verifies them in subsequent version verification stages to avoid selecting the wrong attribution direction too early.

[0043] S3. When the behavioral deviation and attribution results meet the preset evolution trigger conditions, a multi-graph change plan is generated. The multi-graph change plan includes structured modification operations on the nodes and edges of at least one graph in the state graph, behavior graph, side effect graph, and feedback graph, as well as the execution dependency order of the modification operations. Specifically, S3 is executed by the change plan generation module.

[0044] S31. Determine the evolution triggering condition.

[0045] The system initiates evolution when all of the following trigger thresholds are met. Evolution trigger conditions are determined by the following combination of conditions: ; in, Scoring for behavioral deviations This is a threshold for the intensity of deviation, meaning that evolution is only triggered when the behavioral deviation reaches a certain level, thus avoiding overreaction to minor fluctuations. This refers to the current accumulated number of supporting samples, that is, the number of independent evidence samples (such as feedback events, failure records, and manual correction records) that support this round of evolution. To ensure sufficient statistical basis for evolutionary decisions, a minimum support sample threshold is set. The credibility of the evidence is the weighted credibility score of all supporting samples. This serves as a threshold for the credibility of evidence, ensuring that the feedback sources upon which evolutionary decisions are based are sufficiently authoritative. This indicates that the cooldown period has ended. The cooldown period starts from the last time an evolution was triggered or a version was released. It is used to prevent the system from frequently triggering evolutions in a short period of time, which could lead to version instability. This is the logical AND operator, meaning all conditions must be met simultaneously.

[0046] Deviation strength threshold The typical value range is [0.2, 0.5]; minimum support sample threshold Typical values ​​range from 5 to 20; the threshold for the credibility of evidence. Typical values ​​are [0.4, 0.8]; typical cooldown time is set to 1 hour to 24 hours, with the specific value determined based on the sensitivity of the business scenario and the frequency of tasks.

[0047] S32. Generate a multi-graph change plan.

[0048] When the evolution triggering conditions are met, the change plan generation module generates a multigraph evolution plan (MultiGraphEvolutionPlan) based on the deviation object and attribution results. The multigraph evolution plan contains structured modification operations on nodes and edges of at least one graph in the state graph, behavior graph, side effect graph, and feedback graph, as well as the execution dependency order of the modification operations.

[0049] The data structure of the multi-graph change plan is shown in Table 3: Table 3

[0050] The types of structured modification operations include, but are not limited to: edge weight update (increasing or decreasing the selection probability or constraint strength of an edge), node addition (creating new state nodes, behavior nodes, side effect nodes, or feedback nodes in the graph), node splitting (splitting a node into two or more child nodes according to the sample distribution), node merging (merging two or more semantically similar nodes into one node), edge deletion (removing invalid or harmful paths), path reordering (adjusting the execution order of behavior nodes), and abnormal branch generation (adding blocking, manual takeover, compensation, or rollback nodes).

[0051] Different deviations and feedback types correspond to different candidate control actions. The typical correspondence in this embodiment is shown in the table below: Table 4

[0052] S4. Based on the multi-graph change plan, perform structured modification operations in an atomic manner to perform cross-graph dependency verification and reconstruction on the state graph, behavior graph, side effect graph, and feedback graph, and generate candidate world behavior model versions; Specifically, S4 is executed by the multi-graph reconstruction module.

[0053] S41, Cross-graph dependency validation.

[0054] Before formally executing the modification, the multi-graph reconstruction module first performs dependency verification on the multi-graph change plan. The verification includes: (1) whether all source nodes referenced in the change plan exist and are accessible in the current stable version; (2) whether the identifiers of all newly added nodes conflict with existing nodes; (3) whether all edges to be deleted have no valid references from other graphs (for example, the edge to be deleted may still have risk edges from the side effect graph pointing to it); (4) whether there are no circular dependencies in the dependency order; and (5) whether all hard constraints can still be satisfied after the modification.

[0055] The specific rules for cross-graph dependency validation are shown in the table below: Table 5

[0056] S42. Perform structured modification operations.

[0057] The multi-graph reconstruction module executes the modification operations in the multi-graph change plan in the order of dependencies.

[0058] (a) Update of border rights.

[0059] For candidate edges The system updates edge weights based on the current edge weights, positive execution rewards, deviation penalties, side effect risks, feedback reliability, and time decay. Edge weight updates are performed according to the following formula: ; in, For the first wheel edge The weights (i.e., the edge weights in the current stable version) take values ​​in the range [0,1]. This is the attenuation coefficient, with a value in the range (0,1), used to reduce the historical cumulative weight to adapt to current environmental changes; The learning rate is a positive reward, and it is a non-negative real number that controls the step size of the positive feedback in increasing the edge weights. To assess the reliability of the feedback, the value is taken in the range [0,1], and is determined by the type of feedback source, the manual review marker, and the consistency check. Positive performance rewards are calculated based on a combination of positive results such as increased task success rate, reduced costs, and improved user satisfaction. is the deviation penalty coefficient, which is a non-negative real number that controls the weight of the deviation penalty term; This is a deviation penalty, calculated based on the degree of deviation between the current path and the expected path; is the risk penalty coefficient, a non-negative real number, which controls the weight of the risk penalty for side effects; The side effect risk item is calculated based on the risk propagation strength associated with the current edge in the side effect graph; This is a truncation function that limits the calculation results to the interval [0,1].

[0060] Edge weight represents the priority or probability of choosing a certain action edge in a given state. A higher edge weight indicates a greater likelihood that the action path will be selected during the agent's decision-making process. The physical meaning of edge weight update is: a successful path receives a positive reward, thereby increasing the edge weight. Item), the deviation path is penalized, thus reducing the edge weight ( Item), high-risk paths are subject to risk penalties, thereby reducing edge weights ( Item), outdated experience decays over time ( item).

[0061] (ii) Adding new nodes.

[0062] When the new model If the maximum similarity with an existing node is below the new node threshold, and the number of supported samples reaches the threshold and the expected effect passes the initial validation, a new node can be added. The new node score is calculated using the following formula: ; in, For the new model The new score; for With existing nodes The maximum similarity; To support the sample size; For the expected impact; , , These are the weighting coefficients.

[0063] Typical values ​​for the new threshold are [0.6, 0.85]. A node addition operation is performed when the new score exceeds the new threshold. When adding a new behavior node, the system simultaneously establishes the node's preceding state association (adding a corresponding state satisfaction edge to the state graph), potential side effect association (adding a generation edge to the side effect graph pointing to the expected side effect node), and feedback collection association (adding a collection point edge to the feedback graph) to ensure that the new node has complete dependencies in all four graphs.

[0064] (III) Node splitting.

[0065] Node splitting is performed when samples within a node form two or more stable distributions and the distribution differences exceed the splitting threshold. The multi-graph reconstruction module first performs cluster analysis on the associated samples of the target node. When the clustering results show significant separation (inter-cluster distance greater than three times the intra-cluster distance), the original node is split into two child nodes, and the associated edges and cross-graph references of the original node are redistributed according to sample affiliation. The original node is retained as a virtual parent node, storing the list of child nodes and the splitting history.

[0066] (iv) Node merging.

[0067] Node merging is performed when the comprehensive similarity of two nodes across four dimensions—semantic definition, state conditions, behavioral parameters, and feedback distribution—all exceed the merging threshold, and the merged node does not violate any hard constraints. The multi-graph reconstruction module creates a new node, migrates all associated edges and cross-graph references of the two old nodes to the new node, merges the evidence sets of the two nodes, and stores a version mapping pointer to the new node in the old node.

[0068] (v) Delete.

[0069] An edge is marked as a candidate for deletion when its weight remains below the deletion threshold for an extended period (e.g., below 0.1 for 100 consecutive task windows) and it has no critical version dependencies (i.e., no cross-graph references to the edge in historically stable versions). If the edge is not reactivated by any feedback event or successful path within a subsequent cooldown period, deletion is performed.

[0070] (vi) Path rearrangement and abnormal branch generation.

[0071] When the quality benefits of the alternative path exceed the switching threshold, the sequential edges of the behavior graph are updated and the side effect propagation path is recalculated. When the system identifies repeated failures, permission violations, or unrecoverable side effects, a blocking node (preventing entry into the violation path), a manual takeover node (transferring control to manual review), or a compensation rollback node (performing compensation operations and then rolling back to a safe state) is added.

[0072] S43. Generate candidate world behavior model versions.

[0073] After all modifications are completed, the multi-graph reconstruction module encapsulates the modified four graphs, cross-graph mappings, and control rules into a candidate world behavior model version object (WorldBehaviorModelVersion). This version object contains the following core fields, as shown in Table 6: Table 6

[0074] If any hard constraint verification fails during execution (e.g., permission boundaries are breached, or critical task paths are accidentally deleted), the multi-graph reconstruction module will undo all modifications made in this round, restore the four graphs to their state before execution, and record the reason for the failure for subsequent analysis.

[0075] S5. Perform structural consistency verification, replay verification based on historical running data, and shadow running verification without affecting the real environment on the candidate world behavior model version, and generate verification results. Specifically, S5 is executed by the version verification module.

[0076] S51, Structural consistency check.

[0077] The version verification module first performs a structural consistency check on the candidate version, checking the following items: (1) all node references in the candidate four graphs point to valid nodes, and there are no dangling references; (2) the direction of all edges conforms to the graph model definition (e.g., state transition edges must have a clear target state); (3) all source nodes and target nodes referenced in cross-graph mappings exist in the corresponding graphs; (4) all necessary fields in the version object are filled in completely; (5) there are no circular dependencies (e.g., state A depends on state B, and state B depends on state A). If any check fails, the candidate version is directly marked as structurally invalid and will not proceed to the subsequent verification stage.

[0078] S52. Replay verification based on historical operation data.

[0079] The version verification module extracts three types of scenario snapshots from historical running data: (1) historical successful scenarios (tasks completed smoothly, user feedback positive); (2) historical failed scenarios (tasks failed or user feedback negative); and (3) boundary scenarios (extreme conditions such as resources nearing exhaustion, permission criticality, and timeout edge). For each type of scenario, the version verification module re-executes the decision path using the candidate world behavior model version, compares the difference between the candidate version's decision and the actual historical decision, and simulates and evaluates the candidate version's task completion probability, side effect risk, and feedback expectation in that scenario. When the key scenario replay pass rate (i.e., the proportion of scenarios where the simulation result reaches or exceeds the historical best result) reaches a preset threshold, the replay verification is successful.

[0080] The typical threshold for the pass rate of critical scene replays is 90% to 95%, meaning that the candidate version performs no worse than the current stable version in at least 90% of historical critical scenes.

[0081] S53, Shadow Run Verification.

[0082] Under the premise that the actual execution results are not affected in the real environment, the version verification module deploys the candidate world behavior model version to the shadow runtime verification environment. The shadow runtime verification environment receives the same input event stream as the production environment, but the candidate version only performs decision calculations and path simulations, and does not actually execute operations that have side effects on external systems. The shadow runtime verification continuously monitors the stability and anomaly rate of the candidate version within a preset time window (e.g., 24 hours or 100 task instances). When the candidate version's critical task failure rate, side effect anomaly rate, and feedback consistency in the shadow runtime verification all reach the release threshold, the shadow runtime verification passes. If a stopping condition is triggered during the shadow runtime verification process (e.g., the anomaly rate exceeds the preset threshold or an unexpected high-risk side effect path is detected), the shadow runtime verification is immediately terminated and the candidate version is marked as failing.

[0083] S54. Generate verification results.

[0084] After the structural consistency verification, historical replay verification, and shadow run verification all pass, the version verification module calculates the comprehensive verification score of the candidate version according to the following formula. : ; in, The increase in task success rate (the increase in success rate of the candidate version compared to the stable version). For behavior path stability (the inverse variance or consistency rate of behavior paths in candidate versions); This represents the change in side effect risk (the risk increment of the candidate version compared to the stable version). For feedback consistency (the degree to which the percentage of positive feedback shown in the feedback graph closely matches the expected target); This represents the change in resource costs (the increase in unit task cost of the candidate version compared to the stable version). Version drift (the degree of deviation of a candidate version from a stable version in terms of key states and behavior distribution); to These are weighting coefficients, all of which are non-negative real numbers.

[0085] S6. When the verification results meet the preset hard constraints and release conditions, release the candidate world behavior model version as a new stable version and retain the rollback pointer; when the verification results do not meet the hard constraints or release conditions, roll back to the stable version.

[0086] Specifically, S6 is executed by the release rollback module.

[0087] S61, Hard constraint check.

[0088] The release rollback module checks whether candidate versions satisfy all configured hard constraints based on the configured list of hard constraints. Hard constraints include one or more of the following: (1) The failure rate of critical tasks shall not exceed the preset threshold: In replay verification and shadow run verification, the failure rate of tasks marked as "critical" shall not exceed the preset threshold (e.g., 5%). (2) High-risk side effects must not be increased: The frequency or scope of side effects marked as "high-risk" in the side effect chart must not be increased compared to the stable version; (3) Permission and data boundary verification passed: All tool calls and data accesses involved in the behavior paths in the candidate version are within the scope of the permissions granted to the current agent; (4) No dangling nodes or circular conflicts in the four graphs: The four graphs in the candidate version are confirmed to have no dangling references or circular dependencies after structural consistency verification; (5) Key scene playback passed: All historical scenes marked as "key" achieved acceptable results in the playback verification; (6) Paths marked as not automatically modifiable have not been deleted: The system allows manual marking of specific paths (such as compliance approval paths) as locked. Candidate versions must not delete or bypass locked paths.

[0089] S62, release a new version or roll back.

[0090] When the verification results meet the preset hard constraints and comprehensive verification score When the release threshold is exceeded, the release rollback module atomically switches the candidate world behavior model version to the new stable version. This atomic switching ensures that there are no intermediate states during the switch, such as partial updates to four graphs, inconsistent cross-graph mappings, or version reference splits. After the release is complete, the system records the identifier of the current stable version in the version object and saves a rollback pointer (pointing to the parent stable version identifier) ​​for quick recovery in case of subsequent runtime anomalies.

[0091] When the verification results fail to meet any hard constraints or the overall verification score does not reach the release threshold, the release rollback module rejects the candidate version and performs a rollback. The rollback operation restores the world behavior model to the parent stable version, discarding all modifications in the candidate version. The system records the verification report and the reason for failure for subsequent analysis and adjustment of the evolution strategy.

[0092] S7. Update the control parameters based on the actual running results of the new stable version.

[0093] This embodiment also includes a continuous learning step after a new stable version is released, which is executed by the continuous learning module.

[0094] S71. Generate reward values ​​based on the actual performance of the new stable version. After a candidate version is released as a stable version, the continuous learning module continuously monitors its performance in a real production environment. Monitoring metrics include task success rate, side effect incidence and severity, positive and negative user feedback ratio, frequency of manual intervention, and resource consumption level. The continuous learning module calculates a comprehensive reward value based on these metrics. Positive rewards are given when candidate versions improve success rates, reduce the risk of side effects, and decrease manual intervention; negative rewards are given when rollbacks, major side effects, or worsening feedback occur.

[0095] S72. Update control parameters based on reward value. The continuous learning module updates the control parameter vector according to the reward value. Updated: ; in, The learning rate (a non-negative real number, typically ranging from 0.01 to 0.1). The objective function is constructed based on historical evolution samples (used to evaluate the quality of evolutionary decisions under the current combination of control parameters). This indicates that the parameters are projected to a preset safe range to prevent parameter overflow or loss of control. In this embodiment, the objective function... One specific construction method is: take the nearest The average of the real rewards after the release corresponding to each evolutionary decision, i.e. ,in In the first Control parameters are used on each historical evolution sample. The actual running reward value obtained afterwards. This is used to estimate the gradient. The continuous learning module can employ the central finite difference method: for the parameter vector Each component Apply small positive and negative perturbations respectively ,calculate This allows for the acquisition of an approximate gradient vector. In practical deployments, the Synchronous Perturbation Stochastic Approximation (SPSA) method can also be used to estimate the overall gradient through one or two objective function evaluations, thereby reducing computational overhead. The updated control parameters include: weighting coefficients in the behavioral bias calculation ( Weighting coefficients in attribution analysis to ), the threshold value in the evolution triggering conditions ( Cooldown time), learning rate for edge weight update ( ) and verification strategies (replay scene weights and gating thresholds).

[0096] The continuous learning module employs an update strategy with security constraints. For control parameters that directly affect security boundaries, user privacy, or compliance (such as risk penalty coefficients), these constraints are applied. and deviation strength threshold The system has strict upper and lower limits, and updates to the continuous learning module must not cause parameters to exceed these safe ranges. For high-risk changes where the safety boundaries cannot be automatically determined, the system maintains a manual review threshold, and the continuous learning module is not allowed to bypass hard constraints on permissions, security, and compliance.

[0097] If the candidate version fails verification and a rollback is performed, the system can choose to maintain the current control parameters or perform a conservative parameter update based on the running data of the stable version after the rollback. The conservative update uses a smaller learning rate and only adjusts the control parameters directly related to the verification failure.

[0098] This embodiment also provides a self-evolving control system based on a four-graph linkage world behavior model, including: The event acquisition module is used to collect the running events of the intelligent agent and perform standardized processing. The four-graph mapping module is used to map standardized runtime events to state graphs, behavior graphs, side effect graphs, and feedback graphs, establish cross-graph relationships, and generate incremental events in the four graphs. The deviation identification module is used to receive incremental events from four graphs, compare the current running path with the expected path of the stable version, calculate the behavior deviation score, and output the deviation object. The feedback attribution module is used to receive feedback events and node and edge information of the four graphs, calculate attribution scores for feedback events in the feedback graph and candidate nodes and edges in the state graph, behavior graph, and side effect graph, and output the attribution results. The change plan generation module is used to obtain deviation objects and attribution results, and generate multi-graph change plans when the evolution triggering conditions are met. The multi-graph reconstruction module is used to perform structured modification operations atomically based on the multi-graph change plan, perform cross-graph dependency verification and reconstruction, and generate candidate world behavior model versions. The version verification module is used to perform structural consistency verification, replay verification based on historical running data, and shadow running verification without affecting the real environment on the candidate world behavior model version, and generate verification results. The rollback module is used to obtain the verification results. When the verification is successful, a new version is released and the rollback pointer is retained. When the verification fails, the rollback is executed to a stable version. The continuous learning module is used to update control parameters based on the actual running results of the new version or the running data of the stable version after rollback.

[0099] The following is a complete running example of a specific scenario to help understand the overall process of the method described in this embodiment.

[0100] This embodiment applies to a project management agent responsible for managing task scheduling, resource allocation, and milestone progress in a software project. The agent's world behavior model includes the following four figures: The state diagram contains task state nodes such as "Requirements under review", "Under development", "Under testing", and "Closed", as well as transition constraint edges between states; The behavior graph contains behavior nodes such as "Start Development", "Execute Test", "Close Task", and "Request Extension", as well as the triggering order edges of each behavior in different states; The side effect graph contains side effect nodes such as "rework", "delay", and "resource overrun", as well as propagation edges of the side effects generated by the behavior nodes; The feedback graph includes feedback nodes such as user satisfaction ratings, project manager manual correction records, and project delay alerts.

[0101] Step 1: Event Collection and Four-Graph Mapping. During long-term operation, the project management agent repeatedly chose the "Close Task" action under the state combination of "Dependency Task Not Confirmed, Project Phase Approaching Deadline," subsequently resulting in "Rework" and "Delay" side effects. The project manager repeatedly recorded manual corrections in the feedback graph, indicating that "Dependency checks should be added before closing." The event collection module collected these state changes, action selections, side effect occurrences, and manual feedback events. The four-graph mapping module mapped these events to the "Dependency Task Not Confirmed" state node in the state graph, the "Close Task" action node in the action graph, the "Rework" and "Delay" side effect nodes in the side effect graph, and the "Manual Correction" feedback node in the feedback graph, establishing cross-graph associations: state → action precondition association, action → side effect occurrence association, and feedback → action attribution association.

[0102] The above-mentioned event collection and four-graph mapping process corresponds to S1. The collected "state combination + behavior + side effect + feedback" constitute a complete event closed loop.

[0103] Step 2: Behavioral Deviation Identification and Feedback Attribution. At the end of the task window, the deviation identification module compares the agent's actual path in the same state with the expected path in the stable version. In the stable version, the expected behavior in the "dependency task not confirmed" state is "check dependencies before making a decision," while the actual behavior is "directly close the task." The deviation identification module then uses the formula to calculate... ,in (Behavioral sequence differences) and The (feedback difference) item scored highly. Exceeding the set deviation strength threshold The feedback attribution module calculates attribution scores for the "manual correction" event in the feedback graph and the "task closure" node in the behavior graph. The temporal proximity, dependency, and credibility of the evidence are all high. The "Shut Down Task" node was added to the evolutionary candidate set because it exceeded the attribution threshold and the number of supporting samples (5 independent manual corrections) exceeded the minimum support threshold.

[0104] The above-mentioned behavioral deviation identification and feedback attribution process corresponds to S2.

[0105] Step 3: Evolution Triggering and Change Plan Generation. The change plan generation module checks the evolution triggering conditions: Established, Established (5 corrections or 3 thresholds). Established (project manager's manual correction credibility configured to high), cooldown period has passed. All four conditions are met, triggering evolution. The change plan generation module generates a multi-graph change plan, which includes: adding a "Dependency check completed" state node to the state graph; inserting a "Execute dependency check" behavior node to the behavior graph, configuring it as a prerequisite behavior in the "Dependency task not confirmed" state, and adjusting the trigger condition for "Close task" to "Dependency check completed and project phase not due date"; increasing the risk edge weights of "Rework" and "Delay" side effects in the side effect graph; and adding a "Dependency check confirmed" feedback collection point to the feedback graph.

[0106] The above-mentioned evolution triggering and change plan generation process corresponds to S3.

[0107] Step 4: Atomic Reconstruction and Candidate Version Generation. The multi-graph reconstruction module executes the multi-graph change plan according to the dependency order: First, a "Dependency Check Completed" node is added to the state graph. Then, a "Execute Dependency Check" node is added to the behavior graph, and an edge is established from the previous state "Dependency Task Unconfirmed" to this behavior node. Next, the precondition edges of the "Close Task" node are adjusted, and the risk weight of the "Close Task → Rework" edge in the side effect graph is updated synchronously. Finally, a collection point is added to the feedback graph. All modifications are completed in an atomic transaction, generating a candidate world behavior model version. Cross-graph dependency validation passed; no dangling references were found.

[0108] The above atomic reconstruction process corresponds to S4.

[0109] Step 5: Candidate Version Verification. The version verification module verifies the candidate versions. Three verifications were performed. **Structural Consistency Verification:** The integrity of node and edge references in the four graphs was checked; no dangling references were found, and the verification passed. **Historical Replay Verification:** Decision-making was re-simulated using historical project scenarios from the past 6 months (including 10 successful scenarios, 5 failed scenarios, and 3 boundary scenarios). The candidate version performed better than or equal to the stable version in 8 successful scenarios and avoided failure in all 5 failed scenarios. The critical scenario replay pass rate reached 100% (exceeding the preset 90% threshold), and the replay verification passed. **Shadow Run Verification:** The candidate version was deployed to a shadow environment and ran for 72 hours following production requests. No abnormal stop conditions were triggered, and the critical task failure rate decreased from 8% of the original version to 3%, with a reduction in side effect risk of approximately 40%. The shadow run verification passed. **Comprehensive Verification Score:** It exceeds the publication threshold.

[0110] The above verification process corresponds to S5.

[0111] Step 6: Deployment and Rollback Decision. The deployment and rollback module checks hard constraints: the critical task failure rate (3%) is below the preset threshold of 5%; high-risk side effects have not increased; permission boundary checks have passed; all four diagrams are referenced completely; critical scenario replays have passed; and the "compliance approval" locked path has not been deleted. If all hard constraints are met, the deployment and rollback module will proceed. The atom is switched to the new stable version, while retaining the rollback pointer pointing to the old version. .

[0112] The above release process corresponds to S6.

[0113] Step 7: Continuous Learning. After running for one month following the release of the candidate version, the project management agent's project delay rate decreased by approximately 25%, the number of manual interventions decreased by approximately 40%, and project manager satisfaction scores increased. The continuous learning module calculates positive reward values ​​based on these positive results and updates control parameters: increasing the weight of behavioral biases. (Give more importance to compliance of behavioral sequences), and appropriately lower the attribution threshold. (To make the system more sensitive to similar corrections), fine-tune the learning rate of the edge weights. This is to accelerate the accumulation of positive experience. If unexpected issues arise in subsequent versions, the system can quickly revert to the previous version using rollback pointers. .

[0114] The above continuous learning process corresponds to S7.

[0115] Thus, the world behavior model of the project management agent has completed a full feedback graph-driven self-evolution closed loop: from event collection, four-graph mapping, deviation identification and attribution, to evolution triggering, change plan generation, atomic reconstruction, candidate version verification, and then to secure release and continuous learning, the behavior model has been continuously optimized throughout the process while maintaining version traceability and security control.

[0116] Therefore, this invention employs the aforementioned self-evolutionary control method and system for a world behavior model based on four-graph linkage. It transforms operational feedback into structured control signals, achieving atomic reconstruction across the four graphs and ensuring cross-graph dependency consistency and causal chain integrity. Through candidate version replay, shadow runtime verification, and hard constraint gating, it significantly reduces evolutionary risks and improves long-term operational stability. After release, it automatically updates deviation weights, attribution parameters, and trigger thresholds based on actual performance, enabling continuous learning while retaining hard constraints such as permissions and security, without bypassing manual review. This comprehensively improves the agent's behavior correction capabilities, version stability, and result traceability, ensuring that the self-evolutionary process is verifiable, rollbackable, and governable.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A self-evolutionary control method based on a four-graph linkage world behavior model, characterized in that, Includes the following steps: S1. Collect the running events of the intelligent agent, standardize the running events, and map them to the world behavior model composed of state graph, behavior graph, side effect graph and feedback graph, generate four-graph incremental events and establish cross-graph associations; S2. Based on the four graph incremental events, calculate the behavioral deviation of the current running path relative to the expected path of the stable version; Based on the feedback events in the feedback graph, candidate nodes and edges associated with the feedback events are identified in the state graph, behavior graph, and side effect graph. Attribution analysis is then performed on the candidate nodes and edges to generate attribution results. S3. When the behavioral deviation and attribution results meet the preset evolution trigger conditions, a multi-graph change plan is generated. The multi-graph change plan includes structured modification operations on the nodes and edges of at least one graph in the state graph, behavior graph, side effect graph, and feedback graph, as well as the execution dependency order of the modification operations. S4. Based on the multi-graph change plan, perform structured modification operations in an atomic manner to perform cross-graph dependency verification and reconstruction on the state graph, behavior graph, side effect graph, and feedback graph, and generate candidate world behavior model versions; S5. Perform structural consistency verification, replay verification based on historical running data, and shadow running verification without affecting the real environment on the candidate world behavior model version, and generate verification results. S6. When the verification results meet the preset hard constraints and release conditions, release the candidate world behavior model version as a new stable version and retain the rollback pointer; when the verification results do not meet the hard constraints or release conditions, roll back to the stable version.

2. The method according to claim 1, characterized in that, The behavioral deviations described in S2 are calculated using the following formula: ; in, For path Behavioral deviation score; The deviation between the actual state and the expected state; The difference between the actual behavior sequence and the reference path; Considering the risks of side effects and the intensity of transmission; To provide feedback on the discrepancy between the results and the objectives; This is due to performance drift compared to relatively stable historical versions; , , , , These are the weighting coefficients.

3. The method according to claim 1, characterized in that, The attribution analysis described in S2 is performed as follows: For feedback nodes and candidate nodes Calculate the attribution score according to the following formula. : ; in, For time proximity, For dependency degree, For counterfactual discrepancy, For the credibility of the evidence, For the correlation of side effect transmission pathways, to These are the weighting coefficients; When the attribution score exceeds the attribution threshold and the number of supporting samples exceeds the minimum support threshold, the candidate node will be... Add to the evolutionary candidate set.

4. The method according to claim 1, characterized in that, The evolution triggering conditions described in S3 are determined by the following combination of conditions: ; in, Scoring for behavioral deviations This is the threshold for deviation strength. This represents the current cumulative number of supporting samples. The minimum support sample threshold; For the credibility of the evidence, As a threshold for the credibility of evidence; This indicates that the cooldown period has ended. The cooldown period started counting from the last evolution trigger or version release. For logical AND operator.

5. The method according to claim 1, characterized in that, The cross-graph dependency verification and reconstruction described in S4 includes edge weight updates, which are performed according to the following formula: ; in, For the first wheel edge The weights; The attenuation coefficient; The learning rate is positively rewarded. To provide feedback on credibility; To provide positive incentives; This is the deviation penalty coefficient; This is a penalty item for deviations; This is the risk penalty coefficient; Side effects risk item; This is a truncation function.

6. The method according to claim 1, characterized in that, Cross-graph dependency verification and refactoring as described in S4 also includes node addition, with the addition score calculated according to the following formula: ; in, For the new model The new score; for With existing nodes The maximum similarity; To support the sample size; For the expected impact; , , These are the weighting coefficients.

7. The method according to claim 1, characterized in that, The hard constraints described in S6 include one or more of the following, and releasing a new version requires satisfying all configured hard constraints: The failure rate of critical tasks does not exceed a preset threshold; High-risk side effects must not be increased; Permissions and data boundary checks passed; There are no dangling nodes or loop conflicts in the four-graph reference; Key scene playback passed; Paths marked as not automatically modifiable were not deleted.

8. The method according to claim 1, characterized in that, The release of a new stable version in S6 also includes: Reward values ​​are generated based on the actual running results of the new stable version; Based on the reward value, the weight coefficients in the calculation of behavioral bias, the weight coefficients in the attribution analysis, the threshold values ​​in the evolution triggering conditions, and the learning rate or verification strategy for edge weight updates are automatically updated. If the candidate version verification fails and a rollback is performed, the current control parameters remain unchanged or the control parameters are updated based on the running data of the stable version after the rollback.

9. A self-evolving control system based on a four-graph linkage world behavior model, characterized in that, include: The event acquisition module is used to collect the running events of the intelligent agent and perform standardized processing. The four-graph mapping module is used to map standardized runtime events to state graphs, behavior graphs, side effect graphs, and feedback graphs, establish cross-graph relationships, and generate incremental events in the four graphs. The deviation identification module is used to receive incremental events from four graphs, compare the current running path with the expected path of the stable version, calculate the behavior deviation score, and output the deviation object. The feedback attribution module is used to receive feedback events and node and edge information of the four graphs, calculate attribution scores for feedback events in the feedback graph and candidate nodes and edges in the state graph, behavior graph, and side effect graph, and output the attribution results. The change plan generation module is used to obtain deviation objects and attribution results, and generate multi-graph change plans when the evolution triggering conditions are met. The multi-graph reconstruction module is used to perform structured modification operations atomically based on the multi-graph change plan, perform cross-graph dependency verification and reconstruction, and generate candidate world behavior model versions. The version verification module is used to perform structural consistency verification, replay verification based on historical running data, and shadow running verification without affecting the real environment on the candidate world behavior model version, and generate verification results. The rollback module is used to obtain the verification results. When the verification is successful, a new version is released and the rollback pointer is retained. When the verification fails, the rollback is executed to a stable version. The continuous learning module is used to update control parameters based on the actual running results of the new version or the running data of the stable version after rollback.