Emergency decision bidirectional deduction method, device and equipment
By constructing multidimensional feature vectors and causal probing models, and combining virtual intervention testing in a digital twin mirror environment, the problem of inaccurate fault root cause localization in existing technologies has been solved, achieving highly reliable emergency decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE ONLINE SERVICES CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies rely on static topological association or statistical correlation models for fault root cause localization, resulting in a single modeling dimension, difficulty in accurately quantifying the impact between the physical state of the equipment and the upper-level business logic, high false location rate, and lack of physical interpretability and verification reliability.
Collect multi-source data from the target system, construct a unified time series multi-dimensional feature vector, use a causal probing model to analyze anomaly propagation, verify causal relationships through virtual intervention tests in a digital twin mirror environment, and generate fault classification causal paths with probability weights.
It significantly improves the accuracy and verifiability of root cause localization, provides a reliable basis for emergency decision-making, reduces the false location rate, and enhances the interpretability and adaptability of emergency strategies.
Smart Images

Figure CN121980531A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication technology, and in particular to a two-way simulation method, apparatus and device for emergency decision-making. Background Technology
[0002] As enterprises deepen their digital transformation, the complexity, scale, and business dependence of IT systems are increasing, making their stable and continuous operation a key component of core competitiveness. IT operations and maintenance, especially emergency response, as the core line of defense for ensuring business continuity, system stability, and data security, has evolved from basic technical assurance to strategic-level business support. In dealing with sudden failures, efficient emergency response capabilities directly determine the duration of business interruption and the scale of economic losses, driving the accelerated evolution of operations and maintenance technologies towards intelligence, automation, and high reliability.
[0003] Currently, mainstream IT operations and maintenance emergency response technologies are showing a trend of deep integration of intelligence and automation, specifically manifested in two methods: first, automated response and log analysis, which uses tools to centrally collect and correlate logs, automatically triggering fault recovery processes according to preset rules, and using call chain tracing technology to locate abnormal nodes in the microservice architecture; second, AI-driven prediction and diagnosis, which uses machine learning models trained on historical operations and maintenance data to predict hardware failures or performance bottlenecks and automatically generate repair suggestions. However, when facing increasingly complex cross-layer and dynamic fault scenarios, existing technical solutions still have significant shortcomings. Specifically, in the root cause localization stage, existing methods usually rely on static topology correlation or statistical correlation models for analysis. Their modeling dimensions are singular, making it difficult to accurately quantify the impact between the physical state of the equipment and the upper-layer business logic. This results in the root causes being located as suspicious nodes under statistical correlation, containing a large amount of spurious correlation interference, lacking physical interpretability and verification reliability, and having a high false location rate. Summary of the Invention
[0004] This application provides a two-way inference method, apparatus, and equipment for emergency decision-making, which addresses the problem in related technologies that rely on static topological correlation or statistical correlation models for fault root cause localization. These models have a single dimension, making it difficult to accurately quantify the impact between the physical state of the equipment and the upper-level business logic. As a result, the root causes located are mostly suspicious nodes under statistical correlation, containing a large number of spurious correlation interferences, lacking physical interpretability and verification reliability, and having a high false location rate.
[0005] This application provides a two-way simulation method for emergency decision-making, including: Collect device monitoring data, device attribute data, business dependency data, and business traffic data of the target system, and align the collected multi-source data with time axis to construct a multi-dimensional feature vector with a unified time series. The multidimensional feature vectors are input into the causal probing model, and anomaly propagation analysis is performed on the positive and negative influences between the business-related vectors and the equipment-related vectors in the multidimensional feature vectors based on the dual-track spatiotemporal topology sub-network in the causal probing model, so as to obtain an anomaly score list. Furthermore, based on the inverse scenario deduction core subnetwork in the causal probing model, in the digital twin mirror environment of the target system, virtual intervention tests are designed for the candidate root cause nodes determined by the anomaly scoring list, and based on the dynamic impact of the virtual intervention tests on the key indicators of the target system, a fault-level causal path with probability weights is generated.
[0006] This application also provides an emergency decision-making two-way simulation device, including: The feature vector construction module is used to collect device monitoring data, device attribute data, business dependency data and business traffic data of the target system, and to align the collected multi-source data with time axis to construct a multi-dimensional feature vector with a unified time series. The causal probing model is used to perform anomaly propagation analysis on the positive and negative influences between business-related vectors and equipment-related vectors in the multidimensional feature vectors based on a dual-track spatiotemporal topology subnetwork, to obtain an anomaly score list; and, based on the inverse scenario deduction core subnetwork, in the digital twin mirror environment of the target system, to design virtual intervention tests on the candidate root cause nodes determined by the anomaly score list, and to generate a fault-level causal path with probability weights based on the dynamic impact of the virtual intervention tests on the key indicators of the target system.
[0007] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. The processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform steps in the emergency decision-making bidirectional simulation method provided in this application.
[0008] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps in the emergency decision-making bidirectional deduction method provided in this application.
[0009] This application also provides a computer program product that stores instructions that, when executed by a computer, cause the computer to perform the steps in the emergency decision-making two-way deduction method provided in this application.
[0010] The emergency decision-making bidirectional extrapolation method provided in this application collects and time-aligns multi-source data such as device monitoring, attributes, business dependencies, and traffic to construct a unified temporal-series multi-dimensional feature vector, providing a high-quality data foundation with spatiotemporal consistency for subsequent analysis. This vector is then input into a causal probing model, which first uses its dual-track spatiotemporal topology subnetwork to perform anomaly propagation analysis on the positive and negative impacts between device and business vectors, identifying key anomaly nodes and initially achieving quantitative modeling of cross-layer bidirectional impacts. Based on this, the model's inverse hypothetical extrapolation core subnetwork is used to design and execute virtual intervention tests on candidate root cause nodes in the digital twin mirror environment of the target system. This verifies and quantifies the true causal relationship between nodes based on the actual dynamic impact of the intervention on key system indicators. This combination of techniques ultimately generates a fault-level causal path with probability weights, significantly improving the accuracy, verifiability, and interpretability of root cause localization in complex fault scenarios, laying a reliable foundation for subsequent accurate emergency decision-making. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an emergency decision-making two-way simulation method provided for an exemplary embodiment of this application; Figure 2 A schematic diagram of the system architecture involved in the two-way simulation method for emergency decision-making provided as an exemplary embodiment of this application; Figure 3 A schematic diagram of the causal touch model in the emergency decision-making two-way extrapolation method provided as an exemplary embodiment of this application; Figure 4 A schematic diagram of an implementation process of an emergency decision-making two-way simulation method provided as an exemplary embodiment of this application; Figure 5 A schematic diagram of the structure of an emergency decision-making two-way simulation device provided as an exemplary embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The following is a description of the terms used in this application: A Configuration Management Database (CMDB) is a core database used to centrally store and manage all configuration items and their interrelationships in an IT environment. It records detailed attributes, version status, and logical and physical dependencies of hardware and software assets such as servers, network devices, and applications. It provides an accurate and unified data source for IT service management, change management, and impact analysis, serving as the foundational information base for modern IT operations automation and digital twin construction.
[0014] Application Performance Management (APM) is a set of technologies and practices that monitor, diagnose, and manage the performance and availability of software applications to ensure they meet expected service level goals. It uses distributed tracing, code-level diagnostics, and transaction analysis to collect real-time metrics on call chains, response times, and error rates to pinpoint performance bottlenecks, analyze root causes, and optimize user experience and business continuity.
[0015] A Gated Recurrent Unit (GRU) is a variant of a recurrent neural network designed for efficient processing of sequential data. By introducing update and reset gates, it selectively forgets or passes on historical information, solving the vanishing or exploding gradient problems common in traditional recurrent neural networks. In time series modeling, GRUs efficiently capture long-term dependencies and dynamic patterns in data, and are frequently used in scenarios such as equipment performance sequence analysis and fault prediction.
[0016] Digital twin is a technology that uses historical and real-time data of a physical entity or system to construct a high-fidelity, dynamically evolving digital mirror model in virtual space. This model achieves synchronous mapping, simulation analysis, and predictive optimization with the physical entity through data-driven processes. It supports status monitoring, fault simulation, performance extrapolation, and strategy verification in a zero-risk virtual environment, making it a key technology for achieving intelligent operation and maintenance and scientific decision-making.
[0017] Graph Convolutional Networks (GCNs) are deep learning models specifically designed for processing graph-structured data. Borrowing from convolutional neural networks, they learn effective representations of nodes in a graph by aggregating feature information from the node itself and its neighbors, enabling them to capture complex topological dependencies. In business dependency analysis, GCNs can be used to model microservice call chains and identify anomaly propagation patterns and key impact nodes between services.
[0018] The Dual-track Spatiotemporal Topology Network is a dedicated graph neural network architecture designed in this proposal for modeling bidirectional causal effects across layers. This network includes two information propagation paths: a forward spatiotemporal flow and a reverse causal flow. The forward flow models the causal impact of device physical states on business indicators, while the reverse flow captures the adverse effects of business anomalies on underlying devices. Bidirectional anomaly signal propagation and impact intensity calculation between the device layer and the business layer are achieved through cross-layer bridging edges and attention mechanisms.
[0019] The Inverse-State Deduction Core is a key component of the causal probing model in this application, used to perform causal verification in a digital twin environment. Its core idea is to create a mirror copy of the system state and actively apply virtual intervention operations such as positive repair and negative deterioration to candidate root cause nodes. By observing and quantifying the dynamic changes in the distribution of key system indicators before and after the intervention, it distinguishes between strongly causal nodes and pseudo-correlated nodes, generating a reliable causal path with probability weights.
[0020] The Virtual-Real Fusion and Exercise Environment is an isolated, high-fidelity simulation testing environment that uses containerization and other technologies to proportionally reconstruct the production environment topology and inject cloned traffic and simulated faults. Its goal is to support the parallel simulation and effectiveness verification of multiple emergency strategies without affecting the real production system, achieving zero-risk pre-simulation and evaluation of strategies, and overcoming the inherent contradiction between distortion and production interference in traditional exercise environments.
[0021] Hierarchical Reinforcement Learning (HRL) is a reinforcement learning method that solves complex decision-making problems by introducing a hierarchical structure. It decomposes the overall task into subtasks at different time scales or abstract levels, typically including an upper-level meta-policy (responsible for goal decomposition and planning) and lower-level execution policies (responsible for specific action execution). In this proposal, the HRL framework is used to achieve intelligent decomposition of emergency objectives and hierarchical generation and optimization of policies.
[0022] Ecological Cluster Evolutionary Game is a policy optimization mechanism adopted by the L2 agent in this application embodiment, inspired by ecology and evolutionary game theory. It divides the candidate policy population into ecological clusters with different risk preferences, such as conservative, aggressive, and mixed. Through genetic operations such as crossover, mutation, and forced reset, policies evolve within and between clusters. The proportion of each cluster's policy in the population is dynamically adjusted based on its average fitness, thereby driving continuous adaptive optimization of the policy library and avoiding getting trapped in local optima.
[0023] As described in the background section, the fault diagnosis stage of emergency decision-making in related technologies typically relies on static topological association or statistical correlation models. These methods have a single modeling dimension and cannot quantify the bidirectional dynamic causal influence between the physical state of the equipment and the upper-level business logic. This results in the root causes being mostly suspicious nodes under statistical association, containing a large number of spurious correlation interferences. Furthermore, due to the lack of feasible means to actively verify candidate root causes in real-world environments, the diagnostic results lack physical interpretability and verification reliability, leading to a high false-location rate. In the strategy verification and generation stage, existing methods rely on pre-built static rule bases or experience templates. The strategy base updates are lagging, and there is a lack of a decision framework that can integrate high-fidelity environment simulation and intelligent optimization mechanisms. This results in rigid emergency strategies that are difficult to adapt to complex and ever-changing fault scenarios, and insufficient strategy generation efficiency and adaptability.
[0024] To address the systemic technical problems of unverifiable root cause localization and lack of adaptive optimization in strategy generation in related technologies, this application provides a bidirectional inference method based on emergency decision-making. This method first collects multi-source operation and maintenance data of the target system and aligns it along the time axis to construct a multi-dimensional feature vector with a unified time series. Next, the multi-dimensional feature vector is input into a causal probing model. Using the dual-track spatiotemporal topology sub-network in this model, anomaly propagation analysis is performed on the positive and negative influences between device-related vectors and business-related vectors to obtain an anomaly score list. Then, using the inverse hypothetical inference core sub-network in the model, virtual intervention tests, including positive repair, reverse deterioration, and combined interventions, are designed and executed on candidate root cause nodes identified by the anomaly score list in the digital twin mirror environment of the target system. Based on the dynamic impact of the virtual intervention tests on key system indicators, a fault-level causal path with probability weights is generated. This step, by executing proactive intervention tests in an isolated digital twin environment, empirically verifies the causal relationship between nodes, effectively distinguishes strong causal nodes from spurious interference, and significantly improves the accuracy and verifiability of root cause localization.
[0025] Based on the generated high-confidence fault classification causal path, the method provided in this application further inputs it into a virtual-real fusion simulation field. In this field, based on the path, the production environment topology of the target system is reconstructed proportionally using containerization technology. Cloned real business traffic and corresponding simulated faults are injected, and multiple preset emergency response strategy templates, including empirical, conservative, and aggressive strategies, are loaded for parallel simulation. Multiple sets of simulation schemes are output, including three-dimensional evaluation indicators such as business recovery rate, resource consumption ratio, and operational complexity. This achieves pre-validation and quantitative evaluation of the emergency response strategy in a zero-risk, high-fidelity simulation environment.
[0026] Finally, the fault-level causal path and multiple sets of inference schemes are input into a hierarchical reinforcement learning policy generator. This generator uses an L1 agent to decompose the causal path into a priority sequence of emergency sub-objectives. An L2 agent, based on a base policy pool fused from historical policy repositories and inference scheme fragments, uses the three-dimensional fitness defined by the three-dimensional evaluation index as the optimization objective and generates target emergency strategies through an ecological cluster evolutionary game mechanism. This ecological cluster evolutionary game mechanism classifies candidate strategies into three ecological clusters: conservative, aggressive, and hybrid. It drives evolutionary operations such as crossover, mutation, and forced reset based on the three-dimensional fitness, dynamically adjusting the proportion of each ecological cluster according to its average fitness, and ultimately outputs the strategy with the highest fitness.
[0027] As can be seen, the emergency decision-making bidirectional inference method provided in this application systematically solves the problems of inaccurate root cause localization due to the lack of causal verification mechanism and rigid strategy generation due to the lack of high-fidelity inference and adaptive optimization mechanism in the prior art by constructing a complete technical closed loop that connects causal probing, virtual and real inference and intelligent optimization. It realizes the full-process enhancement from accurate fault diagnosis to intelligent strategy generation, and significantly improves the reliability and timeliness of IT operation and maintenance emergency decision-making.
[0028] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0029] Figure 1 This is a flowchart illustrating a two-way simulation method for emergency decision-making, provided as an exemplary embodiment of this application. Figure 1 As shown, the method includes: Step 110: Collect device monitoring data, device attribute data, business dependency data and business traffic data of the target system, and align the collected multi-source data with time axis to construct a multi-dimensional feature vector with a unified time series.
[0030] As can be seen, step 110 is the initial data preparation stage of the two-way simulation method for emergency decision-making. Its core task is to collect key data reflecting the system's operational status from different sources and levels of the target IT system, and to integrate this heterogeneous data into standardized inputs that can be used for subsequent intelligent model analysis through technical means. The specific implementation method of this step is as follows: First, data is collected through distributed probes deployed at different layers. These probes include device-layer probes, network-layer probes, and business-layer probes. Device-layer probes are deployed on physical servers or virtual machines to collect real-time performance metrics such as server CPU utilization, memory usage, and disk I / O. Network-layer probes are deployed on mirrored ports of core switches to collect metrics such as cross-datacenter bandwidth utilization and network packet loss rate. Business-layer probes are integrated into the microservice framework to collect business logic metrics such as microservice call chain response time and transaction success rate. Simultaneously, static device attributes, such as CPU architecture, network bandwidth, and storage type, are extracted from the Configuration Management Database (CMDB); dynamic business dependencies, including call chains between microservices, API dependency graphs, and database transaction relationships, are obtained from the Application Performance Management (APM) system. These data collectively constitute device monitoring data, device attribute data, business dependency data, and business traffic data.
[0031] After data collection, the aforementioned multi-source heterogeneous data needs to be aligned in terms of time axis and spatial topology. Time axis alignment aims to establish a unified time benchmark, aligning configuration change events (such as server expansion) in the CMDB with fluctuations in application performance monitoring metrics at the millisecond level. The system employs a sliding window mechanism to ensure consistency between device status time-series data and business traffic time-series data in the time dimension. Spatial topology alignment maps entities from different data sources to a unified topology structure. For example, it associates service instances recorded in APM with service deployment servers recorded in the CMDB, thereby constructing a mapping relationship between devices and services.
[0032] Building upon this foundation, the system performs feature encoding on the aligned data to construct multi-dimensional feature vectors with a unified time series. For device physical features, the system performs vectorization encoding, for example, mapping CPU architecture to a 128-dimensional embedding vector. For business dependencies, the system constructs a weighted directed graph, where nodes represent service instances, edge weights are calculated from call latency and error rate, and generates a 256-dimensional business feature vector for each node. Finally, these timestamped feature vectors are stored in a time-series database, forming a time-series feature library. This feature library can be sharded and stored according to business domains, and a retrieval structure is established with device identifier and time range as the primary index and business feature tags as the secondary index.
[0033] Therefore, by unifying timeline alignment and spatial topology mapping, the temporal deviations and logical gaps from different monitoring tools and data sources are effectively eliminated, achieving deep integration and structured coding of device physical characteristics and business logic dependencies. The generated multidimensional feature vectors provide spatiotemporally consistent, dimensionally rich, and standardized input data for subsequent causal probing models, fundamentally ensuring that abnormal signals can be accurately and consistently transmitted and analyzed between the device layer and the business layer, laying a reliable data foundation for the entire bidirectional inference process.
[0034] Step 120: Input the multidimensional feature vector into the causal probing model, and perform anomaly propagation analysis on the positive and negative influences between the business-related vectors and the equipment-related vectors in the multidimensional feature vectors based on the dual-track spatiotemporal topology sub-network in the causal probing model, to obtain an anomaly score list.
[0035] As can be seen, the core task of step 120 is to perform in-depth analysis on the multidimensional feature vectors generated in step 110 by constructing a dual-track spatiotemporal topology subnetwork within the causal probing model, so as to quantify the bidirectional anomaly impact between the device layer and the business layer and generate a unified anomaly score list.
[0036] In some exemplary embodiments, to ensure that anomalous signals are fully extracted from their respective domains starting from the raw data before cross-domain fusion, so that the final comprehensive score combines domain expertise with global relevance, making the analysis process more rigorous and reliable, specifically, based on the dual-track spatiotemporal topology sub-network in the causal probing model, anomaly propagation analysis is performed on the positive and negative influences between business-related vectors and device-related vectors in the multidimensional feature vectors, resulting in an anomaly score list, including: Based on the physical layer of the dual-track spatiotemporal topology subnetwork, the device-related vectors are time-series modeled and anomaly aggregated to output a sequence of devices with abnormal signals. Based on the business layer of the dual-track spatiotemporal topology sub-network, graph convolution propagation and anomaly aggregation are performed on business-related vectors to output anomaly scoring sequences of instance nodes. Based on the coupling layer of the dual-track spatiotemporal topology subnetwork, bidirectional propagation and fusion are performed on the abnormal device sequence and the abnormal score sequence of instance node to generate an abnormal score list.
[0037] The dual-track spatiotemporal topology subnetwork is a specially designed heterogeneous graph neural network. Its architecture is clearly divided into a physical layer, a business layer, and a coupling layer to process heterogeneous data containing device-related vectors and business-related vectors. The heterogeneous layered graph structure is the core representation of the data processed by this network. It includes a physical layer subgraph, a business layer subgraph, and cross-layer bridging edges connecting the two layers. The physical layer subgraph uses hardware devices as nodes, and the physical connections between devices as edges. The edge weights can be set to the actual bandwidth utilization. The business layer subgraph uses microservice instances as nodes, and the call relationships between services as edges. The edge weights can be calculated from call latency and error rate. The cross-layer bridging edges establish the deployment relationships between devices and services, for example, by evenly distributing the weights of servers hosting multiple service instances.
[0038] When processing at the physical layer based on a dual-track spatiotemporal topology subnetwork, the network performs temporal modeling and anomaly aggregation on device-related vectors. Specifically, the physical layer processing unit encodes the static attributes of devices into 128-dimensional embedding vectors through its input encoding layer, and combines the device network connectivity relationships and weights to form the initial features of physical layer nodes. Simultaneously, device temporal monitoring indicators (such as CPU and memory sequences) contained in the device-related vectors are input to the GRU temporal modeling layer. The GRU layer captures the periodic and burst patterns of device indicators through its gating mechanism, outputting dynamic temporal features of the devices. Subsequently, the physical layer's output layer, based on the constructed physical layer topology graph, performs graph convolution aggregation on the temporal features output by the GRU layer, calculates the intra-layer anomalous signal strength of each device node, and finally outputs a device sequence identifying anomalous device nodes and their signal anomalies.
[0039] When processing at the business layer based on the dual-track spatiotemporal topology subnetwork, the network performs graph convolutional propagation and anomaly aggregation on business-related vectors. Specifically, the business layer processing unit constructs a business layer dependency graph from dynamic business dependencies (such as microservice call chains) through its input encoding layer, and encodes the business-related vectors into 256-dimensional business feature vectors as the initial features of the business layer nodes. This initial feature and the business layer dependency graph are input together into a multi-hop graph convolutional layer for processing. The graph convolutional layer performs multi-hop information propagation on the business dependency graph through a set convolution order (e.g., 3), covering service nodes within the three-hop neighborhood of the call chain, thereby extracting the anomaly context features of the local call chain. Finally, the output layer of the business layer calculates the intra-layer anomaly score for each service instance based on the result of graph convolutional aggregation, and outputs the anomaly score sequence of the instance nodes.
[0040] When processing in the coupled layer based on the dual-track spatiotemporal topology subnetwork, the network performs bidirectional propagation and fusion of the aforementioned two sequences, which is crucial for achieving bidirectional impact analysis. The coupled layer receives the device sequence of signal anomalies output from the physical layer and the anomaly score sequence of instance nodes output from the service layer as input. Its core operation is accomplished through a cross-attention mechanism. This mechanism calculates the forward propagation weights from physical layer nodes to service layer nodes and the reverse propagation weights from service layer nodes to physical layer nodes, respectively. This realizes a cross-layer propagation mechanism, where the forward propagation path propagates the anomaly signal based on the temporal characteristics of the device's physical state along the direction from the device layer node to the service layer node to model its causal impact on service indicators; the reverse propagation path propagates the anomaly signal based on the contextual characteristics of the service logic in the opposite direction to capture the reverse effect of service anomalies on the underlying devices. Based on the dynamically calculated weights, the anomaly signal and anomaly score are transmitted to the relevant cross-layer nodes along the forward and reverse paths, respectively. Subsequently, for the business layer subgraph, after performing graph convolution operations, the coupling layer performs cross-attention calculations on the feature vectors of the current business layer nodes from the forward propagation path and the corresponding feature vectors from the backward propagation path. This calculation aims to identify potential patterns of dual-track dependencies, and the results are used to update the comprehensive representation of the business layer nodes. Finally, the output layer of the coupling layer fuses the cross-layer anomaly information received by each node with its own intra-layer anomaly information to generate a unified anomaly score list that integrates all nodes in the physical layer and the business layer and reflects the bidirectional cross-layer influence.
[0041] Through an architecture that independently analyzes the physical and business layers and then integrates them bidirectionally via a coupling layer, this system achieves, for the first time in the operations and maintenance field, accurate modeling and quantification of the bidirectional dynamic causal impact between equipment physical status and business logic indicators. It breaks through the limitations of traditional single-layer or unidirectional analysis models, not only capturing the impact of equipment anomalies on business operations but also tracing the reverse effects of business pressure on underlying resources. This significantly reduces root cause misjudgments and spurious correlations caused by incomplete modeling, providing a high-quality, highly interpretable set of candidate root causes for subsequent causal verification.
[0042] In some exemplary embodiments, based on the coupling layer of the dual-track spatiotemporal topology subnetwork, bidirectional propagation and fusion are performed on the device sequence with signal anomalies and the anomaly score sequence of instance nodes to generate an anomaly score list, including: By using the cross-attention mechanism in the coupling layer, the forward propagation weight from the physical layer to the business layer and the backward propagation weight from the business layer to the physical layer are calculated respectively. Based on forward propagation weights, abnormal signals in the sequence of devices with abnormal signals are transmitted to the relevant service layer nodes. Based on the backpropagation weight, the abnormal scores in the abnormal score sequence of the instance node are passed to the relevant physical layer nodes. By integrating cross-layer anomaly information received by each node with its own intra-layer anomaly information, a unified anomaly score list is generated.
[0043] The bidirectional propagation and fusion functionality of the coupling layer can be refined through a cross-attention mechanism. This mechanism first calculates weights: for each pair of device nodes and service nodes with a deployment relationship, it dynamically calculates a forward propagation weight (quantifying the impact of the device anomaly on the service) and a reverse propagation weight (quantifying the stress intensity of the service anomaly on the device). The weight calculation is based on the node's current representation, not fixed values. Then, weighted propagation is performed: based on the forward weights, the abnormal signals of each device node in the sequence of devices with signal anomalies are allocated and transmitted to their associated service nodes according to their weight ratios; simultaneously, based on the reverse weights, the abnormal scores of each service instance in the anomaly score sequence of instance nodes are transmitted in reverse to the physical device node where it is deployed. Finally, feature fusion is performed: after receiving weighted anomaly information from another layer, each node (whether device or service) fuses this cross-layer information with its own intra-layer anomaly information obtained from analysis in the physical or service layer, thereby updating its comprehensive anomaly representation and generating a unified anomaly score list.
[0044] By introducing a cross-attention mechanism to dynamically calculate bidirectional propagation weights, the model can adaptively adjust the importance of cross-layer influences based on the specific fault scenario and the real-time state of nodes, significantly improving the accuracy of characterizing complex and dynamic fault propagation paths. Simultaneously, by clearly distinguishing and integrating intra-layer and cross-layer anomaly information, the model ensures that the final node score reflects both its own health status and the components influenced by related layers, resulting in more comprehensive and accurate analysis results.
[0045] In some exemplary embodiments, the cross-layer propagation mechanism includes: The abnormal signal is propagated along a forward propagation path from the device layer node to the service layer node. Based on the temporal characteristics of the device's physical state, the abnormal signal is propagated to model the causal impact of the abnormal signal on service metrics; and... The abnormal signal is propagated in reverse along the direction from the business layer node to the device layer node. Based on the context characteristics of the business logic, the abnormal signal is propagated to capture the reverse effect of business anomalies on the underlying devices.
[0046] The cross-layer propagation mechanism comprises two opposing and functionally distinct propagation paths. The forward propagation path is defined as a directed information flow from device layer nodes to business layer nodes. This path uses the temporal characteristics of the device's physical state (such as waveform patterns of sudden CPU spikes) as its information carrier, and its design aims to model the causal impact of device resource anomalies on the upper-layer business metrics it supports (such as service response latency). The reverse propagation path is defined as a directed information flow from business layer nodes to device layer nodes. This path uses the contextual characteristics of business logic (such as error propagation patterns on the call chain) as its information carrier, and its design aims to capture the reverse pressure and resource contention caused by anomalies at the business logic level (such as a large number of request timeouts) on the underlying computing, storage, or network devices.
[0047] By formally defining two independent propagation paths—forward and reverse—and their respective information carriers and modeling objectives, the abstract concept of bidirectional influence is endowed with concrete and computable technical connotations. It forces the model to explicitly consider two different causal flows of different directions and properties during design and training, architecturally preventing the model from degenerating into unidirectional analysis, ensuring the built-in and verifiable bidirectional causal analysis capability, and providing a clear logical object for subsequent intervention verification.
[0048] In some exemplary embodiments, graph convolution operations are performed on the service layer subgraph of the heterogeneous hierarchical graph structure in a dual-track spatiotemporal topology subnetwork. After each graph convolution operation, cross-attention is calculated between the current business layer node features from the forward propagation path and the corresponding business layer node feature vectors from the reverse propagation path. Based on the results of cross-attention calculation, the comprehensive representation of the business layer nodes is updated to identify potential dual-track dependency patterns between the device-related vector and the business-related vector.
[0049] In the computation process within the business layer subgraph, this embodiment introduces a more refined feature interaction mechanism. After performing graph convolution operations on the business layer dependency graph to aggregate neighbor information, for each business node, instead of directly using the convolution result, a cross-attention computation process is initiated. The input to this process is two feature vectors: one is the feature of the current business node from the forward propagation path, which incorporates information from associated device anomalies; the other is the feature of the corresponding business node from the reverse propagation path, which reflects the context of the business anomaly propagating within the business layer. The cross-attention mechanism calculates the correlation between these two sets of features and generates a weighted fusion representation accordingly. This fusion representation can identify the extent to which the anomaly of the current business node is caused by underlying device problems (positive impact) and the extent to which it is affected by other abnormal services in the business call chain (reverse impact and intra-layer impact), i.e., it identifies potential dual-track dependency patterns. Finally, the comprehensive representation of the business layer node is updated based on this fusion result.
[0050] Because the update of node features within the business layer is not a simple overlay of bidirectional information, but rather an intelligent fusion through cross-attention, the model can dynamically analyze the contribution of anomaly root causes at the business node level. This distinguishes between equipment-related business bottlenecks and cascading failures between businesses, greatly enhancing the model's ability to identify common causes and chain reaction scenarios in complex failures. This fine-grained pattern recognition ensures that the generated anomaly score list not only identifies the anomalous nodes but also implicitly reveals the anomaly patterns and possible main transmission directions, improving the depth of insight in the analysis results.
[0051] Step 130: Based on the inverse hypothetical core subnetwork in the causal probing model, in the digital twin mirror environment of the target system, design virtual intervention tests for candidate root cause nodes determined by the anomaly scoring list, and generate fault-level causal paths with probability weights based on the dynamic impact of the virtual intervention tests on the key indicators of the target system.
[0052] As can be seen, the core task of step 130 is to conduct active experiments on the preliminary suspected objects generated in step 120 in a secure and isolated simulation environment to empirically verify the causal relationship and quantify its strength. The implementation of this step is based on another core component of the causal probing model—the inverse hypothetical deduction core subnetwork—and is completed in the digital twin mirror environment of the target system.
[0053] First, based on the anomaly score list output in step 120, the nodes with the highest scores are selected as candidate root cause nodes. A fully mirrored digital twin environment will be created for the target system. This environment uses snapshot technology to freeze the current system's device resource status and business traffic status, replicates real-time request data packets, and locks the topology of the configuration management database to ensure that the basic configuration remains unchanged during the simulation, forming a testbed that is consistent with the production environment but completely isolated.
[0054] Subsequently, the core subnetwork of the inverse hypothetical model designs and executes a virtual intervention test in this mirrored environment. The intervention test simulates controlled, proactive changes to the states of candidate root cause nodes. Then, it runs in the mirrored environment for a period of time, inferring and continuously monitoring changes in the distribution of key system indicators before and after the intervention. An intervention effect quantification value is calculated to quantify the strength of the causal relationship. This value compares the differences in indicator distribution before and after the intervention and introduces a time decay factor to emphasize recent effects. Based on the magnitude of this quantification value, the causal strength is graded. Finally, the quantification and grading results of the comprehensive intervention test of the inverse hypothetical model assign a probability weight to each candidate root cause node and its associated fault propagation path, thereby generating a fault-graded causal path with probability weights. This path clarifies the critical chain from the root cause to the fault phenomenon and identifies the causal confidence of each link.
[0055] The above implementation method combines digital twin technology with the experimental approach of proactive intervention, providing a verifiable causal inference method for the IT operations and maintenance field. It overcomes the limitations of traditional AI models that rely solely on historical data and statistical correlation. By directly observing how changes in the cause affect the result through controlled experiments, it effectively distinguishes between genuine strong causal nodes and spurious pseudo-correlation nodes, significantly improving the interpretability and reliability of root cause localization. Furthermore, the generated hierarchical causal paths with probability weights provide a solid and reliable basis for subsequent precise emergency response strategies.
[0056] In some exemplary embodiments, based on the inverse hypothetical inference core subnetwork in the causal probing model, virtual intervention tests are designed for candidate root cause nodes identified from the anomaly scoring list in the digital twin mirror environment of the target system, including: Based on the inverse state deduction core subnetwork in the causal probing model, a system mirror copy containing the current device resource status and business traffic status is created in the digital twin mirror environment of the target system; For candidate root cause nodes identified by the anomaly score list, virtual intervention operations are performed in the system mirror copy. Virtual intervention operations include positive repair, reverse deterioration, and combined intervention. Record the changes in the distribution of key indicators in the system image copy before and after performing virtual intervention operations; Based on the distribution changes, the quantitative value of the intervention effect of the virtual intervention operation is obtained.
[0057] The reverse simulation core sub-network creates a system mirror copy within the digital twin environment, containing the current device resource status and business traffic status. This copy is a complete clone of the production environment at a specific moment. Next, for candidate root cause nodes identified by the anomaly scoring list, specific virtual intervention operations are executed within the system mirror copy. These operations are explicitly categorized into three types: positive repair, such as expanding the database connection pool suspected of failure; negative deterioration, such as artificially limiting the network bandwidth of a server; and combined intervention, i.e., simultaneously performing repair or deterioration operations on multiple related nodes. The system then precisely records the distribution changes of key indicators (such as CPU utilization and service error rate) in the system mirror copy before and after executing these virtual intervention operations. Finally, based on the recorded changes in indicator distribution, the system performs quantitative analysis using a specific algorithm model to calculate a numerical value characterizing the intervention effect—the intervention effect quantification value.
[0058] By employing the above implementation method, virtual intervention testing can be concretized from a concept into a standardized scientific experimental procedure that is repeatable and measurable. By explicitly creating mirror copies, defining three types of intervention operations, recording distribution changes, and quantifying effects, the consistency and objectivity of the causal verification process are ensured. This structured experimental method not only reduces operational complexity but also allows the results of each test to be traceable and reproducible, greatly enhancing the credibility and persuasiveness of the causal inference conclusions.
[0059] In some exemplary embodiments, generating a fault-level causal path with probability weights includes: Based on historical failure case data and features related to candidate root cause nodes, a confusion relationship graph and corresponding feature matrix are constructed. Using the causal determination results of historical cases in the historical failure case data as labels, we perform confounding factor importance analysis on the feature matrix to obtain the confounding factor importance score; The final causal probability weight of candidate root cause nodes is determined by combining the quantitative value of the intervention effect with the importance score of the confounding factor. Based on the final causal probability weight, a fault classification causal path with probability weights is generated.
[0060] Specifically, generating the final probabilistic causal path does not solely rely on the results of real-time intervention tests; it also incorporates historical experience data for comprehensive analysis. First, based on historical failure case data and features related to the current candidate root cause node (such as resource usage patterns and topological location), a confusion relationship graph and corresponding feature matrix are constructed. The confusion relationship graph is used to visualize the potential confusion effects between system components. Then, using the pre-known causal determination results (such as strong causality or spurious correlation) for each historical failure case as annotations, the constructed feature matrix undergoes a confusion factor importance analysis. This analysis calculates the importance score of each confusion factor (such as shared storage devices), quantifying its degree of interference with causal judgment. Next, the quantitative value of the intervention effect obtained from the real-time intervention test is combined with the confusion factor importance scores obtained from the historical analysis. A specific weight fusion algorithm is used to determine the final causal probability weight for each candidate root cause node. This weight simultaneously considers the evidence strength of the real-time experiment and the correction of historical confusion patterns. Finally, based on these fused final causal probability weights, a fault-level causal path with probability weights is generated.
[0061] By incorporating historical failure case data for confounding factor analysis and combining it with real-time intervention evidence, multi-source information fusion and correction for causal judgments were achieved. This effectively overcomes the randomness or scenario-specific bias that may exist in a single experiment, uses historical experience to identify common confounding factors that are prone to misjudgment, and thus corrects and weights the real-time experimental results. This makes the final generated causal path probability weights more robust and comprehensive, significantly reduces the false alarm rate, and improves the ability to trace the causes of complex and hidden failures.
[0062] In some exemplary embodiments, to further perform high-fidelity strategy pre-playing in the virtual-real fusion simulation field, the method provided in this application embodiment further includes: Input the fault classification causal path into the virtual-real fusion simulation field; In the virtual-real fusion simulation field, based on the fault classification causal path, the production environment topology of the target system is reconstructed proportionally through containerization technology, and simulated faults corresponding to the fault classification causal path are injected from the real-time business traffic cloned from the target system. Load multiple preset emergency response strategy templates and perform parallel simulations in the reconstructed simulation environment; Record and analyze the performance indicators of each strategy template during the simulation process, and output multiple simulation schemes that include three-dimensional evaluation indicators such as business recovery degree, resource consumption ratio and operational complexity.
[0063] The virtual-real fusion simulation domain is a simulation testing environment isolated from the real production environment, constructed using lightweight virtualization technology. Its implementation process is as follows: First, the system inputs the fault classification causal path generated in step 130 into this domain. Based on this path, the domain uses containerization technology to compress and reconstruct the production environment topology of the target system at a preset ratio (e.g., 1:100). Simultaneously, it deploys a traffic splitter at the production environment's ingress gateway, clones a portion (e.g., 10%) of real-time business traffic, and injects it into the simulation environment, precisely injecting simulated faults corresponding to the fault classification causal path. For example, it simulates device overload by limiting CPU frequency or simulates service anomalies by forcing APIs to return errors through service mesh rules.
[0064] Next, the field loads various pre-defined emergency response strategy templates from the strategy knowledge base. These templates typically include experience-based strategies based on historical success stories, conservative strategies centered on resource adjustments, and aggressive strategies involving architectural changes. After loading, the field creates an independent, namespace-isolated copy of the simulation environment for each strategy and performs parallel simulations of multiple strategies within this reconstructed high-fidelity simulation environment.
[0065] During the simulation, key performance indicators (KPIs) of each strategy can be monitored and recorded in real time. After the simulation is completed, the recorded data is analyzed, and a set of simulation plans containing three-dimensional evaluation indicators is output for each strategy. These three-dimensional evaluation indicators are: business recovery rate, which measures the degree to which the strategy restores core business functions; resource consumption ratio, which quantifies the additional resource overhead brought about by the strategy execution; and operational complexity, which assesses the difficulty and risk of implementing the strategy.
[0066] By using containerized proportional reconstruction and real traffic cloning technology, a highly realistic fault simulation environment was created without affecting the production system. This allows multiple emergency strategies to be tested in parallel and objectively under zero-risk conditions. The output three-dimensional evaluation indicators provide quantitative and multi-dimensional decision-making basis for subsequent strategy optimization, completely solving the key problems of traditional simulation environments being distorted and unable to reflect real business load and resource competition.
[0067] In some exemplary embodiments, the final policy can also be generated through hierarchical intelligent agent collaboration. Specifically, the method provided in this application embodiment further includes: The fault classification causal path is input into the L1 agent, and decomposed into a priority sequence of emergency sub-objectives; and, The emergency sub-target sequence and the multiple sets of simulation schemes are input into the L2 agent; The L2 agent generates a target emergency strategy based on the base strategy source pool obtained by fusing fragments of the historical strategy library and the inference scheme, and with the three-dimensional fitness defined according to the three-dimensional evaluation index of the inference scheme as the optimization objective, through an ecological cluster evolution game mechanism. The three-dimensional fitness includes business recovery degree, resource consumption ratio and operational complexity.
[0068] The L1 and L2 agents together form a hierarchical reinforcement learning policy generator. The implementation process is as follows: First, the fault classification causal path is input into the L1 agent. The L1 agent analyzes this path, assesses the criticality of each node in the topology (e.g., combining the betweenness centrality of nodes with the priority of the services they carry), and decomposes the global emergency recovery objective into a sequence of emergency sub-objectives with clear priorities. For example, the primary objective might be to quickly reduce the database load to a safe threshold.
[0069] Subsequently, this emergency sub-objective sequence, along with multiple sets of inference schemes output from weight one, is input into the L2 agent. The L2 agent operates based on a dynamic policy resource pool—the base policy source pool. This pool is formed by fusing successful policies from the historical policy library with effective policy fragments from the current inference scheme. The L2 agent uses a comprehensive three-dimensional fitness defined according to the three-dimensional evaluation indicators of the inference scheme (business recovery rate, resource consumption ratio, and operational complexity) as the optimization objective, driving an optimization process called the ecological cluster evolutionary game mechanism, ultimately generating the target emergency policy.
[0070] By adopting a hierarchical intelligent agent architecture, the complex policy generation problem is decomposed into two sub-problems: goal planning and policy search, with L1 and L2 agents each performing their respective duties. This division of labor clearly defines that L1 is responsible for strategic planning based on causal understanding, while L2 is responsible for tactical optimization within a rich policy space. This not only improves decision-making efficiency but also makes the policy generation process both goal-oriented and policy-diverse, providing a flexible intelligent decision-making framework for dealing with complex and ever-changing fault scenarios.
[0071] In some exemplary embodiments, the specific operation process of the ecological cluster evolution game mechanism adopted by the L2 agent may include: The L2 agent constructs a candidate policy tree based on the base policy source pool; The strategies in the candidate strategy tree are classified into conservative ecological clusters, radical ecological clusters, and mixed ecological clusters. Based on three-dimensional fitness, evolutionary operations such as crossover, mutation, and forced reset are performed on strategies within various ecological clusters; Based on the average fitness of strategies within various ecological clusters, the proportion of each ecological cluster in the evolutionary process is dynamically adjusted, and the strategy with the highest fitness during the evolutionary process is output as the target emergency strategy.
[0072] First, the L2 agent constructs a candidate policy tree based on the base policy source pool and the current emergency sub-target sequence, which contains multiple possible policy combination paths. Next, the mechanism categorizes these candidate policies according to their risk and level of innovation, dividing them into a conservative ecological cluster (containing only historically validated policies), a radical ecological cluster (introducing unvalidated policies from new hypothetical scenarios), and a hybrid ecological cluster (a mixture of the former two).
[0073] The mechanism then drives strategy evolution using three-dimensional fitness as a metric. Evolutionary operations include: crossover (exchanging fragments of different strategies to generate new combinations); mutation (randomly modifying local parameters or structure of strategies); and forced reset (forcibly replacing some low-fitness strategies to explore new directions when evolution stalls). In each round of evolution, the system calculates the average fitness of all strategies within each ecological cluster and dynamically adjusts the proportion of each cluster in subsequent evolutionary groups based on this average fitness, achieving survival-of-the-fittest group competition. Finally, the mechanism outputs the strategy with the highest fitness throughout the entire evolutionary history as the target contingency strategy.
[0074] By introducing ecological cluster classification and dynamic proportion adjustment based on average fitness, this mechanism simulates a diverse strategy ecosystem. It allows conservative, mature strategies to coexist and compete with aggressive, novel strategies, leveraging historical experience to ensure the stability of the core strategy while simultaneously exploring innovative approaches to break through local optima. This design endows the strategy optimization process with strong adaptability and global search capabilities, enabling the dynamic evolution of optimal or near-optimal emergency response solutions highly adapted to specific fault scenarios, significantly improving the intelligence level and final effectiveness of strategy generation.
[0075] Figure 2 The diagram shows the overall architecture of the emergency decision-making two-way simulation system provided in this application embodiment. The system architecture is designed as a data-driven, model-coordinated, and closed-loop optimized intelligent decision-making process, mainly comprising four core modules: a data acquisition and alignment module, a causal probing model, a virtual-real simulation field, and a hierarchical reinforcement learning strategy generator. These modules are interconnected through standardized data interfaces, forming a complete technical closed loop from fault perception, root cause analysis, strategy simulation to intelligent generation.
[0076] First, the system's data input is the data acquisition and alignment module. As the foundation of the entire process, this module is responsible for real-time acquisition of key information from heterogeneous data sources. Specifically, it extracts static configuration data such as server hardware attributes and network topology from the Configuration Management Database (CMDB); obtains dynamic business dependency data such as microservice call chains and API dependency graphs from the Application Performance Management (APM) system; and continuously collects real-time monitoring indicators such as server CPU, memory utilization, network bandwidth usage, service response time, and transaction success rate through distributed device, network, and business layer probes. The core function of this module is to perform time-axis alignment and spatial topology mapping on the aforementioned multi-source, heterogeneous operational data, eliminating time-series deviations and logical gaps caused by different acquisition tools and frequencies. By vectorizing device attributes and transforming business dependencies into a weighted directed graph structure, this module ultimately outputs multi-dimensional feature vectors with a unified time series, stored in a time-series feature library, providing standardized, structured, high-quality data input for subsequent analysis. The beneficial effect of this module is that, through deep fusion coding of a unified spatiotemporal benchmark, it lays an accurate and consistent data foundation for subsequent cross-layer causal analysis, solving the core pain points of fragmented and low-integration multi-source operation and maintenance data.
[0077] The collected and processed data is fed into the system's core analysis engine—the causal probing model. This model is the innovative core of the architecture, further divided into two collaborative sub-networks: a dual-track spatiotemporal topology sub-network and a reverse scenario inference core sub-network. The dual-track spatiotemporal topology sub-network receives multi-dimensional feature vectors and, through its internally constructed physical, business, and coupling layers, models and analyzes the positive and negative anomaly impacts between devices and services, outputting a preliminary anomaly score list. This list identifies the anomaly suspicion level of each node (including devices and services) in the system. Subsequently, the reverse scenario inference core sub-network is activated. In the digital twin mirror environment of the target system, it performs virtual intervention tests (such as repair and deterioration operations) on the candidate root cause nodes in the list. By observing and quantifying the dynamic impact of these virtual interventions on key system indicators, this sub-network can empirically verify causal relationships, distinguish between strong causal nodes and spurious interference, and ultimately generate a fault classification causal path with probability weights. The beneficial effect of the causal probing model is that it is the first to integrate bidirectional graph network analysis and digital twin intervention verification in the field of operation and maintenance, which improves the root cause localization from statistical correlation-based speculation to empirical evidence based on controlled experiments, significantly improving the accuracy, interpretability and reliability of fault diagnosis.
[0078] Based on the high-reliability causal path output by the causal probing model, the system enters the strategy pre-playing and optimization phase. The virtual-real simulation module is responsible for this stage. After receiving the fault-level causal path, this module uses containerization technology to proportionally and lightweightly reconstruct the complete topology of the target production environment. Simultaneously, it injects some real business traffic through traffic cloning technology and accurately reproduces the fault scenario described by the path. In this high-fidelity, zero-risk simulation environment, the field loads multiple preset emergency response strategy templates, including empirical strategies, conservative strategies, and aggressive strategies, in parallel for simulation. Through real-time monitoring and post-event analysis, this module outputs quantitative evaluation results for each strategy, including three indicators: business recovery rate, resource consumption ratio, and operational complexity—that is, multiple simulation schemes. The beneficial effect of the virtual-real simulation domain is that it creates a "safety sandbox" highly consistent with the production environment, enabling comprehensive, objective, and quantitative effect verification and comparison of emergency strategies before implementation. This solves the problem of unreliable strategy evaluation results caused by environmental distortion in traditional drills, providing crucial evidence for final decision-making.
[0079] Finally, the system intelligently synthesizes and outputs policies through a hierarchical reinforcement learning policy generator. This generator employs a two-layer agent architecture: the L1 agent receives the fault classification causal path, parses and decomposes it into a set of priority-based emergency sub-objective sequences; the L2 agent receives this sub-objective sequence along with multiple sets of deduction schemes from the virtual-real fusion simulation field. The L2 agent maintains a dynamic base policy source pool formed by fusing fragments of historical successful policies and current deduction schemes. It drives an ecosystem cluster evolutionary game mechanism with the three-dimensional fitness (business recovery rate, resource consumption ratio, and operational complexity) defined by the three-dimensional evaluation metrics of the deduction schemes as the optimization objective. This mechanism divides the policy population into different ecosystem clusters such as conservative, aggressive, and hybrid, evolves them through crossover and mutation operations, and dynamically adjusts their proportions based on the average fitness of each cluster, ultimately outputting the target emergency policy with the highest fitness. The beneficial effect of the hierarchical reinforcement learning policy generator is that it achieves intelligent evolution from historical experience and real-time deduction results to environmentally adaptive policies. Through goal decomposition and ecological competition optimization, this generator can automatically synthesize emergency response plans that achieve the optimal balance between recovery effect, resource cost and execution risk, breaking through the rigid mode of traditional rule base matching and greatly improving the intelligence level and scenario adaptability of strategy generation.
[0080] To more clearly illustrate the technical solutions provided in the embodiments of this application, a specific business scenario is described below. This embodiment takes a microservice system supporting services such as China Mobile APP as an example. The system includes components such as user service, order service, caching service, and payment service, and is distributed and deployed in a multi-cloud environment including a self-built data center in Luoyang and multiple Internet Data Centers (IDCs) in Ningbo, Suzhou, Shandong, and other locations.
[0081] Fault Scenario Description: During peak traffic surges at the beginning or end of the month, a surge in user login requests exhausts the connection pool resources of the Redis cache cluster, leading to cache breakdown. A large number of requests directly access the backend database, causing a significant spike in database connection timeouts and error rates. This fault triggers a chain reaction, resulting in extremely slow responses from core interfaces such as billing inquiries for logged-in users. Frequent refreshes and retries by users further exacerbate the database load and latency, ultimately leading to a near-paralysis of the customer service chain.
[0082] The application process of the technical solution in this scenario is as follows: Fault Detection and Data Preparation: Distributed probes deployed on various servers and microservice nodes continuously collect and monitor metrics. When any key metric is detected, such as server CPU utilization overload or a surge in payment service API error rate that consistently exceeds the dynamic baseline calculated based on historical data, the system automatically triggers a high-level alarm. It then clusters and prioritizes fault scenarios using a multi-dimensional scoring model, generating a processing queue. Subsequently, the system loads the latest full configuration snapshot from the Configuration Management Database (CMDB) and aligns the real-time collected monitoring data with the snapshot in a spatiotemporal correlation, constructing a multi-dimensional feature vector with a unified time series.
[0083] Causal Probing and Root Cause Verification: The multidimensional feature vectors are input into the causal probing model for processing. First, the dual-track spatiotemporal topology sub-network in the model performs bidirectional anomaly propagation analysis on device-related vectors and business-related vectors. The analysis results show that the state of the Redis connection pool exhaustion has an anomaly impact strength score of 0.92 (out of 1.0) on the payment service call delay, initially identifying it as a critical anomaly node. Subsequently, the reverse simulation core sub-network is launched, and a virtual intervention test is performed on the candidate Redis node in the digital twin mirror environment constructed by the system: simulating the expansion of its connection pool capacity by 200%. The simulation verification results show that after implementing this intervention, the error rate of the core transaction service significantly decreased from 45% to 4%. Based on this quantitative value of the intervention effect, the model confirms that the insufficient Redis connection pool resources are a strong causal root cause and assigns it a probability weight greater than 95%, thereby generating a clear fault-level causal path with probability weights.
[0084] Strategy Virtual-Real Simulation and Evaluation: The aforementioned fault classification causal path is input into a virtual-real simulation environment. Based on this path, the environment lightweightly reconstructs the microservice cluster topology of the production environment using containerization technology at a 1:50 scale, and injects 10% cloned real-time business traffic. Simultaneously, fault scenarios are precisely simulated, limiting the maximum number of Redis connections to reproduce cache breakdown. In the environment, the system loads three basic emergency response strategy templates in parallel for simulation: the empirical strategy (directly expanding Redis resources) shows that the transaction success rate can be restored to 85% within 5 minutes, but resource consumption increases by 35%; the conservative strategy (degrading non-core functions) shows that resource consumption only increases by 12%, but the transaction success rate only recovers to 75%; the aggressive strategy (switching some traffic to a backup data center in another region) shows that the transaction success rate can be quickly restored to 98% within 3 minutes, but introduces approximately 20 milliseconds of network latency. Each strategy outputs a three-dimensional evaluation index including business recovery degree, resource consumption ratio, and operational complexity, forming multiple simulation schemes.
[0085] Intelligent Strategy Generation and Execution: Fault classification causal paths and multiple sets of hypothetical scenarios are input into a hierarchical reinforcement learning strategy generator. The L1 agent analyzes the path, decomposing it into a sequence of emergency sub-objectives, such as setting the primary objective as "restoring the core transaction success rate to over 95%" and assigning it a strategic weight of 0.95. The L2 agent, based on a pool of base strategies fused from historical successful strategies and current hypothetical scenario fragments, uses three-dimensional fitness as the optimization objective and performs strategy search and optimization through an ecosystem cluster evolutionary game mechanism. Ultimately, the L2 agent generates a hybrid strategy, such as "switching 50% of traffic to the backup data center while dynamically limiting the remaining traffic." This strategy is quickly validated in a twin field, showing that it can restore the transaction success rate to 98% within 4 minutes while controlling additional resource consumption to 22%. The system then transforms this strategy into a specific set of executable instructions (such as adjusting load balancing configuration and issuing rate limiting rules) and executes it in the production environment, forming a closed-loop decision-making process.
[0086] As can be seen from the above embodiments, the technical solution provided by this application can realize a complete and automated process in a real and complex fault scenario, from accurately locating strong causal root causes, rehearsing the effects of multiple strategies in a zero-risk, high-fidelity environment, to intelligently synthesizing and executing the optimal emergency strategy, which significantly improves the timeliness, accuracy and systematicness of emergency response.
[0087] Figure 3The diagram shown is a structural schematic of the causal probing model in this embodiment. This model is the core of fault root cause localization. It receives a unified temporal multidimensional feature vector from the data acquisition and alignment module as input. Through the collaborative work of its two main internal components—a dual-track spatiotemporal topology subnetwork and a reverse-stressed deduction core subnetwork—it completes the process from multi-source data analysis to generating a high-confidence fault causal path. For example... Figure 3 As shown, the workflow of this model includes: S1, Multi-source Data Input and Feature Construction. Specifically, it collects static device attributes, dynamic business dependencies, and real-time monitoring metrics from heterogeneous data sources such as the target system's Configuration Management Database (CMDB), Application Performance Monitoring (APM) system, and device / service layer probes. Through unified timeline alignment and spatial topology mapping, it eliminates temporal discrepancies and logical breaks between different data sources, providing standardized input for the dual-track spatiotemporal topology sub-network. This includes: Vectorize the physical characteristics of the device (e.g., map the CPU architecture to a 128-dimensional embedding vector).
[0088] Dynamic business dependencies are constructed as a weighted directed graph (nodes represent service instances, and edge weights are calculated from call latency and error rate), and a 256-dimensional business feature vector is generated for each node.
[0089] Generate multidimensional feature vectors with a unified time series and store them in a time series feature library.
[0090] S2, Bidirectional Anomaly Analysis of a Dual-Track Spatiotemporal Topology Network. Multidimensional feature vectors are input into the dual-track spatiotemporal topology sub-network. This sub-network is a hierarchical heterogeneous graph neural network designed to model the bidirectional anomaly impact between the device layer and the service layer. It contains three layers: a physical layer, a service layer, and a coupling layer. Each layer has input encoding, feature processing, and output generation modules.
[0091] The physical layer consists of an input encoding layer, a GRU temporal modeling layer, and an output layer. Its input is a 128-dimensional device embedding vector. Data includes: static device attributes (CPU architecture, network bandwidth, storage type, virtualization identifier, etc.); physical topology relationships (physical device nodes, network connections between devices, connection weights (e.g., actual bandwidth utilization); and device temporal monitoring metrics (CPU utilization, memory utilization, etc.). The data processing is as follows: the input encoding layer encodes static attributes and topology relationships into initial node features. The GRU layer models device temporal metrics, capturing periodic and bursty patterns. The output layer performs graph convolution aggregation based on the physical layer topology graph to calculate the intra-layer anomalous signal strength for each device node. Finally, it outputs a sequence of devices with anomalous signals, identifying the anomalous device nodes and their anomalous strength.
[0092] The business layer consists of an input encoding layer, a multi-hop graph convolutional layer, and an output layer. Its input is a 256-dimensional business feature vector. Data includes: dynamic business dependencies, such as microservice call chains, API dependency graphs, and database transaction associations; and business topology relationships, such as service instance nodes, inter-service call relationships, and call relationship weights (e.g., call latency × error rate). The data processing is as follows: the input encoding layer constructs a business layer dependency graph from the business dependencies and encodes the initial features of the nodes. The graph convolutional layer performs multi-hop information propagation on the dependency graph, extracting abnormal context features from local call chains. The output layer calculates the intra-layer anomaly score for each service instance. The final output is a sequence of anomaly scores for instance nodes, quantifying the anomaly degree of each service instance.
[0093] The coupling layer consists of an input encoding layer, a cross-attention mechanism module, and an output layer. Its inputs are the device sequence of signal anomalies output from the physical layer and the anomaly score sequence of instance nodes output from the service layer. The processing involves calculating the forward propagation weights from the physical layer to the service layer and the backward propagation weights from the service layer to the physical layer using the cross-attention mechanism, achieving bidirectional cross-layer anomaly signal propagation. For the service layer subgraph, after graph convolution, the node features of the forward and backward propagation paths are cross-attentioned to update the node comprehensive representation and identify dual-track dependency patterns. The final output is a unified list of node anomaly scores that integrates anomaly information within each node layer and cross-layer anomaly information.
[0094] Two-way information flow modeling and propagation mechanisms may include: Forward Spatiotemporal Flow: Anomaly signals propagate along the direction from device layer nodes to service layer nodes, based on the temporal characteristics of device physical states, to model their causal impact on service metrics. Impact Strength It can be calculated using the following formula:
[0095] in, For device feature vectors, To serve the feature vector, For a trainable parameter matrix, This is a decay factor that decreases by 0.8 with each additional hop. This indicates vector concatenation.
[0096] Reverse causal flow: Anomaly signals are propagated along the path from business layer nodes to device layer nodes based on the contextual features of business logic to capture the reverse effects of business anomalies on underlying devices. In this reverse propagation, for each business layer node v, the influence weights of its neighboring nodes u are calculated using an attention mechanism. :
[0097] in, and These are the feature vectors of nodes u and v, respectively. This is a trainable parameter matrix.
[0098] Based on attention weights, weighted aggregation of neighbor features is used to update the sum representation of node v. :
[0099] in, This represents the set of neighboring nodes of node v.
[0100] Multi-task collaborative training: The network adopts a multi-task learning framework that includes abnormal state classification, causal strength regression and fault source identification for joint optimization. Its total loss function is the weighted sum of the loss functions of each task, so as to improve the generalization ability and recognition accuracy of the model.
[0101] S3, Core Causal Verification of Inverse Hypothesis. The node anomaly score list output by the dual-track spatiotemporal topology subnetwork serves as input, and is empirically verified and refined by the core subnetwork of inverse hypothesis, specifically including: Candidate root cause selection and mirror environment construction: The node with the highest score is selected as the candidate root cause from the anomaly score list. Anomaly scores of candidate nodes. Calculated by the following formula:
[0102] Among them, device deviation is the Mahalanobis distance between the current metric and the historical baseline; business impact is the weighted sum of downstream service error rates. Subsequently, a system image copy containing the current device resource status and business traffic status is created in the digital twin environment, including: device resource snapshots, such as the instantaneous status of CPU, memory, and disk; business traffic mirrors, which are deep packet copies of real-time requests; and topology locking, used to freeze CMDB configuration to prevent changes during the simulation process, forming an isolated test environment.
[0103] Design and execute virtual intervention tests: Perform three types of virtual operations on candidate root cause nodes: positive intervention, negative intervention, and combined intervention. Positive intervention: restore candidate nodes to a normal state, such as resetting CPU utilization to 110% of the device's historical baseline average. Negative intervention: artificially worsen the state of candidate nodes, such as reducing network bandwidth limits to 10% of the nominal value, and verify whether the deterioration exacerbates the fault. Combined intervention: simultaneously repair multiple related nodes to explore synergistic effects, such as repairing the database connection pool and adding front-end rate limiting.
[0104] Quantifying the Intervention Effect and Causal Grading: In a mirrored environment, the distribution changes of key system indicators (device layer, such as CPU utilization, memory usage, network throughput; business layer, such as service error rate, API response time, transaction throughput) before and after the intervention are simulated and recorded. The causal strength is measured by calculating the quantitative value D of the intervention effect.
[0105] in, and The distributions of key indicators before (true) and after (counterfactual) intervention in time slice t are represented, respectively. γ is the time decay factor (e.g., 0.9), and T is the total extrapolation duration. Based on the D value and indicator recoveries, causal relationships are classified (e.g., strong causation, latent causation, spurious correlation).
[0106] Confusion factor analysis and probability weight generation: Based on historical fault case data, a confusion relationship graph and feature matrix are constructed, and the importance of confusion factors is analyzed, specifically: strong causality (Level 1): D > 0.5 and index recovery rate > 80%; potential causality (Level 2): 0.2 < D ≤ 0.5 and recovery rate 40%~80%; spurious correlation (Level 3): D ≤ 0.2 or recovery rate < 40%, thereby identifying common variables that may interfere with causal judgment. The quantitative value D of the real-time intervention effect and the importance score obtained from historical confusion factor analysis are combined, and a weighted fusion is used to determine the final causal probability weight of each candidate root cause node, thereby generating a fault classification causal path with probability weights.
[0107] For example, a confusion factor is defined as a variable that simultaneously affects both the fault phenomenon and the candidate root cause, such as a database instance shared by multiple services or network bandwidth contention across business domains. A confusion relationship graph is constructed: nodes = system components, edges = confusion intensity. A feature matrix is constructed, where behavior represents historical fault cases, columns represent node features (e.g., resource usage, topological location + environmental features, such as time period, business load), and the target variable is the causal determination result (e.g., strong causality, spurious correlation). The output feature importance ranking is used to identify highly confusing variables, such as shared storage devices with an importance score > 0.7.
[0108] Following the steps outlined above, the causal probing model ultimately generates a fault-level causal path with probability weights. This path clarifies the key transmission chain from the most probable root cause of the fault to the final phenomenon, and assigns empirically validated confidence weights to each causal link in the path, providing a highly reliable basis for subsequent emergency response strategy development.
[0109] Figure 4This is a schematic diagram illustrating a specific implementation process of the two-way emergency decision-making simulation method in this application. The process begins with tiered alarm triggering and, through four core steps—root cause analysis, virtual-real simulation, and intelligent strategy generation—ultimately forms an executable emergency decision-making plan. Figure 4 It clearly demonstrates the complete closed-loop path from data collection, causal analysis, strategy simulation to intelligent decision-making, and the key outputs of each stage, which may include: Step 1, Alarm Triggering and Scene Clustering, this step corresponds to Figure 4 The process of collecting monitoring indicator data and clustering alarm events includes: Data Acquisition and Alarm Generation: Distributed probes deployed at the device, network, and service layers continuously collect and monitor metrics. When any key metric (such as CPU utilization or service error rate) exceeds a dynamic baseline calculated from historical data and remains above it for a preset duration (e.g., 10 seconds), the system automatically triggers an alarm. Alarm levels (P0 / P1 / P2 / P3) are dynamically determined based on the degree of metric deviation (e.g., Z-score ≥ 3), duration, and cross-layer correlation. For example, if both device-level (CPU > 90%) and service-level (error rate > 5%) metrics are abnormal, the alert is escalated to a P0 level emergency alarm.
[0110] Scene clustering and severity scoring: Discrete alarm events that are spatially related within a time window (e.g., 5 minutes) are clustered into unified fault scenarios. The severity of each fault scenario is quantified using a multi-dimensional scoring model.
[0111] Among them, Impact is the sum of SLA weights for affected businesses; Urgency is the rate of metric deterioration; and Certainty is the matching degree of association rules.
[0112] Priority queue division: The processing strategy is determined based on the severity score. Specifically, a score ≥ 7: immediately enters the real-time processing queue and is allocated high-priority computing resources. A score 4 ≤ score < 7: enters the batch processing queue and is processed when the system is idle. A score < 4: judged as an occasional fluctuation, only logged, and not triggered for further analysis.
[0113] For example, during a business peak, the error rate of the order service API surged (Z-score=3.5) and the CPU of the relational database server remained above 95%. The system clustered it into a P0 level failure scenario with a severity score of 8.2, and immediately put it into the real-time processing queue.
[0114] Step 2, Root Cause Analysis and Candidate Set Generation, this step corresponds to Figure 4 The latest configuration snapshot is loaded into the root cause candidate set generation stage to accurately pinpoint the source of the fault, including: Configuration snapshot loading and data association: The system pulls the latest full configuration snapshot (including server attributes, network topology, and service deployment mapping) from the Configuration Management Database (CMDB) and constructs a business-device mapping matrix. Real-time collected monitoring metrics (such as CPU values and API error rates) are dynamically associated with device nodes and service nodes in the graph, ensuring timestamp synchronization.
[0115] Causal probing and anomaly scoring calculation: The correlated data is injected into the causal probing model (its structure is detailed in [link to model]). Figure 3 (and corresponding explanations), through dual-track spatiotemporal topology analysis of the bidirectional impact between equipment and services, and using the inverse state deduction core for causal verification, finally generating fault classification causal paths with probability weights.
[0116] Comprehensive Scoring and Candidate Set Generation: A comprehensive anomaly score is calculated for key nodes. Device anomaly score = Z-score of current indicator deviation from baseline × device criticality weight (core devices = 2.0, edge devices = 1.0). Service impact score = percentage decrease in service SLA × business priority coefficient (core business = 3.0, secondary business = 1.5). The comprehensive score is a weighted sum of the above scores. Nodes ranking in the top 5% of comprehensive scores are selected as candidate root causes, generating a root cause candidate set.
[0117] For example, the causal probing model output shows that the probability weight of the Redis cache cluster connection pool exhaustion affecting payment service latency is 0.95. Its overall anomaly score is calculated to be 9.1 (high device deviation and impact on core payment business), therefore it is included in the root cause candidate set.
[0118] Step 3, Virtual and Real Simulation and Solution Generation, this step corresponds to Figure 4 During the phase of generating multiple simulation scenarios, emergency strategies are rehearsed and evaluated in an isolated environment, specifically including: Simulation environment construction and fault injection: The production environment topology is reconstructed in a lightweight manner using containerization technology on a scale of 1:100 (e.g., 1:100). 10% of real business traffic is injected into the simulation environment by traffic splitting and cloning, and combined with fault injection technology to accurately simulate root cause failure scenarios (e.g., simulating device overload by CPU frequency limiting, and simulating service anomalies by forcing APIs to return errors through service mesh rules).
[0119] Multi-strategy parallel inference and evaluation: Loading three types of basic strategy templates from the strategy knowledge base and performing parallel inference in a simulation environment: Experience-based strategies: Historical success stories, such as "database connection pool expansion + query rate limiting".
[0120] Conservative strategy: Focus on resource adjustments, such as "vertical expansion of 20% + shutdown of non-core functions".
[0121] Aggressive strategies involve architectural changes, such as "fault node isolation + traffic switching across data centers".
[0122] The simulation process incorporates a circuit breaker mechanism, automatically terminating when core metrics deteriorate beyond a threshold. Upon completion, a set of simulation solutions is output for each strategy, including three-dimensional evaluation metrics: business recovery rate, resource consumption ratio, and operational complexity.
[0123] For example, three strategies are simulated in parallel for the aforementioned Redis failure. Scenario A (empirical strategy: scaling up Redis) achieves 85% recovery with 25% resource consumption; Scenario B (conservative strategy: service degradation) achieves 78% recovery with 12% resource consumption; and Scenario C (aggressive strategy: cross-datacenter failover) achieves 92% recovery with 40% resource consumption. The system records and outputs these three simulation scenarios and their three-dimensional metrics.
[0124] Step 4, Intelligent Strategy Synthesis and Decision Output, this step corresponds to Figure 4 From the initial policy pool construction to the formation of emergency decision-making solutions, the final policy is generated through hierarchical agent collaboration, which may include: Base strategy pool construction: Analyze the operation sequence of all deduced schemes, break them down into atomic operation fragments (such as expanding the database connection pool to 150%), and merge them with historical successful strategy fragments to build a base strategy source pool containing hundreds of base strategies.
[0125] L1 Agent: Goal Decomposition: Input the fault classification causal path into the L1 agent. The L1 agent analyzes the criticality of nodes and calculates the node criticality score. :
[0126] in, For the betweenness centrality of nodes, The priority weights for the services they support are assigned. Based on this, the global emergency objective is decomposed into a sequence of priority-based emergency sub-objectives. For example, the primary objective is to reduce the database load to a safe threshold, with a criticality of 0.92.
[0127] L2 Agent: Policy Evolution and Optimization: The L2 agent receives sub-objective sequences and multiple sets of inference schemes. Based on a base policy source pool, it constructs a candidate policy tree and drives the evolutionary game of the ecological cluster with a three-dimensional fitness F as the optimization objective.
[0128] Where R represents business recovery rate, C represents resource consumption ratio, and D represents operational complexity.
[0129] Ecological Cluster Classification and Competition: Candidate strategies are classified into three ecological clusters: conservative ecological clusters (containing only historical strategies), radical ecological clusters (introducing unverified strategies from new proposals), and mixed ecological clusters (a combination of the former two). The proportion of each cluster in the population is also considered. Based on its average fitness Dynamic adjustment, its evolutionary dynamics can be described as follows:
[0130] Where Φ represents the average fitness of all clusters. Clusters with higher fitness will have a higher proportion.
[0131] Strategy evolution operations: Within each ecosystem cluster, strategies are continuously optimized through operations such as crossover (exchanging fragments), mutation (modifying parameters or structure), and forced reset (injecting new fragments when stuck).
[0132] Strategy compression and decision output: Common patterns and key decision rules are extracted from the highly adaptive strategies that have evolved. These are then compressed into interpretable and executable rule templates and finally transformed into natural language instructions containing specific API call sequences, forming emergency decision-making solutions that can be directly issued.
[0133] For example, in the aforementioned example, the L1 agent sets the primary objective as the recovery of core payment services. The L2 agent selects a fragment from the base policy pool and generates a combined policy: 50% of traffic is switched to the backup data center, and dynamic rate limiting is implemented for the remaining traffic. This policy achieves high adaptability during evolution due to its high recoverability (95% recoverability verified in simulation) and controllable resource consumption (22%), and is compressed into a rule template output.
[0134] The emergency decision-making bidirectional extrapolation method provided in this application collects and time-aligns multi-source data such as device monitoring, attributes, business dependencies, and traffic to construct a unified temporal-series multi-dimensional feature vector, providing a high-quality data foundation with spatiotemporal consistency for subsequent analysis. This vector is then input into a causal probing model, which first uses its dual-track spatiotemporal topology subnetwork to perform anomaly propagation analysis on the positive and negative impacts between device and business vectors, identifying key anomaly nodes and initially achieving quantitative modeling of cross-layer bidirectional impacts. Based on this, the model's inverse hypothetical extrapolation core subnetwork is used to design and execute virtual intervention tests on candidate root cause nodes in the digital twin mirror environment of the target system. This verifies and quantifies the true causal relationship between nodes based on the actual dynamic impact of the intervention on key system indicators. This combination of techniques ultimately generates a fault-level causal path with probability weights, significantly improving the accuracy, verifiability, and interpretability of root cause localization in complex fault scenarios, laying a reliable foundation for subsequent accurate emergency decision-making.
[0135] Figure 5This is a schematic diagram of the structure of a data processing apparatus 500 provided for an exemplary embodiment of this application. For example... Figure 5 As shown, the device 500 includes: a receiving module 510 and a generating module 520, wherein: The feature vector construction module 510 is used to collect device monitoring data, device attribute data, business dependency data and business traffic data of the target system, and to align the collected multi-source data with time axis to construct a multi-dimensional feature vector with a unified time series. The causal probing model 520 is used to perform anomaly propagation analysis on the positive and negative influences between business-related vectors and equipment-related vectors in the multidimensional feature vectors based on a dual-track spatiotemporal topology subnetwork, to obtain an anomaly score list; and, based on the inverse scenario deduction core subnetwork, in the digital twin mirror environment of the target system, to design virtual intervention tests on the candidate root cause nodes determined by the anomaly score list, and to generate a fault-level causal path with probability weights based on the dynamic impact of the virtual intervention tests on the key indicators of the target system.
[0136] The emergency decision-making bidirectional simulation device 500 provided in this application collects and time-aligns multi-source data such as device monitoring, attributes, business dependencies, and traffic to construct a unified temporal-series multi-dimensional feature vector, providing a high-quality data foundation with spatiotemporal consistency for subsequent analysis. This vector is then input into a causal probing model, which first uses its dual-track spatiotemporal topology subnetwork to perform anomaly propagation analysis on the positive and negative impacts between device and business vectors, identifying key anomaly nodes and initially achieving quantitative modeling of cross-layer bidirectional impacts. Based on this, the model's inverse simulation core subnetwork is used to design and execute virtual intervention tests on candidate root cause nodes in the digital twin mirror environment of the target system. This verifies and quantifies the true causal relationship between nodes based on the actual dynamic impact of the intervention on key system indicators. This combination of techniques ultimately generates a fault-level causal path with probability weights, significantly improving the accuracy, verifiability, and interpretability of root cause localization in complex fault scenarios, laying a reliable foundation for subsequent accurate emergency decision-making.
[0137] Optionally, the causal probing model 520 is specifically used for: Based on the physical layer of the dual-track spatiotemporal topology subnetwork, the device-related vectors are subjected to time-series modeling and anomaly aggregation, and a sequence of devices with abnormal signals is output. Based on the service layer of the dual-track spatiotemporal topology sub-network, graph convolution propagation and anomaly aggregation are performed on the service-related vectors to output the anomaly scoring sequence of instance nodes; Based on the coupling layer of the dual-track spatiotemporal topology subnetwork, the abnormal device sequence and the abnormal score sequence of the instance node are bidirectionally propagated and fused to generate the abnormal score list.
[0138] Optionally, the causal probing model 520 is specifically used for: The forward propagation weight from the physical layer to the business layer and the backward propagation weight from the business layer to the physical layer are calculated using the cross-attention mechanism in the coupling layer. Based on the forward propagation weight, the abnormal signals in the sequence of devices with abnormal signals are transmitted to the relevant service layer nodes; Based on the backpropagation weights, the abnormal scores in the abnormal score sequence of the instance nodes are passed to the relevant physical layer nodes. By integrating cross-layer anomaly information received by each node with its own intra-layer anomaly information, a unified anomaly score list is generated.
[0139] Optionally, the cross-layer propagation mechanism includes: A forward propagation path is used to propagate the abnormal signal from the device layer node to the service layer node, based on the temporal characteristics of the device's physical state, to model the causal impact of the abnormal signal on service metrics; and, The abnormal signal is propagated in reverse along the direction from the business layer node to the device layer node. Based on the context characteristics of the business logic, the abnormal signal is propagated to capture the reverse effect of business anomalies on the underlying devices.
[0140] Optionally, the causal probing model 520 is specifically used for: In the dual-track spatiotemporal topology subnetwork, graph convolution operation is performed on the service layer subgraph of the heterogeneous hierarchical graph structure; After each graph convolution operation, cross-attention is calculated between the current business layer node features from the forward propagation path and the corresponding business layer node feature vectors from the reverse propagation path. Based on the results of the cross-attention calculation, the comprehensive representation of the business layer node is updated to identify the potential dual-track dependency pattern between the device-related vector and the business-related vector.
[0141] Optionally, the causal probing model 520 is specifically used for: Based on the inverse state deduction core subnetwork in the causal probing model, a system mirror copy containing the current device resource status and service traffic status is created in the digital twin mirror environment of the target system; For candidate root cause nodes identified by the anomaly score list, virtual intervention operations are performed in the system mirror copy. The virtual intervention operations include positive repair, reverse deterioration, and combined intervention. Record the distribution changes of key indicators in the system mirror copy before and after performing the virtual intervention operation; Based on the aforementioned distribution changes, the quantitative value of the intervention effect of the virtual intervention operation is obtained.
[0142] Optionally, the causal probing model 520 is specifically used for: Based on historical failure case data and features related to the candidate root cause nodes, a confusion relationship graph and a corresponding feature matrix are constructed. Using the causal determination results of historical cases in the historical fault case data as labels, the feature matrix is subjected to confusion factor importance analysis to obtain confusion factor importance scores; The final causal probability weight of the candidate root cause node is determined by combining the quantitative value of the intervention effect with the importance score of the confounding factor. Based on the final causal probability weight, the fault classification causal path with probability weight is generated.
[0143] Optionally, the device further includes a deduction module 530, used for: The fault classification causal path is input into the virtual-real fusion simulation field; In the virtual-real fusion simulation field, based on the fault classification causal path, the production environment topology of the target system is reconstructed proportionally using containerization technology, and a portion of the real-time business traffic cloned from the target system and the simulated fault corresponding to the fault classification causal path are injected. Multiple preset emergency response strategy templates are loaded and simulated in parallel within the reconstructed simulation environment; Record and analyze the performance indicators of each strategy template during the simulation process, and output multiple simulation schemes that include three-dimensional evaluation indicators such as business recovery degree, resource consumption ratio and operational complexity.
[0144] Optionally, the apparatus further includes a strategy generation module 540, used for: The fault classification causal path is input into the L1 agent, and decomposed into a priority sequence of emergency sub-objectives; and, The emergency sub-target sequence and the multiple sets of simulation schemes are input into the L2 agent; The L2 agent generates a target emergency strategy based on a base strategy source pool obtained by fusing fragments of the historical strategy library and the inference scheme, and with the three-dimensional fitness defined according to the three-dimensional evaluation index of the inference scheme as the optimization objective, through an ecological cluster evolution game mechanism. The three-dimensional fitness includes business recovery degree, resource consumption ratio and operational complexity.
[0145] Optionally, the strategy generation module 540 is specifically used for: The L2 agent constructs a candidate policy tree based on the base policy source pool; The strategies in the candidate strategy tree are classified into conservative ecological clusters, radical ecological clusters, and mixed ecological clusters. Based on the aforementioned three-dimensional fitness, evolutionary operations involving crossover, mutation, and forced reset are performed on strategies within various ecological clusters; Based on the average fitness of strategies within various ecological clusters, the proportion of each ecological cluster in the evolutionary process is dynamically adjusted, and the strategy with the highest fitness during the evolutionary process is output as the target emergency strategy.
[0146] The 500-degree emergency decision-making two-way simulation device can achieve... Figures 1-4 For details of the method implementation examples, please refer to [link / reference]. Figures 1-4 The two-way simulation method for emergency decision-making shown in the embodiment will not be described in detail here.
[0147] Figure 6 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. For example... Figure 6 As shown, the device includes a memory 61 and a processor 62.
[0148] Memory 61 is used to store computer programs and can be configured to store various other data to support operation on the computing device. Examples of this data include instructions for any application or method used to operate on the computing device, contact data, phone book data, messages, images, videos, etc.
[0149] Processor 62, coupled to memory 61, is used to execute computer programs in memory 61 for: collecting device monitoring data, device attribute data, business dependency data, and business traffic data of the target system; aligning the collected multi-source data along the time axis to construct a multi-dimensional feature vector with a unified time series; inputting the multi-dimensional feature vector into a causal probing model to perform anomaly propagation analysis on the positive and negative influences between business-related vectors and device-related vectors in the multi-dimensional feature vector based on the dual-track spatiotemporal topology subnetwork in the causal probing model, obtaining an anomaly score list; and, based on the inverse scenario deduction core subnetwork in the causal probing model, designing virtual intervention tests for candidate root cause nodes determined by the anomaly score list in the digital twin mirror environment of the target system, and generating a fault-level causal path with probability weights based on the dynamic impact of the virtual intervention tests on the key indicators of the target system.
[0150] The electronic device provided in this application collects and time-aligns multi-source data such as device monitoring, attributes, business dependencies, and traffic to construct a unified temporal-series multi-dimensional feature vector, providing a high-quality data foundation with spatiotemporal consistency for subsequent analysis. This vector is then input into a causal probing model. First, its dual-track spatiotemporal topology sub-network is used to analyze the anomaly propagation of positive and negative influences between device and business vectors, identifying key anomaly nodes and initially achieving quantitative modeling of cross-layer bidirectional influences. Based on this, the model's inverse hypothetical core sub-network is used to design and execute virtual intervention tests on candidate root cause nodes in the digital twin mirror environment of the target system. This verifies and quantifies the true causal relationship between nodes based on the actual dynamic impact of the intervention on key system indicators. This combination of techniques ultimately generates a fault-level causal path with probability weights, significantly improving the accuracy, verifiability, and interpretability of root cause localization in complex fault scenarios, laying a reliable foundation for subsequent accurate emergency decision-making.
[0151] Furthermore, such as Figure 6 As shown, the electronic device also includes other components such as a communication component 63, a display 64, a power supply component 65, and an audio component 66. Figure 6 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 6 The components shown. Additionally, depending on the implementation of the traffic playback device, Figure 6 The components within the dashed box are optional, not mandatory. For example, when an electronic device is implemented as a terminal device such as a smartphone, tablet, or desktop computer, it may include... Figure 6 The components within the dashed box; when the electronic device is implemented as a server-side device such as a conventional server, cloud server, data center, or server array, it may be excluded. Figure 6The component within the dashed box.
[0152] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described data processing method embodiments.
[0153] Accordingly, this application also provides a computer program product, which stores instructions that, when executed by a computer, cause the computer to perform the steps in the data processing method embodiments provided in this application.
[0154] The above Figure 6 The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component may further include a Near Field Communication (NFC) module, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, etc.
[0155] The above Figure 6 The memory in the memory can be implemented by any class of volatile or non-volatile storage devices or combinations thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0156] The above Figure 6 The display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action, but also the duration and pressure associated with the touch or swipe operation.
[0157] The above Figure 6 The power supply component provides power to the various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.
[0158] The above Figure 6The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0159] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0160] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0161] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0162] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0163] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0164] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0165] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other classes of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0166] It should be understood that the training and prediction processes of the AI models involved in the various embodiments of this specification all adhere to multiple legal and compliant principles, including legal data sources, compliant data content, compliant data governance, compliant training objectives and schemes, compliant training processes, compliant training environments and tools, and compliant ethical verification of training results, and comply with the requirements of Article 5 of the Patent Law. Among them: Data Source Legality: All datasets used for AI model training were obtained through legal channels, covering three categories: publicly authorized data, data authorized by partners, and self-collected compliant data. Publicly authorized data originates from compliant data sources following open-source licenses such as Apache 2.0, with complete copyright attribution and authorization scope clearly marked, and no unauthorized open-source code or data reuse exists. Data authorized by partners has been subject to formal data usage agreements, clearly defining the scope, duration, and confidentiality obligations, and possessing a complete authorization chain. For self-collected data involving personal information, strict informed consent procedures have been followed, and anonymization processes (including but not limited to field masking, feature anonymization, and differential privacy technology applications) have been implemented to remove personally identifiable information, fully complying with the requirements of the "Interim Measures for the Administration of Generative Artificial Intelligence Services," the "Personal Information Protection Law," and other relevant laws and regulations.
[0167] Data Content Compliance: The AI model's dataset undergoes multiple screening and cleaning processes to remove all content that may violate social morality or harm public interests. It contains no information that endangers national or public safety, nor does it involve the illegal acquisition or use of genetic resources. For data in sensitive areas (such as healthcare and finance), an additional privacy-preserving computation module (including federated learning and secure multi-party computation technologies) ensures that the data is "usable but not visible," avoiding compliance risks during the original data transmission process and ensuring that the data application scenarios and uses comply with public order and good morals and industry regulatory requirements.
[0168] Data governance compliance: A complete data traceability system is established during the AI model training process to automatically record the source, collection time, annotation process, cleaning rules, and permission allocation of training data, generating traceable compliance reports to ensure that the data is verifiable throughout its entire lifecycle. The dataset annotation process for AI models is completed by a professional human R&D team, clearly defining the proportion of human creative contributions, avoiding reliance on AI-generated data that has not undergone substantial human modification, and complying with the examination requirements for "human main contributions" in AI patent applications.
[0169] Training objectives and scheme compliance: The AI model training objective focuses on intelligent assistance for IT operation and maintenance emergency decision-making. The training scheme and the final output results do not violate any mandatory provisions of laws and administrative regulations, do not harm the public interest or the legitimate rights and interests of others, and do not pose any potential risks of being used for illegal activities, privacy infringement, or public safety disruption. The training scheme and the final output results strictly adhere to the ethical principle of "intelligent for good".
[0170] Compliance of the training process: A closed-loop training framework is adopted to ensure compliance and controllability of the training process. The specific process is as follows: First, training samples are obtained through compliant data sources. After the aforementioned data cleaning and desensitization, they are input into the neural network model to generate preliminary training results. Second, an expert system is introduced to verify the preliminary results. Based on preset rules and human expert experience, the feasibility of the results is evaluated, and outputs that may pose ethical risks or compliance hazards are corrected (such as removing decision logic that violates public order and good morals, and adjusting model parameters that do not comply with safety regulations). Finally, the loss function weights are dynamically optimized based on the feedback from the expert system to strengthen the model's learning of compliant results, avoid overfitting errors or non-compliant labels, and form a closed-loop control of "data input - model training - expert verification - parameter optimization - result feedback" to ensure that the entire training process complies with A5 ethical review requirements.
[0171] Training Environment and Tools Compliance: AI model training is implemented based on nationally licensed chips and a compliant training platform. All open-source frameworks and components used in the training process have obtained their corresponding licenses, and copyright statements and patent citation information are fully retained, with no instances of infringement or reuse. The training environment is constructed using virtual devices (containers / virtual machines) with fixed random seeds and initial parameter configurations to ensure the reproducibility of the training process. Furthermore, through access control and operation log recording, risks such as data leakage and parameter tampering during training are prevented, ensuring the security and compliance of the training process.
[0172] Training results ethical verification and compliance: After the model is trained, it undergoes additional third-party ethical compliance assessment and algorithm filing review to verify that the model output does not violate social morality or harm public interests. For potentially sensitive scenarios (such as public services and intelligent decision-making), a special result verification mechanism is established to ensure that the model always complies with Article 5 of the Patent Law and relevant laws and regulations in practical applications.
[0173] In summary, the data and training process used in the AI model of this specification strictly comply with the relevant provisions of Article 5 of the Patent Law and the Patent Examination Guidelines (2023 Edition), and there are no violations of laws, social ethics, public interests, or illegal use of genetic resources. It fully meets the compliance requirements for patent authorization.
[0174] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0175] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A two-way simulation method for emergency decision-making, characterized in that, include: Collect device monitoring data, device attribute data, business dependency data, and business traffic data of the target system, and align the collected multi-source data with time axis to construct a multi-dimensional feature vector with a unified time series. The multidimensional feature vectors are input into the causal probing model, and anomaly propagation analysis is performed on the positive and negative influences between the business-related vectors and the equipment-related vectors in the multidimensional feature vectors based on the dual-track spatiotemporal topology sub-network in the causal probing model, so as to obtain an anomaly score list. Furthermore, based on the inverse scenario deduction core subnetwork in the causal probing model, in the digital twin mirror environment of the target system, virtual intervention tests are designed for the candidate root cause nodes determined by the anomaly scoring list, and based on the dynamic impact of the virtual intervention tests on the key indicators of the target system, a fault-level causal path with probability weights is generated.
2. The method as described in claim 1, characterized in that, The anomaly propagation analysis, based on the dual-track spatiotemporal topology subnetwork in the causal probing model, examines the positive and negative influences between the business-related vectors and equipment-related vectors in the multidimensional feature vectors, yielding an anomaly score list, including: Based on the physical layer of the dual-track spatiotemporal topology subnetwork, the device-related vectors are subjected to time-series modeling and anomaly aggregation, and a sequence of devices with abnormal signals is output. Based on the service layer of the dual-track spatiotemporal topology sub-network, graph convolution propagation and anomaly aggregation are performed on the service-related vectors to output the anomaly scoring sequence of instance nodes; Based on the coupling layer of the dual-track spatiotemporal topology subnetwork, the abnormal device sequence and the abnormal score sequence of the instance node are bidirectionally propagated and fused to generate the abnormal score list.
3. The method as described in claim 2, characterized in that, The coupling layer based on the dual-track spatiotemporal topology subnetwork performs bidirectional propagation and fusion to generate the anomaly score list, including: The forward propagation weight from the physical layer to the business layer and the backward propagation weight from the business layer to the physical layer are calculated using the cross-attention mechanism in the coupling layer. Based on the forward propagation weight, the abnormal signals in the sequence of devices with abnormal signals are transmitted to the relevant service layer nodes; Based on the backpropagation weights, the abnormal scores in the abnormal score sequence of the instance nodes are passed to the relevant physical layer nodes. By integrating cross-layer anomaly information received by each node with its own intra-layer anomaly information, a unified anomaly score list is generated.
4. The method as described in claim 3, characterized in that, The cross-layer propagation mechanism includes: A forward propagation path is used to propagate the abnormal signal from the device layer node to the service layer node, based on the temporal characteristics of the device's physical state, to model the causal impact of the abnormal signal on service metrics; and, The abnormal signal is propagated in reverse along the direction from the business layer node to the device layer node. Based on the context characteristics of the business logic, the abnormal signal is propagated to capture the reverse effect of business anomalies on the underlying devices.
5. The method as described in claim 4, characterized in that, In the dual-track spatiotemporal topology subnetwork, graph convolution operation is performed on the service layer subgraph of the heterogeneous hierarchical graph structure; After each graph convolution operation, cross-attention is calculated between the current business layer node features from the forward propagation path and the corresponding business layer node feature vectors from the reverse propagation path. Based on the results of the cross-attention calculation, the comprehensive representation of the business layer node is updated to identify the potential dual-track dependency pattern between the device-related vector and the business-related vector.
6. The method as described in claim 1, characterized in that, Based on the inverse hypothetical inference core subnetwork in the causal probing model, in the digital twin mirror environment of the target system, virtual intervention tests are designed for the candidate root cause nodes determined by the anomaly scoring list, including: Based on the inverse state deduction core subnetwork in the causal probing model, a system mirror copy containing the current device resource status and service traffic status is created in the digital twin mirror environment of the target system; For candidate root cause nodes identified by the anomaly score list, virtual intervention operations are performed in the system mirror copy. The virtual intervention operations include positive repair, reverse deterioration, and combined intervention. Record the distribution changes of key indicators in the system mirror copy before and after performing the virtual intervention operation; Based on the aforementioned distribution changes, the quantitative value of the intervention effect of the virtual intervention operation is obtained.
7. The method as described in claim 6, characterized in that, The generation of fault classification causal paths with probability weights includes: Based on historical failure case data and features related to the candidate root cause nodes, a confusion relationship graph and a corresponding feature matrix are constructed. Using the causal determination results of historical cases in the historical fault case data as labels, the feature matrix is subjected to confusion factor importance analysis to obtain confusion factor importance scores; The final causal probability weight of the candidate root cause node is determined by combining the quantitative value of the intervention effect with the importance score of the confounding factor. Based on the final causal probability weight, the fault classification causal path with probability weight is generated.
8. The method as described in claim 1, characterized in that, The method further includes: The fault classification causal path is input into the virtual-real fusion simulation field; In the virtual-real fusion simulation field, based on the fault classification causal path, the production environment topology of the target system is reconstructed proportionally using containerization technology, and a portion of the real-time business traffic cloned from the target system and the simulated fault corresponding to the fault classification causal path are injected. Multiple preset emergency response strategy templates are loaded and simulated in parallel within the reconstructed simulation environment; Record and analyze the performance indicators of each strategy template during the simulation process, and output multiple simulation schemes that include three-dimensional evaluation indicators such as business recovery degree, resource consumption ratio and operational complexity.
9. The method as described in claim 8, characterized in that, Also includes: The fault classification causal path is input into the L1 agent and decomposed to obtain an emergency sub-target sequence with priority; as well as, The emergency sub-target sequence and the multiple sets of simulation schemes are input into the L2 agent; The L2 agent generates a target emergency strategy based on a base strategy source pool obtained by fusing fragments of the historical strategy library and the inference scheme, and with the three-dimensional fitness defined according to the three-dimensional evaluation index of the inference scheme as the optimization objective, through an ecological cluster evolution game mechanism. The three-dimensional fitness includes business recovery degree, resource consumption ratio and operational complexity.
10. The method as described in claim 9, characterized in that, The target emergency strategy generated through the ecological cluster evolution game mechanism includes: The L2 agent constructs a candidate policy tree based on the base policy source pool; The strategies in the candidate strategy tree are classified into conservative ecological clusters, radical ecological clusters, and mixed ecological clusters. Based on the aforementioned three-dimensional fitness, evolutionary operations involving crossover, mutation, and forced reset are performed on strategies within various ecological clusters; Based on the average fitness of strategies within various ecological clusters, the proportion of each ecological cluster in the evolutionary process is dynamically adjusted, and the strategy with the highest fitness during the evolutionary process is output as the target emergency strategy.