Industrial device cascade fault monitoring method based on digital twin model inference

By constructing a digital twin model of Bayesian causal network and spatiotemporal hypergraph structure, the fault propagation path is predicted and control strategies are generated, solving the problem of accurate location and active defense of cascading faults in complex industrial systems and achieving efficient fault prevention and control.

CN122490983APending Publication Date: 2026-07-31BEIJING EASY TIMES DIGITAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING EASY TIMES DIGITAL TECH
Filing Date
2026-04-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately pinpoint the root cause of cascading failures in complex industrial systems and are unable to proactively defend against fault propagation. Traditional digital twin systems lack the ability to extrapolate "hypothetical interventions," resulting in a disconnect between alarms and dynamic blocking strategies, making it impossible to effectively predict fault propagation paths and generate blocking strategies.

Method used

By employing a digital twin model-based approach, a Bayesian causal network model is constructed by acquiring multi-source industrial data. This model is then combined with a spatiotemporal hypergraph structure to predict fault propagation paths, simulate physical interventions, and generate control strategies, thereby achieving closed-loop control throughout the entire process.

Benefits of technology

It realizes closed-loop control of the entire process of cascading faults in industrial equipment, from passive monitoring to active defense, which improves the accuracy and generalization ability of cascading fault prevention and control, and reduces the economic losses of fault prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490983A_ABST
    Figure CN122490983A_ABST
Patent Text Reader

Abstract

This invention provides a method for monitoring cascading faults in industrial equipment based on digital twin model inference. The method includes acquiring multi-source industrial data, constructing a Bayesian causal network model containing leaky mode noise and OR gate logic using a simulation base, predicting the cascading fault propagation path along the model based on a spatiotemporal hypergraph structure when abnormal drift in equipment parameters is detected, identifying candidate intervention devices and performing network truncation simulation physical intervention, calculating the posterior probability of downstream cascading failure, generating a strategy with the goal of minimizing control costs and failure probabilities, and issuing control commands to physical actuators. This method eliminates the reliance on massive amounts of balanced data, accurately tracks fault propagation across physical fields, realizes virtual intervention simulation, generates a minimum-cost blocking strategy, and achieves proactive defense against cascading faults in industrial equipment, reducing economic losses from fault prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial artificial intelligence and equipment fault diagnosis and prediction technology, and in particular to a method for monitoring cascading faults in industrial equipment based on inference from a digital twin model. Background Technology

[0002] In the accelerating process of industrial intelligence, modern industrial systems (such as petrochemical refining and multi-stage water treatment networks) are evolving towards a form with highly tightly coupled physical structures. The core assets within these systems interact frequently with material flows through complex control logic, making the systems highly susceptible to domino effects and cascading failures.

[0003] Currently, academia and industry have widely adopted data-driven fault diagnosis methods for feature extraction and pattern recognition. However, these existing technical solutions mainly suffer from the following problems: existing diagnostic models struggle to accurately pinpoint the root cause in the face of massive concurrent alarms; models relying solely on data-driven approaches and statistical correlations are easily affected by unobserved confounding factors, and their generalization ability drops sharply when faced with rare fault samples or non-stationary environments; traditional digital twin systems lack the ability to extrapolate hypotheses about interventions, making it impossible to perform rigorous "What-if" virtual intervention hypothesis analysis in virtual space, hindering the transition from "passive monitoring" to "active defense"; and existing early warning systems suffer from a severe disconnect between alarms and dynamic blocking strategies. When detecting initial minor anomalies, the system cannot predict fault propagation paths in advance, nor can it automatically generate "minimum cost blocking strategies." Simple and crude shutdowns are not only costly but can sometimes exacerbate system crashes.

[0004] Therefore, how to achieve accurate, efficient, and low-cost cascading fault prediction and active blocking in complex industrial systems is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] This invention provides a method for monitoring cascading faults in industrial equipment based on digital twin model inference, which enables closed-loop control of the entire process from passive monitoring to active defense of cascading faults in industrial equipment, thereby improving the accuracy and generalization ability of cascading fault prevention and control in complex industrial systems.

[0006] On the one hand, this invention provides a method for monitoring cascading faults in industrial equipment based on digital twin model inference, which includes: Multi-source industrial data is acquired and integrated with a simulation base to construct a Bayesian causal network model containing equipment nodes. The Bayesian causal network model is constructed based on a structural intervention algorithm for active learning, and the network nodes contain leaky mode noise or gate logic. When abnormal drift of equipment parameters is detected, the propagation path of cascaded faults is predicted along the Bayesian causal network model based on the spatiotemporal hypergraph structure, wherein a hyperedge envelope in the spatiotemporal hypergraph structure contains at least two equipment nodes with physical coupling relationships. Candidate intervention devices are determined based on the diffusion path. For the candidate intervention devices, the operation of truncating the input causal edges is performed in the Bayesian causal network model to simulate physical intervention, and the cascade failure posterior probability of downstream devices is calculated. Based on the cascading failure posterior probability, a control strategy is generated with the goal of minimizing control costs and minimizing failure probability, and control commands are sent to the corresponding physical actuators to prevent the spread of faults.

[0007] On the other hand, the present invention also provides an industrial equipment cascading fault monitoring system based on digital twin model inference, which includes: Multi-source heterogeneous sensing and twin layer, used to acquire data from IoT sensors and programmable logic controller data acquisition gateways in the field, and to carry a virtual co-simulation base based on cyber-physical systems and original equipment manufacturer models; The causal network dynamic learning layer has a built-in structural intervention active learning module and a leaky mode noise parameter estimation module. It is used to call the structural intervention algorithm to intervene and test the reported data stream and simulation results, and dynamically output and update the directed acyclic topology graph structure with causal connections. The spatiotemporal hypergraph evolution prediction layer includes an anomaly feature drift capture module and a hypergraph attention network inference engine, which are used to receive minute anomaly signals, perform multi-scale temporal and spatial diffusion inference along the causal topology, and output a risk heat map. The causal blocking and closed-loop control layer includes a truncation network virtual intervention simulation module and a multi-objective control optimization solver, which is used to output the minimum cost physical blocking command and penetrate down to establish a communication connection with the field execution mechanism to complete the issuance and execution of physical intervention.

[0008] The present invention provides a method for monitoring cascading faults in industrial equipment based on digital twin model inference. By acquiring multi-source industrial data and fusing it with a simulation base, a Bayesian causal network model embedded with leaky mode noise or gate logic is constructed using a structural intervention algorithm. When abnormal drift of equipment parameters is detected, the spatiotemporal hypergraph structure of multi-physical coupled equipment nodes is used to predict the cascading fault propagation path along the Bayesian causal network model using hyperedge envelopes. Candidate intervention devices are selected based on the propagation path, and physical intervention is simulated by truncating the input causal edges in the Bayesian causal network model. The posterior probability of cascading failure of downstream equipment is calculated, and a control strategy is generated with the goal of minimizing control costs and minimizing failure probability. Control commands are issued to physical actuators to block the spread of faults. This method realizes a closed-loop control of cascading faults in industrial equipment from passive monitoring to active defense, improves the accuracy and generalization ability of cascading fault prevention and control in complex industrial systems, and effectively reduces the economic losses caused by fault prevention and control. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating the industrial equipment cascade fault monitoring method based on digital twin model inference provided in this embodiment of the invention. Figure 2 This is a schematic diagram of the structure of the industrial equipment cascade fault monitoring system based on digital twin model inference provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0012] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0013] Figure 1 This is a flowchart illustrating the industrial equipment cascading fault monitoring method based on digital twin model inference provided in this embodiment of the invention.

[0014] like Figure 1 As shown, the industrial equipment cascading fault monitoring method based on digital twin model inference provided in this embodiment of the invention mainly includes the following steps: 101. Acquire multi-source industrial data and integrate them with the simulation base to construct a Bayesian causal network model containing equipment nodes; In a specific implementation, the Bayesian causal network model is actively constructed based on a structural intervention algorithm, and the network nodes contain missing mode noise OR gate logic. The Bayesian causal network model, based on Bayesian probability, uses device nodes as network nodes and causal relationships between devices as directed edges to construct a directed acyclic graph model, used to represent the causal propagation relationships of faults among industrial equipment. The structural intervention algorithm is an algorithm for actively learning causal relationships. By intervening in device variables in a virtual space, it establishes and eliminates false causal edges, constructing an accurate causal network. Missing mode noise OR gate logic is a logical model embedded in the Bayesian causal network nodes, used to quantify the propagation probability of faults in the time dimension and the uncertainty caused by unobserved external environments to fault propagation.

[0015] In this embodiment, multi-source industrial data, including equipment operating parameters and physical topology relationships, can be collected from IoT sensors, programmable logic controller gateways, and other devices in the industrial field. This multi-source industrial data is then input into a simulation platform formed by fusing a cyber-physical system model, a physical simulation model of the original equipment manufacturer (OEM), and a discrete event simulation model. A structural intervention algorithm is run within the simulation platform to actively learn the causal relationships between devices. Simultaneously, leaky mode noise or gate logic is embedded into each device node of the constructed Bayesian causal network model to complete the model's construction and optimization.

[0016] In a specific implementation, the process of constructing a Bayesian causal network model containing device nodes using a fusion simulation base may include: By integrating real-time sensor data from cyber-physical systems, physical simulation models of individual devices from their original equipment manufacturers, and discrete event simulation models of the entire plant through co-simulation technology; The structural intervention algorithm is used to actively learn the directed acyclic graph, forcibly anchoring selected device variables to preset values ​​in the digital space, and observing the responses of other variables in the network to establish causal edges. By fixing intermediate variables and applying synchronization perturbations to the initial variables, spurious direct causal edges are eliminated, and a Bayesian causal network is constructed. Leakage-mode noise OR gate logic is embedded in network nodes to quantify the probability of fault propagation over time and the leakage uncertainty caused by unobserved external environment.

[0017] In detail, by employing co-simulation technology, real-time sensor data collected by cyber-physical systems, physical simulation models of original equipment manufacturers of various industrial equipment, and discrete event simulation models reflecting the entire plant's production process are deeply integrated to form a simulation base that can accurately map the physical interaction relationships in the industrial field, providing a foundation for the construction of Bayesian causal network models.

[0018] The structural intervention algorithm is run in the digital space corresponding to the simulation base to carry out active learning of the directed acyclic graph. Based on the operating characteristics and monitoring requirements of the industrial equipment, the equipment variable to be intervened is selected and the equipment variable is forcibly anchored to a preset fixed value. At the same time, the state response of other equipment variables in the network is monitored and recorded in real time. Based on the response relationship between variables, the causal edges between equipment nodes are established, and the causal network topology is initially constructed.

[0019] In the initially constructed causal network, the device variable in the middle of the causal edge is selected as the intermediate variable and its value is fixed. Synchronous perturbation operation is applied to the device variable of the initial intervention. The changes in the causal relationship of each device node after the perturbation are observed and analyzed. False direct causal edges that have no actual physical connection and are only related to data are removed from the network, thus completing the accurate construction of the Bayesian causal network.

[0020] Leakage mode noise OR gate logic is embedded in each device node of the completed Bayesian causal network. This logic is used to quantify the probability of fault propagation in the time dimension between device nodes. At the same time, the leakage uncertainty brought by unmonitored external environmental factors to fault propagation is evaluated and quantified, so that the Bayesian causal network model can better fit the actual fault propagation characteristics of industrial sites.

[0021] This embodiment improves the accuracy of the model's analysis in industrial settings by constructing a high-fidelity simulation base and a precise causal network, eliminating false associations, quantifying fault propagation characteristics, and enhancing the model's analytical accuracy in industrial settings.

[0022] 102. When abnormal drift of equipment parameters is detected, the propagation path of cascaded faults is predicted based on the spatiotemporal hypergraph structure along the Bayesian causal network model. In a specific implementation, a hyperedge envelope in the spatiotemporal hypergraph structure contains at least two physically coupled device nodes. Equipment operating parameters from the industrial site can be continuously collected and input into a Bayesian causal network model to monitor changes in these parameters in real time. When abnormal drift in equipment parameters deviating from the normal operating range is detected, based on the pre-constructed spatiotemporal hypergraph structure and the physical coupling relationships between devices, the path of fault propagation from the abnormal device node to other device nodes is analyzed along the causal connections of the device nodes in the Bayesian causal network model, thus predicting the cascading fault propagation path.

[0023] Specifically, the process of predicting the propagation path of cascading failures can include: A spatiotemporal hypergraph structure is used to model the high-order coupling relationships between industrial equipment. In the spatiotemporal hypergraph structure, each hyperedge is used to simultaneously enclose two or more equipment nodes with physical coupling relationships. Based on the aforementioned spatiotemporal hypergraph structure, for multivariate time series data with multiple historical time steps that are continuously input, the hypergraph attention network operator is invoked and combined with the gated loop unit to update the state of the device nodes enclosed by each hyperedge across time steps, so as to continuously perceive the dynamic changes of device parameters. When an abnormal drift in device parameters below the normal alarm threshold is detected based on the updated device node status, the device node with the abnormal drift is taken as the primary lesion, the spatiotemporal hypergraph structure is activated, and a fault propagation probability heatmap spanning physical dimensions is output forward along the directed causal hyperedge to lock the downstream assets in the risk path and obtain the propagation path of the cascaded fault.

[0024] In detail, a spatiotemporal hypergraph structure can be constructed based on the physical coupling relationships such as fluid, thermal, and mechanical relationships between various industrial equipment in the industrial field. Equipment nodes with physical coupling relationships are used as nodes of the hypergraph. By enclosing two or more equipment nodes with higher-order physical coupling relationships with a single hyperedge, the higher-order coupling relationships between industrial equipment can be modeled, allowing the structure to accurately reflect the actual correlation characteristics between equipment.

[0025] Multivariate time series data collected continuously from the industrial site at multiple historical time steps are input into the spatiotemporal hypergraph structure. The hypergraph attention network operator is invoked to aggregate the spatial dimension of the device node features within the hyperedge. At the same time, the gating loop unit is combined to fuse the temporal dimension of the device node states at different time steps, realizing cross-time step updates of the device node states enclosed by each hyperedge. Through continuous state updates, the dynamic changes of device parameters are perceived in real time.

[0026] The updated device node status is compared with the preset normal operation threshold. When a minor abnormal drift in device parameters below the normal alarm threshold is detected, the device node with the abnormal drift is marked as the primary lesion. At the same time, the spatiotemporal hypergraph structure is activated. Based on the directed causal relationship in the Bayesian causal network model, the probability of fault propagation is extrapolated to the downstream device nodes along the directed causal hyperedge in the spatiotemporal hypergraph structure. A fault propagation probability heatmap spanning physical dimensions such as fluid and heat is output. Based on the node connection relationship of high-probability fault propagation in the heatmap, the downstream assets in the risk path are locked, and the propagation path of cascading faults is finally determined.

[0027] In a specific implementation, calling the hypergraph attention network operator and combining it with the gated recurrent unit to update the state of the device nodes enclosed by each hyperedge across time steps includes: Encode the multivariate time series data at each time step into the initial state vector of the corresponding device node; For each hyperedge, the initial state vectors of all device nodes enclosed by the hyperedge at the current time step are aggregated to generate a hyperedge feature vector; The hypergraph attention network operator calculates the attention weights between each device node and other device nodes within its hyperedge based on the feature vectors of each hyperedge, and then performs weighted aggregation on the initial state vector according to the attention weights to generate a spatial aggregated state vector for each device node. The spatial aggregated state vector of each device node and the hidden state vector of the previous time step are input into the gated loop unit. The state fusion of the time dimension is performed through the update gate and reset gate of the gated loop unit, and the updated state vector of each device node at the current time step is output to complete the node state update across time steps.

[0028] In detail, the multivariate time series data collected from the industrial site is processed by vector encoding, and the multivariate time series data at each time step is transformed into the initial state vector of the corresponding equipment node, so as to realize the vectorized representation of the equipment operating parameters and provide a basis for the calculation of node state.

[0029] For each hyperedge in the spatiotemporal hypergraph structure, the initial state vectors of all device nodes enclosed by the hyperedge are collected at the current time step. Feature aggregation processing is performed on all initial state vectors to generate a hyperedge feature vector that can characterize the overall features of the hyperedge.

[0030] The hypergraph attention network operator is invoked, and the feature vectors of each hyperedge are used as input to calculate the attention weight between each device node and other device nodes within its hyperedge. This weight is used to characterize the tightness of the physical coupling relationship between device nodes. Based on the calculated attention weights, the initial state vectors of each device node are weighted and aggregated to generate a spatial aggregated state vector that integrates the association features of nodes within the hyperedge.

[0031] The hidden state vector of each device node in the previous time step is extracted. This hidden state vector is then concatenated with the spatial aggregated state vector of the current time step and input into the gated loop unit. The update gate of the gated loop unit filters the update information of the node state and the reset gate selectively forgets the historical state information, thus completing the state fusion processing in the time dimension. Finally, the updated state vector of each device node in the current time step is output, thereby realizing the cross-time step update of the state of the device nodes enclosed by each hyperedge.

[0032] 103. Based on the diffusion path, determine candidate intervention devices. For the candidate intervention devices, perform the operation of truncating the input causal edges in the Bayesian causal network model to simulate physical intervention, and calculate the cascade failure posterior probability of downstream devices. In a specific implementation process, based on the predicted cascading fault propagation path, device nodes that are on the propagation path and have physical adjustment capabilities can be selected as candidate intervention devices. For the node corresponding to the candidate intervention device, the operation of truncating the input causal edges is performed in the Bayesian causal network model to simulate the physical intervention behavior of the device in virtual space. Based on the model structure after intervention, the posterior probability of cascading failure of downstream key equipment is calculated.

[0033] Specifically, the heatmap includes multiple device nodes and their corresponding fault propagation probability values; the process of determining candidate intervention devices based on the propagation path may include: Along the propagation direction of cascading faults, the device nodes on the propagation path of the cascading faults in the fault propagation probability heatmap are layered according to their topological distance from the source device of the anomaly, forming multiple risk levels. In each risk level, device nodes with physical controllability are selected as candidate intervention devices, wherein physical controllability includes the adjustable capability of the physical actuators corresponding to the nodes. The selected candidate intervention devices are sorted according to their respective risk levels and their failure propagation probability values ​​to generate a candidate intervention device list.

[0034] In detail, the fault propagation probability heatmap contains multiple equipment nodes in the industrial field, and each equipment node is matched with a corresponding fault propagation probability value, which is used to characterize the probability of the node failing due to the influence of the primary lesion.

[0035] When determining candidate intervention devices, the topological distance between each device node on the cascade fault propagation path and the abnormal source device is calculated first, along the direction of cascade fault propagation, starting from the abnormal source device that has experienced abnormal drift. Based on the topological distance, each device node is stratified to form multiple risk levels. The closer the topological distance, the higher the risk level of the device node.

[0036] Physical controllability screening is conducted on equipment nodes in each risk level to determine whether the physical actuators corresponding to the equipment nodes have adjustable capabilities such as speed adjustment and opening adjustment. Equipment nodes with physical controllability are selected as candidate intervention devices for cascade fault prevention and control.

[0037] All candidate intervention devices selected through screening are sorted. First, they are initially sorted according to their respective risk levels, with higher-risk-level candidate intervention devices ranked higher. Within the same risk level, they are then sorted a second time according to the failure propagation probability value corresponding to the device node, from high to low. Finally, a list of candidate intervention devices is generated based on the sorting results, providing a clear basis for device selection in subsequent physical intervention simulations.

[0038] This embodiment improves the rationality of equipment selection and effectively enhances fault blocking efficiency by stratified screening and sorting of candidate intervention devices.

[0039] In a specific implementation, the device nodes on the propagation path of the cascading fault in the fault propagation probability heatmap are layered according to their topological distance from the source device, forming multiple risk levels, including: Based on the directed causal edges in the Bayesian causal network model, a causal propagation tree starting from the abnormal source device is constructed. Calculate the directed path length between each device node in the causal propagation tree and the anomaly source device, and use the directed path length as the topological distance; Based on topological distance, device nodes are divided into a direct association layer, an indirect association layer, and an end-effect layer; The direct association layer corresponds to device nodes with a topological distance of a first preset distance value, the indirect association layer corresponds to device nodes with a topological distance of a second preset distance value to a preset upper limit value, and the end influence layer corresponds to device nodes with a topological distance greater than the preset upper limit value. The candidate intervention devices are preferentially selected from the direct association layer and the indirect association layer. The second preset distance value is greater than the first preset distance value.

[0040] In detail, based on the directed causal edges between device nodes in the Bayesian causal network model, with the abnormal source device that experiences abnormal drift as the root node, a causal propagation tree radiating outward is constructed according to the causal propagation relationship of the fault. This causal propagation tree can clearly characterize the propagation relationship of the fault from the abnormal source device to other device nodes.

[0041] In the constructed causal propagation tree, the directed path length between each device node and the source device of the anomaly is calculated. This length is the number of directed causal edges traversed between the device node and the source device of the anomaly. This directed path length is used as the topological distance between each device node and the source device of the anomaly.

[0042] Based on preset distance classification rules, each device node is divided into three risk levels according to topological distance: the direct association layer, the indirect association layer, and the end-effect layer. Device nodes with a topological distance of the first preset distance value are classified into the direct association layer, device nodes with a topological distance of the second preset distance value to a preset upper limit value are classified into the indirect association layer, and device nodes with a topological distance greater than the preset upper limit value are classified into the end-effect layer, wherein the second preset distance value is greater than the first preset distance value.

[0043] When screening candidate intervention devices, priority is given to screening from the direct and indirect association layers, with the end-effect layer only used as a supplementary screening scope. This ensures that candidate intervention devices can block the fault in the early stages of fault propagation.

[0044] This embodiment quantifies and classifies equipment risk levels, prioritizes high-risk intervention nodes, achieves early fault blocking, and reduces the risk of spread.

[0045] In a specific implementation, the process of truncating input causal edges in the Bayesian causal network model to simulate physical intervention and calculating the posterior probability of cascade failure of downstream devices may include: Obtain the directed acyclic graph structure in the Bayesian causal network model, and determine the first node corresponding to the candidate intervention device and the second node corresponding to the downstream key device; In the directed acyclic graph structure, all input causal edges pointing to the first node are cut off, so that the value of the first node no longer depends on its parent node in the probabilistic graphical model, but is forcibly assigned the set hypothetical intervention value. Based on the directed acyclic graph structure after cutting off the input causal edges, the cascade failure posterior probability is calculated when the second node takes the value of a preset failure state threshold. The cascade failure posterior probability is used to quantitatively characterize the probability change of cascade failure in downstream key equipment after simulated physical intervention.

[0046] In detail, the corresponding directed acyclic graph structure is extracted from the completed Bayesian causal network model. Based on the generated list of candidate intervention devices, the device node corresponding to the candidate intervention device in the directed acyclic graph structure is determined as the first node. At the same time, based on the core production needs of the industrial site, the downstream key equipment on the cascading fault propagation path is determined, and the device node corresponding to it in the directed acyclic graph structure is designated as the second node.

[0047] In the directed acyclic graph structure, a truncation operation is performed on all input causal edges pointing to the first node, completely severing the causal connection between the first node and its upstream parent node. This ensures that the value of the first node is no longer affected by the state of its parent node in the probabilistic graphical model of the Bayesian causal network. At the same time, the value of the first node is forcibly assigned to a set hypothetical intervention value, which is set according to the operational requirements of physical intervention. This simulates the physical intervention behavior of candidate intervention devices in virtual space.

[0048] Using the truncated directed acyclic graph structure after truncation of the input causal edges as the computational basis, and combining it with the Bayesian probability calculation method, the posterior probability of cascading failure under the condition that the value of the second node reaches the preset failure state threshold is used as the computational objective. This probability value can quantitatively reflect the change in the probability of cascading failure of downstream key equipment after the implementation of the hypothetical physical intervention, providing a quantitative indicator for the generation of subsequent control strategies.

[0049] It should be noted that the above calculation of the posterior probability of cascading failure of downstream equipment can be achieved based on the intervention probability formula corresponding to the core operator under the Judea Pearl causal inference framework: 1. Core Formula: Calculation of posterior probability of simulated physical intervention: The core deduction of this invention is based on the following intervention distribution formula: ; 2. Core parameters and Explanation of the function of operators: and (Intervention action): Represents the target intervention device or control variable (e.g., the control command of a variable frequency drive). This represents the physical operation value that the system forcibly sets in the virtual space (e.g., forcibly setting the frequency to 45Hz).

[0050] and (Simulation Objective): This represents the core assets or key state variables downstream of the cascade path (e.g., the vibration intensity of a downstream reactor). This represents the safety threshold status that we are concerned about.

[0051] The essential function of operators: The role of operators is to distinguish between "passive observation" and "active intervention".

[0052] In traditional machine learning or Bayesian networks, conditional probabilities are calculated. This simply means "when we see (observe) the sensor display device". The state is At that time, equipment The state probability. This observation contains a large number of "confounders" and cannot determine the true control effect.

[0053] Here it is introduced An operator represents a system simulating an active physical control behavior (i.e., an agent issuing commands). In terms of algorithm implementation, The operator performs "graph surgery / mutilation": in a directed acyclic graph (DAG), it forcibly cuts off (erases) all pointers to nodes. The input causal edge. This means The state is no longer affected by other upstream variables, but is completely subject to the control commands forcibly issued by the system. .

[0054] Therefore, by calculation It can not only know "what happened", but also accurately answer "what-if" questions, such as "what if the system ignores other current operating conditions and forcibly puts the equipment..." Adjusted to Downstream key equipment What is the probability of a cascading failure? This mathematical operation simulates countless physical trials at low cost in the digital twin space, so that the final output dynamic blocking command is no longer a guess based on historical data, but a quantitative deduction conclusion based on strict causal logic.

[0055] 104. Based on the cascading failure posterior probability, a control strategy is generated with the goal of minimizing control cost and minimizing failure probability, and control commands are sent to the corresponding physical actuators to prevent the spread of faults.

[0056] In a specific implementation process, the process of generating a control strategy with the goal of minimizing control costs and minimizing failure probability, and issuing control commands to the corresponding physical actuators to prevent fault propagation, may include: Establish a multi-objective optimization function, with minimizing the overall control cost and minimizing the probability of line or equipment failure as the constraints; The physical space for finding candidate control devices is narrowed down using an adaptive shortest path algorithm; Within the physical space, the multi-objective optimization function is solved to obtain the optimal solution located at the Pareto front; The optimal solution is converted into specific control commands and sent to the corresponding physical actuators to prevent the spread of the fault.

[0057] In detail, the overall control cost of physical intervention can be quantitatively modeled by combining factors such as equipment operating costs, adjustment operation costs, and capacity loss costs in the industrial site. At the same time, the failure probability of the line or equipment is used as another core quantitative indicator to establish a multi-objective optimization function. This function takes minimizing the overall control cost and minimizing the failure probability of the line or equipment as the two major constraints to achieve multi-objective optimization constraints on the fault blocking strategy.

[0058] By using an adaptive shortest path algorithm, a dynamic weighted topology graph is constructed based on the physical topology of equipment in the industrial site. According to the fault propagation probability heatmap and fault propagation characteristics, the boundary and pruning conditions of the path search are dynamically adjusted to narrow down the physical space of candidate regulating equipment, eliminate equipment nodes with no intervention value, and focus the solution of the multi-objective optimization function on the range of high-value candidate regulating equipment.

[0059] The equipment parameters within the physical space of the narrowed candidate adjustment equipment are used as input to a multi-objective optimization function. A multi-objective optimization algorithm is used to solve the function, and solutions that cannot improve the performance of one objective without reducing the performance of another are selected, forming a set of optimal solutions located at the Pareto front. The optimal solution that best fits the production needs of the industrial site is selected from this set.

[0060] The optimal solution obtained is transformed into specific control commands that can be recognized and executed by physical actuators. These commands contain specific parameters and operational requirements for equipment adjustment. The control commands are then sent to the corresponding physical actuators via industrial communication networks. The physical actuators then perform the specific adjustment operations according to the commands, thereby effectively preventing the spread of cascading faults.

[0061] In a specific implementation, the process of narrowing down the physical space for finding candidate adjustment devices using an adaptive shortest path algorithm may include: Construct a dynamic weighted physical topology graph with industrial equipment as nodes and physical connections as edges. The weight of each edge is a dynamic weight function that integrates fault propagation sensitivity, physical distance, real-time operating condition factors, and controllability cost. Starting with the detected abnormal source device, the initial search boundary is defined according to the fault propagation probability heatmap output by the spatiotemporal hypergraph structure. A forward search is performed in the dynamic weighted physical topology graph, and the search path is pruned in real time based on the cumulative dynamic weight threshold, the fault propagation probability threshold, and the physical controllability conditions of the device. A reverse search is performed simultaneously from the key downstream equipment, and the intersection area of ​​the forward search and reverse search paths is determined as the narrowed physical space of the candidate regulating equipment.

[0062] In detail, a physical topology graph is constructed by using various industrial equipment in the industrial site as nodes and the physical connections between the equipment, such as fluid, mechanical, and thermal connections, as edges. A dynamic weight function is configured for each edge of this topology graph. This function integrates multiple core factors such as fault propagation sensitivity, physical distance between equipment, real-time operating conditions in the industrial site, and controllability costs of the equipment. The weight values ​​of the edges are dynamically adjusted according to the real-time changes of each factor to form a dynamically weighted physical topology graph.

[0063] Starting with the detected abnormal source device as the starting point of the forward search, the initial search boundary is defined based on the fault propagation probability heatmap output by the spatiotemporal hypergraph structure and the magnitude of the fault propagation probability. The forward search is performed in the direction of fault propagation in the dynamically weighted physical topology graph. During the search, the cumulative dynamic weight value is calculated in real time. Search paths with cumulative dynamic weight values ​​exceeding the cumulative dynamic weight threshold, device fault propagation probability below the fault propagation probability threshold, and devices that do not have physical controllability are pruned in real time, and invalid search paths are eliminated.

[0064] Simultaneously, starting from the critical downstream equipment on the cascading fault propagation path, a reverse search is performed in the dynamically weighted physical topology graph. The search direction is opposite to the fault propagation direction, and the pruning conditions for the reverse search are consistent with those for the forward search. After the forward and reverse searches are completed, the intersection region of the paths obtained from the two searches is extracted, and this region is determined as the narrowed-down candidate regulating device physical space.

[0065] This embodiment improves the efficiency of subsequent optimization and the accuracy of intervention by dynamically narrowing the search range of candidate devices and eliminating invalid nodes.

[0066] In a specific implementation process, a timeliness adjustment factor is introduced into the dynamic weight function. The timeliness adjustment factor is monotonically increasing with the cumulative time after the fault occurs, which is used to simulate the gradual narrowing effect of the intervention window during the fault propagation process. During the forward search process, the time difference between the current search time and the initial time of the fault is used as input to dynamically calculate the cumulative dynamic weight threshold, so that the search space gradually shrinks as the intervention window period narrows. When the time difference exceeds the preset intervention failure time, the cumulative dynamic weight threshold is set to infinity, the search is terminated and the devices on the fault propagation path are no longer included in the candidate adjustment device physical space.

[0067] In detail, a timeliness adjustment factor is introduced into the dynamic weight function corresponding to each edge of the dynamic weighted physical topology graph. The value of this factor is monotonically increasing with the cumulative time after the fault occurs. As time goes by after the fault occurs, the value of the timeliness adjustment factor continues to increase, thereby simulating the effect of the intervention window gradually narrowing during the propagation of cascading faults in industrial sites, so that the dynamic weight function can better fit the time characteristics of fault propagation.

[0068] During the forward search of the dynamically weighted physical topology graph, the time difference between the current search time and the initial fault time is recorded in real time. This time difference is used as an input parameter and substituted into a preset threshold calculation model to dynamically calculate the cumulative dynamic weight threshold. As the time difference increases, the intervention window gradually narrows, and the value of the cumulative dynamic weight threshold decreases accordingly. This allows the search space of the forward search to gradually shrink as the intervention window narrows, enabling the search process to focus on device nodes with greater immediate intervention value.

[0069] An intervention failure time is preset. When the time difference between the current search time and the fault initiation time exceeds the intervention failure time, it is determined that the intervention operation on the fault propagation path can no longer play an effective blocking role. The cumulative dynamic weight threshold is set to infinity, the forward search in that direction is immediately terminated, and the equipment on the fault propagation path is stopped from being included in the candidate adjustment equipment physical space, so as to avoid invalid equipment screening and deduction.

[0070] This embodiment combines timeliness factors to dynamically adjust the search space, improving the timeliness of screening and algorithm efficiency, and ensuring the actual prevention and control value of intervention equipment.

[0071] The following are some specific examples to illustrate this: Example 1: Application Scenario and Composition of Rotating Machinery Resonance Cascade Blocking Based on Variable Frequency Fine-Tuning: This embodiment is applied to a large-scale fluid pipeline system, which includes upstream fluid source equipment, rigid pipelines, and downstream large centrifugal pumps. Each device is equipped with a high-frequency vibration sensor, and the downstream centrifugal pump is controlled by a variable frequency drive (VFD).

[0072] Implementation Process: When upstream equipment experiences minor high-frequency, minute abnormal vibrations due to slight bearing wear, these vibrations have not yet reached the single-machine alarm threshold. The spatiotemporal hypergraph network of this invention captures this weak feature and predicts, along with the physical topology, that if this vibration frequency propagates along the fluid or pipeline, it will overlap with the inherent mechanical frequency of the downstream large centrifugal pump, easily triggering destructive physical resonance. At this point, the system utilizes a causal digital twin space... The operator performs counterfactual deduction to calculate the resonance probability at different VFD frequencies.

[0073] Effect: The system, through a multi-objective optimization algorithm, determined that a complete shutdown would be too costly. Ultimately, the system automatically sent a command to the variable frequency drive to pre-adjust the speed of the downstream centrifugal pump by two percent. This operation, without interrupting production line operation and with only minimal capacity fluctuations, cleverly avoided the abnormal excitation frequency transmitted from the source, successfully cutting off the cascading propagation path of mechanical resonance from a physical perspective, and preventing catastrophic consequences such as downstream bearing breakage or blade fracture.

[0074] Example 2: Application Scenario and Composition of Pressure Cascade Interruption and Dynamic Valve Reconfiguration in Complex Chemical Fluid Pipeline Networks: This embodiment is applied to a complex multi-stage petrochemical refining pipeline network. The system includes multi-stage reactors, transfer pumps, and multiple intelligent regulating and control valves (such as valve A, valve B, bypass valve C, etc.), with pressure and flow sensors deployed at each node.

[0075] Implementation Process: The system detected a slight internal leak or jamming in upstream valve A, causing a minor pressure parameter drift in a local pipeline section. In traditional systems, this is usually ignored until the downstream core reactor shuts down due to material shortage or overpressure. The causal digital twin system of this invention treats this drift as the "primary lesion." The system initiates a What-if analysis, extrapolating multiple timelines in parallel: "If the current situation remains unchanged, what is the probability that the fault will spread to core equipment C?" and "If the opening of valve B is fine-tuned to 60% at this time, or the bypass circuit is cut off in advance, how much will the posterior probability decrease?"

[0076] Effects: Based on the simulation results, the system automatically optimizes and issues commands before a fault (such as pump cavitation or reactor material shortage) actually occurs, precisely adjusting the opening of the relay intelligent regulating valve B to 60%. This embodiment effectively absorbs pressure fluctuations transmitted from upstream, ensuring the stable operation of the core reactor and transforming passive post-event maintenance into proactive online topology reconfiguration.

[0077] Example 3: Flexible load reduction blocking of "one-to-many" cooling water systems in distributed factories Application Scenarios and Components: This embodiment is applied to a distributed manufacturing plant with high-order physical coupling (one-to-many). The system consists of a main cooling water circulation pump cluster and dozens of parallel, independent chemical reaction vessels supplied by it.

[0078] Implementation Process: A slight performance degradation occurred in the main cooling water pump, resulting in a 5% decrease in the total cooling water flow rate. Since a hyperedge encloses the main pump and all downstream reactors, the spatiotemporal hypergraph model quickly calculated that under this "one-control-many" structure, the probability of all reactors simultaneously triggering a high-temperature cascade alarm increases dramatically. The system accurately assessed the leakage probability caused by external ambient temperature using a LeakyNoisy-OR node.

[0079] Effects: In response to the risk of global overload, the system utilizes a non-dominated sorting genetic algorithm (NSGA-II) to balance minimizing overall control costs with minimizing equipment failure probability. The system's final shutdown command does not shut down the main pump (which would paralyze the entire plant), but rather automatically cuts off or reduces the operating load of two non-critical reactors (partial load shedding). This ensures that the remaining high-value core reactors have sufficient cooling water, preventing a cascading disaster of system-wide overload failure with minimal localized capacity concessions.

[0080] In summary, this embodiment has the following effects: 1. Breaking down data silos and imbalances significantly improves diagnostic robustness in complex environments: Because existing purely data-driven models (such as CNNs) heavily rely on large, balanced datasets and are easily influenced by unobserved confounding factors, they can fall into a trap of "strong correlation but weak causality." To address this technical deficiency, this invention innovatively introduces a Structural Intervention Algorithm (SIA) into the digital twin foundation. By implementing a dual active intervention mechanism of forced variable anchoring and synchronous perturbation in the digital space, spurious direct causal edges are eliminated. Therefore, the causal topology graph derived in this invention has extremely high reliability, completely eliminating the dependence on rare fault samples from real production lines, and greatly improving the model's generalization ability and prediction accuracy in non-stationary industrial environments and even under zero-fault sample conditions.

[0081] 2. A fundamental leap has been achieved from "passive alarm shutdown" to "low-cost active physical defense": Traditional industrial early warning systems suffer from a structural flaw: a severe disconnect between alarm and physical containment strategies. Faced with cascading failures, they often resort to reactive, reactive alarms or drastic, costly power outages, resulting in significant economic losses and triggering secondary stress spikes. To address this deficiency, this invention utilizes a spatiotemporal hypergraph to precisely pinpoint the path of weak anomalies across physical fields, and introduces… The operator performs counterfactual simulations of the truncated network. This allows the system to virtually rehearse the consequences of millions of control strategies. Therefore, combined with multi-objective optimization algorithms (such as NSGA-II), this invention can automatically generate and execute dynamic blocking commands with "minimum control cost," such as fine-tuning inverter speed and bypass unloading. This perfectly avoids production interruptions caused by downtime while maximizing equipment safety (minimizing the probability of failure), achieving an extremely flexible and economical proactive intelligent defense.

[0082] Based on the same general inventive concept, this invention also protects an industrial equipment cascade fault monitoring system based on digital twin model inference. The industrial equipment cascade fault monitoring system based on digital twin model inference provided by this invention will be described below. The industrial equipment cascade fault monitoring system based on digital twin model inference described below can be referred to in correspondence with the industrial equipment cascade fault monitoring method based on digital twin model inference described above.

[0083] Figure 2 This is a schematic diagram of the structure of the industrial equipment cascade fault monitoring system based on digital twin model inference provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the industrial equipment cascade fault monitoring system based on digital twin model inference in this embodiment includes a multi-source heterogeneous sensing and twin layer 21, a causal network dynamic learning layer 22, a spatiotemporal hypergraph evolution prediction layer 23, and a causal blocking and closed-loop control layer 24.

[0084] Among them, the multi-source heterogeneous sensing and twin layer 21 is used to acquire data from IoT sensors and programmable logic controller data acquisition gateways on site, and carries a virtual co-simulation base based on cyber-physical systems and original equipment manufacturer models. The dynamic learning layer 22 of the causal network has a built-in structure intervention active learning module and a leaky mode noise parameter estimation module. It is used to call the structure intervention algorithm to intervene and test the reported data stream and simulation results, and dynamically output and update the directed acyclic topology graph structure with causal connections. The spatiotemporal hypergraph evolution prediction layer 23 includes an anomaly feature drift capture module and a hypergraph attention network inference engine, which is used to receive small anomaly signals, perform multi-scale temporal and spatial diffusion inference along the causal topology, and output a risk heat map. The causal blocking and closed-loop control layer 24 includes a truncation network virtual intervention simulation module and a multi-objective control optimization solver, which is used to output the minimum cost physical blocking command and penetrate down to establish a communication connection with the field execution mechanism to complete the issuance and execution of physical intervention.

[0085] Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. The processor 310, communication interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions from the memory 330 to execute a method for monitoring cascaded faults in industrial equipment based on digital twin models.

[0086] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0087] It should be noted that all relevant information that may be involved in the various embodiments of the present invention is processed in strict accordance with the requirements of laws and regulations, following the principles of legality, legitimacy, and necessity, based on the reasonable purpose of the business scenario, and is information that users actively provide or generate during the use of the product / service, as well as information obtained with user authorization.

[0088] The information processed by this invention may vary depending on the specific product / service scenario and should be based on the specific scenario in which the user uses the product / service. This may involve user account information, device information, or other related information. This invention will treat the relevant information and its processing with the utmost diligence.

[0089] This invention places great emphasis on the security of relevant information and has adopted reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent unauthorized access, public disclosure, use, modification, damage or loss of relevant information.

[0090] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0091] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for monitoring cascading faults in industrial equipment based on digital twin model inference, characterized in that, include: Multi-source industrial data is acquired and integrated with a simulation base to construct a Bayesian causal network model containing equipment nodes. The Bayesian causal network model is constructed based on a structural intervention algorithm for active learning, and the network nodes contain leaky mode noise or gate logic. When abnormal drift of equipment parameters is detected, the propagation path of cascaded faults is predicted along the Bayesian causal network model based on the spatiotemporal hypergraph structure, wherein a hyperedge envelope in the spatiotemporal hypergraph structure contains at least two equipment nodes with physical coupling relationships. Candidate intervention devices are determined based on the diffusion path. For the candidate intervention devices, the operation of truncating the input causal edges is performed in the Bayesian causal network model to simulate physical intervention, and the cascade failure posterior probability of downstream devices is calculated. Based on the cascading failure posterior probability, a control strategy is generated with the goal of minimizing control costs and minimizing failure probability, and control commands are sent to the corresponding physical actuators to prevent the spread of faults.

2. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 1, characterized in that, The integrated simulation platform constructs a Bayesian causal network model containing device nodes, including: By integrating real-time sensor data from cyber-physical systems, physical simulation models of individual devices from their original equipment manufacturers, and discrete event simulation models of the entire plant through co-simulation technology; The structural intervention algorithm is used to actively learn the directed acyclic graph, forcibly anchoring selected device variables to preset values ​​in the digital space, and observing the responses of other variables in the network to establish causal edges. By fixing intermediate variables and applying synchronization perturbations to the initial variables, spurious direct causal edges are eliminated, and a Bayesian causal network is constructed. Leakage-mode noise OR gate logic is embedded in network nodes to quantify the probability of fault propagation over time and the leakage uncertainty caused by unobserved external environment.

3. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 1, characterized in that, Based on the spatiotemporal hypergraph structure, the propagation path of cascaded faults predicted along the Bayesian causal network model includes: A spatiotemporal hypergraph structure is used to model the high-order coupling relationships between industrial equipment. In the spatiotemporal hypergraph structure, each hyperedge is used to simultaneously enclose two or more equipment nodes with physical coupling relationships. Based on the aforementioned spatiotemporal hypergraph structure, for multivariate time series data with multiple historical time steps that are continuously input, the hypergraph attention network operator is invoked and combined with the gated loop unit to update the state of the device nodes enclosed by each hyperedge across time steps, so as to continuously perceive the dynamic changes of device parameters. When an abnormal drift in device parameters below the normal alarm threshold is detected based on the updated device node status, the device node with the abnormal drift is taken as the primary lesion, the spatiotemporal hypergraph structure is activated, and a fault propagation probability heatmap spanning physical dimensions is output forward along the directed causal hyperedge to lock the downstream assets in the risk path and obtain the propagation path of the cascaded fault.

4. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 3, characterized in that, The heatmap contains multiple device nodes and their corresponding fault propagation probability values; Candidate intervention devices identified based on the diffusion path include: Along the propagation direction of cascading faults, the device nodes on the propagation path of the cascading faults in the fault propagation probability heatmap are layered according to their topological distance from the source device of the anomaly, forming multiple risk levels. In each risk level, device nodes with physical controllability are selected as candidate intervention devices, wherein physical controllability includes the adjustable capability of the physical actuators corresponding to the nodes. The selected candidate intervention devices are sorted according to their respective risk levels and their failure propagation probability values ​​to generate a candidate intervention device list.

5. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 4, characterized in that, The device nodes on the propagation path of the cascading fault in the fault propagation probability heatmap are layered according to their topological distance from the source device, forming multiple risk levels, including: Based on the directed causal edges in the Bayesian causal network model, a causal propagation tree starting from the abnormal source device is constructed. Calculate the directed path length between each device node in the causal propagation tree and the anomaly source device, and use the directed path length as the topological distance; Based on topological distance, device nodes are divided into a direct association layer, an indirect association layer, and an end-effect layer; The direct association layer corresponds to device nodes with a topological distance of a first preset distance value, the indirect association layer corresponds to device nodes with a topological distance of a second preset distance value to a preset upper limit value, and the end influence layer corresponds to device nodes with a topological distance greater than the preset upper limit value. The candidate intervention devices are preferentially selected from the direct association layer and the indirect association layer. The second preset distance value is greater than the first preset distance value.

6. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 3, characterized in that, Calling the hypergraph attention network operator and combining it with a gated recurrent unit to update the state of the device nodes enclosed by each hyperedge across time steps includes: Encode the multivariate time series data at each time step into the initial state vector of the corresponding device node; For each hyperedge, the initial state vectors of all device nodes enclosed by the hyperedge at the current time step are aggregated to generate a hyperedge feature vector; The hypergraph attention network operator calculates the attention weights between each device node and other device nodes within its hyperedge based on the feature vectors of each hyperedge, and then performs weighted aggregation on the initial state vector according to the attention weights to generate a spatial aggregated state vector for each device node. The spatial aggregated state vector of each device node and the hidden state vector of the previous time step are input into the gated loop unit. The state fusion of the time dimension is performed through the update gate and reset gate of the gated loop unit, and the updated state vector of each device node at the current time step is output to complete the node state update across time steps.

7. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 1, characterized in that, Performing the operation of truncating input causal edges in the Bayesian causal network model to simulate physical intervention, and calculating the posterior probability of cascade failure of downstream devices includes: Obtain the directed acyclic graph structure in the Bayesian causal network model, and determine the first node corresponding to the candidate intervention device and the second node corresponding to the downstream key device; In the directed acyclic graph structure, all input causal edges pointing to the first node are cut off, so that the value of the first node no longer depends on its parent node in the probabilistic graphical model, but is forcibly assigned the set hypothetical intervention value. Based on the directed acyclic graph structure after cutting off the input causal edges, the cascade failure posterior probability is calculated when the second node takes the value of a preset failure state threshold. The cascade failure posterior probability is used to quantitatively characterize the probability change of cascade failure in downstream key equipment after simulated physical intervention.

8. The method for monitoring cascading faults in industrial equipment based on digital twin model inference according to claim 1, characterized in that, The control strategy is generated with the goal of minimizing control costs and failure probability, and control commands are issued to the corresponding physical actuators to prevent fault propagation, including: Establish a multi-objective optimization function, with minimizing the overall control cost and minimizing the probability of line or equipment failure as the constraints; The physical space for finding candidate control devices is narrowed down using an adaptive shortest path algorithm; Within the physical space, the multi-objective optimization function is solved to obtain the optimal solution located at the Pareto front; The optimal solution is converted into specific control commands and sent to the corresponding physical actuators to prevent the spread of the fault.

9. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 8, characterized in that, The adaptive shortest path algorithm is used to narrow down the physical space for finding candidate control devices, including: Construct a dynamic weighted physical topology graph with industrial equipment as nodes and physical connections as edges. The weight of each edge is a dynamic weight function that integrates fault propagation sensitivity, physical distance, real-time operating condition factors, and controllability cost. Starting with the detected abnormal source device, the initial search boundary is defined according to the fault propagation probability heatmap output by the spatiotemporal hypergraph structure. A forward search is performed in the dynamic weighted physical topology graph, and the search path is pruned in real time based on the cumulative dynamic weight threshold, the fault propagation probability threshold, and the physical controllability conditions of the device. A reverse search is performed simultaneously from the key downstream equipment, and the intersection area of ​​the forward search and reverse search paths is determined as the narrowed physical space of the candidate regulating equipment.

10. The industrial equipment cascading fault monitoring method based on digital twin model inference according to claim 9, characterized in that, The dynamic weighting function introduces a timeliness adjustment factor, which is monotonically increasing with the cumulative time after the fault occurs, and is used to simulate the gradual narrowing effect of the intervention window during the fault propagation process. During the forward search process, the time difference between the current search time and the initial time of the fault is used as input to dynamically calculate the cumulative dynamic weight threshold, so that the search space gradually shrinks as the intervention window period narrows. When the time difference exceeds the preset intervention failure time, the cumulative dynamic weight threshold is set to infinity, the search is terminated and the devices on the fault propagation path are no longer included in the candidate adjustment device physical space.