Fault causal diagram constrained digital twin counterfactual fault diagnosis method and system

By constructing a digital twin and fault cause-effect graph of the electromechanical system, generating virtual fault samples and performing consistency screening of propagation chains, the modeling difficulties and cross-condition adaptability problems of fault diagnosis in complex electromechanical systems are solved, and high-accuracy and interpretable fault diagnosis is achieved.

CN122432848APending Publication Date: 2026-07-21HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
Filing Date
2026-06-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing fault diagnosis methods for complex electromechanical systems suffer from problems such as difficulty in modeling, high cost of parameter identification, insufficient adaptability across operating conditions, lack of closed-loop feedback between diagnostic results and digital twin simulation results, and insufficient output capability for unknown faults, making it difficult to achieve accurate and stable fault diagnosis under complex operating conditions.

Method used

By constructing a digital twin of the electromechanical system and a fault cause-effect graph, virtual fault samples are generated. Combined with representation decomposition and propagation chain consistency score screening, a diagnostic model is trained to achieve fault category identification, fault source node location, and fault propagation chain diagnosis. The propagation chain deviation is used to update parameters, forming a data closed-loop coupling.

Benefits of technology

It improves the generalization ability, propagation chain interpretation ability, and online adaptive ability of fault diagnosis under complex working conditions, and enhances the accuracy and interpretability of diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432848A_ABST
    Figure CN122432848A_ABST
Patent Text Reader

Abstract

The application discloses a kind of fault causal diagram constraint digital twin counterfactual fault diagnosis method and system, belong to fault diagnosis and intelligent operation and maintenance technical field.The method constructs electromechanical system digital twin and fault causal diagram, according to the working condition corresponding to real monitoring sample and candidate fault source hypothesis generation virtual fault sample;Mechanism invariant factor, working condition domain factor and fault discrimination factor are obtained by representation decomposition, only change working condition domain factor generation and screening counterfactual sample, combine spread reasoning and output fault category, fault source node, fault propagation chain or unknown fault mark, and according to the deviation of deduction propagation chain and diagnostic propagation chain, jointly update digital twin, diagnostic model and graph parameters.The application improves the generalization ability of fault diagnosis under complex working condition, propagation chain explainability and long-term online adaptive ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault diagnosis and intelligent operation and maintenance technology, and particularly to a fault diagnosis method and system based on a fault causal graph-constrained digital twin counterfactual fault diagram. More specifically, it relates to a fault diagnosis method and system that, by constructing a digital twin of an electromechanical system and a fault causal graph, and combining representation decomposition, constrained counterfactual sample generation, propagation chain consistency score screening, and propagation chain deviation-driven collaborative updating, achieves fault category identification, fault source node location, fault propagation chain diagnosis, and unknown fault labeling under complex operating conditions. Background Technology

[0002] Electromechanical systems (EMS) are widely used in industrial equipment, intelligent manufacturing, rail transportation, energy and power, robotics, servo drives, and automated production lines. These systems typically feature complex structural coupling, a wide range of operating conditions, long fault propagation chains, significant aliasing of monitoring signals, and the interplay between fault mechanisms and operating disturbances. During long-term operation, EMS are susceptible to load fluctuations, temperature changes, environmental disturbances, component aging, parameter drift, and control deviations, leading to anomalies in state variables such as vibration, current, temperature rise, pressure, displacement, or response time. Therefore, accurate, stable, and interpretable fault diagnosis is crucial for EMS.

[0003] Existing fault diagnosis methods mainly include mechanistic model-based methods and data-driven methods. Mechanism-based methods rely on relatively accurate dynamic models, thermal models, control models, or energy transfer models. However, when the system structure is complex, the coupling relationship is strong, or the operating conditions are variable, they suffer from difficulties in modeling, high parameter identification costs, and insufficient adaptability across operating conditions. Although data-driven methods can achieve fault identification through signal processing, feature learning, and deep networks, there are often significant data distribution shifts between different operating conditions. This can easily lead to misjudging differences in operating conditions as differences in faults, resulting in decreased diagnostic performance under unseen or heterogeneous operating conditions.

[0004] To alleviate the problems of insufficient samples and operational condition transfer, existing technologies have begun to incorporate methods such as digital twins, transfer learning, graph structure learning, and counterfactual augmentation. For example, virtual fault samples are generated through digital twin models to supplement the lack of real fault samples; fault representation is enhanced through graph structure relationships; and the training space is expanded through counterfactual samples. However, existing technologies still have the following shortcomings: First, the virtual samples generated by digital twins and subsequent diagnostic models often only reach the level of sample supplementation, lacking integrated coupling and utilization around the propagation chain; Second, the correlation constraints between fault propagation structure, fault source location, and propagation timing are not tight enough, making it difficult to simultaneously support representation learning, propagation inference, and reverse causation under the same graph structure; Third, the counterfactual sample generation process usually lacks joint constraints on fault semantic preservation and the rationality of the propagation chain, easily introducing training samples inconsistent with the real fault mechanism; Fourth, there is a lack of closed-loop feedback between diagnostic results and digital twin inference results, making it difficult for digital twin parameters, diagnostic model parameters, and fault causal graph parameters to be continuously and collaboratively updated with system operation; Fifth, the joint output capability for unknown faults, fault source nodes, and propagation chains is still insufficient, making it difficult to balance diagnostic accuracy, interpretability, and long-term online adaptability.

[0005] Therefore, it is necessary to provide a fault diagnosis method and system based on a fault cause-effect graph-constrained digital twin counterfactual model to achieve closed-loop coupling of data among real monitoring samples, virtual fault samples, counterfactual samples, fault cause-effect graphs, and digital twins, thereby improving the generalization ability, propagation chain interpretation ability, and online adaptive ability of fault diagnosis under complex working conditions. Summary of the Invention

[0006] The purpose of this invention is to overcome the aforementioned shortcomings in the existing technology and provide a method and system for counterfactual fault diagnosis using a fault causal graph-constrained digital twin. By constructing a digital twin of the electromechanical system and a fault causal graph, virtual fault samples are generated based on the operating conditions corresponding to the real monitoring samples and the hypotheses of candidate fault sources. The real monitoring samples and virtual fault samples are input into a representation decomposition network to obtain mechanism-invariant factors, operating condition factors, and fault discrimination factors. Based on the same fault causal graph, propagation reasoning and source fault tracing are performed on the fault discrimination factors. While keeping the mechanism-invariant factors and fault discrimination factors unchanged, only the operating condition factors are replaced or interpolated to generate counterfactual samples, which are then screened based on the propagation chain consistency score. On this basis, a diagnostic model is trained, and the parameters of the digital twin, diagnostic model, and fault causal graph are jointly updated according to the deviation between the diagnostic propagation chain and the propagation chain inferred from the digital twin. This enables fault category identification, fault source node localization, fault propagation chain diagnosis, and unknown fault labeling under complex operating conditions.

[0007] To achieve the above objectives, this invention provides a counterfactual fault diagnosis method using a fault causal graph-constrained digital twin, comprising: constructing a digital twin of an electromechanical system and a fault causal graph; generating a virtual fault sample set based on the operating conditions corresponding to real monitoring samples and candidate fault source hypotheses, wherein the virtual fault sample set includes virtual fault samples with fault labels and operating condition labels; inputting the real monitoring samples and the virtual fault samples in the virtual fault sample set into a representation decomposition network to obtain mechanism-invariant factors, operating condition domain factors, and fault discrimination factors; and performing propagation inference and source fault tracing on the same fault causal graph based on the fault discrimination factors to obtain the fault discrimination factors of each node, the forward propagation result, and the posterior probability of the source fault; and only replacing or interpolating the operating condition domain factors from samples with different operating conditions, while keeping the mechanism-invariant factors and fault discrimination factors of the replaced samples unchanged, to generate counterfactual samples. This method involves feeding the counterfactual samples back into the representation decomposition network, and filtering the counterfactual samples based on the deviation between the back-in fault discrimination factor and the original fault discrimination factor, as well as the consistency score of the propagation chain of the counterfactual samples relative to the fault causal graph. A diagnostic model is trained based on real monitoring samples, virtual fault samples in the virtual fault sample set, and the filtered counterfactual samples. The sample to be diagnosed is input into the diagnostic model, which outputs the fault category, fault source node, fault propagation chain, or unknown fault marker. The operating condition domain factor of the sample to be diagnosed, along with the fault category and fault source node, is input into the digital twin to obtain the inferred propagation chain. The model parameters of the digital twin, the model parameters of the diagnostic model, and the propagation weight parameters and propagation delay parameters in the fault causal graph are jointly updated based on the deviation between the inferred propagation chain and the fault propagation chain.

[0008] To achieve the above objectives, this invention provides a fault causal graph-constrained digital twin counterfactual fault diagnosis system, comprising: a digital twin and sample generation module, used to construct a digital twin of an electromechanical system and generate a virtual fault sample set based on the operating conditions corresponding to real monitoring samples and candidate fault source hypotheses, wherein the virtual fault sample set includes virtual fault samples with fault labels and operating condition labels; a fault causal graph and representation decomposition module, used to construct a fault causal graph, mapping real monitoring samples and virtual fault samples in the virtual fault sample set to mechanism-invariant factors, operating condition factors, and fault discrimination factors, and performing propagation inference and source fault tracing on the same fault causal graph based on the fault discrimination factors to obtain the fault discrimination factors, forward propagation results, and posterior probability of source faults for each node; and a constrained counterfactual generation and filtering module, used to replace or interpolate only those from non- The system generates counterfactual samples by keeping the operating condition domain factors of the same operating condition samples unchanged, while maintaining the mechanism invariant factors and fault discrimination factors of the replaced samples unchanged. The counterfactual samples are then selected based on the deviation between the fault discrimination factors obtained from the feedback and the original fault discrimination factors, as well as the consistency of the propagation chain corresponding to the forward propagation results. The diagnosis and collaborative update module is used to train a diagnostic model based on real monitoring samples, virtual fault samples in the virtual fault sample set, and the selected counterfactual samples. It outputs fault categories, fault source nodes, fault propagation chains, or unknown fault markers. The operating condition domain factors of the sample to be diagnosed, as well as the fault categories and fault source nodes, are input into the digital twin to obtain the inferred propagation chain. Based on the deviation between the inferred propagation chain and the fault propagation chain, the model parameters of the digital twin, the model parameters of the diagnostic model, and the propagation weight parameters and propagation delay parameters in the fault causal graph are jointly updated.

[0009] Compared with the prior art, the present invention has at least the following beneficial effects:

[0010] 1. This invention constructs a digital twin and generates virtual fault samples based on the operating conditions and candidate fault source hypotheses corresponding to the real monitoring samples. This enables the virtual samples and real samples to form a unified data entry point in the subsequent representation decomposition, propagation inference and model training processes, thereby improving the operating condition coverage and fault coverage of the training samples.

[0011] 2. Based on the fault discrimination factor, the present invention performs propagation reasoning and source fault tracing on the same fault causal graph to obtain the fault discrimination factor, forward propagation result and source fault posterior probability of each node. This enables the same graph structure to simultaneously support fault propagation modeling, source fault location and propagation chain reasoning, thereby improving the structural consistency of fault diagnosis, the propagation chain interpretation capability and the source tracing capability.

[0012] 3. This invention generates counterfactual samples by replacing or interpolating only the operating condition domain factors while keeping the mechanism-invariant factors and fault discrimination factors unchanged. This allows for explicit separation of operating condition changes and fault semantics in the representation space. The counterfactual samples are then screened by the fault discrimination factor bias and propagation chain consistency scores obtained from the feedback, thereby reducing the interference of invalid counterfactual samples on model training and improving the generalization diagnostic capability under heterogeneous operating conditions.

[0013] 4. This invention obtains a propagation chain by inputting the operating condition domain factors of the sample to be diagnosed, as well as the fault category and fault source node obtained from the diagnosis, into a digital twin. Based on the deviation between the propagation chain and the diagnostic propagation chain, the parameters of the digital twin, the diagnostic model, and the fault cause-effect graph are jointly updated, so that real monitoring, propagation reasoning, and virtual simulation form a closed-loop coupling, thereby improving the system's adaptive ability to changes in operating conditions, parameter drift, and system aging in long-term online operation scenarios.

[0014] 5. This invention can not only output fault categories, but also fault source nodes, fault propagation chains, or unknown fault markers, thereby improving diagnostic accuracy while enhancing the interpretability and practical engineering application value of diagnostic results.

[0015] 6. This invention organically couples fault cause-effect graphs, digital twins, representation decomposition, and restricted counterfactual generation, avoiding the problem of existing technologies where modules are separated and simply spliced ​​together. This makes it more conducive to achieving fault diagnosis with robustness, interpretability, and sustainable update capabilities in complex electromechanical systems. Attached Figure Description

[0016] Figure 1 This is a flowchart of the fault cause-effect graph constraint digital twin counterfactual fault diagnosis method provided by the present invention.

[0017] Figure 2 This is a block diagram of the fault cause-effect graph constrained digital twin counterfactual fault diagnosis system provided by the present invention. Detailed Implementation

[0018] The present invention will be further described in detail below with reference to preferred embodiments. It should be understood that the following specific embodiments are only for illustrating the present invention and are not intended to limit the scope of protection of the present invention. The fault cause-effect graph constrained digital twin counterfactual fault diagnosis method in this embodiment is applicable to scenarios such as motor drive systems, servo actuators, robot electromechanical systems, pump systems, compressor systems, transmission systems, and rail transit electromechanical equipment that require fault category identification, fault source node location, fault propagation chain diagnosis, and unknown fault marking under complex working conditions.

[0019] In this specific embodiment, the same letters represent the same meaning, and different letters represent different meanings; the same terms represent the same technical features, and different terms represent different technical features. The superscript " "" indicates the actual number of monitored samples, indicated by the superscript " "" indicates the quantity corresponding to the virtual fault sample, indicated by the superscript " "" indicates the quantity corresponding to the counterfactual sample, indicated by the superscript " "" indicates the quantity corresponding to the mechanism-invariant factor, indicated by the superscript " "" indicates the quantity corresponding to the operating condition domain factor, indicated by the superscript " "" indicates the quantity corresponding to the fault discrimination factor, indicated by the superscript " " indicates the quantity corresponding to the reconstructed sample, indicated by the superscript " "Indicates the updated value; subscript " " represents the sample number, subscript " "and" "Indicates the node sequence number, subscript" "Indicates the candidate fault source node number, subscript " "" indicates the corresponding quantity in the teacher diagnostic model, with the subscript " "Indicates the corresponding quantity of the student diagnostic model, symbol" Used to represent general input samples. To avoid ambiguity, parameters with the same letter but different subscripts are explained separately.

[0020] Combination Figure 1 The method of this embodiment includes steps S1 to S7.

[0021] S1. Digital twin construction, virtual fault sample generation, and fault cause-effect graph construction.

[0022] This step specifically includes: acquiring the structural topology parameters, energy transfer parameters, control response parameters, operating condition parameters, and multi-source monitoring data of the electromechanical system, and constructing a digital twin; then, based on the operating conditions corresponding to the real monitoring samples and the hypothesis of candidate fault sources, generating virtual fault samples with fault labels and operating condition labels from the digital twin; and constructing a fault cause-effect graph with propagation weight parameters and propagation delay parameters.

[0023] 1. Definition of Real Monitoring Samples and Operating Conditions

[0024] The actual monitoring sample set can be represented as: , In the formula, For the actual monitoring sample set, For the first One real monitoring sample, For the first Fault label corresponding to each real monitoring sample For the first The operating condition label corresponding to each real monitoring sample This represents the total number of actual monitored samples.

[0025] In one implementation, the real monitoring sample It consists of one or more channels of signal selected from vibration channel, current channel, temperature channel, pressure channel, displacement channel, speed channel, or control feedback channel; the operating condition label Used to characterize one or more operating conditions, such as speed, load, temperature, environmental disturbances, or task mode.

[0026] 2. Construction of Digital Twins

[0027] The digital twin is preferably constructed using a combination of mechanistic modeling and data calibration. Specifically, an initial digital twin is first established based on the equipment structure diagram, control logic, energy flow relationships, and fault evolution mechanisms; then, the model parameters of the digital twin are calibrated using real monitoring samples, so that the inferred samples generated by the digital twin gradually approximate the real monitoring samples.

[0028] In this embodiment, a digital twin can be represented as: , In the formula, For the first A virtual fault sample, For generating digital twins, The set of model parameters for a digital twin. For the first Operating condition labels corresponding to each virtual fault sample For the first The fault label corresponding to each virtual fault sample For the first The candidate fault source node label corresponding to each virtual fault sample.

[0029] In a preferred embodiment, the digital twin includes at least a component-level model, a subsystem-level model, and a system-level model; wherein, the component-level model is used to describe the local state evolution of bearings, gears, motor windings, sensors, valve bodies, or power devices, the subsystem-level model is used to describe the coupling behavior of transmission chains, control loops, execution units, or energy transfer units, and the system-level model is used to describe the overall response of the whole machine under different operating conditions and different fault source assumptions.

[0030] 3. Virtual Fault Sample Generation

[0031] The virtual fault sample set can be represented as: , In the formula, For a set of virtual fault samples, For the first A virtual fault sample, For the first The fault label corresponding to each virtual fault sample For the first Operating condition labels corresponding to each virtual fault sample For the first The candidate fault source node labels corresponding to each virtual fault sample For the first The propagation chain label corresponding to each virtual fault sample This represents the total number of virtual fault samples.

[0032] Among them, the propagation chain tag Digital twins can be used to label a given working condition. Fault Labels and candidate fault source node labels The time-series simulation results are generated to characterize the edge sequence, propagation direction, and propagation order of the fault from the source node to the other nodes.

[0033] 4. Construction of Fault Cause-Effect Graph

[0034] A cause-effect graph can be represented as: , In the formula, This is a cause-and-effect graph of the fault. For a set of nodes, Let be a set of directed edges. To propagate the set of weight parameters, For the set of propagation delay parameters, For nodes To the node The propagation weight parameter, For nodes To the node The propagation delay parameter.

[0035] The set of nodes It includes at least component nodes, state variable nodes, monitoring point nodes, and fault event nodes; the directed edge set It includes at least structural influence edges, energy propagation edges, control effect edges, and fault-inducing edges. For any edge consisting of nodes... Pointing to node The edges whose properties are determined by at least the propagation weight parameters. and propagation delay parameters Sure.

[0036] 5. Training methods for digital twins

[0037] Regarding model training methods, the preferred approach for digital twins is "mechanism initialization + data calibration." Specifically, the initial parameters of the digital twin are first determined based on the device's mechanism model. Then, iterative calibration is performed using the differences between real monitoring samples and digital twin-derived samples. One of the following methods can be used: stochastic gradient descent, momentum gradient descent, Adam (adaptive moment estimation), or AdamW (adaptive moment estimation with weight decay). Update the training process; during training, learning rate decay, early stopping mechanisms, and gradient clipping can be combined to improve training stability.

[0038] In this step, the actual monitoring sample set Provides realistic working conditions; the digital twin is labeled with working condition tags. Fault Labels and candidate fault source node labels Generate virtual fault samples for input And simultaneously output the propagation chain tag. Cause-effect graph This provides a unified graph structure foundation for subsequent fault discrimination factor calculation, forward propagation reasoning, backward causal reasoning, counterfactual screening, and graph parameter updating. The purpose of this step is to establish a unified starting point for "real monitoring—virtual simulation—propagation structure." Its beneficial effects are twofold: firstly, the digital twin improves the coverage of fault and operational condition samples; secondly, the fault causal graph explicitly introduces propagation paths, propagation strengths, and propagation sequences, providing support for consistency constraints and closed-loop updates in subsequent steps around the propagation chain.

[0039] S2. Characterization decomposition network construction and three-factor extraction

[0040] This step specifically includes: inputting real monitoring samples and virtual fault samples into the representation decomposition network to obtain mechanism-invariant factors, operating condition domain factors, and fault discrimination factors; among which, the mechanism-invariant factors are used to characterize the mechanism response that exists stably across operating conditions, the operating condition domain factors are used to characterize the domain shift caused by changes in operating conditions, and the fault discrimination factors are used to characterize the fault category, fault source node, and related information on propagation status.

[0041] 1. Definition of Representation Decomposition Network

[0042] In this embodiment, the characterization decomposition network includes a mechanistic encoder, a condition-domain encoder, a fault encoder, and a decoder. For any sample... The three-factor extraction and reconstruction process satisfies: , In the formula, As the mechanism invariant factor, Mechanistic encoder; For operating condition domain factors, For operating condition domain encoders; For fault discrimination factors, The encoder is faulty; To reconstruct the sample, For decoders.

[0043] 2. Reconstructing Constraints

[0044] To ensure the invariance of the mechanism factor Operating condition factors and fault discrimination factor For the original sample The joint representation capability is introduced to construct a reconstruction loss: , In the formula, To reconstruct the loss, For the input sample, To reconstruct the sample, It is a 2-norm.

[0045] 3. Correspondence between the three factors and nodes in the fault cause-effect graph

[0046] In one implementation, fault discrimination factors can be determined based on a preset mapping relationship between monitoring channels and graph nodes. Mapped to the fault discrimination factor of each node, denoted as ;in, Represents a node The corresponding fault discrimination factor, Represents a set of nodes The total number of nodes in the system.

[0047] 4. Training methods for representation decomposition networks

[0048] The representation decomposition network preferably employs a two-stage training approach. In the first stage, a mixture of real monitoring samples and virtual fault samples is used as input to pre-train the mechanistic encoder, operating condition encoder, fault encoder, and decoder, enabling the three-factor decomposition and reconstruction process to initially converge. In the second stage, graph constraints, counterfactual samples, and diagnostic constraints from subsequent steps are introduced to jointly fine-tune the pre-trained parameters. During training, a mini-batch training method is preferred, and the Adam (adaptive moment estimation) optimizer is used to update the network parameters.

[0049] In this step, real monitoring samples and virtual fault samples are used as unified inputs, and are mapped to mechanism-invariant factors respectively through a characterization decomposition network. Operating condition factors and fault discrimination factor Among them, the fault discrimination factor It is then fed into the subsequent fault cause-effect graph constraint module, and the operating condition factor. It is then fed into the subsequent restricted counterfactual generation module, with the mechanism invariant factor. Fault discrimination factor These factors serve as fixed inputs for generating counterfactual samples. The purpose of this step is to decouple the three factors for subsequent propagation reasoning and counterfactual construction. Its beneficial effects are: by explicitly distinguishing between mechanistic information, operating condition information, and fault information, it provides a direct data foundation for subsequent restricted counterfactual construction that "only replaces or interpolates operating condition domain factors," and reduces the interference of operating condition changes on fault identification.

[0050] S3. Fault discrimination factor calculation, forward propagation constraint, backward causation constraint, and graph Laplace regularization based on the same fault cause-effect graph.

[0051] This step specifically includes: based on the fault discrimination factor obtained in step S2. In the same fault cause-effect graph The fault discrimination factor, forward propagation result, and source fault posterior probability are calculated and then graph Laplace regularization, forward propagation constraint, and reverse attribution constraint are applied to them.

[0052] 1. Graph Laplace Regularization

[0053] To ensure that the fault discrimination factors of adjacent nodes in the fault cause-effect graph maintain local smoothness in the graph structure, a graph Laplacian regularization loss is introduced: , In the formula, For the graph Laplacian regularization loss, The feature matrix is ​​composed of each fault discrimination factor. For matrix The transpose of the matrix, For the graph Laplace matrix, For degree matrix, It is an adjacency matrix. For matrix trace operation, where the adjacency matrix The preferred choice is the set of directed edges in the fault cause-effect graph. and the set of propagation weight parameters The degree matrix is ​​determined jointly. For adjacency matrix The corresponding angle matrix.

[0054] 2. Forward Propagation Constraints

[0055] To ensure that the fault discrimination factors remain consistent along the propagation direction of the fault cause-effect graph, a forward propagation constraint is introduced: , In the formula, Forward propagation constraint loss, Set of directed edges One of the edges in, For the edge activation coefficient, For nodes Fault discrimination factor For nodes Fault discrimination factor Forward propagation mapping function, For nodes To the node The propagation weight parameter, For nodes To the node The propagation delay parameter.

[0056] In one implementation, the forward propagation mapping function It can be implemented using graph convolution propagation units, graph attention propagation units, or differentiable state propagation units.

[0057] 3. Backward abduction constraint

[0058] To infer the fault source node based on anomaly observations, a reverse cause-finding constraint is introduced: , In the formula, To constrain losses through reverse attribution, For the set of candidate fault source nodes, For nodes The posterior probability of the source fault. For anomaly observation set and cause-effect graph The obtained nodes The source fault estimation results, This is a set of anomalous observations.

[0059] In one implementation, the set of abnormal observations Composed of abnormal monitoring points, abnormal state variables, or abnormal output deviations exceeding a preset threshold in the current sample; inverse cause function. It can be implemented using reverse graph reasoning, Bayesian posterior inference, or graph message backpropagation.

[0060] 4. Consistency score of the propagation chain

[0061] To assess the reasonableness of the propagation chain of counterfactual samples and online output results in subsequent steps, the propagation chain consistency score is defined as: , , In the formula, For the sample The corresponding set of propagation edges, The threshold for edge activation coefficients. For the propagation error threshold, It is a norm 2. For the propagation edge set The number of sides in For the sample relative to cause-effect graph The consistency score of the propagation chain. For the sample Middle node To the node The estimated propagation time difference.

[0062] 5. Training method for graph constraint module

[0063] In this step, the graph constraint module and the representation decomposition network are jointly trained. Specifically, in each training round, the representation decomposition network first outputs the fault discrimination factor, and then the graph constraint module calculates the graph Laplacian regularization loss. Forward propagation constraint loss and reverse abduction constraint loss The aforementioned losses are then propagated back to the corresponding parameters of the mechanistic encoder, the operating condition encoder, the fault encoder, and the graph constraint module.

[0064] In this step, based on the fault discrimination factor In the same fault cause-effect graph By performing propagation reasoning and source fault tracing, fault discrimination factors are obtained. Forward propagation results and posterior probability of source failure And calculate the propagation chain consistency score. The consistency score of this propagation chain will be used in subsequent step S4 for counterfactual sample screening and in step S7 for graph parameter update gating. The purpose of this step is to simultaneously embed the same fault causal graph into the three stages of fault representation smoothing, fault propagation reasoning, and fault source node tracing. Its beneficial effect is that subsequent counterfactual generation, diagnostic output, and online updates are no longer isolated modules, but rather unified around the same propagation chain with consistency constraints, thereby improving structural interpretability and source tracing capabilities.

[0065] S4. Restricted counterfactual sample generation and screening with only replacement or interpolation of operating condition domain factors.

[0066] This step specifically includes: replacing or interpolating only the operating condition domain factors from samples of different operating conditions, while keeping the mechanism invariant factors and fault discrimination factors of the replaced samples unchanged, to generate counterfactual samples; then feeding the counterfactual samples back into the characterization decomposition network, and filtering the counterfactual samples based on the deviation between the back-input fault discrimination factors and the original fault discrimination factors and the consistency score of the propagation chain.

[0067] 1. Restricted Counterfactual Sample Generation

[0068] Interpolating the operating condition domain factors from samples of different operating conditions satisfies: , In the formula, For the replaced or interpolated operating condition domain factor, These are the interpolation coefficients. The operating condition domain factor corresponding to the first operating condition sample. The operating condition domain factor corresponding to the second operating condition sample. As a counterfactual sample, For decoder, The mechanism invariance factor of the replaced sample. The fault discrimination factor of the replaced sample.

[0069] In a preferred embodiment, the two operating condition samples used to generate counterfactual samples have different operating condition labels; more preferably, the two operating condition samples have the same fault label or the same candidate fault source node label, so as to make the restricted counterfactual generation process more stable.

[0070] 2. Counterfactual sample feedback and consistency deviation calculation

[0071] The counterfactual sample Feed back the characterization decomposition network and calculate the counterfactual consistency bias: , In the formula, For counterfactual consistency bias, The encoder is faulty. Counterfactual sample The fault discrimination factor obtained after the feedback is The fault discrimination factor of the original sample. The weighting coefficient for the consistency score item in the propagation chain. Counterfactual sample relative to cause-effect graph The consistency score of the propagation chain.

[0072] 3. Counterfactual Sample Selection Rules

[0073] When the counterfactual consistency deviation satisfies: When, the counterfactual samples are retained for subsequent training; when At that time, the counterfactual samples are removed or their training weights are reduced. In the formula, This is the counterfactual screening threshold.

[0074] 4. Training methods for restricted counterfactual modules

[0075] The constrained counterfactual generation and filtering module in this step is not trained independently, but jointly with the representation decomposition network and the graph constraint module. Specifically, during training, the working condition domain factor is first generated from the current batch of samples. Then perform replacement or interpolation operations to generate And obtain counterfactual samples through the decoder. The counterfactual samples are then fed back into the characterization decomposition network to calculate... and will As one of the constraints in backpropagation.

[0076] 5. The data coupling relationship, function, and beneficial effects of this step.

[0077] In this step, the operating condition domain factor From step S2, mechanism invariant factor and fault discrimination factor Also from step S2, the propagation chain consistency score This comes from step S3. In other words, the generation of counterfactual samples is driven by three factors, while the selection of counterfactual samples depends simultaneously on the fault discrimination factor bias obtained from the feedback and the consistency score of the propagation chain on the same fault causal graph. The purpose of this step is to expand the operating condition space without changing the fault semantics. Its beneficial effects are: it can expand the training sample distribution using operating condition domain factors from different operating condition samples, and it can avoid introducing invalid samples inconsistent with the true fault mechanism through the dual constraints of fault discrimination factor maintenance and propagation chain consistency score, thereby improving the generalization ability under heterogeneous operating conditions.

[0078] S5. Diagnostic model training based on real monitoring samples, virtual fault samples, and filtered counterfactual samples.

[0079] This step specifically includes: training a diagnostic model based on real monitoring samples, virtual fault samples, and screened counterfactual samples; preferably, the diagnostic model includes a teacher diagnostic model and a student diagnostic model, and physical information constraints, unknown fault labeling constraints, and knowledge distillation constraints are introduced during the training process.

[0080] 1. Physical Information Constraints

[0081] To ensure that the intermediate representations and output results of the diagnostic model conform to the underlying mechanisms of the equipment, physical information constraints are introduced: , In the formula, For physical information constraint loss, The number of sampling points is the physical constraint. For the first The system state variables corresponding to each sampling point The first derivative of the system state quantity is... The second derivative of the system state variables is given. For the first The control input quantity corresponding to each sampling point A set of physical mechanism parameters, It is the residual function corresponding to the dynamic equation, energy conservation equation, control response equation, or state constraint equation.

[0082] 2. Unknown fault labeling constraints

[0083] To achieve unknown fault labeling, a rejection score is constructed: , In the formula, To avoid rating, The distance term weighting coefficient is the fault discrimination factor. To reconstruct the weight coefficients of the loss term, The weighting coefficient for the counterfactual consistency deviation term. Fault discrimination factor With the known fault category center set Distance metric between Given the known fault category center set, To reconstruct the loss, This is known as counterfactual consistency bias.

[0084] when When, output an unknown fault flag; when When the known fault result is displayed, In the formula, This is the rejection threshold.

[0085] 3. Teacher Diagnostic Model Training

[0086] In a preferred embodiment, the total loss of the teacher diagnostic model can be expressed as: , In the formula, The total loss of the teacher diagnostic model. These are the weighting coefficients for the classification loss term. To diagnose the classification loss of the teacher diagnostic model, To reconstruct the weight coefficients of the loss term, The weight coefficients of the graph Laplace regularization term are... These are the weighting coefficients for the forward propagation constraint terms. For the weighting coefficients of the reverse attribution constraint term, These are the weighting coefficients for the physical information constraint term. The weight coefficients for the unknown fault labeling constraint term. Based on rejection scoring Constructed unknown fault labeling constraint loss.

[0087] When training the teacher diagnostic model, it is first pre-trained using real monitoring samples and virtual fault samples, and then fine-tuned by introducing screened counterfactual samples, so that the model can simultaneously have the ability to fit real samples, cover virtual samples, and generalize counterfactual samples.

[0088] 4. Knowledge Distillation and Student Diagnostic Model Training

[0089] In this embodiment, the student diagnostic model is obtained through knowledge distillation, and the knowledge distillation loss satisfies:

[0090] , In the formula, For knowledge distillation loss, For the weighting coefficients of the categorical distillation items, The weighting coefficients for the distillation term of the fault discrimination factor. The weighting coefficient for the distillation item in the rejection score is... For distillation temperature parameters, This represents the Kullback-Leibler divergence (relative entropy). For the Softmax function, The pre-classification output values ​​of the teacher diagnostic model. The pre-classification output values ​​of the student diagnostic model. The fault discrimination factor output by the teacher diagnostic model. The fault discrimination factor output by the student diagnostic model. The rejection score output by the teacher diagnostic model. The rejection score output by the student diagnostic model.

[0091] The total loss of the student diagnostic model can be expressed as: , In the formula, Diagnose the total loss of the model for students. Weight coefficients for the classification loss term in the student model. Diagnose the classification loss of the model for students. This represents the weighting coefficient for the knowledge distillation loss term.

[0092] 5. Training methods for diagnostic models

[0093] Regarding model training methods, a phased training strategy is preferred:

[0094] Phase 1: Training the digital twin and the representation decomposition network to stabilize the output of the digital twin and achieve initial convergence of the three-factor decomposition and reconstruction;

[0095] Phase 2: Integrating Graph Laplace regularization, forward propagation constraints, backward attribution constraints, counterfactual sample screening constraints, physical information constraints, and unknown fault labeling constraints into the teacher diagnostic model to train the teacher diagnostic model;

[0096] Phase 3: Fix the teacher's diagnostic model and train the student's diagnostic model through knowledge distillation;

[0097] Phase 4: During online operation, the student diagnostic model is incrementally updated based on step S7.

[0098] In this step, real monitoring samples, virtual fault samples, and filtered counterfactual samples are used together as training inputs for the diagnostic model; fault discrimination factors Reconstruction loss Counterfactual consistency bias Graph constraint loss Forward propagation constraint loss Reverse attribution constraint loss and physical information constraint loss These factors collectively constitute the training constraints of the diagnostic model. The purpose of this step is to integrate the intermediate results generated in the preceding steps into the diagnostic model training. Its beneficial effects are: through the synergistic effect of multi-source training samples and multiple constraints, it can improve both the ability to identify fault categories and the ability to label unknown faults, as well as the model's deployment adaptability.

[0099] S6, Online Fault Diagnosis Output

[0100] This step specifically includes: inputting the sample to be diagnosed into the diagnostic model, outputting the fault category, fault source node, fault propagation chain or unknown fault marker, and calculating the diagnostic confidence.

[0101] 1. Diagnostic confidence

[0102] In this embodiment, the diagnostic confidence level satisfies: , In the formula, For diagnostic confidence, For the Softmax function, The output values ​​before classification for the student diagnostic model.

[0103] 2. Output of fault type, fault source node, and fault propagation chain.

[0104] When the rejection score is satisfied When the fault is known, the fault category is output as the pre-classification output value from the student diagnostic model. It is determined that the source node of the fault is determined by the set of posterior probabilities of the source fault. The candidate fault source node with the highest probability is determined, and the fault propagation chain is formed by the set of propagation edges. Sure.

[0105] Specifically, the fault source node can be denoted as: , In the formula, The output fault source node, This represents the independent variable that maximizes the objective function.

[0106] Fault propagation chains can be represented by a set of propagation edges. Output the corresponding directed edge sequence.

[0107] When the rejection score is satisfied When this happens, an unknown fault flag is output.

[0108] 3. The data coupling relationship, function, and beneficial effects of this step.

[0109] In this step, the student diagnostic model outputs the pre-classification output values. Fault discrimination factor , set of posterior probabilities of source faults and propagation edge set Together, they constitute the final diagnostic result; among which, diagnostic confidence... Consistency score of propagation chain This will be further used in step S7 for gating updates. The purpose of this step is to transform the trained diagnostic capabilities into online output capabilities. Its beneficial effects are: it can not only output fault categories, but also fault source nodes and fault propagation chains, and output unknown fault markers when necessary, thereby improving the interpretability and engineering applicability of the system.

[0110] S7. Joint update of digital twin, diagnostic model and fault cause-effect graph based on the deviation between the inference propagation chain and the fault propagation chain.

[0111] This step specifically includes: inputting the operating condition domain factor of the sample to be diagnosed, as well as the fault category and fault source node, into the digital twin to obtain the inference propagation chain, and jointly updating the model parameters of the digital twin, the model parameters of the diagnostic model, and the propagation weight parameters and propagation delay parameters in the fault cause-effect graph based on the deviation between the inference propagation chain and the fault propagation chain.

[0112] 1. Generation of inference samples and inference propagation chains

[0113] For the sample to be diagnosed, its operating condition factor Fault type and fault source node Inputting a digital twin generates corresponding inference samples. and the corresponding deductive propagation chain; among which, This indicates the operating condition label corresponding to the sample to be diagnosed. This indicates the fault source node label corresponding to the sample to be diagnosed.

[0114] 2. Digital twin parameter updates and diagnostic model parameter updates

[0115] Joint update satisfies: , , In the formula, For the updated set of parameters of the digital twin model, The parameter set of the digital twin model before the update. Update the step size for parameters of the digital twin model. For parameters Operator for finding gradient The actual monitoring sample corresponding to the sample to be diagnosed. For digital twins in working condition labels and fault source node labels The generated inference samples, For the updated diagnostic model parameter set, This is the set of parameters for the diagnostic model before the update. To update the step size of the diagnostic model parameters, For parameters Operator for finding gradient This represents the total loss of the current diagnostic model.

[0116] In this embodiment, when the update object is the teacher diagnostic model, Desirable When the updated object is a student diagnostic model, Desirable .

[0117] 3. Loss from Map Parameter Correction and Map Parameter Update

[0118] The corrections to the propagation weight parameters and propagation delay parameters in the fault cause-effect graph satisfy the following: , In the formula, To correct the loss for the graph parameters, For a high-confidence propagation edge set, The nodes are estimated from high-confidence samples. To the node The propagation weight parameter, For the nodes in the current cause-effect graph of the fault To the node The propagation weight parameter, The nodes are estimated from high-confidence samples. To the node The propagation delay parameter, For the nodes in the current cause-effect graph of the fault To the node The propagation delay parameter, This is the weighting coefficient for the time delay term.

[0119] Correspondingly, the updates to the propagation weight parameters and propagation delay parameters satisfy: ,, In the formula, For the updated node To the node The propagation weight parameter, For the updated node To the node The propagation delay parameter, To propagate the step size of the weight parameter update, To update the propagation delay parameter step size, Correcting the loss for graph parameters to the propagation weight parameters The partial derivatives, Correcting the loss for the graph parameters on the propagation delay parameter The partial derivatives of .

[0120] 4. Gating conditions for updating graph parameters

[0121] Furthermore, updates to the propagation weight parameters and propagation delay parameters are performed only if the following conditions are met: , In the formula, For diagnostic confidence, For diagnostic confidence threshold, For real monitoring samples relative to cause-effect graph The consistency score of the propagation chain. This is the threshold for the consistency score of the propagation chain.

[0122] 5. Joint update training method

[0123] The joint update in this step is preferably implemented using an online incremental approach. Specifically, when a newly arrived sample to be diagnosed meets the preset update conditions, a deduction propagation chain is first generated using its operating condition domain factor, fault category, and fault source node; then, the deduction propagation chain is compared with the fault propagation chain, and the parameters of the digital twin model are adjusted accordingly. Diagnostic model parameters and graph parameters Perform one or more rounds of small-step updates. To prevent online drift, sliding window, exponentially weighted average, or periodic freeze strategies can be used to control the update frequency and magnitude.

[0124] In this step, the operating condition domain factor From step S2, the fault category and fault source node are from step S6, and the inference sample is derived. The propagation chain originates from the digital twin, the fault propagation chain originates from step S6, and the consistency score of the propagation chain is... From step S3. The above quantities collectively determine the update direction and magnitude of the digital twin parameters, diagnostic model parameters, and fault cause-effect graph parameters. The purpose of this step is to establish a closed-loop feedback path between the real diagnostic results and the virtual simulation results. Its beneficial effects are: it enables the digital twin to continuously approximate the real system, allows the diagnostic model to continuously adapt to operating condition drift and equipment aging, and gradually brings the propagation weight parameters and propagation delay parameters in the fault cause-effect graph closer to the real propagation law, thereby improving the stability and adaptability under long-term online operating conditions.

[0125] Corresponding to the methods described above, such as Figure 2 As shown, the present invention also provides a fault diagnosis system based on a fault causal graph-constrained digital twin counterfactual model. The system includes a digital twin and sample generation module, a fault causal graph and representation decomposition module, a constrained counterfactual generation and screening module, and a diagnosis and collaborative update module.

[0126] The digital twin and sample generation module is used to execute step S1, which involves constructing a digital twin of the electromechanical system and generating virtual fault samples with fault labels and operating condition labels based on the operating conditions corresponding to the real monitoring samples and the hypothesis of candidate fault sources. The fault cause-effect graph and representation decomposition module is used to execute steps S2 and S3, which involves constructing a fault cause-effect graph, mapping the real monitoring samples and virtual fault samples to mechanism-invariant factors, operating condition domain factors, and fault discrimination factors, and calculating the fault discrimination factors, forward propagation results, and posterior probability of source faults for each node on the same fault cause-effect graph based on the fault discrimination factors. The restricted counterfactual generation and screening module is used to execute step S4, which involves only replacing or interpolating the operating condition domain factors from samples with different operating conditions while maintaining the mechanism-invariant factors and fault discrimination factors of the replaced samples. The fault discrimination factor remains unchanged, and counterfactual samples are generated. The counterfactual samples are then selected based on the deviation between the fault discrimination factor obtained from the feedback and the original fault discrimination factor, as well as the consistency score of the propagation chain corresponding to the forward propagation result. The diagnosis and collaborative update module is used to execute steps S5 to S7, that is, to train the diagnosis model based on real monitoring samples, virtual fault samples and selected counterfactual samples, output fault category, fault source node, fault propagation chain or unknown fault marker, and input the operating condition domain factor of the sample to be diagnosed, as well as the fault category and fault source node, into the digital twin to obtain the inferred propagation chain. Based on the deviation between the inferred propagation chain and the fault propagation chain, the model parameters of the digital twin, the model parameters of the diagnosis model, and the propagation weight parameters and propagation delay parameters in the fault cause-effect graph are jointly updated.

[0127] In one implementation, the system can be deployed on an industrial field server, an edge computing terminal, an embedded controller, or a cloud-edge collaborative platform; wherein, the teacher diagnostic model can be deployed on the server side or in the cloud, the student diagnostic model can be deployed on the edge or an embedded terminal, and the digital twin and the fault cause-effect graph can be deployed on any one or more of the edge, server side, or cloud.

[0128] This invention, through the coordinated implementation of steps S1 to S7, forms a complete technical chain: "generation of real monitoring samples and virtual fault samples—three-factor characterization decomposition—propagation inference and source fault tracing on the same fault causal graph—restricted counterfactual generation by only replacing or interpolating factors in the operating condition domain—counterfactual sample feedback screening—diagnostic model training and online output—joint updating of the digital twin, diagnostic model, and fault causal graph based on the deviation between the inference propagation chain and the fault propagation chain." This technical chain enables fault category identification, fault source node localization, fault propagation chain diagnosis, and unknown fault labeling under complex operating conditions, and improves the generalization ability, interpretability, and long-term online adaptive capability of the diagnostic system.

[0129] Comparative experiment

[0130] To verify the fault causal graph-constrained digital twin counterfactual fault diagnosis method and system described in this invention, their ability to identify fault categories, locate fault source nodes, diagnose fault propagation chains, mark unknown faults, adapt to long-term online conditions, and deploy in a lightweight manner under complex operating conditions, the inventors built a servo drive electromechanical system test platform for verification.

[0131] The servo drive electromechanical system includes a servo motor, coupling, reduction gear, rolling bearing, load actuation unit, driver, and controller. The acquired multi-source monitoring signals include vibration signals, current signals, temperature signals, speed signals, and control feedback signals. The sampling frequency is set to 12.8 kHz (kilohertz), and each sample contains 2048 sampling points. The acquired multi-source signals are sliced ​​according to a unified time window to form a real monitoring sample. And assign fault labels based on the corresponding fault status and operating conditions. and working condition labels .

[0132] In this embodiment, six types of known faults and two types of unknown faults are defined. The six types of known faults are: bearing outer ring fault, bearing inner ring fault, rolling element fault, gear wear fault, shaft misalignment fault, and winding partial short circuit fault. The two types of unknown faults are: coupling looseness fault and rotor eccentricity fault. Corresponding to the six types of known faults, multiple candidate fault source node labels are further defined. These correspond to different initial fault nodes among bearing nodes, gear nodes, shaft nodes, winding nodes, and coupling nodes, respectively.

[0133] Five operating conditions are set, denoted as operating conditions. Operating conditions Operating conditions Operating conditions and working conditions Among them, operating conditions To operating conditions As a training condition, the working condition and working conditions As unseen test conditions, each condition is determined by a combination of speed, load, and ambient temperature. A real monitoring sample set is constructed based on actual collected samples. And using the aforementioned digital twins to label different working conditions Fault Labels and candidate fault source node labels The following steps are performed: parameter perturbation, fault injection, and timing simulation to generate a virtual fault sample set. Simultaneously output propagation chain labels .

[0134] In one specific implementation, the digital twin includes a component-level model, a subsystem-level model, and a system-level model. The component-level model describes the local state evolution of bearings, gears, motor windings, couplings, and sensors; the subsystem-level model describes the coupling behavior of the transmission chain, actuators, and control loops; and the system-level model describes the overall response of the entire machine under different operating conditions and different fault source nodes. Preferably, an initial digital twin model is first established based on the equipment structure diagram and control logic, and then the parameters of the digital twin are adjusted using the deviation between real monitoring samples and projected samples. Perform calibration.

[0135] In this embodiment, the fault cause-effect graph The set of nodes in Includes component nodes, state variable nodes, monitoring point nodes, and fault event nodes; a set of directed edges. Includes structural influence edges, energy propagation edges, control effect edges, and fault-inducing edges; and a set of propagation weight parameters. and propagation delay parameter set Corresponding to the set of directed edges respectively The propagation strength and timing attributes of each edge in the fault cause-effect graph. The initial structure of the fault cause-effect graph is determined by the equipment topology, energy transfer relationships, and control dependencies, and the initial propagation weight parameters. and propagation delay parameters The results are provided by mechanistic analysis and then updated online using subsequent high-confidence samples.

[0136] In this embodiment, the actual monitoring sample set is used. and virtual fault sample set Simultaneously input to the characterization decomposition network. The characterization decomposition network includes a mechanistic encoder. Operating condition domain encoder Fault encoder and decoder Output the mechanism invariant factor respectively Operating condition factors Fault discrimination factor and reconstructed samples The training process employs a two-stage approach: the first stage utilizes only the reconstruction loss. The representation decomposition network and the basic classifier head are pre-trained using classification loss; in the second stage, graph Laplacian regularization loss is introduced. Forward propagation constraint loss Reverse attribution constraint loss Counterfactual consistency bias Physical information constraint loss and unknown fault labeling constraint loss Joint fine-tuning is performed. The optimization algorithm uses the Adam (Adaptive Moment Estimator) optimizer, with an initial learning rate set to... The batch size is set to 64, the number of training rounds is set to 150, and early stopping is triggered when the validation set metric no longer improves for 15 consecutive rounds.

[0137] In the process of generating constrained counterfactual samples, for two samples with different operating conditions in the same batch, the mechanism invariance factor is preferably selected. and fault discrimination factor Nodes originating from the same fault category or the same candidate fault source node, with the operating condition domain factors denoted as follows: and .according to Generate replacement or interpolation of the operating condition domain factor. ,in, The interpolation coefficient is preferably between 0.2 and 0.8 in this embodiment. Then, according to... Generate counterfactual samples The counterfactual sample Feedback fault encoder And combined with the consistency score of the propagation chain Calculate the counterfactual consistency bias Only retain those that meet the requirements. Counterfactual samples are used in subsequent diagnostic model training.

[0138] During the diagnostic model training phase, teacher and student diagnostic models are prioritized. The teacher diagnostic model employs a deeper network structure to achieve stronger representation capabilities, while the student diagnostic model uses a lightweight network structure to meet edge deployment requirements. The teacher diagnostic model is first trained using real monitoring samples, virtual fault samples, and filtered counterfactual samples, and then trained based on knowledge distillation loss. The pre-classification output values, fault discrimination factors, and rejection scores from the teacher diagnostic model are distilled and fed into the student diagnostic model. After training, the student diagnostic model is deployed at the edge to perform online fault diagnosis.

[0139] During the online operation phase, the sample to be diagnosed is input into the student diagnostic model, which outputs the fault category, fault source node, fault propagation chain, and diagnostic confidence. Alternatively, it can output an unknown fault marker. Then, the operating condition factor corresponding to the sample to be diagnosed will be... The output fault category and fault source node are input into the digital twin to obtain the simulation sample. And the corresponding deduction propagation chain, and adjust the digital twin parameters based on the deviation between the deduction propagation chain and the fault propagation chain. Diagnostic model parameters and fault cause-effect graph parameters Perform a joint update. Preferably, only when the diagnostic confidence level is... And the consistency score of the propagation chain Only when necessary should the graph parameters be updated to reduce the risk of erroneous updates.

[0140] This embodiment is used to verify the data closed-loop coupling relationship between real monitoring samples, virtual fault samples, counterfactual samples, fault causal graphs, and digital twins, to prove that the fault causal graph-constrained digital twin counterfactual fault diagnosis method provided by the present invention can not only improve the fault category identification capability under complex working conditions, but also improve the fault source node location and fault propagation chain diagnosis capability, and enhance the adaptive update capability under long-term online operation conditions.

[0141] Comparative Example

[0142] To verify the actual contribution of each key technical feature of the present invention, the following comparative examples are set up.

[0143] Comparative Example 1

[0144] Comparative Example 1 uses only real monitoring samples and virtual fault samples to train the ordinary fault classification network. It does not construct a fault cause-effect graph, does not perform the representation decomposition of mechanism invariant factors, operating condition factors and fault discrimination factors, does not perform restricted counterfactual sample generation and screening, and does not perform joint updates between digital twins, diagnostic models and fault cause-effect graphs.

[0145] In this comparative example, the virtual fault samples are used only as supplementary training samples and are not coupled with propagation chain constraints, fault source node localization, and online update mechanisms.

[0146] Comparative Example 2

[0147] Comparative Example 2, based on Comparative Example 1, introduces a representation decomposition network to decompose the input sample into mechanism-invariant factors, operating condition factors, and fault discrimination factors. However, it does not construct a fault causal graph, calculate fault discrimination factors, perform forward propagation constraints and backward causal constraints, or generate restricted counterfactual samples.

[0148] In this comparative example, although the explicit separation of mechanism information, operating condition information and fault information was achieved, structural constraints related to the propagation chain have not yet been formed.

[0149] Comparative Example 3

[0150] Comparative Example 3, based on Comparative Example 2, further introduces a fault cause-effect graph and performs graph Laplace regularization, forward propagation constraints, and backward causal constraints based on the fault cause-effect graph. However, it does not perform restricted counterfactual sample generation and screening that only replaces or interpolates operating condition domain factors, nor does it perform joint updates driven by propagation chain bias.

[0151] In this comparative example, the fault causal graph has been involved in fault representation learning and source fault localization, but it lacks a restricted counterfactual enhancement mechanism centered on operating condition factors, making it difficult to further expand the training space under unseen operating conditions.

[0152] Comparative Example 4

[0153] Comparative Example 4 employs the same training process as the method of this invention, including constructing a digital twin, a fault cause-effect graph, a representation decomposition network, generating and screening constrained counterfactual samples, physical information constraints, and unknown fault labeling constraints, and distilling to obtain a student diagnostic model, but without performing joint updates between the digital twin, the diagnostic model, and the fault cause-effect graph.

[0154] In this comparative example, the training phase already has a structure that is quite similar to that of the present invention, but the online operation phase lacks a closed-loop update mechanism driven by the deviation between the inference propagation chain and the fault propagation chain.

[0155] Test metrics

[0156] To comprehensively evaluate the performance of the method of this invention and each comparative example, the following test indicators were set:

[0157] 1. Generalization ability metrics for heterogeneous operating conditions: These are evaluated using precision, macro-average F1 score, and unseen operating condition average recall. Among them, F1 (harmonic mean) is used to comprehensively reflect precision and recall.

[0158] 2. Indicators for unknown fault labeling capability and propagation chain interpretation capability: The unknown fault labeling accuracy, known / unknown fault separation AUROC (area under the receiver operating characteristic curve), fault source node localization accuracy, and propagation chain consistency score are used for evaluation.

[0159] 3. Online adaptive capability metrics: The initial accuracy, the accuracy after 30 days of continuous operation, and the accuracy after adaptive updates are used for evaluation.

[0160] 4. Lightweight deployment capability metrics: These are evaluated using accuracy, number of parameters, model size, and single-sample inference time.

[0161] Experimental Results and Analysis

[0162] 1. Verification of generalization ability under heterogeneous operating conditions

[0163] The method of the present invention and Comparative Examples 1 to 3 are all applied under the following conditions. To operating conditions On-the-job training, in work conditions where no training was participated in and working conditions The above test was conducted, and the experimental results are shown in Table 1.

[0164] Table 1 Comparison of generalization performance under heterogeneous operating conditions

[0165] Comparative Example 1 87.4 86.1 84.8 Comparative Example 2 91.2 90.5 89.3 Comparative Example 3 93.0 92.4 91.6 The method of the invention 96.7 96.0 95.3

[0166] As shown in Table 1, Comparative Example 1, which only uses virtual fault samples as supplementary data and lacks three-factor representation decomposition and propagation chain constraints, exhibits a significant decrease in diagnostic performance under unseen operating conditions. Comparative Example 2, by introducing a representation decomposition network, shows improvements in both accuracy and macro-average F1 score, indicating that explicit separation of mechanism information, operating condition information, and fault information helps reduce the interference of operating condition deviation on fault identification. Comparative Example 3, by further introducing a fault causal graph, continues to improve performance, demonstrating that embedding the propagation structure into the fault representation learning and source fault localization process can improve the structural consistency of fault identification factors.

[0167] Compared to Comparative Example 3, the method of this invention generates restricted counterfactual samples by only replacing or interpolating the operating condition domain factors, and further expands the training space corresponding to unseen operating conditions by screening samples based on the consistency scores of fault discrimination factors and propagation chains. This improves the accuracy from 93.0% to 96.7%, the macro-average F1 score from 92.4% to 96.0%, and the average recall rate for unseen operating conditions from 91.6% to 95.3%. This demonstrates that the method of this invention has stronger generalization fault diagnosis capabilities under complex and heterogeneous operating conditions.

[0168] 2. Verification of unknown fault labeling capabilities and propagation chain interpretation capabilities

[0169] In this experiment, the two types of unknown faults were only included in the test set and not used in training. The unknown fault labeling accuracy, known / unknown fault separation AUROC, fault source node localization accuracy, and propagation chain consistency score were used as evaluation indicators. The experimental results are shown in Table 2.

[0170] Table 2 Comparison of Unknown Fault Marking Capability and Propagation Chain Explanation Capability

[0171] Comparative Example 1 76.8 81.5 72.4 0.63 Comparative Example 2 83.1 87.0 80.6 0.72 Comparative Example 3 88.5 91.6 89.4 0.84 The method of the invention 94.2 96.3 95.8 0.93

[0172] As shown in Table 2, Comparative Example 1, lacking constraints on fault cause-effect graphs, posterior probability inference of fault source nodes, and unknown fault labeling, is prone to misclassifying unknown faults as known faults and has weak output capabilities for fault source nodes and propagation chains. Comparative Example 2 improves feature representation through representation decomposition, thus improving the accuracy of unknown fault labeling and fault source node localization. However, due to the absence of a fault cause-effect graph, the consistency score of the propagation chain remains low.

[0173] In Comparative Example 3, after introducing a fault cause-effect graph, the accuracy of fault source node localization increased from 80.6% to 89.4%, and the propagation chain consistency score increased from 0.72 to 0.84. This demonstrates that using the same fault cause-effect graph for both forward propagation constraints and backward causation constraints significantly enhances the fault propagation chain reasoning ability and source fault localization ability. Building upon this, the method of this invention further utilizes restricted counterfactual samples for screening and enhancement, and combines a rejection scoring mechanism to mark unknown faults. This improves the accuracy of unknown fault marking to 94.2%, the known / unknown separation AUROC to 96.3%, the fault source node localization accuracy to 95.8%, and the propagation chain consistency score to 0.93. This indicates that the present invention not only improves the safe handling capability of unknown faults but also enhances the rationality and interpretability of the propagation chain output.

[0174] 3. Online adaptive capability verification

[0175] In this experiment, online monitoring data from 30 consecutive days of operation were selected to simulate data distribution changes caused by system aging, load drift, and environmental changes. The method of this invention was compared with Comparative Example 4 to verify the effectiveness of the joint update mechanism based on the deviation between the extrapolation propagation chain and the fault propagation chain. The experimental results are shown in Table 3.

[0176] Table 3 Comparison of Online Adaptive Capabilities

[0177] Comparative Example 4 96.1 89.7 91.3 The method of the invention 96.7 94.0 95.4

[0178] As shown in Table 3, without joint updates, the accuracy of Comparative Example 4 decreased from 96.1% to 89.7% after 30 days of continuous operation. This indicates that relying solely on static parameters obtained through offline training is insufficient to adapt to long-term system aging, changes in operating conditions, and parameter drift. The method of this invention generates a propagation chain based on the operating condition domain factors of the sample to be diagnosed, the output fault category, and the fault source node. Then, it jointly updates the digital twin parameters based on the deviation between the propagation chain and the fault propagation chain. Diagnostic model parameters and the propagation weight parameters in the cause-effect graph of the fault. and propagation delay parameters The system maintained an accuracy rate of 94.0% after 30 days of continuous operation, and recovered to 95.4% after adaptive updates, significantly outperforming Comparative Example 4. This demonstrates that the present invention possesses better long-term online operational stability and adaptive capability.

[0179] 4. Validation of lightweight deployment capabilities

[0180] In this experiment, the teacher diagnostic model and the student diagnostic model were compared to verify the lightweight deployment effect brought about by knowledge distillation. The experimental results are shown in Table 4.

[0181] Table 4 Comparison of Lightweight Deployment Capabilities

[0182] Teacher diagnostic model 97.3 12.6 48.7 34.5 Student diagnostic model 96.7 2.1 8.4 7.9

[0183] As shown in Table 4, after knowledge distillation, the accuracy of the student diagnostic model decreased by only 0.6 percentage points compared to the teacher diagnostic model, but the number of parameters decreased from 12.6M to 2.1M, the model size decreased from 48.7MB to 8.4MB, and the single-sample inference time decreased from 34.5ms to 7.9ms. This indicates that while maintaining high fault diagnosis performance, fault source node localization capability, and unknown fault labeling capability, the student diagnostic model is more suitable for deployment in edge devices, embedded terminals, and online monitoring devices.

[0184] Experimental conclusions

[0185] As can be seen from Tables 1 to 4, the method of the present invention has at least the following advantages compared to the comparative examples:

[0186] Firstly, by constructing a digital twin and generating virtual fault samples with fault labels, operating condition labels, and propagation chain labels based on the operating conditions corresponding to the real monitoring samples and the hypothesis of candidate fault sources, a unified data entry point for real monitoring samples and virtual fault samples is achieved, thereby improving the operating condition coverage and fault coverage of training samples.

[0187] Secondly, by decomposing real monitoring samples and virtual fault samples into mechanism-invariant factors, operating condition factors, and fault discrimination factors, and performing propagation reasoning and source fault tracing on the same fault causal graph, the fault discrimination factors, forward propagation results, and source fault posterior probabilities of each node are obtained. This achieves unified constraints on fault propagation structure, fault source node location, and propagation chain output, thereby improving the explanatory power of fault propagation chain and the ability to trace the source.

[0188] Third, by generating counterfactual samples by only replacing or interpolating the operating condition domain factors while keeping the mechanism-invariant factors and fault discrimination factors unchanged, and then combining the fault discrimination factor deviation and propagation chain consistency scores obtained from the feedback for screening, the effective separation of operating condition changes and fault semantics is achieved, thereby improving the fault generalization diagnosis capability under unseen operating conditions and heterogeneous operating conditions.

[0189] Fourth, by inputting the operating condition domain factors of the sample to be diagnosed, as well as the fault category and fault source node obtained from the diagnosis, into the digital twin, a propagation chain is obtained. Based on the deviation between the propagation chain and the fault propagation chain, the parameters of the digital twin, the parameters of the diagnostic model, and the propagation weight parameters and propagation delay parameters in the fault cause-effect graph are jointly updated, a closed-loop coupling between real monitoring, propagation reasoning, and virtual simulation is realized, thereby improving the stability and adaptability of the system during long-term online operation.

[0190] Fifth, by obtaining student diagnostic models through knowledge distillation, the system can significantly reduce the number of parameters, model size and inference latency while maintaining high fault diagnosis performance, thus making it more suitable for edge deployment and engineering applications.

[0191] Therefore, the above embodiments, comparative examples, and experimental results fully demonstrate that the fault cause-effect graph constrained digital twin counterfactual fault diagnosis method and system of the present invention can achieve the technical effects described in the specification and has good engineering application value.

[0192] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, or combinations made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A fault diagnosis method using a fault cause-effect graph-constrained digital twin counterfactual model, characterized in that, include: A digital twin and a fault cause-effect graph of the electromechanical system are constructed, and a virtual fault sample set is generated based on the operating conditions corresponding to the real monitoring samples and the hypothesis of candidate fault sources. The virtual fault sample set includes virtual fault samples with fault labels and operating condition labels. The real monitoring samples and virtual fault samples from the virtual fault sample set are input into the representation decomposition network to obtain mechanism-invariant factors, operating condition factors, and fault discrimination factors. Based on the fault discrimination factors, propagation inference and source fault tracing are performed on the same fault causal graph to obtain the fault discrimination factors, forward propagation results, and posterior probabilities of the source faults at each node. Only the operating condition factors from different operating condition samples are replaced or interpolated, while keeping the mechanism-invariant factors and fault discrimination factors of the replaced samples unchanged, to generate counterfactual samples. The counterfactual samples are fed back into the representation decomposition network, and the results are determined based on the deviation between the fed-back fault discrimination factors and the original fault discrimination factors, and the counterfactual samples relative to the fault. The counterfactual samples are screened based on the consistency score of the propagation chain in the fault causal graph. A diagnostic model is trained based on real monitoring samples, virtual fault samples in the virtual fault sample set, and the screened counterfactual samples. The sample to be diagnosed is input into the diagnostic model, which outputs the fault category, fault source node, fault propagation chain, or unknown fault marker. The operating condition domain factor of the sample to be diagnosed, as well as the fault category and fault source node, are input into the digital twin to obtain the inferred propagation chain. The model parameters of the digital twin, the model parameters of the diagnostic model, and the propagation weight parameters and propagation delay parameters in the fault causal graph are jointly updated based on the deviation between the inferred propagation chain and the fault propagation chain.

2. The fault causal graph-constrained digital twin counterfactual fault diagnosis method according to claim 1, characterized in that, Each virtual fault sample in the virtual fault sample set and the fault cause-effect graph satisfy the following: , In the formula, For a set of virtual fault samples, For the first A virtual fault sample, For the first The fault label corresponding to each virtual fault sample For the first Operating condition labels corresponding to each virtual fault sample For the first The candidate fault source node labels corresponding to each virtual fault sample For the first The propagation chain label corresponding to each virtual fault sample The total number of virtual fault samples. This is a cause-and-effect graph of the fault. For a set of nodes, Let be a set of directed edges. To propagate the set of weight parameters, For the set of propagation delay parameters, For nodes To the node The propagation weight parameter, For nodes To the node The propagation delay parameter; the node set includes at least component nodes, state variable nodes, monitoring point nodes and fault event nodes, and the directed edge set includes at least structural influence edges, energy propagation edges, control action edges and fault induction edges.

3. The fault cause-effect graph constrained digital twin counterfactual fault diagnosis method according to claim 2, characterized in that, The output of the characterization decomposition network and the graph Laplacian regularization satisfy: , In the formula, For the input sample, As the mechanism invariant factor, For operating condition domain factors, For fault discrimination factors, To reconstruct the sample, For mechanism-invariant factor extraction function, This is the factor extraction function for the operating condition domain. This is the fault discrimination factor extraction function. For decoding function, For the graph Laplacian regularization loss, The feature matrix is ​​composed of fault discrimination factors of each node. For the graph Laplace matrix, For degree matrix, It is an adjacency matrix. This is for matrix trace operations.

4. The fault cause-effect graph constrained digital twin counterfactual fault diagnosis method according to claim 2, characterized in that, The forward propagation result and the posterior probability of the source fault are constrained by forward propagation constraints and backward attribution constraints, respectively, wherein the forward propagation constraints and the backward attribution constraints satisfy the following: , In the formula, Forward propagation constraint loss, For the edge activation coefficient, For nodes Fault discrimination factor For nodes Fault discrimination factor Forward propagation mapping function, For nodes To the node The propagation weight parameter, For nodes To the node The propagation delay parameter, To constrain losses through reverse attribution, For the set of candidate fault source nodes, For nodes The posterior probability of the source fault. For the set of anomalous observations, This is a cause-and-effect graph of the fault. For anomaly observation set and cause-effect graph The obtained nodes The source fault estimation results.

5. The fault cause-effect graph constrained digital twin counterfactual fault diagnosis method according to claim 1, characterized in that, The generation and screening of the counterfactual samples satisfy the following: , when The counterfactual sample is retained when... The counterfactual samples are removed or their training weights are reduced, where, For the replaced or interpolated operating condition domain factor, As the first operating condition factor, For the second operating condition domain factor, These are the interpolation coefficients. As a counterfactual sample, For decoding function, As the mechanism invariant factor, For fault discrimination factors, For counterfactual consistency bias, This is the fault discrimination factor extraction function. The weighting coefficient for the consistency score item in the propagation chain. Counterfactual sample relative to cause-effect graph The consistency score of the propagation chain. This is the counterfactual screening threshold.

6. The fault causal graph-constrained digital twin counterfactual fault diagnosis method according to claim 5, characterized in that, The unknown fault marker is determined based on a rejection score, which satisfies the following: , in , when Output unknown fault flag when The known fault result is output in time, where, To avoid rating, , and Here, C represents the weighting coefficients, and C is the known set of fault category centers. Fault discrimination factor With the known fault category center set Distance metric between To reconstruct the loss, For the input sample, To reconstruct the sample, For counterfactual consistency bias, To set a rejection threshold, It is a 2-norm.

7. The fault cause-effect graph constrained digital twin counterfactual fault diagnosis method according to claim 1, characterized in that, Physical information constraints are introduced when training the diagnostic model, satisfying the following: ,, In the formula, For physical information constraint loss, The number of sampling points is the physical constraint. For the first The system state variables corresponding to each sampling point The first derivative of the system state variable is given by [the first derivative of the ... second derivative]. The second derivative of the system state variable is given by [the second derivative of the second derivative of the second derivative of the third derivative of the second derivative of the third derivative of the fourth derivative of the fifth derivative of the sixth derivative of the fifth derivative of the sixth derivative of the fifth derivative of the sixth derivative of the seventh ... For the first The control input quantity corresponding to each sampling point A set of physical mechanism parameters, It is the residual function corresponding to the dynamic equation, energy conservation equation, control response equation, or state constraint equation.

8. The fault causal graph-constrained digital twin counterfactual fault diagnosis method according to claim 1, characterized in that, The diagnostic model includes a teacher diagnostic model and a student diagnostic model. The student diagnostic model is obtained through knowledge distillation, and the knowledge distillation satisfies the following: , In the formula, For knowledge distillation loss, , and These are the weighting coefficients. For distillation temperature parameters, For Kullback-Leibler divergence, For the Softmax function, The pre-classification output values ​​of the teacher diagnostic model. The pre-classification output values ​​of the student diagnostic model. The fault discrimination factor output by the teacher diagnostic model. The fault discrimination factor output by the student diagnostic model. The rejection score output by the teacher diagnostic model. The rejection score output by the student diagnostic model. It is a 2-norm.

9. The fault cause-effect graph constrained digital twin counterfactual fault diagnosis method according to claim 2, characterized in that, The joint update of the digital twin, the diagnostic model, and the fault cause-effect graph satisfies: , Updates to the propagation weight and propagation delay parameters are performed only if the following conditions are met: , In the formula, and These are the parameters of the digital twin model before and after the update, respectively. Update the step size for parameters of the digital twin model. For real monitoring samples, For operating condition labels, For the fault source node label, For digital twins in working condition labels and fault source node labels The generated inference samples, Indicates the parameter The gradient operator; and These are the diagnostic model parameters before and after the update, respectively. To update the step size of the diagnostic model parameters, The total loss of the diagnostic model, Indicates the parameter The gradient operator; To correct the loss for the graph parameters, For a high-confidence propagation edge set, The nodes are estimated from high-confidence samples. To the node The propagation weight parameter, The nodes are estimated from high-confidence samples. To the node The propagation delay parameter, The weighting coefficient for the time delay term. For the updated node To the node The propagation weight parameter, For the updated node To the node The propagation delay parameter, To propagate the step size of the weight parameter update, To update the propagation delay parameter step size, For diagnostic confidence, For diagnostic confidence threshold, Correcting the loss for graph parameters to the propagation weight parameters The partial derivatives, Correcting the loss for the graph parameters on the propagation delay parameter The partial derivatives; This is a cause-and-effect graph of the fault. For real monitoring samples relative to cause-effect graph The consistency score of the propagation chain. This is the threshold for the consistency score of the propagation chain.

10. A fault diagnosis system based on a fault cause-effect graph-constrained digital twin counterfactual model, characterized in that, include: The digital twin and sample generation module is used to construct a digital twin of the electromechanical system and generate a virtual fault sample set based on the operating conditions and candidate fault source hypotheses corresponding to the real monitoring samples. The virtual fault sample set includes virtual fault samples with fault labels and operating condition labels. The fault cause-effect graph and representation decomposition module is used to construct a fault cause-effect graph, mapping the real monitoring samples and the virtual fault samples in the virtual fault sample set to mechanism-invariant factors, operating condition factors, and fault discrimination factors. Based on the fault discrimination factors, propagation reasoning and source fault tracing are performed on the same fault cause-effect graph to obtain the fault discrimination factors, forward propagation results, and posterior probability of the source fault for each node. The restricted counterfact generation and filtering module is used to replace or interpolate only the operating condition domain factors from different operating condition samples while keeping the mechanistic invariant factors and fault discrimination factors of the replaced samples unchanged, generate counterfact samples, and filter the counterfact samples based on the deviation between the fault discrimination factors obtained from the feedback and the original fault discrimination factors and the consistency score of the propagation chain of the counterfact samples relative to the fault causal graph. The diagnosis and collaborative update module is used to train a diagnostic model based on real monitoring samples, virtual fault samples in the virtual fault sample set, and screened counterfactual samples. It outputs fault categories, fault source nodes, fault propagation chains, or unknown fault markers. It inputs the operating condition domain factors of the sample to be diagnosed, as well as the fault categories and fault source nodes, into a digital twin to obtain a deduced propagation chain. Based on the deviation between the deduced propagation chain and the fault propagation chain, it jointly updates the model parameters of the digital twin, the model parameters of the diagnostic model, and the propagation weight parameters and propagation delay parameters in the fault causal graph.