Apparatus fault diagnosis method based on causal residual learning and invariant risk minimization

By employing causal residual learning and invariant risk minimization, the interference of operating condition changes is eliminated, a causal residual sequence is constructed and dynamically weighted and fused, solving the problem of insufficient diagnostic accuracy under varying operating conditions in existing technologies, and achieving stable fault feature extraction and cross-domain generalization.

CN122333119BActive Publication Date: 2026-08-04HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2026-06-05
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing equipment fault diagnosis methods cannot effectively adaptively assess the importance of different time scales and physical channels when faced with complex vibration signals, resulting in insufficient diagnostic accuracy and cross-domain generalization performance under varying operating conditions, and lack of dynamic characterization of instantaneous fault characteristics and evolution trend characteristics.

Method used

We employ a causal residual learning and invariant risk minimization approach to eliminate interference caused by changes in operating conditions through causal structure discovery, construct causal residual sequences, and utilize a multi-timescale, multi-channel attention network for recalibration and dynamic weighted fusion. Combined with invariant risk minimization constraints, we perform model pre-training to ensure stable fault feature extraction under different operating conditions.

Benefits of technology

It achieves stable fault identification under varying operating conditions, improves diagnostic accuracy and cross-domain generalization performance, effectively filters out environmental confounding factors, ensures comprehensive characterization and adaptive fusion of fault features, and overcomes the limitations of static feature fusion in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333119B_ABST
    Figure CN122333119B_ABST
Patent Text Reader

Abstract

This invention relates to the field of fault diagnosis technology and discloses a device fault diagnosis method based on causal residual learning and invariant risk minimization. This method acquires multi-channel vibration signals of the target mechanical equipment during operation. Based on causal structure discovery and backdoor adjustment criteria, the vibration signals are decontaminated to eliminate interference components caused by changes in operating conditions, constructing a causal residual sequence that only characterizes the physical mechanism of the equipment fault. The sequence is input into a multi-timescale, multi-channel attention network to achieve recalibration of the causal residual sequence; hierarchical temporal features at different time resolutions are extracted and dynamically weighted and fused using a time-scale attention mechanism to obtain fault feature representations. These fault feature representations are input into a pre-trained fault classifier to output the final fault diagnosis result for the target mechanical equipment. This invention improves the model's cross-domain generalization performance and diagnostic accuracy under unknown operating conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault diagnosis technology, specifically a device fault diagnosis method based on causal residual learning and invariant risk minimization, as well as a computer terminal and computer-readable storage medium for applying this method. Background Technology

[0002] The safety and reliability of industrial equipment are paramount, making precise equipment health diagnostics essential to prevent unplanned downtime. A common approach to fault diagnosis involves analyzing vibration signals collected from operating equipment to identify fault types. However, with the rapid development of intelligent maintenance technologies, data-driven methods have become the mainstream paradigm for fault diagnosis. Deep learning models achieve high-precision fault classification by automatically extracting hierarchical fault features from condition monitoring data.

[0003] To simultaneously capture the inherent high-frequency transient pulses and low-frequency evolution trends in complex vibration signals, feature extraction methods based on multi-scale architectures have become a key research direction in this field. Existing technologies typically employ several approaches: first, constructing parallel multi-scale convolutional neural networks to mine hierarchical features of the signal using convolutional kernels of different sizes; second, developing end-to-end multi-scale learning methods to simultaneously extract spatiotemporal features of data to reduce information loss; and third, introducing multimodal, multi-level fusion mechanisms or graph convolutional cascade frameworks to overcome the model's dependence on a large number of training samples and address diagnostic challenges under complex operating conditions such as non-stationary rotational speeds. However, most of these existing methods rely on static feature fusion strategies and lack adaptive evaluation mechanisms for the importance of different time scales and physical channels. Therefore, they cannot dynamically recalibrate attention weights based on transient fault characteristics, a problem that urgently needs to be addressed. Summary of the Invention

[0004] To address the technical problems existing in the prior art, this invention provides a device fault diagnosis method based on causal residual learning and invariant risk minimization. This method can eliminate the interference of environmental confounding factors, comprehensively characterize the transient pulse characteristics and evolution trend characteristics of faults, and improve the cross-domain generalization performance and diagnostic accuracy of the model under unknown operating conditions.

[0005] To achieve the above objectives, the present invention provides the following technical solution: This invention discloses a device fault diagnosis method based on causal residual learning and invariant risk minimization, comprising steps S1-S4.

[0006] S1. Acquire multi-channel vibration signals of the target mechanical equipment during operation.

[0007] S2. Based on the causal structure discovery and backdoor adjustment criteria, the multi-channel vibration signal is de-complexed to eliminate interference components caused by changes in operating conditions and construct a causal residual sequence that only characterizes the physical mechanism of equipment failure.

[0008] S3. Input the causal residual sequence into a multi-timescale multi-channel attention network, and recalibrate the causal residual sequence based on the channel attention mechanism; based on the recalibrated causal residual sequence, extract the hierarchical temporal features at different time resolutions through a multi-branch structure with time alignment characteristics, and use the time-scale attention mechanism to perform dynamic weighted fusion to obtain the fault feature representation.

[0009] S4. Input the fault feature representation into the pre-trained fault classifier and output the final fault diagnosis result of the target mechanical equipment; wherein, the multi-timescale multi-channel attention network and the fault classifier are jointly pre-trained based on historical multi-working condition data, and an invariant risk minimization constraint is introduced into the optimization objective of the pre-training to enable the model to eliminate false correlations under specific working conditions and learn stable causal consistency features across working conditions.

[0010] As a further improvement to the above scheme, the specific process of step S2 includes: Using a causal discovery algorithm, the causal dependency between operating condition variables and multi-channel vibration signals is analyzed from historical observation data, and a directed acyclic graph structure is derived, thereby identifying the causal parent node set of the multi-channel vibration signals. Based on the directed acyclic graph structure, the operating condition variables in the causal parent node set are identified as mixed operating condition variables, and the mixed operating condition variables are used as control variables in the backdoor adjustment criterion. A generalized additive model is used to smoothly fit the nonlinear mapping between the mixed operating condition variables and the multi-channel vibration signal to estimate the predictable mixed components driven by the change in operating condition. The predictable confounding components are eliminated by regression in the multi-channel vibration signal, and only the residual variance that cannot be explained by the confounding variables of the operating conditions is retained, thereby reconstructing the causal residual sequence.

[0011] As a further improvement to the above scheme, in step S2, causal topology identification is performed based on the SURD framework. The specific process is as follows: Based on Shannon entropy and mutual information theory, the Shannon entropy of the future state of the target variable is decomposed into four independent components: redundant causality, unique causality, co-causality, and causal leakage. Specifically, multiple subsets of observed variables are constructed for the input operating condition variables and multi-channel vibration signals. By calculating the expected value of the specific mutual information increment derived from each subset of observed variables under the given target variable future state, the corresponding redundant causality, unique causality, and co-causality are quantified respectively. Based on the quantized components, the directed acyclic graph structure between the operating condition variables and multi-channel vibration signals is inferred.

[0012] As a further improvement to the above scheme, in step S3, the multi-timescale multi-channel attention network includes short-timescale branches, medium-timescale branches and long-timescale branches arranged in parallel; wherein, a prefix truncation operation is used to perform window-level time alignment on the one-dimensional causal residual sequence to ensure that the receptive fields of each branch present a nested inclusion relationship in the time domain, so that the features extracted at different time scales can characterize the same fault event.

[0013] As a further improvement to the above scheme, step S3, which involves extracting and fusing time-series features at different time resolutions, includes the following specific steps: One-dimensional convolutional neural networks set up in parallel in each time scale branch are used to extract fault features from the input signals of corresponding lengths. Among them, the short time scale branch uses small convolutional kernels to capture high-frequency transient impacts, while the medium and long time scale branches use progressively larger convolutional kernels to capture low-frequency evolution trends and global temporal dependence features, thus jointly constituting hierarchical temporal features. A time-scale attention network is constructed, and the network is used to perform global correlation analysis on the hierarchical temporal features to dynamically infer the adaptive importance weights of each scale branch feature for the current fault sample. The hierarchical time-series features are weighted and summed using the adaptive importance weights to obtain the fault feature representation.

[0014] As a further improvement to the above scheme, step S3, the specific process of recalibrating the causal residual sequence includes: For the one-dimensional causal residual sequence input at each time scale branch, global average pooling is used to extract the global average feature vectors of acceleration, velocity and displacement channels; The global average feature vector is nonlinearly transformed using a two-layer fully connected network, and dynamic attention coefficients for each physical channel are generated using an activation function. The dynamic attention coefficients are multiplied element-wise by channel with the corresponding original causal residual input sequence to obtain the recalibrated causal residual sequence, which is then used as the input to the subsequent convolutional neural network.

[0015] As a further improvement to the above scheme, in step S4, the joint objective optimization function expression in the pre-training phase is: ; In the formula, For the joint objective optimization function in the pre-training phase; This is a collection of training environments that include various historical operating condition data. For set A single specific source domain operating environment; For the model in a specific environment The risks of experience; For fixed virtual classifier parameters; These are the classifier parameters of the fault classifier; This is a feature extractor constructed from the multi-timescale, multi-channel attention network. These are the weighting coefficients; This indicates a fixed virtual classifier. In a specific environment The gradient norm obtained by taking the derivative below; For sample index, The total number of samples; Index for time-scale branches These represent the short, medium, and long time scale branches, respectively. For channel index, Total number of channels; For the first In the nth sample The weights of each branch; For the first In the nth sample Channel weights.

[0016] As a further improvement to the above scheme, in step S1, the multi-channel vibration signal is synchronously acquired by sensors deployed on the target mechanical equipment within the same sampling time window, specifically including discrete one-dimensional acceleration signal, velocity signal and displacement signal.

[0017] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the device fault diagnosis method based on causal residual learning and invariant risk minimization as described above.

[0018] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the device fault diagnosis method based on causal residual learning and invariant risk minimization as described above.

[0019] Compared with the prior art, the beneficial effects of the present invention are: 1. The equipment fault diagnosis method disclosed in this invention achieves stable fault identification under varying operating conditions by integrating a comprehensive diagnostic mechanism that combines causal residual construction, multi-timescale multi-channel attention networks, and invariant risk minimization constraints. This invention removes environmental interference components at the signal source level through causal discovery and backdoor adjustment, and introduces invariant risk minimization constraints during model pre-training, forcing the model to learn stable causal consistency characteristics across operating conditions. This effectively avoids the model fitting spurious correlations caused by external operating conditions such as speed and load. Simultaneously, by adaptively and dynamically weighting multi-source heterogeneous signals through channel attention and time-scale attention, it achieves a comprehensive characterization of fault transient pulses and evolutionary trends, improving the model's cross-domain generalization performance and diagnostic accuracy under unknown operating conditions.

[0020] 2. This invention employs a generalized additive model to smoothly fit the nonlinear dependency between confounding variables of operating conditions and multi-channel vibration signals. Furthermore, it utilizes a backdoor adjustment criterion to regress and eliminate predictable confounding components, thus achieving explicit decoupling between environmental confounding factors and fault characteristics. This strategy effectively filters out trend interference introduced by operating condition fluctuations, enabling the reconstructed causal residual sequence to accurately characterize independent causal changes caused by the fault itself, providing a cleaner and more reliable data foundation for subsequent feature extraction.

[0021] 3. This invention achieves strict time synchronization of features across multiple time scales on the same fault event by introducing a prefix-truncated time alignment strategy to construct the inputs of each network branch. This strategy enables the receptive fields of branches with different time scales (long, medium, and short) to exhibit a nested inclusion relationship in the time domain, avoiding the cross-event aliasing effect caused by time misalignment splicing, and ensuring the time consistency of high-frequency transient impacts and low-frequency modulation trends in multi-time scale representation.

[0022] 4. This invention achieves adaptive fusion of heterogeneous physical signals and multi-resolution temporal features by constructing a dual dynamic weighting mechanism that includes physical channel attention and temporal scale attention. This mechanism can automatically infer and strengthen key physical channel features with high signal-to-noise ratios based on the characteristics of the specific fault sample, while dynamically adjusting the aggregation weights of different temporal scale features in the overall representation. This overcomes the lack of flexibility in traditional static fixed fusion strategies, enabling the model to robustly select the optimal physical channel and temporal scale representations.

[0023] 5. This invention introduces regularization constraints of channel distribution entropy and time scale distribution entropy into the pre-training joint optimization objective function, effectively incentivizing the diversity of model feature exploration. This mechanism prevents the attention network from degenerating into trivial solutions that rely solely on a single redundant channel or a single time scale during training, guiding the model to fully explore and utilize the complementary information between different physical sensors and different temporal resolutions, further ensuring the effectiveness and comprehensiveness of the fused features. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the cause-effect graph in Embodiment 1 of the present invention.

[0025] Figure 2 This is a flowchart of the equipment fault diagnosis method based on causal residual learning and invariant risk minimization in Embodiment 1 of the present invention.

[0026] Figure 3 This is a directed acyclic graph of the operating condition confusion effect in Embodiment 1 of the present invention.

[0027] Figure 4 This is a structural causal model diagram of the IRM principle in an embodiment of the present invention.

[0028] Figure 5 This is the laboratory mechanical vibration analysis and fault simulation test bench in Embodiment 1 of the present invention.

[0029] Figure 6 The average F1 score is the average score of each method under the four bearing fault labels (Normal: normal bearing, IF: inner ring fault, OF: outer ring fault, BF: ball fault) in Embodiment 1 of the present invention.

[0030] Figure 7 This is a visualization of the classification performance of different models in Embodiment 1 of the present invention; Figure 7 In the diagram, (a) represents the method of this invention, (b) represents the CDDG method, (c) represents the DGNIS method, (d) represents the CCDG method, (e) represents the CCN method, and (f) represents the IEDGNet method.

[0031] Figure 8 The diagnostic accuracy of different methods in each case in Embodiment 1 of the present invention is shown.

[0032] Figure 9 These are the ROC curves of different methods in Embodiment 1 of the present invention; Figure 9In the diagram, (a) represents the method of this invention, (b) represents the CDDG method, (c) represents the DGNIS method, (d) represents the CCDG method, (e) represents the CCN method, and (f) represents the IEDGNet method; the horizontal axis FPR represents the false positive rate, the vertical axis TPR represents the true positive rate, AUC represents the area under the ROC curve, N: normal bearing, I: inner ring fault, O: outer ring fault, and B: ball fault.

[0033] Figure 10 This is a schematic diagram of the structure of the computer terminal in Embodiment 2 of the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Example 1 This embodiment provides a device fault diagnosis method based on causal residual learning and invariant risk minimization, which can be applied to the fault diagnosis of rotating machinery (e.g., bearings).

[0036] First, we will introduce causal discovery and causal graphs.

[0037] This invention employs a structural causal model to characterize the dependencies between operating condition variables, sensor signals, and fault states. The structural causal model assumes that each variable... From its parent node set Correlation function and independent noise Co-generation: (1) In the formula, d The total number of variables. Directed acyclic graphs in graph structures correspond to conditional independence relations and can be learned from multivariate observation data using constraint-based causal discovery algorithms (such as the Peter-Clark (PC) algorithm). The PC algorithm starts with a complete graph and progressively removes edges and directed edges through conditional independence tests, ultimately obtaining a causal graph or equivalence class representation consistent with the observation distribution. For example... Figure 1 The diagram shown is a cause-and-effect graph. Figure 1 (a) in the model is a structural causal model. Figure 1 (b) in the graph is a directed acyclic graph. Figure 1 (c) in the graph is a partially directed acyclic graph.

[0038] In fault diagnosis scenarios, this embodiment models operating condition variables and multi-channel vibration signals together as nodes in a causal graph. Using a causal discovery algorithm, this embodiment infers a directed acyclic graph structure that characterizes the relationship between operating condition variables and sensor channels. Based on this inferred causal structure, this embodiment uses an effective regression method to calculate direct causal residuals only based on the corresponding parent node set, thereby eliminating predictable components driven by operating conditions at a precise signal level.

[0039] Next, we will introduce the minimization of invariant risk.

[0040] Traditional Empirical Risk Minimization (ERM) minimizes the average risk of all training samples: (2) (3) in, This represents the set of training environments, where each environment... e This corresponds to a different data distribution determined by a specific combination of sampling rate, rotational speed, and load conditions. , This refers to cross-entropy loss. Since empirical risk minimization aims only to minimize the overall average loss across all environments, it indiscriminately fits all significant correlations in the training data, including spurious features caused by changes in operating conditions.

[0041] Invariant Risk Minimization (IRM) aims to explore invariant predictive relationships among multiple environments: it seeks a representation Φ and a classifier w such that the same classifier is optimal in all environments simultaneously. Its optimization objective is shown in equation (4): (4) This constraint fundamentally compresses the mapping space at the functional level, excluding feature combinations that depend on specific environmental decision boundaries, and retaining only those that can be mapped through a unified function. A stable pattern that can be explained in all environments. Therefore, this constraint is consistent with the causal constraint, that is, causality remains unchanged under interventionist changes in environmental conditions.

[0042] In cross-domain fault diagnosis scenarios, let the input space be... (Represents a multi-channel signal of length L), the tag space is This embodiment defines an environment set. Each of these environments Corresponding to a specific operating condition, this invention systematically solves two key fundamental problems that hinder reliable practical diagnostics: Problem I: Generalization failure caused by spurious associations and environmental confounding factors. Standard diagnostic models typically approximate the conditional distribution P(Y|X) by minimizing empirical risk. From a strictly causal perspective, the observed signal can be directly generated by the standard structural equation (5): (5) Where Y represents the fault state, E represents the operating condition, and N represents noise. Since E directly affects X, prediction models trained based on empirical risk minimization may encode spurious correlations dependent on the environment, leading to instability in the prediction mechanism, for example: (6) Therefore, when the operating conditions originate from the source environment Transform into an unseen target domain environment When this happens, the joint distribution will change, that is: (7) This could cause the model to fail.

[0043] Therefore, the first objective of this invention is to explicitly block interference paths at the signal level. E→X Construct causal residual sequences Specifically, learning a residual construction operator. , making (8) Based on the decontamination of the residuals, this embodiment further employs invariant risk minimization to learn the causal invariant mapping: (9) This allows the prediction mechanism to remain stable under different environments: (10) Problem II: Limitations of static feature fusion across time scales and sensor channels. Given a multi-channel input... It typically contains high-frequency transient components. With low-frequency modulation components And satisfy A fixed window length L cannot simultaneously satisfy both conflicting requirements.

[0044] To model information across multiple time scales, this embodiment introduces a set of time scale operators. and construct (11) Its corresponding multi-timescale representation is However, most existing methods use sample-independent weights for static fusion, for example: (12) A similar fixed weight is used in channel fusion, which implicitly assumes that the importance of each time scale and sensor channel remains constant under different samples and operating conditions.

[0045] Therefore, the second objective of this invention is to learn an adaptive dynamic weighting function for samples, namely, time-scale attention weights. With channel attention weights And construct an adaptive fusion representation (13) in This represents the feature after channel reweighting at the s-th time scale. This allows the model to dynamically adjust the contribution of features between temporal granularity and sensor modality for each specific sample.

[0046] Please see Figure 2 The diagnostic method includes steps S1-S4.

[0047] S1. Acquire multi-channel vibration signals of the target mechanical equipment during operation.

[0048] In step S1, the multi-channel vibration signal is synchronously acquired by sensors deployed on the target mechanical equipment within the same sampling time window, specifically including discrete one-dimensional acceleration signal, velocity signal and displacement signal.

[0049] S2. Based on the causal structure discovery and backdoor adjustment criteria, the multi-channel vibration signal is de-complexed to eliminate interference components caused by changes in operating conditions and construct a causal residual sequence that only characterizes the physical mechanism of equipment failure.

[0050] S3. Input the causal residual sequence into a multi-timescale multi-channel attention network, and recalibrate the causal residual sequence based on the channel attention mechanism; based on the recalibrated causal residual sequence, extract the hierarchical temporal features at different time resolutions through a multi-branch structure with time alignment characteristics, and use the time-scale attention mechanism to perform dynamic weighted fusion to obtain the fault feature representation.

[0051] S4. Input the fault feature representation into the pre-trained fault classifier and output the final fault diagnosis result of the target mechanical equipment; wherein, the multi-timescale multi-channel attention network and the fault classifier are jointly pre-trained based on historical multi-working condition data, and an invariant risk minimization constraint is introduced into the optimization objective of the pre-training to enable the model to eliminate false correlations under specific working conditions and learn stable causal consistency features across working conditions.

[0052] It should be noted that the training and testing samples are of the same type, and residual and multi-timescale multi-channel attention mechanisms were also implemented during historical training.

[0053] The following is a detailed breakdown of the fault diagnosis framework of this invention, which consists of three integrated core modules: a causal residual construction module, used to eliminate clutter interference and decouple the feature components containing fault information from the original vibration signal; a causal invariant feature learning module based on invariant risk minimization, which introduces an IRM regularization term to ensure that the feature representation learned from the causal residual sequence remains stable under different operating conditions; and a multi-timescale multi-channel attention network module, which acts as a deep feature extractor and achieves adaptive fusion through a dual attention mechanism to capture hierarchical temporal features and heterogeneous physical channel features.

[0054] (1) Causal residual component module The initial stage aims to construct causal residual sequences, thereby explicitly decoupling fault-sensitive components from environmental confounding factors. This process is achieved through two consecutive steps: causal topology identification and residual generation.

[0055] 1) Causal Topology Identification Based on the SURD Framework: In the causal discovery stage, this invention adopts the SURD framework proposed by Martínez-Sánchez et al. This method is naturally applicable to the causal decomposition task of multivariate time series data and can effectively handle the inherent complex characteristics of bearing data. The SURD framework is based on Shannon entropy and mutual information theory, decomposing the entropy of the future state of the target variable into four independent components: unique causality, redundant causality, co-causality, and causal leakage. Its mathematical form is as follows: (14) Among them, it means Shannon entropy of the future state of the target variable. , and They represent subsets of observed variables, respectively. arrive Redundant causality, unique causality, and co-causality. Furthermore, Used to quantify causal leakage, reflecting the causal impact of unobserved variables; while C represents the set consisting of all observed subsets containing two or more variables. Redundant causality, unique causality, and co-causality are mathematically defined as the expected value of their corresponding specific mutual information increments, in the following forms: (15) in , , Represents a specific future state of the target variable. Under the condition of a subset of observed variables The specific redundancy, synergy, or unique information increments derived from it, and This represents the probability of that future state occurring.

[0056] First, the causal residual learning module implements model-driven decontamination processing during the input phase. Based on the physical mechanism of signal generation, such as... Figure 3 As shown, this invention models the observed vibration signal as a linear superposition of two different components: one is the fault-sensitive component originating from the fault mechanism, and the other is the mixed environment component caused by changes in operating conditions. Its specific expression is shown in equation (16): (16) In the formula, This represents the causal contribution of fault label Y to signal X (i.e., the fault component), which reflects the driving effect of the fault on the vibration signal. This represents the confounding contribution of the operating condition variable Z to the signal X, reflecting the interference of external factors such as operating conditions on the vibration signal; ε This is the random error term.

[0057] By combining causal parent nodes and a health status reference window, a causal regression elimination strategy is employed to suppress condition-driven components in the original signal and reconstruct a causal residual sequence focused on the fault mechanism. This process aims to achieve cross-domain fault feature alignment, thereby significantly assisting the classifier in constructing domain-independent decision boundaries. This residualization strategy fundamentally aims to decouple the confounding effects of variable Z from the original signal, ensuring that only causal components driven by the fault itself are extracted.

[0058] 2) Backdoor adjustment and hybridization elimination based on generalized additive model: Since the modulation effect of operating variables such as rotational speed and load on the vibration channel is essentially nonlinear, continuous, and smooth, this invention uses a generalized additive model to model this hybrid term. The generalized additive model framework has significant advantages: it uses an interpretable additive smoothing function to characterize complex nonlinear dependencies, and at the same time, through the roughness penalty term and the automatic selection mechanism of smoothing parameters, it achieves an optimal balance between fitting accuracy and model complexity. Its mathematical form can be described by equation (17): (17) In the formula, For the target response variable, For the intercept term, For smoothing functions, For the first j The predictor variables at time 1 t The value of , This is the error term.

[0059] To avoid overfitting the model to transient pulse impacts, and thus ensure... This invention employs a penalized spline method, which can only capture stable and inherent nonlinear dependencies in a healthy state. The objective function (18) is minimized to achieve the desired result. Parameter estimation: (18) In the formula, the first term measures the goodness of fit, and the second term is determined by the smoothing parameter. Weighted roughness penalty term. Smoothing parameter. Generalized cross-validation or unbiased risk estimation methods are typically used to automatically select the cross-validation criterion. For example, the commonly used generalized cross-validation criterion can be written as equation (19): (19) In the formula, n is the number of observed samples. Indicates model bias. This represents the effective degrees of freedom. The selection process for this parameter ensures the function... Without overfitting to the health state reference window, the operating condition components are analyzed to the greatest extent possible.

[0060] Therefore, the desired causal residual signal can be obtained. : (20) (twenty one) The core of calculating causal residual signals lies in analyzing the fluctuation patterns of the observed signal X under the condition of controlling the confounding variable Z. This process aims to separate the statistical dependence explained by Z from the vibration signal, retaining only the residual variance that cannot be explained by Z. Such residuals can effectively characterize the independent causal driving effect generated by the fault Y. From a methodological perspective, this algorithm is an engineering implementation of the backdoor adjustment formula, and its core logic strictly follows the backdoor criterion in causal inference, which can be specifically expressed by equation (22) as follows: (twenty two) The core contribution of this method lies in constructing a transformation mechanism based on residualization, effectively transforming the abstract backdoor adjustment paradigm in causal inference into a feasible signal processing flow for fault diagnosis scenarios. By employing a generalized additive model to model and regress the interference of exogenous confounding variables Z in the causal parent node set, the extracted residual signal... It can accurately characterize the independent causal changes caused by faults. This process not only strictly follows the confounding control criteria of causal diagrams, but also achieves effective decoupling of environmental confounding factors and fault characteristics in vibration signals, thereby breaking down the barrier between theoretical rigor and engineering practicality.

[0061] Although residual generation methods based on causal graphs can filter out operating condition trends and implicitly encode domain knowledge, potential environment-dependent variations still remain in the signal. Therefore, invariant risk minimization learns and obtains domain-invariant causal representations from causal residuals by constraining the feature extraction process.

[0062] (2) Causal invariant feature learning module based on invariant risk minimization 1) Theoretical Construction of the Causal Invariance Principle: After acquiring the causal residual signal rich in fault information, this module aims to extract feature representations that remain unchanged under various operating conditions. For example... Figure 4 As shown, formally defined, each different working condition is abstracted into an independent environment. e During the training phase, it is assumed that a dataset covering multiple source environments is available; these source environments are denoted as... The core objective of this invention is to train a feature extractor and classifier that not only exhibits excellent performance in the aforementioned known environments, but also in unknown target environments. It demonstrates strong generalization ability.

[0063] This invention employs the invariant risk minimization paradigm proposed by Arjovsky et al. Formally, within the framework of this invention, let Φ( Let denot Φ be the feature extractor, whose function is to map the input causal residual signal into latent feature representations; let u denote the parameters of the final classifier used for fault category prediction. The core principle of the invariant risk minimization paradigm is to find a feature representation Φ such that there exists a unique and invariant classifier parameter. This allows for the simultaneous achievement of optimal classification performance across all training environments. This crucial requirement constitutes a fundamental structural invariance constraint on the final decision function.

[0064] In multi-category fault diagnosis scenarios, the invariance principle makes the following assumption: given the extracted feature representations, the conditional probability distribution of fault labels must remain constant across various operating conditions. Mathematically, this constraint can be expressed as: (twenty three) In the formula, e and e′ represent different working conditions; and These are the original input data sampled from environment e and environment e′, respectively. and These are the actual fault labels corresponding to the above input data; This is a feature extraction function that maps the input to the latent feature space. For the specific values ​​of the latent feature vector, i.e. . Extract features for a given When, the conditional probability distribution of the fault category. Let be the support set of the distribution, and let represent the range of values ​​for which the probability is nonzero.

[0065] 2) Solvable Optimization Based on Gradient Penalty Relaxation: Invariant Risk Minimization (IRM) aims to learn a feature representation Φ such that a single classifier w based on this representation can synchronously reach its optimum in all training environments. This requirement imposes strict structural constraints, effectively compressing the solution space of the decision function. However, since directly satisfying this hard constraint is computationally difficult, this invention adopts a relaxation strategy based on gradient penalty, namely the IRMv1 algorithm. Based on the first-order stationarity condition of optimality, this embodiment improves the solution space of the decision function by fixing the virtual classifier. This constraint is achieved by penalizing the gradient norm in each environment. This mechanism forces the model to discard features that depend on environment-specific pseudo-correlation, retaining only causal features that maintain a stable association with the fault label in different domains. Finally, the overall training objective function of this invention is constructed as shown in equation (24): (twenty four) make Let represent the sample-by-sample loss function, and further define it as follows: (25) In the formula, s It is a linear weighted score of the fault features by the classifier. g Let be the mapping function from logarithmic odds to probability. Based on this, the environment... e Experience risk can be expressed as: (26) By leveraging the regularization effect of the gradient penalty term, the feature extractor Φ is forced to discard features that depend on specific environmental perturbations. Essentially, the core objective of minimizing invariant risk (ensuring the classifier parameters w are simultaneously optimal across all environments) is inherently consistent with exploring universally applicable causal mechanisms for failures. By formalizing the core principle that invariance implies causality as an optimization constraint, IRM effectively compresses the function space using cross-environment label information, guiding the model to learn the true causal mapping between failures and observed effects. This deep constraint on the function space significantly surpasses the effects achievable solely through signal-level residual processing.

[0066] (3) Multi-timescale multi-channel attention network module The core module of the framework proposed in this invention is a multi-timescale, multi-channel attention network, which aims to model the differential feature contributions generated by different temporal resolutions and heterogeneous physical channels. By performing adaptive weighting on these multi-view inputs, the network can effectively learn fault features with strong discriminative power.

[0067] 1) Construction of time-aligned input based on prefix truncation strategy: Specifically, the input of this network consists of three sets of time-synchronized one-dimensional time series, namely acceleration... ,speed and displacement The discrete-time index is denoted as t∈{1,…,T}, where T represents the uniform length of the sampling time window (T=2048 in this embodiment). To achieve multi-source information fusion, this invention constructs a unified three-channel input representation, defined as shown in equation (27): (27) To ensure strict alignment between the extracted multi-timescale features and the same time interval of a specific fault event, this invention employs a prefix truncation-based strategy to construct the network input. This method achieves window-level time alignment, ensuring that the input signals of all timescale branches originate from the same continuous sampling sequence, and that this sequence lies within a unified time interval. Therefore, the receptive fields of each timescale branch exhibit a nested inclusion relationship in the time domain, thereby ensuring that the features extracted at different timescales can consistently represent the same fault event. The mathematical form of this relationship is defined as shown in equation (28): (28) The prefix truncation operator is formally defined as follows: ( In this way, the present invention can effectively avoid the cross-event aliasing effect caused by time misalignment splicing, thereby ensuring the temporal consistency of fault modes in multi-timescale representation. The mathematical expression of this relationship is shown in equation (29): (29) By relying on strict time alignment constraints, the model can acquire highly complementary feature representations that simultaneously cover high-frequency transient impacts and low-frequency modulation trends. As a result, the model significantly improves the discriminative efficiency across time scales, while also enhancing diagnostic accuracy and generalization ability.

[0068] 2) Hierarchical temporal feature fusion based on time-scale attention mechanism: In actual engineering implementation, this network architecture consists of three parallel feature extraction branches, denoted as {S, M, L}. The backbone network of each branch adopts a one-dimensional convolutional neural network, and its formal definition is shown in equation (30): (30) In the formula, f τ ( ) represents the convolutional feature extractor corresponding to the branch at time scale τ, θ τ Let C represent the set of learnable parameters of the one-dimensional convolutional neural network corresponding to the τ-th time-scale branch.τ T represents the number of feature channels. τ This indicates the timing length after pooling operations.

[0069] Specifically, the short timescale (S) branch employs a limited local receptive field, exhibiting superior performance in capturing high-frequency transient shocks; the medium timescale (M) branch is used to extract feature patterns at medium timescales; in contrast, the long timescale (L) branch processes the complete 2048-point observation window, thus effectively characterizing the low-frequency evolution trends and global temporal dependencies associated with persistent faults.

[0070] After feature extraction is completed in each parallel branch, three different feature vectors are obtained, denoted as follows: , and Formally, the aforementioned eigenvectors are stacked to construct a unified multi-time-scale feature matrix. Its definition is shown in equation (31): (31) Subsequently, an attention network was constructed to dynamically infer the adaptive importance weights of each timescale component. Using the obtained weight vector, the model efficiently aggregates discrete multi-timescale features through a weighted summation mechanism, ultimately obtaining a unified feature representation with strong discriminative ability for fault modes. Its specific mathematical form is shown in Equation (32): (32) in This represents the weight of the τ-th branch in the k-th sample. Meanwhile, to reduce the risk of the attention mechanism degenerating into a trivial solution (i.e., all weights concentrated on a single branch), this invention introduces a time-scale distribution entropy regularization term into the objective function. This regularization term aims to incentivize the model to maintain effective exploration diversity of features across multiple time scales, avoiding convergence to a suboptimal local optimum dominated by a single time scale. Its rigorous mathematical formalization is as follows: (33) Compared with traditional static fixed fusion strategies, this dynamic weighting paradigm has greater flexibility in feature aggregation, enabling the model to intelligently identify and focus on the optimal time scale representation for different fault types.

[0071] 3) Heterogeneous Physical Channel Fusion Based on Channel Attention Mechanism: In the parallel branch of multi-timescale temporal feature processing, this invention introduces a multi-channel physical attention mechanism, which is used to adaptively fuse information from three different physical signal channels (acceleration, velocity, and displacement). Specifically, at the input stage of each timescale branch, the model learns a physical importance weight vector through a channel attention network. The input sequence is then recalibrated using this weight. Its mathematical expression is as follows: (34) In the formula, It is the global average feature vector of the physical channels in the k-th sample. Represents a non-linear activation function. The first-layer learnable weight matrix for physical channel attention. The second-layer learnable weight matrix represents the attention of the physical channels, where ⊙ denotes element-wise multiplication by channel. Intuitively, the learned weights... It can adaptively enhance key physical domain components (which contain the most diagnostically valuable information of the current sample) while effectively suppressing irrelevant interference introduced by low-information channels.

[0072] Subsequently, the feature sequences were physically recalibrated. The convolutional backbone network f is input to the corresponding branch τ In (·), to facilitate deep feature extraction. The rigorous mathematical formalization of this operation is shown in Equation (35): (35) Adopting a similar design approach, to mitigate the risk of channel weights being overly concentrated on a single physical signal, this invention introduces a channel distribution entropy regularization term into the objective function. This regularization term aims to incentivize the model to maintain exploration diversity along the physical channel dimension and ensure that the fused features can fully utilize the complementary advantages of acceleration, velocity, and displacement signals. Its specific expression is as follows: (36) in This represents the weight of the c-th channel in the k-th sample. This feature recalibration process adaptively strengthens the physical signal domain that has core diagnostic significance for the current critical fault stage. Different fault modes exhibit highly significant characteristics in the critical physical domain: impact faults induce obvious acceleration spikes, thus giving the acceleration channel a higher weight; while low-frequency degradation modes have higher observability in the velocity or displacement domain. By endowing the diagnostic network with adaptive autonomous weight adjustment capabilities, this innovative mechanism effectively mitigates the potential risk of decision bias caused by a single redundant channel, thereby ensuring robust dynamic selection of the optimal physical channel with the highest signal-to-noise ratio for the target fault characteristics.

[0073] In summary, this invention constructs a time-domain multi-timescale multi-channel network. Signals with three window lengths are processed in parallel to characterize fault modes at short, medium, and long time scales, respectively. Simultaneously, causal residual acceleration, residual velocity, and residual displacement are used as multi-channel inputs. A compressed-excitation channel attention module is employed to learn the importance of each physical channel, and a time-scale attention module is used to achieve weighted fusion across time scales. By adaptively reinforcing time-scale-channel features with high diagnostic information content and stability, the proposed network effectively improves diagnostic accuracy while maintaining a lightweight design.

[0074] By combining all the key components mentioned above, this invention derives the overall optimization objective function by aggregating the Empirical Risk Minimization (ERM) term, the Invariance Risk Penalty term, and the Double Entropy Regularization term. The design goal of this composite loss function is to maximize the inherent discriminative diversity of the overall features across multiple time scales and channels while enhancing the model's cross-domain generalization robustness. Its mathematical form is defined as shown in equation (37): (37) In the formula, For the joint objective optimization function in the pre-training phase; This is a collection of training environments that include various historical operating condition data. For set A single specific source domain operating environment; For the model in a specific environment The risks of experience; For fixed virtual classifier parameters; These are the classifier parameters of the fault classifier; This is a feature extractor constructed from the multi-timescale, multi-channel attention network. These are the weighting coefficients; This indicates a fixed virtual classifier. In a specific environment The gradient norm obtained by taking the derivative below; For sample index, The total number of samples; Index for time-scale branches These represent the short, medium, and long time scale branches, respectively. For channel index, Total number of channels; For the first In the nth sample The weights of each branch; For the first In the nth sample Channel weights.

[0075] To verify the effectiveness of the diagnostic method proposed in this invention, the following experiments are also provided in this embodiment: A. Data Description and Proposed Framework Validation Laboratory test bench bearing dataset: such as Figure 5 As shown, the laboratory test bench bearing dataset was collected from a laboratory mechanical vibration analysis and fault simulation test bench, which consists of a drive motor, bearings, sensors, and other components. The dataset contains four sample types: outer ring fault, inner ring fault, rolling element fault, and normal state. Data was collected at speeds of 500–1000 rpm, and operating condition data were also collected to construct a directed acyclic graph (DAG) to generate causal residual sequences.

[0076] As shown in Table 1, this laboratory dataset summarizes the types of bearing failures. In this dataset, Normal indicates that the bearing is in a normal state, OF indicates outer ring failure, IF indicates inner ring failure, and BF indicates ball failure.

[0077] Table 1: Laboratory Datasets

[0078] To verify the effectiveness of the proposed framework, this invention uses accuracy, precision, recall, F1 score, confusion matrix, and ROC curve to evaluate the model performance.

[0079] (38)

[0080] Among them, TP, FP, TN and FN correspond to the English terms true positive, false positive, true negative and false negative, respectively.

[0081] To verify the superiority of the proposed framework, this embodiment conducted cross-operating condition fault diagnosis experiments on a laboratory dataset. Specifically, this embodiment used measured data under five speed conditions, randomly selecting four conditions as the source domain for training, and using the remaining domain for generalization performance testing. Here, A represents the 500 r / min condition, B represents the 625 r / min condition, C represents the 750 r / min condition, D represents the 875 r / min condition, and E represents the 1000 r / min condition. The experimental results were compared with several representative methods, including DGNIS, CCN, IEDGNet, CDDG, and CCDG, to highlight the excellent generalization performance of the proposed method.

[0082] Depend on Figure 6As can be seen, the average F1 scores obtained by the proposed method for the four fault types under different generalization tasks are 0.9360, 0.8466, 0.8585, and 0.8878, respectively. However, this method does not achieve optimal performance on all single fault labels. For example, on the normal bearing label, the CDDG method achieves a score of 0.9744, which is superior to the method of this invention.

[0083] Figure 7 The confusion matrix obtained in the tenth experiment under Case 5 is shown, and it is clear that the method of the present invention has the best classification performance (diagnostic performance). Figure 8 The average accuracy of the six comparison methods is shown across five cases.

[0084] Depend on Figure 9 As can be seen from the ROC curve, the proposed method maintains a high recall rate while having a low false positive rate.

[0085] As shown in Table 2, it can be seen that the accuracy and standard deviation (%) of the diagnostic results of the method proposed in this invention are higher than those of the current best method. Ten experiments were conducted for each case, and the average value and standard deviation were obtained.

[0086] Table 2: Diagnostic results of the method of the present invention and existing methods in different cases

[0087] Example 2 This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the fault diagnosis method as described in Embodiment 1.

[0088] like Figure 10 As shown, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to at least one processor 101. This embodiment does not limit the specific connection medium between the processor 101 and the memory 102. Figure 10 The example shown is the connection between processor 101 and memory 102 via bus 100. Bus 100 is... Figure 10 The connections between other components are shown in bold lines and are for illustrative purposes only, not as limiting information. Bus 100 can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 10 The bus is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. Alternatively, the processor 101 may also be called a controller; there is no restriction on the name.

[0089] In this embodiment, the memory 102 stores instructions that can be executed by at least one processor 101. The at least one processor 101 can execute the aforementioned method by executing the instructions stored in the memory 102.

[0090] The processor 101 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 102 and calling data stored in memory 102, the processor can perform various functions and process data, thereby monitoring the device as a whole.

[0091] In one possible design, processor 101 may include one or more processing units. Processor 101 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 101. In some embodiments, processor 101 and memory 102 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.

[0092] Processor 101 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the fault diagnosis method disclosed in Embodiment 1 can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules in processor 101.

[0093] Memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 102 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 102 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In this embodiment, memory 102 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0094] By designing and programming the processor 101, the code corresponding to the fault diagnosis method described in the foregoing embodiments can be embedded into the chip, thereby enabling the chip to execute the code during operation. Figure 2 The steps of the fault diagnosis method shown are as follows. How to design and program the processor 101 is a technique well-known to those skilled in the art and will not be described further here.

[0095] Example 3 This embodiment provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the fault diagnosis method as described in Embodiment 1.

[0096] The computer-readable storage medium may include flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., provided on the computer device. Of course, the storage medium may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory is typically used to store the operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various types of data that have been output or will be output.

[0097] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for equipment fault diagnosis based on causal residual learning and invariant risk minimization, characterized in that, include: S1. Acquire multi-channel vibration signals of the target mechanical equipment during operation; S2. Based on the causal structure discovery and backdoor adjustment criteria, the multi-channel vibration signal is de-complexed to eliminate interference components caused by changes in operating conditions and construct a causal residual sequence that only characterizes the physical mechanism of equipment failure. S3. Input the causal residual sequence into a multi-timescale multi-channel attention network, and recalibrate the causal residual sequence based on the channel attention mechanism; based on the recalibrated causal residual sequence, extract the hierarchical temporal features at different time resolutions through a multi-branch structure with time alignment characteristics, and use the time-scale attention mechanism to perform dynamic weighted fusion to obtain the fault feature representation. S4. Input the fault feature representation into the pre-trained fault classifier and output the final fault diagnosis result of the target mechanical equipment; wherein, the multi-timescale multi-channel attention network and the fault classifier are jointly pre-trained based on historical multi-working condition data, and an invariant risk minimization constraint is introduced into the optimization objective of the pre-training to enable the model to eliminate false correlations under specific working conditions and learn stable causal consistency features across working conditions.

2. The equipment fault diagnosis method based on causal residual learning and invariant risk minimization according to claim 1, characterized in that, The specific process of step S2 includes: Using a causal discovery algorithm, the causal dependency between operating condition variables and multi-channel vibration signals is analyzed from historical observation data, and a directed acyclic graph structure is derived, thereby identifying the causal parent node set of the multi-channel vibration signals. Based on the directed acyclic graph structure, the operating condition variables in the causal parent node set are identified as mixed operating condition variables, and the mixed operating condition variables are used as control variables in the backdoor adjustment criterion. A generalized additive model is used to smoothly fit the nonlinear mapping between the mixed operating condition variables and the multi-channel vibration signal to estimate the predictable mixed components driven by the change in operating condition. The predictable confounding components are eliminated by regression in the multi-channel vibration signal, and only the residual variance that cannot be explained by the confounding variables of the operating conditions is retained, thereby reconstructing the causal residual sequence.

3. The equipment fault diagnosis method based on causal residual learning and invariant risk minimization according to claim 2, characterized in that, In step S2, causal topology identification is performed based on the SURD framework. The specific process is as follows: Based on Shannon entropy and mutual information theory, the Shannon entropy of the future state of the target variable is decomposed into four independent components: redundant causality, unique causality, co-causality, and causal leakage. Specifically, multiple subsets of observed variables are constructed for the input operating condition variables and multi-channel vibration signals. By calculating the expected value of the specific mutual information increment derived from each subset of observed variables under the given target variable future state, the corresponding redundant causality, unique causality, and co-causality are quantified respectively. Based on the quantized components, the directed acyclic graph structure between the operating condition variables and multi-channel vibration signals is inferred.

4. The equipment fault diagnosis method based on causal residual learning and invariant risk minimization according to claim 1, characterized in that, In step S3, the multi-timescale multi-channel attention network includes parallel short-timescale branches, medium-timescale branches, and long-timescale branches; wherein, a prefix truncation operation is used to perform window-level time alignment on the one-dimensional causal residual sequence to ensure that the receptive fields of each branch present a nested inclusion relationship in the time domain, so that the features extracted at different time scales can characterize the same fault event.

5. The equipment fault diagnosis method based on causal residual learning and invariant risk minimization according to claim 4, characterized in that, In step S3, the specific process of extracting and fusing hierarchical temporal features at different time resolutions includes: One-dimensional convolutional neural networks set up in parallel in each time scale branch are used to extract fault features from the input signals of corresponding lengths. Among them, the short time scale branch uses small convolutional kernels to capture high-frequency transient impacts, while the medium and long time scale branches use progressively larger convolutional kernels to capture low-frequency evolution trends and global temporal dependence features, thus jointly constituting hierarchical temporal features. A time-scale attention network is constructed, and the network is used to perform global correlation analysis on the hierarchical temporal features to dynamically infer the adaptive importance weights of each scale branch feature for the current fault sample. The hierarchical time-series features are weighted and summed using the adaptive importance weights to obtain the fault feature representation.

6. The equipment fault diagnosis method based on causal residual learning and invariant risk minimization according to claim 5, characterized in that, Step S3, the specific process of recalibrating the causal residual sequence includes: For the one-dimensional causal residual sequence input at each time scale branch, global average pooling is used to extract the global average feature vectors of acceleration, velocity and displacement channels; The global average feature vector is nonlinearly transformed using a two-layer fully connected network, and dynamic attention coefficients for each physical channel are generated using an activation function. The dynamic attention coefficients are multiplied element-wise by channel with the corresponding original causal residual input sequence to obtain the recalibrated causal residual sequence, which is then used as the input to the subsequent convolutional neural network.

7. The equipment fault diagnosis method based on causal residual learning and invariant risk minimization according to claim 4, characterized in that, In step S4, the joint objective optimization function expression in the pre-training phase is: In the formula, For the joint objective optimization function in the pre-training phase; This is a collection of training environments that include various historical operating condition data. For set A single specific source domain operating environment; For the model in a specific environment The risks of experience; For fixed virtual classifier parameters; These are the classifier parameters of the fault classifier; This is a feature extractor constructed from the multi-timescale, multi-channel attention network. These are the weighting coefficients; This indicates a fixed virtual classifier. In a specific environment The gradient norm obtained by taking the derivative below; For sample index, The total number of samples; Index for time-scale branches These represent the short, medium, and long time scale branches, respectively. For channel index, Total number of channels; For the first In the nth sample The weights of each branch; For the first In the nth sample Channel weights.

8. The equipment fault diagnosis method based on causal residual learning and invariant risk minimization according to claim 1, characterized in that, In step S1, the multi-channel vibration signal is synchronously acquired by sensors deployed on the target mechanical equipment within the same sampling time window, specifically including discrete one-dimensional acceleration signal, velocity signal and displacement signal.

9. A computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the equipment fault diagnosis method based on causal residual learning and invariant risk minimization as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the equipment fault diagnosis method based on causal residual learning and invariant risk minimization as described in any one of claims 1 to 8.