Rolling bearing fault diagnosis method based on maximum independence causal decoupling

By constructing a causal decoupling network based on a maximum independence autoencoder and introducing non-causal nuclear norm Wasserstein distance alignment, the problem of insufficient generalization performance of causal mechanisms in rolling bearing fault diagnosis is solved, and high-accuracy cross-device fault diagnosis is achieved.

CN121901608APending Publication Date: 2026-04-21BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Current cross-domain generalization methods based on causal mechanisms are insufficient in decoupling causal features and generalization performance in rolling bearing fault diagnosis, resulting in low stability and efficiency when analyzing complex fault types and operating conditions.

Method used

A causal decoupling network based on a maximum independence autoencoder is constructed. The Hilbert-Schmidt independence criterion and the non-causal nuclear norm Wasserstein distance alignment are introduced to extract causal features and improve the generalization ability of the model, thereby achieving end-to-end fault diagnosis.

Benefits of technology

It achieves high-accuracy rolling bearing fault diagnosis under complex working conditions, with an average diagnostic accuracy of 95.89%. It overcomes the distribution differences caused by different equipment and working conditions and has the ability to diagnose faults across equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901608A_ABST
    Figure CN121901608A_ABST
Patent Text Reader

Abstract

The invention discloses a rolling bearing fault diagnosis method based on maximum independence causal decoupling, and belongs to rolling bearing cross-equipment fault diagnosis methods. According to the method, a deep causal learning model MICDAE is constructed, a maximum independence auto-encoder is adopted as a causal decoupling network, a Hilbert-Schmidt independence criterion is introduced into the encoder for extracting causal features to maximize independence between the extracted features and domain tags thereof, and the causal feature extraction capability of the encoder is improved. In addition, alignment based on NWD between non-causal features is introduced in training of the MICDAE network, and the generalization ability of the model is further improved by fully utilizing the decoupling features. An original vibration signal of the rolling bearing is directly used as input in the test stage of the causal decoupling model, so that the end-to-end characteristic is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a cross-equipment fault diagnosis method for rolling bearings based on domain generalization technology aligned with the maximum independence causal decoupling self-encoder (MICDAE) and the non-causal nuclear norm Wasserstein distance (NWD). Background Technology

[0002] Rolling bearings are among the most critical components in rotating machinery systems, and their health is closely related to the safe and stable operation of the equipment. However, due to their frequent operation under high loads and high speeds, rolling bearings are prone to various failures. If potential failures in rolling bearings are not detected in time, the entire equipment operation will face the risk of failure, potentially leading to equipment downtime or damage, and even personal injury or death, posing a serious threat to life and property safety.

[0003] Currently, deep learning methods are widely used in bearing fault diagnosis, achieving high accuracy. However, traditional deep learning methods rely on sufficient labeled fault samples and the assumption that training and test samples follow the same distribution. In real-world scenarios, rolling bearing operating conditions are varied and unpredictable, resulting in significant differences between the distributions of test and training samples, which severely reduces the diagnostic accuracy of traditional deep learning models. To address this issue, fault diagnosis methods based on domain generalization (DG) have emerged. DG can train a model using a limited number of samples (the source domain) under a finite number of operating conditions. This model can then generalize well to any unknown operating condition (the target domain), overcoming the difference in data distribution between the two domains, i.e., domain shift. Furthermore, compared to artificially induced rolling bearing faults in experimental platforms, bearing faults in actual equipment occur less frequently, and labeled fault samples are scarce, often insufficient to support DG on that equipment. Therefore, cross-equipment DG is typically performed using the experimental platform dataset and a small amount of visible labeled data from actual equipment as the source domain, and the unknown operating conditions of the actual equipment as the target domain. Cross-device data generation (DG) based on causal learning is a common approach. It reveals the causal relationship between fault data and fault types by constructing a structural causal model. Due to the consistency of the rolling bearing fault generation mechanism, this causal relationship is invariant across rolling bearings in different devices. Therefore, common fault features across devices can be defined as causal features, while features specific to each device can be defined as non-causal features. A deep learning network constrained by causal aggregation loss can decouple these two aspects. However, current causal mechanism-based methods have shortcomings in causal feature decoupling and generalization performance, resulting in low stability and efficiency when analyzing complex fault types and operating conditions.

[0004] To address the aforementioned issues, this invention constructs a deep causal learning model, MICDAE, employing a maximum independence autoencoder as the causal decoupling network. The Hilbert-Schmidt independence criterion (HSIC) is introduced into the encoder extracting causal features to maximize the independence between the extracted features and their domain labels, thereby improving its ability to extract causal features. Furthermore, alignment based on non-causal feature dimensionality (NWD) is introduced during the training of the MICDAE network, fully utilizing the decoupled features to further enhance the model's generalization ability. In the testing phase, this causal decoupling model directly uses the original vibration signal of the rolling bearing as input, thus exhibiting end-to-end characteristics.

[0005] In summary, cross-equipment fault diagnosis of rolling bearings based on MICDAE and non-causal NWD alignment is an innovative research problem with significant research and application value. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of current cross-domain generalization methods based on causal mechanisms in terms of causal feature decoupling and generalization performance when applied to rolling bearing fault diagnosis. To address these issues, a causal decoupling network based on a maximum independence autoencoder is innovatively constructed, and non-causal NWD alignment is introduced, providing a high-accuracy, end-to-end rolling bearing fault diagnosis method with practical application value.

[0007] like Figure 3 As shown, the steps of this method are as follows:

[0008] S1 collects source domain data;

[0009] S2 uses the sliding window method to divide the collected rolling bearing vibration signals into samples, generating several one-dimensional time series samples.

[0010] The samples are labeled with class and domain labels according to the fault and operating condition of the signal, which will be used for the training of the deep learning model.

[0011] S3 inputs source domain samples into the model in batches, and uses a causal encoder and a domain encoder to decouple the samples causally, obtaining causal and non-causal features. The two feature vectors are then concatenated and input into the decoder to obtain the reconstructed samples. Input the causal features into the classifier to obtain the classification result. Inputting non-causal features into the classifier yields...

[0012] S4 calculates the causal aggregation loss based on causal and non-causal features and their class labels, and is derived from the input samples and... Calculate the reconstruction loss, and calculate the independence loss from the causal features and their domain labels, and from the class labels of the causal features and their domain labels. Calculate the classification loss by Calculate the NWD alignment loss and update the model parameters based on these losses;

[0013] S5 takes test set samples from unknown operating conditions as input and tests the trained model to obtain the fault diagnosis accuracy.

[0014] Return to S3 until the number of iterations reaches the preset maximum value, then proceed to S6;

[0015] S6 selects the model with the highest accuracy as the optimal model to complete the cross-equipment fault diagnosis task for rolling bearings.

[0016] This invention overcomes the distribution differences caused by varying equipment and operating conditions by training a high-generalization-performance deep transfer learning model using dual-condition data from both the same and different equipment, thus achieving fault diagnosis under unknown operating conditions. In practical applications on real equipment, it can utilize a sufficient amount of labeled data from experimental platforms as an aid, while employing a limited amount of known labeled data from actual equipment to train a model capable of generalizing to fault diagnosis under unknown operating conditions. Furthermore, the performance of this invention was verified by comparing it with other domain generalization methods on the task of cross-equipment rolling bearing fault diagnosis; the experimental results are available in […]. Figure 4 The average accuracy rate for each diagnostic task can reach 95.89%. Attached Figure Description

[0017] Figure 1 This is a detailed schematic diagram of MICDAE.

[0018] Figure 2 This provides detailed configuration instructions for the MICDAE network architecture.

[0019] Figure 3 This is a flowchart of the present invention.

[0020] Figure 4 This presents the experimental results comparing the present invention with other domain generalization models.

[0021] Figure 5 The results are from ablation experiments targeting the loss function. Detailed Implementation

[0022] This invention proposes a cross-equipment fault diagnosis method for rolling bearings based on MICDAE and non-causal NWD alignment, which is described below in conjunction with the appendix. Figure 1 , 2 Sections 3 and 4 detail specific embodiments of the present invention:

[0023] Step S1: Collect source domain data. Using an accelerometer, collect a rolling bearing dataset (referred to as source domain one) from the experimental platform. Combine this with various fault vibration signals of the rolling bearing collected from the device under diagnostic conditions (referred to as source domain two), which will be used for subsequent model training.

[0024] Step S2, sample partitioning. To obtain a sufficient number of fault samples for subsequent model training, the sliding window method is used to partition the one-dimensional time-series fault signal y of the rolling bearing in source domain one and source domain two: the sliding window size is set to be equal to the length l of the required samples, so that the sliding window slides from left to right on y with a step s, and the signal intercepted in the sliding window after each slide is taken as a sample. If the length of the signal is L, the number of samples N is calculated by equation (1):

[0025]

[0026] Based on the required N and l, the corresponding step size s can be calculated. Finally, according to the fault type and operating condition type of the rolling bearing, class labels and domain labels are added to several samples obtained by the sliding window method.

[0027] Step S3 involves inputting samples from source domain one and source domain two into the model in batches for feature extraction and classification. The model structure proposed in this invention can be found in the appendix. Figure 1 It consists of four parts: causal encoder E C Domain encoder E D Decoder D and classifier F C E C E D They are convolutional autoencoders with the same structure, and the samples are input into E. C E D Decoupling is performed to obtain causal and non-causal features separately. Based on the principle of causal learning, under the consistent causal mechanism followed by rolling bearing faults, causal features should include fault information that remains unchanged between rolling bearings in different devices (i.e., between domains), while non-causal features should include domain offset information unrelated to this causal mechanism. The decoder D is a deconvolutional network of the encoder, with encoder E... C E D The outputs of classifier F are concatenated along the channel dimension and used as input to reconstruct the classifier, outputting the reconstructed sample. C It is a multilayer perceptron composed of fully connected layers, which performs the final classification of faults based on causal feature vectors containing fault-related information.

[0028] Detailed configuration of the network architecture of the deep learning model proposed in this invention can be found in the appendix. Figure 2 Two encoders E C E DThe structure is identical, consisting of five batch-normalized one-dimensional convolutional layers (denoted as Conv1d+BN in the diagram): the first convolutional layer uses wide-kernel convolution to reduce high-frequency noise interference and extract fault features in the low-frequency band of the signal; the second to fifth convolutional layers are ordinary convolutional layers, with the kernel size decreasing layer by layer, to gradually capture detailed features through convolution operations. Each convolutional layer has a one-dimensional max-pooling layer (Maxpool1d) to downsample the features and reduce computation. The decoder D consists of deconvolutional layers (Conv1d+BN) corresponding to the above five convolutional layers, and each layer is preceded by an inverse operation of the max-pooling layer for upsampling (Unsample) to ensure that the reconstructed samples are the same size as the input samples. Classifier F C It consists of two fully connected (Linear) layers that map feature vectors to a specific number of fault categories. The Output Size column in the table represents the size of the output tensor of this layer: batch size × channel size × timing length.

[0029] Step S4: Based on the extracted features, classification results, and sample class and domain labels, calculate the loss function, and update the model parameters with the goal of minimizing the loss function. Let the vector set in the same batch be... Let represent 2n input samples from two domains, where d represents the domain label of the sample, i is the sample index, and its maximum value is the batch size n. Then the causal features of this vector set are: Non-causal characteristics Reconstructed Samples Where || represents the vector concatenation operation; the classifier output The calculation process for each part of the total loss function is as follows:

[0030] Step S401, calculate the causal aggregation loss L ca Due to the complexity of the components of rolling bearing vibration signals, decoupling causal and non-causal features is difficult. Therefore, a causal aggregation loss is used in the causal encoder E... C He Domain Encoder E D Regularization is applied. To ensure that causal and non-causal features fully contain their respective components while reducing irrelevant components, the causal aggregation loss aims to promote the aggregation of causal features of samples with the same fault class, while increasing the distance between causal features of samples with different classes; for non-causal features, it promotes aggregation within the domain and separation between domains. Based on this, the aggregation loss of causal features is calculated using equations (2) and (3) respectively. Aggregation loss with non-causal features

[0031]

[0032] In equation (2), These represent the i1th and i2th samples with domain labels d1 and d2, respectively. The causal characteristics, Indicates sample The class label is 1, and {·} represents the indicator function, which equals 1 if the condition is true and 0 otherwise. The first term of equation (2) represents the inner product and summation of the causal features of all domain pairs d1,d2 and sample pairs i1,i2, but only considers sample pairs with the same class label; the second term only considers causal feature pairs with different class labels. Thus, by minimizing this loss function, the first term maximizes the aggregation of causal features of the same class, and the second term minimizes the aggregation of causal features of different classes. In equation (3), The non-causal features of the i1th and i2th samples with the same domain label d are represented. Let d1 and d2 represent the non-causal features of the i1th and i2th samples with domain labels d1 and d2, respectively. The first term of equation (3) calculates the inner product of non-causal feature pairs with the same domain label and sums them to maximize the aggregation degree of non-causal features within the same domain; the second term takes non-causal feature pairs with different domain labels to minimize the aggregation degree of non-causal features across different domains. Finally, as shown in equation (4), the two terms are added together to obtain the total causal aggregation loss L. ca :

[0033]

[0034] Step S402, calculate the reconstruction loss L r In MICDAE, the reconstruction loss is used to constrain the training of the autoencoder network by minimizing the difference between the input sample and the reconstructed sample, thereby reducing the information loss of the model and improving its noise resistance. The reconstruction loss is calculated by equation (5):

[0035]

[0036] Where ||·||² represents the L2 norm. The L2 norm of the vector difference measures the difference between two vectors; the smaller the absolute value of the L2 norm, the smaller the difference. Step S403: Calculate the independence loss L. hsic Causal encoder E C To learn domain-independent features from multiple domains, the dependency between extracted causal features and their corresponding domain labels should be minimized as much as possible. HSIC is used as a metric to measure this dependency and is integrated into the causal encoder, constraining the training process by minimizing it. For variables X and Y, HSIC is the Hilbert-Schmidt norm of the covariance of their projections onto the high-dimensional feature space; a smaller HSIC indicates a weaker dependency. For causal features... We map it to the regenerating kernel Hilbert space with the domain label d, and the independence loss can be calculated by the biased empirical estimate of HSIC, as shown in Equation (6):

[0037]

[0038] Where tr(·) is the trace of the matrix. (representing a 2n×2n real matrix) represent Gaussian kernel matrices for causal features and domain labels, respectively; Let I be the centered matrix, where I represents the identity matrix, 1 2n 1 2n T This represents a 2n×2n all-one matrix.

[0039] Step S404, calculate the classification loss L cl Alignment loss L with NWD nwd The causal features, after encoding, already contain sufficient fault-related information and are input into F. C Fault classification is performed. This is done to train the model's classification ability and reduce causal features. Classification results With class tags The difference between them is taken as the objective, and the cross-entropy loss is calculated as the classification loss, as shown in Equation (7):

[0040]

[0041] To address the non-causal components in causal decoupling, this invention proposes a Non-Causal Feature (NWD) alignment mechanism. While training the model's classification ability, alignment reduces the differences between extracted features from different domains, enabling the model to learn inter-domain invariant features and thus achieve transfer learning. Non-causal features of samples from different domains. The domain-related differences contained within the features are more pronounced than those of undecoupled features, excluding domain-independent causal features, making them more suitable for alignment training. This invention introduces NWD, which is output through a classifier. nuclear norm number measurement The difference is addressed by reducing the NWD between source domain one and source domain two (domain labels d=1,2 respectively), thereby further enhancing the domain generalization performance of the causal decoupling network. The NWD alignment loss between source domain data is shown in Equation (8):

[0042]

[0043] in‖·‖ * This represents the nuclear norm.

[0044] Step S405: Calculate the total loss, which serves as the objective function for model training, and update the model parameters. Combining equations (2) to (8), the total loss function can be obtained, as shown in equation (9):

[0045] Loss = L cl +αL nwd +γL r +γL ca +λL HSIC (9)

[0046] α, β, γ, and λ are weight parameters used to balance the proportion of each loss function. They are adjusted according to the characteristics of the dataset during the experiment to achieve the best training results.

[0047] During the training process of this method, the Adam optimizer is used to iteratively update the model parameters. It adaptively adjusts the learning rate of each parameter by calculating the first and second moments of the gradient of the loss function in order to efficiently minimize the loss function.

[0048] Step S5: Test the model. Input test set samples from the target domain (unknown operating conditions), and output the diagnostic results of the fault types as the performance index of this model. After testing, return to S3 to continue training until the number of training iterations reaches the preset maximum value.

[0049] Step S6: Select the model with the highest accuracy as the optimal model to complete the cross-equipment fault diagnosis task of rolling bearings.

[0050] To verify the performance of this invention and its improvement in the loss function, we conducted comparative experiments with other domain generalization methods and ablation experiments on the loss function of this model. The publicly available Case Western Reserve University bearing dataset (hereinafter referred to as the CWRU dataset) was used as the source domain. It provides ample labeled data for four operating conditions: A (load: 0 hp, speed: 1797 rpm), B (1 hp, 1772 rpm), C (2 hp, 1750 rpm), and D (3 hp, 1730 rpm). Fault state types include normal (N), inner race fault (I), outer race fault (O), and ball fault (B). The dataset collected from the mechanical failure simulation test bench manufactured by SpectraQuest (hereinafter referred to as the MFS dataset) was used as the target actual equipment. This dataset can simulate common mechanical equipment failures and is used for fault diagnosis research. The faulty bearing model used was KR-12K, with an inner diameter of 0.75 inches and an outer diameter of 1.8501 inches. To support multi-device fault diagnosis testing, we collected bearing fault data under four different operating conditions: O (1200 rpm), P (1800 rpm), Q (2400 rpm), and R (2400 rpm). By replacing different faulty bearings, we simulated the same four fault types (N, I, O, and B) using a test bench and acquired actual bearing vibration signals. We used the P-condition data from the MFS dataset as the second source domain to simulate a small amount of labeled data under known operating conditions on the actual equipment; the other three operating conditions were used as the target domain to simulate data under unknown operating conditions on the actual equipment, and were not used for training but for model testing.

[0051] In the data partitioning stage S2, we set the sliding window size (i.e. the time length of the sample) l = 1024, the number of samples in each working condition N = 800, and combined with the length L of the original time signal, we calculate the sliding window step size D according to formula (1).

[0052] For the hyperparameters of the training phases S3 to S4 of the proposed model, after preliminary experiments, we set them to the values ​​most conducive to model training. The batch size of the two source domain input samples was set to 64; the learning rate was set to 0.0002; and the maximum number of iterations was set to 100. In order to reasonably set the weight parameters α, β, γ and λ of each part of the loss function in equation (9), we analyzed the impact of the changes of each parameter on the model performance through multiple experiments, and finally set each parameter in the range with the highest model diagnostic accuracy, with values ​​of α = 0.5, β = 1, γ = 1, and λ = 0.5.

[0053] To verify the superior performance of the rolling bearing cross-equipment fault diagnosis method of the present invention compared with other domain generalization methods, as shown in the appendix... Figure 4As shown, comparative experiments were conducted on 12 cross-device transfer learning tasks. The experimental results demonstrate that the proposed method outperforms the comparative methods in terms of average diagnostic accuracy across all target domains, proving its superiority.

[0054] To verify the effectiveness of this invention in loss function design, ablation experiments were conducted. The experimental results are attached. Figure 5 As shown, we conducted experiments with data O and R as the target domains, and compared the diagnostic accuracy of the model using the full loss function with that of the model using a loss function with one part removed but with the same structure. This further confirms that each part of the loss function makes a significant contribution to improving model performance.

[0055] In summary, this method can complete the task of cross-equipment fault diagnosis of rolling bearings using one-dimensional timing signals with good accuracy.

Claims

1. A method for diagnosing rolling bearing faults based on maximum independence causal decoupling, characterized in that: S1 collects source domain data; S2 uses the sliding window method to divide the collected rolling bearing vibration signals into samples, generating several one-dimensional time series samples. It also adds class and domain labels to the samples according to the fault and working condition types of the signals for subsequent training of deep learning models. S3 inputs source domain samples into the model in batches, and uses a causal encoder and a domain encoder to decouple the samples causally, obtaining causal and non-causal features. The two feature vectors are then concatenated and input into the decoder to obtain the reconstructed samples. Input the causal features into the classifier to obtain the classification result. Inputting non-causal features into the classifier yields... S4 calculates the causal aggregation loss based on causal and non-causal features and their class labels, and is derived from the input samples and... Calculate the reconstruction loss, and calculate the independence loss from the causal features and their domain labels, and from the class labels of the causal features and their domain labels. Calculate the classification loss by Calculate the NWD alignment loss and update the model parameters based on these losses; S5 takes a test set sample from an unknown working condition as input to test the trained model, obtains the fault diagnosis accuracy, and returns to S3. This process continues until the number of iterations reaches the preset maximum value, then proceeds to S6. S6 selects the model with the highest accuracy as the optimal model to complete the cross-equipment fault diagnosis task for rolling bearings.

2. The method according to claim 1, characterized in that: Step S1: Collect source domain data; collect rolling bearing datasets from the experimental platform using an accelerometer, denoted as source domain one, and collect various fault vibration signals of the rolling bearings from the device to be diagnosed under certain working conditions, denoted as source domain two.

3. The method according to claim 2, characterized in that: Step S2, divide the samples; use the sliding window method to divide the one-dimensional timing fault signal y of the rolling bearing in source domain one and source domain two: set the sliding window size to be equal to the length l of the required sample, so that the sliding window slides from left to right on y with a step s, and take the signal intercepted in the sliding window after each slide as a sample; if the length of the signal is L, the number of samples N is calculated by equation (1): Based on the required N and l, the corresponding step size s is calculated; finally, according to the fault type and operating condition type of the rolling bearing, class labels and domain labels are added to several samples obtained by the sliding window method.

4. The method according to claim 2, characterized in that: Step S3: Input samples from source domain 1 and source domain 2 into the model in batches for feature extraction and classification; the model consists of four parts: causal encoder E C Domain encoder E D Decoder D and classifier F C E C E D They are convolutional autoencoders with the same structure, and the samples are input into E. C E D Decoupling is performed to obtain causal and non-causal features respectively. Based on the principle of causal learning, under the consistent causal mechanism followed by rolling bearing faults, causal features should include fault information that remains unchanged between rolling bearings in different devices (i.e., between domains), while non-causal features should include domain offset information unrelated to this causal mechanism. Decoder D is the deconvolutional network of the encoder, with encoder E... C E D The output of classifier F is concatenated along the channel dimension and used as input to reconstruct the reconstructed sample; classifier F C It is a multilayer perceptron composed of fully connected layers, which performs the final classification of faults based on causal feature vectors containing fault-related information; Network architecture of deep learning models: two encoders E C E D The structure is the same, consisting of five batch-normalized one-dimensional convolutional layers: the first convolutional layer uses a wide kernel; the second to fifth convolutional layers are ordinary convolutional layers, with the kernel size decreasing layer by layer, to gradually capture detailed features through convolution operations; each convolutional layer has a one-dimensional max-pooling layer to downsample the features and reduce computation; the decoder D consists of deconvolutional layers corresponding to the above five convolutional layers, and each layer is preceded by an inverse max-pooling layer for upsampling; the classifier F... C It consists of two fully connected layers that map feature vectors to a specific number of fault categories.

5. The method according to claim 2, characterized in that: Step S4: Based on the extracted features, classification results, and sample class and domain labels, calculate the loss function, and update the model parameters with the goal of minimizing the loss function; assuming a vector set in the same batch... Let d represent 2n input samples from two domains, where d represents the domain label of the sample, i is the sample index, and its maximum value is the batch size n; then the causal features of this vector set are: Non-causal characteristics Reconstructed Samples Where || represents the vector concatenation operation; the classifier output The calculation process for each part of the total loss function is as follows: Step S401, calculate the causal aggregation loss L ca Due to the complexity of the components of rolling bearing vibration signals, decoupling causal and non-causal features is difficult; therefore, causal aggregation loss is used in the causal encoder E... C He Domain Encoder E D Regularization is applied; to ensure that causal and non-causal features fully contain the corresponding components and reduce irrelevant components, the causal aggregation loss aims to promote the aggregation of causal features of samples of the same fault class, while increasing the distance between causal features of samples of different classes; for non-causal features, it promotes their aggregation within the domain and separation between domains; based on this, the aggregation loss of causal features is calculated by equations (2) and (3) respectively. Aggregation loss with non-causal features In equation (2), These represent the i1th and i2th samples with domain labels d1 and d2, respectively. The causal characteristics, Indicates sample The class label, 1{·} represents the indicator function, which is equal to 1 if the condition is true, otherwise equal to 0; the first term of equation (2) represents the inner product and summation of the causal features of all domain pairs d1,d2 and sample pairs i1,i2, but only considers sample pairs with the same class label; the second term only considers causal feature pairs with different class labels; thus, by minimizing this loss function, the first term maximizes the aggregation degree of causal features of the same class, and the second term minimizes the aggregation degree of causal features of different classes; in equation (3), The non-causal features of the i1th and i2th samples with the same domain label d are represented. Let i and j represent the non-causal features of the i1th and i2th samples with domain labels d1 and d2, respectively. The first term of equation (3) calculates the inner product of non-causal feature pairs with the same domain label and sums them to maximize the aggregation degree of non-causal features within the same domain; the second term calculates the aggregation degree of non-causal features with different domain labels to minimize the aggregation degree of non-causal features across different domains; finally, as shown in equation (4), the two terms are added together to obtain the total causal aggregation loss L. ca : Step S402, calculate the reconstruction loss L r In MICDAE, the reconstruction loss is used to minimize the difference between the input sample and the reconstructed sample to constrain the training of the autoencoder network, thereby reducing the information loss of the model and improving its noise resistance; the reconstruction loss is calculated by equation (5): Where ||·||2 represents the L2 norm; the L2 norm of the difference between vectors can be used to measure the difference between the two vectors. The smaller the absolute value of the L2 norm, the smaller the difference between the two vectors. Step S403, calculate the independence loss L hsic ; Causal encoder E C To learn domain-independent features from multiple domains, the dependency between the extracted causal features and their corresponding domain labels should be minimized as much as possible. HSIC is used as a metric to measure this dependency and is integrated into the causal encoder, constraining the training process by minimizing it. For variables X and Y, HSIC is the Hilbert-Schmidt norm of the covariance of their projections onto the high-dimensional feature space; the smaller the HSIC, the weaker the dependency between them. For causal features... Mapping it to the regenerating kernel Hilbert space with the domain label d, the independence loss can be calculated by the biased empirical estimate of HSIC, as shown in Equation (6): Where tr(·) is the trace of the matrix. (representing a 2n×2n real matrix) represent Gaussian kernel matrices for causal features and domain labels, respectively; Let I be the centered matrix, where I represents the identity matrix, 1 2n 1 2n T Represents a 2n×2n all-one matrix; Step S404, calculate the classification loss L cl Alignment loss L with NWD nwd The causal features, after being encoded, already contain sufficient fault-related information and are input into F. C Fault classification is performed; to train the model's classification ability and reduce causal features. Classification results With class tags The difference between them is taken as the objective, and the cross-entropy loss is calculated as the classification loss, as shown in Equation (7): To address the non-causal component in causal decoupling, a Non-Causal Feature (NWD) alignment mechanism is proposed. While training the model's classification ability, alignment reduces the differences between extracted features from different domains, enabling the model to learn inter-domain invariant features and thus achieve transfer learning. This mechanism is used to align non-causal features from samples in different domains. The domain-related differences contained within the features are more pronounced than those of undecoupled features, excluding domain-independent causal features, making them more suitable for alignment training; NWD is introduced, and the classifier outputs... nuclear norm number measurement The difference is addressed by reducing the NWD between source domain one and source domain two (domain labels d=1,2 respectively) to further enhance the domain generalization performance of the causal decoupling network; the NWD alignment loss between source domain data is shown in Equation (8): in‖·‖ * Represents the nuclear norm; Step S405: Calculate the total loss as the objective function for model training and update the model parameters; combining equations (2) to (8), the total loss function can be obtained, as shown in equation (9): Loss=L cl +αL nwd +βL r +γL ca +λL HSIC (9) Where α, β, γ, and λ are weight parameters used to balance the weight of each loss function; The Adam optimizer is used to iteratively update the model parameters. It adaptively adjusts the learning rate of each parameter by calculating the first and second moments of the gradient of the loss function in order to efficiently minimize the loss function.

6. The method according to claim 1, characterized in that: Step S5, test the model; input test set samples from the target domain, output the diagnostic results of the fault type as the performance index of this model; after testing, return to S3 to continue training until the number of training iterations reaches the preset maximum value; Step S6: Select the model with the highest accuracy as the optimal model to complete the cross-equipment fault diagnosis task of rolling bearings.