Rolling bearing fault diagnosis method based on variable exploration invariant learning

Through the variable exploration invariant learning (VEIL) method, the variable exploration module and the invariant learning loss function are constructed, and the invariant parameters are optimized, which solves the problem of low accuracy caused by the distribution difference in rolling bearing fault diagnosis under variable working conditions, and achieves high-precision fault diagnosis.

CN120561771APending Publication Date: 2025-08-29SICHUAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510707682.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Under variable working conditions, in rolling bearing fault diagnosis, the distribution difference between the source domain samples and the target domain samples is large, resulting in a low accuracy rate of fault diagnosis, which is difficult for the existing technology to effectively solve this problem.

Method used

The variable exploration invariant learning (VEIL) method is adopted, and the model is sparsely trained by constructing the variable exploration module, distinguishing and optimizing the invariant parameters and variable parameters, constructing the invariant learning loss function, and using the SAM optimizer to optimize the loss function, and constructing a subnet with only invariant parameters for troubleshooting.

Benefits of technology

It improves the model's external generalization performance and the accuracy of fault diagnosis, and can perform high-precision fault diagnosis on target domain samples of unknown operating conditions under variable operating conditions, reducing the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561771A_ABST
    Figure CN120561771A_ABST
Patent Text Reader

Abstract

The invention discloses a rolling bearing fault diagnosis method based on variable exploration invariant learning, and the method comprises the following steps: inputting different distributed source domain samples into an initial parameter model, and carrying out the sparse training of the model through a mask; training samples extracted from a plurality of different distributions are input into the model, invariant parameters and variable parameters which are distinguished by masks are optimized, and the variable exploration process is completed. Taking a parameter model which only contains an invariant parameter after variable exploration as a sub-network to carry out invariant learning, constructing a loss function of invariant learning, inputting a training sample into an invariant learning network to obtain a cross entropy of a prediction probability as a main loss function, and introducing an invariant constraint to guide the model to learn an invariant feature; and optimizing the invariant learning loss function by using an SAM optimizer. And inputting a to-be-tested sample with unknown distribution into the trained VEIL network to obtain a prediction label of the to-be-tested sample, so that the fault diagnosis process of the rolling bearing is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rolling bearing fault diagnosis, and in particular relates to a rolling bearing fault diagnosis method based on variable exploration invariant learning. Background Art

[0002] Rolling bearings are critical tribological components used to reduce friction and ensure smooth operation. They are widely used in nearly all types of rotating machinery, such as transmissions, wind turbines, rolling mills, and train cars. The proper operation of rotating machinery depends largely on the health of rolling bearings. As core components of mechanical systems, rolling bearings have a significantly higher probability of failure than other transmission components when subjected to extreme operating conditions such as high temperatures, high speeds, and heavy loads. Industrial data indicates that nearly half of mechanical system failures are caused by bearing anomalies, highlighting the importance and urgency of monitoring the health of rolling bearings. To reduce maintenance costs and avoid casualties, timely and accurate fault diagnosis is essential before immeasurable losses occur.

[0003] However, in real-world applications, monitoring equipment often operates under varying conditions, such as changes in load, speed, and temperature. These variations lead to significant distribution differences between source and target domain samples. Source domain samples are typically acquired under specific operating conditions, while target domain samples are acquired under different operating conditions. This distribution difference negatively impacts the performance of fault diagnosis models, as the features and patterns trained on the source domain may not be applicable to the target domain. This can lead to high false positives or missed diagnoses in real-world applications, impacting the safety and reliability of equipment. Therefore, research on how to reduce the distribution difference between source and target domain samples, or employing techniques such as transfer learning to improve the generalization capabilities of fault diagnosis models, is an important and pressing issue in rolling bearing fault diagnosis. Overcoming these challenges can achieve more accurate fault identification and prediction, ensuring the proper operation of equipment.

[0004] In order to solve the problem of low fault diagnosis accuracy caused by the large difference in distribution of source domain samples and target domain samples under variable working conditions, a large number of related studies have emerged in recent years. For example, Fang et al. proposed an intelligent fault diagnosis method for rolling bearings based on deep transfer learning. They used a sliding window to preprocess the original time domain vibration signal to construct a sample dataset, and adopted a deep residual shrinkage network to filter out sample noise and automatically extract data features. Finally, the marginal distribution and conditional distribution between domains were aligned through the maximum mean difference and the local maximum mean difference, weakening the difference in data feature distribution between different working conditions. Xu et al. proposed an intelligent fault diagnosis method for bearings based on time-varying order spectrum and multi-scale domain adaptive network. They used the good noise robustness and high time-frequency aggregation characteristics of synchronous compressed wave packet transform to eliminate the influence of speed fluctuations and reduce the distribution offset of bearing data under time-varying speed. They also established a global-local feature fusion model to extract domain-invariant features closely related to the bearing fault state from the time-varying order spectrum. Li et al. developed a rolling bearing fault diagnosis method based on multi-domain information fusion and multi-kernel maximum mean difference. They constructed a multi-domain information feature extractor based on a multi-attention mechanism to capture important features of vibration signals in different transform domains. They then performed feature fusion and used MK-MMD to reduce the distribution distance between the source and target domains to obtain domain-invariant features. These methods are all based on difference-based domain adaptation, using various means to reduce the distribution difference between the source and target domains, thereby improving the adaptability and generalization performance of the model. However, these methods are often accompanied by increased computational complexity and parameter adjustment challenges, which affect the accuracy of fault diagnosis.

[0005] In response to these problems, the present invention proposes a distribution out-generalization method - Variable Exploration Invariant Learning (VEIL), which is used to solve the problem of large distribution differences between source domain samples and target domain samples in rolling bearing fault diagnosis under variable working conditions. In the proposed VEIL, distribution knowledge is used to find parameters that are sensitive to distribution changes (i.e., variable parameters), and relatively stable parameters in different distributions are regarded as invariant parameters, so as to find a sub-network that is resistant to distribution changes; then invariant learning is used to train the sub-network to optimize the invariant parameters, so that the model learns invariant features and improves the distribution out-generalization performance. The above characteristics of VEIL enable it to achieve high-precision rolling bearing fault diagnosis when faced with the problem of large distribution differences between source domain samples and target domain samples under variable working conditions. Summary of the Invention

[0006] The purpose of the present invention is to solve the above problems and provide a method for pruning the model by constructing a variable exploration module to obtain a sub-network containing only invariant parameters, which greatly reduces the computational complexity. The sub-network is then used for invariant learning to enable the model to learn invariant features, thereby solving the problem of low rolling bearing fault diagnosis accuracy caused by large differences in the distribution of source domain samples and target domain samples under variable working conditions, thereby realizing a rolling bearing fault diagnosis method based on variable exploration invariant learning that can accurately and stably diagnose rolling bearing faults.

[0007] To solve the above technical problems, the technical solution of the present invention is: a rolling bearing fault diagnosis method based on variable exploration invariant learning, comprising the following steps:

[0008] S1. First, source domain samples with different distributions are input into the initial parameter model, and the model is trained in a sparse manner using masks, which effectively reduces the complexity of the model, improves training efficiency and computational efficiency, and enhances the robustness of the model.

[0009] S2. Then, the training samples extracted from multiple different distributions are input into the model, and the invariant parameters and variable parameters separated by the mask are optimized. The optimization process uses the Adam optimizer to calculate the gradient, and uses this to sort the degree to which the parameters are affected by the distribution. The invariant parameters with the least relevance to the classification task are eliminated, and the variable parameters with the least influence of the distribution are used as new invariant parameters, thereby achieving the purpose of dynamically updating the mask until the parameter optimization target is minimized and the variable exploration process is completed;

[0010] S3. Then, the parameter model containing only the invariant parameters after completing the variable exploration is used as a sub-network for invariant learning. The loss function of invariant learning is constructed. The cross entropy of the predicted probability obtained by inputting the training samples into the invariant learning network is used as the main loss function. Invariance constraints are introduced to guide the model to learn invariant features.

[0011] S4. Use the SAM optimizer to optimize the invariant learning loss function to minimize the loss and further improve the out-of-distribution generalization performance of the model;

[0012] S5. Input the unknown distribution of the test sample into the trained VEIL network to obtain the predicted label of the test sample, thus completing the rolling bearing fault diagnosis process.

[0013] Furthermore, the variable exploration in S2 is to extract multiple data sets from n distributions, that is, ε={e1,…e m}, where each distribution Contains n samples and class labels Therefore, for each sample from distribution e, we can assign a distribution index And denote the data points as (x,y,d); In addition, let ε unseen is an unknown distribution, and the samples collected from it are used to evaluate the generalization performance of invariant learning; let is a parameterized model with parameters θ∈Θ. The features extracted by this model are expressed as The goal of variable exploration is to prune a sparse subnetwork from an over-parameterized model. During the pruning process, we need to find and exclude variable parameters that change with the distribution, and retain the invariant parameters that remain stable when the distribution changes, thereby improving the performance of out-of-distribution generalization.

[0014] Furthermore, in the process of exploring the variables in S2, in order to obtain a good initialization, the model will be sparse, and a binary mask (mask) m is usually applied for element-by-element multiplication, that is, m°θ. Such a mask m can be learned through optimization or obtained according to certain standards, such as sensitivity analysis, weight value, Fisher information, or even random initialization; by setting the sparsity ratio You can decide how many parameters to exclude from sparse training; choose Random Initialization to start variant exploration from a pre-trained model with a randomly initialized mask m.

[0015] Furthermore, during the variable exploration process in S2, there are two optimization objectives:

[0016]

[0017] Where θ inv and θ var It represents the invariant parameters and variable parameters distinguished by mask m, which are expressed as Equation (3) and Equation (4). h and g are two fully connected layers that map the extracted features to the class label space respectively. and distribution space Intuitively, the goal is a classification task, which attempts to make predictions based on label information, while Trying to distinguish each distribution based on distribution information is not helpful for subsequent invariant learning.

[0018] Furthermore, the optimization process in S2 includes the following sub-steps:

[0019] S21. Using the Adam (Adaptive Moment Estimation) optimizer, in order to avoid conflicts when optimizing constant parameters and variable parameters, two Adam optimizers are used, namely Adam1 and Adam2, to include constant parameters and variable parameters respectively; in addition, Adam1 will include the category prediction head h, while Adam2 will include the distribution prediction head g; in minimizing In the process, the gradient amplitude can be used To find the parameters most relevant to the loss function, that is, the constant parameters; similarly, by minimizing The gradient magnitude can be found Larger parameters, i.e. variable parameters, are sensitive to distribution information and cannot help extract invariant features; these parameters can then be sorted according to the size of the gradient to show the degree to which they are activated by the corresponding target;

[0020] S22. Next, in order to dynamically improve the entire variable exploration process, the mask m is updated after each iteration, and the invariant parameter with the smallest activation degree is excluded, while the variable parameter with the smallest activation degree is used as the invariant parameter. The specific process is as follows:

[0021]

[0022] Where ArgTopK(v,k) is used to find the index of the top k largest elements in the array or list v, and m[·] represents the index m. In addition, to determine how many parameters need to be swapped, the cosine annealing function is used, which is expressed as follows:

[0023]

[0024] Where t and T represent the current iteration round and the total iteration round respectively, and the hyperparameter α determines the maximum value. The new mask obtained by equations (5) and (6) finds parameters that are less affected by distribution information and more closely related to the learning task, further improving the performance of the model.

[0025] S23, will eventually and Once the optimization is completed, that is, the variable exploration is completed, the final invariant parameters can be used as the pruned sub-network for the subsequent invariant learning.

[0026] Furthermore, the invariant learning in S3 is a machine learning method that aims to improve the robustness of the model to various transformations by capturing the immutable features of the data. In classification tasks, invariant learning focuses on the properties of local or global features that remain unchanged in different environments or working conditions. Its goal is to enable the model to recognize and utilize these invariant features, thereby avoiding being sensitive to small changes in the input data.

[0027] Furthermore, the invariant learning in S3 includes the following steps: selecting a five-layer convolutional neural network as the network structure for invariant learning, because the structure of the convolutional neural network is naturally spatially invariant, and a pooling layer can be used to reduce the sensitivity to position changes. After obtaining the updated mask m, the invariant parameters are used as a subnetwork for invariant learning, and the invariant learning loss function is constructed:

[0028]

[0029] In the formula Represents the loss calculated to promote the model to learn invariant feature representations in different environments given the input x and label y. Indicates further processing of the output of the parameterized model f, converting the output into classification probability, is the cross entropy loss, and its specific expression is:

[0030]

[0031] In the formula It is a regularization loss that can be used to constrain the complexity of the model, thereby guiding the model to learn the invariant features, that is, the invariance constraint. The specific expression is:

[0032]

[0033] λ is a parameter used to control cross entropy loss and regularization loss. By adjusting λ, we can balance the model's ability to fit the training data and maintain model regularization.

[0034] Furthermore, the SAM optimizer is used in S4 to optimize the invariant learning loss function. Although SAM achieves good results in the same distribution generalization task, its performance in the out-of-distribution generalization task is quite limited. This is because when using SAM to optimize the out-of-distribution generalization task, the variable parameters are subjected to incorrect perturbations, which promotes the fitting of false features. Specifically, in the out-of-distribution generalization problem, the invariant features and false features correspond to the invariant parameters θ, respectively. inv and variable parameter θ var , applying perturbations to the invariant parameters can enhance the extraction of invariant features, while applying perturbations to the variable parameters will promote the combination of false features and label information; therefore, after eliminating the variable parameters and retaining only the invariant parameters, SAM can be used to optimize the loss function, and the optimization goal is:

[0035]

[0036] Furthermore, the SAM aims to calculate an optimal parameter perturbation In a neighborhood with a radius of ρ, the loss value can be maximized By applying ∈ * , loss change is called sharpness, which indicates the flatness of the learned loss function; intuitively, flatter loss functions usually exhibit better generalization properties, because a slight shift applied in the input space does not significantly change the loss value; To apply SAM to VEIL, one only needs to calculate the optimal perturbation ∈ * The updated mask is then applied to parameter perturbations, a process that not only removes variable parameters but also reduces the computational burden of SAM.

[0037] The beneficial effects of the present invention are:

[0038] 1. The rolling bearing fault diagnosis method based on variable exploration invariant learning provided by the present invention, in VEIL, uses a mask to perform sparse training on the initial model to reduce the complexity of the model, constructs a variable exploration module, uses a mask to distinguish variable parameters and invariant parameters, and then uses two Adam optimizers to optimize and replace the variable parameters and invariant parameters, and updates the mask. After completing the optimization goal, a sub-network containing only invariant parameters is obtained. The invariant parameters in such a sub-network are less affected by the distribution and have good generalization ability, which is also conducive to the subsequent invariant learning of invariant features that are not affected by the distribution.

[0039] 2. The VEIL of the present invention utilizes a subnetwork for invariant learning, constructs an invariant learning loss function, and introduces invariance constraints to guide the model to learn invariant representations. Finally, the SAM optimizer is used to optimize the invariant learning loss function. Although SAM's performance in out-of-distribution generalization tasks is quite limited, the variable exploration module eliminates variable parameters in the model, making its application in out-of-distribution generalization tasks feasible. SAM's perturbations on invariant parameters enhance the extraction of invariant features, thereby improving the model's out-of-distribution generalization capabilities.

[0040] 3. VEIL's aforementioned characteristics enable models trained using source domain samples from diverse operating conditions to achieve excellent classification and generalization performance. This allows VEIL to maintain high-precision fault diagnosis for target domain samples from unknown operating conditions, even when faced with significant distribution discrepancies between source and target domain samples under varying operating conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of the rolling bearing fault diagnosis method based on the VEIL method of the present invention;

[0042] Figure 2 This is a schematic diagram of the Adam optimizer principle of the present invention;

[0043] Figure 3 This is a physical picture of the Paderborn bearing data test bench of the present invention;

[0044] Figure 4 This is a physical picture of the accelerated life test device of the present invention;

[0045] Figure 5 It is a graph of CRIC values ​​when the present invention optimizes invariant learning using different methods;

[0046] Figure 6 is the ROC curve diagram of the normal sample of the present invention;

[0047] Figure 7 is the ROC curve diagram of the inner race fault sample of the present invention;

[0048] Figure 8 is the ROC curve diagram of the outer race fault sample of the present invention;

[0049] Figure 9 is the t-SNE plot of VEIL of the present invention;

[0050] Figure 10 is the t-SNE graph of the SVM of the present invention;

[0051] Figure 11 is the t-SNE diagram of MLDA of the present invention;

[0052] Figure 12 is the t-SNE diagram of the DTJM of the present invention;

[0053] Figure 13 t-SNE diagram of CAADG of the present invention.

[0054] Explanation of the accompanying symbols: 1. Electric motor; 2. Torque measurement shaft; 3. Rolling bearing test module; 4. Flywheel; 5. Load motor. DETAILED DESCRIPTION

[0055] The present invention will be further described below with reference to the accompanying drawings and specific embodiments:

[0056] like Figure 1 As shown, the rolling bearing fault diagnosis method based on variable exploration invariant learning provided by the present invention includes the following steps:

[0057] S1. First, source domain samples with different distributions are input into the initial parameter model, and the model is trained in a sparse manner using masks, which effectively reduces the complexity of the model, improves training efficiency and computational efficiency, and enhances the robustness of the model.

[0058] In this example, distribution refers to working conditions, and different distributions mean different working conditions. In the transfer learning scenario, source domain samples refer to data samples obtained from the source domain. The source domain refers to a field with rich annotated information or prior knowledge. These samples carry the characteristics, patterns, and knowledge of the field, that is, training samples. The initial parameter model used in this article is the Kaiming (He) initialization model commonly used in convolutional neural networks. Sparse training is a commonly used training technique in this field.

[0059] S2. Then, the training samples extracted from multiple different distributions are input into the model, and the invariant parameters and variable parameters distinguished by the mask are optimized. The optimization process uses the Adam optimizer to calculate the gradient, and uses this to sort the degree to which the parameters are affected by the distribution. The invariant parameters with the least correlation with the classification task are excluded, and the variable parameters with the least influence of the distribution are used as new invariant parameters, thereby achieving the purpose of dynamically updating the mask until the parameter optimization target is minimized and the variable exploration process is completed.

[0060] The variable exploration in step S2 is to extract multiple data sets from m distributions, that is, ε = {e1, ...e m}, where each distribution Contains n samples and class labels Therefore, for each sample from distribution e, we can assign a distribution index And denote the data points as (x,y,d); In addition, let ε unseen is an unknown distribution, and the samples collected from it are used to evaluate the generalization performance of invariant learning. is a parameterized model with parameters θ∈Θ. The features extracted by this model are expressed as The goal of variable exploration is to prune a sparse subnetwork from an over-parameterized model. During the pruning process, we need to find and exclude variable parameters that change with the distribution, and retain the invariant parameters that remain stable when the distribution changes, thereby improving the performance of out-of-distribution generalization. Represents space, d represents distribution index, belonging to Used to identify which distribution the sample comes from. represents an m-dimensional vector space, f θ Represents a parameterized model, and the mapping relationship is That is, input one-dimensional feature x, output D-dimensional feature representation, θ represents the model parameters, belonging to the parameter space Θ, Z: the feature representation extracted by the model, belonging to That is, M-dimensional feature vector.

[0061] During the variable exploration process in step S2, in order to obtain a good initialization, the model is sparse, usually applying a binary mask (mask) m to perform element-by-element multiplication, that is, Such a mask m can be learned through optimization or obtained according to certain criteria, such as sensitivity analysis, weight value, Fisher information, or even random initialization; by setting the sparsity ratio You can decide how many parameters to exclude from sparse training; choose Random Initialization to start variant exploration from a pre-trained model with a randomly initialized mask m.

[0062] During the variable exploration process in S2, there are two optimization goals:

[0063]

[0064] Where θ inv and θ var It represents the invariant parameters and variable parameters distinguished by mask m, which are expressed as Equation (3) and Equation (4). h and g are two fully connected layers that map the extracted features to the class label space respectively. and distribution space Intuitively, the goal is a classification task, which attempts to make predictions based on label information, while Trying to distinguish each distribution based on distribution information is not helpful for subsequent invariant learning.

[0065] The optimization process in step S2 includes the following sub-steps:

[0066] S21. Using the Adam (Adaptive Moment Estimation) optimizer, in order to avoid conflicts when optimizing constant parameters and variable parameters, two Adam optimizers are used, namely Adam1 and Adam2, to include constant parameters and variable parameters respectively. In addition, Adam1 will include the category prediction head h, while Adam2 will include the distribution prediction head g; in minimizing In the process, the gradient amplitude can be used To find the parameters most relevant to the loss function, i.e., the constant parameters. Similarly, by minimizing The gradient magnitude can be found Larger parameters, i.e. variable parameters, are sensitive to distribution information and cannot help extract invariant features. These parameters can then be sorted according to the size of the gradient to show the degree to which they are activated by the corresponding target.

[0067] S22. Next, in order to dynamically improve the entire variable exploration process, the mask m is updated after each iteration, and the invariant parameter with the smallest activation degree is excluded, while the variable parameter with the smallest activation degree is used as the invariant parameter. The specific process is as follows:

[0068]

[0069] Where ArgTopK(v,k) is used to find the index of the top k largest elements in the array or list v, and m[·] represents the index m. In addition, to determine how many parameters need to be swapped, the cosine annealing function is used, which is expressed as follows:

[0070]

[0071] Where t and T represent the current iteration round and the total iteration round respectively, and the hyperparameter α determines the maximum value. The new mask obtained by equations (5) and (6) finds parameters that are less affected by distribution information and more closely related to the learning task, further improving the performance of the model.

[0072] S23, will eventually and Once the optimization is completed, that is, the variable exploration is completed, the final invariant parameters can be used as the pruned sub-network for the subsequent invariant learning.

[0073] S3. Then, the parameter model that only contains invariant parameters after completing variable exploration is used as a sub-network for invariant learning. The loss function of invariant learning is constructed. The cross entropy of the predicted probability obtained by inputting the training samples into the invariant learning network is used as the main loss function, and invariance constraints are introduced to guide the model to learn invariant features.

[0074] Invariant learning in step S3 is a machine learning method that aims to improve the model's robustness to various transformations by capturing the immutable characteristics of the data. In classification tasks, invariant learning focuses on the properties of local or global features that remain unchanged across different environments or working conditions. Its goal is to enable the model to recognize and utilize these invariant features, thereby avoiding sensitivity to small changes in the input data. In this embodiment, the various transformations include rotation, translation, and scaling.

[0075] The invariant learning in step S3 includes the following steps: selecting a five-layer convolutional neural network as the network structure for invariant learning, because the structure of the convolutional neural network is naturally spatially invariant, and a pooling layer can be used to reduce the sensitivity to position changes. After obtaining the updated mask m, the invariant parameters are used as a subnetwork for invariant learning, and the invariant learning loss function is constructed:

[0076]

[0077] In the formula Represents the loss calculated to promote the model to learn invariant feature representations in different environments given the input x and label y. Indicates further processing of the output of the parameterized model f, converting the output into classification probability, is the cross entropy loss, and its specific expression is:

[0078]

[0079] In the formula It is a regularization loss that can be used to constrain the complexity of the model, thereby guiding the model to learn the invariant features, that is, the invariance constraint. The specific expression is:

[0080]

[0081] λ is a parameter used to control cross entropy loss and regularization loss. By adjusting λ, we can balance the model's ability to fit the training data and maintain model regularization (such as invariance).

[0082] S4. Use the SAM optimizer to optimize the invariant learning loss function to minimize the loss and further improve the out-of-distribution generalization performance of the model.

[0083] In step S4, the invariant learning loss function is optimized using the SAM optimizer. Although SAM achieves good results in the in-distribution generalization task, its performance in the out-of-distribution generalization task is quite limited. This is because when using SAM to optimize the out-of-distribution generalization task, the variable parameters are incorrectly perturbed, which promotes the fitting of false features. Specifically, in the out-of-distribution generalization problem, the invariant features and false features correspond to the invariant parameters θ inv and variable parameter θ var , applying perturbations to the invariant parameters can enhance the extraction of invariant features, while applying perturbations to the variable parameters will promote the combination of false features and label information. Therefore, after eliminating the variable parameters and retaining only the invariant parameters, SAM can be used to optimize the loss function. The optimization goal is:

[0084]

[0085] SAM aims to calculate an optimal parameter perturbation In a neighborhood with a radius of ρ, the loss value can be maximized By applying ∈ * , loss change is called sharpness, which indicates the flatness of the learned loss function. Intuitively, a flatter loss function usually exhibits better generalization properties, because a slight shift in the input space does not significantly change the loss value. To apply SAM to VEIL, we only need to calculate the optimal perturbation ∈ * The updated mask is then applied to the parameter perturbations. This process not only removes variable parameters but also reduces the computational burden of SAM. The VEIL network is trained until convergence with the invariant learning loss function.

[0086] S5. Input the unknown distribution of the test sample into the trained VEIL network to obtain the predicted label of the test sample, thus completing the rolling bearing fault diagnosis process.

[0087] In order to verify the performance of the proposed VEIL method in rolling bearing fault diagnosis, the present invention uses a rolling bearing vibration signal dataset. The dataset is collected from a modular test bench, such as Figure 3 As shown, the modular test bench includes an electric motor 1, a torque measuring shaft 2, a rolling bearing test module 3, a flywheel 4 and a load motor 5 connected in sequence. The electric motor 1, the torque measuring shaft 2, the rolling bearing test module 3, the flywheel 4 and the load motor 5 constitute the existing Paderborn bearing data test bench.

[0088] The rolling bearing fault types in the dataset are divided into three types: normal, inner race fault, and outer race fault. These faults are either caused by human intervention or naturally generated. Specifically, a manual electric engraving machine is used to damage the rolling bearing and run it on a modular test platform to collect the corresponding fault sample data. Faults generated under natural operating conditions are generated by accelerated life tests. Figure 4 As shown, the accelerated life test bench consists of a motor and a bearing box connected together. It accelerates the reduction of the bearing service life by applying a higher radial force to the bearing inside the bearing box, and provides unfavorable conditions of improper bearing lubrication to promote the formation of damage, thereby collecting naturally occurring rolling bearing failure sample data.

[0089] After building the test platform, the bearings were operated at different speeds, load torques, and radial forces. Bearing data was sampled at a 64kHz frequency to obtain artificially damaged and naturally damaged bearing samples with varying damage levels under different operating conditions. A subset of these samples was selected to validate the proposed method, as shown in Table 1.

[0090] Table 1 Experimental working conditions

[0091] Working conditions Source of damage Speed ​​(rpm) Load torque (Nm) Radial force (N) Damage level C1 Man-made damage 1500 0.7 1000 1 C2 Man-made damage 900 0.7 1000 1 C3 Man-made damage 1500 0.1 1000 1 C4 Man-made damage 1500 0.7 400 1 C5 Natural damage 1500 0.7 1000 1

[0092] In real-world industrial scenarios, the operating environment of rolling bearings is typically filled with noise, which can interfere with sensor signals. Therefore, to make model training more realistic, we added Gaussian white noise to all operating condition samples in Table 1, and set the signal-to-noise ratio of the vibration signal to noise to -2dB.

[0093] The VEIL parameters were set as follows: sparsity ratio R = 60%, initial learning rate of 1e-3 for both Adam optimizers, and hyperparameter α = 0.2. The five-layer convolutional neural network with invariant learning was configured as shown in Table 2. The hyperparameter λ = 0.6 was used to balance cross-entropy loss and regularization loss. For SAM, adaptive SAM was not used in this paper, and the perturbation amplitude was set to ρ = 0.05. Once set, the VEIL parameters were not changed and were used in subsequent experiments.

[0094] Table 2 Five-layer convolutional neural network configuration table

[0095]

[0096]

[0097] To verify that VEIL can effectively address the low rolling bearing fault diagnosis accuracy issue caused by the significant distribution discrepancy between source and target domain samples under variable operating conditions, this experiment randomly sampled 300 rolling bearing samples (100 normal rolling bearing samples, 100 inner race fault samples, and 100 outer race fault samples) from operating condition C1 at a 1:1:1 ratio. Similarly, 300 samples were sampled from operating conditions C2 and C3 using the same ratio and method. These samples from conditions C1, C2, and C3 served as source domain samples with different distributions for training VEIL. Furthermore, 300 samples from operating condition C4, using the same ratio and method, were sampled as target domain samples with an unknown distribution to test VEIL's performance.

[0098] Then, according to the rolling bearing fault diagnosis process based on the VEIL method, source domain samples are input into the parameterized VEIL. After the VEIL is trained, target domain samples are input for fault diagnosis. The proposed VEIL method is compared with traditional machine learning methods: support vector machines (SVM), deep learning methods: multi-layer domain adaptation (MLDA), transfer learning methods: discriminative transfer joint matching (DTJM), and domain generalization-related methods: correlation-aware adversarial domain generalization (CAADG). To ensure that these methods are well-suited for the rolling bearing fault diagnosis task, the parameters of SVM, MLDA, DTJM, and CAADG are optimized using k-fold cross-validation. The optimized parameters are as follows: for SVM, the penalty factor is C = 0.50, and the network learning rate is α = 2.2×e -2 For CAADG, the network learning rate α is 0.001, the momentum is set to m = 0.9, and the weight decay is set to 5×10 -4 The hyperparameters for the adversarial and distribution matching losses are γ = σ = 0.1. For MLDA, the network learning rate is α = 0.001, and the penalty coefficient is λ = 0.70. For DTJM, the subspace dimension is set to d = 20, and the regularization parameter is λ = 1. The final experimental results are the average of 25 experiments to avoid random errors. The diagnostic accuracy of each method for the three fault types and the average diagnostic accuracy are shown in Table 3.

[0099] Table 3 Diagnostic accuracy and average accuracy of each method for different faults under working condition C4 (%)

[0100]

[0101]

[0102] Table 3 shows that VEIL's diagnostic accuracy and average accuracy for different faults are higher than those of SVM, MLDA, DTJM, and CAADG. Therefore, VEIL uses human-induced damage fault samples from different operating conditions as source domain training samples and can accurately diagnose faults in target domain samples from unknown operating conditions.

[0103] In actual production, rolling bearing failures are often caused by natural fatigue from long-term operation. Therefore, to test VEIL's diagnostic performance for fault samples caused by natural damage, 300 samples were randomly sampled from condition C1:1:1 (100 normal rolling bearing samples, 100 inner race fault samples, and 100 outer race fault samples). Similarly, 300 samples were sampled from conditions C2 and C3 using the same ratio and method. The samples from conditions C1, C2, and C3 were used as source domain samples with different distributions. Furthermore, 300 samples were sampled from condition C5 using the same ratio and method as the target domain samples with an unknown distribution. Fault diagnosis was then performed according to the rolling bearing fault diagnosis process based on variable exploration-invariant learning described above, completing the entire rolling bearing fault diagnosis process. The other four methods maintained the same network parameters and structure and re-diagnosed the samples from condition C5. The final experimental results are the average of 25 experiments to avoid random errors. The diagnostic accuracy of each method for the three fault types and the average diagnostic accuracy are shown in Table 4.

[0104] Table 4 Diagnostic accuracy and average accuracy of each method for different faults under working condition C5 (%)

[0105]

[0106] Comparing Tables 3 and 4, we can see that the average diagnostic accuracy of the five methods (SVM, MLDA, DTJM, CAADG, and VEIL) for samples in operating condition C5 is lower than that for samples in operating condition C4. This is primarily due to the fact that the source domain training samples only contain fault samples caused by human damage, resulting in better learning and memory of the features of these samples. Of course, it is also possible that due to the different sources of damage, samples with human damage and natural damage exhibit different characteristic representations, which can affect the model's learning effect and classification ability. Overall, considering the scarcity of fault samples caused by natural damage in actual industrial scenarios and the relatively small decrease in fault diagnosis accuracy in the experimental results, the VEIL model maintains excellent diagnostic performance, verifying that VEIL can effectively address the problem of low rolling bearing fault diagnosis accuracy caused by the large distribution difference between source and target domain samples under variable operating conditions. In this example, the samples in operating condition C5 are fault samples caused by natural damage, while the samples in operating condition C4 are fault samples caused by human damage.

[0107] VEIL's invariant performance evaluation:

[0108] Generally speaking, the effectiveness of invariant learning is evaluated by comparing its classification accuracy on cross-distribution test data with that of ERM. However, this comparison may be affected by the specific experimental data used, leading to potential bias. Therefore, the covariate-shift representation invariance criterion (CRIC) is introduced as a quantitative criterion for robustly evaluating invariant performance. The smaller the CRIC value, the more invariant the representation learned by the representation method, that is, the easier it is to learn invariant features that remain stable under different working conditions. For the specific method of calculating the CRIC value, please refer to the literature. VEIL uses SAM to optimize invariant learning. In order to verify that SAM can be applied to invariant learning, SAM is replaced with the classic invariant learning optimization methods Empirical Risk Minimization (ERM), Invariant Risk Minimization (IRM), and Distributionally Robust Optimization (DRO), and their CRIC values ​​are calculated respectively, represented by Q. The final calculation results are all taken as the negative logarithm of 10, that is, -log 10 Q, convenient expression, the specific results are as follows Figure 5 shown.

[0109] Depend on Figure 5 As can be seen, ERM, because it aims to minimize the average loss over the training samples, overfits to the training samples and therefore fails to capture invariant patterns. Its invariant performance is significantly worse than that of the other three optimization methods. SAM, on the other hand, has a significantly lower CRIC value than IRM and DRO. This means that the invariant performance of the invariant learning model optimized using SAM in this paper is superior to that of IRM and DRO.

[0110] The quality of out-of-distribution generalization performance generally depends on the model's generalization ability on unknown distributions, that is, it depends on the model's performance on data sets with unknown distributions. In the task of rolling bearing fault diagnosis, it can be said to be the diagnostic performance of fault samples under unknown working conditions and the ability to distinguish fault characteristics.

[0111] The Receiver Operator Characteristic (ROC) curve and the area under the ROC curve (AUC) are introduced to measure a model's classification ability. Generally speaking, the closer the ROC-AUC is to 1, the better the model's classification ability is. The closer the ROC-AUC is to 0.5, the worse the model's classification ability is. A ROC-AUC of 0.5 indicates that the model's classification ability is random. The ROC curve is primarily used for binary classification problems. When used for multi-classification problems, the ROC curve can be extended using the "one-vs-rest (OvR)" or "one-vs-one (OvO)" method. In the OvR method, one of the classes is considered positive and the remaining classes are considered negative. The true positive rate and false positive rate are calculated based on this. The ROC curve for each class can be derived by analogy. According to the method of testing the fault diagnosis accuracy in the previous article, source domain samples are extracted, and target domain samples are extracted from working condition C5. Then the whole fault diagnosis process is carried out. The experimental results are averaged over 25 experiments. The ROC curve of each fault type is calculated respectively. Similarly, the ROC curves of SVM, MLDA, DTJM, and CAADG are calculated for comparison. Figure 6 、 Figure 7 and Figure 8 In addition, in order to observe the classifiability of the feature representations of different fault types learned by the model, t-SNE is used to intuitively represent the fault features extracted by VEIL and the other four methods, as shown in Figure 2. Figures 9-13 .

[0112] Depend on Figure 6-Figure 8 It can be seen that the ROC-AUC of the method VEIL proposed in this chapter is significantly greater than that of SVM, MLDA, DTJM, and CAADG, indicating that the classification ability of VEIL is relatively good. Figures 9-13 It can be seen that the fault features extracted by VEIL are clearly differentiated from each other, and fault features of the same category are well clustered together, showing better classification than the fault features extracted by the other four methods. Combined with a comparison of the diagnostic accuracy and average accuracy of different rolling bearing faults using each method, it can be seen that the VEIL model can generalize well to unknown distributions and has excellent out-of-distribution generalization performance.

[0113] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A rolling bearing fault diagnosis method based on variable exploration invariant learning, characterized in that: The following steps are involved: S1. First, source domain samples with different distributions are input into the initial parameter model, and the model is trained in a sparse manner using masks, which effectively reduces the complexity of the model, improves training efficiency and computational efficiency, and enhances the robustness of the model. S2. Then, the training samples extracted from multiple different distributions are input into the model, and the invariant parameters and variable parameters separated by the mask are optimized. The optimization process uses the Adam optimizer to calculate the gradient, and uses this to sort the degree to which the parameters are affected by the distribution. The invariant parameters with the least relevance to the classification task are eliminated, and the variable parameters with the least influence of the distribution are used as new invariant parameters, thereby achieving the purpose of dynamically updating the mask until the parameter optimization target is minimized and the variable exploration process is completed; S3. Then, the parameter model containing only the invariant parameters after completing the variable exploration is used as a sub-network for invariant learning. The loss function of invariant learning is constructed. The cross entropy of the predicted probability obtained by inputting the training samples into the invariant learning network is used as the main loss function. Invariance constraints are introduced to guide the model to learn invariant features. S4. Use the SAM optimizer to optimize the invariant learning loss function to minimize the loss and further improve the out-of-distribution generalization performance of the model; S5. Input the unknown distribution of the test sample into the trained VEIL network to obtain the predicted label of the test sample, thus completing the rolling bearing fault diagnosis process.

2. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1 is characterized in that: The variable exploration in S2 is to extract multiple data sets from m distributions, that is, ε = {e1, ...e m }, where each distribution Contains n samples and class labels Therefore, for each sample from distribution e, we can assign a distribution index And denote the data points as (x,y,d); In addition, let ε unseen is an unknown distribution, from which samples are collected to evaluate the generalization performance of invariant learning; let f θ : is a parameterized model with parameters θ∈Θ. The features extracted by this model are expressed as The goal of variable exploration is to prune a sparse subnetwork from an over-parameterized model. During the pruning process, we need to find and exclude variable parameters that change with the distribution, and retain the invariant parameters that remain stable when the distribution changes, thereby improving the performance of out-of-distribution generalization.

3. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1 is characterized in that: During the variable exploration process in S2, in order to obtain a good initialization, the model is sparse, and a binary mask (mask) m is usually applied to perform element-by-element multiplication, that is, Such a mask m can be learned through optimization or obtained according to certain criteria, such as sensitivity analysis, weight value, Fisher information, or even random initialization; by setting the sparsity ratio You can decide how many parameters are excluded from sparse training; Choose Random Initialization to start variant exploration from a pre-trained model with a randomly initialized mask m.

4. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1, characterized in that: In the process of variable exploration in S2, there are two optimization goals: Where θ inv and θ var It represents the invariant parameters and variable parameters distinguished by mask m, which are expressed as Equation (3) and Equation (4). h and g are two fully connected layers that map the extracted features to the class label space respectively. and distribution space Intuitively, the goal is a classification task, which attempts to make predictions based on label information, while Trying to distinguish each distribution based on distribution information is not helpful for subsequent invariant learning.

5. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1, characterized in that: The optimization process in S2 includes the following steps: S21. Using the Adam (Adaptive Moment Estimation) optimizer, in order to avoid conflicts when optimizing constant parameters and variable parameters, two Adam optimizers are used, namely Adam1 and Adam2, to include constant parameters and variable parameters respectively; in addition, Adam1 will include the category prediction head h, while Adam2 will include the distribution prediction head g; in minimizing In the process, the gradient amplitude can be used To find the parameters most relevant to the loss function, i.e. the invariant parameters; Likewise, by minimizing The gradient magnitude can be found Larger parameters, i.e. variable parameters, are sensitive to distribution information and cannot help extract invariant features; these parameters can then be sorted according to the size of the gradient to show the degree to which they are activated by the corresponding target; S22. Next, in order to dynamically improve the entire variable exploration process, the mask m is updated after each iteration, and the invariant parameter with the smallest activation degree is excluded, while the variable parameter with the smallest activation degree is used as the invariant parameter. The specific process is as follows: Where ArgTopK(v,k) is used to find the index of the top k largest elements in the array or list v, and m[·] represents the index m. In addition, to determine how many parameters need to be swapped, the cosine annealing function is used, which is expressed as follows: Where t and T represent the current iteration round and the total iteration round respectively, and the hyperparameter α determines the maximum value. The new mask obtained by equations (5) and (6) finds parameters that are less affected by distribution information and more closely related to the learning task, further improving the performance of the model. S23, will eventually and Once the optimization is completed, that is, the variable exploration is completed, the final invariant parameters can be used as the pruned sub-network for the subsequent invariant learning.

6. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1, characterized in that: The invariant learning in S3 is a machine learning method that aims to improve the robustness of the model to various transformations by capturing the immutable features of the data. In classification tasks, invariant learning focuses on the properties of local or global features that remain unchanged in different environments or working conditions. Its goal is to enable the model to recognize and utilize these invariant features, thereby avoiding sensitivity to small changes in the input data.

7. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1, characterized in that: The invariant learning in S3 includes the following steps: selecting a five-layer convolutional neural network as the network structure for invariant learning, because the structure of the convolutional neural network is naturally spatially invariant, and a pooling layer can be used to reduce sensitivity to position changes. After obtaining the updated mask m, the invariant parameters are used as a subnetwork for invariant learning, and the invariant learning loss function is constructed: In the formula Represents the loss calculated to promote the model to learn invariant feature representations in different environments given the input x and label y. Indicates further processing of the output of the parameterized model f, converting the output into classification probability, is the cross entropy loss, and its specific expression is: In the formula It is a regularization loss that can be used to constrain the complexity of the model, thereby guiding the model to learn the invariant features, that is, the invariance constraint. The specific expression is: λ is a parameter used to control cross entropy loss and regularization loss. By adjusting λ, we can balance the model's ability to fit the training data and maintain model regularization.

8. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1, characterized in that: In S4, the SAM optimizer is used to optimize the invariant learning loss function. Although SAM has achieved good results in the same distribution generalization task, its performance in the out-of-distribution generalization task is quite limited. This is because when using SAM to optimize the out-of-distribution generalization task, the variable parameters are subjected to incorrect perturbations, which promotes the fitting of false features. Specifically, in the out-of-distribution generalization problem, the invariant features and false features correspond to the invariant parameters θ, respectively. inv and variable parameter θ var , applying perturbations to the invariant parameters can enhance the extraction of invariant features, while applying perturbations to the variable parameters will promote the combination of false features and label information; therefore, after eliminating the variable parameters and retaining only the invariant parameters, SAM can be used to optimize the loss function, and the optimization goal is:

9. The rolling bearing fault diagnosis method based on variable exploration invariant learning according to claim 1, characterized in that: The SAM aims to calculate an optimal parameter perturbation In a neighborhood with a radius of ρ, the loss value can be maximized By applying ∈ * , loss change Known as sharpness, this indicates the flatness of the learned loss function; Intuitively, flatter loss functions usually exhibit better generalization properties, since a slight shift applied in the input space does not significantly change the loss value; to apply SAM to VEIL, one only needs to calculate the optimal perturbation ∈ * The updated mask is then applied to parameter perturbations, a process that not only removes variable parameters but also reduces the computational burden of SAM.