Rolling bearing domain generalization fault diagnosis method and device, medium and equipment

Through time-frequency conversion and mixed feature networks, the discriminant and generalized features of rolling bearings are extracted, and the problem of insufficient fusion of time-domain and frequency-domain information in the prior art is solved, and the accuracy and generalization ability of fault diagnosis are improved.

CN120213462AActive Publication Date: 2025-06-27CIVIL AVIATION UNIV OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510555760.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-27
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing rolling bearing domain generalization fault diagnosis method cannot fully integrate the time domain and frequency domain information of vibration signals, resulting in inaccurate fault diagnosis.

Method used

A rolling bearing domain generalization fault diagnosis method is proposed, time-domain and frequency domain signals are obtained through time-frequency conversion, and a hybrid feature network is constructed, including IBN network and Mixstyle network, differentiating and generalizing features are extracted, and then spliced ​​and entered into the fault classifier for diagnosis.

Benefits of technology

This method effectively extracts the distinctive features and generalization features of rolling bearings, improves the accuracy of fault diagnosis and the generalization capabilities of the model, and is suitable for practical applications under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120213462A_ABST
    Figure CN120213462A_ABST
Patent Text Reader

Abstract

The invention discloses a rolling bearing domain generalization fault diagnosis method, and relates to the technical field of domain generalization fault diagnosis, and the method comprises the following steps: obtaining a time domain signal of a rolling bearing vibration signal, carrying out the time-frequency conversion of the time domain signal, and obtaining a frequency domain signal; constructing a fault diagnosis model; the fault diagnosis model comprises a mixed feature network and a fault classifier; the hybrid feature network comprises an IBN network and a Mixstack network; jointly inputting the time domain signal and the frequency domain signal into a mixed feature network, and extracting distinguishing features in the frequency domain signal through an IBN network; generalization features in the time domain signals are extracted through a Mixstack network; the distinguishing features extracted through the IBN network and the generalization features extracted through the MixStyle network are spliced to form a high-dimension feature vector; and inputting the spliced feature signals into a fault classifier for fault diagnosis. The method can effectively extract the distinguishing features and generalization features of the rolling bearing, and improves the diagnosis accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of fault diagnosis, and particularly to a rolling bearing domain generalization fault diagnosis method. Background Art

[0002] In the fault diagnosis of rolling bearings, traditional deep learning models rely on the training data and validation data to follow similar distributions. However, in actual industrial scenarios, the rotational speed and load conditions of rolling bearings change dynamically, resulting in significant differences in data distributions and limited model generalization capabilities. Researchers have explored more flexible and efficient domain generalization (DG) methods. Currently, researchers mainly achieve domain generalization through data processing, optimizing network architectures, and designing optimization strategies to cope with dynamic and complex working condition changes and ensure reliable diagnostic capabilities on unknown data. Therefore, this invention needs to make full use of vibration data and further study how to design a network architecture that can effectively extract generalization features while taking into account discriminative features.

[0003] For existing DG diagnostic methods, in the processing of the obtained vibration signals, the domain generalization method cannot fully integrate the time-domain and frequency-domain information of the vibration signals, resulting in inaccurate fault diagnosis. Summary of the Invention

[0004] A rolling bearing domain generalization fault diagnosis method provided by an embodiment of this specification can effectively extract the discriminative features and generalization features of rolling bearings and improve the accuracy of diagnosis. The method includes the following steps:

[0005] Obtain the time-domain signal of the rolling bearing vibration signal, perform time-frequency conversion on the time-domain signal to obtain the frequency-domain signal;

[0006] Construct a fault diagnosis model; the fault diagnosis model includes a hybrid feature network and a fault classifier; the hybrid feature network includes an IBN network and a Mixstyle network;

[0007] Jointly input the time-domain signal and the frequency-domain signal into the hybrid feature network, extract the discriminative features in the frequency-domain signal through the IBN network; extract the generalization features in the time-domain signal through the Mixstyle network; splice the discriminative features extracted through the IBN network and the generalization features extracted through the MixStyle network to form a high-dimensional feature vector;

[0008] Input the spliced feature signal into the fault classifier for fault diagnosis.

[0009] Preferably, it further includes the optimization of the feature extractor parameters, including the following steps:

[0010] According to different fault types, confirm n working conditions e, and the empirical risk of each working condition e is

[0011] Construct the IRM expression; adjust the weights of the invariance constraints in the IRM expression using cross-entropy loss. The initial weights are:

[0012]

[0013] The adjusted weights are:

[0014]

[0015] In the formula, η is the update rate;

[0016] Give more invariance constraints to the working conditions with worse performance to obtain the improved invariance constraints;

[0017] The improved IRM formula is:

[0018]

[0019] In the formula, λ represents the balance parameter, which is used to adjust the weight between the empirical risk and the invariance constraint, represents the gradient of the loss function with respect to w when w = 1.0, ‖·‖ 2 represents the quadratic norm of the gradient, w e represents the weight of the invariance constraint, Φ represents the feature extractor, and w represents the fault classifier, is all source domain working conditions;

[0020] In actual mini-batch optimization, use an unbiased estimate to replace the invariance constraint. The formula is:

[0021]

[0022] In the formula, and are two random mini-batch samples from working condition e, with a size of b; l is a loss function, L wIRM is the improved invariance constraint;

[0023] Set the empirical risk to the cross-entropy loss of the classifier. The formula is:

[0024]

[0025] The formula for the joint optimization objective is:

[0026] L loss = L c + λL wIRM ;

[0027] The optimization strategy is:

[0028]

[0029] In the formula, l0 is the learning rate, is the current feature extractor parameter, is the updated feature extractor parameter, is the current feature extractor parameter, is the updated feature extractor parameter.

[0030] Preferably, the time-frequency conversion of the time-domain signal includes the following steps:

[0031] The phase-frequency characteristic is obtained through Fourier transform, and the expression of the phase value is:

[0032]

[0033] In the formula, R(x 1d )I(x 1d ) are the real part and the imaginary part of F(X 1d ), respectively, and the Fourier transform is calculated by the fast Fourier transform algorithm.

[0034] Preferably, the construction of the hybrid feature network includes the following steps:

[0035] Combine the IN network and the BN network, add the IN network to the shallow layer of the CNN network, and retain the BN network in the deep layer. Combine and apply instance normalization and batch normalization in the CNN network; for the given input feature f(x), the processes of instance normalization and batch normalization are as follows:

[0036]

[0037] Among them, μ in and σ in are the mean and standard deviation calculated during instance normalization, while μ bn and σ bn are the statistics of batch normalization;

[0038] Through weighted combination, the final output feature is:

[0039] f output =γ·f bn +(1 - γ)·f in ;

[0040] Among them, γ is a hyperparameter in the interval [0,1], which is used to adjust the balance between instance normalization and batch normalization;

[0041] Apply the MixStyle network to the shallow layer of the CNN network. In the shallow layer, MixStyle mixes the working condition features of the existing samples to generate the sample features under the new working condition;

[0042] For a given input sample x i and x j , calculate the feature representations f(x i ) and f(x j ), and extract the mean and standard deviation from them, that is:

[0043] μ i , σ i = Mean(f(x i ), Std(f(x i ));

[0044] μ j , σ j = Mean(f(x j ), Std(f(x j ));

[0045] By mixing these statistical information, generate new mean and standard deviation:

[0046] μ mix = αμ i + (1 - λ)μ j

[0047] σ mix = ασ i + (1 - λ)σ j ;

[0048] Where α is a hyperparameter between 0 and 1, used to adjust the mixing ratio of the samples;

[0049] Use the mixed mean and standard deviation to perform style transformation on the features, so as to obtain a new feature representation f min :

[0050]

[0051] In the mixed feature extraction network, the IBN network is used to extract the discriminative features in the phase-frequency signal, and the Mixstyle network is used to extract the generalization features in the time-domain signal. After splicing the two, they are input into the fault classifier.

[0052] The present invention also provides a rolling bearing domain generalization fault diagnosis system, and the system includes:

[0053] A processor;

[0054] A memory, on which a computer program that can run on the processor is stored;

[0055] Wherein, when the computer program is executed by the processor, it implements the steps of the rolling bearing domain generalization fault diagnosis method.

[0056] The present invention also provides a computer-readable storage medium, on which a data processing program is stored. When the data processing program is executed by a processor, the steps of the rolling bearing domain generalization fault diagnosis method are implemented.

[0057] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:

[0058] The rolling bearing fault diagnosis method proposed by the present invention makes full use of the characteristics of vibration signals to provide multi-dimensional support for the diagnosis task at the signal level; a novel hybrid feature network is designed, and the IBN network is used to extract discriminative features from phase-frequency information to improve the accuracy of the diagnosis task, and the MixStyle network is used to extract generalization features from time-domain signals to enhance the generalization ability of the model.

[0059] Secondly, aiming at the characteristic differences of different source domains, the method proposes an improved IRM optimization method. By introducing the classification cross-entropy loss to assign different weights to the IRMs of different source domains and focusing on the source domains with poor performance, it ensures that the source domain with the weakest performance can also obtain optimized performance. This method significantly improves the diagnostic generalization ability of the model in a multi-source heterogeneous environment and provides technical support for practical applications under complex working conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0061] Figure 1 is a schematic flow chart of a rolling bearing domain generalization fault diagnosis method provided by the present invention;

[0062] Figure 2 is a model of a rolling bearing domain generalization fault diagnosis method provided by the present invention;

[0063] Figure 3 are the amplitude-frequency characteristics and phase-frequency characteristics extracted by a signal processing module provided by the present invention; wherein, Figure 3 (a) is the amplitude spectrum; Figure 3 (b) is the phase spectrum;

[0064] Figure 4 is a hybrid feature extractor provided by the present invention;

[0065] Figure 5 is a confusion matrix of some generalization experiments in the experimental settings provided by the present invention; wherein, Figure 5 (a), (b), (c), and (d) are the confusion matrices of the generalization experiments;

[0066] Figure 6 This is the T-SNE visualization result in the experimental setup provided by the present invention; wherein, Figure 6 (a), (b), (c), and (d) are the T-SNE visualization results;

[0067] Figure 7 This is the comparative experimental result on the data of the bearing vibration test bench in the experimental setup provided by the present invention;

[0068] Figure 8 This is the confusion matrix and the T-SNE visualization result on the data of the bearing vibration test bench in the experimental setup provided by the present invention; wherein, Figure 8 (a), (b), (c), and (d) are the confusion matrices, Figure 8 (e), (f), (g), and (h) are the T-SNE visualization results. Detailed implementation manners

[0069] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0070] In the fault diagnosis of rolling bearings, the collection of fault signals is very important. It is necessary to comprehensively, reasonably, and fully utilize the vibration signals to provide multi-dimensional feature support for the diagnosis task at the signal level. Since the commonly used feature network extractors usually cannot take into account the recognition of feature signals, the deep model will perform poorly when facing complex diagnosis tasks.

[0071] The existing DG diagnosis methods mainly achieve domain generalization by data processing, optimizing the network architecture, and designing optimization strategies to cope with dynamic and complex working condition changes. They only enhance the diversity of data, but do not solve the distribution shift problem at the feature level. The optimization strategies proposed in network optimization will also lead to a decline in the performance of the diagnosis model.

[0072] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.

[0073] Figure 1 This is a schematic flowchart of a rolling bearing domain generalization fault diagnosis method in the present invention, Figure 2 This is a model diagram, which specifically includes the following steps:

[0074] S101: The phase-frequency characteristics obtained by Fourier transform contain rich high-level semantics and are less sensitive to operating condition changes. Figure 3 As shown in the figure, through the visualization of the original signal, it can be seen that the information contained in the phase-frequency characteristic far exceeds the amplitude-frequency characteristic. Therefore, the phase-frequency characteristic contains rich invariant information, and invariant features can be extracted from the phase-frequency characteristic. The phase-frequency characteristic can be obtained by the following formula:

[0075]

[0076] Where n is the index and N is the length of the signal. The Fourier transform can be efficiently calculated using the Fast Fourier Transform algorithm. The expression for the phase value is as follows, where R(x 1d ) and I(x 1d ) are F(X 1d )'s real and imaginary parts.

[0077] The time domain signal directly reflects the dynamic changes and time characteristics of the system, which is very important for fault diagnosis. Compared with the frequency domain features, the time domain signal can retain more details of the original data and better describe the changes of the rolling bearing under different working conditions. Therefore, the time domain signal contains richer generalization information, and the generalization features are extracted from the time domain signal.

[0078] S102: Figure 4 As shown in Figure 1, the hybrid feature extractor consists of two modules: the IBN network and the MixStyle network. The IBN network extracts invariant features from the phase-frequency signal, and the MixStyle network extracts generalized features from the time-domain signal, thereby realizing the fusion of distinguishing features and generalizing features.

[0079] IBN: The features extracted by the shallow layer of the CNN network are more about the characteristics of working condition changes, while the features extracted by the deep layer contain higher-level semantic information. IBN filters out information related to the environment while retaining the distinguishing features, improving the accuracy of the fault truth, and eliminating the impact of working condition changes in the shallow network. The IBN network structure is as follows Figure 4 shown.

[0080] The core idea of ​​the IBN network is to perform double normalization on the feature map to achieve flexible adaptation to features in different domains. Specifically, the IBN network applies instance normalization and batch normalization after each convolutional layer, so that the deep learning network can better capture the discriminative features in the task, thereby improving the generalization ability. The parameters of the IBN network are shown in Table 1. This process can be described by the following formula:

[0081] For a given input feature map f(x), the process of instance normalization and batch normalization is as follows:

[0082]

[0083] Among them, μ in and σ in are the mean and standard deviation calculated during instance normalization respectively, while μ bn and σ bn are the statistics of batch normalization.

[0084] Through weighted combination, the final output feature is expressed as:

[0085] f output = γ·f bn +(1 - γ)·f in ;

[0086] Among them, γ is a hyperparameter in the range of [0, 1], which is used to adjust the balance between instance normalization and batch normalization.

[0087] Table 1 Detailed parameters of the IBN network

[0088] Name Input / Output / Convolution Kernel / Stride / Padding Normalization Method Cov (1,4,8,1,1) IBN Cov (4,16,8,1,1) IBN Cov (16,32,8,1,1) BN Cov (32,64,8,1,1) BN Cov (64,128,8,1,1) BN AdaptiveMaxPool 4 /

[0089] The features extracted by the MixStyle shallow neural network contain more working condition transformation features. Therefore, the MixStyle network is applied to the shallow layer of the CNN network. The structural design of the MixStyle network is as Figure 4 shown. In the shallow layer, MixStyle generates sample features under new working conditions by mixing the working condition features of existing samples, which is beneficial for the deep model to extract richer generalization features. The detailed parameters of the MixStyle network are shown in Table 2; this method is described through the following steps:

[0090] First, for the given input samples x i and x j , MixStyle calculates their feature representations f(x i ) and f(x i ), and extracts the mean and standard deviation from them, that is:

[0091] μ i , σ i = Mean(f(x i )), Std(f(x i ));

[0092] Then, by mixing these statistical information, new mean and standard deviation are generated:

[0093] μ mix = αμ i +(1 - λ)μ j ;

[0094] σmix = ασ i + (1 - λ)σ j ;

[0095] Where α is a hyperparameter between 0 and 1, used to adjust the mixing ratio of the samples. Then, the mean and standard deviation of the mixture are used to perform style transformation on the features, thereby obtaining a new feature representation f mix :

[0096]

[0097] After the phase-frequency signal is feature-extracted by the IBN network, a 1*2560-dimensional vector is generated. After the time-domain signal is feature-extracted by the MixStyle network, a 1*2560-dimensional vector is also generated. In the present invention, the two are concatenated to form a 1*5120-dimensional vector and then input into a fault classifier for fault diagnosis.

[0098] Table 2 Detailed parameters of the MixStyle network

[0099] Name Input / Output / Convolution Kernel / Stride / Padding Normalization Method Cov (1,4,8,1,1) BN MixStyle / / Cov (4,16,8,1,1) BN MixStyle / / Cov (16,32,8,1,1) BN Cov (32,64,8,1,1) BN Cov (64,128,8,1,1) BN AdaptiveMaxPool 4 /

[0100] S103: In IRM, assuming that the data comes from multiple working conditions and each working condition has a different distribution, by optimizing a constrained problem, IRM requires the model to maintain the same optimal prediction performance in all working conditions. The optimization objective is as follows:

[0101]

[0102] Where: represents the empirical risk in working condition e, and the constraint condition of the formula is:

[0103]

[0104] Since the above formula is a two-layer optimization problem, it is simplified to a single-variable optimization problem;

[0105]

[0106] Where Φ becomes the entire invariant predictor, w = 1.0 is a scalar and fixed "virtual" classifier, the gradient norm penalty is used to measure the optimality of the virtual classifier in each environment, and λ ∈ [0, ∞) is a regularization parameter, used to balance the prediction ability (empirical risk minimization term) and the invariance of the predictor 1·Φ(x).

[0107] However, there is a problem with the above formula. In the fault diagnosis of rolling bearings, due to the different data distribution offsets under various working conditions, the way of treating data under various working conditions equally in the above formula often cannot achieve the optimal effect. Therefore, the present invention uses cross-entropy loss to adjust the weight of the invariance constraint, giving more invariance constraints to the working conditions with worse performance, so as to ensure that the deep model performs excellently under all working conditions, and then improve the generalization ability of the deep model. The specific improvements are as follows. Suppose there are n working conditions, and the empirical risk of each working condition e in the current mini-batch is In the initial state, the weight w of the domain e is set to:

[0108]

[0109] For each working condition e sampled in the current mini-batch, its weight update formula is:

[0110]

[0111] where: η is the update rate, used to control the amplitude of the update;

[0112] After obtaining the updated weight, multiply the weight by the invariance constraint to obtain the improved IRM:

[0113]

[0114] where: represents the empirical risk in working condition e. λ is a balance parameter, used to adjust the weight between the empirical risk and the invariance constraint represents the gradient of the loss function with respect to w when w = 1.0, ‖·‖ 2 represents the quadratic norm of the gradient, w e represents the weight of the invariance constraint, Φ represents the feature extractor, and w represents the fault classifier. is all source domain working conditions. However, in actual mini-batch optimization (such as the stochastic gradient descent algorithm), it is difficult to directly calculate the invariance constraint Therefore, the present invention uses the unbiased estimate of to replace it, and the formula is as follows:

[0115]

[0116] where, and are two random mini-batch samples from working condition e, with size b, and l is a loss function, L wIRM is the improved invariance constraint.

[0117] The present invention takes the empirical risk Set as the cross-entropy loss of the classifier:

[0118]

[0119] Wherein, p(·) and p(·|·) respectively represent the marginal probability distribution and the conditional probability distribution.

[0120] The joint optimization objective of the proposed IBNMixNet is expressed by the following formula:

[0121] L loss = L c + λL wIRM ;

[0122] Wherein, L wIRM is the improved invariance constraint, and λ is a hyperparameter for adjusting the relationship between L c and L IRM . The optimization strategy is as follows:

[0123]

[0124] Wherein, l0 is the learning rate, is the current feature extractor parameter, is the updated feature extractor parameter, is the current feature extractor parameter, is the updated feature extractor parameter.

[0125] Based on the rolling bearing fault diagnosis method of the present invention, the characteristics of the vibration signal are fully utilized to provide multi-dimensional support for the diagnosis task at the signal level; a novel hybrid feature extraction network is designed, the IBN network is used to extract discriminative features from the phase-frequency information to improve the accuracy of the diagnosis task, and the MixStyle network is used to extract generalization features from the time-domain signal to enhance the generalization ability of the model; aiming at the characteristic differences of different source domains, an improved IRM optimization method is proposed: by introducing the classification cross-entropy loss to assign different weights to the IRMs of different source domains, focusing on the source domains with poor performance, so as to ensure that the source domain with the weakest performance can also obtain optimized performance. This method significantly improves the diagnostic generalization ability of the model in a multi-source heterogeneous environment and provides technical support for practical applications under complex working conditions.

[0126] The above is the rolling bearing fault diagnosis method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides an optimization idea for the parameters of the feature extractor, including the following steps:

[0127] According to different fault types, confirm multiple working conditions;

[0128] Construct the IRM expression:

[0129]

[0130] The cross entropy loss is used to adjust the weight of the invariance constraints in the IRM expression. The weight in the initial state is:

[0131]

[0132] The adjusted weights are:

[0133]

[0134] Where: η is the update rate, used to control the update amplitude;

[0135] More invariance constraints are given to the working conditions with worse performance to obtain improved invariance constraints; unbiased estimates of improved non-deformation constraints are obtained based on the improved invariance constraints;

[0136] The improved IRM formula is:

[0137]

[0138] in: represents the empirical risk in condition e, λ is a balance parameter used to adjust the weight between the empirical risk and the invariance constraint, It means that when w = 1.0, the gradient of the loss function with respect to w, ‖·‖ 2 represents the quadratic norm of the gradient, w e represents the weight of the invariance constraint, Φ represents the feature extractor, w represents the fault classifier, For all source domain conditions.

[0139] In actual mini-batch optimization, unbiased estimation is used instead of invariant constraints, and the formula is:

[0140]

[0141] in, and are two random mini-batches of size b from case e and l is a loss function, L wIRM For improved invariance constraints;

[0142] The empirical risk is set as the cross entropy loss of the classifier, which is:

[0143]

[0144] The joint optimization objective is:

[0145] L loss =L c +λL wIRM ;

[0146] It is shown that the specific optimization strategy is as follows:

[0147]

[0148] where l0 is the learning rate, is the current feature extractor parameter, is the updated feature extractor parameter, is the current feature extractor parameter, is the updated feature extractor parameter.

[0149] In this embodiment, specific experimental verifications are given, including specifically:

[0150] 1: Dataset introduction

[0151] A: CWRU dataset: The CWRU dataset is one of the commonly used bearing fault datasets. The basic situation of the CWRU dataset is shown in Table 3. This dataset includes acceleration signals collected from the drive end and the fan end. Under each collection point, it is divided into four different working conditions according to different loads and speeds. Each working condition includes one normal data and three fault data (inner ring, outer ring, rolling element), and the damage diameters are 0.007, 0.014, 0.021, and 0.028 inches respectively. Eight domains and seven fault categories are selected for the generalization experiment of rolling bearing fault diagnosis.

[0152] Table 3 Details of the CWRU dataset

[0153] Domain Name Rotational Speed Load Collection Location Fault Type A 1797 0HP Drive End 0123456 B 1772 1HP Drive End 0123456 C 1750 2HP Drive End 0123456 D 1730 3HP Drive End 0123456 E 1797 0HP Fan End 0123456 F 1772 1HP Fan End 0123456 G 1750 2HP Fan End 0123456 H 1730 3HP Fan End 0123456

[0154] B: HIT aero-engine bearing fault dataset: The HIT dataset is one of the typical representatives of aero-engine bearing fault datasets. The basic situation of the HIT dataset is shown in Table 5. In the HIT dataset, it is divided into 28 working conditions according to different high and low pressure speeds of the aero-engine. Each working condition includes a normal condition, an inner ring fault, and two outer ring faults. The detailed fault conditions are shown in Table 4. In order to increase the experimental difficulty and test the generalization ability of the IBNMixNet model, four working conditions with relatively large differences in high and low pressure speeds are selected for the generalization experiment.

[0155] C: Data of the bearing vibration test bench: The third set of experimental data comes from the bearing vibration test bench in the experimental center, with the model number TYS1-8. Five types of faults and one normal type were designed in the experiment. The fault types are: normal, inner ring fault, outer ring fault, ball fault, outer-inner ball fault, and outer-inner ring fault. Nine working conditions were designed according to different rotational speeds and axial loads, and the working condition settings are shown in Table 6. To increase the experimental difficulty, the designed axial load changes greatly, and the vibration signal distribution shifts more severely, so as to better verify the excellent performance of the proposed model.

[0156] Table 4 Fault types of the HIT aeroengine bearing dataset

[0157] Number Fault Location Damage Depth Damage Length 1 Outer Ring 0.5 0.5 2 Inner Ring 0.5 0.5 3 Inner Ring 0.5 1.0

[0158] Table 5 The 28 working conditions included in the HIT aeroengine bearing dataset

[0159]

[0160] Table 6 The 9 working conditions included in the data of the bearing vibration test bench

[0161]

[0162] Table 7 Ablation experiment design

[0163]

[0164]

[0165] Table 8 Range of key parameters in the comparative experiment

[0166]

[0167] 2: Experimental setup

[0168] To ensure the fairness of the experimental results, the same main structure is adopted in the experiment. The experiment is carried out on a computer equipped with an NVDIA GeForce RTX 4060Ti GPU, a 13th Gen Intel(R) Core(TM) i7 CPU, and 32GB of memory. To verify the effectiveness of the IBNMixNet method, extensive experiments are designed. The experiments are divided into two parts: ablation experiments and comparative experiments. The ablation experiment design is shown in Table 7, and the comparative experiment parameter ranges are shown in Table 8. The comparative methods are introduced as follows:

[0169] 1): ERM: The core idea is to minimize the empirical risk of all source domains, thus achieving domain generalization.

[0170] 2): DCORAL: Align the covariance between groups of source domain data to reduce distribution differences and thus achieve domain generalization.

[0171] 3): DCC: Align the MMD metrics between groups of source domain data, align the features of each source domain, and achieve domain generalization.

[0172] 4): CNN-C: Introduce center loss into CNN to improve the generalization ability of the deep diagnosis model.

[0173] 5): DRO: This method combines DRO with regularization to improve the generalization ability of over-parameterized models on the worst-case population while maintaining the average accuracy.

[0174] 6): VREx: Improve the generalization ability by reducing the risk differences during the source domain training process.

[0175] 7): MixStyle: Generate samples of new styles through feature mixing to improve the generalization ability.

[0176] 3: Experimental results on the CWRU dataset

[0177] A: Ablation experiment results: The experimental results of various ablation models on the CWRU dataset are shown in Table 9, and the confusion matrices of some generalization experiments are as Figure 5 shown. First, it can be seen from the table that the M6 method obtains the highest average accuracy, and in most generalization tasks, the M6 method also obtains the optimal performance. By comparing M1 and M2, it is found that simply combining IRM and MixStyle cannot improve the model performance because the method of treating each source domain equally cannot guarantee the improvement of the generalization ability of the deep diagnosis model. Combining with M3, it can be seen that the improved IRM shows more superior performance due to fully considering the source domains with worse performance and obtains the highest accuracy in two generalization tasks, but the overall performance is still inferior to M6. By comparing M4 with M1 and M5, it is found that using a hybrid feature extractor obtains better results than a single feature extractor, which proves the importance of considering both generalization features and discriminative feature extraction. By comparing M6 and M4, it is found that optimizing the deep diagnosis model using the improved IRM obtains better results.

[0178] Table 9 Ablation experiment results on the CWRU dataset

[0179] Invisible Domain 0 1 2 3 4 5 6 7 Average M1 88.28 86.72 100 97.66 83.59 90.62 95.31 88.28 91.31 M2 89.06 86.72 100 92.97 88.28 87.5 94.35 91.41 91.29 M3 93.75 80.47 100 98.44 85.16 94.53 99.22 91.41 92.87 M4 92.91 89.06 100 92.97 85.94 89.06 99.22 89.84 92.38 M5 77.34 85.94 100 94.53 81.25 81.25 87.5 77.34 85.64 M6 95.24 90.62 100 100 90.48 94.05 100 91.67 95.26

[0180] B: Comparative experiment results: On the CWRU dataset, a comparative experiment was conducted between M6 and other well-known domain generalization methods. The experimental results are shown in Table 10, and the confusion matrix is as Figure 5As shown. It can be seen from the table that the M6 method achieved the highest average accuracy rate, which was 3% - 25% higher than other methods. In eight generalization tasks, the M6 method achieved the highest accuracy rate in six generalization tasks. It can be seen from the T-SNE graph that the M6 method achieved the best clustering effect compared to other methods, demonstrating the powerful domain generalization diagnosis ability of the M6 method.

[0181] Table 10 Comparative experiment results on the CWRU dataset

[0182] Invisible Domain 0 1 2 3 4 5 6 7 Average ERM 81.25 72.66 92.19 83.59 62.5 67.19 74.22 68.75 75.29 DCORAL 89.84 82.03 98.44 92.19 83.59 92.97 100 91.41 91.31 DCC 76.56 76.56 81.25 76.56 85.16 92.97 98.44 94.53 85.25 CNN-C 83.59 71.09 89.84 85.16 67.19 74.22 75.78 65.62 75.56 DRO 71.09 78.12 82.03 78.12 57.03 63.28 70.31 63.28 70.41 VREx 97.66 86.72 100 88.28 82.03 93.75 78.91 77.34 88.09 MixStyle 88.28 86.72 100 97.66 83.59 90.62 95.31 88.28 91.31 M6 95.24 90.62 100 100 90.48 94.05 100 91.67 95.26

[0183] C: Hyperparameter analysis: The present invention analyzes the hyperparameter λ on the CWRU dataset, and the experimental results are shown in Table 11. It can be seen from the table that when λ > 1, the accuracy rate in each generalization task shows a downward trend, and when λ < 1, the performance in each generalization task also shows a downward trend. When λ = 1, the deep diagnosis model achieves the optimal effect. Therefore, λ = 1 is selected in the experiments of this article.

[0184] Table 11 Hyperparameter analysis

[0185] Invisible Domain 0 1 2 3 4 5 6 7 Average λ = 0.5 83.59 77.34 100 98.44 85.94 92.19 100 88.28 90.72 λ = 1.0 91.41 90.62 100 100 90.62 93.75 100 91.41 94.73 λ = 2.0 89.06 85.16 100 98.44 88.28 90.62 99.22 89.06 92.48 λ = 4.0 89.06 82.03 100 99.22 85.16 90.62 100 87.5 91.70 λ = 7.0 87.50 82.03 98.44 99.22 83.59 90.62 100 87.5 91.11

[0186] 4: Experimental results on the HIT aeroengine bearing fault dataset

[0187] A: Ablation experiment results: The experimental results of various ablation models on the HIT dataset are shown in Table 12. Similarly, it can be seen from the table that the M6 method achieved the highest average accuracy rate, and also achieved the optimal effect in the vast majority of generalization tasks. The average accuracy rate was 1% - 6% higher than other deep diagnosis models. By comparing M1, M2, and M3, it was found that the improved IRM achieved the highest accuracy rate, once again demonstrating the excellent performance of the improved IRM. By comparing M4, M5, and M1, it was found that the effect of hybrid feature extraction was better than that of single feature extraction, once again demonstrating the significance of taking into account both generalization features and discriminative features.

[0188] Table 12 Ablation experiment results on the HIT dataset

[0189] Invisible Domain 0 1 2 4 Average M1 89.06 99.22 93.36 71.88 88.38 M2 92.19 97.27 96.48 70.7 89.16 M3 95.7 100 96.09 68.75 90.14 M4 91.8 99.22 89.45 71.48 87.99 M5 87.5 100 82.42 71.48 85.35 M6 94.95 100 98.44 71.48 91.22

[0190] B: Comparative experiment results: Two groups of generalization experiments were carried out on the HIT dataset. The first group of experiments was to generalize three domains to an unseen target domain, as shown in Table 13, and the second group of experiments was to generalize two domains to two unseen target domains, as shown in Table 14. The experimental results are shown in the table, and the T-SNE visualization graph is as Figure 6As shown. In the first set of generalization experiments, method M6 achieved the highest average accuracy, which was 1% - 6% higher than other methods. Among the four generalization tasks, M6 obtained the highest accuracy in three of them. In the second set of generalization experiments, it still achieved the highest average accuracy, which was 2% - 5% higher than other methods. Through the T-SNE graph, it can be found that by fusing the extraction of discriminative features and generalization features, the M6 model reduced the intra-class distance between various faults and increased the inter-class distance, and M6 obtained the best clustering effect. The results of the above two sets of generalization experiments further prove that the proposed method has strong domain generalization ability.

[0191] Table 13 Comparison experiment results on the HIT dataset (1)

[0192] Invisible Domain 0 1 2 4 Average ERM 89.54 98.05 91.02 71.88 87.62 DCORAL 80.47 99.22 86.33 71.48 84.38 DDC 84.38 96.48 95.31 70.7 86.72 CNN-C 90.62 100 96.88 71.88 89.85 DRO 74.61 88.28 85.53 71.48 79.98 VREx 72.27 83.59 80.47 71.09 76.86 MixStyle 89.06 99.22 93.36 71.88 88.38 M6 94.53 100 98.44 71.48 91.11

[0193] Table 14 Comparison experiment results on the HIT dataset (2)

[0194] Invisible Domain 01 23 13 04 Average ERM 75.36 70.12 79.11 75.99 75.14 DCORAL 74.81 71.88 71.68 76.56 73.73 DDC 80.06 67.78 72.47 79.11 74.85 CNN-C 80.45 68.95 72.66 75.39 74.36 DRO 74.61 68.95 71.88 76.17 72.90 VREx 69.14 70.51 71.88 75.20 71.68 MixStyle 73.83 69.73 77.93 75.78 74.32 M6 80.86 73.05 81.45 76.18 77.88

[0195] 5: Experimental results of the bearing vibration test bench data

[0196] A: Ablation experiment results: The ablation experiment results on the bearing vibration test bench data in the experimental center are shown in Table 15. It is found from the experimental results that the proposed method achieved the highest average accuracy, which was 2% - 5% higher than other methods. By comparing M1, M2, and M3, it is concluded that the improved IRM performed best in terms of accuracy, verifying its superior performance again. After further comparing M4, M5, and M1, the present invention found that the effect of hybrid feature extraction was significantly better than that of single feature extraction, which once again emphasized the importance of simultaneously paying attention to generalization features and discriminative features.

[0197] B: Comparison experiment results: The ablation experiment results on the bearing vibration test bench data in the experimental center are as Figure 7 shown, and the confusion matrices and TSNE visualization graphs of some comparison experiments are as Figure 8 shown. It can be clearly observed from these experimental results that the proposed method still maintained excellent performance in the generalization diagnosis effect, which was 5% - 8% higher than other methods. The analysis of the confusion matrix further confirmed the high classification accuracy of the method of the present invention, significantly reducing the misdiagnosis rate. Through the visualization of the T-SNE graph, it is observed that the proposed method more effectively completed the clustering of fault samples by taking into account both generalization features and discriminative features, achieving a larger inter-class interval and a smaller intra-class interval, thus obtaining better diagnostic performance.

[0198] Table 15 Ablation experiment results on the bearing vibration test bench dataset

[0199]

[0200]

[0201] The above is the rolling bearing domain generalization fault diagnosis method provided by an embodiment of this embodiment. Based on the same idea, this embodiment also provides a corresponding rolling bearing domain generalization fault diagnosis system. For the specific limitations of the rolling bearing domain generalization fault diagnosis system, reference can be made to the limitations of the rolling bearing domain generalization fault diagnosis method in the above text, which will not be elaborated here. Each module in the above rolling bearing domain generalization fault diagnosis system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0202] This embodiment also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 provided rolling bearing domain generalization fault diagnosis method.

[0203] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0204] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A rolling bearing domain generalized fault diagnosis method, characterized in that: include: Obtaining a time domain signal of a rolling bearing vibration signal, performing time-frequency conversion on the time domain signal, and obtaining a frequency domain signal; Constructing a fault diagnosis model; the fault diagnosis model includes a hybrid feature network and a fault classifier; the hybrid feature network includes an IBN network and a Mixstyle network; The time domain signal and the frequency domain signal are jointly input into the hybrid feature network, and the discriminative features in the frequency domain signal are extracted through the IBN network; the generalized features in the time domain signal are extracted through the Mixstyle network; the discriminative features extracted through the IBN network and the generalized features extracted through the MixStyle network are concatenated to form a high-dimensional feature vector; The spliced ​​feature signals are input into the fault classifier for fault diagnosis.

2. The rolling bearing domain generalized fault diagnosis method according to claim 1, characterized in that: It also includes the optimization of feature extractor parameters, including the following steps: According to different fault types, n working conditions e are confirmed, and the empirical risk of each working condition e is Construct an IRM expression; adjust the weights of the invariance constraints in the IRM expression using the cross entropy loss. The weights in the initial state are: The adjusted weights are: Where η is the update rate; Give more invariance constraints to the working conditions with worse performance to obtain improved invariance constraints; The improved IRM formula is: In the formula, λ represents the balance parameter, which is used to adjust the weight between the empirical risk and the invariance constraint. It means that when w = 1.0, the gradient of the loss function with respect to w, ‖·‖ 2 represents the quadratic norm of the gradient, w e represents the weight of the invariance constraint, Φ represents the feature extractor, w represents the fault classifier, For all source domain conditions; In actual mini-batch optimization, unbiased estimation is used instead of invariant constraints, and the formula is: In the formula, and are two random mini-batches of size b from condition e; l is a loss function, L wIRM For improved invariance constraints; The empirical risk is set as the cross entropy loss of the classifier, which is: The formula for the joint optimization objective is: L loss =L c +λ L wIRM ; The optimization strategy is: In the formula, l0 is the learning rate, is the current feature extractor parameter, are the updated feature extractor parameters, is the current feature extractor parameter, are the updated feature extractor parameters.

3. The rolling bearing domain generalized fault diagnosis method according to claim 1, characterized in that: The time-frequency conversion of the time domain signal comprises the following steps: The phase-frequency characteristic is obtained by Fourier transform, and the expression of the phase value is: In the formula, R(x 1d )I(x 1d ) are F(X 1d ), and the Fourier transform is calculated by the fast Fourier transform algorithm.

4. The rolling bearing domain generalized fault diagnosis method according to claim 1, characterized in that: The construction of the hybrid feature network, The following steps are involved: Combining the IN network and the BN network, the IN network is added to the shallow layer of the CNN network, and the BN network is retained in the deep layer. Instance normalization and batch normalization are combined and applied in the CNN network; for a given input feature f(x), the process of instance normalization and batch normalization is as follows: Among them, μ in and σ in are the mean and standard deviation calculated when the instance is normalized, and μ bn and σ bn is the batch normalized statistic; Through weighted combination, the final output features are: f output =γ·f bn +(1-γ)·f in ; Among them, γ is a hyperparameter in the interval [0,1], which is used to adjust the balance between instance normalization and batch normalization; The MixStyle network is applied to the shallow layer of the CNN network. In the shallow layer, MixStyle generates sample features under new working conditions by mixing the working condition features of existing samples. For a given input sample x i and x j , calculate the feature representation f(x i ) and f(x j ), and extract the mean and standard deviation from it, namely: m i ,s i =Mean(f(x i )),Std(f(x i )); m j ,s j =Mean(f(x j )),Std(f(x j )); By blending these statistics, a new mean and standard deviation are generated: m mix =am i +(1-λ)μ j s mix =as i +(1-λ)σ j ; Among them, α is a hyperparameter between 0 and 1, which is used to adjust the mixing ratio of samples; Use the mixed mean and standard deviation to perform style conversion on the features to obtain a new feature representation f mix : In the hybrid feature extraction network, the distinguishing features in the phase-frequency signal are extracted by the IBN network, and the generalized features in the time domain signal are extracted by the Mixstyle network, and the two are spliced ​​and input into the fault classifier.

5. A rolling bearing domain generalized fault diagnosis system, characterized in that: The system comprises: processor; a memory having stored thereon a computer program executable on the processor; Wherein, when the computer program is executed by the processor, the steps of the rolling bearing domain generalized fault diagnosis method as described in any one of claims 1 to 4 are implemented.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a data processing program, and when the data processing program is executed by a processor, the steps of the rolling bearing domain generalized fault diagnosis method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Improved JGSA algorithm-based rolling bearing fault diagnosis method under variable working conditions

    CN113869451A

  • Multi-source domain bearing fault diagnosis method based on hybrid convolution

    CN117349749A

  • Bearing fault diagnosis method based on domain adversarial network and feature enhancement

    CN119046660A

  • Variable working condition bearing fault diagnosis method based on domain adversarial hybrid graph convolutional network

    CN119397362A