A Fault Diagnosis Method for Bearing Systems Based on a Deep Time-Frequency Domain Adaptive Network
By designing a deep time-frequency domain adaptive network, combining the spectrum neural layer and Sinkhorn divergence, the problem of insufficient cross-domain generalization capability in bearing fault diagnosis is solved, and higher fault diagnosis accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510164343.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The existing deep learning models lack cross-domain generalization capabilities in bearing fault diagnosis, especially in different equipment and operating conditions, and lack of time-frequency domain characteristics for special learning architectures, resulting in insufficient accuracy and robustness of cross-domain fault diagnosis.
The design is based on the deep time-frequency domain adaptive network, and the Sinkhorn divergence is constructed, combined with the time-domain and spectrum feature extraction, and the spectrum is realized end-to-end fault diagnosis is achieved. Spectral convolution, spectrum pooling and spectrum batch normalization are used to process the spectrum information, and domain adaptation is performed by iteratively calculating the Sinkhorn divergence.
It improves the cross-domain robustness and accuracy of bearing system fault diagnosis, can effectively identify fault patterns in different fields and conditions, simplifies processing flow and improves the adaptability and accuracy of the model.
Smart Images

Figure CN119719958B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bearing system fault diagnosis, and in particular to a bearing system fault diagnosis method based on a deep time-frequency domain adaptive network. Background Art
[0002] Rotating machinery is an essential component of many industries, including manufacturing, transportation, and energy. Bearings are essential components of rotating machinery and are subject to large loads at high speeds, which makes them susceptible to malfunctions and failures. Therefore, bearing fault diagnosis is critical to the safe, reliable, and stable operation of rotating equipment.
[0003] In recent years, fault diagnosis has made significant progress due to the development of intelligent technology, especially the increasing popularity of deep learning. A large number of studies have explored the applicability of various deep learning architectures, including convolutional networks (CNNs), recurrent convolutional networks (RNNs), Transformers, etc. in the field of fault diagnosis. However, traditional deep learning has poor cross-domain generalization ability because it assumes the consistency of data distribution, which is not always guaranteed in practice. Insufficient labeled data of the target object and changes in operating conditions further limit the applicability of these methods and hinder the use of historical data. This limitation is particularly important in fault diagnosis because mechanical equipment may exhibit different degradation modes and fault characteristics. Therefore, conducting cross-domain fault diagnosis research has greater practical significance.
[0004] Deep transfer learning techniques, especially domain adaptation methods, have been shown to be successful in improving fault diagnosis in cross-domain scenarios. Although domain adaptation methods have significantly improved the cross-domain generalization of fault diagnosis by reducing the differences in feature distributions, some challenges still exist. For example, a unique feature of time series data for fault diagnosis is that feature shifts occur in both the time and frequency domains. While existing models use traditional structures such as convolutional neural networks to process time domain signals or time-frequency transformed data, these structures lack dedicated learning architectures for spectral representation. This limitation can limit the model's ability to learn physical features from the spectrum. Second, few studies have explored domain invariance and reducing feature shifts in the spectral and time domains. In complex variable scenarios, data mining in the time domain feature space is difficult to handle cross-domain generalization, resulting in negative transfer. Therefore, it is necessary to design a learning architecture for time-frequency spectrum representation that achieves domain adaptation in both the time and spectral spaces to further improve the cross-domain generalization of fault diagnosis. Summary of the invention
[0005] In order to overcome the shortcomings and deficiencies of the prior art, the present invention provides a bearing system fault diagnosis method based on a deep time-frequency domain adaptive network.
[0006] A fault diagnosis method for a bearing system based on a deep time-frequency domain adaptive network, the method comprising:
[0007] Step S1, dataset preparation: Collect multi-domain vibration signals of bearings with different fault modes, and establish a cross-domain dataset, including the labeled source dataset D s , the unlabeled target dataset D t and the test dataset D test ;
[0008] Step S2, model initialization: Construct a deep time-frequency domain adaptive network, and initialize the deep time-frequency domain adaptive network with parameters θ, and then perform hyperparameter tuning according to the selected dataset;
[0009] Step 3S, model parameter optimization: Train the deep time-frequency domain adaptive network with the cross-domain training sets D s and D t ;
[0010] Step S4, test verification: Output the classification result of the optimized deep time-frequency domain adaptive network for the test set D test , and count and analyze the results.
[0011] Further, the constructing of the deep time-frequency domain adaptive network includes: obtaining a spectral representation by using the discrete Fourier transform, spectral neural layer calculation, and calculating the Sinkhorn divergence through an iterative data solution method.
[0012] Further, for the obtaining of the spectral representation by using the discrete Fourier transform, for a given time series sample x ∈ R C×T containing C channels and t time points, perform a one-dimensional discrete Fourier transform on each channel to obtain f = F(x) ∈ C C×T , and the expression is:
[0013]
[0014] where k represents the k-th frequency index and c represents the c-th channel. f (c) represents the discrete Fourier transform result of the signal x (c) , and F(x (c) ) is also the discrete Fourier transform result of the signal x (c) , x (c) represents the original signal of the c-th channel, t represents each discrete time point in the time series, T represents the total number of samples of the signal x (c) , e is the natural constant Euler number, here used as the base of the exponential function in the Fourier transform, j is the imaginary unit, 2π is the coefficient in the exponential function, and C represents the total number of channels.
[0015] Furthermore, the spectral neural layer calculation includes: spectral convolution, spectral pooling, spectral activation, and spectral batch normalization;
[0016] The spectral convolution uses the complex convolution type. Given a complex input vector x = u + iv ∈ C, a complex filter matrix W = A + iB ∈ C is used for convolution, where u and v ∈ R are real vectors, and A and B ∈ R are real matrices. The calculation expression of the complex convolution is:
[0017] W * x = (A * u - B * v) + i(B * u - A * v)
[0018] Among them, W * x represents the product between the vector x and the matrix W, the symbol * represents the convolution operation, W is a complex matrix containing the real part and the imaginary part, and x is a complex vector; A and B represent the linear transformations acting on u and v respectively, u and v represent the real part and the imaginary part of the complex vector x respectively, and i represents the imaginary unit.
[0019] The matrix form expression is:
[0020]
[0021] Among them, represents the real part of the complex expression W * x above, represents the imaginary part of the complex expression W * x above.
[0022] The spectral pooling usually operates based on the importance of frequency components, retaining the main frequency information. Its spectral pooling process expression is:
[0023]
[0024] Among them, Y is the output after pooling, Pool represents the pooling operation, x(t) is the time-domain signal, and f is the frequency.
[0025] The spectral activation uses a complex activation function based on ReLU, that is, the zReLU activation function. Given a complex input x ∈ C, zReLU only retains the values in the first quadrant. The expression is:
[0026]
[0027] Among them, θ x represents the phase of x, which satisfies the Cauchy - Riemann equation, except for any point in the set in.
[0028] The spectral batch normalization is to avoid the vanishing gradient through batch normalization and is also used to accelerate training. Given a mini-batch where x i∈C, this task is regarded as whitening a two-dimensional vector, and the expression is:
[0029]
[0030] where, represents the data obtained after normalization, E[X] represents the expected value of matrix X, and V represents the covariance matrix, and the expression is:
[0031]
[0032] where, V rr and V ri and V ir and V ii respectively represent the specific elements in the covariance matrix V, and respectively represent the real part and the imaginary part of matrix X, Cov represents the covariance operation symbol, which ensures that has a standard complex distribution with mean μ = 0, covariance Γ = 1 and pseudo-covariance Γ = 0; the expressions for the mean, covariance and pseudo-covariance are respectively:
[0033]
[0034] where, μ represents the expected value or mean of the data and Γ represents the covariance, Π represents the pseudo-covariance, and i represents the imaginary unit.
[0035] Spectral batch normalization has two parameters, including a shift parameter β and a scaling parameter γ, both of which are complex. The scaling parameter has three degrees of freedom to ensure that the variances of two original principal components are both 1. The translation parameter β = β r +iβ i has two trainable values corresponding to the real part and the imaginary part. Spectral batch normalization SBN is defined as:
[0036]
[0037] In the formula, SBN represents the result of the batch normalization operation, β is the shift parameter, and γ represents the scaling parameter. Usually, in this application, γ ri =β r =β i = 0.
[0038] Furthermore, the Sinkhorn divergence is calculated by the iterative data solution method, and the specific steps are as follows: First, initialize two input distributions, namely the feature data distribution and the fault mode distribution of the bearing system. Then, perform optimization calculations through the Sinkhorn divergence, gradually adjust these distributions to minimize them under a certain metric, and finally obtain the calculated Sinkhorn divergence to measure the difference between the two distributions. This regularization is known to be equivalent to restricting the search space and has the following scaling form:
[0039] P0 = diag(u)Kdiag(v)
[0040] where P0 represents the value obtained through matrix transformation, diag(u) is the operation of diagonal matrixizing the vector u, the matrix K = exp(-C / η) is the Gibbs kernel, and diag(v) is the operation of diagonal matrixizing the vector v. Note that exp(·) here represents element-wise operation. Starting from the selected initial values, the fixed points are solved by the following iterative method:
[0041]
[0042] where u (l+1) represents the updated value of the vector u in the (l + 1)-th iteration, v (l+1) represents the updated value of the vector v in the (l + 1)-th iteration, u (l) and v (l) represent the current values in the l-th iteration respectively, u s represents the result after non-linearly transforming the vector u, u t represents the transposed form of the vector u, t is the time step, s represents the index value, and K T represents the transposed matrix of the matrix K.
[0043] After L iterations, the optimal pairing is P (L) = diag(u (L) )Kdiag(v (L) ). Therefore, the final divergence is:
[0044]
[0045] where represents the calculation result of the final divergence, η represents the regularization parameter, C is a matrix, P (L) also represents a matrix, L is the current layer number of the matrix, c ij represents the element in the i-th row and j-th column of the matrix C, n s represents the number of samples related to u s , and n t represents the number of samples related to u tThe relevant number of samples, and respectively represent the i-th and j-th vector elements of the L-th layer in vectors u s and u t , and K ij represents the element at the i-th row and j-th column in matrix K.
[0046] To achieve the task objective of fault diagnosis, the output features of the deep time-frequency domain adaptive network are mapped to each fault category, and a fully connected layer with a SoftMax activation function is used as a classifier to output the class probabilities.
[0047] Furthermore, the objective of the deep time-frequency domain adaptive network is to minimize the structural risk of the network on the target domain, and the expression is:
[0048]
[0049] where θ * represents the optimal parameter in the optimization process, θ is the network parameter, ε t represents the objective function, Pr[·] represents the data distribution in the training dataset, D represents the dataset containing the input feature x and the corresponding label y, t is the time step, c(·) is the loss function, g(x) represents the predicted value generated by the model for the input x, and y is the target variable that the model needs to predict during training.
[0050] The objective function of the deep time-frequency domain adaptive network simultaneously minimizes the empirical risk on the source domain and the difference between domains. The former is manifested as the classification error based on supervised learning, that is:
[0051]
[0052] where represents the calculated value, n s is the number of samples, C represents the total number of categories, is the calculation symbol of the cross-entropy loss function, i and s are index variables respectively, is the i-th sample x i in the dataset of the s-th subset, i represents the s-th subset of the i-th sample y.
[0053] The expression of the objective function of the deep time-frequency domain adaptive network is:
[0054]
[0055] where θ * represents the optimal parameter in the optimization process, θ is the network parameter, argmin represents the minimization operation, is the objective function, is the calculation result of the final divergence, η represents the regularization parameter, and λ represents the trade-off coefficient.
[0056] Beneficial effects:
[0057] The present invention proposes a bearing system fault diagnosis method based on a deep time-frequency domain adaptive network. This method first designs a spectral feature learning structure containing multiple spectral neural layers to process complex spectral values. Secondly, the time-domain feature extraction branch is combined with the spectral feature extraction branch to form a time-frequency domain adaptive network; among them, the spectral representation process of the time-frequency domain adaptive network is completed by using the regularized optimal transport distance, that is, the Sinkhorn divergence. Finally, an end-to-end fault diagnosis architecture based on the time-frequency domain adaptive network is proposed. This architecture aligns the time-frequency domain from both the time domain and the spectral space through learning representations, thereby achieving more robust and accurate performance. This method extracts more discriminative features from the time domain and the spectral domain. The spectral feature branch of the deep time-frequency domain adaptive network includes operations such as spectral convolution and spectral pooling to process complex spectral representations. The designed deep time-frequency domain adaptive method performs source-target alignment on the extracted time-frequency domain spectral representations and uses the Sinkhorn divergence to learn more domain-invariant representations. Applying the deep time-frequency domain adaptive network to fault diagnosis, this method improves the robustness and accuracy in cross-domain fault diagnosis scenarios. Description of the drawings
[0058] Figure 1 is the flowchart of the method steps of the present invention;
[0059] Figure 2 is the spectral representation diagram of the present invention including one-dimensional discrete Fourier transform;
[0060] Figure 3 is the spectral convolution diagram of the present invention including the complex convolution type;
[0061] Figure 4 is the one-dimensional and two-dimensional spectral pooling diagram of the present invention including complex convolution operations;
[0062] Figure 5 is the detailed structure diagram of the deep time-frequency domain adaptive network of the present invention. Specific implementation manners
[0063] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following further describes the present application in detail with reference to the drawings and specific embodiments.
[0064] As Figure 1As shown in the figure, a fault diagnosis method for a bearing system based on a deep time-frequency domain adaptive network, the method comprising:
[0065] Step S1, data set preparation: Collect multi-domain vibration signals of bearings with different fault modes, and establish a cross-domain data set, including the labeled source data set D s , the unlabeled target data set D t and the test data set D test ;
[0066] Specifically, this step includes in the first step of the fault diagnosis method, first collecting the bearing vibration signals with different fault modes. The diversity of the data set is the key to the success of this method. Therefore, the collected vibration signals should cover different fault conditions, such as normal operation, slight wear, severe damage, etc. These signals come from multiple domains and need to be collected under different working conditions to ensure the extensiveness and representativeness of the data. The collected signals will constitute a cross-domain data set, including: the source data set D s , which is the labeled training data set, containing the vibration signals and corresponding labels under different fault modes. The basic data for training the algorithm. The target data set D t , which is the unlabeled data set, simulating the situation of lack of labels in actual applications, and is usually used to test the generalization ability of the algorithm. The test data set D test , this data set contains the bearing vibration signals for model verification, verifying the classification accuracy of the deep time-frequency domain adaptive network.
[0067] Step S2, model initialization: Construct a deep time-frequency domain adaptive network, and initialize the deep time-frequency domain adaptive network with parameters θ, and then tune the hyperparameters according to the selected data set;
[0068] Specifically, this step includes in the model initialization stage, first constructing a deep time-frequency domain adaptive network, which is used to learn the features of vibration signals from both the time and frequency domains simultaneously. The network structure usually includes a time-frequency feature extraction layer, a fusion layer, a classifier, etc. During the initialization process, the network is initialized with parameters θ. These parameters control the structure and weights of the network. On this basis, in order to ensure that the algorithm can be optimized on different data sets, hyperparameter tuning is performed. The purpose of hyperparameter tuning is to select appropriate parameters such as learning rate, regularization coefficient, number of training epochs, etc., so as to ensure that the model does not overfit or underfit during the training process. This step is of great significance for subsequent model training and optimization.
[0069] Step 3S, model parameter optimization: Use the cross-domain training sets D s and D t to train the deep time-frequency domain adaptive network;
[0070] Specifically, this step includes using the cross-domain training set D during the model parameter optimization phase. s and D t to train the deep time-frequency domain adaptive network. The training objective is to enable the network to learn useful features from the source dataset and have good generalization ability on the target dataset. To achieve this goal, a cross-domain training method is adopted, using the labeled information of the source dataset and the unlabeled information of the target dataset to enhance the learning ability of the model. This process is carried out through iterative optimization, continuously adjusting the weights and biases in the network to minimize the loss function (such as cross-entropy loss) to achieve accurate fault diagnosis.
[0071] Step S4, Test and Validation: Output the classification results of the optimized deep time-frequency domain adaptive network for the test set D test and statistically analyze the results.
[0072] Specifically, this step includes entering the test and validation phase after the model training is completed. At this time, the trained deep time-frequency domain adaptive network will be used to classify the test set D test By statistically analyzing the classification results of the test data, the performance of the model in actual applications can be evaluated. The test metrics include classification accuracy, recall rate, F1 value, etc., aiming to comprehensively evaluate the effectiveness of the model in the fault diagnosis task. In this step, further optimization and adjustment can also be made according to the performance of the model to ensure that the model can be stably applied to actual bearing fault detection.
[0073] Constructing a deep time-frequency domain adaptive network includes: obtaining a spectral representation using the discrete Fourier transform, spectral neural layer calculation, and calculating the Sinkhorn divergence through an iterative data solution method.
[0074] The deep time-frequency domain adaptive network proposed in this application consists of (1) a feature extractor g(·) and (2) a classifier c(·). The structure is characterized in that the feature extractor contains two branches in the time domain and the frequency domain, namely the time domain g τ (·) and the frequency domain g s (·), which respectively learn feature representations from the time domain and the frequency domain. Among them, the time domain branch uses a traditional one-dimensional convolutional neural network, including four one-dimensional convolutional layers and two max-pooling layers, and each convolutional layer is followed by a BN layer and a ReLU activation function.
[0075] Meanwhile, the frequency-domain branch uses the 1D frequency-domain convolutional network designed in this application, which includes four 1D convolutional layers and two frequency-domain pooling layers, and the frequency-domain BN layer and the zReLU activation function are also placed therein. To construct a frequency-domain feature extractor corresponding to the time-domain branch, this application proposes the following spectral representation, spectral neural layer, and calculates the Sinkhorn divergence through an iterative data solution method.
[0076] The spectral representation is obtained by using the discrete Fourier transform. The first step of frequency-domain learning is to obtain the spectral representation by using the discrete Fourier transform. For a given time-series sample x ∈ R containing C channels and t time points C×T , as Figure 2 shown, the figure shows a signal processing process based on a deep time-frequency domain adaptive network, especially for the processing of vibration signals in bearing system fault diagnosis. The content in the figure involves the operation of converting the input signal from the time domain to the frequency domain. The specific details include: 1). The input signal is a vibration signal obtained through multiple channels (each channel may correspond to a different sensor or data source). These signals change over time and contain different time-frequency information, which can reveal the working state or potential fault modes of the device. In the figure, these signals are arranged at time points T in different channels (expressed as C). 2). On the left side of the figure, the input signal is presented as time-domain data, where each row represents the data of one channel, and the vertical arrangement represents the change of the signal at time point T. At this time, the signal is mainly processed in the time dimension, but the information contained therein is far more than that, so further processing is required to extract the frequency-domain features. 3). In the figure, the process of performing the Fourier transform on the input time-domain signal is represented by the symbol F(·). The Fourier transform converts the time-domain signal into a frequency-domain signal, which can capture the characteristics of the signal from the frequency perspective. On the right side of the figure, each column shows the distribution of the signal at different frequencies. After the Fourier transform, the time-domain information of the signal is converted into frequency-domain information, which helps to extract the content that is very sensitive to frequency features in the fault diagnosis process. 4). After being converted to the frequency domain, the signal not only contains time information but also can express frequency components, which is crucial for analyzing mechanical faults such as bearing damage or wear. The frequency-domain signal can clearly reveal the characteristics of different frequency bands and help analyze the type and severity of the faults. Through the deep time-frequency domain adaptive network, the network can simultaneously learn the time-domain and frequency-domain information to more comprehensively capture the operating state of the bearing system. 5). In the signal after the frequency-domain conversion, the network further extracts and optimizes the key features and finally applies them to the fault diagnosis task. Through deep learning, the network can accurately identify different fault modes and achieve accurate classification and prediction. Performing a 1D discrete Fourier transform on each channel gives f = F(x) ∈ C C×T , that is
[0077]
[0078] In the formula, k represents the k-th frequency index, c represents the c-th channel, and f (c) represents the discrete Fourier transform result of the signal x (c) , and F(x (c) ) is also the discrete Fourier transform result of the signal x (c) . x (c) represents the original signal of the c-th channel, t represents each discrete time point in the time series, T represents the total number of samples of the signal x (c) . e is the natural constant Euler's number, which is used here as the base of the exponential function in the Fourier transform. j is the imaginary unit, 2π is the coefficient in the exponential function, and C represents the total number of channels.
[0079] Spectrum neural layer calculation: The spectrum neural layer designed in this application includes spectrum convolution, spectrum pooling, spectrum activation, and spectrum batch normalization. The spectrum convolution adopts a complex convolution type to retain as much information as possible in the frequency domain. Given a complex input vector x = u + iv ∈ C, this application convolves it with a complex filter matrix W = A + iB ∈ C, where u and v ∈ R are real vectors, and A and B ∈ R are real matrices. The complex convolution calculation is as follows:
[0080] W * x = (A * u - B * v) + i(B * u - A * v)
[0081] where W * x represents the product between the vector x and the matrix W, the symbol * represents the convolution operation, W is a complex matrix containing real and imaginary parts, x is a complex vector; A and B represent linear transformations acting on u and v respectively, u and v represent the real and imaginary parts of the complex vector x respectively, and i represents the imaginary unit.
[0082] As Figure 3 shown, it demonstrates a complex matrix and vector operation process. The core task is to perform specific transformations on the signal to extract useful features. This process involves operations on complex matrices and vectors. Through this method, the algorithm can perform complex conversions between the time domain and the frequency domain, helping to improve the performance of the fault diagnosis model. The following is a detailed description of this process: 1). Initialization of complex matrices and vectors. At the Figure 3 top, the matrices A and B are two complex matrices, where A represents the real part matrix and B represents the imaginary part matrix, and i is the imaginary unit. This matrix will participate in subsequent calculations and be used to operate on the input signal. The vector u is the real part of the input signal, and v is the imaginary part of the input signal. The complex vector provides rich time-domain and frequency-domain features by combining the real part u and the imaginary part v. 2). At the Figure 3 second layer is the multiplication of two complex matrices and vectors, mainly as the product of the matrix and the vector, used to extract features in the time-domain signal. 3). At theFigure 3 For the last layer, the complex result is calculated, and finally a complex result is obtained, where the real part is A*u - B*v and the imaginary part is B*u - A*v. The result of this step is to combine the multiplication results of all the previous matrices and vectors through subtraction and addition, and finally obtain a complex signal. The real part and the imaginary part respectively represent different levels of features in the signal, and these features are used for subsequent fault diagnosis. Its matrix form expression is:
[0083]
[0084] Among them, represents the real part of the complex expression W*x above, represents the imaginary part of the complex expression W*x above.
[0085] The spectrum pooling usually operates based on the importance of frequency components, retaining the main frequency information, and its spectrum pooling process expression is:
[0086]
[0087] Among them, Y is the output after pooling, Pool represents the pooling operation, x(t) is the time-domain signal, and f is the frequency.
[0088] In addition, the spectrum pooling designed in this application can be regarded as a low-pass filter as shown in Figure 4 , which mainly has two benefits: 1). Due to the uneven distribution of information density with respect to frequency, more information can be retained under the same size. 2). The represented scale can be reduced more flexibly without suffering a large reduction like traditional pooling. Figure 4 The upper half of [] shows the process of dimension transformation of the matrix, and the lower half is a two-dimensional matrix transformation. Through this two-dimensional matrix representation, multiple feature dimensions of the signal can be considered simultaneously. At this step, the elements in the matrix are compressed or transformed through certain operations (such as pooling, convolution, etc.) into the target matrix. Finally, the final result is the features extracted in the height and width directions. This process extracts the core features in the signal through matrix transformation and provides valuable information for subsequent calculations and fault diagnosis. Through the extraction of this high-dimensional information, the algorithm can identify potential patterns or anomalies in the signal, thereby helping to perform more accurate fault diagnosis. Thus, Figure 4The presented processing process performs deep learning and processing on the time-domain and frequency-domain information in the signal through multiple matrix transformations, compressions, and feature extractions. Each matrix transformation is to ensure that the signal features can be effectively extracted and provide accurate information for subsequent classification or prediction. This process combines various operations such as matrix transposition, compression, and dimension adjustment, optimizing the signal representation method, and can effectively mine potential fault patterns and features in complex signals, thereby improving the accuracy and efficiency of fault diagnosis.
[0089] In addition, the present application adopts a ReLU-based complex activation function for the spectral activation of neuron units, namely the zReLU activation function. Given a complex input x ∈ C, zReLU only retains the values in the first quadrant, expressed as follows:
[0090]
[0091] where θx represents the phase of x, which satisfies the Cauchy-Riemann equations, except for the points in the set any point in.
[0092] Finally, batch normalization is used to avoid gradient vanishing and accelerate training at the same time. Given a mini-batch where xi ∈ C, the present application regards this task as whitening a two-dimensional vector, expressed as follows:
[0093]
[0094] where, represents the data obtained after normalization, E[X] represents the expected value of matrix X, V represents the covariance matrix, and the expression is:
[0095]
[0096] where, V rr , V ri , V ir and V ii respectively represent the specific elements in the covariance matrix V, and respectively represent the real part and the imaginary part of matrix X, Cov represents the covariance operation symbol, which ensures that has a standard complex distribution with mean μ = 0, covariance Γ = 1, and pseudo-covariance Γ = 0; the mean, covariance, and pseudo-covariance expressions are respectively:
[0097]
[0098] where, μ represents the expected value or mean of the data , Γ represents the covariance, Π represents the pseudo-covariance, and i represents the imaginary unit.
[0099] Spectral batch normalization has two parameters, including a shift parameter β and a scale parameter γ, both of which are complex. The scale parameter has three degrees of freedom to ensure that the variances of two original principal components are both 1. The translation parameter β = β r + iβ i has two trainable values corresponding to the real part and the imaginary part. The spectral batch normalization SBN can be defined as:
[0100]
[0101] where SBN represents the result of the batch normalization operation, β is the shift parameter, and γ represents the scale parameter. Usually, in this application, γ ri = β r = β i = 0.
[0102] The Sinkhorn divergence is calculated through an iterative data solution. The Sinkhorn divergence provides a computationally efficient way to approximate optimal transport and makes the loss function differentiable. This regularization is known to be equivalent to restricting the search space and has the following scaling form:
[0103] P0 = diag(u)Kdiag(v)
[0104] where P0 represents the value obtained through matrix transformation, diag(u) is the operation of diagonal matrixizing the vector u, the matrix K = exp(-C / η) is the Gibbs kernel, and diag(v) is the operation of diagonal matrixizing the vector v. Note that exp(·) here represents element-wise operation. Starting from the selected initial values, the fixed points are solved through the following iterative manner:
[0105]
[0106] where u (l+1) represents the updated value of the vector u in the (l + 1)-th iteration, v (l+1) represents the updated value of the vector v in the (l + 1)-th iteration, u (l) and v (l) represent the current values in the l-th iteration respectively, u s represents the result after non-linearly transforming the vector u, u t represents the transposed form of the vector u, t is the time step, s represents the index value, and K T represents the transposed matrix of the matrix K.
[0107] After L iterations, the optimal pairing is P (L) = diag(u (L) )Kdiag(v(L) )。Therefore, the final divergence is as follows:
[0108]
[0109] Among them, represents the calculation result of the final divergence, η represents the regularization parameter, C is a matrix, and P (L) also represents a matrix, L is the current layer number of the matrix, and c ij represents the element in the i-th row and j-th column of matrix C, and n s represents the number of samples related to u s and n t represents the number of samples related to u t respectively. and respectively represent the i-th and j-th vector elements in the L-th layer of vectors u s and u t , and K ij represents the element in the i-th row and j-th column of matrix K.
[0110] To achieve the task objective of fault diagnosis, the output features of the deep time-frequency domain adaptive network are mapped to each fault category, and a fully connected layer with a SoftMax activation function is used as a classifier to output class probabilities.
[0111] In addition, the classifier of this application is a fully connected layer with a SoftMax activation function, and the detailed structure of the deep time-frequency domain adaptive network is as Figure 5 shown. Figure 5 shows the specific implementation process of a deep time-frequency domain adaptive network for bearing system fault diagnosis. The core of this method is to perform fault diagnosis by extracting time-domain and frequency-domain features, combining the Sinkhorn divergence and the classifier. Figure 5It mainly includes: 1). Input of source domain and target domain data, which are vibration signals collected from different bearing fault modes. In practical applications, source domain data is usually labeled data, while target domain data is unlabeled data. The input data will be processed by different feature extraction modules. 2). Time-domain feature extraction. The time-domain feature extractor consists of multiple convolutional layers, ReLU activation functions, and batch normalization layers. Through these layers, the model can extract meaningful time-domain features from the original signal. For example, the convolutional layer extracts local features of the signal, the ReLU activation function introduces non-linear transformation, and the batch normalization layer helps accelerate the training process and prevent overfitting. 3). Frequency-domain feature extraction. This part is similar to the time-domain feature extractor and also consists of convolutional layers, ReLU activation functions, and batch normalization layers. In the frequency domain, the changes in the signal can more intuitively reflect the frequency characteristics. Especially when dealing with bearing fault signals, frequency-domain features can often reveal subtle fault modes. 4). Fusion of time-domain and frequency-domain features. This process combines feature information from different sources to enhance the comprehensive discrimination ability of the model, thereby improving the accuracy of fault diagnosis. 5). The fused features are further calculated by Sinkhorn divergence. In this process, the model attempts to minimize the Sinkhorn divergence to reduce the distribution difference between the source domain and the target domain, thereby improving the classification performance on the target domain. 6). Classifier and classification loss. After feature fusion and adjustment through Sinkhorn divergence, the final feature vector will be input into the classifier for prediction of fault modes. The classifier is usually a fully connected layer with a SoftMax activation function, which is responsible for mapping the extracted features to the labels of fault types. Thus, Figure 5 shows a complete deep time-frequency domain adaptive network framework, which includes multiple modules such as time-domain and frequency-domain feature extraction, feature fusion, Sinkhorn divergence calculation, classifier, and classification loss. This method combines the time domain and the frequency domain, and uses the feature differences between the source domain and the target domain for transfer learning, thereby improving the accuracy of bearing system fault diagnosis. In practical applications, this algorithm can effectively address the challenges brought by data distribution differences in different fields and ensure efficient diagnostic capabilities on unlabeled target domain data. The goal of the network is to minimize the structural risk of the network on the target domain, which is expressed as follows:
[0112]
[0113] where θ * represents the optimal parameters in the optimization process, θ is the network parameter, ε tLet \(f\) denote the objective function, \(Pr[\cdot]\) denote the data distribution in the training dataset, \(D\) denote the dataset containing the input feature \(x\) and the corresponding label \(y\), \(t\) be the time step, \(c(\cdot)\) be the loss function, \(g(x)\) denote the predicted value generated by the model for the input \(x\), and \(y\) be the target variable that the model aims to predict during training.
[0114] Therefore, its objective function simultaneously minimizes the empirical risk on the source domain and the difference between domains. The former is manifested as the classification error based on supervised learning, that is:
[0115]
[0116] where, denotes the calculated value, \(n\) s is the number of samples, \(C\) represents the total number of classes, is the calculation symbol of the cross-entropy loss function, \(i\) and \(s\) are index variables respectively, is the \(s\)-th subset of the \(i\)-th sample \(x\) in the dataset i and denotes the \(s\)-th subset of the \(i\)-th sample \(y\) i in the dataset.
[0117] Finally, the objective function of the network described in this application is expressed as follows:
[0118]
[0119] where, \(\theta\) * denotes the optimal parameter in the optimization process, \(\Theta\) is the network parameter, \(argmin\) represents the minimization operation, is the objective function, is the calculation result of the final divergence, \(\eta\) represents the regularization parameter, and \(\lambda\) represents the trade-off coefficient.
[0120] The present invention proposes a bearing system fault diagnosis method based on a deep time-frequency domain adaptive network. This method first designs a spectrum feature learning structure containing multiple spectrum neural layers to process complex spectrum values. Secondly, it combines the time-domain feature extraction branch and the spectrum feature extraction branch to form a time-frequency domain adaptive network; among them, the regularized optimal transport distance, namely the Sinkhorn divergence, is adopted to complete the time-frequency spectrum representation process of the time-frequency domain adaptive network. This method has the following advantages:
[0121] 1. Joint extraction of time-frequency domain features: This method combines the time-domain and frequency-domain feature extraction branches to form a time-frequency domain adaptive network. This method can simultaneously utilize time-domain and spectrum information, extract more discriminative features, and improve the accuracy of fault diagnosis.
[0122] 2. Adaptive time-frequency spectrum representation: By introducing the Sinkhorn divergence (regularized optimal transport distance), this method can effectively align the time-frequency spectrum representations of the source and target domains, ensuring the adaptability and robustness of the model on different data sources.
[0123] 3. Spectrum feature processing: The designed spectrum feature learning structure includes operations such as spectral convolution and spectral pooling, which can effectively process complex spectrum representations and further improve the performance of fault diagnosis.
[0124] 4. End-to-end architecture: This method proposes an end-to-end fault diagnosis architecture based on a time-frequency domain adaptive network, which can automatically learn the features extracted from the time domain and frequency domain without manual feature design, simplifying the processing flow.
[0125] 5. Cross-domain robustness: This method can maintain high robustness and accuracy in cross-domain fault diagnosis scenarios, especially suitable for fault diagnosis between different operating conditions or devices.
[0126] 6. Domain-invariant feature learning: By introducing the Sinkhorn divergence, the model can learn domain-invariant features, enabling the model to maintain consistent performance when facing data from different sources.
[0127] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various equivalent changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalent scope.
Claims
1. A fault diagnosis method for a bearing system based on a deep time-frequency domain adaptive network, characterized in that , The method includes: Step S1, Dataset Preparation: Collect multi-domain vibration signals of bearings with different fault modes and establish a cross-domain dataset, including the labeled source dataset D s , the unlabeled target dataset D t and the test dataset D test ; Step S2, Model initialization: Construct a deep time-frequency domain adaptive network, initialize the deep time-frequency domain adaptive network with parameters θ, and then prepare hyperparameter tuning according to the selected dataset; Step 3S, Model Parameter Optimization: Use the cross-domain training set D s and D t to train the deep time-frequency domain adaptive network; Step S4, Test and Verification: Output the classification results of the optimized deep time-frequency domain adaptive network for the test set D test and count and analyze the results; The construction of the deep time-frequency domain adaptive network includes: obtaining a spectral representation using the discrete Fourier transform, spectral neural layer calculation, and calculating the Sinkhorn divergence through an iterative data solution; The spectral neural layer calculation includes: spectral convolution, spectral pooling, spectral activation, and spectral batch normalization; The spectral convolution adopts the complex convolution type. Given a complex input vector , a complex filter matrix is used for convolution, where u and v ∈ R are real vectors, and A and B ∈ R are real matrices. The calculation expression of the complex convolution is as follows: Among them, represents the product x between a vector W and a matrix The symbol W denotes the convolution operation, W is a complex matrix containing real and imaginary parts, x is a complex vector; A and B represent linear transformations acting on u and v respectively, u and v represent the real and imaginary parts of the complex vector x respectively, i denotes the imaginary unit; The matrix form expression is: Among them, represents the real part of the complex expression , and represents the imaginary part of the complex expression ; The spectral pooling operates based on the importance of frequency components, retaining the main frequency information. The expression for the spectral pooling process is: Among them, Y is the output after pooling, and Pool represents the pooling operation. x ( t ) is the time-domain signal. f is the frequency. The spectral activation uses a complex activation function based on ReLU, that is, the zReLU activation function. Given a complex input x ∈ C, zReLU only retains the values in the first quadrant. The expression is: Among them, θ x represents the phase of x, satisfying the Cauchy-Riemann equations, except for any point in the set ; The spectral batch normalization avoids gradient vanishing through batch normalization and is used to accelerate training. Given a mini-batch , where x i ∈ C, this task is regarded as whitening a two-dimensional vector, and the expression is: Among them, represents the data obtained after normalization, and E X represents the expected value of the matrix X , and V represents the covariance matrix, and the expression is: Among them, V rr , V ri , V ir and V ii respectively represent the specific elements in the covariance matrix V . and respectively represent the real and imaginary parts of the matrix X . Cov represents the covariance operation symbol, which is used to ensure that has a standard complex distribution with a mean μ = 0, a covariance Γ = 1, and a pseudo-covariance Γ = 0; the mean, covariance, and pseudo-covariance expressions are respectively: Among them, represents the expected value or mean of the data, represents the covariance, represents the pseudo-covariance, i represents the imaginary unit; Spectral batch normalization has two parameters, including a shift parameter β and a scale parameter , both of which are complex. The scale parameter has three degrees of freedom to ensure that the variances of two original principal components are both 1. The translation parameter has two trainable values corresponding to the real and imaginary parts. Spectral batch normalization SBN is defined as: Among them, SBN represents the result of the batch normalization operation, β is the shift parameter, represents the scaling parameter, and is initialized .
2. The fault diagnosis method for a bearing system based on a deep time-frequency domain adaptive network according to claim 1, characterized in that, The spectral representation is obtained by using the discrete Fourier transform on a given time series sample containing C channels and t time points , and a 1D discrete Fourier transform is performed on each channel to obtain , and the expression is: Among them, k represents the k-th frequency index, and c represents the c-th channel. f (c) represents the signal x (c) of the discrete Fourier transform result. F ( x (c) ) is also the signal x (c) of the discrete Fourier transform result. x (c) represents the c -th channel's original signal. t represents each discrete time point in the time series. T represents the signal x (c) of the total number of samples. e is the natural constant Euler's number, used here as the base of the exponential function in the Fourier transform. j is the imaginary unit, and 2π is the coefficient in the exponential function. C represents the total number of channels.
3. The fault diagnosis method of a bearing system based on a deep time-frequency domain adaptive network according to claim 1, characterized in that For calculating the Sinkhorn divergence through an iterative data solution, initialize two input distributions, namely the feature data distribution of the bearing system and the fault mode distribution, and perform optimization calculations through the Sinkhorn divergence. Gradually adjust these distributions so that they are minimized under a certain metric to obtain the calculated Sinkhorn divergence, which is used to measure the difference between the two distributions. Regularization is known to be equivalent to restricting the search space and has a scaling form of: Among them, P 0 represents the value obtained through matrix transformation, diag ( u ) performs the operation of diagonal matrixization on the vector u , the matrix is the Gibbs kernel, diag ( v ) performs the operation of diagonal matrixization on the vector v , exp(‧) represents element-wise operation, starting from the selected initial value, the fixed point is solved by the following iterative method: Among them, u (l+1) represents a vector u at the update value in the l +1-th iteration, v (l+1) represents a vector v at the update value in the l +1-th iteration, u (l) and v (l) respectively represent the current values in the l -th iteration, u s represents the result after performing a non-linear transformation on the vector u ; u t represents the transposed form of the vector u ; t is the time step, s represents the index value, K T represents the transposed matrix of the matrix K ; After L iterations, the optimal pairing is , and the final divergence is: Among them, represents the calculation result of the final divergence, η represents the regularization parameter, C is a matrix, P (L) also represents a matrix, L is the current layer number of the matrix, c ij represents the matrix C in the i row and j column element, n s represents the number of samples related to u s ; n t represents the number of samples related to u t ; and respectively represent the u s and u t in the L layer and the i th and j th vector elements, K ij represents the matrix K in the i row and j column element; Map the output features of the deep time-frequency domain adaptive network to each fault category, and use a fully connected layer with a SoftMax activation function as a classifier to output class probabilities.
4. The fault diagnosis method for a bearing system based on a deep time-frequency domain adaptive network according to claim 1, characterized in that The goal of the deep time-frequency domain adaptive network is to minimize the structural risk of the network on the target domain. The expression is: Among them, represents the optimal parameter during the optimization process, θ is the network parameter, represents the objective function, Pr[·] represents the data distribution in the training dataset, D represents the dataset that contains the input features x and the corresponding labels y ; t is the time step, c (·) is the loss function, g( x ) represents the predicted value generated by the model for the input x ; y is the target variable that the model needs to predict during training; The objective function of the deep time-frequency domain adaptive network simultaneously minimizes the empirical risk on the source domain and the difference between domains. The former is manifested as the classification error based on supervised learning, that is: Among them, represents the calculated value, n s is the number of samples, C represents the total number of categories, is the calculation symbol of the cross-entropy loss function, i and s are index variables respectively, is the i th sample in the dataset x i of the s th subset, represents the i th sample y i of the s th subset; The expression for the objective function of the deep time-frequency domain adaptive network is: Among them, represents the optimal parameter in the optimization process, θ is the network parameter, and argmin represents the minimization operation, is the objective function, is the calculation result of the final divergence, η represents the regularization parameter, λ represents the trade-off coefficient.
Citation Information
Patent Citations
Bearing fault diagnosis method and device adopting self-attention domain adaptive graph convolutional network
CN118410395A
Cross-domain fault sample translation generation method
CN118503802A