Rolling bearing cross-domain fault diagnosis method and device and storage medium
Through the cross-domain fault diagnosis method of rolling bearings based on transfer learning, the deep migration network model and wavelet time-frequency graph feature extraction are used to solve the problems of data scarcity and noise interference, improving the accuracy and generalization ability of fault diagnosis, and adapting to fault diagnosis under different working conditions.
Patent Information
- Application Number
- CN202510370286.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art faces the problems of scarce data, noise interference and low accuracy and generalization capabilities of fault diagnosis under complex working conditions in rolling bearing fault diagnosis.
A cross-domain fault diagnosis method based on transfer learning is adopted. By dividing the fault data into source domain data and target domain data, the deep migration network model is used for training, combining nearest neighbor data smoothing and wavelet time-frequency graph feature extraction, feature extraction and classification are performed, and the transfer learning is enhanced by norms.
It significantly improves the generalization ability and diagnostic accuracy of the model, reduces data labeling costs, improves the fault diagnosis ability under different operating conditions, reduces noise interference, and enhances the cross-domain learning ability of the model.
Smart Images

Figure CN120296506A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis, and particularly to a cross-domain fault diagnosis method, device and storage medium for rolling bearings. Background Art
[0002] The importance of rolling bearings for motors cannot be ignored. In the industrial and mechanical fields, motors are ubiquitous power sources, which are used to drive various devices and mechanical systems. Rolling bearings are one of the key components inside motors, and they play a crucial role in maintaining the normal operation and performance of motors. Health monitoring and fault prediction of them are an important application research direction in the field of fault diagnosis. In recent years, with the rapid progress of computing power and algorithms, intelligent fault diagnosis methods have shown significant potential in modeling the relationship between mechanical data and health status. Intelligent fault diagnosis is regarded as an important technology to ensure the stable and safe operation of mechanical equipment. Its application can effectively improve the operation efficiency, reduce the maintenance cost, and more importantly, effectively prevent the occurrence of safety accidents. In the field of fault diagnosis of rolling bearings, intelligent fault diagnosis methods have also been widely used in practice. Currently, the fault diagnosis methods for rolling bearings are divided into signal processing-based methods and artificial intelligence-based methods.
[0003] Traditional signal processing is a method that uses mathematical and engineering techniques to process the signals collected by sensors in order to extract useful information and features for fault diagnosis and monitoring. For the fault diagnosis of rolling bearings, traditional signal processing methods usually include techniques such as signal filtering, spectrum analysis, time-domain analysis, amplitude spectrum analysis, order analysis, etc., to detect and analyze the faults of rolling bearings from the signals collected by vibration, sound or other sensors. Traditional signal processing methods rely on mathematical theories and signal processing techniques, such as Fourier transform, wavelet transform, etc., to convert the original signals into frequency-domain or time-domain features. Traditional methods rely on predefined feature extraction and selection methods, which may not be able to effectively adapt to complex and changeable fault situations, and their generalization ability may be poor, and continuous adjustment and optimization are required. The success of this method depends to a large extent on the quality of feature engineering, and the selection of features is often based on the subjective judgment of domain experts, which may ignore some important features. If appropriate features are not selected, it may lead to a decline in fault diagnosis performance. Fault diagnosis methods based on artificial intelligence use artificial intelligence techniques such as machine learning and deep learning to automatically learn features and patterns from the data collected by sensors to diagnose the faults of rolling bearings. These methods usually involve using a large number of labeled data samples to train the model so that the model can make intelligent decisions and predictions. The success of deep learning depends on the independent and identically distributed assumption, that is, the data distributions of the training set and the test set are the same; however, in the real world, this assumption is often difficult to meet. In addition, fault diagnosis methods based on deep learning require a large amount of labeled data, but large-scale data annotation is an expensive and time-consuming operation. In real industrial scenarios, the diversity of production batches, structural designs, and operating conditions of bearings will result in different distributions of training data and test data. It is very difficult to collect enough labeled data to support the training of deep learning networks. Different data distributions will lead to insufficient generalization of deep learning models, resulting in a decline in fault diagnosis accuracy.
[0004] In addition, the field of fault diagnosis usually faces the problem of data scarcity, especially for some rare fault situations. At the same time, most current intelligent fault diagnosis methods still use one-dimensional signal data as training samples, which cannot fully utilize the performance of neural networks. At the same time, since the original vibration signals may be very complex, containing combinations of multiple frequencies and amplitudes, it makes it difficult for the model to directly learn information about faults from them. In addition, the original vibration signals may be affected by noise and other interferences, and these factors may make it difficult for the model to distinguish the true fault features.
[0005] The existing patent CN115705396A discloses a rolling bearing fault diagnosis method, which includes the following steps: First, obtain the vibration signal of the rolling bearing to be diagnosed and perform time-frequency conversion to obtain the frequency-domain signal; then input the frequency-domain signal into the constructed fault diagnosis model to diagnose the fault of the rolling bearing to be diagnosed; the fault diagnosis model includes a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a primary capsule layer, and a classification capsule layer connected in sequence. This method does not solve the problem that when directly using vibration signals to input into a neural network in the prior art, it may be difficult to capture periodic or transient features in the vibration signals, resulting in the model overfitting noise or the model performance degrading.
[0006] The existing patent CN106404396A discloses a rolling bearing fault diagnosis method, which includes obtaining the vibration signal of the rolling bearing; using the EMD empirical mode decomposition algorithm to decompose the vibration signal to obtain the decomposed signal; using PeakVue to extract the peak value, perform envelope detection, and perform time-frequency domain conversion on the decomposed signal to obtain the frequency spectrum signal; calculate the fault frequency of the rolling bearing; compare the frequency spectrum signal with the fault frequency to obtain the fault diagnosis result of the rolling bearing. In this method, the traditional EMD algorithm used may exhibit mode mixing when processing non-stationary signals, resulting in inaccurate signal features after decomposition, and further leading to a decrease in the fault diagnosis accuracy.
[0007] In summary, neither of the above two existing patents solves the problems of low accuracy and generalization ability in fault diagnosis in the face of data scarcity, noise interference, and complex working conditions in the field of fault diagnosis. Summary of the Invention
[0008] Based on the above technical problems, the present invention proposes a rolling bearing cross-domain fault diagnosis method, device, and storage medium to solve the problems of low accuracy and generalization ability in fault diagnosis in the face of data scarcity, noise interference, and complex working conditions in the field of fault diagnosis.
[0009] A rolling bearing cross-domain fault diagnosis method includes:
[0010] Collect the fault data of the rolling bearing and divide the fault data into source domain data and target domain data;
[0011] Based on the source domain data and target domain data, determine the corresponding source domain dataset and target domain dataset respectively;
[0012] Use the source domain dataset and target domain dataset to train the deep transfer network model;
[0013] Perform fault diagnosis on the rolling bearing through the trained deep transfer network model.
[0014] Further, collect the fault data of the rolling bearing, including:
[0015] Collect the fault data of the rolling bearing with different fault types under different working conditions. The fault data includes vibration acceleration data.
[0016] Further, the fault types include one or more of inner race fault, rolling element fault and outer race fault.
[0017] Further, before respectively determining the corresponding source domain dataset and target domain dataset based on the source domain data and the target domain data, it further includes:
[0018] Perform preprocessing operations on the source domain data and the target domain data respectively. The preprocessing operations include one or more of removing the DC component, smoothing the neighboring data, and data augmentation operations.
[0019] Further, the smoothing of the neighboring data corresponds to Formula 1, where S s (t) k is the data of the k-th sampling point after smoothing the neighboring data, L is the window length of the neighboring data smoothing algorithm, S s′ (t) is the source domain data or the target domain data, i represents the i-th data point of the signal, and j represents the symmetric index offset relative to the center point k in the current sliding window.
[0020] Further, it further includes:
[0021] Do not perform the smoothing of the neighboring data on the starting sampling point of the signal and the k sampling points after the starting sampling point, and the ending sampling point and the k sampling points before the ending sampling point, where ceil() represents the floor operation.
[0022] Further, the data augmentation operation includes:
[0023] Use the preset sliding window to perform overlapping sampling on the source domain data and the target domain data respectively. The number of sampling points of the overlapping signal satisfies the expression where f s is the sampling frequency of the signal, K is the number of sampling points of the signal to be segmented, T s is the time resolution, T f is the frequency resolution of the signal to be segmented, and n is the rotational speed of the rolling bearing.
[0024] Further, respectively determining the corresponding source domain dataset and target domain dataset based on the source domain data and the target domain data includes:
[0025] Perform continuous wavelet transform on the source domain data and the target domain data respectively, extract the time-frequency domain features and generate the corresponding wavelet time-frequency diagrams.
[0026] Generate corresponding source domain dataset and target domain dataset based on the wavelet time-frequency diagram.
[0027] Furthermore, use the source domain dataset and the target domain dataset to train the deep transfer network model, including:
[0028] Input the source domain dataset and the target domain dataset into the deep transfer network model;
[0029] Use the deep transfer network model to extract features from the source domain dataset and the target domain dataset respectively to obtain source domain data features and target domain data features;
[0030] Determine the corresponding classification loss based on the source domain data features;
[0031] Determine the transfer loss based on the target domain data features;
[0032] Determine the total loss of the deep transfer network model according to the classification loss and the transfer loss;
[0033] Train the deep transfer network model based on the total loss.
[0034] Furthermore, the deep transfer network model includes: a feature extraction module, a transfer learning module, and a classification module.
[0035] Furthermore, use the deep transfer network model to extract features from the source domain dataset and the target domain dataset respectively, including:
[0036] Use the feature extraction module of the deep transfer network model to extract features from the source domain dataset and the target domain dataset respectively.
[0037] Furthermore, the feature extraction module of the deep transfer network model uses ResNet18 as the backbone network.
[0038] Furthermore, determine the corresponding classification loss based on the source domain data features, including:
[0039] Based on the source domain data features, determine the classification loss through the classification module of the deep transfer network model.
[0040] Furthermore, determine the transfer loss based on the target domain data features, including:
[0041] Based on the target domain data features, determine the transfer loss through the transfer learning module of the deep transfer network model.
[0042] Furthermore, based on the source domain data features, determine the classification loss through the classification module of the deep transfer network model, including:
[0043] Based on the source domain data features, the classification loss is determined using the cross-entropy loss function of the classification module. The cross-entropy loss function is represented by Equation (2), Equation (2), where, is the classification loss, R is the number of rows of the source domain output classification matrix P(X s ) and C is the number of columns of the source domain output classification matrix P(X s ). y i,j is the true class label of the element, and P(X s ) i,j is the element in the i-th row and j-th column of P(X s ).
[0044] Furthermore, based on the target domain data features, the transfer loss is determined through the transfer learning module of the deep transfer network model, including:
[0045] Based on the target domain data features, the norm-enhanced transfer loss function of the transfer learning module is used to determine the transfer loss. The norm-enhanced transfer loss function is represented by Equation (3), Equation (3), where, is the transfer loss, B t is the batch size of the input samples, P(X t ) is the output classification matrix of the target domain, X t is the target domain input sample, and n is the number of samples.
[0046] Furthermore, according to the classification loss and the transfer loss, the total loss of the deep transfer network model is determined, including:
[0047] The sum of the classification loss and the transfer loss is used as the total loss of the deep transfer network model.
[0048] Furthermore, the total loss is represented by Equation (4), Equation (4) is,
[0049]
[0050] where γ ∈ [0.1, 0.5, 1, 1.5, 2], δ represents the classification loss weight factor, and γ represents the trade-off factor.
[0051] Furthermore, using the source domain dataset and the target domain dataset to train the deep transfer network model also includes:
[0052] Dividing the target domain dataset into a target domain training set and a target domain test set;
[0053] Using the target domain training set and the source domain dataset to train the deep transfer network model;
[0054] Using the target domain test set to evaluate the effect of the trained deep transfer network model.
[0055] Furthermore, it also includes: using the source-only method and the deep transfer method DCORAL to compare and verify the effectiveness of the trained deep transfer network model.
[0056] A cross-domain fault diagnosis device for rolling bearings, which is used to execute according to the above method. The device includes:
[0057] An acquisition module, which is used to acquire the fault data of the rolling bearing and divide the fault data into source domain data and target domain data;
[0058] A determination module, which is used to determine the corresponding source domain data set and target domain data set based on the source domain data and the target domain data respectively;
[0059] A training module, which is used to train the deep transfer network model by using the source domain data set and the target domain data set;
[0060] A diagnosis module, which is used to perform fault diagnosis on the rolling bearing through the trained deep transfer network model.
[0061] A computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein the computer program can execute the above method when being run by an electronic device.
[0062] A computer program product, including a computer program, which implements the steps of the above method when being executed by a processor.
[0063] An electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the method through the computer program.
[0064] Based on the above technical solutions, the present invention has at least the following beneficial effects:
[0065] 1. The present invention adopts a cross-domain fault diagnosis method for rolling bearings based on transfer learning, trains a deep transfer network model by using a source domain data set and a target domain data set, transfers the acquired knowledge from the source domain to the target domain, can effectively solve problems such as data scarcity and domain differences in fault diagnosis, extracts the features of the target domain data in a targeted manner by using the feature representations of the source domain data and the model, thereby significantly improving the generalization ability and diagnosis accuracy of the model. Using the trained deep transfer network model for fault diagnosis of rolling bearings can reduce the data annotation cost and accelerate the convergence process of the model, providing support for accurate fault diagnosis.
[0066] 2. The present invention uses a nearest neighbor data smoothing algorithm for preprocessing, which can effectively reduce the interference of noise in the bearing vibration signal. The wavelet time-frequency diagram is used as the input of the deep migration network. Compared with directly inputting the bearing time-domain vibration signal, using the wavelet time-frequency diagram as input can enable the neural network to better capture the time-frequency characteristics of the signal. The feature extraction method of the wavelet time-frequency diagram is more abstract and universal than directly inputting the vibration signal, and can better adapt to the input signals under different working conditions, thereby improving the generalization ability and cross-domain fault diagnosis ability of the deep migration network model.
[0067] 3. The fault diagnosis method used in the present invention performs transfer learning from the perspective of structural adaptation through norm-enhanced transfer loss. Compared with the data distribution adaptation method (MMD, CORAL, etc.), it does not require additional distance metrics, and in the process of domain alignment, it can improve the prediction diversity and prediction discriminability of the diagnostic model for the target domain data, and has strong cross-domain learning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0069] Figure 1 A flow chart of a rolling bearing cross-domain fault diagnosis method according to an embodiment of the present invention;
[0070] Figure 2 A schematic diagram of using a sliding window for data enhancement in one embodiment of the present invention;
[0071] Figure 3 A wavelet time-frequency diagram generated for one embodiment of the present invention;
[0072] Figure 4 This is a schematic diagram of the structure of the ResNet18 residual block;
[0073] Figure 5 Schematic diagram of the dual-stream structure of the deep migration network model;
[0074] Figure 6 A schematic diagram of the overall structure of a deep migration network model constructed for an embodiment of the present invention;
[0075] Figure 7 This is a t-SNE schematic diagram of an A→B migration task using the method of the present invention in one embodiment of the present invention;
[0076] Figure 8 A schematic diagram of a rolling bearing cross-domain fault diagnosis device according to an embodiment of the present invention;
[0077] Figure 9 It is a block diagram of a computer system for an electronic device implementing the embodiments of the present invention;
[0078] Figure 10 It is a schematic diagram of an electronic device for cross - domain fault diagnosis of rolling bearings according to an embodiment of the present invention. Detailed implementation manners
[0079] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0080] The following further describes the present invention in detail with specific embodiments, and these embodiments should not be construed as limiting the scope claimed by the present invention.
[0081] Embodiment
[0082] To solve the problems of low accuracy and generalization ability of fault diagnosis in the face of data scarcity, noise interference, and complex working conditions in the field of fault diagnosis, the present invention proposes a cross - domain fault diagnosis method, device, and storage medium for rolling bearings.
[0083] As Figure 1 shows a flowchart of a cross - domain fault diagnosis method for rolling bearings according to an embodiment of the present invention, and the process includes the following steps:
[0084] S1, collect the fault data of the rolling bearing, and divide the fault data into source - domain data and target - domain data.
[0085] Furthermore, collecting the fault data of the rolling bearing includes: collecting the fault data of rolling bearings with different fault types under different working conditions, and the fault data includes vibration acceleration data. Further, the fault types include one or more of inner - race faults, rolling - element faults, and outer - race faults.
[0086] It should be understood that the fault data of rolling bearings of different fault types under different working conditions in this step can be obtained by actual sensors or by obtaining existing public data sets. The CWRU bearing data set obtained in this embodiment is a famous public benchmark data set provided by Case Western Reserve University. This data set provides vibration signals of 6205 bearings under different working conditions and different fault types. The 6205 bearing is a widely used motor bearing. In this embodiment, the vibration signal of the 6205 bearing at the drive end at 12KHz is selected as a verification case. Nine fault classes are combined with 3 fault types and 3 fault sizes. The 3 fault classes are inner race fault (IRF), ball fault (BF), and outer race fault (ORF). The 3 fault sizes are 0.007 inch (S: small), 0.014 inch (M: medium), and 0.021 inch (L: large). The combined fault classes are SIRF, SBF, SORF, MIRF, MBF, MORF, LIRF, LBF, and LORF. Ten different samples composed of 9 fault classes and normal classes under working conditions of 3HP, 0H, and 2HP are used. Specific data examples are shown in Table 1.
[0087] Table 1 Data set settings under different working conditions
[0088]
[0089] S2, respectively determine the corresponding source domain data set and target domain data set based on the source domain data and the target domain data.
[0090] Furthermore, before respectively determining the corresponding source domain data set and target domain data set based on the source domain data and the target domain data, it further includes: respectively performing preprocessing operations on the source domain data and the target domain data. The preprocessing operations include one or more of removing the DC component, smoothing the neighboring data, and data augmentation operations.
[0091] The preprocessing operations adopted in this embodiment include removing the DC component, smoothing the neighboring data, and data augmentation operations. Specifically, in this embodiment, the DC component is removed from the source domain data and the target domain data respectively to remove the DC component in the signal and retain the AC part of the signal, which helps to better analyze and diagnose the dynamic characteristics and fault information in the signal; then, the neighboring data smoothing process is performed based on the signal after removing the DC component; finally, the data augmentation operation is performed on the data after the neighboring data smoothing process.
[0092] Taking the source domain data as an example, the formula for removing the DC component is where S s′ (t) represents the data after removing the DC component, Denote the source domain data as $N$, where $N$ represents the source domain data There are a total of $N$ data points in it.
[0093] In order to reduce the impact of noise in the signal on the performance of the depth migration network, perform neighborhood data smoothing on the signals after removing the DC component in the source domain and the target domain. Given the window length of neighborhood data smoothing as $L$, calculate the average value of the data within the window, and use this average value as the value of the data point at the center of the window. Then move the window by 1 sampling point and continue this operation. The core idea of neighborhood data smoothing is to use the average value of the data points within the window to approximate the trend of the signal within this window, thereby smoothing out the influence of random noise.
[0094] Furthermore, the neighborhood data smoothing process corresponds to Formula 1 where $S$ s $(t)$ k is the data of the $k$-th sampling point after neighborhood data smoothing, $L$ is the window length of the neighborhood data smoothing algorithm, and the length of $L$ satisfies $L = 2n + 1$, $n\in Z$, $S$ s′ $(t)$ is the source domain data or the target domain data, $i$ represents the $i$-th data point of the signal, and $j$ represents the symmetric index offset relative to the center point $k$ in the current sliding window. In this embodiment, $S$ s′ $(t)$ is the data after removing the DC component.
[0095] Furthermore, since the sampling points near the starting sampling point and the ending sampling point of the signal sometimes do not meet the requirements of the neighborhood data smoothing algorithm, the present invention still uses the original data for these signal sampling points. Specifically, in this embodiment, the starting sampling point of the signal and the $k$ sampling points after the starting sampling point, the ending sampling point and the $k$ sampling points before the ending sampling point are not subjected to neighborhood data smoothing, where ceil() represents the floor operation.
[0096] Signals usually contain different local features in different time periods. Using a sliding window can capture the changes of the signal at different time scales, which helps the model better understand the dynamic properties of the signal. Perform data augmentation on the source domain data and the target domain data respectively. By using a sliding window to perform signal segmentation, more samples can be generated, increasing the amount of training data.
[0097] Furthermore, the data augmentation operation includes: using a preset sliding window to perform overlapping sampling on the source domain data and the target domain data respectively, and the number of sampling points of the overlapping signal satisfies the expression where $f$ s is the sampling frequency of the signal, $K$ is the number of sampling points of the signal to be segmented, $T$ s is the time resolution, $T$ fis the frequency resolution of the signal to be segmented, and n is the rotational speed of the rolling bearing.
[0098] Taking the source domain data as an example, to sample using a sliding window of length L, the sliding window moves a distance of D each time, and there is partial overlap between the signal obtained by overlapping sampling and the signal after it. The samples of the overlapping signal are represented represents the m-th overlapping sampling sample. At the same time, to ensure that each sample during overlapping sampling contains at least one week's worth of fault feature information of the bearing rotation, the number of sampling points in the sample should be no less than the number of sampling points within one week of the bearing rotation time. In this embodiment, as Figure 2 shown, sample using a sliding window of length 2048, the sliding window moves a distance of 350 each time, and there is partial overlap between the signal obtained by overlapping sampling and the signal after it. For each type of fault of the bearing under each working condition, 200 signal segments are obtained through overlapping sampling, totaling 2000 signal samples.
[0099] Further, based on the source domain data and the target domain data, the corresponding source domain dataset and target domain dataset are determined respectively, including the following two steps:
[0100] (1) Perform continuous wavelet transform on the source domain data and the target domain data respectively, extract time-frequency domain features, and generate corresponding wavelet time-frequency diagrams.
[0101] The following will take the source domain data as an example to introduce the process of performing continuous wavelet transform on the data in this embodiment, extracting time-frequency domain features, and generating corresponding wavelet time-frequency diagrams. Specifically, use the wavelet analysis toolbox PyWavelets of python to perform continuous wavelet transform on the samples after data enhancement to obtain a wavelet coefficient matrix. The specific process of continuous wavelet transform is:
[0102]
[0103] where W(a,b) is the wavelet coefficient matrix, a is the scale factor, b is the time shift factor, * represents the conjugate function, and ψ(t) is the wavelet basis function.
[0104] Further, the present invention selects the complex Morlet wavelet as the wavelet basis function. The complex Morlet wavelet is a commonly used wavelet basis for continuous wavelet transform, which can achieve good localization characteristics in both the time domain and the frequency domain. This enables it to accurately locate transient and periodic features in the signal in terms of time and frequency, and is very effective for capturing the impact signals and vibration period changes generated by rolling bearing faults. At the same time, it has an exponentially decaying oscillation form, which is similar to the decay of the transient impact generated when the rolling bearing fails. Therefore, it is often used for fault feature extraction of rolling bearings, and its expression is:
[0105]
[0106] Among them, B is the bandwidth of the complex Morlet wavelet, and C is the central frequency of the complex Morlet wavelet. It should be understood that in other embodiments, other functions can be selected as the wavelet basis function according to the actual situation.
[0107] Furthermore, the actual frequency of the signal has the following relationship with the scale factor a: Among them, F c represents the central frequency of the wavelet basis function, f s represents the sampling frequency of the original signal, and F a represents the frequency corresponding to the signal scale. According to the sampling theorem, in order to make the frequency range of the wavelet time-frequency diagram be (0, 0.5f s ), and the scale range be (2F c , ∞), the length of the scale sequence is selected to be 256 in this embodiment.
[0108] Taking the data set of working condition A in Table 1 as an example in this embodiment, based on the wavelet coefficient matrix obtained by the above continuous wavelet transform, the imshow() function in the Matplotlib plotting library of Python is used to generate a wavelet time-frequency diagram from the wavelet coefficient matrix and save it as an RGB image with a pixel size of 224*224. The generated wavelet time-frequency diagram is as Figure 3 shown, Figure 3 in which the wavelet time-frequency diagrams corresponding to 9 fault classes and the normal type are shown. It should be understood that the method of generating a wavelet time-frequency diagram from the wavelet coefficient matrix is not limited to the manner and image format mentioned in the above embodiment.
[0109] (2) Generate the corresponding source domain data set and target domain data set based on the wavelet time-frequency diagram.
[0110] Specifically, in this embodiment, the ImageFolder class in the Torchvision library is used to generate the corresponding source domain data set and target domain data set from the generated wavelet time-frequency diagram. Then, the ToTensor class in the Torchvision library is used to normalize the wavelet time-frequency diagram to obtain tensors with a numerical range of [0, 1], and these tensors form a labeled source domain data set and an unlabeled target domain data set where N s represents the number of samples in the source domain data set, and N t represents the number of samples in the target domain data set.
[0111] S3. Use the source domain data set and the target domain data set to train the deep transfer network model.
[0112] The deep transfer network model includes: a feature extraction module, a transfer learning module, and a classification module.
[0113] Further, training the deep transfer network model using a source domain dataset and a target domain dataset includes:
[0114] S301, inputting the source domain dataset and the target domain dataset into the deep transfer network model.
[0115] S302, using the deep transfer network model to respectively extract features from the source domain dataset and the target domain dataset to obtain source domain data features and target domain data features.
[0116] Further, using the deep transfer network model to respectively extract features from the source domain dataset and the target domain dataset includes: using the feature extraction module of the deep transfer network model to respectively extract features from the source domain dataset and the target domain dataset. Specifically, in this embodiment, the PyTorch library is used to train the deep transfer network model, and the feature extraction module of the deep transfer network model uses ResNet18 as the backbone network. ResNet18 mainly consists of Figure 4 the two types of residual blocks shown. In addition, the deep transfer network model uses Figure 5 the two-stream structure shown. The two-stream structure of the deep transfer network model makes full use of the knowledge of the source domain and the data of the target domain, so as to achieve better performance on the target task. When training the deep transfer network model, the target domain dataset and the source domain dataset are input simultaneously. Further, in this embodiment, the output dimension of the deep transfer network model is 32 * 512, and the data dimension output after passing through the bottleneck layer is 32 * 512. Subsequently, the feature extraction module is used to respectively extract features from the source domain dataset and the target domain dataset to obtain source domain data features and target domain data features. Figure 6 It is a schematic diagram of the overall structure of the deep transfer network model constructed in an embodiment of the present invention.
[0117] The target domain dataset and the source domain dataset share part of the network structure and calculate losses respectively. Specifically, the classification loss is calculated according to step S303 using the data features of the source domain, and the transfer loss is calculated according to step S304 using the data features of the target domain.
[0118] S303, determining the corresponding classification loss based on the source domain data features.
[0119] Further, determining the corresponding classification loss based on the source domain data features includes: determining the classification loss through the classification module of the deep transfer network model based on the source domain data features.
[0120] Specifically, based on the source domain data features, this embodiment determines the classification loss through the classification module of the deep transfer network model, including the following process: Based on the source domain data features, the cross-entropy loss function of the classification module is used to determine the classification loss, and the cross-entropy loss function is represented by Formula 2. Formula 2, where, is the classification loss, R is the number of rows of the source domain output classification matrix P(X s ), C is the number of columns of the source domain output classification matrix P(X s ), y i,j is the true class label of the element, and P(X s ) i,j is the element in the i-th row and j-th column of P(X s ).
[0121] S304. Determine the transfer loss based on the target domain data features.
[0122] Furthermore, determining the transfer loss based on the target domain data features includes: Based on the target domain data features, the transfer learning module of the deep transfer network model is used to determine the transfer loss.
[0123] The core problem of using domain adaptation for transfer learning is to find a regularization term with excellent prediction discriminability and prediction diversity. Currently, commonly used transfer regularization terms mostly use explicit or implicit distance metrics to reduce the data distribution difference between the source domain and the target domain. However, if the network structure itself can handle the data distribution difference, there is no need for an additional distance metric and it is simpler and easier to use. Therefore, the present invention uses batch kernel norm maximization as the transfer regularization term. The transfer learning module uses the Norm-Enhanced Transfer Loss (NET Loss) as the core method. In the case of insufficient labels, the difference between the labeled data and the unlabeled data may cause a high-density data region to appear near the decision boundary of a specific task, resulting in a decrease in the certainty of prediction. Currently, the method of minimizing Shannon entropy is commonly used to improve the prediction discriminability, but this method will cause the minority classes to be more likely to be classified into the majority classes. There is a strict inverse monotonicity between the Frobenius norm of a matrix and Shannon entropy, and we can achieve the same effect by maximizing the Frobenius norm. Let the dimension of the batch output matrix A of the network be B*C, where B is the batch size and C is the number of classes, then the upper bound of ∥A∥ F is:
[0124]
[0125] where, ∥A∥ F is the Frobenius norm, j ∈ [0, C].
[0126] The prediction diversity can be measured by the number of categories in the batch output matrix A. Assume A m , A n are the row vectors of the batch output matrix A. When the predicted categories of the two are different, A m and A n can be regarded as linearly independent; when the predicted categories of the two are the same and ∥A∥ F is close to , the two can be regarded as linearly related. The more the number of linearly independent row vectors, the richer the prediction diversity of A. The matrix rank is the sum of the number of linearly independent row vectors in the matrix. Therefore, the prediction diversity can be improved by maximizing the rank of matrix A, rank(A).
[0127] Calculating the rank of a positive definite matrix is an NP-hard non-convex problem. Relevant theories show that when ∥A∥ F ≤1, the convex envelope of rank(A) is the nuclear norm ∥A∥ * . The calculation formula of the nuclear norm of a matrix is as follows: where D = min{B, C}, and σ k represents the k-th largest singular value of matrix A. When ∥A∥ F is close to the upper bound, the convex envelope of rank(A) becomes and is proportional to ∥A∥ * . Therefore, the prediction diversity can be increased by maximizing ∥A∥ * .
[0128] The range relationship between ∥A∥ * and ∥A∥ F is expressed in the following form:
[0129]
[0130] The above relational expression shows that ∥A∥ * and ∥A∥ F constrain each other. Increasing ∥A∥ * will lead to an increase in ∥A∥ F . Therefore, the discriminability of the prediction can be increased by maximizing ∥A∥ * .
[0131] To sum up, maximizing ∥A∥ * can improve the discriminability and diversity of the prediction. Therefore, the present invention introduces batch nuclear norm maximization as a transfer regularization term. Assume X t is the input sample of the target domain, B t is the batch number of the input samples, the output classification matrix of the target domain is P(X t ), ∥.∥ *Denote the nuclear norm of the matrix, then the NET transfer loss function can be expressed as Equation 3:
[0132]
[0133] Specifically, based on the target domain data features, this embodiment determines the transfer loss through the transfer learning module of the deep transfer network model, including the following process: Based on the target domain data features, use the norm-enhanced transfer loss function of the transfer learning module to determine the transfer loss. The norm-enhanced transfer loss function is represented by Equation 3, Equation 3, where, is the transfer loss, B t is the batch number of input samples, P(X t ) is the output classification matrix of the target domain, X t is the input sample of the target domain, and n is the number of samples.
[0134] S305. Determine the total loss of the deep transfer network model according to the classification loss and the transfer loss.
[0135] Furthermore, determining the total loss of the deep transfer network model according to the classification loss and the transfer loss includes: taking the sum of the classification loss and the transfer loss as the total loss of the deep transfer network model. It should be understood that in other embodiments of the present invention, other operations may also be performed on the classification loss and the transfer loss to obtain the total loss of the deep transfer network model.
[0136] Specifically, the total loss in this embodiment is represented by Equation 4, and Equation 4 is,
[0137]
[0138] where γ ∈ [0.1, 0.5, 1, 1.5, 2], δ represents the classification loss weight factor, and γ represents the trade-off factor. The trade-off factor in this embodiment is 1.
[0139] S306. Train the deep transfer network model based on the total loss.
[0140] When the total loss value reaches the convergence state, the training of the deep transfer network model is completed.
[0141] Further, training the deep transfer network model using the source domain dataset and the target domain dataset further includes: before step S301, dividing the target domain dataset into a target domain training set and a target domain test set; subsequently, using the target domain training set and the source domain dataset to train the deep transfer network model, this process corresponding to steps S302 to S303 above; finally, using the target domain test set to evaluate the effect of the trained deep transfer network model. Specifically, in this embodiment, the StratifiedShuffleSplit class in the sklearn library is used to stratify and divide the target domain dataset into a training set and a test set at a ratio of 9:1 to ensure that the number of samples of each category in the test set is balanced, and then the Dataloader class in the Pytorch library is used to load data to train the deep transfer network model and train the deep transfer network model.
[0142] Further, it further includes: using the source-only method and the deep transfer method DCORAL to conduct a comparative verification on the effect of the trained deep transfer network model.
[0143] To verify the effect of the trained deep transfer network model, this embodiment uses the source-only method and the deep transfer method DCORAL to compare the effects of the trained deep transfer network model. Specifically, in this embodiment, three working conditions are combined to obtain a total of 6 transfer tasks to verify the cross-domain diagnosis ability of the present invention between different working conditions. Using A→B to represent the transfer task, A represents the source domain, and B represents the target domain. The method of the present invention is compared with the source-only method without using transfer learning and the deep transfer method DCORAL at the same time, and all comparison methods use the same hyperparameters. Set the number of training iterations epoch to 50, the learning rate to 0.001, the weight decay to 0.001, and use the SGD optimizer. To prevent overfitting during the training process, an early stopping method with an interval of 30 is used. To accelerate the convergence speed of the model and reduce model oscillation, this embodiment uses learning rate decay, where x represents the current epoch and lr i represents the initial learning rate. To reduce the randomness of the results, each method is trained 4 times under 4 different random number seeds [42, 84, 126, 168] to calculate the average accuracy and variance, and the deep transfer network model with the highest diagnostic accuracy is saved.
[0144] The specific experimental results are shown in Table 2. The method proposed by the present invention achieved an average accuracy of 98.87% in 6 transfer tasks of the Case Western Reserve University dataset. Compared with other comparison methods, the method of the present invention has a higher diagnostic accuracy. This shows that the method of the present invention has fully learned domain-invariant features and has strong cross-domain generalization ability.
[0145] Table 2 Average accuracy ± variance of different methods
[0146]
[0147] In addition, to visually demonstrate the performance of the method proposed in the present invention, in this embodiment, taking the A→B migration task as an example, a confusion matrix is used to evaluate the fault diagnosis ability of different methods, and the visualization image of t-Stochastic Neighbor Embedding (t-SNE) is drawn using the TSNE class in sklearn. t-SNE can map the high-dimensional representation of the model for the dataset to a low-dimensional space and observe the distribution of data points in the low-dimensional space, which helps to analyze whether the model can separate data points of different categories and whether there are overlapping regions between categories. The specific visualization results are as Figure 7 shown. The results of the confusion matrix show that the method proposed in the present invention has the highest fault diagnosis accuracy. The t-SNE results show that the boundaries of fault samples of different categories are clear, while the distances between fault samples of the same category are compact. The visualization results of other comparison methods show that the fault diagnosis ability for the inner race fault of the bearing is weak, and there is confusion for the inner race faults with different degrees of faults. The visualization results show that the method of the present invention fully learns domain-invariant features, realizes the transfer of knowledge from the source domain to the target domain, and shows good performance in the cross-domain fault diagnosis scenario.
[0148] S4. Fault diagnosis of the rolling bearing is performed through the trained deep transfer network model.
[0149] To achieve the same purpose as the above method, the present invention also proposes a rolling bearing cross-domain fault diagnosis device.
[0150] As Figure 8 shown in the schematic diagram of a rolling bearing cross-domain fault diagnosis device according to an embodiment of the present invention, the device includes: a collection module 81, a determination module 82, a training module 83, and a diagnosis module 84.
[0151] The collection module 81 is used to collect the fault data of the rolling bearing and divide the fault data into source domain data and target domain data.
[0152] The determination module 82 is used to respectively determine the corresponding source domain dataset and target domain dataset based on the source domain data and the target domain data.
[0153] The training module 83 is used to train the deep transfer network model using the source domain dataset and the target domain dataset.
[0154] The diagnosis module 84 is used to perform fault diagnosis on the rolling bearing through the trained deep transfer network model.
[0155] As an alternative solution, the above-mentioned device is also used for: collecting fault data of rolling bearings with different fault types under different working conditions, where the fault data includes vibration acceleration data.
[0156] As an alternative solution, the fault types in the above-mentioned device include one or more of inner ring faults, rolling element faults, and outer ring faults.
[0157] As an alternative solution, the above-mentioned device is also used for: before respectively determining the corresponding source domain data set and target domain data set based on the source domain data and target domain data, it further includes:
[0158] Performing preprocessing operations on the source domain data and target domain data respectively, where the preprocessing operations include one or more of removing the DC component, nearest neighbor data smoothing, and data augmentation operations.
[0159] As an alternative solution, the above-mentioned device is also used for: the corresponding formula for nearest neighbor data smoothing
[0160] One, where S s (t) k is the data of the k-th sampling point after nearest neighbor data smoothing, L is the window length of the nearest neighbor data smoothing algorithm, S s′ (t) is the source domain data or target domain data, i represents the i-th data point of the signal, and j represents the symmetric index offset relative to the center point k in the current sliding window.
[0161] As an alternative solution, the above-mentioned device is also used for: not performing nearest neighbor data smoothing on the starting sampling point of the signal and the k sampling points after the starting sampling point, and the ending sampling point and the k sampling points before the ending sampling point, where ceil() represents the floor operation.
[0162] As an alternative solution, the above-mentioned device is also used for: data augmentation operations, including:
[0163] Using a preset sliding window to perform overlapping sampling on the source domain data and target domain data respectively, and the number of sampling points of the overlapping signal satisfies the expression where f s is the sampling frequency of the signal, K is the number of sampling points of the signal to be segmented, T s is the time resolution, T f is the frequency resolution of the signal to be segmented, and n is the rotational speed of the rolling bearing.
[0164] As an alternative solution, the above-mentioned device is also used for: respectively determining the corresponding source domain data set and target domain data set based on the source domain data and target domain data, including:
[0165] Perform continuous wavelet transform on the source domain data and the target domain data respectively, extract time-frequency domain features, and generate corresponding wavelet time-frequency diagrams.
[0166] Generate corresponding source domain datasets and target domain datasets based on the wavelet time-frequency diagrams.
[0167] As an optional solution, the above device is also used for: training a deep transfer network model using the source domain datasets and the target domain datasets, including:
[0168] Input the source domain datasets and the target domain datasets into the deep transfer network model;
[0169] Use the deep transfer network model to perform feature extraction on the source domain datasets and the target domain datasets respectively to obtain source domain data features and target domain data features;
[0170] Determine the corresponding classification loss based on the source domain data features;
[0171] Determine the transfer loss based on the target domain data features;
[0172] Determine the total loss of the deep transfer network model according to the classification loss and the transfer loss;
[0173] Train the deep transfer network model based on the total loss.
[0174] As an optional solution, the above device is also used for: the deep transfer network model, including: a feature extraction module, a transfer learning module, and a classification module.
[0175] As an optional solution, the above device is also used for: performing feature extraction on the source domain datasets and the target domain datasets respectively using the deep transfer network model, including:
[0176] Use the feature extraction module of the deep transfer network model to perform feature extraction on the source domain datasets and the target domain datasets respectively.
[0177] As an optional solution, the feature extraction module of the above device uses ResNet18 as the backbone network.
[0178] As an optional solution, the above device is also used for: determining the corresponding classification loss based on the source domain data features, including:
[0179] Based on the source domain data features, determine the classification loss through the classification module of the deep transfer network model.
[0180] As an optional solution, the above device is also used for: determining the transfer loss based on the target domain data features, including:
[0181] Based on the characteristics of the target domain data, determine the transfer loss through the transfer learning module of the deep transfer network model.
[0182] As an alternative solution, the above device is also used to: based on the characteristics of the source domain data, determine the classification loss through the classification module of the deep transfer network model, including:
[0183] Based on the characteristics of the source domain data, use the cross-entropy loss function of the classification module to determine the classification loss, and the cross-entropy loss function is represented by Formula 2, Formula 2, where, is the classification loss, R is the number of rows of the source domain output classification matrix P(X s ) and C is the number of columns of the source domain output classification matrix P(X s ), y i,j is the true class label of the element, and p(x s ) i,j is the element in the i-th row and j-th column of p(X s ).
[0184] As an alternative solution, the above device is also used to: based on the characteristics of the target domain data, determine the transfer loss through the transfer learning module of the deep transfer network model, including:
[0185] Based on the characteristics of the target domain data, use the norm-enhanced transfer loss function of the transfer learning module to determine the transfer loss, and the norm-enhanced transfer loss function is represented by Formula 3, Formula 3, where, is the transfer loss, B t is the number of batches of input samples, p(x t ) is the output classification matrix of the target domain, x t is the input sample of the target domain, and n is the number of samples.
[0186] As an alternative solution, the above device is also used to: determine the total loss of the deep transfer network model according to the classification loss and the transfer loss, including:
[0187] Take the sum of the classification loss and the transfer loss as the total loss of the deep transfer network model.
[0188] As an alternative solution, the above device is also used to: the total loss is represented by Formula 4, and Formula 4 is,
[0189]
[0190] where γ ∈ [0.1, 0.5, 1, 1.5, 2], δ represents the classification loss weight factor, and γ represents the trade-off factor.
[0191] As an alternative, the above-mentioned device is also used for: training a deep transfer network model using a source domain dataset and a target domain dataset, and further includes:
[0192] Dividing the target domain dataset into a target domain training set and a target domain test set;
[0193] Training the deep transfer network model using the target domain training set and the source domain dataset;
[0194] Evaluating the effect of the trained deep transfer network model using the target domain test set.
[0195] As an alternative, the above-mentioned device is also used for: further includes: comparing and verifying the effect of the trained deep transfer network model using the source-only method and the deep transfer method DCORAL.
[0196] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0197] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0198] According to one aspect of the present application, there is provided a computer program product, which includes a computer program.
[0199] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0200] Figure 9 Schematically shows a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.
[0201] It should be noted that Figure 9 The computer system 1100 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0202] Such as Figure 9As shown, the computer system 1100 includes a central processing unit 1101 (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1102 (ROM) or a program loaded from a storage section 1108 into a random access memory 1103 (RAM). In the random access memory 1103, various programs and data required for system operation are also stored. The central processing unit 1101, the read-only memory 1102, and the random access memory 1103 are connected to each other via a bus 1104. An input / output interface 1105 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1104.
[0203] The following components are connected to the input / output interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a local area network card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from it can be installed into the storage section 1108 as needed.
[0204] Specifically, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1109, and / or installed from the removable medium 1111. When the computer program is executed by the central processing unit 1101, various functions defined in the system of the present application are executed.
[0205] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1109, and / or installed from the removable medium 1111. When the computer program is executed by the central processing unit 1101, various functions provided by the embodiment of the present application are executed.
[0206] According to another aspect of the embodiments of the present application, an electronic device for cross-domain fault diagnosis of a rolling bearing is further provided. In this embodiment, the electronic device is taken as an example of a terminal device for illustration. As Figure 10 shown, the electronic device includes a memory 1202 and a processor 1204. A computer program is stored in the memory 1202, and the processor 1204 is configured to execute the steps in any of the above method embodiments through the computer program.
[0207] Optionally, in this embodiment, the above electronic device may be at least one network device among multiple network devices in a computer network.
[0208] Optionally, in this embodiment, the above processor may be configured to execute the methods in the embodiments of the present application through a computer program.
[0209] Optionally, those of ordinary skill in the art can understand that Figure 10 the structure shown is only schematic, Figure 10 and it does not limit the structure of the above electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown Figure 10 in the figure, or have a different configuration from that shown Figure 10 in the figure.
[0210] Among them, the memory 1202 can be used to store software programs and modules, such as program instructions / modules corresponding to a cross-domain fault diagnosis method and device for a rolling bearing in the embodiments of the present application. The processor 1204 executes various functional applications and data processing by running the software programs and modules stored in the memory 1202, that is, implements the above cross-domain fault diagnosis method for a rolling bearing. The memory 1202 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1202 may further include a memory remotely disposed relative to the processor 1204, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 1202 can specifically but not limitedly be used to store target domain and source domain data information. As an example, as Figure 10 shown, the above memory 1202 may include but are not limited to the acquisition module 81, the determination module 82, the training module 83, and the diagnosis module 84 in the above cross-domain fault diagnosis device for a rolling bearing. In addition, it may further include but are not limited to other module units in the above device, which will not be elaborated in this example.
[0211] Optionally, the above-described transmission device 1206 is used to receive or send data via a network. Specific examples of the above network may include a wired network and a wireless network. In one example, the transmission device 1206 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one example, the transmission device 1206 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0212] In addition, the above-described electronic device further includes: a display 1208, which is used to display the above-described fault diagnosis data; and a connection bus 1210, which is used to connect each module component in the above-described electronic device.
[0213] In other embodiments, the above-described terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes in the form of network communication. Among them, the nodes may form a peer-to-peer network, and any form of computing device, such as electronic devices like servers and terminals, can become a node in the blockchain system by joining the peer-to-peer network.
[0214] According to one aspect of the present application, there is provided a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the rolling bearing cross-domain fault diagnosis method provided in the above various optional implementation manners.
[0215] Optionally, in this embodiment, the above-described computer-readable storage medium may be set to store the methods for executing the embodiments of the present application.
[0216] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing the relevant hardware of the terminal device through a program. The program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.
[0217] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0218] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in the storage medium and includes several instructions to enable one or more electronic devices to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0219] In the above embodiments of the present application, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0220] In the several embodiments provided by the present application, it should be understood that the disclosed application program can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0221] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0222] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0223] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0224] In summary, from the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects:
[0225] 1. The present invention adopts a cross-domain fault diagnosis method for rolling bearings based on transfer learning. By using a source domain dataset and a target domain dataset to train a deep transfer network model, the knowledge obtained is transferred from the source domain to the target domain, which can effectively solve problems such as data scarcity and domain differences in fault diagnosis. By using the feature representations of the source domain data and model, the features of the target domain data are extracted in a targeted manner, thereby significantly improving the generalization ability and diagnostic accuracy of the model. Applying the trained deep transfer network model to the fault diagnosis of rolling bearings can reduce the data annotation cost and accelerate the convergence process of the model, providing support for accurate fault diagnosis.
[0226] 2. The present invention uses a nearest neighbor data smoothing algorithm for preprocessing, which can effectively reduce the interference of noise in the bearing vibration signal. The wavelet time-frequency diagram is used as the input of the deep transfer network. Compared with directly inputting the bearing time-domain vibration signal, using the wavelet time-frequency diagram as the input can enable the neural network to better capture the time-frequency characteristics of the signal. The feature extraction method of the wavelet time-frequency diagram is more abstract and general than directly inputting the vibration signal, and can better adapt to the input signals under different working conditions, improving the generalization ability and cross-domain fault diagnosis ability of the deep transfer network model.
[0227] 3. The fault diagnosis method used in the present invention conducts transfer learning from the perspective of structural adaptation through norm-enhanced transfer loss. Compared with data distribution adaptation methods (such as MMD, CORAL, etc.), it does not require additional distance metrics, and can improve the prediction diversity and discriminability of the diagnostic model for target domain data during the process of domain alignment, and has strong cross-domain learning ability.
[0228] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0229] It should be noted that in the description of this specification, the descriptions referring to the reference terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
Claims
1. A cross-domain fault diagnosis method for rolling bearings, characterized in that, Including: Collecting fault data of a rolling bearing and dividing the fault data into source domain data and target domain data; Respectively determining corresponding source domain data sets and target domain data sets based on the source domain data and the target domain data; Training a deep transfer network model by using the source domain data set and the target domain data set; Performing fault diagnosis on the rolling bearing through the trained deep transfer network model.
2. The method according to claim 1, characterized in that, Collecting fault data of a rolling bearing, including: Collecting fault data of rolling bearings with different fault types under different working conditions, where the fault data includes vibration acceleration data.
3. The method according to claim 2, wherein The fault types include one or more of inner ring faults, rolling element faults, and outer ring faults.
4. The method according to claim 1, wherein Before respectively determining corresponding source domain data sets and target domain data sets based on the source domain data and the target domain data, it further includes: Performing preprocessing operations on the source domain data and the target domain data respectively, where the preprocessing operations include one or more of removing DC components, smoothing adjacent data, and data augmentation operations.
5. The method according to claim 4, characterized in that The nearest neighbor data smoothing process corresponds to Formula 1, where S s (t) k is the data of the k-th sampling point after nearest neighbor data smoothing, L is the window length of the nearest neighbor data smoothing algorithm, S s′ (t) is the source domain data or the target domain data, i represents the i-th data point of the signal, and j represents the symmetric index offset relative to the center point k in the current sliding window.
6. The method according to claim 5, wherein It further includes: Do not perform nearest neighbor data smoothing on the starting sampling point of the signal, the k sampling points after the starting sampling point, the ending sampling point, and the k sampling points before the ending sampling point, where ceil() represents the ceiling operation.
7. The method according to claim 4, wherein The data augmentation operation includes: Overlapping sampling is respectively performed on the source domain data and the target domain data by using a preset sliding window. In order to improve the calculation efficiency of subsequent continuous wavelet transform, the number of sampling points of the overlapping signal satisfies the expression N = 2 k , where k is a positive integer.
8. The method according to any one of claims 1 to 7, characterized in that Respectively determining corresponding source domain data sets and target domain data sets based on the source domain data and the target domain data, including: Performing continuous wavelet transform on the source domain data and the target domain data respectively, extracting time-frequency domain features and generating corresponding wavelet time-frequency diagrams; Generating corresponding source domain data sets and target domain data sets based on the wavelet time-frequency diagrams.
9. The method according to claim 8, characterized in that Training a deep transfer network model by using the source domain data set and the target domain data set, including: Inputting the source domain data set and the target domain data set into the deep transfer network model; Using the deep transfer network model to respectively extract features from the source domain data set and the target domain data set to obtain source domain data features and target domain data features; Determining a corresponding classification loss based on the source domain data features; Determining a transfer loss based on the target domain data features; Determining the total loss of the deep transfer network model according to the classification loss and the transfer loss; Training the deep transfer network model based on the total loss.
10. The method according to claim 9, wherein The deep transfer network model includes: a feature extraction module, a transfer learning module, and a classification module.
11. The method according to claim 10, characterized in that, Using the deep transfer network model to respectively extract features from the source domain data set and the target domain data set, including: Using the feature extraction module of the deep transfer network model to respectively extract features from the source domain data set and the target domain data set.
12. The method according to claim 11, wherein The feature extraction module of the deep transfer network model uses ResNet18 as the backbone network.
13. The method according to claim 10, wherein Determining a corresponding classification loss based on the source domain data features, including: Determining the classification loss based on the source domain data features through the classification module of the deep transfer network model.
14. The method according to claim 10, characterized in that Determining a transfer loss based on the target domain data features, including: Determining the transfer loss based on the target domain data features through the transfer learning module of the deep transfer network model.
15. The method according to claim 13, wherein Determining the classification loss based on the source domain data features through the classification module of the deep transfer network model, including: Based on the source domain data features, the classification loss is determined using the cross-entropy loss function of the classification module. The cross-entropy loss function is represented by Formula 2, and the Formula 2 is Among them, is the classification loss, R is the number of rows of the source domain output classification matrix P(X s ), C is the number of columns of the source domain output classification matrix P(X s ), y i,j is the true class label of the element, P(X s ) i,j is the element in the i-th row and j-th column of P(X s ).
16. The method according to claim 14, characterized in that, Based on the target domain data features, the transfer loss is determined through the transfer learning module of the deep transfer network model, including: Based on the characteristics of the target domain data, the transfer loss is determined by using the norm-enhanced transfer loss function of the transfer learning module. The norm-enhanced transfer loss function is represented by Formula 3, and the Formula 3, where, is the transfer loss, B t is the batch number of input samples, P(X t ) is the output classification matrix of the target domain, X t is the input sample of the target domain, and n is the number of samples.
17. The method according to any one of claims 9 to 16, characterized in that, Based on the classification loss and the transfer loss, the total loss of the deep transfer network model is determined, including: The sum of the classification loss and the transfer loss is used as the total loss of the deep transfer network model.
18. The method according to claim 17, wherein The total loss is represented by Formula 4, and the Formula 4 is where γ ∈ [0.1, 0.5, 1, 1.5, 2], δ represents the classification loss weight factor, and γ represents the trade-off factor.
19. The method according to claim 9, characterized in that, Training the deep transfer network model using the source domain dataset and the target domain dataset further includes: Dividing the target domain dataset into a target domain training set and a target domain test set; Training the deep transfer network model using the target domain training set and the source domain dataset; Evaluating the effect of the trained deep transfer network model using the target domain test set.
20. The method according to claim 19, characterized in that, It further includes: Using the source-only method and the deep transfer method DCORAL to conduct a comparative verification of the effect of the trained deep transfer network model.
21. A rolling bearing cross-domain fault diagnosis device, characterized in that It includes: An acquisition module, configured to acquire the fault data of the rolling bearing and divide the fault data into source domain data and target domain data; A determination module, configured to respectively determine the corresponding source domain dataset and target domain dataset based on the source domain data and the target domain data; A training module, configured to train the deep transfer network model using the source domain dataset and the target domain dataset; A diagnosis module, configured to perform fault diagnosis on the rolling bearing through the trained deep transfer network model.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein the computer program can be executed by an electronic device to perform the method described in any one of Claims 1 to 20.
23. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method described in any one of Claims 1 to 20 are implemented.
24. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of Claims 1 to 20 through the computer program.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method
CN106404396A
Cited By
PEM electrolytic cell fault diagnosis method and domain alignment network model architecture
CN121144968A
Bearing fault diagnosis method and system based on transfer learning and interpretable analysis
CN121880966A
A Bearing Fault Diagnosis Method and System Based on Transfer Learning and Interpretable Analysis
CN121880966B