Rotary machinery cross-domain fault diagnosis method based on transfer learning

Through feature extraction and loss function optimization based on transfer learning, the accuracy and generalization problems in the middle and span working conditions of rotary mechanical fault diagnosis are solved, and efficient fault identification and diagnosis are achieved to adapt to complex industrial environments.

CN120493060APending Publication Date: 2025-08-15LANZHOU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510561654.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the fault diagnosis of rotary machinery, the existing technology has problems with insufficient generalization ability due to the lack of data labels and the multi-scale nature of complex fault data, especially in cross-working conditions, it is difficult to achieve high accuracy and stable fault identification.

Method used

Using a transfer learning-based method, by constructing a feature extraction module, improving the local maximum mean difference function and weighting factor, combining multi-scale convolutional blocks and self-attention mechanisms, effective alignment and feature extraction of source and target domain data are achieved, pseudo-labels are generated and model loss functions are optimized, and diagnostic accuracy and generalization performance are improved.

Benefits of technology

The accuracy and stability of rotary machinery fault diagnosis is significantly improved under cross-work conditions, reducing misjudgment and misjudgment, and enhancing the model's adaptability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493060A_ABST
    Figure CN120493060A_ABST
Patent Text Reader

Abstract

The invention discloses a rotating machine cross-domain fault diagnosis method based on transfer learning. The method comprises the steps that a sensor is used for collecting source domain data and target domain data of a rotating machine; based on a feature extraction module, improving a local maximum mean value function, a weighting factor and a classifier, and constructing a fault diagnosis model; inputting the source domain data into a feature extraction module and a classifier, and obtaining an optimal weight and source domain classification loss through back propagation; inputting target domain data into the feature extraction module and the classifier after weight migration through weight migration to obtain a pseudo tag and classification loss of a target domain; obtaining domain alignment loss by using an improved local maximum mean value function; adjusting by using a weighting factor to obtain total loss; and inputting the target domain test data into the trained network model to obtain a fault result. According to the method, the fault features of the signals can be effectively extracted, and higher fault diagnosis accuracy and better generalization performance are shown under the cross-working condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rotating machinery fault diagnosis, and more specifically to a rotating machinery cross-domain fault diagnosis method based on transfer learning. Background Art

[0002] Rotating machinery is widely used in a variety of complex and critical industrial applications, including wind power systems, aerospace, transportation, and manufacturing. However, due to the complexity and diversity of its operating environments, rotating machinery is prone to failure. Failure to promptly and effectively diagnose and repair faults can lead to equipment downtime, production interruptions, and even serious safety incidents, resulting in significant economic losses and potential safety hazards.

[0003] Traditional fault diagnosis methods often rely on manual feature extraction and expert knowledge, which often leads to limitations and instability when faced with different operating environments and fault types. In recent years, with the development of deep learning technology, more and more researchers have begun exploring the use of deep learning methods for rotating machinery fault diagnosis. However, most deep learning-based rotating machinery fault diagnosis methods require a large amount of labeled data to train the fault diagnosis model to achieve effective fault diagnosis results. However, in actual industrial processes, the data obtained is often unlabeled, which poses a challenge to deep learning fault diagnosis methods.

[0004] Transfer learning methods, by learning domain-invariant features between the source and target domains, can achieve good fault identification results in unlabeled data and are widely used in rotating machinery fault diagnosis problems with unlabeled data. In "Open-set domain generalization for fault diagnosis through data augmentation and a dual-level weighted mechanism," Jian et al. demonstrated excellent transfer learning under unknown operating conditions by constructing a multi-domain hybrid, a GAN with an auxiliary classifier, and a dynamic adversarial learning method with a dual-layer attention mechanism. In "Dynamic weighted adversarial domain adaptation network with sparse representation denoising module for rotating machinery fault diagnosis," Niu et al. designed a sparse representation denoising module and a weighted adversarial domain adaptation strategy to assign greater weight to source domain samples with high relevance to the target domain, reducing negative transfer and achieving good fault classification results for imbalanced data.

[0005] The above method obtains domain-invariant features of source and target domain data through feature extraction structure, and reduces the distribution difference between the source and target domains through difference measurement function. Although it has achieved good results, the generalization ability in complex fault identification is insufficient due to the multi-scale complexity of fault data. Summary of the Invention

[0006] In view of this, the present invention provides a cross-domain fault diagnosis method for rotating machinery based on transfer learning, which can solve the problems mentioned in the above background technology.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] An embodiment of the present invention provides a cross-domain fault diagnosis method for rotating machinery based on transfer learning, comprising the following steps:

[0009] S1: Use sensors to collect one-dimensional vibration signals of rotating machinery under different operating conditions, and preprocess them into source domain data and target domain data;

[0010] S2: Construct a fault diagnosis model, including a feature extraction module, an improved local maximum mean difference function, a weighting factor, and a classifier;

[0011] S3: Inputting the source domain data into the feature extraction module and the classifier, optimizing the weights through back propagation, and obtaining the optimal weights and the classification loss of the source domain data;

[0012] S4: Through weight migration, the target domain data is input into the migrated feature extraction module and classifier to generate pseudo labels of the target domain and calculate the classification loss of the target domain data;

[0013] S5: Use the improved local maximum mean difference function to align the domain alignment features of the source domain and the target domain to obtain the domain alignment loss;

[0014] S6: Adjust the weights of the target domain classification loss and domain alignment loss based on the weighting factor, calculate the total loss and optimize the model;

[0015] S7: Input the target domain test data into the trained model and output the fault diagnosis results.

[0016] Furthermore, the step S1 includes:

[0017] S11: Acquire one-dimensional vibration data of various fault states of rotating machinery under different working conditions through sensors;

[0018] S12: Use a sliding window to perform overlapping sampling on the one-dimensional vibration data signal to generate a source domain dataset and a target domain dataset; wherein the sliding window size is 1028, the step size is 64, and the target domain data is divided into a training set and a test set in a ratio of 7:3.

[0019] Furthermore, the feature extractor includes: a multi-scale convolution block, a fusion weight block and an improved self-attention mechanism; the processing process is as follows:

[0020] S21. Data is input to the multi-scale convolution block, and convolution layers with different convolution kernel sizes are used to extract multi-scale features of the three branches;

[0021] S22, the multi-scale features of the three branches are input into a fusion weight block to obtain weight coefficients corresponding to the three branches;

[0022] S23. The multi-scale features of the three branches are input into the improved self-attention mechanism to obtain the output of the feature extractor.

[0023] Furthermore, the multi-scale convolution block includes: two convolution layers of the first branch, two convolution layers of the second branch, and two convolution layers of the third branch; the processing process is as follows:

[0024] S211, the data is input into the first convolutional layer of the first, second and third branches respectively to obtain shallow features corresponding to the three branches;

[0025] S212. The shallow features corresponding to the three branches are respectively input into the second convolutional layer of the corresponding branch to obtain multi-scale features of the three branches.

[0026] Furthermore, the fusion weight block includes: a fusion layer, a global average pooling layer, a fully connected layer and a softmax layer; the processing process is as follows:

[0027] S221. The multi-scale features of the three branches are input to the fusion layer to obtain fusion features.

[0028] S222, the fused features are input to the global average pooling layer to obtain a feature representation of the fused features;

[0029] S223. The feature representation of the fused feature is input to the fully connected layer to obtain an enhanced feature representation;

[0030] S224. The strong feature representation is input to the Softmax layer to obtain weight coefficients corresponding to the three branches.

[0031] Furthermore, the improved self-attention mechanism includes: a linear transformation layer, a weighted layer, an overlay layer, and an output layer; the processing process is as follows:

[0032] S231, the multi-scale features of the three branches are respectively input into the linear transformation layer to obtain the query matrix, key matrix, and value matrix of the three branches;

[0033] S232: Input the query matrix, key matrix, value matrix of each of the three branches and the weight coefficients corresponding to the three branches into the weighted layer respectively to obtain the weighted query matrix, key matrix, value matrix of each of the three branches;

[0034] S233: Input the weighted query matrix, key matrix, and value matrix of each of the three branches into the superposition layer to obtain a fused query matrix, key matrix, and value matrix;

[0035] S234. The fused query matrix, key matrix, and value matrix are input to the output layer to obtain the output of the feature extractor.

[0036] Furthermore, the improved local maximum mean difference function includes the following:

[0037] The covariance of the source domain data is determined to be Δx1, the covariance of the target domain data is Δx2, and the variance difference Δ between the source domain and the target domain is as follows:

[0038] Δ=||Δx1-Δx2||

[0039] According to the variance difference Δ, a dynamic Gaussian kernel function is constructed The expression is as follows:

[0040]

[0041] Among them, x i represents the source domain data sample, y i represents the target domain data sample, ||x i -y i || represents sample x i and y i The Euclidean distance of

[0042] According to the dynamic Gaussian kernel function The expression of the improved local maximum mean difference function is as follows:

[0043]

[0044] Among them, C represents the number of categories, n s and n t represent the source domain dataset and the target domain dataset respectively, and Respectively represent the source domain features and target domain features output by the i-th sample in the model, and They represent the weights of the i-th sample and the j-th sample in the source domain belonging to category c, and They represent the weights of the i-th sample and the j-th sample in the target domain belonging to category c respectively.

[0045] Furthermore, the weighting factors include the following:

[0046] S61: Calculate the KL divergence of the probability distribution of the source domain and the target domain;

[0047] S62: Normalize the KL divergence to obtain a dynamic factor to adjust the weight ratio of classification loss and domain alignment loss.

[0048] Furthermore, the total loss function is as follows:

[0049] L=L1+1·L2+(1-1)·L3

[0050] Among them, L represents the total loss of the fault diagnosis model, L1 represents the classification loss of the source domain data, L2 represents the classification loss of the target domain, L3 represents the domain alignment loss, and l represents the weighting factor.

[0051] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:

[0052] Improved fault diagnosis accuracy: This invention utilizes transfer learning technology to effectively extract fault characteristics from rotating machinery signals and demonstrates higher fault diagnosis accuracy across a wide range of operating conditions. This means the system can accurately identify and diagnose mechanical faults in diverse operating environments, reducing false positives and missed detections.

[0053] Enhanced generalization performance: By improving the local maximum mean function and weighting factors, this paper effectively aligns source and target domain data, thereby improving the model's generalization capabilities across diverse operating conditions. This allows the model to not only perform well under specific conditions but also adapt to a variety of complex and changing industrial environments, ensuring the stability and reliability of fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0055] Figure 1 Flowchart of the cross-domain fault diagnosis method for rotating machinery based on transfer learning provided by the present invention.

[0056] Figure 2 This is a complete technical principle diagram from data input to fault diagnosis result output provided by the present invention.

[0057] Figure 3 This is a structural diagram of the feature extraction module provided by the present invention.

[0058] Figure 4a-4l To verify the confusion matrix diagram of the present invention in 12 migration tasks in the embodiment. DETAILED DESCRIPTION

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0060] Reference Figure 1-2 As shown, the embodiment of the present invention discloses a cross-domain fault diagnosis method for rotating machinery based on transfer learning, which includes the following steps S1 to S7. Figure 2 The core idea of cross-domain transfer learning is intuitively presented, which is to solve the problem of missing target domain data labels through source domain knowledge transfer and dynamic domain alignment, and ultimately achieve high-precision and strong generalization fault diagnosis.

[0061] S1: Use sensors to collect one-dimensional vibration signals of rotating machinery under different operating conditions, and pre-process the data into source domain data and target domain data. The source domain data refers to, for example, the vibration data collected by the sensor of a wind turbine at a specific speed and load (including clear fault labels: such as "bearing inner ring crack"). It is used as the basic data for training the model and teaching the fault diagnosis model to recognize the characteristics of different faults. The target domain data is a new problem to be solved by the present invention, such as the same generator with a different operating condition (different speed / load), or another generator with a similar structure, but its vibration data has no fault label.

[0062] The present invention allows the fault diagnosis model to transfer the experience learned based on source domain data under one machine working condition to another machine with different working conditions, and can accurately diagnose even if the new machine (target domain data) does not have a fault label.

[0063] The step S1 specifically includes:

[0064] S11. Use sensors to obtain one-dimensional vibration data of various fault states of rotating machinery under different working conditions.

[0065] S12. Use a sliding window to perform overlapping sampling on the one-dimensional vibration data signal to obtain a source domain data set and a target domain data set.

[0066] Specifically, the source domain data and target domain data collected by the sensor are sampled using an overlapping method, with a sliding window size of 1028 and a step size of 64, and the target domain data are divided into a training set and a test set in a ratio of 7:3.

[0067] S2: Construct a fault diagnosis model, including a feature extraction module, an improved local maximum mean difference function, a weighting factor, and a classifier;

[0068] Among them, the structure diagram of the feature extraction module is as follows Figure 3 As shown in the figure, the input data is input into the multi-scale convolution layer for multi-scale feature extraction. The specific operation is as follows: the input data is input into three branches for feature extraction, where the first branch has two convolution layers with a convolution kernel size of 1×3. Each convolution layer first passes through a batch normalization layer and then uses a ReLU activation function. The second branch has two convolution layers with a convolution kernel size of 1×5. Each convolution layer first passes through a batch normalization layer and then uses a ReLU activation function. The third branch has two convolution layers with a convolution kernel size of 1×7. Each convolution layer first passes through a batch normalization layer and then uses a ReLU activation function (where the first convolution layer is used to extract shallow features). Through convolution kernels of different sizes, three multi-scale features s1, s2 and s3 of the input data are obtained, and their mathematical expressions are as follows.

[0069] s1=ψ(η(f 1×3 (ψ(η(f 1×3 (x))))))

[0070] s2=ψ(η(f 1×5 (ψ(η(f 1×5 (x))))))

[0071] s3=ψ(η(f 1×7 (ψ(η(f 1×7 (x))))))

[0072] Among them, ψ(·) represents the Relu activation function, η(·) represents the batch normalization layer, and f 1×3 (·),f 1×5 (·) and f 1×7 (·) denotes convolution kernels of sizes 1×3, 1×5, and 1×7, respectively.

[0073] Furthermore, the three multi-scale features s1, s2, and s3 are input into the fusion weight block to obtain the weight coefficients of the three branches. The specific operation is as follows: the three multi-scale features are added element by element to obtain the fusion feature, and then passed through a global average pooling layer, a fully connected layer, and a softmax layer to obtain the weight coefficients α1, α2, and α3 of the three branches. The mathematical expressions are as follows:

[0074]

[0075] in, represents the Softmax function, represents the fully connected layer, τ(·) represents the global average pooling layer, Represents the fusion feature after element-by-element addition of s1, s2, and s3, i∈[1,2,3].

[0076] Furthermore, the three multi-scale features s1, s2, s3 and the weight coefficients α1, α2, α3 of the three branches are input into the improved self-attention mechanism to obtain the output of the feature extractor. The specific operation is: the three multi-scale features s1, s2, s3 are linearly transformed respectively to obtain the query matrix Q1, key matrix K1, value matrix V1 of the first branch, the query matrix Q2, key matrix K2, value matrix V2 of the second branch, and the query matrix Q3, key matrix K3, value matrix V3 of the third branch. The mathematical expressions are as follows:

[0077]

[0078] Among them, Q i , K i and V i They represent the Q, K, and V matrices generated by the linear transformation of the three branch features, respectively. si represents the three multi-scale features s1, s2, and s3. and Respectively represent the parameter matrices corresponding to the linear transformation, i∈[1,2,3].

[0079] The query matrix Q of the three branches i , key matrix K i , value matrix V i and the weight coefficient α of the corresponding branch i Multiply to get the adjusted query matrix Bond Matrix Value Matrix Then adjust the query matrix of the three branches Bond Matrix Value Matrix Adding them together, we get the fusion query matrix Q, fusion key matrix K and fusion value matrix V, whose mathematical expressions are as follows:

[0080]

[0081] Furthermore, by fusing the query matrix Q, the fusion key matrix K, and the fusion value matrix V, the output of the feature extraction module is calculated using the formula of the self-attention mechanism. Its mathematical expression is as follows:

[0082]

[0083] Among them, K T represents the transposed matrix of K, d k represents the dimension of the key matrix K.

[0084] Furthermore, the improved local maximum mean difference function dynamically adjusts the bandwidth of the traditional high-speed sum function based on the covariance difference of the inter-domain data, thereby improving the accuracy and robustness of the difference measurement. If the covariance of the source domain data is Δx1 and the covariance of the target domain data is Δx2, the variance difference Δ between the source domain and the target domain is as follows:

[0085] Δ=||Δx1-Δx2||

[0086] According to the variance difference Δ, a dynamic Gaussian kernel function is proposed The expression is as follows:

[0087]

[0088] Among them, x i represents the source domain data sample, y i represents the target domain data sample, ||x i -y i || represents sample x i and y i The Euclidean distance of .

[0089] According to the dynamic Gaussian kernel function The expression of the improved local maximum mean difference function is as follows:

[0090]

[0091] Among them, C represents the number of categories, n s and n t represent the source domain dataset and the target domain dataset respectively, and Respectively represent the source domain features and target domain features output by the i-th sample in the model, and They represent the weights of the i-th sample and the j-th sample in the source domain belonging to category c, and They represent the weights of the i-th sample and the j-th sample in the target domain belonging to category c respectively.

[0092] Furthermore, the specific principle of the weighting factor is as follows: assuming that the source domain data and the target domain data are probability distributions N1 and N2, the KL divergence of each batch N1 and N2 is Its mathematical expression is as follows:

[0093]

[0094] Among them, N represents the batch size, C represents the number of categories, and P i,j represents the probability of the i-th sample in the j-th category in the source domain probability distribution N1, Q i,j It represents the probability of the i-th sample in the j-th category in the target domain probability distribution N2.

[0095] The Normalize and get the dynamic factor l, whose mathematical expression is as follows:

[0096]

[0097] Furthermore, the classifier mainly consists of two fully connected layers and a Softmax layer, which outputs the probabilities of different faults and finally obtains the fault classification results.

[0098] S3: Input the source domain data into the feature extraction module and classifier, and obtain the optimal weights and source domain classification loss through back propagation.

[0099] The source domain data is input into the feature extractor to obtain the domain-invariant features of the source domain data. The domain-invariant features are then input into the classifier. Back propagation is performed by minimizing the cross entropy loss function to obtain the optimal weight and source domain classification loss L1. The mathematical expression is as follows:

[0100]

[0101] Where C represents the fault category, y i represents the source domain label, Represents the probability of each category in the source domain.

[0102] S4: Through weight transfer, the target domain data is input into the feature extraction module and classifier after weight transfer to obtain the pseudo label of the target domain and the classification loss of the target domain.

[0103] The target domain data is input into the feature extraction module after weight transfer to obtain the domain-invariant features of the target domain data. The domain-invariant features of the target domain are then input into the classifier after weight transfer, and the category with the highest probability is generated as the pseudo label of the target domain through the Softmax function. By minimizing the cross entropy loss function and performing back propagation, we can obtain the target domain classification loss L2, which is expressed mathematically as follows:

[0104]

[0105] Where C represents the fault category, represents the target domain pseudo label, p i Represents the probability of each category in the target domain.

[0106] S5: Use the improved local maximum mean function to align the domain alignment features of the source domain and the target domain to obtain the domain alignment loss.

[0107] After obtaining the domain-invariant features of the source domain data and the target domain data, the improved local maximum mean function is used for inter-domain alignment to obtain the domain alignment loss L3, which is mathematically expressed as follows:

[0108] L3=D(F1,F2)

[0109] Among them, L3 represents the domain alignment loss, F1 represents the source domain invariant features obtained by the source domain data through the feature extractor, and F2 represents the target domain invariant features obtained by the target domain data through the feature extractor.

[0110] S6: Use the weighting factor to adjust the target domain classification loss and domain alignment loss to get the total loss.

[0111] After obtaining the domain-invariant features of the source domain data and the target domain data, the KL divergence is used to calculate the weighting factor l, and the weighting factor l is used to weight the target domain classification loss and the domain alignment loss to obtain the total loss L of the entire fault diagnosis model. Its mathematical expression is as follows:

[0112] L=L1+1·L2+(1-1)·L3

[0113] Among them, L represents the total loss of the fault diagnosis model, L1 represents the classification loss of the source domain data, L2 represents the classification loss of the target domain data, L3 represents the domain alignment loss, and l represents the dynamic factor.

[0114] S7: Input the target domain test data into the trained network model to obtain the failure results of the target domain task.

[0115] Input the training sets of source domain data and target domain data into the constructed fault diagnosis model, use the Adam optimization algorithm to update the parameters, and obtain the trained fault diagnosis model after reaching the set iteration batch;

[0116] The test set of the target domain is input into the trained fault diagnosis model to obtain the final fault diagnosis result.

[0117] To verify the effectiveness of the present invention, this embodiment uses the Case Western Reserve University (CWRU) bearing dataset to perform cross-domain rotating machinery fault diagnosis.

[0118] The bearing under test in the CWRU is an SKF6205, and the sampling frequency is 12kHz. In this verification example, the data from the four operating conditions (0HP, 1HP, 2HP, and 3HP) are labeled as Dataset A0, Dataset A1, Dataset A2, and Dataset A3, respectively. Five fault states (0.007-Ball, 0.007-Inner, 0.014-Ball, 0.014-Inner, and 0.021-Outer) were selected for experimental verification based on the damage severity and location. The number of sampling points in each segment was set to 2048, and 300 samples were selected for each fault type, with 210 samples used for training and 90 samples used for testing.

[0119] This verification embodiment uses one of the working conditions as the source domain data and the other three working conditions as the target domain data, with a total of 12 groups of migration tasks. Taking A0-A1 as an example, it means that the A0 data set is the source domain and the A1 data set is the target domain. In order to comprehensively evaluate the fault diagnosis effect of the proposed method, this verification embodiment introduces three indicators: recall rate Recall, precision rate Precision and F1 value. Recall rate Recall represents the proportion of all samples that are actually positive examples that are correctly predicted as positive examples by the model. Precision represents the proportion of positive examples that are actually positive examples among those predicted by the model. The F1 value is the harmonic mean of precision and recall. The F1 value is between 0 and 1. The higher the F1 value, the better the performance of the model. The mathematical expressions of Recall, Precision and F1 values are:

[0120]

[0121] Among them, TP represents true positive examples, FN represents false negative examples, and FP represents false positive examples.

[0122] The experimental results of different migration tasks of the present invention are shown in Table 1.

[0123] Table 1 Indicators of different migration tasks of the present invention

[0124]

[0125]

[0126] As can be seen from Table 1, in the A0-A2, A1-A0, A1-A2, A2-A1, A2-A3, A3-A1, and A3-A2 migration tasks, the four indicators are all greater than or equal to 97%, indicating that the present invention performs very well in these migration tasks, can effectively cope with the distribution differences between the source and target domains, and demonstrates strong adaptability. In contrast, the indicators of A0-A1, A0-A3, A1-A3, A2-A0, and A3-A0 are all higher than 91%. Although all indicators remain at a high level above 91%, there is still a certain performance gap. This difference may be due to the more significant distribution offset between the source and target domains. Despite this, the present invention still maintains good stability and reliability in these tasks.

[0127] In order to more accurately illustrate the effectiveness of the present invention in all migration tasks, this verification example uses the confusion matrix method to visualize the diagnosis results of 12 migration tasks, such as Figure 4a-4l As shown (the vertical axis true label represents the real label, and the horizontal axis predicted label represents the predicted label). Figure 4b 、 Figure 4e 、 Figure 4l It can be seen that among A0-A2, A1-A2 and A3-A2, there is a type of fault that is misclassified to varying degrees. Figure 4d 、 Figure 4h 、 Figure 4j and Figure 4k It can be seen that in A1-A0, A2-A1, A3-A0 and A3-A1, there are two types of faults that are misclassified to varying degrees. Figure 4a 、 Figure 4c 、 Figure 4i It can be seen that in A0-A1, A0-A3 and A2-A3, there are three types of faults that are misclassified to varying degrees. Figure 4f 、 Figure 4g As can be seen, there are four types of faults A1-A3 and A2-A0 that are misclassified to varying degrees. This indicates that when facing different migration tasks, the performance of the present invention varies due to the varying degrees of difference between the source and target domains.

[0128] Figure 4a-4l To verify the confusion matrix diagram of the present invention in 12 migration tasks in the embodiment. Figure 4a is the confusion matrix of A0-A1; Figure 4b is the confusion matrix of A0-A2; Figure 4c is the confusion matrix of A0-A3; Figure 4d is the confusion matrix of A1-A0; Figure 4eis the confusion matrix of A1-A2; Figure 4f is the confusion matrix of A1-A3; Figure 4g is the confusion matrix of A2-A0; Figure 4h is the confusion matrix of A2-A1; Figure 4i is the confusion matrix of A2-A3; Figure 4j is the confusion matrix of A3-A0; Figure 4k is the confusion matrix of A3-A1; Figure 4l is the confusion matrix of A3-A2.

[0129] The proposed method performs well in homologous distribution or mildly offset scenarios, verifying the effectiveness of the transfer learning framework.

[0130] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0131] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A cross-domain fault diagnosis method for rotating machinery based on transfer learning, characterized in that: The following steps are involved: S1: Use sensors to collect one-dimensional vibration signals of rotating machinery under different operating conditions, and preprocess them into source domain data and target domain data; S2: Construct a fault diagnosis model, including a feature extraction module, an improved local maximum mean difference function, a weighting factor, and a classifier; S3: Inputting the source domain data into the feature extraction module and the classifier, optimizing the weights through back propagation, and obtaining the optimal weights and the classification loss of the source domain data; S4: Through weight migration, the target domain data is input into the migrated feature extraction module and classifier to generate pseudo labels of the target domain and calculate the classification loss of the target domain data; S5: Use the improved local maximum mean difference function to align the domain alignment features of the source domain and the target domain to obtain the domain alignment loss; S6: Adjust the weights of the target domain classification loss and domain alignment loss based on the weighting factor, calculate the total loss and optimize the model; S7: Input the target domain test data into the trained model and output the fault diagnosis results.

2. A cross-domain fault diagnosis method for rotating machinery based on transfer learning according to claim 1, characterized in that: The step S1 comprises: S11: Acquire one-dimensional vibration data of various fault states of rotating machinery under different working conditions through sensors; S12: Use a sliding window to perform overlapping sampling on the one-dimensional vibration data signal to generate a source domain dataset and a target domain dataset; wherein the sliding window size is 1028, the step size is 64, and the target domain data is divided into a training set and a test set in a ratio of 7:

3.

3. The method for cross-domain fault diagnosis of rotating machinery based on transfer learning according to claim 1, characterized in that: The feature extractor includes: a multi-scale convolution block, a fusion weight block and an improved self-attention mechanism; the processing process is as follows: S21. Data is input to the multi-scale convolution block, and convolution layers with different convolution kernel sizes are used to extract multi-scale features of the three branches; S22, the multi-scale features of the three branches are input into a fusion weight block to obtain weight coefficients corresponding to the three branches; S23. The multi-scale features of the three branches are input into the improved self-attention mechanism to obtain the output of the feature extractor.

4. A cross-domain fault diagnosis method for rotating machinery based on transfer learning according to claim 3, characterized in that: The multi-scale convolution block includes: two convolution layers of the first branch, two convolution layers of the second branch, and two convolution layers of the third branch; the processing process is as follows: S211, the data is input into the first convolutional layer of the first, second and third branches respectively to obtain shallow features corresponding to the three branches; S212. The shallow features corresponding to the three branches are respectively input into the second convolutional layer of the corresponding branch to obtain multi-scale features of the three branches.

5. The method for cross-domain fault diagnosis of rotating machinery based on transfer learning according to claim 3 is characterized in that: The fusion weight block includes: fusion layer, global average pooling layer, fully connected layer and softmax layer; the processing process is as follows: S221. The multi-scale features of the three branches are input to the fusion layer to obtain fusion features. S222, the fused features are input to the global average pooling layer to obtain a feature representation of the fused features; S223. The feature representation of the fused feature is input to the fully connected layer to obtain an enhanced feature representation; S224. The strong feature representation is input to the Softmax layer to obtain weight coefficients corresponding to the three branches.

6. The method for cross-domain fault diagnosis of rotating machinery based on transfer learning according to claim 3, characterized in that: The improved self-attention mechanism includes: a linear transformation layer, a weighted layer, an overlay layer, and an output layer; the processing process is as follows: S231, the multi-scale features of the three branches are respectively input into the linear transformation layer to obtain the query matrix, key matrix, and value matrix of the three branches; S232: Input the query matrix, key matrix, value matrix of each of the three branches and the weight coefficients corresponding to the three branches into the weighted layer respectively to obtain the weighted query matrix, key matrix, value matrix of each of the three branches; S233: Input the weighted query matrix, key matrix, and value matrix of each of the three branches into the superposition layer to obtain a fused query matrix, key matrix, and value matrix; S234. The fused query matrix, key matrix, and value matrix are input to the output layer to obtain the output of the feature extractor.

7. The method for cross-domain fault diagnosis of rotating machinery based on transfer learning according to claim 1, characterized in that: The improved local maximum mean difference function includes the following: The covariance of the source domain data is determined to be Δx1, the covariance of the target domain data is Δx2, and the variance difference Δ between the source domain and the target domain is as follows: Δ=|Δx1-Δx2| According to the variance difference Δ, a dynamic Gaussian kernel function is constructed The expression is as follows: Among them, x i represents the source domain data sample, y i represents the target domain data sample, ||x i -y i || represents sample x i and y i The Euclidean distance of According to the dynamic Gaussian kernel function The expression of the improved local maximum mean difference function is as follows: Among them, C represents the number of categories, n s and n t represent the source domain dataset and the target domain dataset respectively, and Respectively represent the source domain features and target domain features output by the i-th sample in the model, and They represent the weights of the i-th sample and the j-th sample in the source domain belonging to category c, and They represent the weights of the i-th sample and the j-th sample in the target domain belonging to category c respectively.

8. The method for cross-domain fault diagnosis of rotating machinery based on transfer learning according to claim 1, characterized in that: The weighting factors include the following: S61: Calculate the KL divergence of the probability distribution of the source domain and the target domain; S62: Normalize the KL divergence to obtain a dynamic factor to adjust the weight ratio of classification loss and domain alignment loss.

9. The method for cross-domain fault diagnosis of rotating machinery based on transfer learning according to claim 1, characterized in that: The total loss function is as follows: L=L1+1·L2+(1-1)·L3 Among them, L represents the total loss of the fault diagnosis model, L1 represents the classification loss of the source domain data, L2 represents the classification loss of the target domain, L3 represents the domain alignment loss, and l represents the weighting factor.

Citation Information

Cited By

  • Intelligent flutter monitoring method and system for blisk machining

    CN120950904A