Mutual information guide domain self-adaption method and equipment for intelligent migration diagnosis of rotating machinery, and medium
By employing a mutual information-guided domain adaptive method, and utilizing an enhanced convolutional neural network and a mutual information neural estimator, the problems of scarce labeled samples and distribution offset in time-varying scenarios in rotating machinery fault diagnosis are solved, achieving more accurate fault diagnosis results.
Patent Information
- Application Number
- CN202511538892.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-01-30
AI Technical Summary
Existing unsupervised adaptive domain methods rely on massive labeled source domain samples in rotating machinery fault diagnosis. Furthermore, methods with few samples are prone to distribution shifts and overfitting, making it difficult to resolve the contradiction between scarce labeled source domain samples and the demand for massive labeled monitoring data. The data distribution shift problem is particularly prominent in time-varying scenarios.
We employ a mutual information-guided domain adaptation method, which constructs an enhanced convolutional neural network with enhanced local channel attention modules and a mutual information neural estimator. Combined with a joint maximum mean-variance difference measure method based on multidimensional statistical alignment, we build an unsupervised transfer training framework based on average teachers. This framework utilizes a small number of labeled source domain samples and a large number of unlabeled samples to achieve feature extraction and label association.
It effectively improves the performance of rotating machinery fault diagnosis, solves the problems of difficult time-varying data distribution alignment and scarce source domain labeled samples, and achieves more accurate cross-domain fault diagnosis.
Smart Images

Figure CN121434591A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mechanical fault diagnosis, in particular to a mutual information guided domain adaptation method for intelligent migration diagnosis of rotating machinery, equipment and medium. BACKGROUND
[0002] Fault diagnosis is considered as one of the key technologies to ensure the safe and reliable operation of rotating machinery. With the advantage of adaptive feature extraction, deep learning-based fault diagnosis methods have been extensively studied in the past. However, in actual scenarios, there are often problems of data labeling difficulty and different distribution of training and testing data, which greatly limits the effective and extensive application of deep learning-based fault diagnosis methods in actual industrial scenarios.
[0003] Unsupervised domain adaptation is an effective method to solve the above-mentioned diagnosis between data with different distributions. By applying cross-domain distribution alignment technology, the rich labeled source domain is migrated to the unlabeled and differently distributed target domain, and the distribution difference between the source domain and the target domain is alleviated.
[0004] Existing unsupervised domain adaptation methods are divided into unsupervised domain adaptation methods with sufficient labeled source domain and unsupervised domain adaptation methods with few samples of source domain according to the difference in data availability. The former uses a large number of labeled source domain samples to integrate domain adaptation methods such as adversarial training and domain distribution alignment measurement to reduce the difference between domains, thereby realizing unsupervised domain adaptation. However, the unsupervised domain adaptation method with sufficient labeled source domain relies on a large number of labeled source domain samples, and when the number of source domain labeled samples is limited or scarce, it may be difficult to have high diagnostic accuracy. The latter fully considers the difficulty of data labeling in actual scenarios and only uses a very small number of labeled source domain samples as the basis for knowledge migration. However, the few labeled source domain samples are very easy to cause distribution deviation and overfitting phenomenon, and at the same time, a large number of unlabeled samples collected are ignored, wasting the feature information contained in these samples.
[0005] In summary, the existing technology has at least the following disadvantages: 1) the above-mentioned unsupervised domain adaptation method with sufficient labeled source domain relies on a large number of labeled source domain samples for domain adaptation. 2) The current unsupervised domain adaptation method with few samples is prone to distribution deviation and overfitting, and also wastes valuable unlabeled samples. 3) The limitation of domain deviation caused by the distribution deviation of data in the time-varying scene brings challenges to the unsupervised domain adaptation method. Based on this, the existing unsupervised domain adaptation method cannot solve the contradiction between the scarcity of labeled source domain samples and the demand for a large number of labeled monitoring data, especially considering the significant domain deviation problem under time-varying conditions. SUMMARY
[0006] To solve the above problems, the application provides a mutual information guided domain adaptive method, equipment and medium for intelligent migration diagnosis of rotating machinery.
[0007] To achieve the above object, the application provides the following solutions. In a first aspect, the application provides a mutual information guided domain adaptive method for intelligent migration diagnosis of rotating machinery, comprising: Obtaining original vibration signals of rotating machinery under different fault states and preprocessing; the original vibration signals cover various health states and working conditions of source domains and target domains; Constructing a mutual information guided domain adaptive network; the domain adaptive network includes an enhanced convolutional neural network with an enhanced local channel attention module and a mutual information neural estimator for multi-dimensional variable mutual information estimation; Constructing a joint maximum mean variance difference metric method for multi-dimensional statistical quantity alignment; Taking the domain adaptive network as a teacher model and a student model respectively, constructing an average teacher based unsupervised migration training framework; Using the joint maximum mean variance difference metric method, constructing an inter-domain difference loss; Determining a mutual information loss based on a lower bound of mutual information calculated by the mutual information neural estimator under a set parameter; Using an average teacher unsupervised migration training strategy, combining a classification loss, the inter-domain difference loss and the mutual information loss to train and update the student model in the unsupervised migration training framework, and using an exponential moving average strategy to train and update the teacher model in the unsupervised migration training framework to obtain an unsupervised migration diagnosis model; Inputting the preprocessed original vibration signals into the unsupervised migration diagnosis model to obtain a migration diagnosis result.
[0008] Optionally, the enhanced convolutional neural network with an enhanced local channel attention module is composed of four feature extraction modules, three enhanced local channel attention modules, two fully connected layers, an exponential linear unit and a SoftMax layer. The enhanced local channel attention modules are arranged between the four feature extraction modules; the fourth feature extraction module is connected with the first fully connected layer through a flattening operation; the first fully connected layer, the exponential linear unit, the second fully connected layer and the SoftMax layer are connected in sequence.
[0009] Optionally, the four feature extraction modules are each composed of a one-dimensional convolutional layer, a maximum pooling layer, a batch normalization layer and an exponential linear unit.
[0010] Optionally, the enhanced local channel attention module consists of a compressed excitation branch, a residual branch, and an enhanced local attention branch; The compressed excitation branch is used to compress the features extracted by the feature extraction module in order to obtain the importance of the excitation channel; The enhanced local attention branch is used to perform local attention enhancement processing on the features extracted by the feature extraction module in order to obtain enhanced local attention importance; The residual branch is used to multiply the features extracted by the feature extraction module with the importance of the enhanced local attention and the importance of the excitation channel channel one channel at a time to obtain the output feature map.
[0011] Optionally, the mutual information neural estimator consists of a projection head and a mutual information neural estimator body.
[0012] Optionally, the method for measuring the joint maximum mean variance difference of the multidimensional statistics alignment includes the joint maximum mean difference and the maximum variance difference; The calculation process for the joint maximum mean difference is expressed as follows: ; The calculation process for the maximum variance difference is expressed as follows: ; In the formula, Indicates the joint maximum mean difference. They represent the first i Extraction features of the source domain sample and the first j Extracting features from samples in the target domain. They represent the first i The predicted label of the first source domain sample and the first j Predicted labels for each target domain sample. M Indicates the number of samples in the source domain. k ( ) represents the Gaussian kernel function. N Indicates the number of samples in the target domain. , They represent the first j Extraction features of the source domain sample and the first i Extracting features from samples in the target domain. , They represent the first j Extraction features of the source domain sample and the first i Predicted labels for each target domain sample. For the maximum variance difference, Var ( The ) represents the variance of the extracted features in the RKHS. k 1( ) represents the t-student kernel function.
[0013] Alternatively, the inter-domain difference loss can be expressed as: ; ; ; In the formula, for and The joint maximum mean difference between them This represents the joint maximum mean difference loss. and These are the features extracted and predicted labels for the enhanced target domain samples by the student model, respectively. and These represent the features extracted and predicted labels for the teacher model on enhanced labeled source domain samples, respectively. To enhance the feature extraction of labeled source domain samples for the teacher model. The student model enhances the extraction features of target domain samples. The maximum variance difference between them. This represents the loss due to the maximum variance. This represents the loss due to inter-domain differences.
[0014] Optionally, the loss function formed by the classification loss, the inter-domain difference loss, and the mutual information loss is expressed as: ; In the formula, Indicates the total loss. Represents classification loss. Indicates mutual information loss. Indicates the loss due to inter-domain differences. and These are all hyperparameters that balance the various losses. e It is the natural logarithm.
[0015] Secondly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the mutual information guided domain adaptive method for intelligent migration diagnosis of rotating machinery provided above.
[0016] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the mutual information guided domain adaptive method for intelligent migration diagnosis of rotating machinery provided above.
[0017] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a mutual information-guided domain adaptation method, device, and medium for intelligent migration diagnosis of rotating machinery. By constructing a mutual information-guided domain adaptation network, it can effectively extract discriminative fault-related features by synergistically utilizing a small number of labeled source domain samples and a large number of unlabeled source domain samples, thus solving the problems of time-varying operating condition data distribution shift and the inherent contradiction between source domain sample labeling and method requirements faced by existing unsupervised domain adaptation methods. By setting an enhanced local channel attention module in the enhanced convolutional neural network, channel weights can be adaptively learned, driving the network to focus on the most discriminative channel, thereby improving feature extraction capabilities. By setting a mutual information neural estimator for multidimensional variable mutual information estimation, feature-label correlation can be strengthened, and then the enhanced convolutional neural network can be guided to maximize the mutual information between features and labels, constructing a strongly dependent feature-label correspondence system. By constructing a joint maximum mean-variance difference measure method for multidimensional statistical alignment, the higher-order distribution differences of mean and variance can be minimized simultaneously, effectively reducing domain shift and enhancing the extraction capability of domain-invariant features, achieving better domain distribution alignment. By combining classification loss, inter-domain difference loss, and mutual information loss for model update, unsupervised transfer diagnostic models can be endowed with excellent feature extraction and generalization capabilities, thereby establishing strong correlations between features and labels to achieve accurate cross-domain fault diagnosis. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a mutual information-guided domain adaptive method for intelligent migration diagnostics of rotating machinery, provided in an embodiment of this application; Figure 2 This is a schematic diagram of a domain adaptive network structure provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a mutual information neural estimator provided in an embodiment of this application; Figure 4 A schematic diagram of the architecture design of an ELCA module provided in an embodiment of this application; Figure 5 A schematic diagram comparing the diagnostic accuracy of different methods provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] In one exemplary embodiment, this application provides a mutual information-guided domain adaptive method for intelligent migration diagnosis of rotating machinery. This method is executed by a computer device, specifically by a terminal or server alone, or by both. In this embodiment, the method is described using a server as an example. Figure 1 As shown, the method includes: Step 100: Acquire the raw vibration signals of the rotating machinery under different fault conditions and perform preprocessing. The raw vibration signals cover various health states and operating conditions in both the source and target domains.
[0023] Step 101: Construct a mutual information-guided domain adaptive network. The domain adaptive network includes an enhanced convolutional neural network with an enhanced local channel attention module and a mutual information neural estimator (MINE) for estimating mutual information of multidimensional variables. The structure of the enhanced convolutional neural network is as follows: Figure 2 As shown, the structure of the mutual information neural estimator is as follows: Figure 3 As shown.
[0024] Step 102: Construct a joint maximum mean variance difference measure method with multidimensional statistics aligned.
[0025] Step 103: Using the domain adaptive network as the teacher model and the student model respectively, construct an unsupervised transfer training framework based on the average teacher.
[0026] Step 104: Construct the inter-domain difference loss using the joint maximum mean variance difference measure method.
[0027] Step 105: Determine the mutual information loss based on the lower bound of mutual information calculated by the mutual information neural estimator under set parameters.
[0028] Step 106: Using the average teacher's unsupervised transfer training strategy, the student model in the unsupervised transfer training framework is trained and updated by combining classification loss, inter-domain difference loss and mutual information loss, and the teacher model in the unsupervised transfer training framework is trained and updated by using the exponential moving average strategy, so as to obtain the unsupervised transfer diagnostic model.
[0029] Step 107: Input the preprocessed original vibration signal into the unsupervised transfer diagnosis model to obtain the transfer diagnosis results.
[0030] By implementing steps 100-107 above, this application can effectively improve fault diagnosis performance by utilizing a domain adaptive network guided by mutual information, addressing the problems in current unsupervised domain adaptive methods such as difficulty in aligning time-varying data distributions, the inherent contradiction between the scarcity of source domain labeled samples and actual needs.
[0031] In another exemplary embodiment of this application, in practical applications, step 100 described above can be performed using multiple sensors to collect the original time-series vibration acceleration signals (i.e., original vibration signals) of rotating machinery, covering various health states and operating conditions in both the source and target domains. The preprocessing method can be: dividing the original time-series vibration acceleration signals of the rotating machinery collected by multiple sensors into equal-length segments, performing a fast Fourier transform, and then taking the absolute value to obtain the frequency domain spectrum signal of the rotating machinery collected by multiple sensors.
[0032] In another exemplary embodiment of this application, step 101 is essentially to construct an initialization network guided by mutual information, and the implementation process is mainly described in the following two parts (1) and (2).
[0033] (1) Enhanced Convolutional Neural Network. A one-dimensional enhanced convolutional neural network based on the enhanced local channel attention mechanism (ELCA) was designed. Specifically, the enhanced convolutional neural network consists of four feature extraction modules, three enhanced local channel attention modules (referred to as ELCA modules), two fully connected layers (Dense), one exponential linear unit (ELU), and one SoftMax layer.
[0034] like Figure 2As shown, an ELCA module is set up between the four feature extraction modules. The fourth feature extraction module is connected to the first fully connected layer (Dense) through a flattening operation. The first fully connected layer (Dense), the exponential linear unit (ELU), the second fully connected layer (Dense), and the SoftMax layer are connected sequentially. Each of the four feature extraction modules consists of a one-dimensional convolutional layer (Covn1D), a max pooling layer, a batch normalization layer, and an exponential linear unit (ELU). Figure 2 The structure shown introduces a newly designed ELCA module, which has been proven in subsequent experiments to have excellent channel weight extraction and feature extraction capabilities. Each convolutional layer is followed by a batch normalization layer, which accelerates network convergence. Unlike typical networks that use the ReLU activation function, this network employs the ELU activation function, which optimizes the network's learning and feature extraction capabilities.
[0035] Based on the above description, the augmented convolutional neural network is essentially composed of four one-dimensional convolutional layers (Covn1D), four max pooling layers, four batch normalization layers, four exponential linear units (ELU), three ELCA modules, one fully connected layer (Dense), one ELU, one fully connected layer, and one SoftMax layer.
[0036] The ELCA module accepts feature maps generated from previous convolutional layers, batch normalization layers, exponential linear units, and max pooling layers. The architecture of the ELCA module is as follows: Figure 4 As shown, the ELCA module consists of a compressed excitation branch (SE branch), a residual branch, and an enhanced local attention branch (ELA branch).
[0037] Assuming the feature map is the input to the ELCA module, where L The length of the feature map, C This refers to the dimension (number of channels) of the feature map. In the compressed excitation branch, the feature map... Map First, a global average pooling (GAP) operation is performed to generate compressed features along the channel dimension. These compressed features are then processed through a compressed fully connected layer (FC), a rectified linear function (ReLU), an activated fully connected layer (FC), and a sigmoid function to obtain the primary channel importance. E , represented as: .
[0038] In the formula,h This represents the h-th data point of the compressed feature. B Indicates the number of feature maps in the batch. and These represent the features after the compressed fully connected layer and the features after the activated fully connected layer, respectively. r Indicates the compression ratio. Represents the linear rectified function. This represents the Sigmoid function. Following this, the importance will be... E and L Multiplying vectors of dimension 1 with all elements equal to 1 yields the importance of the excitation channels. , represented as: .
[0039] in, The per-element penalty for broadcasting, l L The above indicates L A vector of all 1s.
[0040] In the enhanced local attention module, feature maps Map It undergoes channel-level averaging, followed by depthwise separable convolution (DSConv1D), batch normalization (BN), and sigmoid transformation to enhance local attention importance. W L , represented as: .
[0041] in, This represents a depthwise separable convolution with a kernel length of k, and BN(·) represents batch normalization. In the residual branch, the original input feature map is output, and finally, it is combined with the local attention importance enhancement. W L and the importance of incentive channels W SE Perform channel-by-channel multiplication to obtain the final output feature map. Y , represented as: .
[0042] in, This indicates channel-by-channel multiplication. X This represents the original input feature map.
[0043] Based on the above description, in this application, the compression excitation branch is used to compress the features extracted by the feature extraction module to obtain the importance of the excitation channels. The enhanced local attention branch is used to perform local attention enhancement processing on the features extracted by the feature extraction module to obtain enhanced local attention importance. The residual branch is used to multiply the features extracted by the feature extraction module with the enhanced local attention importance and the excitation channel importance channel by channel to obtain the output feature map.
[0044] Clearly, by inputting feature maps into the ELCA module, not only can the importance of different channels be distinguished, but channel-level attention recalibration can also be adaptively performed based on the global average and learnable weights. This enables the enhanced convolutional neural network to achieve effective noise separation while preserving diagnostically critical low-amplitude signal components.
[0045] (2) Mutual Information Neural Estimator (MINE).
[0046] Considering the impact of large data distribution deviations under time-varying operating conditions on fault diagnosis, this application designs a mutual information neural estimator for multidimensional variable mutual information estimation. This mutual information neural estimator consists of a projection head and a mutual information neural estimator body. The overall mutual information estimation process is expressed as follows: .
[0047] .
[0048] in, The parameter is Mutual information neural estimator, The variable represented by parameter is estimated by a mutual information neural estimator. P and Q Mutual information The lower bound, Indicates about P and Q joint distribution Expectations Represents about variables P and Q joint distribution Expectations * indicates the distribution of the variable *.
[0049] For example, with As the input feature vector of MINE, As the input prediction labels for MINE, after being projected by the projection head, both will be projected into the same d-dimensional space, represented as: .
[0050] .
[0051] in, The projection head represents the eigenvector. This represents the projection head for predicting labels. d z The dimension of the feature vector. d y This represents the dimension of the predicted label vector. Indicates the process Figure 3 The feature vector output by the projection head in the process is a post-processed vector. Indicates the process Figure 3 The predicted label post-processing vector output after the projection head in the image.
[0052] Positive sample pairs were then constructed for the mutual information neural estimator. and negative sample pairs The construction process is represented as follows: .
[0053] .
[0054] in, Indicates will The shuffled random sample pairs have n One negative sample.
[0055] The mutual information neural estimator is then initialized with the following parameters: In the k-th update of the mutual information neural estimator, the mutual information is calculated by assigning scores to positive and negative sample pairs, and is represented as: .
[0056] .
[0057] in, The score of the feature vector-predicted label sample pair represents the feature vector. The score for the feature-predicted label sample pair. This represents the parameters of the mutual information neural estimator after the (k-1)th update. By calculating the above scores, the lower bound of the estimated mutual information can be obtained. , represented as: .
[0058] After obtaining the estimated lower bound of mutual information, the mutual information neural estimator is updated. The update process is represented as follows: .
[0059] in, This represents the learning rate of the mutual information neural estimator. Indicates parameters The update gradient, This represents the parameter that is updated during the k-th iteration. This represents the parameters updated during the (k-1)th iteration. The total number of updates, K, of the mutual information neural estimator is set manually; updates stop when k=K.
[0060] In another exemplary embodiment of this application, considering the limitations of traditional distance measurement methods in measuring the distance between data with high similarity, this application designs a joint maximum mean variance difference measure (JMMVD) method based on multidimensional statistical alignment in step 102 above.
[0061] In this embodiment, the JMMVD method consists of two parts: Joint Maximum Mean Difference (JMMD) and Maximum Variance Difference (MVD), expressed as follows: .
[0062] in, Indicates the maximum variance difference. Indicates the joint maximum mean difference. They represent the first i Extraction features of the source domain sample and the first j Extracting features from samples in the target domain. They represent the first i The predicted label of the first source domain sample and the first j Predicted labels for each target domain sample.
[0063] For example, the calculation process for the joint maximum mean difference is expressed as follows: .
[0064] For example, the calculation process for the maximum variance difference is expressed as follows: .
[0065] In the formula, M Indicates the number of samples in the source domain. k ( ) represents the Gaussian kernel function. N Indicates the number of samples in the target domain. , They represent the first j Extraction features of the source domain sample and the first i Extracting features from samples in the target domain. , They represent the first j Extraction features of the source domain sample and the firsti Predicted labels for each target domain sample. For the maximum variance difference, Var ( The ) represents the variance of the extracted features in the RKHS. k 1( ) represents the t-student kernel function.
[0066] Among them, Gaussian kernel function , represented as: In the formula, For hyperparameters, x, y This is the input variable used for the Gaussian kernel function example operation.
[0067] The t-student kernel function can be expressed as: .
[0068] In the formula, is the hyperparameter of the t-student kernel function.
[0069] In another exemplary embodiment of this application, considering the problem of insufficient training in traditional simple gradient update methods and the problem of training divergence easily caused by adversarial training, this application constructs an unsupervised transfer training framework based on average teachers in step 103 above, which guides the student model update through the teacher model to achieve better training results. For example, combined with the description of steps 104-106, the unsupervised transfer training framework based on average teachers can be divided into training data augmentation, loss function construction, and updating of the student and teacher models. Based on this, the updating process of the teacher and student models includes: 1) Training data augmentation.
[0070] To enable the teacher and student models to extract more fault features, all training data X undergoes random data augmentation in each training epoch, as follows: .
[0071] in, Indicates that it follows a Gaussian distribution noise, To conform to the distribution Random scaling factor, This indicates the enhanced sample.
[0072] 2) Construction of loss function.
[0073] The unsupervised transfer training strategy based on average teachers effectively updates the student model by combining multiple loss functions. Therefore, the loss function can be expressed as: .
[0074] in, Represents classification loss. Indicates mutual information loss. This represents the inter-domain difference loss (i.e., the inter-domain difference loss between the source domain and the target domain). To balance the hyperparameters of each loss, express The maximum value, express The minimum value, express The fastest growing point Indicates the number of training sessions. Indicates control Hyperparameters that determine the rate of change.
[0075] For example, classification loss It is calculated in the following way: .
[0076] .
[0077] .
[0078] in, The loss for the student model's prediction of labeled samples in the source domain. The consistency loss between the teacher model and the student model predictions. and Let represent the predicted labels of the student model for labeled samples in the source domain and the predicted labels for unlabeled samples in the source domain, respectively. and Let represent the predicted labels of the teacher model for labeled samples in the source domain and the predicted labels for unlabeled samples in the source domain, respectively. Indicates the total number of samples in the source domain. Indicates the total number of categories. Indicates the first i The true labels of the source domain samples Indicates the first i Predicted labels for each source domain sample.
[0079] For example, mutual information loss For parameters The lower bound of mutual information, calculated by the mutual information neural estimator, is obtained by the following formula: .
[0080] For example, inter-domain difference loss The following formula is used to calculate: .
[0081] .
[0082] .
[0083] In the formula, for and The joint maximum mean difference between them This represents the joint maximum mean difference loss. and These are the features extracted and predicted labels for the enhanced target domain samples by the student model, respectively. and These represent the features extracted and predicted labels for the teacher model on enhanced labeled source domain samples, respectively. To enhance the feature extraction of labeled source domain samples for the teacher model. The student model enhances the extraction features of target domain samples. The maximum variance difference between them. This represents the loss due to the largest variance difference.
[0084] 3) Update of student and teacher models.
[0085] To enhance the stability of model updates while learning more discriminative fault features, an exponential moving average (EMA) strategy is introduced to provide more stable updates to the teacher model. This stable teacher model then guides the student model to learn more key fault features. For example, the update process for both the student and teacher models at the e-th update is as follows: .
[0086] .
[0087] in, The learning rate for the student model. Represents the gradient of the student model. and These are the parameters before and after the teacher model update, respectively. The attenuation coefficient of the EMA. and These are the parameters before and after the student model update.
[0088] Based on the above description, after the teacher and student models have been trained, an unsupervised transfer diagnostic model can be obtained. At this point, the frequency domain spectral signals of the test set in the target domain will be input into the unsupervised transfer diagnostic model, and they will only pass through a frozen, enhanced convolutional neural network, where the classifier module will output the diagnostic classification results.
[0089] In summary, the method provided in this application first uses multiple sensors to collect raw time-series vibration acceleration signals of mechanical signals. Then, it performs a Fast Fourier Transform and takes the absolute value to obtain frequency domain feature samples of the fault data. Next, it constructs a domain adaptive network guided by mutual information. Enhanced convolutional neural networks are then constructed separately. A mutual information neural estimator is also built. Furthermore, a joint maximum mean-variance difference measure method based on multidimensional statistical alignment is designed. By measuring inter-domain differences at the mean and variance levels, a more accurate inter-domain distance is obtained. Subsequently, an unsupervised transfer training framework based on average teachers is constructed, achieving a balance between stable updates and powerful feature extraction capabilities by guiding student model updates through a teacher model. Finally, the final transfer diagnostic result is output. The frequency domain feature samples of the target domain test set are input into the trained student model, and they are only processed through the enhanced convolutional neural network to output the final diagnostic classification result. This application fully considers the impact of different working conditions in industrial scenarios on diagnosis, providing a new method for fault diagnosis and important technical support for the safe and reliable operation of equipment. Inter-domain distance refers to the distance between the data distributions of the source domain and the target domain. Estimating and minimizing this distance is one of the core aspects of unsupervised transfer learning diagnostics. The constructed enhanced convolutional neural network and estimator are the foundation for unsupervised transfer learning training; only by training them can unsupervised transfer learning diagnostics be achieved.
[0090] In another exemplary embodiment of this application, such as Figure 5 As shown, the MIGDAN (Mutual Informationguided Domain Adaptation Network) method is the method proposed in this application. DANN is a domain adversarial neural network, DAN is a domain adaptive network, DCTLN is a deep convolutional transfer learning network, DCMADA is a deep convolutional multi-adversarial network, and GADAN is a gradient-aligned domain adversarial network. The diagnostic accuracy of the proposed MIGDAN method in 12 different transfer diagnostic scenarios was evaluated. The MIGDAN method achieved the highest diagnostic accuracy and best diagnostic stability in all transfer diagnostic tasks, with an average diagnostic accuracy of 98.60% (standard deviation 0.85%). Compared with advanced methods such as DAN, DANN, DCTLN, DCMADA, and GADAN, its average diagnostic accuracy is improved by 19.99%, 20.24%, 13.19%, 9.80%, and 16.74%, respectively, fully verifying the effectiveness and superiority of the method provided in this application.
[0091] Furthermore, to verify the diagnostic capability of the MIGDAN method proposed in this application, ablation experiments were conducted. Table 1 lists the ablation experiment results under various transfer tasks. For the key modules in the MIGDAN method, MIGDAN-SEBlock (replacing the ELCA module with SEBlock), MIGDAN-woELCA (removing all ELCA modules), MIGDAN-woMINE (removing the mutual information neural estimator), and MIGDAN-MMD (replacing JMMVD with maximum mean difference) were constructed respectively. Compared to the MIGDAN-wo-ELCA baseline model, the MIGDAN method provided in this application leads in accuracy by 20.58%. Compared to MIGDAN-SEBlock, the MIGDAN method provided in this application leads in accuracy by 6.13%, verifying the effectiveness of the ELCA module. Compared to the traditional SEBlock, the ELCA module achieves more precise channel weight allocation, allowing the enhanced CNN to focus its attention on more discriminative fault features, thereby improving cross-domain diagnostic accuracy. Compared to MIGDAN-wo-MINE, the MIGDAN method exhibits significant advantages in both average diagnostic accuracy and stability, consistently maintaining superior performance across all cross-domain diagnostic tasks. This result demonstrates that the Mutual Information Neural Estimator (MINE) provided in this application effectively guides the enhanced convolutional neural network to establish tighter feature-label associations, enabling it to learn features more dependent on labels, ultimately improving cross-domain diagnostic performance. The MIGDAN method provided in this application achieves a 2.95% improvement in average diagnostic accuracy compared to MIGDAN-MMD, and its diagnostic performance is more stable. This result validates the effectiveness of the Joint Maximum Mean Variance Distribution (JMMVD) metric. Compared to the traditional MMD metric, which only utilizes the mean to calculate domain differences, JMMVD enhances the alignment capability of cross-domain data distributions, thereby enabling the network to achieve more accurate cross-domain diagnostic results.
[0092] Table 1 Ablation experimental results under various migration tasks
[0093] By introducing a novel ELCA module, the MIGDAN method provided in this application guides the enhanced CNN to focus on more discriminative channels and features, achieving more accurate cross-domain fault diagnosis with the enhanced feature extraction capability.
[0094] Based on the above description, the method provided in this application has the following advantages compared to the prior art: 1. The MIGDAN method proposed in this application can achieve very accurate unsupervised transfer diagnosis when source domain labeled samples are scarce.
[0095] 2. The introduced mutual information neural estimator can maximize the extraction of mutual information between features and labels, strengthening the feature-label correlation. Furthermore, by guiding the enhanced convolutional neural network to maximize the mutual information between features and labels, a strongly dependent feature-label correspondence system (i.e., a tighter feature-label relationship) is constructed, prompting the network to extract more fault features.
[0096] 3. The designed ELCA module can achieve more effective feature extraction and channel attention extraction, guiding the network to focus on more important features and overcoming the limitation of traditional attention mechanisms in extracting low-amplitude features.
[0097] 4. The designed unsupervised transfer training strategy based on average teachers combines stable student model updates with powerful feature extraction capabilities, resulting in better training efficiency.
[0098] 5. An unsupervised mean teacher training strategy that integrates classification loss, mutual information loss and JMMVD loss is proposed, which endows the method provided in this application with excellent feature extraction and generalization capabilities, thereby establishing a strong correlation between features and labels to achieve accurate cross-domain fault diagnosis.
[0099] 5. This application constructs a transfer diagnosis task based on time-varying operating conditions and a scenario with scarce source domain labels, and completes comprehensive experimental verification on two highly challenging mechanical fault datasets. Through rigorous comparison with state-of-the-art methods and ablation experiments, the superiority of the method provided in this application is fully demonstrated.
[0100] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores mutual information guidance domain adaptive data for intelligent migration diagnostics of rotating machinery. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a mutual information guidance domain adaptive method for intelligent migration diagnostics of rotating machinery.
[0101] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0102] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0103] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0104] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0105] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0106] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (RRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0107] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0108] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0109] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A mutual information guided domain adaptation method for intelligent migration diagnosis of rotating machinery, characterized in that, The method comprises the following steps: obtaining original vibration signals of a rotating machine in different fault states and preprocessing the original vibration signals; the original vibration signals cover various health states and working conditions of source domains and target domains; constructing a mutual information guided domain adaptive network; the domain adaptive network comprises an enhanced convolutional neural network with an enhanced local channel attention module and a mutual information neural estimator for multi-dimensional variable mutual information estimation; constructing a joint maximum mean variance difference measurement method with multi-dimensional statistical quantity alignment; constructing an average teacher based unsupervised transfer training framework with the domain adaptive network as a teacher model and a student model respectively; constructing an inter-domain difference loss by using the joint maximum mean variance difference measurement method; determining a mutual information loss based on a lower bound of mutual information calculated by the mutual information neural estimator under a set parameter; training and updating the student model in the unsupervised transfer training framework by using an average teacher unsupervised transfer training strategy in combination with a classification loss, the inter-domain difference loss and the mutual information loss, and training and updating the teacher model in the unsupervised transfer training framework by using an exponential moving average strategy, to obtain an unsupervised transfer diagnosis model; inputting the preprocessed original vibration signals into the unsupervised transfer diagnosis model to obtain a transfer diagnosis result.
2. The mutual information guided domain adaptation method for intelligent migrating diagnosis of rotating machinery according to claim 1, characterized in that, The enhanced convolutional neural network with the enhanced local channel attention module is composed of four feature extraction modules, three enhanced local channel attention modules, two fully connected layers, an exponential linear unit and a SoftMax layer; the enhanced local channel attention modules are arranged between the four feature extraction modules; the fourth feature extraction module is connected with the first fully connected layer through a flattening operation; the first fully connected layer, the exponential linear unit, the second fully connected layer and the SoftMax layer are connected in sequence.
3. The mutual information guided domain adaptation method for intelligent migrating diagnosis of rotating machinery according to claim 2, characterized in that, Each of the four feature extraction modules is composed of a one-dimensional convolutional layer, a maximum pooling layer, a batch normalization layer and an exponential linear unit.
4. The mutual information guided domain adaptation method for intelligent migrating diagnosis of rotating machinery according to claim 2, characterized in that, The enhanced local channel attention module is composed of a compression excitation branch, a residual branch and an enhanced local attention branch; the compression excitation branch is used for compressing the features extracted by the feature extraction module to obtain excitation channel importance; the enhanced local attention branch is used for performing local attention enhancement processing on the features extracted by the feature extraction module to obtain enhanced local attention importance; the residual branch is used for multiplying the features extracted by the feature extraction module with the enhanced local attention importance and the excitation channel importance channel by channel to obtain an output feature map.
5. The mutual information guided domain adaptation method for intelligent migrating diagnosis of rotary machines according to claim 1, wherein, The mutual information neural estimator is composed of a projection head and a mutual information neural estimator body.
6. The mutual information guided domain adaptation method for intelligent migrating diagnosis of rotary machines according to claim 1, wherein, The joint maximum mean variance difference measurement method with multi-dimensional statistical quantity alignment comprises joint maximum mean difference and maximum variance difference; the calculation process of the joint maximum mean difference is represented as: ; the calculation process of the maximum variance difference is represented as: ; In the formula, Indicates the joint maximum mean difference. They represent the first i Extraction features of the source domain sample and the first j Extracting features from samples in the target domain. They represent the first i The predicted label of the first source domain sample and the first j Predicted labels for each target domain sample. M Indicates the number of samples in the source domain. k ( ) represents the Gaussian kernel function. N Indicates the number of samples in the target domain. , They represent the first j Extraction features of the source domain sample and the first i Extracting features from samples in the target domain. , They represent the first j Extraction features of the source domain sample and the first i Predicted labels for each target domain sample. For the maximum variance difference, Var ( The ) represents the variance of the extracted features in the RKHS. k 1( ) represents the t-student kernel function.
7. The mutual information guided domain adaptation method for intelligent migrating diagnosis of rotary machines according to claim 1, wherein, the inter-domain difference loss is represented as: ; ; ; In the formula, is and The joint maximum mean difference between represents the joint maximum mean difference loss, and The extracted features and predicted labels of the student model on the enhanced target domain samples are and The extracted features and predicted labels of the teacher model on the enhanced labeled source domain samples are The extracted features of the teacher model on the enhanced labeled source domain samples and the extracted features of the student model on the enhanced target domain samples The maximum variance difference between represents the maximum variance difference loss, is the domain difference loss.
8. The mutual information guided domain adaptation method for intelligent migrating diagnosis of rotary machines according to claim 1, wherein, the loss function formed by the classification loss, the inter-domain difference loss and the mutual information loss is represented as: ; wherein, denotes the total loss, denotes the classification loss, denotes the mutual information loss, denotes the inter-domain discrepancy loss, and are hyper-parameters balancing the respective losses, e is the natural logarithm.
9. A computer device comprising: A memory, a processor, and a computer program stored on the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the mutual information guided domain adaptation method for intelligent migration diagnosis of rotating machinery according to any one of claims 1-8.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the mutual information guided domain adaptation method for intelligent migration diagnosis of rotating machinery according to any one of claims 1-8.