Fault diagnosis method based on multi-mode cross-sensor Transform
By adopting the multimodal cross-sensor Transformer method in fault diagnosis, using channel embedding and self-supervised learning technology, the problem of insufficient single modal data is solved, and efficient fusion and feature extraction of multimodal data is achieved, which significantly improves the accuracy and robustness of fault diagnosis.
Patent Information
- Application Number
- CN202510143718.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
AI Technical Summary
The existing fault diagnosis methods rely on single modal data, fail to fully utilize the complementarity and synergy of multimodal data, and obtaining high-quality labeled data is expensive, which limits the training and application of deep learning models.
The fault diagnosis method based on multimodal transsensor Transformer is adopted to convert multi-sensor data into high-dimensional features through channel embedding, and the correlation between time and frequency domain features is captured using self-supervised learning, and the model training and prediction process is optimized through uncertainty weighted losses and Sharple value weighted output.
It significantly improves the accuracy and robustness of fault diagnosis, effectively integrates multimodal data, reduces dependence on a large amount of labeled data, reduces data acquisition costs, and enhances the application value of the model in actual industrial environments.
Smart Images

Figure CN120086679A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a fault diagnosis method based on a multi-modal cross-sensor Transformer. Background Art
[0002] With the rapid development of modern industry, the complexity and intelligence level of large-scale mechanical equipment are constantly increasing. These devices are prone to various faults during long-term operation in complex environments, resulting in production losses and even safety accidents. Therefore, intelligent fault diagnosis technology is of great significance in ensuring the normal operation of equipment, improving production efficiency, and guaranteeing industrial safety. Traditional fault diagnosis methods mainly rely on signal processing technologies such as Fourier transform and wavelet transform. Although these methods are effective in some cases, they usually require professional knowledge and have limitations in terms of adaptability and efficiency, making it difficult to cope with complex industrial environments and diverse fault types.
[0003] In recent years, deep learning technology has been widely applied in the field of fault diagnosis. Deep learning models can automatically learn features from raw data, have powerful pattern recognition capabilities and adaptability, and are suitable for processing large-scale datasets and real-time diagnosis tasks. Commonly used deep learning models include convolutional neural networks, recurrent neural networks, and Transformers, etc. Among them, the Transformer architecture has gradually been applied in the field of fault diagnosis due to its advantages in feature extraction and adaptability. The Transformer effectively captures long-range dependencies in data through the self-attention mechanism, is suitable for processing time series data, and can better identify the dynamic behavior and feature patterns of fault signals.
[0004] However, most of the existing fault diagnosis methods rely on single-modal data, such as time-domain or frequency-domain data, and fail to fully utilize the complementarity and synergy of multi-modal data. In addition, it is often difficult and costly to obtain a large amount of high-quality labeled data in practical applications, which limits the training and application of deep learning models. To solve these problems, multi-modal learning and self-supervised learning technologies have emerged. Multi-modal learning integrates information from different data modalities (such as time domain and frequency domain) to extract a more comprehensive feature set, thereby improving the accuracy and robustness of the prediction model. Self-supervised learning solves the problem of insufficient labeled data by learning latent representations from a large amount of unlabeled data, enabling the model to learn effective feature representations without supervision.
[0005] In summary, in order to improve the accuracy and robustness of fault diagnosis, there is an urgent need for an intelligent fault diagnosis method that can effectively integrate multi-modal data and use self-supervised learning for feature extraction. Summary of the Invention
[0006] The object of the present invention is to provide a fault diagnosis method based on a multi-modal cross-sensor Transformer, which converts multi-sensor data into high-dimensional features through channel embedding, captures the correlation between time-domain and frequency-domain features using self-supervised learning, and optimizes the model training and prediction processes through uncertainty-weighted loss and Shapley value-weighted output, realizing the efficient fusion and feature extraction of multi-modal data and significantly improving the accuracy and robustness of fault diagnosis.
[0007] To achieve the above object, the present invention provides a fault diagnosis method based on a multi-modal cross-sensor Transformer, including the following steps:
[0008] Step S1, in the data preprocessing stage, perform Fourier transform on the input multi-sensor time-domain data to obtain frequency-domain data;
[0009] Step S2, in the training stage, for each sample, first calculate the embedding vectors of the multi-modal data; secondly, input the multi-modal embedding vectors into the time encoder and frequency encoder respectively to extract features; then, enhance the multi-modal features to generate two different enhanced views; finally, calculate the time contrast loss, frequency-domain contrast loss and time-frequency contrast loss, and dynamically adjust the weights of the loss function according to the uncertainty weighting mechanism;
[0010] Step S3, in the fine-tuning stage, use the labeled data to fine-tune the parameters of the classifier while keeping the encoder parameters unchanged;
[0011] Step S4, in the deployment stage, use the multi-modal output weighting method based on the Shapley value to integrate the outputs of different modalities and calculate the final classification result.
[0012] Preferably, in step S2, calculating the embedding vectors of the multi-modal data, the specific process is as follows:
[0013] The data of each sensor is regarded as an independent channel, and the data of each channel is regarded as a token. These tokens are mapped into a vector space of a fixed dimension through a linear transformation, and then passed through a Dropout layer to increase the generalization ability of the model, as follows:
[0014]
[0015] Among them, represents the time-domain data of the k-th sensor of the i-th sample; is the vector after time-domain embedding; represents the frequency-domain data of the k-th sensor of the i-th sample; is the vector after frequency-domain embedding.
[0016] Preferably, in step S2, the multi-modal embedding vectors are respectively input into the time encoder and the frequency encoder to extract features as follows:
[0017]
[0018] Among them, G T and G F are the time domain encoder and the frequency domain encoder respectively; and are the feature representations in the time domain and the frequency domain respectively.
[0019] Preferably, in step S2, the multi-modal features are enhanced to generate two different enhanced views as follows:
[0020]
[0021] Among them, and are the two views after time domain enhancement; and are the two views after frequency domain enhancement;
[0022] Then, and are input into the projection head for non-linear mapping to obtain features and
[0023] Preferably, in step S2, the time contrast loss, the frequency domain contrast loss and the time-frequency contrast loss are calculated. The specific process is as follows:
[0024] Step S241: For each time series sample Take it and its corresponding non-linearly transformed feature as a positive sample pair, and the non-linear transformation of other samples as a negative sample pair. Then the time contrast loss L T is as follows:
[0025]
[0026] Among them, represents the cosine similarity between x and y; N is the number of samples; is the indicator function, which takes the value of 0 when j = i and 1 when j ≠ i; τ is the temperature coefficient;
[0027] Step S242: For each frequency domain sample Take it and its corresponding non-linearly transformed feature as a positive sample pair, and the non-linear transformation of other samples as a negative sample pair. Then the frequency domain contrast loss LF , as follows:
[0028]
[0029] Step S243, time-frequency contrast loss L C , as follows:
[0030]
[0031] where L CT represents the time-domain cross-contrast loss;
[0032]
[0033] where L CF represents the frequency-domain cross-contrast loss;
[0034] L C = αL CT +(1 - α)L CF ;
[0035] where α is the weight coefficient;
[0036] Step S244, total loss function L total , as follows:
[0037] L total = ω T L T + ω F L F + ω C L C ;
[0038] where ω T , ω F and ω C are the weight coefficients of the time contrast loss, frequency contrast loss, and time-frequency contrast loss, respectively.
[0039] Preferably, in step S2, the weights of the loss function are dynamically adjusted according to the uncertainty weighting mechanism, and the specific process is as follows:
[0040] Introduce three learnable uncertainty parameters α T , α F and α C , corresponding to the weights of the time contrast loss, frequency contrast loss, and time-frequency contrast loss respectively; use the softmax function to normalize these uncertainty parameters to ensure that the sum of all weights is 1, while keeping the scale of the total loss unchanged and allowing the model to dynamically adjust the weights according to the importance of the features. The specific formula is as follows:
[0041]
[0042] During the training process, the uncertainty parameters are optimized by the gradient descent method, and the model automatically adjusts the weights of the losses of each modality according to the feature importance in the current training state.
[0043] Preferably, in step S4, in the deployment stage, a multi-modal output weighting method based on the Shapley value is used to integrate the outputs of different modalities and calculate the final classification result. The specific process is as follows:
[0044] Step S41: First, calculate the time-domain Shapley value That is, the increase in the model's prediction ability after adding time-domain features when only frequency-domain features are available, as follows:
[0045]
[0046] where γ{T,F} is the model's prediction when using both time-domain and frequency-domain features; γ{F} is the model's prediction when only using frequency-domain features;
[0047] Then, calculate the frequency-domain Shapley value That is, the increase in the model's prediction ability after adding frequency-domain features when only time-domain features are available, as follows:
[0048]
[0049] Step S42: Calculate the weight of each modality in the model output according to the Shapley value of each modality, as follows:
[0050]
[0051] where κ T represents the time-domain weight; κ F represents the frequency-domain weight;
[0052] Step S43: During the inference process, the final prediction of the model is the weighted sum of the time-domain output and the frequency-domain output, as follows:
[0053]
[0054] where and are the time-domain and frequency-domain outputs respectively.
[0055] Therefore, the present invention adopts the above-mentioned fault diagnosis method based on multi-modal cross-sensor Transformer, and the beneficial effects are as follows:
[0056] (1) Through the channel embedding technology, the present invention embeds the time-domain and frequency-domain data from different sensors into a high-dimensional space respectively, retains the long-term dependence of the data and the variable relationship across sensors, and solves the deficiencies of traditional fault diagnosis methods in multi-modal data fusion and feature extraction;
[0057] (2) The present invention adopts a multi-modal self-supervised learning strategy. Through time contrast loss, frequency contrast loss and time-frequency contrast loss, it captures the dynamic features of time series and the periodic features of the frequency domain simultaneously, and enhances the model's ability to identify fault signals;
[0058] (3) The present invention introduces an uncertainty weighted loss mechanism to dynamically adjust the weights of different modal losses, enabling the model to focus on learning the most informative features, and improving the accuracy and robustness of diagnosis;
[0059] (4) In the model deployment stage, the present invention adopts a multi-modal output weighting method based on Shapley values to fairly allocate the contribution of different modalities to the model output, ensuring that key modalities have a greater impact on the prediction results and further improving the accuracy of prediction;
[0060] (5) The present invention can effectively integrate multi-modal data, make full use of its complementarity and synergy effect, significantly improve the accuracy and robustness of fault diagnosis. At the same time, the application of self-supervised learning reduces the dependence on a large amount of labeled data, reduces the data acquisition cost, and enhances the application value of the model in the actual industrial environment.
[0061] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0062] Figure 1 is a flowchart of a fault diagnosis method based on multi-modal cross-sensor Transformer of the present invention;
[0063] Figure 2 is a model framework diagram in an embodiment of the present invention. Detailed Embodiments
[0064] The technical solution of the present invention will be further described below through the drawings and embodiments.
[0065] As Figure 1 shown, a fault diagnosis method based on multi-modal cross-sensor Transformer of the present invention includes the following steps:
[0066] Step S1: In the data preprocessing stage, perform Fourier transform on the input multi-sensor time-domain data to obtain frequency-domain data.
[0067] Step S2. In the training stage, for each sample, first calculate the embedding vectors of the multimodal data; second, input the multimodal embedding vectors into the temporal encoder and the frequency encoder respectively to extract features; then, enhance the multimodal features to generate two different enhanced views; finally, calculate the temporal contrast loss, the frequency-domain contrast loss, and the temporal-frequency contrast loss, and dynamically adjust the weights of the loss function according to the uncertainty weighting mechanism.
[0068] Step S21. Calculate the embedding vectors of the multimodal data.
[0069] The data of each sensor is regarded as an independent "channel", and the data of each channel is regarded as a token. These tokens are mapped into a vector space of a fixed dimension through a linear transformation (such as a fully connected layer), and then passed through a Dropout layer to increase the generalization ability of the model. This process can be expressed by the following formula:
[0070]
[0071] where represents the time-domain data of the k-th sensor of the i-th sample; is the vector after time-domain embedding; represents the frequency-domain data of the k-th sensor of the i-th sample; is the vector after frequency-domain embedding.
[0072] Step S22. Input the multimodal embedding vectors into the temporal encoder and the frequency encoder respectively to extract features, as follows:
[0073]
[0074] where G T and G F are the temporal encoder and the frequency encoder respectively; and are the temporal and frequency-domain feature representations respectively.
[0075] Step S23. Enhance the multimodal features to generate two different enhanced views, as follows:
[0076]
[0077] where and are the two views after temporal enhancement; and are the two views after frequency-domain enhancement.
[0078] Then, and The input projection head performs non-linear mapping to obtain features and
[0079] Step S24: Calculate the time contrast loss, frequency domain contrast loss, and time-frequency contrast loss, and dynamically adjust the weights of the loss function according to the uncertainty weighting mechanism.
[0080] Step S241: For each time series sample Take it and the corresponding non-linearly transformed feature As a positive sample pair, and the non-linear transformation of other samples As a negative sample pair, then the time contrast loss L T , is as follows:[[]]
[0081]
[0082] Where Represents the cosine similarity between x and y; N is the number of samples; Is the indicator function, taking the value of 0 when j = i and 1 when j ≠ i; τ is the temperature coefficient.
[0083] Step S242: For each frequency domain sample Take it and the corresponding non-linearly transformed feature As a positive sample pair, and the non-linear transformation of other samples As a negative sample pair, then the frequency domain contrast loss L F , is as follows:[[]]
[0084]
[0085] Step S243: The time-frequency contrast loss L C , is as follows:[[]]
[0086]
[0087] Where, L CT Represents the time domain cross-pair loss.
[0088]
[0089] Where, L CF Represents the frequency domain cross-contrast loss.
[0090] L C = αL CT +(1 - α)L CF ;
[0091] Where, α is the weight coefficient.
[0092] Step S244: The total loss function Ltotal , as follows:
[0093] L total = ω T L T + ω F L F + ω C L C ;
[0094] where ω T , ω F and ω C are the weight coefficients of the temporal contrast loss, frequency contrast loss, and time-frequency contrast loss, respectively.
[0095] Step S245: Dynamically adjust the weights of the loss function according to the uncertainty weighting mechanism. The specific process is as follows:
[0096] Introduce three learnable uncertainty parameters α T , α F and α C , corresponding to the weights of the temporal contrast loss, frequency contrast loss, and time-frequency contrast loss respectively; use the softmax function to normalize these uncertainty parameters to ensure that the sum of all weights is 1. This can keep the scale of the total loss unchanged while allowing the model to dynamically adjust the weights according to the importance of the features. The specific formula is as follows:
[0097]
[0098] During the training process, optimize these uncertainty parameters by gradient descent. The model will automatically adjust the weights of each modal loss according to the feature importance in the current training state.
[0099] Step S3: In the fine-tuning stage, use the labeled data to fine-tune the parameters of the classifier while keeping the encoder parameters unchanged.
[0100] Step S4: In the deployment stage, use the multi-modal output weighting method based on the Shapley value to integrate the outputs of different modalities and calculate the final classification result.
[0101] Step S41: Using the multi-modal output weighting method based on the Shapley value to integrate the outputs of different modalities and calculate the final classification result means measuring the contribution of each modality to the model prediction by calculating the Shapley value of each modality and allocating weights according to these contribution degrees. This method can effectively solve the problem that directly averaging the multi-modal outputs may lead to dilution of key information. The specific process is as follows:
[0102] First, calculate the Shapley value in the time domain That is, in the case of only frequency-domain features, the increase in the model's prediction ability after adding time-domain features is as follows:
[0103]
[0104] where γ{T,F} is the model prediction when using both time-domain and frequency-domain features; γ{F} is the model prediction when using only frequency-domain features.
[0105] Then, calculate the frequency-domain Shapley value That is, in the case of only time-domain features, the increase in the model's prediction ability after adding frequency-domain features is as follows:
[0106]
[0107] Step S42: Calculate the weight of each modality in the model output according to the Shapley value of each modality, as follows:
[0108]
[0109] where κ T represents the time-domain weight; κ F represents the frequency-domain weight.
[0110] Step S43: During the inference process, the final prediction of the model is the weighted sum of the time-domain output and the frequency-domain output, as follows:
[0111]
[0112] where and are the outputs of the time domain and the frequency domain, respectively.
[0113] Embodiment
[0114] In this embodiment, the pump equipment of a chemical plant is taken as the object, aiming to realize the intelligent diagnosis of its fault state. The architecture of the model is as Figure 2 shown.
[0115] First, install multiple sensors at the key parts of the pump equipment, including vibration sensors and acoustic emission sensors, to collect the operation data of the equipment in real time. The data collection frequency is set to 5.12 kHz to ensure that the subtle changes during the operation of the equipment can be captured. The collected data includes time-domain signals and frequency-domain signals. For subsequent processing, perform a fast Fourier transform on the time-domain signals to convert them into frequency-domain signals for analyzing the characteristics of the equipment in both time and frequency dimensions.
[0116] The collected time-domain and frequency-domain data are input into the channel embedding module. This module treats the data of each sensor as an independent channel and embeds the data into a high-dimensional space through the self-attention mechanism. Specifically, for the time-domain data, the data at different time points are regarded as different channels. Through linear transformation and dropout operations, they are mapped into a vector space of a fixed dimension to generate embedding vectors. The frequency-domain data is also embedded in a similar way. This process preserves the long-term dependencies of the data and the variable relationships across sensors, laying a foundation for subsequent feature extraction and fault diagnosis.
[0117] The embedded data is fed into the multi-modal self-supervised learning module. This module contains a time encoder and a frequency encoder, which respectively extract features from the embedding vectors in the time domain and the frequency domain. Through two dropout operations, two different views are generated to increase the diversity of the data. Then, the time contrast loss, the frequency contrast loss, and the time-frequency contrast loss are calculated. The time contrast loss is used to capture the dynamic features of the time series, the frequency contrast loss is used to identify the periodic features in the frequency domain, and the time-frequency contrast loss promotes the model's comprehensive understanding of time-frequency information by comparing the similarities between the time-domain features and the frequency-domain features. By optimizing these contrast losses, the model can learn more discriminative feature representations, providing strong support for fault diagnosis.
[0118] During the model training process, an uncertainty-weighted loss mechanism is introduced. By introducing learnable uncertainty parameters and normalizing them using the softmax function, the weights of the time contrast loss, the frequency contrast loss, and the time-frequency contrast loss are dynamically adjusted. The model automatically adjusts the weights of the losses for each modality according to the feature importance in the current training state, making the training process pay more attention to the most informative features. The Adam optimizer is used to update the model parameters. After multiple iterative trainings, the model gradually converges and can accurately identify the fault features of the device.
[0119] After the model training is completed, in the deployment stage, a multi-modal output weighting method based on the Shapley value is adopted. First, the Shapley values of the time-domain and frequency-domain features for the model prediction are calculated to measure their contribution degrees. Then, the weights of each modality are calculated according to the Shapley values to ensure that the key modalities have a greater impact on the model output. In the actual diagnosis process, the outputs of the model in the time domain and the frequency domain are weighted and fused to obtain the final fault diagnosis result. By comparing with the actual fault status, the accuracy and reliability of the method of the present invention are verified, and it can effectively identify the fault types of the device, providing strong technical support for the maintenance and management of the device.
[0120] Therefore, the present invention adopts the above-mentioned fault diagnosis method based on multi-modal cross-sensor Transformer, converts multi-sensor data into high-dimensional features through channel embedding, uses self-supervised learning to capture the correlation of features in the time domain and frequency domain, and optimizes the model training and prediction processes through uncertainty weighted loss and Shapley value weighted output, realizes the efficient fusion and feature extraction of multi-modal data, and significantly improves the accuracy and robustness of fault diagnosis.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A fault diagnosis method based on multimodal cross-sensor Transformer, characterized in that: The following steps are involved: Step S1: In the data preprocessing stage, Fourier transform is performed on the input multi-sensor time domain data to obtain frequency domain data; Step S2: In the training phase, for each sample, the embedding vector of the multimodal data is first calculated; Secondly, the multimodal embedding vectors are input into the time encoder and frequency encoder to extract features. Then, the multimodal features are enhanced to generate two different sets of enhanced views. Finally, the time contrast loss, frequency domain contrast loss and time-frequency contrast loss are calculated, and the weights of the loss function are dynamically adjusted according to the uncertainty weighting mechanism. Step S3: In the fine-tuning stage, the parameters of the classifier are fine-tuned using labeled data, while the encoder parameters remain unchanged; Step S4: In the deployment phase, a multimodal output weighting method based on Shapley value is used to integrate the outputs of different modalities and calculate the final classification result.
2. A fault diagnosis method based on a multimodal cross-sensor Transformer according to claim 1, characterized in that: In step S2, the embedding vector of the multimodal data is calculated. The specific process is as follows: The data of each sensor is regarded as an independent channel, and the data of each channel is regarded as a token. These tokens are mapped to a vector space of fixed dimension through linear transformation, and then pass through the Dropout layer to increase the generalization ability of the model, as shown below: in, represents the time domain data of the kth sensor of the ith sample; is the vector after embedding in the time domain; represents the frequency domain data of the kth sensor for the i-th sample; is the vector after embedding in frequency domain.
3. A fault diagnosis method based on a multimodal cross-sensor Transformer according to claim 2, characterized in that: In step S2, the multimodal embedding vector is input into the time encoder and frequency encoder to extract features, as shown below: Among them, G T and G F They are time domain encoder and frequency domain encoder respectively; and They are the feature representations in time domain and frequency domain respectively.
4. A fault diagnosis method based on a multimodal cross-sensor Transformer according to claim 3, characterized in that: In step S2, the multimodal features are enhanced to generate two different sets of enhanced views, as shown below: in, and These are the two views after time domain enhancement; and These are two views after frequency domain enhancement; Then, and Input projection head for nonlinear mapping to obtain features and 5. A fault diagnosis method based on a multimodal cross-sensor Transformer according to claim 4, characterized in that: In step S2, the time contrast loss, frequency domain contrast loss and time-frequency contrast loss are calculated. The specific process is as follows: Step S241: for each time series sample Compare it with the corresponding nonlinear transformed features As a positive sample pair, nonlinear transformation with other samples As a negative sample pair, the temporal contrast loss L T , as shown below: in, represents the cosine similarity between x and y; N is the number of samples; is the indicator function, which takes the value of 0 when j=i and takes the value of 1 when j≠i; τ is the temperature coefficient; Step S242: for each frequency domain sample Compare it with the corresponding nonlinear transformed features As a positive sample pair, nonlinear transformation with other samples As a negative sample pair, the frequency domain contrast loss L F , as shown below: Step S243: time-frequency contrast loss L C , as shown below: Among them, L CT represents the cross-contrast loss in the time domain; Among them, L CF represents the cross-contrast loss in the frequency domain; L C =αL CT +(1-α)L CF ; Among them, α is the weight coefficient; Step S244, total loss function L total , as shown below: L total =ω T L T +oh F L F +oh C L C ; Among them, ω T ,ω F and ω C They are the weight coefficients of time contrast loss, frequency contrast loss and time-frequency contrast loss respectively.
6. A fault diagnosis method based on a multimodal cross-sensor Transformer according to claim 5, characterized in that: In step S2, the weight of the loss function is dynamically adjusted according to the uncertainty weighting mechanism. The specific process is as follows: Introducing three learnable uncertainty parameters α T , α F and α C , corresponding to the weights of time contrast loss, frequency contrast loss, and time-frequency contrast loss respectively; these uncertainty parameters are normalized using the softmax function to ensure that the sum of all weights is 1, while keeping the scale of the total loss unchanged, while allowing the model to dynamically adjust the weights according to the importance of the features. The specific formula is as follows: During the training process, the uncertainty parameters are optimized through the gradient descent method, and the model automatically adjusts the weights of each modal loss according to the feature importance in the current training state.
7. The fault diagnosis method based on multimodal cross-sensor Transformer according to claim 1 is characterized in that: In step S4, during the deployment phase, a multimodal output weighting method based on Shapley value is used to integrate the outputs of different modalities and calculate the final classification result. The specific process is as follows: Step S41: First, calculate the time domain Shapley value That is, when there are only frequency domain features, the increase in model prediction ability after adding time domain features is as follows: Among them, γ{T,F} is the model prediction when both time domain and frequency domain features are used; γ{F} is the model prediction when only frequency domain features are used; Then, calculate the frequency domain Shapley value That is, when there are only time domain features, the increase in model prediction ability after adding frequency domain features is as follows: Step S42: Calculate the weight of each mode in the model output according to its Shapley value, as shown below: Among them, κ T represents the time domain weight; κ F represents the frequency domain weight; Step S43: During the inference process, the model's final prediction is the weighted sum of the time domain output and the frequency domain output as follows: in, and They are the output in time domain and frequency domain respectively.
Citation Information
Patent Citations
Time sequence equipment fault diagnosis method based on comparison self-supervised learning
CN116070128A
Bearing diagnosis method and system based on multi-modal and multi-scale fusion network
CN117763494A
Multi-scale feature fusion Vision-Transform model method for rolling bearing fault diagnosis
CN118817309A
Bearing fault diagnosis method and device, computer equipment and storage medium
CN118940104A
Bearing fault diagnosis method based on time-frequency domain contrast learning under small sample
CN119249229A
Cited By
Photovoltaic fault diagnosis method and system based on multi-mode adaptive weighting
CN122173886A