Rotary machinery fault diagnosis method based on self-supervised time-frequency comparison fusion network
By adopting a self-supervised time-frequency comparison fusion network in rotating machinery fault diagnosis and using a multi-level contrast learning strategy to integrate time-frequency characteristics, the existing methods rely on labeled data and lack of multi-scale understanding are solved, and more efficient fault diagnosis effect is achieved.
Patent Information
- Application Number
- CN202510100188.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
Existing rotary machinery fault diagnosis methods rely on a large amount of labeled data, are difficult to obtain, and are overly dependent on single feature extraction, lack in-depth understanding of multi-scale dependence on fault signals, and fail to fully explore complementary information between the time domain and the frequency domain characteristics.
Using a method based on self-supervised time-frequency comparison and fusion network, the multi-level comparison learning strategy is used to integrate time-domain and frequency domain features, deeply explore their complementary information, and build a multi-scale comparison learning framework to improve the representation ability and diagnostic performance of the model.
It significantly improves the accuracy and efficiency of fault diagnosis, especially in complex working conditions and diversified signal environments, its diagnostic effect far exceeds that of existing methods and has broad application prospects and practical value.
Smart Images

Figure CN119939516A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a rotating machinery fault diagnosis method based on a self-supervised time-frequency comparison fusion network, and belongs to the field of fault diagnosis. Background Art
[0002] Existing methods for rotating machinery fault diagnosis usually rely on a large amount of labeled data. However, in practical applications, these labeled data are usually difficult to obtain, which limits the effectiveness of the methods. Therefore, how to use unlabeled data to improve model performance has become a key research issue. Self-supervised learning, as a learning method that does not require a large amount of labeled data, has shown significant application prospects in many fields in recent years. By designing pre-training tasks, self-supervised learning can extract effective features from unlabeled data and perform well in downstream tasks. However, existing methods mostly focus on time domain features, ignore the complementarity of frequency domain features, and there is relatively little research on deep-level interaction and fusion of features. As an important method of self-supervised learning, contrastive learning shows significant potential in scenarios where labeled data is scarce. By constructing contrast tasks (such as positive and negative sample comparison), contrastive learning can automatically extract useful features from unlabeled data and promote feature fusion. It has become a research hotspot in the field of fault diagnosis.
[0003] In recent years, self-supervised contrastive learning has been widely used in rotating machinery fault diagnosis. Some researchers use contrastive learning to align time and frequency domain features in latent space to explore physical consistency and achieve accurate fault diagnosis. Some researchers have proposed a time-frequency self-contrastive learning method that generates contrast pairs in the time and frequency domains to enhance feature extraction and use cross-time-frequency interaction strategies to guide feature fusion. Other researchers have proposed a time-frequency contrastive learning method that relies only on positive sample pairs, which significantly improves the effect of bearing fault diagnosis when there is insufficient labeled data. Although self-supervised learning has made some progress in fault diagnosis, it still faces the following problems: (1) Existing methods overly rely on single feature extraction and lack a deep understanding of the multi-scale dependence of fault signals. In addition, most of these methods are based on traditional feature extraction networks (such as 1D-CNN), which makes it difficult to achieve comprehensive feature fusion. (2) Although time-frequency feature fusion has been proven to be effective in fault diagnosis, current self-supervised learning methods fail to fully exploit the complementary information between time and frequency domain features. The fused features are not trained through contrastive learning, resulting in the potential diagnostic basis not being fully utilized, which limits the generalization ability and diagnostic performance of the model. (3) Existing contrastive learning architectures mainly focus on instance-level or feature-level contrast, ignoring the core goal of generating more compact cluster distributions. This deficiency limits the feature expression ability and diagnostic performance of the model. Summary of the invention
[0004] In view of the above challenges, the present invention provides a rotating machinery fault diagnosis method based on a self-supervised time-frequency contrast fusion network. Through a multi-level contrast learning strategy, it effectively integrates time domain and frequency domain features, deeply mines their complementary information, and solves the problem of limited model representation ability caused by insufficient feature interaction and fusion. The present invention significantly improves the fault diagnosis effect.
[0005] The technical solution of the present invention is: a rotating machinery fault diagnosis method based on a self-supervised time-frequency contrast fusion network, and the specific steps of the method are as follows:
[0006] Step 1, collect rotating machinery fault data set;
[0007] Step 2: Clean and normalize the collected data set, and divide it into independent working condition subsets according to working conditions; then, divide the data set into an unlabeled training set, a labeled fine-tuning set, and a test set in a ratio of 8:1:1;
[0008] Step 3: Perform fast Fourier transform on the divided data set to convert the time signal into a frequency signal. Then, in the unlabeled training set, the time signal and the frequency signal are processed by the data enhancement library to obtain the time-frequency sample pair and its enhanced version.
[0009] Step 4, build a self-supervised pre-training model; build a self-supervised pre-training model through feature extraction network and multi-level contrastive learning strategy; among them, construct a pre-training target loss function through multi-level contrastive learning strategy, including instance level, cluster level and fusion level contrastive loss function; enhance the similarity of features of the same instance, expand the difference of features of different instances, and guide similar samples to approach their respective cluster centers;
[0010] Step 5: Use the unlabeled training set processed in Step 3 as input, and then perform iterative training according to the pre-training target loss function designed in Step 4 until the loss function converges or reaches the maximum number of iterations, and finally save the pre-training model;
[0011] Step 6: Fine-tune the pre-trained model, including loading the pre-trained model parameters, using the cross entropy loss function, and obtaining the fine-tuning target loss function through label guidance.
[0012] Step 7: Use the divided labeled fine-tuning set to train the fine-tuned pre-trained model, optimize the objective function through the labels, iterate until convergence or the maximum number of iterations is reached, and save the trained fine-tuned pre-trained model;
[0013] Step 8. Load the parameters of the trained and fine-tuned pre-trained model, input the divided test set samples, select the category corresponding to the maximum value according to the category probability output by the model to obtain the predicted label, and then obtain the diagnosis result.
[0014] Furthermore, in Step 1, the rotating machinery fault data set includes a plurality of sample pairs, and each pair of samples includes a vibration signal and its corresponding state label.
[0015] Furthermore, in the Step 2, the normalization process adopts Z-score normalization, and the truncation method is used to ensure the balance of the number of samples of each category; secondly, different operating condition subsets correspond to different operating conditions, and the fault information contained in the vibration signal is different, but the number of categories and labels remain consistent; the data set division steps are: the first step is to randomly select 80% of the sample pairs from each operating condition subset and remove the labels to form an unlabeled training set; the second step is to randomly select 10% from the remaining 20% labeled data as a labeled fine-tuning set; the third step is to use the remaining 10% as a labeled test set.
[0016] Furthermore, in Step 3, the process of applying the data enhancement library to the time signal and the frequency signal respectively is as follows:
[0017] For a given one-dimensional vibration signal V, let X t ≡V, then represents the i-th time series sample in the unlabeled training set, and n represents the number of samples in the mini-batch; the temporal signal enhancement process is as follows:
[0018]
[0019] in, represents the enhanced time domain signal, Indicates the enhanced A t represents the time domain enhancement library, T j represents the combined operation of temporal data enhancement, T a , T b and T c They represent the data enhancement operations of jitter, scaling and permutation, respectively, and α, β, and γ are the enhancement parameters;
[0020] Using Fast Fourier Transform The time domain signal Convert to frequency domain signal The frequency domain signal enhancement process is as follows:
[0021]
[0022] Among them, A f represents the frequency domain enhancement library, F j represents the frequency domain data enhancement combination operation, which consists of adding frequency components (F a ) and frequency component removal operation (F b ), ω, θ and ∈ are frequency domain enhancement parameters, and the enhanced frequency domain signal is expressed as Indicates the enhanced F a and F b They refer to the data enhancement operations of adding and removing frequency components respectively. The enhanced frequency domain signal is expressed as
[0023] Furthermore, the Step 4 includes: extracting time-frequency features from the time and frequency signals using a time encoder and a frequency encoder with the same structure but independent parameters;
[0024] For the original timing signal and the enhanced time domain signal The encoding is performed using a temporal encoder, where D is the signal dimension; the encoding process is defined as follows:
[0025]
[0026] in, represents the time-domain coded representation of the temporal encoder output, l is the feature dimension, is the i-th temporal encoding representation in the mini-batch, represents the temporal enhancement coded representation of the temporal encoder output, is the i-th temporal enhancement encoding representation in the mini-batch; Time_Enocer(,) represents the temporal encoder, θ t are the parameters of the time encoder;
[0027] The frequency signal is processed using the transformer frequency encoder to obtain frequency embedding representation and frequency enhancement coding representation. The process is expressed as follows:
[0028]
[0029] in, represents the frequency domain coded representation of the frequency encoder output, is the ith frequency-encoded representation in the mini-batch; represents the frequency domain enhanced coded representation of the frequency encoder output, is the i-th frequency enhanced encoding representation in the mini-batch; Frequency_Enocer(,) represents the frequency encoder, θ f is the parameter of the frequency encoder.
[0030] Furthermore, in the Step 4, a pre-training target loss function is constructed through a multi-level contrastive learning strategy, including: instance-level, cluster-level and fusion-level contrastive loss functions are defined as follows:
[0031] The time-frequency feature encoding result is expressed as After linear projection to the time-frequency space, the corresponding instance-level embedding representation is obtained Subsequently, these projected instance-level embeddings are represented To operate, represents the original time-frequency embedding view, where represents the original embedding representation of the i-th time-frequency view in the mini-batch, let represents the enhanced time-frequency embedding view, where represents the enhanced embedding representation of the i-th time-frequency view in the mini-batch; in a training dataset of size n, each original embedding view Its corresponding enhanced embedded view Positive pairing and negative pairing with other samples; original embedding view Instance-level contrast loss It is expressed as:
[0032]
[0033] Among them, τ I represents the instance-level temperature parameter, s(·) represents the cosine similarity between samples, considering the original embedding view and enhanced embedded views The symmetry between , the instance-level contrast loss is further expressed as:
[0034]
[0035] in, Represents the original embedded view The instance-level contrast loss is Represents an enhanced embedded view Instance-level contrast loss;
[0036] By constructing two pseudo classifiers with different parameters for the time domain and frequency domain encoding representations, we can obtain the time-frequency pseudo labels. Respectively represent the original view category assignment The i-th column and enhanced view category assignment The i-th column of , n is the number of samples, k is the number of clusters; the original view category assignment Cluster-level contrastive loss It is expressed as:
[0037]
[0038] Among them, τ c is the clustering level temperature parameter. In the clustering process, in order to prevent degenerate solution, the cross entropy constraint is introduced:
[0039]
[0040] in Respectively represent the original view category assignment and Enhanced View Category Assignment Considering the symmetry between the original view and the enhanced view, the cluster-level contrast loss is expressed as:
[0041]
[0042] in, Indicates the original view category assignment The cluster-level contrast loss is Indicates enhanced view category assignment The cluster-level contrast loss is is the cross entropy loss function;
[0043] Enhanced embedding representation in time domain Frequency domain enhanced embedding representation Fusion is performed using time domain perception and frequency domain perception respectively; fusion features dominated by time domain perception The generated representation is as follows:
[0044]
[0045] Among them, Attn is the attention network, which is constructed by the feed-forward neural network FFN, [,] represents the horizontal concatenation of matrices or vector sets along the rows, is the common enhanced embedding representation after concatenation, is the fusion weight vector calculated by Attn, λ t ,λ f They represent the probability distribution of the time domain and frequency domain fusion weights calculated by softmax, τ is the scaling factor, It is the time-frequency perception weight dominated by time domain perception. Representation of time-domain enhanced embedding representation and frequency domain enhanced embedding representation A set of, d is the vector dimension;
[0046] Fusion features dominated by frequency domain perception The generated representation is as follows:
[0047]
[0048] in, The fusion weight vector is calculated by weight sharing Attn, λ t ,λ f They represent the probability distribution of the time domain and frequency domain fusion weights calculated by softmax, respectively. It is the time-frequency perception weight dominated by frequency domain perception;
[0049] Align features using momentum contrast method;
[0050] Denote the parameters of the attention fusion module as θ fu , the parameters of the momentum model are denoted by θ mom , the update process is expressed as:
[0051]
[0052] Where m∈[0,1] is the momentum coefficient;
[0053] Through the momentum model M mom Generate time- and frequency-dominated fusion features and calculate time-frequency embedding features The fused features are used as queries and the features generated by the momentum model are used as keys, which are expressed as follows:
[0054]
[0055] Among them, the fusion features dominated by time domain As a query Frequency-Domain-Dominated Fusion Features As another query q f , key k t and k f Respectively represent the momentum model M mom The generated time domain and frequency domain momentum fusion features;
[0056] {q t ,q f ,k t ,k f} as input to calculate the time-frequency contrast loss; momentum contrast loss It is expressed as:
[0057]
[0058] in, represents the i-th time domain query feature, represents the jth frequency domain momentum feature, and each time domain query feature The corresponding frequency domain momentum characteristics Positive pairing, and negative pairing with other samples, B represents the mini-batch size, τ m is the L2-normalized temperature hyperparameter, L t and L f denote the contrast loss of the time domain and frequency domain branches respectively;
[0059] Where B represents the mini-batch size, τ m is the temperature hyperparameter of the L2 normalization.
[0060] Construct a public pseudo classifier CPC to represent the time-frequency fusion feature Input a public pseudo classifier to generate a public pseudo label Then, the pseudo labels generated by combining the enhanced time-frequency representation Together we construct the confidence target distribution Q:
[0061]
[0062] The confident target assignment with the highest probability is chosen as the target assignment; highly confident assignments are boosted and instances at cluster boundaries are further blurred by:
[0063]
[0064] As an auxiliary target distribution to guide the clustering results, q ij represents the i-th row and j-th column of the confidence target distribution Q, p ij As an auxiliary target distribution to guide the clustering results, the optimization objective is expressed as:
[0065]
[0066] in, Public pseudo labels The i-th row and j-th column of , KL(||) represents the Kullback-Leibler divergence, is the objective loss function optimized with high confidence;
[0067] The domain loss function is calculated based on the constructed target loss function, and the total loss is obtained by joint training through the domain loss function and the fusion level comparison guidance;
[0068] Combining the above loss functions, by expressing the time domain and frequency domain representation Using Formula 6 and Formula 9, we get the following two domain loss functions:
[0069]
[0070] in, Represent the intra-domain objective loss functions in the time domain branch and the frequency domain branch respectively;
[0071] Therefore, the contrast loss in the time-frequency domain It is expressed as:
[0072]
[0073] Contrast loss in time-frequency domain Jointly trained with the proposed fusion-level contrastive guidance, which includes momentum contrastive loss and high confidence guide Total pre-training loss It is expressed as follows:
[0074]
[0075] Where λ1 and λ2 are loss weight coefficients.
[0076] Furthermore, the Step 5 includes:
[0077] The unlabeled training set is processed by a random selection method and loaded into the data loader in batches. Each batch contains four inputs: time-frequency signal pairs and enhanced time-frequency signal pairs. The sample pairs are sent to the encoder network for feature extraction, and then the parameters in the encoder network are iteratively updated according to the pre-training objective function until the pre-training objective loss function converges or the maximum number of iterations is reached. Finally, the parameters of the pre-training model are saved.
[0078] Furthermore, in Step 6, the objective loss function of the fine-tuning model is It is composed of the cross entropy loss function and the target loss function as follows:
[0079]
[0080] Where N is the total number of samples, C is the total number of categories, and y i,c is the true label of the i-th sample for category c, is the predicted probability of the i-th sample for category c.
[0081] Furthermore, the Step 7 includes:
[0082] The labeled fine-tuning set and its true label input are input into the fine-tuned pre-trained model, and compared with the true label according to the fine-tuning objective loss function. The parameters of the fine-tuning model encoder network and linear classifier are iteratively updated until convergence or the maximum number of iterations is reached. After fine-tuning is completed, the parameters of the fine-tuned pre-trained model are saved.
[0083] Furthermore, in Step 8, the fine-tuning model is a fine-tuning model with frozen parameters after training, and the final fault diagnosis result is an average value after the model has been trained for multiple times.
[0084] In the Step 4, the feature extraction network uses an encoder adapted from Transformer to perform feature encoding, and customizes a time encoder and a frequency encoder with the same structure but independent parameters for the timing and frequency signals respectively.
[0085] In the Step 4, the multi-level contrastive learning strategy referred to is essentially an objective function consisting of contrastive losses at the instance level, clustering level, and feature fusion level. By optimizing the objective function, the model can learn general time-frequency features.
[0086] In the above Step 5, in order to ensure the generalization ability of the model, the data should be processed and divided into batches by a random selection method each time during training. After multiple rounds of training, the final pre-trained model is saved.
[0087] In Step 6, the fine-tuning model has different overall structures and training tasks from the pre-training model, but the model parameter structures correspond. The fine-tuning model optimizes the classifier through labels to adapt to downstream fault diagnosis tasks.
[0088] In Step 7, the target loss functions used by the fine-tuning model and the pre-training model are different. The fine-tuning model uses the cross entropy loss function to optimize the model parameters by comparing the predicted labels with the true labels; while the pre-training model uses multiple contrast loss functions to optimize the encoding network and learn better feature representations.
[0089] Based on the existing self-supervised contrastive learning framework, the method of the present invention proposes a novel multi-level contrastive learning strategy, which significantly enhances the robustness of the model to multi-scale features and effectively copes with cross-domain differences caused by working condition changes and signal noise. The present invention introduces a fusion-level contrast mechanism, cleverly combines instance-level and cluster-level contrast constraints, further optimizes the feature representation of samples, and ensures that the model can effectively capture key information at different scales. By fusing the complementary features of the time domain and the frequency domain, the model can more comprehensively characterize the fault signal and improve the ability to identify complex fault modes. In addition, in order to further improve the diagnostic ability of the model, the present invention introduces a high-confidence guidance mechanism, which uses time-frequency complementary information to strengthen the clustering allocation of samples, making the feature distribution of similar faults more compact. This mechanism improves the separability of fault types and significantly improves the diagnostic performance. The present invention significantly improves the accuracy and efficiency of rotating machinery fault diagnosis, especially when facing complex working conditions and diversified signals, its diagnostic effect far exceeds the existing methods, and has broad application prospects and practical value.
[0090] The beneficial effects of the present invention are:
[0091] 1. This method uses a large amount of unlabeled data for feature learning, combines it with a small amount of labeled data for fine-tuning, and adopts a multi-level contrastive learning strategy, including instance-level, cluster-level, and fusion-level contrastive learning. It improves model performance from three levels: sample, category, and feature fusion, and constructs a multi-scale contrastive learning framework, which significantly improves the fault diagnosis effect.
[0092] 2. The present invention proposes a time-frequency momentum contrast fusion method to deeply mine the time-frequency complementary information, and fully utilize the time-frequency mutual information through a high-confidence guidance mechanism, enhance the complementarity of time domain and frequency domain features, and comprehensively consider the multi-scale dynamic dependency of fault signals, thereby improving the clustering ability of the model, fault diagnosis effect, diagnostic performance and generalization ability.
[0093] 3. Based on the existing self-supervised contrastive learning framework, this method proposes a novel multi-level contrastive learning strategy, which significantly enhances the robustness of the model to multi-scale features and effectively copes with cross-domain differences caused by operating condition changes and signal noise, thereby improving the accuracy and generalization ability of fault diagnosis.
[0094] 4. The present invention introduces a fusion-level comparison mechanism, which cleverly combines instance-level and cluster-level comparison constraints to further optimize the feature representation of samples and ensure that the model can effectively capture key information at different scales.
[0095] 5. By integrating the complementary features of the time domain and the frequency domain, the model of the present invention can more comprehensively characterize the fault signal and improve the ability to identify complex fault modes.
[0096] 6. The present invention introduces a high-confidence guidance mechanism and uses time-frequency complementary information to strengthen the clustering distribution of samples, making the characteristic distribution of similar faults more compact, improving the separability of fault types, and significantly improving the accuracy of diagnosis.
[0097] 7. The present invention significantly improves the accuracy and efficiency of rotating machinery fault diagnosis, especially when faced with complex working conditions and diversified signals. Its diagnostic effect far exceeds that of existing methods and has broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 is a flow chart of a method according to an embodiment of the present invention;
[0099] Figure 2 A flowchart of iterative updating of the method of the present invention;
[0100] Figure 3 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0101] Example 1: Figure 1-Figure 3 As shown, a rotating machinery fault diagnosis method based on a self-supervised time-frequency contrast fusion network, the specific steps of the method are as follows:
[0102] Step 1, collect a rotating machinery fault data set; the rotating machinery fault data set includes multiple sample pairs, each pair of samples contains a vibration signal and its corresponding state label.
[0103] Step 2. Clean and normalize the collected data set, and divide the data set; the normalization process uses Z-score standardization, and the truncation method is used to ensure the balance of the number of samples in each category; the steps of data set division are as follows: the first step is to randomly select 80% of the sample pairs from each working condition subset and remove the labels to form an unlabeled training set; the second step is to randomly select 10% from the remaining 20% labeled data as a labeled fine-tuning set; the third step is to use the remaining 10% as a labeled test set.
[0104] Step 3: Perform fast Fourier transform on the divided data set to convert the time signal into a frequency signal. Then, in the unlabeled training set, the time signal and the frequency signal are processed by the data enhancement library to obtain the time-frequency sample pair and its enhanced version.
[0105] In the Step 3, the original vibration signal of each data set is regarded as a time series signal, converted into a frequency signal through fast Fourier transform, and combined with the time series signal to form a time-frequency sample pair to form a new data set; wherein, only the unlabeled training set data is enhanced, while the fine-tuning set and the test set are not enhanced;
[0106] Furthermore, in Step 3, the process of applying the data enhancement library to the time signal and the frequency signal respectively is as follows:
[0107] For a given one-dimensional vibration signal V, let X t ≡V, then represents the i-th time series sample in the unlabeled training set, and n represents the number of samples in the mini-batch; the temporal signal enhancement process is as follows:
[0108]
[0109] in, represents the enhanced time domain signal, Indicates the enhanced A t represents the time domain enhancement library, T j represents the combined operation of temporal data enhancement, T a , T band T c They represent the data enhancement operations of jitter, scaling and permutation, respectively, and α, β, and γ are the enhancement parameters;
[0110] Using Fast Fourier Transform The time domain signal Convert to frequency domain signal The frequency domain signal enhancement process is as follows:
[0111]
[0112] Among them, A f represents the frequency domain enhancement library, F j represents the frequency domain data enhancement combination operation, which consists of adding frequency components (F a ) and frequency component removal operation (F b ), ω, θ and ∈ are frequency domain enhancement parameters, and the enhanced frequency domain signal is expressed as Indicates the enhanced F a and F b They refer to the data enhancement operations of adding and removing frequency components respectively. The enhanced frequency domain signal is expressed as
[0113] Step 4: Build a self-supervised pre-training model. Use a feature extraction network and a multi-level contrastive learning strategy to build a self-supervised pre-training model. The pre-training target loss function is constructed through a multi-level contrastive learning strategy, including instance-level, cluster-level, and fusion-level contrastive loss functions.
[0114] In order to fully extract the time-frequency features, the present invention designs a Transformer-based encoder network, which focuses on mining the multi-scale characteristics of one-dimensional vibration signals and realizes the efficient fusion of time-frequency features.
[0115] Furthermore, the Step 4 includes: extracting time-frequency features from the time and frequency signals using a time encoder and a frequency encoder with the same structure but independent parameters;
[0116] For the original timing signal and the enhanced time domain signal The encoding is performed using a temporal encoder, where D is the signal dimension; the encoding process is defined as follows:
[0117]
[0118] in, represents the time-domain coded representation of the temporal encoder output, l is the feature dimension, is the i-th temporal encoding representation in the mini-batch, represents the temporal enhancement coded representation of the temporal encoder output, is the i-th temporal enhancement encoding representation in the mini-batch; Time_Enocer(,) represents the temporal encoder, θ t are the parameters of the time encoder;
[0119] The frequency signal is processed using the transformer frequency encoder to obtain frequency embedding representation and frequency enhancement coding representation. The process is expressed as follows:
[0120]
[0121] in, represents the frequency domain coded representation of the frequency encoder output, is the ith frequency-encoded representation in the mini-batch; represents the frequency domain enhanced coded representation of the frequency encoder output, is the i-th frequency enhanced encoding representation in the mini-batch; Frequency_Enocer(,) represents the frequency encoder, θ f is the parameter of the frequency encoder.
[0122] Furthermore, in the Step 4, a pre-training target loss function is constructed through a multi-level contrastive learning strategy, including: instance-level, cluster-level and fusion-level contrastive loss functions are defined as follows:
[0123] In instance-level contrastive learning, in order to enhance the robustness of the model, the instance-level loss does not directly act on the encoding result. Instead, the time-frequency feature encoding result is represented as After linear projection to the time-frequency space, the corresponding instance-level embedding representation is obtained Subsequently, these projected instance-level embeddings are represented To operate, represents the original time-frequency embedding view, where represents the original embedding representation of the i-th time-frequency view in the mini-batch, let represents the enhanced time-frequency embedding view, where represents the enhanced embedding representation of the i-th time-frequency view in the mini-batch; in a training dataset of size n, each original embedding view Its corresponding enhanced embedded view Positive pairing and negative pairing with other samples; original embedding view Instance-level contrast loss It is expressed as:
[0124]
[0125] Among them, τ I represents the instance-level temperature parameter, s(·) represents the cosine similarity between samples, considering the original embedding view and enhanced embedded views The symmetry between , the instance-level contrast loss is further expressed as:
[0126]
[0127] in, Represents the original embedded view The instance-level contrast loss is Represents an enhanced embedded view Instance-level contrast loss;
[0128] Second, cluster-level contrastive learning can encourage similar samples to cluster and help the model learn more discriminative representations at the class level. Cluster-level constraints are also not directly applied to the output of the encoder. Instead, it is the result of applying clustering Specifically, the present invention constructs two pseudo classifiers with different parameters for the time domain and frequency domain coding representations respectively to obtain time-frequency pseudo labels; Respectively represent the original view category assignment The i-th column and enhanced view category assignment The i-th column of , n is the number of samples, k is the number of clusters; the original view category assignment Cluster-level contrastive loss It is expressed as:
[0129]
[0130] Among them, τ c is the clustering level temperature parameter. In the clustering process, in order to prevent degenerate solution, the cross entropy constraint is introduced:
[0131]
[0132] in Respectively represent the original view category assignment and Enhanced View Category Assignment Considering the symmetry between the original view and the enhanced view, the cluster-level contrast loss is expressed as:
[0133]
[0134] in, Indicates the original view category assignment The cluster-level contrast loss is Indicates enhanced view category assignment The cluster-level contrast loss is is the cross entropy loss function;
[0135] In order to make full use of the complementary information of different views, this paper proposes a fine-grained instance-level attention method for automatically perceiving view fusion weights. Frequency domain enhanced embedding representation Fusion is performed using time domain perception and frequency domain perception respectively; fusion features dominated by time domain perception The generated representation is as follows:
[0136]
[0137] Among them, Attn is the attention network, which is constructed by the feed-forward neural network FFN, [,] represents the horizontal concatenation of matrices or vector sets along the rows, is the common enhanced embedding representation after concatenation, is the fusion weight vector calculated by Attn, λ t ,λ f They represent the probability distribution of the time domain and frequency domain fusion weights calculated by softmax, τ is the scaling factor, It is the time-frequency perception weight dominated by time domain perception. Representation of time-domain enhanced embedding representation and frequency domain enhanced embedding representation A set of, d is the vector dimension;
[0138] Fusion features dominated by frequency domain perception The generated representation is as follows:
[0139]
[0140] in, The fusion weight vector is calculated by weight sharing Attn, λ t ,λ f They represent the probability distribution of the time domain and frequency domain fusion weights calculated by softmax, respectively. It is the time-frequency perception weight dominated by frequency domain perception;
[0141] Align features using momentum contrast method;
[0142] In order to effectively align features, this paper introduces a momentum comparison method. The momentum model is a copy of the attention fusion network, which adopts a momentum update strategy instead of traditional back propagation and gradient update. The parameters of the attention fusion module are represented as θ fu , the parameters of the momentum model are denoted by θmom , the update process is expressed as:
[0143]
[0144] Where m∈[0,1] is the momentum coefficient;
[0145] Similar to the attention fusion module method, the momentum model M mom Generate time- and frequency-dominated fusion features and calculate time-frequency embedding features The fused features are used as queries and the features generated by the momentum model are used as keys, which are expressed as follows:
[0146]
[0147] Among them, the fusion features dominated by time domain As the query q t , frequency domain dominated fusion features As another query q f , key k t and k f Respectively represent the momentum model M mom The generated time domain and frequency domain momentum fusion features;
[0148] {q t ,q f ,k t ,k f} as input to calculate the time-frequency contrast loss; momentum contrast loss It is expressed as:
[0149]
[0150] in, represents the i-th time domain query feature, represents the jth frequency domain momentum feature, and each time domain query feature The corresponding frequency domain momentum characteristics Positive pairing, and negative pairing with other samples, B represents the mini-batch size, τ m is the L2-normalized temperature hyperparameter, L t and L f denote the contrast loss of the time domain and frequency domain branches respectively;
[0151] Where B represents the mini-batch size, τ m is the temperature hyperparameter of the L2 normalization.
[0152] The present invention selects the fusion feature dominated by the time domain As the final feature representation, it is used as the input of the subsequent high confidence guidance module. In order to make full use of the complementary information of time and frequency, a target guidance mechanism emphasizing high confidence instances is designed and introduced. A public pseudo classifier CPC is constructed to represent the time-frequency fusion feature. Input a public pseudo classifier to generate a public pseudo label Then, the pseudo labels generated by combining the enhanced time-frequency representation Together we construct the confidence target distribution Q:
[0153]
[0154] This ensures that high-confidence instances are emphasized. The confident target assignment with the highest probability is selected as the target assignment; highly confident assignments are enhanced and instances at cluster boundaries are further blurred by the following operations:
[0155]
[0156] As an auxiliary target distribution to guide the clustering results, q ij represents the i-th row and j-th column of the confidence target distribution Q, p ij As an auxiliary target distribution to guide the clustering results, the optimization objective is expressed as:
[0157]
[0158] in, Public pseudo labels The i-th row and j-th column of , KL(||) represents the Kullback-Leibler divergence, is the objective loss function optimized with high confidence;
[0159] The domain loss function is calculated based on the constructed target loss function, and the total loss is obtained by joint training through the domain loss function and the fusion level comparison guidance;
[0160] Combining the above loss functions, by expressing the time domain and frequency domain representation Using Formula 6 and Formula 9, we get the following two domain loss functions:
[0161]
[0162] in, Represent the intra-domain objective loss functions in the time domain branch and the frequency domain branch respectively;
[0163] Therefore, the contrast loss in the time-frequency domain It is expressed as:
[0164]
[0165] Contrast loss in time-frequency domain Jointly trained with the proposed fusion-level contrastive guidance, which includes momentum contrastive loss and high confidence guide Total pre-training loss It is expressed as follows:
[0166]
[0167] Where λ1 and λ2 are loss weight coefficients.
[0168] Step 5: Use the unlabeled training set processed in Step 3 as input, and then perform iterative training according to the pre-training target loss function designed in Step 4 until the loss function converges or reaches the maximum number of iterations, and finally save the pre-training model;
[0169] Furthermore, the Step 5 includes:
[0170] The unlabeled training set is processed by a random selection method and loaded into the data loader in batches. Each batch contains four inputs: time-frequency signal pairs and enhanced time-frequency signal pairs. The sample pairs are sent to the encoder network for feature extraction, and then the parameters in the encoder network are iteratively updated according to the pre-training objective function until the pre-training objective loss function converges or the maximum number of iterations is reached. Finally, the parameters of the pre-training model are saved.
[0171] Step 6: Fine-tune the pre-trained model, including loading the pre-trained model parameters, using the cross entropy loss function, and obtaining the fine-tuning target loss function through label guidance.
[0172] The fine-tuning model structure is similar to the pre-training model, and is mainly composed of an encoder network and a linear classifier. The encoder network structure is the same as the pre-training model, and the linear classifier is the common pseudo classifier in the above pre-training;
[0173] Furthermore, in Step 6, the objective loss function of the fine-tuning model is It is composed of the cross entropy loss function and the target loss function as follows:
[0174]
[0175] Where N is the total number of samples, C is the total number of categories, and y i,c is the true label of the i-th sample for category c, is the predicted probability of the i-th sample for category c.
[0176] Step 7. Perform supervised fine-tuning learning: Use the divided labeled fine-tuning set to train the fine-tuned pre-trained model, optimize the objective function through labels, iterate until convergence or the maximum number of iterations is reached, and save the trained fine-tuned pre-trained model;
[0177] Furthermore, the Step 7 includes:
[0178] The labeled fine-tuning set and its true label input are input into the fine-tuned pre-trained model, and compared with the true label according to the fine-tuning objective loss function. The parameters of the fine-tuning model encoder network and linear classifier are iteratively updated until convergence or the maximum number of iterations is reached. After fine-tuning is completed, the parameters of the fine-tuned pre-trained model are saved.
[0179] Step 8. Perform rotating machinery fault diagnosis: load the parameters of the trained and fine-tuned pre-trained model, input the divided test set samples, select the category corresponding to the maximum value according to the category probability output by the model to obtain the predicted label, and then obtain the diagnosis result.
[0180] Furthermore, in Step 8, the fine-tuning model is a fine-tuning model with frozen parameters after training, and the final fault diagnosis result is an average value after the model has been trained for multiple times.
[0181] The present invention proposes a self-supervised intelligent fault diagnosis method (TF-CFN) based on time-frequency contrast fusion. This method uses a specially designed Transformer encoder to extract effective features from complex time and frequency domain signals and promote deep feature fusion. Based on the existing self-supervised contrast learning framework, a multi-level contrast learning strategy is innovatively introduced. This strategy not only enhances the robustness of the model to multi-scale features, but also effectively deals with cross-domain differences caused by working condition changes or noise, thereby improving the accuracy and generalization ability of fault diagnosis. In order to fully explore the complementary information of time domain and frequency domain features, the framework designs a novel fusion-level contrast mechanism. Through the time-frequency momentum contrast fusion method, the time-frequency mutual information is focused on, and the multi-scale dynamic dependency of the fault signal is fully considered, thereby improving the diagnostic performance and generalization ability of the model. In addition, the framework introduces a high-confidence guidance mechanism, which significantly improves the accuracy of fault diagnosis by utilizing time-frequency complementary information to achieve more compact clustering assignments.
[0182] This algorithm is essentially a classification task. To explain the steps of the algorithm more intuitively, Figure 2 Describe the iterative update process of the entire framework:
[0183] (1) Start training: Start training after the model is built.
[0184] (2) Data input: Input the unlabeled dataset into the pre-trained model. The unlabeled data here includes four sample pairs: original sample pairs in the time domain and frequency domain, and sample pairs in the time domain and frequency domain after data enhancement.
[0185] (3) Pre-training: Iteratively update the parameters of the pre-training model through the designed objective function. The objective function here is: in is the intra-domain contrast loss in the time domain and frequency domain, respectively, given by instance-level formula (6) And the cluster-level contrast loss formula (9) The main updated parameters of the model are the time encoder parameters θ t and frequency encoder parameter θ f , and common pseudo-classifier parameters.
[0186] (4) Optimize the pre-training objective function: Use the optimizer to optimize the model parameters until the target loss function converges or the maximum number of iterations is reached. After the iteration is completed, save the pre-training model.
[0187] (5) Supervised fine-tuning: Load the pre-trained model parameters obtained in the previous step, and optimize the fine-tuned model through label guidance to adapt to the downstream task. The objective function here is the cross entropy loss function shown in formula (24).
[0188] (6) Optimize the fine-tuning objective function: Similar to pre-training, iteratively update this fine-tuning objective loss function, update the fine-tuning model parameters, and save the fine-tuning model.
[0189] (7) Fault diagnosis: Load the fine-tuned model and input the test samples to obtain the diagnosis results.
[0190] In order to illustrate the effect of the present invention, the technical solution of the present invention is further described below through specific embodiments:
[0191] 1. Simulation conditions;
[0192] The present invention uses pycharm software to perform experimental simulation. The experiment was conducted on the fault diagnosis dataset CWRU (containing one-dimensional vibration signals and labels), and the experiment included two classification tasks: (1) fault diagnosis, and (2) cross-operating fault diagnosis. The parameters in the experiment were set as follows: the learning rate was 3e-4, the batch size was 64, the maximum number of iterations was 100, the optimizer used the AdamW optimizer, the momentum coefficient m was 0.9999, and the temperature parameter τ of the instance-level and cluster-level contrast loss was I , τ C The temperature hyperparameter τ is 0.5 and 1.0 respectively, L-2 normalized m is 0.2.
[0193] 2. Simulation content;
[0194] The method proposed in the present invention is based on a rotating machinery fault diagnosis method of a self-supervised time-frequency contrast fusion network (TF-CFN) and is compared with the existing self-supervised contrast learning rotating machinery fault diagnosis method. The specific comparison methods are as follows: (1) Fault Diagnosis Method Based on Momentum Contrast (MoCo), (2) Fault diagnosis method based on SimCLR (SimCLR), (3) Fault diagnosis method based on Temporal and Contextual Contrasting (TS-TCC), (4) Fault diagnosis method based on time-frequency consistency (TF-C), (5) Fault diagnosis method based on the SSPCL contrastive learning framework (SSPCL), (6) Fault diagnosis method based on the time-frequency contrastive learning framework (CLFT); (7) Fault diagnosis methodbased on the time-frequency Siamese network network, TS-TFSIAM); (8) Fault diagnosis method based on time-frequency alignment and interaction (TFAI).
[0195] 3. Simulation results;
[0196] The simulation experiment gives the experimental results of the comparison method and the method proposed in the present invention under the data set CWRU. In the present invention, the CWRU data set is divided into four subsets with different working conditions: CWRU_0, CWRU_1, CWRU_2, and CWRU_3. Each data subset has 10 categories, including health, outer ring fault, inner ring fault, and ball fault status. This embodiment performs traditional fault diagnosis tasks on the four working condition subsets, and then cross-working condition fault diagnosis is performed on the four working condition subsets. Note that the processing method of the remaining data sets is the same as the setting of the present invention.
[0197] In this simulation experiment, four widely used indicators are used to measure the performance of the TF-CFN method proposed in this invention and other comparative methods, namely, accuracy (ACC), precision (Precision), recall (Recall) and F1 score (F1). Given a test sample set, the above four indicators can be defined as follows:
[0198] (1) Accuracy (ACC) is the ratio of the number of samples correctly predicted by the model to the total number of samples, which measures the overall prediction effect of the model. The formula is:
[0199]
[0200] Among them (the same below): TP (True Positive) represents the number of samples that are actually positive and correctly predicted as positive; TN (True Negative) represents the number of samples that are actually negative and correctly predicted as negative; FP (False Positive) represents the number of samples that are actually negative but incorrectly predicted as positive; FN (False Negative) represents the number of samples that are actually positive but incorrectly predicted as negative.
[0201] (2) Precision is the ratio of samples predicted to be positive to samples that are actually positive, which measures the accuracy of the model in predicting positive results.
[0202]
[0203] (3) Recall is the proportion of actual positive samples that are correctly predicted as positive, which measures the model's ability to identify positive samples.
[0204]
[0205] (4) F1 score (F1) is the harmonic average of precision and recall, taking into account the balance between the two. The F1 score will be higher when both precision and recall are high, which is suitable for class imbalance problems.
[0206]
[0207] Through these four indicators, the performance of the classification model can be comprehensively evaluated, including the accuracy of the overall prediction, the precision of the positive class prediction, and the ability to identify the positive class.
[0208] Tables 1 and 2 show the corresponding performance indicators of the TF-CFN method proposed in the present invention and other comparative methods in fault diagnosis tasks.
[0209] Table 1 Performance evaluation of all methods on fault diagnosis tasks on CWRU_0 and CWRU_1 datasets
[0210]
[0211]
[0212] Table 2 Performance evaluation of all methods on fault diagnosis tasks on CWRU_2 and CWRU_3 datasets
[0213]
[0214] Table 3 shows the corresponding performance indicators of the TF-CFN method proposed in this invention and other comparative methods in the cross-operating fault diagnosis task. For example, pre-training on the CWRU_0 subset and fine-tuning test on the CWRU_1 subset are performed. The purpose of this scenario is to evaluate the domain adaptability and generalization ability of the proposed method.
[0215] Table 3 Performance evaluation of all methods on the cross-condition fault diagnosis task on the CWRU dataset
[0216]
[0217]
[0218] It can be seen from Table 1, Table 2 and Table 3 that all index values of the TF-CFN method proposed in the present invention in the two fault diagnosis tasks under the CWRU dataset are higher than those of other comparison methods, which further proves the superiority of the TF-CFN method proposed in the present invention in fault diagnosis and its strong generalization ability.
[0219] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A rotating machinery fault diagnosis method based on a self-supervised time-frequency contrast fusion network, characterized in that: The specific steps of the method are as follows: Step 1, collect rotating machinery fault data set; Step 2: Clean and normalize the collected data set and divide the data set; Step 3: Perform fast Fourier transform on the divided data set to convert the time signal into a frequency signal; Then, in the unlabeled training set, the data augmentation library is applied to the time signal and the frequency signal to obtain the time-frequency sample pairs and their enhanced versions; Step 4: Build a self-supervised pre-training model. Use a feature extraction network and a multi-level contrastive learning strategy to build a self-supervised pre-training model. The pre-training target loss function is constructed through a multi-level contrastive learning strategy, including instance-level, cluster-level, and fusion-level contrastive loss functions. Step 5: Use the unlabeled training set processed in Step 3 as input, and then perform iterative training according to the pre-training target loss function designed in Step 4 until the loss function converges or reaches the maximum number of iterations, and finally save the pre-training model; Step 6: Fine-tune the pre-trained model, including loading the pre-trained model parameters, using the cross entropy loss function, and obtaining the fine-tuning target loss function through label guidance. Step 7: Use the divided labeled fine-tuning set to train the fine-tuned pre-trained model, optimize the objective function through the labels, iterate until convergence or the maximum number of iterations is reached, and save the trained fine-tuned pre-trained model; Step 8. Load the parameters of the trained and fine-tuned pre-trained model, input the divided test set samples, select the category corresponding to the maximum value according to the category probability output by the model to obtain the predicted label, and then obtain the diagnosis result.
2. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: In Step 1, the rotating machinery fault data set includes a plurality of sample pairs, and each pair of samples includes a vibration signal and its corresponding state label.
3. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: In Step 2, the normalization process uses Z-score normalization, and the truncation method is used to ensure the balance of the number of samples of each category; The steps of data set division are as follows: First, 80% of sample pairs are randomly selected from each working condition subset and the labels are removed to form an unlabeled training set; In the second step, 10% of the remaining 20% labeled data is randomly selected as the labeled fine-tuning set; in the third step, the remaining 10% is used as the labeled test set.
4. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: In Step 3, the process of applying the data enhancement library to the time signal and the frequency signal respectively is as follows: For a given one-dimensional vibration signal V, let X t ≡V, then represents the i-th time series sample in the unlabeled training set, and n represents the number of samples in the mini-batch; the temporal signal enhancement process is as follows: in, represents the enhanced time domain signal, Indicates the enhanced A t represents the time domain enhancement library, T j represents the combined operation of temporal data enhancement, T a , T b and T c They represent the data enhancement operations of jitter, scaling and permutation, respectively, and α, β, and γ are the enhancement parameters; Using Fast Fourier Transform The time domain signal Convert to frequency domain signal The frequency domain signal enhancement process is as follows: Among them, A f represents the frequency domain enhancement library, F j represents the frequency domain data enhancement combination operation, which consists of adding frequency components (F a ) and frequency component removal operation (F b ), ω, θ and ∈ are frequency domain enhancement parameters, and the enhanced frequency domain signal is expressed as Indicates the enhanced F a and F b They refer to the data enhancement operations of adding and removing frequency components respectively. The enhanced frequency domain signal is expressed as 5. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: The Step 4 includes: extracting time-frequency features from the time and frequency signals using a time encoder and a frequency encoder with the same structure but independent parameters; For the original timing signal and the enhanced time domain signal The encoding is performed using a temporal encoder, where D is the signal dimension; the encoding process is defined as follows: in, represents the time-domain coded representation of the temporal encoder output, l is the feature dimension, is the i-th temporal encoding representation in the mini-batch, represents the temporal enhancement coded representation of the temporal encoder output, is the i-th temporal enhancement encoding representation in the mini-batch; Time_Enocer(,) represents the temporal encoder, θ t are the parameters of the time encoder; The frequency signal is processed using the transformer frequency encoder to obtain frequency embedding representation and frequency enhancement coding representation. The process is expressed as follows: in, represents the frequency domain coded representation of the frequency encoder output, is the ith frequency-encoded representation in the mini-batch; represents the frequency domain enhanced coded representation of the frequency encoder output, is the i-th frequency enhanced encoding representation in the mini-batch; Frequency_Enocer(,) represents the frequency encoder, θ f is the parameter of the frequency encoder.
6. The method for fault diagnosis of rotating machinery based on self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: In the Step 4, a pre-training target loss function is constructed through a multi-level contrastive learning strategy, including: instance-level, cluster-level and fusion-level contrastive loss functions are defined as follows: The time-frequency feature encoding result is expressed as After linear projection to the time-frequency space, the corresponding instance-level embedding representation is obtained Subsequently, these projected instance-level embeddings are represented To operate, represents the original time-frequency embedding view, where represents the original embedding representation of the i-th time-frequency view in the mini-batch, let represents the enhanced time-frequency embedding view, where represents the enhanced embedding representation of the i-th time-frequency view in the mini-batch; in a training dataset of size n, each original embedding view Its corresponding enhanced embedded view Positive pairing and negative pairing with other samples; original embedding view Instance-level contrast loss It is expressed as: Among them, τ I represents the instance-level temperature parameter, s(·) represents the cosine similarity between samples, considering the original embedding view and enhanced embedded views The symmetry between , the instance-level contrast loss is further expressed as: in, Represents the original embedded view The instance-level contrast loss is Represents an enhanced embedded view Instance-level contrast loss; By constructing two pseudo classifiers with different parameters for the time domain and frequency domain encoding representations, we can obtain the time-frequency pseudo labels. and Respectively represent the original view category assignment The i-th column and enhanced view category assignment The i-th column of , n is the number of samples, k is the number of clusters; the original view category assignment Cluster-level contrast loss It is expressed as: Among them, τ c is the clustering level temperature parameter. In the clustering process, in order to prevent degenerate solution, the cross entropy constraint is introduced: in Respectively represent the original view category assignment and Enhanced View Category Assignment Considering the symmetry between the original view and the enhanced view, the cluster-level contrast loss is expressed as: in, Indicates the original view category assignment The cluster-level contrast loss is Indicates enhanced view category assignment The cluster-level contrast loss is is the cross entropy loss function; Enhanced embedding representation in time domain Frequency domain enhanced embedding representation Fusion is performed using time domain perception and frequency domain perception respectively; fusion features dominated by time domain perception The generated representation is as follows: Among them, Attn is the attention network, which is constructed by the feed-forward neural network FFN, [,] represents the horizontal concatenation of matrices or vector sets along the rows, is the common enhanced embedding representation after concatenation, is the fusion weight vector calculated by Attn, λ t ,λ f They represent the probability distribution of the time domain and frequency domain fusion weights calculated by softmax, τ is the scaling factor, It is the time-frequency perception weight dominated by time domain perception. Representation of time-domain enhanced embedding representation and frequency domain enhanced embedding representation A set of, d is the vector dimension; Fusion features dominated by frequency domain perception The generated representation is as follows: in, The fusion weight vector is calculated by weight sharing Attn, λ t ,λ f They represent the probability distribution of the time domain and frequency domain fusion weights calculated by softmax, respectively. It is the time-frequency perception weight dominated by frequency domain perception; Align features using momentum contrast method; Denote the parameters of the attention fusion module as θ fu , the parameters of the momentum model are denoted by θ mom , the update process is expressed as: Where m∈[0,1] is the momentum coefficient; Through the momentum model M mom Generate time- and frequency-dominated fusion features and calculate time-frequency embedding features and The fused features are used as queries and the features generated by the momentum model are used as keys, which are expressed as follows: Among them, the fusion features dominated by time domain As the query q t , frequency domain dominated fusion features As another query q f , key k t and k f Respectively represent the momentum model M mom The generated time domain and frequency domain momentum fusion features; {q t ,q f ,k t ,k f } as input to calculate the time-frequency contrast loss; momentum contrast loss It is expressed as: in, represents the i-th time domain query feature, represents the jth frequency domain momentum feature, and each time domain query feature The corresponding frequency domain momentum characteristics Positive pairing, and negative pairing with other samples, B represents the mini-batch size, τ m is the L2-normalized temperature hyperparameter, L t and L f denote the contrast loss of the time domain and frequency domain branches respectively; Where B represents the mini-batch size, τ m is the temperature hyperparameter of the L2 normalization. Construct a public pseudo classifier CPC to represent the time-frequency fusion feature Input a public pseudo classifier to generate a public pseudo label Then, the pseudo labels generated by combining the enhanced time-frequency representation and Together we construct the confidence target distribution Q: The confident target assignment with the highest probability is chosen as the target assignment; highly confident assignments are boosted and instances at cluster boundaries are further blurred by: As an auxiliary target distribution to guide the clustering results, q ij represents the i-th row and j-th column of the confidence target distribution Q, p ij As an auxiliary target distribution to guide the clustering results, the optimization objective is expressed as: in, Public pseudo labels The i-th row and j-th column of , KL(||) represents the Kullback-Leibler divergence, is the objective loss function optimized with high confidence; The domain loss function is calculated based on the constructed target loss function, and the total loss is obtained by joint training through the domain loss function and the fusion level comparison guidance; Combining the above loss functions, by expressing the time domain and frequency domain representation Using Formula 6 and Formula 9, we get the following two domain loss functions: in, and Represent the intra-domain objective loss functions in the time domain branch and the frequency domain branch respectively; Therefore, the contrast loss in the time-frequency domain It is expressed as: Contrast loss in time-frequency domain Jointly trained with the proposed fusion-level contrastive guidance, which includes momentum contrastive loss and high confidence guide Total pre-training loss It is expressed as follows: Where λ1 and λ2 are loss weight coefficients.
7. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: The Step 5 includes: The unlabeled training set is processed by a random selection method and loaded into the data loader in batches. Each batch contains four inputs: time-frequency signal pairs and enhanced time-frequency signal pairs. The sample pairs are sent to the encoder network for feature extraction, and then the parameters in the encoder network are iteratively updated according to the pre-training objective function until the pre-training objective loss function converges or the maximum number of iterations is reached. Finally, the parameters of the pre-training model are saved.
8. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: In Step 6, the objective loss function of the fine-tuning model is It is composed of the cross entropy loss function and the target loss function as follows: Where N is the total number of samples, C is the total number of categories, and y i,c is the true label of the i-th sample for category c, is the predicted probability of the i-th sample for category c.
9. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: The Step 7 includes: The labeled fine-tuning set and its true label input are input into the fine-tuned pre-trained model, and compared with the true label according to the fine-tuning objective loss function. The parameters of the fine-tuning model encoder network and linear classifier are iteratively updated until convergence or the maximum number of iterations is reached. After fine-tuning is completed, the parameters of the fine-tuned pre-trained model are saved.
10. The rotating machinery fault diagnosis method based on the self-supervised time-frequency contrast fusion network according to claim 1 is characterized in that: In Step 8, the fine-tuning model is a fine-tuning model with frozen parameters after training, and the final fault diagnosis result is the average value of the model after multiple trainings.
Citation Information
Cited By
Electromechanical equipment fault diagnosis method based on data enhancement and self-supervision time-frequency comparison
CN120763624A
Mechanical fault diagnosis method based on prototype-driven double-view-angle collaborative comparison fusion network
CN120952053A
Boundary-guided rolling bearing semi-supervised fault diagnosis method and system
CN121051573A
Hydrogen leakage positioning method, device, equipment and medium
CN121412650A