Carrying equipment bearing fault diagnosis method and device based on self-supervised learning

By using the self-supervised learning TFDDCF model, combined with an improved convolutional neural network and a Transformer encoder, the problem of insufficient labeled data in bearing fault diagnosis is solved, achieving high-precision and efficient fault identification, and is suitable for bearing diagnosis in complex dynamic systems.

CN121301745APending Publication Date: 2026-01-09GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524420.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing intelligent diagnostic methods have low accuracy in identifying bearing faults when labeled data is lacking, cannot effectively adapt to bearing vibration signals in complex dynamic systems, and suffer from severe noise interference.

Method used

A self-supervised learning approach is adopted, which is pre-trained using the TFDDCF model. The TF contrastive loss function and the TF fusion loss function are used to extract potential fault features from unlabeled data. The improved convolutional neural network and Transformer encoder are combined to fuse and align the time domain and frequency domain features, and the target diagnostic model is constructed and fine-tuned.

Benefits of technology

It improves the accuracy and efficiency of bearing fault diagnosis, reduces noise interference, and realizes intelligent fault diagnosis in the case of scarce labeled data, with fast, accurate and stable diagnostic capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301745A_ABST
    Figure CN121301745A_ABST
Patent Text Reader

Abstract

The invention discloses a carrying equipment bearing fault diagnosis method and device based on self-supervised learning, and the method comprises the steps: collecting vibration signals during the operation of a carrying equipment bearing, and enabling the vibration signals to comprise time domain data signals of different fault states of the bearing; preprocessing the vibration signals of the bearing to obtain final time domain and frequency domain representations of the vibration signals under the training set, the fine tuning set and the test set; constructing a TFDDCF model to pre-train the time domain representation and the frequency domain representation, and introducing a TF comparison loss function and a TF fusion loss function into the TFDDCF model to pre-train and supervise the domain representation and the frequency domain representation; and constructing a target diagnosis model MTFDDCF based on the TFDDCF model to carry out learning training and fine tuning, inputting the test set into the fine-tuned target diagnosis model MTFDDCF to extract features in the test signal, and classifying different types of features. According to the method, noise interference can be effectively reduced, the capacity of extracting unlabeled data feature information is high, the structure is simple, the accuracy is high, and intelligent carrying equipment bearing fault diagnosis can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bearing fault diagnosis of carrying equipment, and in particular to a bearing fault diagnosis method and device based on self-supervised learning of carrying equipment. BACKGROUND

[0002] As one of the key components in the transportation industry, bearing of carrying equipment often faces various complex loads during operation, such as vertical and axial loads caused by vehicle vibration, impact load between bogie and wheel, and road unevenness. Therefore, maintaining the good state of the bearing is crucial to ensure the safe operation of the transportation equipment. Under such complex working conditions, the bearing of carrying equipment may have multiple failure modes, including wear, spalling, cracking, overheating and corrosion. If the bearing fails, it will directly affect the stability and safety of the transportation equipment, and thus pose a great risk to personnel safety and cargo transportation. Therefore, real-time fault monitoring and diagnosis of the bearing of carrying equipment to ensure timely detection of faults and take effective measures has become an important means to ensure transportation safety.

[0003] Early bearing fault diagnosis relies on signal processing and machine learning methods, but such methods require a lot of manual operation and cannot well adapt to bearing vibration signals in complex dynamic systems. In recent years, with the rapid progress of artificial intelligence, deep learning algorithms have been successfully applied to the field of bearing fault diagnosis and have achieved remarkable research results. However, deep learning methods require a large amount of labeled data for training, and the complexity and nonlinearity of bearing fault signals in actual working conditions and the high cost of obtaining labeled data make existing intelligent fault diagnosis techniques face severe challenges in practical applications. SUMMARY

[0004] The present application aims to provide a bearing fault diagnosis method and device based on self-supervised learning of carrying equipment, which can effectively reduce noise interference, extract unlabelled data feature information, and has the advantages of simple structure, high accuracy, and intelligent bearing fault diagnosis of carrying equipment. In order to achieve the above purpose, the present application adopts the following technical solutions: According to one aspect of the present application, a bearing fault diagnosis method based on self-supervised learning of carrying equipment is provided, which comprises the following steps: Step S01: Collecting vibration signals of the bearing of carrying equipment during operation, the vibration signals including time domain data signals of different fault states of the bearing; Step S02: Normalize the vibration signal of the bearing, and then divide the normalized bearing vibration signal into a training set, a fine-tuning set, and a test set. The training set consists of unlabeled data, and the fine-tuning set is divided into labeled fine-tuning set and unlabeled fine-tuning set according to the proportion. Perform fast Fourier transform on the vibration signals under the training set, fine-tuning set, and test set, and then enhance the vibration signals under each dataset to obtain the final time domain and frequency domain representation of the vibration signals under the training set, fine-tuning set, and test set. Step S03: Construct a TFDDCF model to pre-train the time-domain and frequency-domain representations. The TFDDCF model consists of a dual encoder and a TF fusion encoder. The time-domain and frequency-domain representations of unlabeled signals in the training set are divided into fixed-length signal segments, which are simultaneously input into the dual encoder to extract time-domain and frequency-domain features. These segments are then fed into the TF fusion encoder for further feature extraction and fusion. The dual encoder consists of an improved convolutional neural network (ICNN), a linear embedding layer, a peak sparse attention mechanism (PSA), time-domain and frequency-domain encoders, a projection head, and a prediction head. The time-domain and frequency-domain encoders (which have identical network structures) are both composed of L stacked Transformer encoder layers. Each Transformer encoder layer includes a multi-head self-attention block (MSA) and a layer normalization function. The TF fusion encoder consists of a peak sparse attention (ICNN), a temporal and frequency domain encoder, a projection head, and a prediction head. Specifically, a multi-head fusion attention module (MFA) is added between the MSA and FFN in the temporal encoder to form a new temporal encoder. The MFA has the same structure as the MSA, but the difference is that the MFA accepts both temporal and frequency inputs simultaneously, thereby realizing a multi-head cross-attention mechanism between the temporal and frequency modes. Step S04: In the TFDDCF model, the TF contrastive loss function and the TF fusion loss function are introduced to supervise the pre-training of the time domain and frequency domain representations. The TF contrastive loss function is used to supervise the pre-training process of the dual encoder, and the TF fusion loss function is used to supervise the pre-training process of the TF fusion encoder. In each pre-training stage, the similarity and common features in the time domain and frequency domain are effectively mined, so as to learn the potential fault feature representation in the unlabeled signal. Step S05: After pre-training, construct the target diagnostic model MTFDDCF based on the TFDDCF model. Since there is a deviation between the learned potential fault feature representation and the actual fault, fine-tune the target diagnostic model MTFDDCF using a small amount of labeled data in the fine-tuning set. Correct the mapping relationship between the features learned by the target diagnostic model MTFDDCF and the corresponding fault categories, so that the target diagnostic model MTFDDCF has a more accurate understanding of the fault features corresponding to different fault categories of bearings. Step S06: Input the test set into the fine-tuned target diagnostic model MTFDDCF, extract features from the test signal through the target diagnostic model MTFDDCF, classify different types of features, and output the diagnostic accuracy.

[0005] In the preferred embodiment of the above scheme, in step S02, the vibration signal of the bearing is normalized using a maximum-minimum normalization method to standardize the vibration signals of various bearing faults, so that the amplitude distribution of the time-domain signal is between 0 and 1. The maximum-minimum normalization method satisfies: ,(1); in, f n,z Representing the z The first of the columns n 1 eigenvalue, Represents the normalized eigenvalues. f z Indicates the first z All feature values ​​of the column.

[0006] The enhancement processing of vibration signals from various datasets includes the following steps: Constructing a data augmentation module that includes Gaussian signal-to-noise ratio (SNR) noise processing, translation, scaling, and time reverse processing, inputting vibration signals from various datasets into the data augmentation module and sequentially performing Gaussian SNR, translation, scaling, and time reverse processing to obtain enhanced time-domain data signals, thereby improving the data augmentation module's adaptability to data diversity and invariance.

[0007] In a further preferred embodiment of the above scheme, step S03, constructing the TFDDCF model for pre-training the time-domain and frequency-domain representations, includes the following sub-steps: Step S301: Divide the time-domain and frequency-domain representations of unlabeled signals in the training set into signal segments of fixed length. Treat the time-domain and frequency-domain signal segments within the same time-domain segment as corresponding time-frequency signal pairs. Input the time-domain and frequency-domain signals simultaneously into an improved convolutional neural network (ICNN) through dual channels for preliminary feature extraction. Further process the signal using a linear projection embedding layer. Add a Class embedding before each initially extracted feature segment, and then add a Position embedding to each feature segment and Class embedding to obtain an embedding sequence. The specific process is as follows: Assume a given segment of length... D Time-domain signal sequence It satisfies: (2); (3); (4); (5); (6); in, This represents the embedding sequence that will be input into the temporal encoder. Represents a time-domain sequence of fixed length after segmentation; , R 1×l Represents a single row l A vector of columns whose elements come from the real number field. R , MaxP ( T ,k, s () represents the max pooling operation, with a pooling kernel size of k Step size is s , Represents the ReLU activation function; BN (·) indicates the normalization operation. conv ( T , W ) indicates a convolution operation using kernel W. Drop ( T , p () indicates a Dropout operation, with a drop probability set to 1. p ; T 1. T 2. T c Representing time-domain signal sequences respectively T The output of convolutional layer 1, convolutional layer 2, and dropout layer; input data fragment After ICNN , Indicates category embedding, Represents positional embedding, X This indicates that the image has passed through a linear projection layer. Step S302: Input the embedded sequence into the PSA mechanism, and its output result Adding it to the original input fragment forms a residual connection, the formation process of which satisfies: (7); (8); in, This represents the output result after each feature segment is input into the PSA mechanism. Represents the query matrix. Represents the bond matrix. Representative value matrix, The dimension representing each attention head. This indicates that the terms are multiplied one by one. Indicates taking the first e Larger scores are set to negative infinity, while other scores are set to negative infinity. Soft(·) indicates that the Softmax function is applied to the sparsified attention scores. This represents the linear transformation matrix of the output. For time-domain embedded sequences; frequency signals f Simultaneously, the input is fed into a dual-encoder structure consisting of L stacked Transformer encoders, and the embedded sequence is obtained according to the processing method of the time-domain signal input to the encoder. Then embed the time domain sequence The input to the time-domain encoder satisfies: (9); (10); in, and Represents embedded sequence The time-domain output after passing through the time-domain encoder; a similar process yields... and Represents embedded sequence Frequency domain output after frequency encoder; Step S303: ... and As input to the TF fusion encoder, and The class embedding in the projection head serves as the input to the projection head and is ultimately processed by the prediction head to obtain the temporal query. and frequency query It satisfies: (11); in, and The projection heads represent the time domain and frequency domain, respectively. and These represent the time-domain and frequency-domain prediction heads, respectively, both of which are constructed from feedforward networks (FFNs). Step S304: Align the time-domain and frequency-domain features using the momentum comparison model to fully learn the potential fault characteristics in the vibration signal; the momentum comparison model employs a momentum update strategy for backpropagation and gradient update, wherein the update process of the momentum update strategy is expressed as follows: (12); in, m The momentum coefficient is represented by the parameters of the dual encoder as follows: The parameters of the momentum ratio model are expressed as follows: , Combined with temporal embedded sequences The training process continuously updates the momentum ratio calculation model. Its update calculation process satisfies: (13); in, For time domain key , For frequency key; for time key and frequency key As input to the TF contrastive loss function, and combined with time-domain queries and frequency query The TF contrastive loss function is calculated together to align the time-domain and frequency-domain features during pre-training, thus fully learning the potential fault features in the signal. The momentum model is only used during pre-training and not in fault diagnosis. This approach ensures the alignment of time-domain and frequency-domain features during pre-training, fully learning the potential fault features in the signal and improving training efficiency. Step S305: The ICNN of the TF fusion encoder is used to extract features from the output of the dual encoders, and the extracted features are filtered using PSA. The filtered features are then merged with the original input and input into the time-domain and frequency-domain encoders. The ICNN feature extraction and peak sparse attention filtering of the extracted features satisfy the following conditions: (14); (15); (16); (17); Step S306, output the time domain Class Embedded The input is fed into the projection head, and then passes through the prediction head to obtain the binary output logit z, where the binary output satisfies: (18); Step S307: The binary output is fed into the TFDDCF model, which supervises the pre-training process of the dual encoder and the TF fusion encoder by referencing the TF contrastive loss function and the TF fusion loss, respectively.

[0008] In a further preferred embodiment of the above scheme, step S04, in which the TFDDCF model supervises the pre-training process of the dual encoder and the TF fusion encoder by referring to the TF contrastive loss function and the TF fusion loss function respectively, includes the following steps: Step S401, TF contrastive loss function is used and dot product sum and The dot product is used to measure the feature similarity between time-domain and frequency-domain signals, thereby maximizing the alignment of TF features of positive sample pairs in the time-domain and frequency-domain signals, and minimizing the similarity features between negative sample pairs in the time-domain and frequency-domain signals; the similarity calculation process satisfies the following formula: (19); (20); (twenty one); Where B represents the small batch size. It is used for Normalized hyperparameters, Indicates the time-domain signal loss. Indicates frequency domain signal loss. Indicates time-frequency contrast loss; Step S402: Guide the mutual matching optimization of time domain and frequency domain features, and define the output logit of positive sample pairs as 1 and the logit of negative sample pairs as 0. Step S403: The negative pairs of samples with high feature similarity in the similarity calculation of the TF contrastive loss function are taken as hard negative samples. The hard negative samples are represented by a negative frequency sequence sampled based on the feature similarity distribution. Here, the time-domain segment... Hard negative frequency samples Sampled as: (twenty two); P (·) represents the probability mass function of the feature similarity distribution; Fusion loss of TF fusion loss function satisfy: (twenty three); in, sig (·) represents the sigmoid function. From the front via TF fusion encoder Obtained from From time-domain hard negative pairs Obtained from Then from the frequency domain hard negative pair Obtained in; Step S404: The total loss ζ of the TFDDCF model is measured by combining the TF contrastive loss function and the TF fusion loss function. The total loss ζ satisfies the following expression: ; (twenty four); Where β represents the weight parameter assigned to the TF contrastive loss function and the TF fusion loss function. , These represent the time-frequency contrast loss and the time-frequency fusion loss, respectively.

[0009] Based on another aspect of the present invention, a bearing fault diagnosis device for transport equipment based on a self-supervised learning algorithm is also proposed. The fault diagnosis device includes a bearing vibration signal acquisition module, a signal preprocessing module, a model pre-training module, a data fine-tuning module, and a fault diagnosis module. The bearing vibration signal acquisition module, the signal preprocessing module, the model pre-training module, and the bearing fault diagnosis module are electrically connected. The bearing vibration signal acquisition module installs accelerometers in the vertical, radial, and axial directions on the bearing housing to acquire the bearing vibration signal. The signal preprocessing module is responsible for normalizing the acquired bearing vibration signal and performing time-domain enhancement-frequency-domain transformation on the normalized bearing vibration signal to obtain the time-domain and frequency-domain representations of the signal. The time-frequency domain data is divided into a training set, a fine-tuning set, and a test set. Only the fine-tuning set contains a small amount of labeled data, while the training set and the test set contain unlabeled data. The model pre-training module is responsible for inputting the time-frequency domain signals from the bearing training set into the constructed TFDDCF model to learn the potential fault features in the signals and obtain common features in the time and frequency domain signals. The data fine-tuning module, based on the potential fault feature representations obtained after learning and training in the TFDDCF model, constructs a target diagnostic model MTFDDCF using the TFDDCF model and fine-tunes the target diagnostic model MTFDDCF using a small amount of labeled data. The fault diagnosis module inputs the test set data into the fine-tuned target diagnostic model MTFDDCF for diagnosis, realizing fault classification of bearings when labeled data is scarce.

[0010] In a further preferred embodiment of the above scheme, the TFDDCF model consists of a dual encoder and a TF fusion encoder. The time-frequency signal is first processed by the dual encoder, and then the independent features in the time and frequency domain signals are learned after alignment in the time and frequency domains. Subsequently, the output of the dual encoder is used as the input of the TF fusion encoder for processing. Through the fusion of the time and frequency domains, the common features in the time and frequency domain signals are learned. In summary, because the present invention adopts the above-described technical solution, the present invention has the following beneficial technical effects: (1) This invention improves the self-supervised learning method and uses predefined agent tasks to learn the latent feature representation of unlabeled data. It can be used for intelligent fault diagnosis of bearings in various equipment when labeled data is scarce, which greatly improves the efficiency and accuracy of fault diagnosis. (2) In this invention, the time domain and frequency domain features of the signal are extracted simultaneously, which increases the comprehensiveness and complementarity of the information; at the same time, a feature extraction structure based on ICNN-Transformer is proposed, and a peak sparse attention mechanism is designed to effectively extract key features and improve computational efficiency. (3) In this invention, two pre-proxy tasks are designed: time-frequency comparison and time-frequency fusion, corresponding to the dual encoder and TF fusion encoder in TFDDCF. A time-frequency comparison loss function and a time-frequency fusion loss function are constructed to supervise the pre-training process of TFDDCF. During the time-frequency feature comparison and fusion process, potential fault features of unlabeled data are learned, providing a foundation for subsequent diagnosis; (3) The present invention can effectively reduce noise interference, has a strong ability to extract feature information of unlabeled data, has a simple structure and high accuracy, can realize intelligent transportation equipment bearing fault diagnosis, and has significant advantages of fast diagnosis speed, high diagnosis accuracy and high stability. It can meet the actual diagnostic application needs of bearings when labeled data is scarce, and also has broad application potential in other fields. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating the execution steps of a self-supervised learning-based method for diagnosing bearing faults in transport equipment, as described in this invention.

[0013] Figure 2 This is a schematic diagram of the dual encoder operation process of the present invention.

[0014] Figure 3 This is a schematic diagram of the ICNN structure of the present invention.

[0015] Figure 4 This is a schematic diagram of the peak sparse attention mechanism structure of the present invention.

[0016] Figure 5 This is a schematic diagram of the TF fusion encoder workflow of the present invention.

[0017] Figure 6 This is a schematic diagram of the TF contrastive loss function of the present invention.

[0018] Figure 7 This is a schematic diagram of the TF fusion loss function of the present invention.

[0019] Figure 8 This is a schematic diagram of the fault diagnosis process of the present invention.

[0020] Figure 9 This is a structural block diagram of a bearing fault diagnosis device for transport equipment based on a self-supervised learning algorithm provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, preferred embodiments will be listed below with reference to the accompanying drawings to provide a clear and complete description of the invention. However, it should be noted that the examples in the specification are preferred examples, not all examples. Furthermore, many details listed in the specification are merely to provide the reader with a deep understanding of the key issues of this invention; the invention can be implemented even without these specific details.

[0022] Please see Figure 1 This invention provides a method for diagnosing bearing faults in transport equipment based on self-supervised learning, comprising the following steps: Step S01: Collect vibration signals of the bearings of the transport equipment during operation, wherein the vibration signals include time-domain data signals of different fault states of the bearings; Step S02: Normalize the bearing vibration signal, then divide the normalized bearing vibration signal into a training set, a fine-tuning set, and a test set. The training set consists of unlabeled data. The fine-tuning set is divided into labeled and unlabeled sets proportionally. A small amount of fine-tuning set data is labeled for the model fine-tuning stage. The test set data is used to verify the model. Subsequently, perform a Fast Fourier Transform on the vibration signals in the training set, fine-tuning set, and test set. Then, enhance the vibration signals in each dataset to obtain the final time-domain and frequency-domain representations of the vibration signals in the training set, fine-tuning set, and test set. The enhancement processing of the vibration signals in each dataset includes the following steps: constructing a Gaussian signal-to-noise ratio (SNR) noise set. The data augmentation module performs noise reduction, translation, scaling, and time reversal processing. Vibration signals from various datasets are input into the data augmentation module and sequentially subjected to Gaussian signal-to-noise ratio, translation, scaling, and time reversal to obtain augmented time-domain data signals. After augmentation of the time-domain data, the frequency-domain data and the augmented time-domain data are used as two types of data input to the subsequent model. Data augmentation can diversify the original signal and improve the model's adaptability to data diversity and invariance. Step S03: Construct a TFDDCF model to pre-train the temporal and frequency domain representations. The TFDDCF model consists of a dual encoder and a TF fusion encoder. The dual encoder comprises an improved convolutional neural network (ICNN), a linear embedding layer, a peak sparse attention (PSA) mechanism, temporal and frequency domain encoders, a projection head, and a prediction head. Both the temporal and frequency domain encoders consist of L stacked Transformer encoder layers. Each Transformer encoder layer includes a multi-head self-attention block (MSA), layer normalization, and a feedforward network (FFN), as shown in the figure. Figure 2 As shown; the TF fusion encoder consists of ICNN, peak sparse attention, time-domain and frequency-domain encoders, projection head and prediction head; the time-domain and frequency-domain representations of unlabeled signals in the training set are divided into signal segments of fixed length, and simultaneously input into the dual encoder to extract time-domain and frequency-domain features, and then sent into the TF fusion encoder for feature extraction and fusion again; Step S04: In the TFDDCF model, the TF contrastive loss function and the TF fusion loss function are introduced to supervise the pre-training of the time domain and frequency domain representations. The TF contrastive loss function is used to supervise the pre-training process of the dual encoder, and the TF fusion loss function is used to supervise the pre-training process of the TF fusion encoder. In each pre-training stage, the similarity and common features in the time domain and frequency domain are effectively mined, so as to learn the potential fault feature representation in the unlabeled signal. Step S05: After pre-training, construct the target diagnostic model MTFDDCF based on the TFDDCF model. Since there is a deviation between the learned potential fault feature representation and the actual fault, fine-tune the target diagnostic model MTFDDCF using a small amount of labeled data in the fine-tuning set. Correct the mapping relationship between the features learned by the target diagnostic model MTFDDCF and the corresponding fault categories, so that the target diagnostic model MTFDDCF has a more accurate understanding of the fault features corresponding to different fault categories of bearings. Step S06: Input the vibration signals in the test set into the fine-tuned target diagnostic model MTFDDCF, extract the features in the test signals through the target diagnostic model MTFDDCF, classify different types of features, and output the diagnostic accuracy. If the diagnostic accuracy of the target diagnostic model MTFDDCF on the test set is greater than 95%, then the diagnostic accuracy is used for fault diagnosis of bearings of transport equipment.

[0023] In this invention, in step S01, vibration signals of seven bearing states at different speeds are collected using a motor bearing test bench for transport equipment. The corresponding bearing fault types are numbered, such as the vibration signal types of the seven bearing fault states {inner ring fault, outer ring fault, rolling element fault, inner ring and outer ring fault, inner ring and rolling element fault, outer ring and rolling element fault, inner ring and outer ring fault and rolling element fault} being numbered {0,1,2,3,4,5,6} respectively. The normalized vibration signals are divided into a training set, a fine-tuning set, and a test set, which are used in subsequent stages. The training set consists of unlabeled data (1800 samples per fault type), the fine-tuning set consists of a small amount of labeled data (30 samples per fault type), and the test set consists of unlabeled data (200 samples per fault type). The acquired vibration signals are subjected to a Fast Fourier Transform to obtain frequency domain data. The time domain data is input into a data enhancement module to increase the diversity and generalization ability of the data. In this invention, step S02, which normalizes the vibration signal of the bearing, specifically includes the following steps: The vibration signal of the bearing is normalized using a maximum-minimum normalization method to standardize the vibration signals of various bearing faults, ensuring that the amplitude of the time-domain signal is distributed between 0 and 1. The maximum-minimum normalization method satisfies the following: ,(1); in, f n,z Representing the z The first of the columns n 1 eigenvalue, Represents the normalized eigenvalues. f z Indicates the first z All feature values ​​of the column; In step S03, the TFDDCF model is constructed to pre-train the time-domain and frequency-domain representations. The process of inputting the time-domain and frequency-domain representation signals into the TFDDCF model for pre-training includes the following steps: S301: The time-domain and frequency-domain representations of unlabeled signals in the training set are divided into fixed-length signal segments. Time-domain and frequency-domain signal segments within the same time segment are considered as corresponding time-frequency signal pairs. The time-domain and frequency-domain signals are simultaneously input into the improved convolutional neural network (ICNN) through dual channels for preliminary feature extraction. Its structure is as follows: Figure 2Therefore, the improved convolutional neural network ICNN is responsible for the initial feature extraction and combines the PSA mechanism to select important features, thereby reducing computational complexity and enhancing the model's ability to process long sequences. The temporal encoder and frequency encoder extract features of the signal in different domains, respectively. It is worth noting that the temporal encoder and frequency encoder have the exact same network structure; after the initial feature extraction, further processing is performed through a linear projection embedding layer. A learnable class embedding is added before each initially extracted feature segment. This class embedding is a trainable parameter used to represent the sequence information of the input signal and is optimized during model training. Sequence information refers to data with temporal or spatial order, such as temporal signals and frequency signals. Its final output is used to calculate the representation of the TF contrastive loss. Simultaneously, a position embedding is added to each feature segment and class embedding to obtain an embedding sequence (i.e., adding a position embedding next to each feature segment and a position embedding next to each class embedding, such as...). Figure 2 As shown), this allows the model to capture sequence information during training. In this process, the dual encoder aligns the time-domain and frequency-domain signals, and TFDDCF learns the latent fault feature representation in the signal during the time-frequency domain alignment stage. The specific process is as follows: Assuming a given segment of length... D Time-domain signal sequence It satisfies: (2); (3); (4); (5); (6); in, This represents the embedding sequence that will be input into the temporal encoder. This represents a time-domain sequence of fixed length after segmentation. , R 1×l Represents a single row l A vector of columns whose elements come from the real number field. R , MaxP ( T ,k, s () represents the max pooling operation, with a pooling kernel size of k Step size is s , Represents the ReLU activation function. BN (·) indicates the normalization operation. conv ( T ,W ) indicates a convolution operation using kernel W. Drop ( T , p () indicates a Dropout operation, with a drop probability set to 1. p ; T 1. T 2. T c Representing time-domain signal sequences respectively T The structure of the improved convolutional neural network ICNN, after passing through the outputs of convolutional layer 1, convolutional layer 2, and the Dropout layer, can be found in [link to documentation]. Figure 3 As shown; Input data fragment After ICNN , Indicates category embedding, Represents positional embedding, X This indicates that the image has passed through a linear projection layer. SS302: Inputting embedded sequences into the PSA mechanism; its structure and workflow can be found in [link to documentation]. Figure 4 Its output results Adding the processed signal to the original input segment creates a residual connection (the processed signal is added to the unprocessed signal, which is the original signal). The formation process satisfies the following: (7); (8); in, This represents the output result after each feature segment is input into the PSA mechanism. Represents the query matrix. Represents the bond matrix. Representative value matrix, The dimension representing each attention head. This indicates that the terms are multiplied one by one. Indicates taking the first e Larger scores are set to negative infinity, while other scores are set to negative infinity. Soft(·) indicates that the Softmax function is applied to the sparsified attention scores. This represents the linear transformation matrix of the output. For temporal embedding sequences; frequency signal f Simultaneously, the input is fed into a dual-encoder structure consisting of L stacked Transformer encoders, and the embedded sequence is obtained according to the processing method of the time-domain signal input to the encoder. Then embed the time domain sequence The input to the time-domain encoder satisfies: (9); (10); in, and Represents embedded sequence The time-domain output after passing through the time-domain encoder; a similar process yields... and Represents embedded sequence The frequency domain output is obtained after passing through a frequency encoder; the encoder consists of L Transformer layers. Simultaneously, a similar process is performed to obtain... and Represents embedded sequence The output after the frequency encoder; Step S303: ... and As input to the TF fusion encoder, and The first Class embedding in the algorithm serves as the input to the projection head and is ultimately processed by the prediction head to obtain the temporal query. and frequency query It satisfies: (11); in, and The projection heads represent the time domain and frequency domain, respectively. and These represent the time-domain and frequency-domain prediction heads, respectively, both of which are constructed from feedforward networks (FFNs). Step S304: Align time-domain and frequency-domain features using a momentum comparison model to fully learn potential fault characteristics in the vibration signal. The momentum comparison model addresses the difficulty in effectively aligning time-domain and frequency-domain features. Its structure is basically the same as that of a dual encoder in processing embedded sequences, only lacking a prediction block. Simultaneously, the input time-domain and frequency-domain signals to different channels must be aligned as much as possible between the two types of data, aligning a segment of signal from the time domain to the frequency domain, thus achieving the goal of having both time-domain and frequency-domain representations of the same signal segment. The structure of the embedded sequence can be found in [reference needed]. Figure 2 The momentum model employs a momentum update strategy (the momentum update process corresponds to the time-frequency domain alignment process), thus eliminating the need to consider normal backpropagation and gradient updates. The momentum model is only used during pre-training and not in bearing fault diagnosis. This approach ensures alignment of time-domain and frequency-domain features during pre-training, fully learning potential fault features in the signal and improving training efficiency. The momentum comparison model uses a momentum update strategy for backpropagation and gradient updates, where the momentum update strategy's update process is expressed as follows: (12); in,m The momentum coefficient is represented by the parameters of the dual encoder as follows: The parameters of the momentum ratio model are expressed as follows: Combined with temporal embedded sequences The training process continuously updates the momentum ratio calculation model. Its update calculation process satisfies: (13); in, For time domain key , For frequency key; for time key and frequency key As input to the TF contrastive loss function, and combined with time-domain queries and frequency query Together, we calculate the TF contrastive loss function to obtain the alignment of time-domain and frequency-domain features during pre-training, thus fully learning the potential fault features in the signal; Step S305: The workflow of the TF fusion encoder is as follows: The improved convolutional neural network (ICNN) of the TF fusion encoder is used to extract features from the output of the dual encoders. The extracted features are then filtered using PSA. Subsequently, the filtered features are merged with the original input and input into the time-domain and frequency-domain encoders. The TF fusion encoder is used for deep fusion of time-frequency dual-domain features, combining the outputs of the dual encoders... and As input to this encoder in the TF fusion encoder, the projection head is the output of the fused time-domain and frequency-domain processing. Figure 5 As shown, the time-domain and frequency-domain encoders in the TF fusion encoder differ from those in the standard encoder. Specifically, a multi-head fusion attention module (MFA) is added between the MSA and FFN in the time-domain encoder to form a new time-domain encoder. The MFA has the same structure as the MSA, but unlike the MSA, the MFA accepts both time-domain and frequency inputs simultaneously, thus implementing a multi-head cross-attention mechanism between the time-domain and frequency modes, as shown below. Figure 5As shown, the workflow of the TF fusion encoder is as follows: First, feature extraction is performed using ICNN, and the features are then filtered using the PSA mechanism. Subsequently, the filtered features are merged with the original input and fed into the time-domain and frequency-domain encoders. To achieve deep fusion of time-frequency domain features, the component encoders in the TF fusion encoder have information exchange capabilities. Furthermore, the TF fusion encoder adopts a parallel structure, unlike the traditional Transformer encoder-decoder structure. This parallel design aims to treat time-domain and frequency modes as equally important, expecting them to be seamlessly integrated at the same learning speed. The parallel cross-attention mechanism effectively promotes the fusion of deep time-domain and frequency features; the improved convolutional neural network ICNN for feature extraction and peak sparse attention for filtering the extracted features respectively satisfy the following: (14); (15); (16); (17); in, The frequency domain characteristics of the merged data are represented. This represents the temporal characteristics of the merged data; Step S306, output the time domain Class Embedded The input is fed into the projection head, then passes through the prediction head, resulting in the binary output logit z, where the binary output satisfies the result. z : (18); Step S307: The binary output is fed into the TFDDCF model, which supervises the pre-training process of the dual encoder and the TF fusion encoder by referencing the TF contrastive loss function and the TF fusion loss, respectively.

[0024] In this invention, in step S4, the TF contrastive loss function and the TF fusion loss function are introduced into the TFDDCF model to supervise the training of the dual encoder and the TF fusion encoder, respectively, to learn the potential fault features in the unlabeled data. The TFDDCF model supervises the pre-training process of the dual encoder and the TF fusion encoder by referring to the TF contrastive loss function and the TF fusion loss function, respectively, including the following steps: Step S401, TF contrastive loss function is used and dot product sum and The dot product is used to measure the feature similarity between time-domain and frequency-domain signals, thereby maximizing the alignment of TF features of positive sample pairs in the time-domain and frequency-domain signals and minimizing the similarity features between negative sample pairs in the time-domain and frequency-domain signals. TF contrastive loss function and TF fusion loss function are introduced into the dual encoder and TF fusion encoder in TFDDCF, respectively, to supervise the training of the dual encoder and TF fusion encoder. Time-domain and frequency-domain features can be accurately aligned in the dual encoder, fully mining the similar features in the time-domain and frequency-domain signals. Furthermore, time-domain and frequency-domain features can be effectively fused in the TF fusion encoder, thereby mining common features in the time-domain and frequency-domain. The goal of the TF contrastive loss function is to maximize the feature similarity within each positive sample pair. A positive sample pair is a pair of signals composed of time-domain and frequency-domain signals at a certain moment. Conversely, a negative sample pair is a pair of time-frequency domain signals at different moments, i.e., mismatched or incorrectly matched signals. Therefore, the feature similarity between negative sample pairs should be minimized. At the same time, since the features within a negative sample pair represent different fault information, it is necessary to minimize the similarity between features. The workflow is as follows: Figure 6 As shown, the loss function utilizes and dot product sum and The similarity is measured by the dot product; the similarity calculation process satisfies the following formula: (19); (20); (twenty one); Where B represents the small batch size. It is used for Normalized hyperparameters, Indicates the time-domain signal loss. Indicates frequency domain signal loss. Indicates time-frequency contrast loss; Although the dataset used in the pre-training process is unlabeled, the objects to be optimized by the TFDDCF model are generated during the calculation of contrastive loss. This is achieved by maximizing the alignment of TF features between positive sample pairs, while distancing the features between negative sample pairs. This strategy allows the TFDDCF model to learn potential key fault feature representations from unlabeled data during pre-training, and the alignment of TFs also facilitates TF fusion in the next step of the TFDDCF model. Step S402 guides the mutual matching and optimization of time-domain and frequency-domain features, defining the output logit of positive sample pairs as 1 and the logit of negative sample pairs as 0; to extract key fault features from the fused features, a TF fusion loss function is constructed to supervise the training of the TF fusion encoder, and its working principle is as follows.Figure 7 As shown, the TF fusion loss function is optimized by guiding the mutual matching of time-domain and frequency-domain features, achieving correct feature fusion and matching. In other words, the output logits of a positive sample pair is defined as 1, i.e., true (the output of the positive sample pair is true), and the logits of a negative sample pair is defined as 0, i.e., false (the output of the negative sample pair is false). Furthermore, to further improve the effectiveness of the loss function, a hard negative sampling strategy is used to sample hard negative sample pairs for calculating the TF fusion loss.

[0025] Step S403: Negative pairs with high feature similarity in the similarity calculation of the TF contrastive loss function are taken as hard negative samples. A negative frequency sequence is sampled based on the feature similarity distribution to represent hard negative samples. Hard negative sample pairs are negative pairs with high feature similarity in the TF contrastive loss, but still have some differences. Although identifying hard negative sample pairs is more difficult, it is more likely to learn key feature representations in the process. Specifically, for each time-domain segment, a negative frequency sequence is sampled based on the feature similarity distribution as a hard negative sample; the higher the feature similarity, the greater the probability of it being sampled. Wherein, the time-domain segment... Hard negative frequency samples Sampled as: (twenty two); P (·) represents the probability mass function of the feature similarity distribution; the fusion loss of the TF fusion loss function. satisfy: (twenty three); in, sig (·) represents the sigmoid function. From the front via TF fusion encoder Obtained from From time-domain hard negative pairs Obtained from Then from the frequency domain hard negative pair The TF fusion loss function guides the TF fusion encoder to predict whether features match. Therefore, the fusion of time and frequency domain features enhances prediction. Conversely, the TF fusion loss function further enhances the advantages of the TFDDCF model in feature fusion. This TFDDCF model, by optimizing the TF fusion loss, achieves deep time-frequency interaction and can learn more fault feature representations from unlabeled data. Step S404: During the training and optimization of the parameters of the dual encoder and the TF fusion encoder in the TFDDCF model, the total loss ζ of the TFDDCF model is measured by combining the TF contrastive loss function and the TF fusion loss function. The total loss ζ satisfies the following expression: ; (twenty four); Where β represents the weight parameter assigned to the TF contrastive loss function and the TF fusion loss function. , These represent the time-frequency contrast loss and the time-frequency fusion loss, respectively. The smaller the loss value, the better the performance of the method during training.

[0026] In this invention, step S6, verifying the accuracy of the diagnostic effect of the MTFDDCF model, includes the following steps: To verify that the diagnostic model can effectively identify the fault categories of the bearings of the transport equipment, the signals in the test set are input into the finely tuned MTFDDCF. The MTFDDCF first extracts the features from the test signals, classifies different types of features, and outputs the diagnostic accuracy. If the diagnostic accuracy of the MTFDDCF on the test set is greater than 95%, it is used for fault diagnosis of the bearings of the transport equipment; Figure 8 The flowchart shown is for the fault diagnosis process of bearings in transport equipment: First, a set of bearing fault signals is collected from the test bench and labeled in a 1:3 ratio. One-third of the signals (1 / 3 of the fault signal set) are labeled, while the remaining two-thirds are left unlabeled. Preliminary processing of the signals is then performed. Next, during the pre-training process of TFDDCF, the unlabeled test dataset is used for pre-training to mine potential fault features in the data. Then, the target diagnostic model MTFDDCF is constructed and initialized. Simultaneously, a small amount of labeled data is used to fine-tune the MTFDDCF model to meet the fault diagnosis task of bearings in transport equipment. Finally, the constructed target model is used to diagnose and classify the test data. The parameters of the dual encoder and TF fusion encoder in the TFDDCF model are shown in Table 1. Table 1. Structural parameters of dual encoders and TF fusion encoders in the TFDDCF model Wherein, PSA represents peak sparse attention mechanism, MSA represents multi-head self-attention block, FFN represents feedforward network, and MFA represents multi-head fusion attention module.

[0027] Furthermore, such as Figure 9 As shown, this invention also proposes a fault diagnosis device for bearings of transport equipment based on a self-supervised learning algorithm, see [link to relevant documentation]. Figure 9The system includes a bearing vibration signal acquisition module 901, a signal preprocessing module 902, a model pretraining module 903, a data fine-tuning module 904, a fault diagnosis module 905, and a fault diagnosis result verification module 906. The bearing vibration signal acquisition module 901, the signal preprocessing module 902, the model pretraining module 903, the data fine-tuning module 904, the bearing fault diagnosis module 905, and the fault diagnosis result verification module 906 are electrically connected. The bearing vibration signal acquisition module 901 has acceleration sensors installed on the bearing housing in the vertical, radial and axial directions to acquire the vibration signal of the bearing. The signal preprocessing module 902 is responsible for normalizing the obtained bearing vibration signal, and performing time-domain enhancement-frequency domain transformation on the normalized bearing vibration signal to obtain the time-domain and frequency-domain representations of the signal. The processed time-domain-frequency-domain data is divided into training set, fine-tuning set and test set. Only the fine-tuning set contains a few labeled data, while the training set and test set contain unlabeled data. The model pre-training module 903 is responsible for inputting the time-frequency domain signals from the bearing training set into the constructed TFDDCF model to learn the potential fault features in the signals and obtain the common features in the time and frequency domain signals. The data fine-tuning module 904 constructs a target diagnostic model MTFDDCF based on the potential fault feature representation obtained after learning and training in the TFDDCF model, and fine-tunes the target diagnostic model MTFDDCF using a small amount of labeled data. The fault diagnosis module 905 inputs the test set data into the fine-tuned target diagnosis model MTFDDCF for diagnosis, thereby realizing fault classification of bearings when labeled data is scarce.

[0028] The fault diagnosis result verification module 906 is used to verify diagnostic results with a diagnostic accuracy greater than 95% on the test set, and is then used for fault diagnosis of bearings in transport equipment. Meanwhile, to verify the effectiveness of the proposed self-supervised learning algorithm-based bearing fault diagnosis method for transport equipment, this method is compared with other self-supervised learning methods and semi-supervised deep learning methods, including TFAI, TFPred, DDTLN, SSMN, and ISSML. A set of examples containing fault data under seven different operating conditions is used for verification; Table 2 shows the results of each method under different performance indicators.

[0029] Table 2. Performance of each method under different performance indicators According to the data in Table 2, the MTFDDCF method outperforms other comparative methods in terms of average accuracy at different speeds. The MTFDDCF method achieves the highest average accuracy at 2100 rpm, reaching 99.37%, which is 2.6% higher than TFAI, 1.03% higher than TFPred, 1.78% higher than DDTLN, 1.5% higher than SSMN, and 2.34% higher than ISSMN. Examples demonstrate that the proposed method and device for fault diagnosis of bearings in transport equipment based on a self-supervised learning algorithm can effectively perform condition detection and fault diagnosis.

[0030] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that for those skilled in the art, any improvements, equivalent substitutions, and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for fault diagnosis of bearings in transport equipment based on self-supervised learning, characterized in that, The bearing fault diagnosis method includes the following steps: Step S01: Collect vibration signals of the bearings of the transport equipment during operation, wherein the vibration signals include time-domain data signals of different fault states of the bearings; Step S02: Normalize the vibration signal of the bearing, and then divide the normalized bearing vibration signal into a training set, a fine-tuning set, and a test set. The training set consists of unlabeled data, and the fine-tuning set is divided into labeled fine-tuning set and unlabeled fine-tuning set according to the proportion. Perform fast Fourier transform on the vibration signals under the training set, fine-tuning set, and test set, and then enhance the vibration signals under each dataset to obtain the final time domain and frequency domain representation of the vibration signals under the training set, fine-tuning set, and test set. Step S03: Construct a TFDDCF model to pre-train the time-domain and frequency-domain representations. The TFDDCF model consists of a dual encoder and a TF fusion encoder. The time-domain and frequency-domain representations of unlabeled signals in the training set are divided into fixed-length signal segments, which are simultaneously input into the dual encoder to extract time-domain and frequency-domain features. These segments are then fed into the TF fusion encoder for further feature extraction and fusion. The dual encoder consists of an improved convolutional neural network, a linear embedding layer, a peak sparse attention mechanism, a time-domain and frequency-domain encoder, a projection head, and a prediction head. Both the time-domain and frequency-domain encoders consist of L stacked Transformer encoder layers. Each Transformer encoder layer includes a multi-head self-attention block, layer normalization, and a feedforward network. The TF fusion encoder consists of peak sparse attention, a time-domain and frequency-domain encoder, a projection head, and a prediction head. Step S04: In the TFDDCF model, the TF contrastive loss function and the TF fusion loss function are introduced to supervise the pre-training of the time domain and frequency domain representations. The TF contrastive loss function is used to supervise the pre-training process of the dual encoder, and the TF fusion loss function is used to supervise the pre-training process of the TF fusion encoder. In each pre-training stage, the similarity and common features in the time domain and frequency domain are effectively mined, so as to learn the potential fault feature representation in the unlabeled signal. Step S05: After pre-training, construct the target diagnostic model MTFDDCF based on the TFDDCF model. Since there is a deviation between the learned potential fault feature representation and the actual fault, fine-tune the target diagnostic model MTFDDCF using a small amount of labeled data in the fine-tuning set. Correct the mapping relationship between the features learned by the target diagnostic model MTFDDCF and the corresponding fault categories, so that the target diagnostic model MTFDDCF has a more accurate understanding of the fault features corresponding to different fault categories of bearings. Step S06: Input the test set into the fine-tuned target diagnostic model MTFDDCF, extract features from the test signal through the target diagnostic model MTFDDCF, classify different types of features, and output the diagnostic accuracy.

2. The method for diagnosing bearing faults in transport equipment based on self-supervised learning as described in claim 1, characterized in that, In step S02, the vibration signal of the bearing is normalized using a maximum-minimum normalization method to standardize the vibration signals of various bearing faults, ensuring that the amplitude of the time-domain signal is distributed between 0 and 1. The maximum-minimum normalization method satisfies the following: ,(1); in, f n,z Representing the z The first of the columns n 1 eigenvalue, Represents the normalized eigenvalues. f z Indicates the first z All feature values ​​of the column.

3. The method for diagnosing bearing faults in transport equipment based on self-supervised learning as described in claim 1, characterized in that, The enhancement processing of vibration signals from each dataset includes the following steps: A data augmentation module is constructed that includes Gaussian signal-to-noise ratio (SNR) processing, translation, scaling, and time-domain reversal processing. Vibration signals from various datasets are input into the data augmentation module and sequentially processed with Gaussian SNR, translation, scaling, and reversal to obtain the augmented time-domain data signal.

4. The method for diagnosing bearing faults in transport equipment based on self-supervised learning as described in claim 1, characterized in that, Step S03 includes the following sub-steps: Step S301: Divide the time-domain and frequency-domain representations of unlabeled signals in the training set into signal segments of fixed length. Treat the time-domain and frequency-domain signal segments within the same time-domain segment as corresponding time-frequency signal pairs. Input the time-domain and frequency-domain signals simultaneously into an improved convolutional neural network (ICNN) through dual channels for preliminary feature extraction. Further process the signal using a linear projection embedding layer. Add a Class embedding before each initially extracted feature segment, and then add a Position embedding to each feature segment and Class embedding to obtain an embedding sequence. The specific process is as follows: Assume a given segment of length... D Time-domain signal sequence It satisfies: (2); (3); (4); (5); (6); in, This represents the embedding sequence that will be input into the temporal encoder. This represents a time-domain sequence of fixed length after segmentation. , R 1×l Represents a single row l A vector of columns whose elements come from the real number field. R , MaxP ( T ,k, s () represents the max pooling operation, with a pooling kernel size of k Step size is s , This represents the ReLU activation function. BN (·) indicates the normalization operation. conv ( T , W ) indicates a convolution operation using kernel W. Drop ( T , p () indicates a Dropout operation, with a drop probability set to 1. p ; T 1. T 2. T c Representing time-domain signal sequences respectively T The output of convolutional layer 1, convolutional layer 2, and dropout layer; input data fragment After ICNN , Indicates category embedding, Represents positional embedding, X This indicates that the image has passed through a linear projection layer. Step S302: Input the embedded sequence into the PSA mechanism, and its output result Adding it to the original input fragment forms a residual connection, the formation process of which satisfies: (7); (8); in, This represents the output result after each feature segment is input into the PSA mechanism. Represents the query matrix. Represents the bond matrix. Representative value matrix, The dimension representing each attention head. This indicates that the terms are multiplied one by one. Indicates taking the first e Larger scores are set to negative infinity, while other scores are set to negative infinity. Soft(·) indicates that the Softmax function is applied to the sparsified attention scores. This represents the linear transformation matrix of the output. For temporal embedding sequences; frequency signal f Simultaneously, the input is fed into a dual-encoder structure consisting of L stacked Transformer encoders, and the embedded sequence is obtained according to the processing method of the time-domain signal input to the encoder. Then embed the time domain sequence The input to the time-domain encoder satisfies: (9); (10); in, and Represents embedded sequence The time-domain output after passing through the time-domain encoder; a similar process yields... and Represents embedded sequence Frequency domain output after frequency encoder; Step S303: ... and As input to the TF fusion encoder, and The class embedding in the projection head serves as the input to the projection head and is ultimately processed by the prediction head to obtain the temporal query. and frequency query It satisfies: (11); in, and The projection heads represent the time domain and frequency domain, respectively. and These represent the time-domain and frequency-domain prediction heads, respectively, both of which are constructed from feedforward networks (FFNs). Step S304: Align the time-domain and frequency-domain features using the momentum comparison model to fully learn the potential fault characteristics in the vibration signal; the momentum comparison model employs a momentum update strategy for backpropagation and gradient update, wherein the update process of the momentum update strategy is expressed as follows: (12); in, m The momentum coefficient is represented by the parameters of the dual encoder as follows: The parameters of the momentum ratio model are expressed as follows: , Combined with temporal embedded sequences The training process continuously updates the momentum ratio calculation model. Its update calculation process satisfies: (13); in, For time domain key , For frequency keys; Time Domain Key and frequency key As input to the TF contrastive loss function, and combined with time-domain queries and frequency query Together, we calculate the TF contrastive loss function to obtain the alignment of time-domain and frequency-domain features during pre-training, thus fully learning the potential fault features in the signal; Step S305: The ICNN of the TF fusion encoder is used to extract features from the output of the dual encoders, and the extracted features are filtered using PSA. The filtered features are then merged with the original input and input into the time-domain and frequency-domain encoders. The ICNN feature extraction and peak sparse attention filtering of the extracted features satisfy the following conditions: (14); (15); (16); (17); Step S306, output the time domain Class Embedded The input is fed into the projection head, and then passes through the prediction head to obtain the binary output logit z, where the binary output satisfies: (18); Step S307: The binary output is fed into the TFDDCF model, which supervises the pre-training process of the dual encoder and the TF fusion encoder by referencing the TF contrastive loss function and the TF fusion loss, respectively.

5. The method for diagnosing bearing faults in transport equipment based on self-supervised learning as described in claim 1, characterized in that, In step S04, the TFDDCF model supervises the pre-training process of the dual encoder and the TF fusion encoder by using the TF contrastive loss function and the TF fusion loss function respectively, including the following steps: Step S401, TF contrastive loss function is used and dot product sum and The dot product is used to measure the feature similarity between time-domain and frequency-domain signals, thereby maximizing the alignment of TF features of positive sample pairs in time-domain and frequency-domain signals, and minimizing the similarity features between negative sample pairs in time-domain and frequency-domain signals. The similarity calculation process satisfies the following formula: (19); (20); (21); Where B represents the small batch size. It is used for Normalized hyperparameters, Indicates the time-domain signal loss. Indicates frequency domain signal loss. Indicates time-frequency contrast loss; Step S402: Guide the mutual matching optimization of time domain and frequency domain features, and define the output logit of positive sample pairs as 1 and the logit of negative sample pairs as 0. Step S403: The negative pairs of samples with high feature similarity in the similarity calculation of the TF contrastive loss function are taken as hard negative samples. The hard negative samples are represented by a negative frequency sequence sampled based on the feature similarity distribution. Here, the time-domain segment... Hard negative frequency samples Sampled as: (22); P (·) represents the probability mass function of the feature similarity distribution; Fusion loss of TF fusion loss function satisfy: (23); in, sig (·) represents the sigmoid function. From the front via TF fusion encoder Obtained from From time-domain hard negative pairs Obtained from Then from the frequency domain hard negative pair Obtained; Step S404: The total loss ζ of the TFDDCF model is measured by combining the TF contrastive loss function and the TF fusion loss function. The total loss ζ satisfies the following expression: ; (24); Where β represents the weight parameter assigned to the TF contrastive loss function and the TF fusion loss function. , These represent the time-frequency contrast loss and the time-frequency fusion loss, respectively.

6. The method for diagnosing bearing faults in transport equipment based on self-supervised learning as described in claim 1, characterized in that, If the diagnostic accuracy of the target diagnostic model MTFDDCF on the test set is greater than 95%, it can be used for fault diagnosis of bearings in transport equipment.

7. A fault diagnosis device for bearings of transport equipment based on a self-supervised learning algorithm, characterized in that, A fault diagnosis device for processing the self-supervised learning-based bearing fault diagnosis method for transport equipment as described in any one of claims 1 to 6, the fault diagnosis device comprising a bearing vibration signal acquisition module, a signal preprocessing module, a model pretraining module, a data fine-tuning module, and a fault diagnosis module, wherein the bearing vibration signal acquisition module, the signal preprocessing module, the model pretraining module, and the bearing fault diagnosis module are electrically connected. The bearing vibration signal acquisition module has accelerometers installed on the bearing housing in the vertical, radial and axial directions to acquire the vibration signals of the bearing. The signal preprocessing module is responsible for normalizing the obtained bearing vibration signal, and performing time-domain enhancement-frequency domain transformation on the normalized bearing vibration signal to obtain the time-domain and frequency-domain representations of the signal. The processed time-domain-frequency-domain data is divided into training set, fine-tuning set and test set. Only the fine-tuning set contains a few labeled data, while the training set and test set contain unlabeled data. The model pre-training module is responsible for inputting the time-frequency domain signals from the bearing training set into the constructed TFDDCF model to learn the potential fault features in the signals and obtain the common features in the time and frequency domain signals. The data fine-tuning module constructs a target diagnostic model MTFDDCF based on the potential fault feature representation obtained after learning and training in the TFDDCF model, and fine-tunes the target diagnostic model MTFDDCF using a small amount of labeled data. The fault diagnosis module inputs the test set data into the fine-tuned target diagnosis model MTFDDCF for diagnosis, realizing fault classification of bearings when labeled data is scarce.

8. A bearing fault diagnosis device for transport equipment based on a self-supervised learning algorithm according to claim 7, characterized in that, The TFDDCF model consists of a dual encoder and a TF fusion encoder. The time-frequency signal is first processed by the dual encoder, and then the independent features in the time and frequency domain signals are learned after alignment in the time and frequency domains. Subsequently, the output of the dual encoder is used as the input of the TF fusion encoder for processing. Through the fusion of the time and frequency domains, the common features in the time and frequency domain signals are learned.