A Rotating Machinery Fault Detection Method Based on Self-Supervised Learning
Through the self-supervised learning rotary mechanical fault detection method, combined with multi-scale convolutional fusion and dynamic weighting strategies, the problems of label dependence and feature information ignorance in the existing technology are solved, and efficient fault detection is achieved to adapt to actual industrial scenarios.
Patent Information
- Application Number
- CN202510490073.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art relies on a large number of comprehensive fault tags in rotary mechanical fault detection, which is cost-effective to obtain, ignore the multi-scale feature information of the timing signal, and cannot fully extract local and global information in the timing signal, making it difficult to adapt to complex nonlinear characteristics.
Using a rotating mechanical fault detection method based on self-supervised learning, the Kolmogorov–Arnold Networks (KAN) network, Reversible Instance Normalization (RevIN) normalization module, linear layer and improved bidirectional state space model architecture are used to perform feature extraction and timing modeling through multi-scale convolutional fusion module and dynamic weighting strategy to realize self-supervised fault detection.
While reducing training costs, it maintains high fault detection performance, improves the model's ability to analyze the timing characteristics of vibration signals, and improves the accuracy and adaptability of fault detection.
Smart Images

Figure CN120030483B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of fault detection, and specifically relates to a method for rotating machinery fault detection based on self-supervised learning. Background Art
[0002] Rotating machinery is widely used in modern industrial fields such as aerospace, energy, and manufacturing. Due to long-term high-load operation, rotating machinery is prone to failures, such as bearing wear, gear cracks, etc. If these failures are not detected in time, it may lead to a decline in equipment performance, an increase in operating costs, and even production safety accidents. Therefore, it is crucial to develop efficient fault diagnosis technologies.
[0003] Traditional fault diagnosis methods mainly include signal processing, physical models, and data-driven methods. Signal processing methods are effective for simple faults, but have limited effects when facing complex non-linear problems; although physical models can simulate the mechanical operating state, they are computationally complex and difficult to adapt to the non-linear characteristics in industrial environments; with the rapid development of deep learning technologies, data-driven methods have emerged continuously, which can automatically extract vibration signal features and improve the accuracy of fault diagnosis. In data-driven fault diagnosis methods, supervised and semi-supervised deep learning methods have the problem of label dependence.
[0004] Chinese Patent Application No. CN2021101911539 discloses a fault detection method based on numerical simulation. It collects the vibration response signals of the equipment, and uses methods such as similarity search and simulation verification to compare the sampled fault features with the sample features in the fault library to find the most similar fault type. Then it uses numerical simulation to verify whether this type of fault can generate the vibration signal of the sampled sample, so as to conduct fault detection of rotating machinery. Then it mainly relies on the sample features in the fault library and the sampled signals for feature comparison, but the process of establishing the fault library itself is relatively costly, and it is difficult to obtain comprehensive and representative real fault samples in actual industrial scenarios. It is difficult to find corresponding fault features in the existing fault library for some compound faults or new faults that are inconvenient to classify, and the time and economic costs of establishing the fault library are relatively high.
[0005] In summary, in actual industrial scenarios, the fault diagnosis of the existing technology needs to rely on a large number of comprehensive fault labels, which has the problem of too high acquisition costs, and the existing technology ignores the multi-scale feature information of time series signals, unable to comprehensively extract local and global information in time series signals, and at the same time it is easy to ignore the problem that the vibration signal itself has complex time series characteristics, making it difficult to obtain comprehensive and sufficient annotation information. Summary of the Invention
[0006] To solve the above technical problems, the present application provides a self-supervised learning-based rotating machinery fault detection method, which proposes a self-supervised fault detection method that does not rely on any fault label information and can maintain good detection performance while reducing training costs.
[0007] To achieve the above object, the present application is implemented through the following technical solutions:
[0008] The present application is a self-supervised learning-based rotating machinery fault detection method, characterized in that: the rotating machinery fault detection method is implemented through a rotating machinery fault detection model, and the rotating machinery fault detection model includes a data preprocessing module, a Kolmogorov–Arnold Networks network (i.e., KAN network), a Reversible InstanceNormalization normalization module (i.e., RevIN normalization module), a linear layer, and an improved bidirectional state space model architecture, where the improved bidirectional state space model architecture includes a multi-scale convolution fusion module, a state space model, and dynamic fusion weights , the multi-scale convolution fusion module includes a time-step convolution block, a feature convolution block, and a dynamic weight strategy, and performs feature extraction on the forward path and the reverse path. Specifically, the rotating machinery fault detection method includes the following steps:
[0009] Step 1: Collect the mechanical vibration signals of the device, and perform data preprocessing on the collected mechanical vibration signals through the data preprocessing module to obtain preprocessed data samples , divide the data of the preprocessed mechanical vibration signals into a training set and a test set, and all the mechanical vibration signals in the training set are of the normal category;
[0010] Step 2: Input the data samples preprocessed in Step 1 into the rotating machinery fault detection model. The mechanical vibration signals in each data sample have a time-step dimension and a feature dimension. The mechanical vibration signals in the data sample first pass through the KAN network to increase the feature dimension, and then pass through the RevIN normalization module for normalization processing. The normalized data samples are respectively sent to two identical linear layers to increase the feature dimension again to obtain a forward dimension-increased result and a dimension-increased result , the forward dimension-increased result is used to be fed into the multi-scale convolution fusion module in the forward and reverse paths, and the dimension-increased result is used for residual connection with the result after the subsequent state space model performs time series modeling;
[0011] Step 3: The forward dimension-increased result in Step 2 The time-step convolutional block and the feature convolutional block of the multi-scale convolutional fusion module are respectively input for feature extraction, and the results of the parallel time-step convolutional block and the feature convolutional block are weighted and fused by combining the dynamic weight strategy, fusing the time-step dimension and the feature dimension to obtain the result of the multi-scale convolutional fusion module. Among them, the forward-path feature extraction process is: the forward upsampling result is directly input into the time-step convolutional block and the feature convolutional block of the multi-scale convolutional module for feature extraction, and the dynamic fusion weights in the multi-scale convolutional fusion module on the forward path are used to perform weighted fusion on the results of the two parallel time-step convolutional blocks and the feature convolutional blocks to obtain the result of the multi-scale convolutional fusion module on the forward path . The reverse-path feature extraction process is: the forward upsampling result is flipped along the time-step dimension and then sent into the time-step convolutional block and the feature convolutional block of the multi-scale convolutional fusion module, and the dynamic fusion weights in the multi-scale convolutional fusion module on the reverse path are used to perform weighted fusion on the results of the two parallel time-step convolutional blocks and the feature convolutional blocks, and the fused result is flipped again to obtain the result of the multi-scale convolutional fusion module on the reverse path ;
[0012] Step 4: The result of the multi-scale convolutional fusion module on the forward path is input into the state space model on the forward path for temporal modeling, and the modeling result is connected with the upsampling result in Step 2 as a residual connection to obtain the modeling result of the forward-path state space model branch. The result of the multi-scale convolutional fusion module on the reverse path is input into the state space model on the reverse path for temporal modeling, and the modeling result is connected with the upsampling result in Step 2 as a residual connection to obtain the modeling result of the reverse-path state space model branch. The dynamic weight strategy is used to perform weighted fusion on the modeling result of the forward-path state space model branch and the modeling result of the reverse-path state space model branch to obtain the bidirectional modeling result after weighted fusion by the dynamic fusion weights ;
[0013] Step 5: The bidirectional modeling result after weighted fusion by the dynamic fusion weights is gradually downsampled through the linear layer and the KAN network, and is inverse-normalized through the RevIN normalization module to restore to the same dimension as the input data sample , and finally the reconstructed sample is output;
[0014] Step 6: Use mean square error as loss function to train the rotating machinery fault detection model. During the test, use the threshold derived from the Youden index as the fault judgment criterion. When the reconstructed sample output in step 5 is The input data sample of step 2 When the mean square error loss is greater than the threshold, it is judged as a fault. When the reconstructed sample output in step 5 The data sample from step 2 When the mean square error loss is less than the threshold, it is considered normal.
[0015] A further improvement of the present application is that: Step 2 converts the data sample preprocessed in Step 1 into The dimension is increased through the KAN network, and then normalized through the RevIN normalization module. The normalized result is further increased through the linear layer to obtain the positive dimension increase result. And the result of dimension upgrading The calculation formula is:
[0016]
[0017] in, To use the linear layer dimension increase operation, is the RevIN normalization operation, This is the operation of dimension increase using the KAN network.
[0018] A further improvement of the present application is that the feature convolution block in the multi-scale convolution fusion module is used to convolve the feature dimension, perform a one-dimensional convolution operation along the feature dimension, and the convolution kernel of the feature convolution block scans different features to explore the interaction mechanism between multi-dimensional features, and outputs the feature dimension convolution result as :
[0019]
[0020] Among them, when performing feature extraction on the forward path, The input data of the multi-scale convolution fusion module is the result of the forward dimension increase. , when doing feature extraction on the reverse path, The result of positive dimension upgrading Dimensionality increase result by flipping along the time step dimension , is the flipping operation along the time step dimension, It is the feature convolution operation;
[0021] The time step convolution block in the multi-scale convolution fusion module acts on the time step dimension and needs to exchange the input data of the multi-scale convolution fusion module. For the feature dimension and time step dimension, then perform a one-dimensional convolution operation along the time step dimension. By sliding the convolution kernel of the time step convolution block in the time dimension, local patterns of the features on the time step convolution block changing over time are extracted. Finally, the convolution result in the time step dimension is output as :
[0022]
[0023] Among them, represents the operation of swapping the time step dimension and the feature dimension. is the time step convolution operation. When performing feature extraction in the forward path, is the input data of the multi-scale convolution fusion module, that is, the forward dimension elevation result. When performing feature extraction in the reverse path, is the forward dimension elevation result. The dimension elevation result flipped along the time step dimension. ;
[0024] Through the dynamic weight strategy, the convolution result of the extracted feature dimension and the convolution result of the time step dimension are weighted and summed to obtain the result of the multi-scale convolution fusion module in the forward path:
[0025]
[0026] The result of the multi-scale convolution fusion module in the reverse path is:
[0027]
[0028] Among them, is the activation function. is the dynamic fusion weight in the multi-scale convolution fusion module in the forward path. is the dynamic fusion weight in the multi-scale convolution fusion module in the reverse path. represents the operation of flipping along the time step dimension.
[0029] A further improvement of this application is that: Step 4 specifically includes the following steps:
[0030] Step 4.1, input the result of the multi-scale convolution fusion module in the forward path into the state space model in the forward path for time series modeling, capture the dependence of the current state on the historical state, and make a residual connection between the modeled result and the dimension elevation result in Step 2:
[0031]
[0032] Among them, is the modeling result of the forward path state space model branch, is the modeling operation of the state space model, is the activation function;
[0033] Step 4.2: Input the result of the multi-scale convolution fusion module on the reverse path into the state space model of the reverse path for temporal modeling and make a residual connection with the dimensionality increase result :
[0034]
[0035] Among them, is the modeling result of the reverse path state space model branch;
[0036] Step 4.3: Use the dynamic fusion weight to weight and then add the modeling result of the forward path state space model branch and the modeling result
[0037]
[0038] Among them, is the two-way modeling result after weighted fusion by the dynamic fusion weight , is the activation function.
[0039] A further improvement of this application is that: the final output reconstruction sample in step 5 is:
[0040]
[0041] Among them, represents the inverse normalization operation of the RevIN normalization module, is the dimensionality reduction operation using the KAN network, represents the dimensionality reduction operation using the linear layer.
[0042] A further improvement of this application is that: in step 6, the mean square error is used as the loss function to train the rotating machinery fault detection model, which is expressed as:
[0043] Among them, represents the th mechanical vibration signal sample after preprocessing, represents the final output reconstruction sample, is the total number of data samples in the training set.
[0044] The beneficial effects of the present application are as follows:
[0045] Through the self-supervised fault detection model based on signal reconstruction, the present application can maintain high fault detection performance while having a relatively low training cost, solving the problems of the existing technology's dependence on labels and high detection costs in fault detection solutions.
[0046] By combining the multi-scale convolution fusion method, the present application solves the problem that the existing technology ignores the multi-scale feature information of time series signals, and combines the dynamic weight strategy to integrate information in different dimensions, effectively improving the feature extraction ability of the model.
[0047] The present application proposes an improved bidirectional state space model architecture, which solves the problem that the existing technology ignores the complex time series characteristics of vibration signals itself, improves the model's analysis ability of the time series characteristics of vibration signals, improves the model's reconstruction ability of normal signals, and further improves the accuracy of fault detection.
[0048] The present application is more suitable for actual industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is the flowchart of the detection method of the present application.
[0050] Figure 2 is the schematic diagram of the rotating machinery fault detection model of the present application.
[0051] Figure 3 is the improved bidirectional state space model architecture diagram of the present application.
[0052] Figure 4 is the loss scatter plot on the visualization test set of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will disclose the embodiments of the present invention in the form of diagrams. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0054] Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other; and all model trainings in the comparison experiment table are programmed using the Python language and completed on an Intel Xeon Gold 6230R processor and an NVIDIA GeForce RTX 3090 graphics card.
[0055] The embodiments of the present invention will be disclosed below with reference to the drawings. For the sake of clarity, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0056] As Figure 1 shown, this application is a method for rotating machinery fault detection based on self-supervised learning. The rotating machinery fault detection method is implemented through a rotating machinery fault detection model. As Figure 2 shown, the rotating machinery fault detection model includes a data preprocessing module, a Kolmogorov–Arnold Networks network, i.e., a KAN network, a ReversibleInstance Normalization normalization module, i.e., a RevIN normalization module, a linear layer, and an improved bidirectional state space model architecture. The improved bidirectional state space model architecture includes a multi-scale convolutional fusion module, a state space model, and dynamic fusion weights , and the multi-scale convolutional fusion module includes a time-step convolutional block, a feature convolutional block, and a dynamic weight strategy, and performs feature extraction on the forward path and the reverse path. Specifically, the rotating machinery fault detection method includes the following steps:
[0057] Step 1: Collect the mechanical vibration signals of the device, and perform data preprocessing on the collected mechanical vibration signals through the data preprocessing module to obtain preprocessed data samples , that is, use a sliding window to segment the collected mechanical vibration signals, perform a fast Fourier transform operation on the mechanical vibration signals within the window to map the mechanical vibration signals within the window from the time domain to the frequency domain, and perform normalization processing on the mechanical vibration signals. Divide the data of the normalized mechanical vibration signals into a training set and a test set. All the training sets are mechanical vibration signals under normal working conditions;
[0058] Step 2: Input the data samples preprocessed in Step 1 into the rotating machinery fault detection model. The mechanical vibration signals in each data sample have a time-step dimension and a feature dimension. The mechanical vibration signals in the data sample first pass through the KAN network to increase the feature dimension, and then pass through the RevIN normalization module for normalization processing. The normalized data samples are respectively fed into two identical linear layers to increase the feature dimension again, and the forward dimension-increasing result fed to the multi-scale convolutional fusion module is obtained and the dimension-increasing result used for residual connection with the result after the subsequent state space model performs time series modeling , and the calculation formula is:
[0059]
[0060] Among them, is the operation of increasing the dimension using a linear layer, is the RevIN normalization operation, is the operation of increasing the dimension using the KAN network.
[0061] Step 3: Input the forward dimension-increasing result in Step 2 into the temporal convolutional block and the feature convolutional block of the multi-scale convolutional fusion module respectively for feature extraction, and combine the dynamic fusion strategy to perform weighted fusion on the results of the parallel temporal convolutional block and the feature convolutional block, fusing the temporal dimension and the feature dimension to obtain the result of the multi-scale convolutional fusion module. Among them, the forward path feature extraction process is: Input the forward dimension-increasing result directly into the temporal convolutional block and the feature convolutional block of the multi-scale convolutional module for feature extraction, and use the dynamic fusion weights in the multi-scale convolutional fusion module on the forward path to perform weighted fusion on the results of the two parallel temporal convolutional block and the feature convolutional block to obtain the result of the multi-scale convolutional fusion module on the forward path . The reverse path feature extraction process is: Flip the forward dimension-increasing result along the temporal dimension, then send it into the temporal convolutional block and the feature convolutional block of the multi-scale convolutional fusion module, and use the dynamic fusion weights in the multi-scale convolutional fusion module on the reverse path to perform weighted fusion on the results of the two parallel temporal convolutional block and the feature convolutional block, and flip the fusion result again to obtain the result of the multi-scale convolutional fusion module on the reverse path . Specifically,
[0062] The feature convolutional block in the multi-scale convolutional fusion module is used to perform convolution on the feature dimension, perform a one-dimensional convolution operation along the feature dimension. The convolutional kernel of the feature convolutional block scans different features to explore the interaction mechanism between multi-dimensional features, and the output of the feature dimension convolution result is :
[0063]
[0064] Among them, when performing feature extraction on the forward path, is the input data of the multi-scale convolutional fusion module, that is, the forward dimension-increasing result . When performing feature extraction on the reverse path, is the forward dimension-increasing result after flipping along the temporal dimension , is the operation of flipping along the temporal dimension, is the feature convolution operation;
[0065] The temporal convolutional block in the multi-scale convolutional fusion module acts on the temporal dimension, and it is necessary to swap the feature dimension and the temporal dimension of the input data of the multi-scale convolutional fusion module , then perform a one-dimensional convolutional operation along the temporal dimension, and extract the local pattern of the features on the temporal convolutional block changing with time by the sliding of the convolutional kernel of the temporal convolutional block in the time dimension. Finally, the convolutional result of the temporal dimension is output as :
[0066]
[0067] Among them, represents the operation of swapping the temporal dimension and the feature dimension, is the temporal convolutional operation. When performing feature extraction on the forward path, is the input data of the multi-scale convolutional fusion module, that is, the forward upsampling result of the linear layer . When performing feature extraction on the reverse path, is the forward upsampling result The upsampling result flipped along the temporal dimension ;
[0068] In order to effectively integrate the output representations of the temporal convolution and the feature dimension convolution, through the dynamic weight strategy, the feature dimension convolution result extracted and the temporal dimension convolution result are weighted and summed to obtain the result of the multi-scale convolutional fusion module on the forward path:
[0069]
[0070] The result of the multi-scale convolutional fusion module on the reverse path is:
[0071]
[0072] Among them, is the activation function, is the dynamic fusion weight in the multi-scale convolutional fusion module on the forward path, is the dynamic fusion weight in the multi-scale convolutional fusion module on the reverse path, represents the operation of flipping along the temporal dimension. The dynamic fusion weight in the multi-scale convolutional fusion module on the forward path and the dynamic fusion weight in the multi-scale convolutional fusion module on the reverse path are used as trainable parameters and are obtained by model training. The goal of weighted fusion is to adaptively balance the information contributions of the time dimension and the feature dimension according to the task requirements.
[0073] Result of the multi-scale convolution fusion module on the forward path Input the state space model of the forward path for temporal modeling, and combine the modeling result with the dimensionality increase result in Step 2 Perform a residual connection to obtain the modeling result of the forward path state space model branch. The result of the multi-scale convolution fusion module on the reverse path Input the state space model of the reverse path for temporal modeling, and combine the modeling result with the dimensionality increase result in Step 2 Perform a residual connection to obtain the modeling result of the reverse path state space model branch. Use the dynamic weight strategy for the modeling result of the forward path state space model branch and the modeling result of the reverse path state space model branch Perform weighted fusion to obtain the two-way modeling result after weighted fusion with the dynamic weight strategy . The specific steps are as follows:
[0074] Step 4.1: Input the result of the multi-scale convolution fusion module on the forward path into the state space model of the forward path for temporal modeling to capture the dependence of the current state on historical states. The result after modeling is connected with the dimensionality increase result in Step 2 to perform a residual connection:
[0075]
[0076] where, is the modeling result of the forward path state space model branch, is the state space model modeling operation, is the activation function;
[0077] Step 4.2: Input the result of the multi-scale convolution fusion module on the reverse path into the state space model of the reverse path for temporal modeling and connect it with the dimensionality increase result to perform a residual connection:
[0078]
[0079] where, is the modeling result of the reverse path state space model branch;
[0080] Step 4.3: Use the dynamic fusion weight to perform weighted addition on the modeling result of the forward path state space model branch and the modeling result of the reverse path state space model branch :
[0081]
[0082] where, It is the result of bidirectional modeling after weighted fusion with dynamic fusion weights and is the activation function. For the activation function.
[0083] Step 5: The result of bidirectional modeling after weighted fusion with dynamic fusion weights is gradually reduced in dimension through a linear layer and a KAN network, and is de-normalized through a RevIN normalization module to restore to the same dimension as the input data sample and finally the reconstructed sample is output :
[0084]
[0085] Among them, represents the de-normalization operation of the RevIN normalization module, is the dimension reduction operation using the KAN network, represents the dimension reduction operation using the linear layer.
[0086] Step 6: Use the mean square error as the loss function to train the rotating machinery fault detection model. During testing, use the threshold derived from the Youden index as the fault determination criterion. When the mean square error loss between the reconstructed sample output in Step 5 and the input data sample in Step 2 is greater than the threshold, it is determined as a fault. When the mean square error loss between the reconstructed sample output in Step 5 and the data sample in Step 2 is less than the threshold, it is determined as normal. Specifically, when using the mean square error as the loss function to train the rotating machinery fault detection model, it is expressed as:
[0087] Among them, represents the th mechanical vibration signal sample after preprocessing, represents the finally output reconstructed sample, is the total number of data samples in the training set.
[0088] This application uses the rolling bearing fault dataset of Jiangnan University to verify the effectiveness of the proposed self-supervised fault detection method. The fault data is provided based on the fault diagnosis experiment of a three-phase induction motor (Mitsubishi SB-JR). The rated power of the motor is 3.7 kW, and the speed range is controlled from 400 to 800 r / min by adjusting the input voltage. The dataset covers vibration signals in four states: normal state, outer race fault, inner race fault, and rolling element fault. The faults are artificially created by wire cutting technology. The vibration data at a speed of 600 r / min is used for the experiment. All the data used is transformed from the time domain to the frequency domain using a sliding window method with a length of 1024. After normalization and standardization, the number of samples in the randomly sampled training set is 1246, the number of samples in the test set is 438, the length of each sample is 512, and the feature dimension is 1. Among them, all the samples in the training set are normal samples. The ratio of normal samples to abnormal samples in the test set is 1:1. The abnormal samples of each category are randomly and evenly sampled from each abnormal class. The samples in the training set and the test set do not overlap with each other.
[0089] In Table 1, the method of this application is compared with advanced methods in recent years. The results of these benchmark methods are obtained by re-experimenting according to the open-source code. Table 1 shows the four indicators of accuracy, precision, recall, and F1 score of each method on the Jiangnan University dataset. Compared with other methods in the table, the method proposed in this application reaches the optimal or sub-optimal performance in multiple indicators, which proves the effectiveness of the proposed method.
[0090] Table 1 Comparative experiment results on the Jiangnan University dataset
[0091]
[0092] In addition, a visualized loss scatter plot is also drawn. In Figure 4 , the green dots are normal signals, the blue dots are fault signals, and the horizontal dashed line represents the fault threshold for classification. The sample loss points above the threshold are fault samples, and those below the threshold are normal samples. From Figure 4 , it can be seen that this application can reconstruct normal signal samples well, and the reconstruction loss for fault samples is relatively large, indicating good fault detection performance.
[0093] The above description is only for the implementation manner of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for rotating machinery fault detection based on self-supervised learning, characterized in that: The rotating machinery fault detection method is implemented through a rotating machinery fault detection model, and the rotating machinery fault detection model includes an improved bidirectional state space model architecture. The improved bidirectional state space model architecture includes a multi-scale convolution fusion module, a state space model, and a dynamic weight strategy. Specifically, the rotating machinery fault detection method includes the following steps: Step 1: Collect the mechanical vibration signals of the device and perform preprocessing to obtain data samples , and divide the data samples into a training set and a test set, where the training set consists entirely of mechanical vibration signals of the normal category; Step 2: Input the data sample into the rotating machinery fault detection model. The mechanical vibration signal in the data sample first increases the feature dimension and then undergoes normalization processing. The normalized data sample is respectively fed into two identical linear layers to further increase the feature dimension, obtaining the forward dimension-increasing result fed to the multi-scale convolution fusion module and the dimension-increasing result for subsequent residual connection ; Step 3: The multi-scale convolution fusion module extracts features from the forward and reverse directions of the forward dimensionality increase result in Step 2 to obtain the result of the multi-scale convolution fusion module on the forward path and the result of the multi-scale convolution fusion module on the reverse path; Step 4: Perform temporal modeling on the results of the multi-scale convolution fusion module on the forward path and the results of the multi-scale convolution fusion module on the reverse path respectively to obtain the modeling results of the forward path state space model branch and the modeling results of the reverse path state space model branch. Use the dynamic weight strategy to perform weighted fusion on the modeling results of the forward path state space model branch and the modeling results of the reverse path state space model branch to obtain the bidirectional modeling results; Step 5. Dimensionality reduction is performed on the two-way modeling results obtained in Step 4, and then inverse normalization is carried out to finally output the reconstructed samples ; Step 6: Use the mean square error as the loss function to train the rotating machinery fault detection model. During testing, perform fault detection by comparing the mean square error loss of the data sample with the fault threshold.
2. The method for rotating machinery fault detection based on self-supervised learning according to claim 1, wherein: The rotating machinery fault detection model includes a data preprocessing module, a Kolmogorov–Arnold Networks (KAN) network, a Reversible Instance Normalization (RevIN) normalization module, a linear layer, and an improved bidirectional state space model architecture. The improved bidirectional state space model architecture includes a multi-scale convolutional fusion module, a state space model, and dynamic fusion weights. The multi-scale convolutional fusion module includes a time-step convolutional block, a feature convolutional block, and a dynamic weight strategy, and performs feature extraction on the forward path and the reverse path.
3. The method for rotating machinery fault detection based on self-supervised learning according to claim 2, characterized in that: The specific content of step 2 is as follows: the data samples preprocessed in step 1 are input into the rotating machinery fault detection model. The mechanical vibration signals in each data sample have a time step dimension and a feature dimension. The mechanical vibration signals in the data samples first pass through the KAN network to increase the feature dimension, and then pass through the RevIN normalization module for normalization processing. The normalized data samples are respectively fed into two identical linear layers to increase the feature dimension again, and the forward dimension-increasing results fed to the multi-scale convolution fusion module are obtained and the dimension-increasing results for residual connection with the results after the subsequent state space model time series modeling : The calculation formula is as follows: , Among them, is the operation of increasing the dimension using a linear layer, is the RevIN normalization operation, is the operation of increasing the dimension using the KAN network.
4. A method for rotating machinery fault detection based on self-supervised learning according to claim 2, characterized in that: Step 3 takes the positive dimensionality increase result in Step 2 and inputs it into the temporal convolution block and the feature convolution block of the multi-scale convolution fusion module respectively for feature extraction, and combines the dynamic weight strategy to perform weighted fusion on the results of the parallel temporal convolution block and feature convolution block, fusing the temporal dimension and the feature dimension to obtain the result of the multi-scale convolution fusion module. Among them, the feature extraction process of the forward path is: taking the positive dimensionality increase result and directly inputting it into the temporal convolution block and the feature convolution block of the multi-scale convolution module for feature extraction, and using the dynamic fusion weight in the multi-scale convolution fusion module on the forward path to perform weighted fusion on the results of the two parallel temporal convolution blocks and feature convolution blocks to obtain the result of the multi-scale convolution fusion module on the forward path . The feature extraction process of the reverse path is: flipping the positive dimensionality increase result along the temporal dimension, then sending it into the temporal convolution block and the feature convolution block of the multi-scale convolution fusion module, and using the dynamic fusion weight in the multi-scale convolution fusion module on the reverse path to perform weighted fusion on the results of the two parallel temporal convolution blocks and feature convolution blocks, and flipping the fusion result again to obtain the result of the multi-scale convolution fusion module on the reverse path . Specifically, it includes the following steps: Step 3.1: The feature convolution block in the multi-scale convolution fusion module is used to perform convolution on the feature dimension, and perform a one-dimensional convolution operation along the feature dimension. The convolution kernel of the feature convolution block scans different features, and the output feature dimension convolution result is :[[]]END]] , Among them, when performing feature extraction on the forward path, is the input data of the multi-scale convolution fusion module, that is, the forward dimensionality increase result When performing feature extraction on the reverse path, is the forward dimensionality increase result The dimensionality increase result flipped along the time step dimension , is the operation of flipping along the time step dimension, is the feature convolution operation; Step 3.2: The temporal convolutional block in the multi-scale convolutional fusion module acts on the temporal dimension, and it is necessary to swap the feature dimension and the temporal dimension of the input data of the multi-scale convolutional fusion module, and then perform a one-dimensional convolutional operation along the temporal dimension. By sliding the convolutional kernel of the temporal convolutional block in the time dimension, the local pattern of the features on the temporal convolutional block changing with time is extracted, and finally the convolutional result of the temporal dimension is output as : By sliding the convolutional kernel of the temporal convolutional block in the time dimension, the local pattern of the features on the temporal convolutional block changing with time is extracted, and finally the convolutional result of the temporal dimension is output as : , Among them, represents an operation that exchanges the time step dimension and the feature dimension, is a time step convolution operation. When performing feature extraction in the forward path, is the input data of the multi-scale convolution fusion module, that is, the forward dimension increase result , when performing feature extraction in the reverse path, is the forward dimension increase result is the dimension increase result flipped along the time step dimension ; Step 3.3: Through the dynamic weight strategy, perform weighted summation on the convolution results of the extracted feature dimensions and the convolution results of the time step dimensions to obtain the result of the multi-scale convolution fusion module on the forward path: , The result of the multi-scale convolution fusion module on the reverse path is: , Among them, is the activation function, is the dynamic fusion weight in the multi-scale convolution fusion module on the forward path, is the dynamic fusion weight in the multi-scale convolution fusion module on the reverse path, represents the operation of flipping along the time step dimension.
5. The method for rotating machinery fault detection based on self-supervised learning according to claim 4, wherein: In step 4, the result of the multi-scale convolution fusion module on the forward path is input into the state space model of the forward path for temporal modeling, and the modeling result is combined with the result of dimensionality increase in step 2 to make a residual connection, obtaining the modeling result of the forward path state space model branch. The result of the multi-scale convolution fusion module on the reverse path is input into the state space model of the reverse path for temporal modeling, and the modeling result is combined with the result of dimensionality increase in step 2 to make a residual connection, obtaining the modeling result of the reverse path state space model branch. The dynamic weight strategy is used to perform weighted fusion on the modeling result of the forward path state space model branch and the modeling result of the reverse path state space model branch to obtain the dynamic fusion weight The bidirectional modeling result after weighted fusion, specifically including the following steps: Step 4.1: The result of the multi-scale convolution fusion module in the forward path is input into the state space model of the forward path for temporal modeling to capture the dependence of the current state on historical states. The result after modeling is residually connected to the result of dimensionality increase in Step 2 as follows: , Among them, is the modeling result of the forward path state space model branch, is the modeling operation of the state space model, is the activation function; Step 4.2: The result of the multi-scale convolution fusion module on the reverse path is input into the state space model of the reverse path for temporal modeling and is subjected to residual connection with the result of dimensionality increase as follows: , Among them, is the modeling result of the reverse path state space model branch; Step 4.3, use the dynamic fusion weight for the modeling result of the forward path state space model branch and the modeling result of the reverse path state space model branch to perform weighted addition: , Among them, is the result of bidirectional modeling after weighted fusion by dynamic fusion weights , and is the activation function.
6. The method for rotating machinery fault detection based on self-supervised learning according to claim 5, wherein: In step 5, through the dynamic fusion weight The bidirectional modeling results after weighted fusion are gradually reduced in dimension through a linear layer and a KAN network, and are denormalized through a RevIN normalization module to restore to the same dimension as the input data sample and finally the reconstructed sample is output : , Among them, represents the denormalization operation of the RevIN normalization module, is the dimensionality reduction operation using the KAN network, represents the dimensionality reduction operation using a linear layer.
7. A method for rotating machinery fault detection based on self-supervised learning according to claim 1, characterized in that: In step 6, the mean squared error is used as the loss function to train the rotating machinery fault detection model. During testing, the threshold derived from the Youden index is used as the fault determination criterion. When the mean squared error loss between the reconstructed sample output in step 5 and the input data sample in step 2 is greater than the threshold, it is determined as a fault. When the mean squared error loss between the reconstructed sample output in step 5 and the data sample in step 2 is less than the threshold, it is determined as normal. Among them, using the mean squared error as the loss function to train the rotating machinery fault detection model is expressed as: , Among them, represents the th mechanical vibration signal sample after preprocessing, represents the final output reconstructed sample, is the total number of data samples in the training set.
Citation Information
Patent Citations
Multi-scale feature fusion gearbox fault diagnosis method based on self-attention mechanism
CN116010900A
Walking rail insulation fault positioning method based on multi-scale convolution fusion one-dimensional Transform
CN118501629A