Rotary machinery fault detection method based on self-supervised learning

By adopting a self-supervised learning method in rotary mechanical fault detection, combining KAN network, RevIN normalization module and multi-scale convolutional fusion module, the problem of relying on high-cost labels and ignoring multi-scale features in the existing technology is solved, and efficient fault detection is achieved.

CN120030483AActive Publication Date: 2025-05-23NANJING UNIV OF POSTS & TELECOMM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510490073.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-23
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The prior art relies on a large number of comprehensive fault tags in rotary mechanical fault detection, which is costly to obtain, and ignores the multi-scale feature information of timing signals, making it impossible to fully extract local and global information.

Method used

Using a rotating mechanical fault detection method based on self-supervised learning, fault detection without relying on fault labels is achieved through the Kolmogorov–Arnold Networks (KAN) network, Reversible Instance Normalization (RevIN) normalization module, multi-scale convolutional fusion module and improved bidirectional state space model architecture.

Benefits of technology

While reducing training costs, it maintains high fault detection performance, and can fully extract multi-scale feature information of timing signals, improving the model's feature extraction ability and the accuracy of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030483A_ABST
    Figure CN120030483A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of fault detection, and discloses a rotating machine fault detection method based on self-supervised learning, the method uses a KAN network to transform feature dimensions, introduces a RevIN normalization module to eliminate negative effects of inconsistent signal distribution, and provides an improved bidirectional state space model architecture, and the multi-scale convolution fusion module extracts local and global feature information, feeds the local and global feature information to the state space model of the corresponding path to perform time sequence modeling, performs dynamic weight fusion on a multi-scale convolution result through a dynamic weight strategy, and adaptively balances contribution of the bidirectional state space model in the time sequence modeling. The method does not need any fault annotation data, can maintain good fault detection performance while the training cost is extremely low, and is more in line with the actual industrial scene requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of fault detection, and specifically relates to a rotating machinery fault detection method based on self-supervised learning. Background Art

[0002] Rotating machinery is widely used in modern industrial fields such as aerospace, energy, and manufacturing. Rotating machinery is prone to failures such as bearing wear and gear cracks due to long-term high-load operation. If these failures are not detected in time, they may cause equipment performance degradation, increased operating costs, and even production safety accidents. Therefore, it is crucial to develop efficient fault diagnosis technology.

[0003] Traditional fault diagnosis methods mainly include signal processing, physical models, and data-driven methods. Signal processing methods are effective for simple faults, but have limited effects when facing complex nonlinear problems. Although physical models can simulate the operating state of machinery, they are complex to calculate and difficult to adapt to the nonlinear characteristics in industrial environments. With the rapid development of deep learning technology, data-driven methods continue to emerge, which can automatically extract vibration signal features and improve fault diagnosis accuracy. Among data-driven fault diagnosis methods, supervised and semi-supervised deep learning methods have label dependency problems.

[0004] Chinese patent application number CN2021101911539 discloses a fault detection method based on numerical simulation, which collects the vibration response signal of the equipment, uses similarity search and simulation verification methods to compare the sampled fault characteristics with the sample characteristics in the fault library, finds the most similar fault type, and then uses numerical simulation to verify whether this type of fault can generate the vibration signal of the sampled sample, so as to perform fault detection of rotating machinery. Then it mainly relies on the sample characteristics in the fault library and the sampled signal for feature comparison, but the process of establishing a fault library itself is relatively expensive. It is difficult to obtain comprehensive and representative real fault samples in actual industrial scenarios. Some complex faults or new faults that are inconvenient to classify are difficult to find corresponding fault characteristics in the existing fault library. The time and economic cost of establishing a fault library are relatively high.

[0005] In summary, in real industrial scenarios, fault diagnosis in existing technologies needs to rely on a large number of comprehensive fault labels, which have the problem of high acquisition costs. In addition, existing technologies ignore the multi-scale feature information of time series signals and cannot fully extract local and global information from time series signals. At the same time, they also tend to ignore the problem that vibration signals themselves have complex time series characteristics, making it difficult to obtain comprehensive and sufficient labeling information. Summary of the invention

[0006] In order to solve the above technical problems, the present application provides a rotating machinery fault detection method based on self-supervised learning, which proposes a self-supervised fault detection method that does not rely on any fault label information, and can reduce training costs while maintaining good detection performance.

[0007] In order to achieve the above objectives, this application is implemented through the following technical solutions:

[0008] The present application is a rotating machinery fault detection method based on self-supervised learning, characterized in that: the rotating machinery fault detection method is implemented by a rotating machinery fault detection model, the rotating machinery fault detection model includes a data preprocessing module, a Kolmogorov–Arnold Networks network, namely a KAN network, a Reversible InstanceNormalization normalization module, namely a RevIN normalization module, a linear layer and an improved bidirectional state space model architecture, wherein the improved bidirectional state space model architecture includes a multi-scale convolution fusion module, a state space model and a dynamic fusion weight The multi-scale convolution fusion module includes a time-step convolution block, a feature convolution block and a dynamic weight strategy, and performs feature extraction on the forward path and the reverse path. Specifically, the rotating machinery fault detection method includes the following steps:

[0009] Step 1: Collect the mechanical vibration signal of the equipment, and perform data preprocessing on the collected mechanical vibration signal through the data preprocessing module to obtain the preprocessed data sample , dividing the preprocessed mechanical vibration signal data into a training set and a test set, wherein the training set is entirely mechanical vibration signals of normal categories;

[0010] Step 2: The data samples preprocessed in step 1 The rotating machinery fault detection model is input. The mechanical vibration signal in each data sample has a time step dimension and a feature dimension. The mechanical vibration signal in the data sample first passes through the KAN network to increase the feature dimension, and then passes through the RevIN normalization module for normalization processing. The normalized data samples are respectively sent to two identical linear layers to increase the feature dimension again, and the positive dimension increase result is obtained. And the result of dimension upgrading , the forward dimension-raising result Used to feed the multi-scale convolution fusion modules in the forward and reverse paths, the dimension-upgraded results Used to make residual connection with the results of subsequent state space model time series modeling;

[0011] Step 3: Upgrade the result of the forward dimension in step 2 The time step convolution block and feature convolution block of the multi-scale convolution fusion module are respectively input for feature extraction, and the results of the parallel time step convolution block and feature convolution block are weighted fused in combination with the dynamic weight strategy to fuse the time step dimension and feature dimension to obtain the result of the multi-scale convolution fusion module. The forward path feature extraction process is as follows: the forward dimension upscaling result Directly input the time step convolution block and feature convolution block of the multi-scale convolution module for feature extraction, and pass the dynamic fusion weights in the multi-scale convolution fusion module on the forward path The results of the two parallel time-step convolution blocks and the feature convolution blocks are weighted fused to obtain the result of the multi-scale convolution fusion module on the forward path. The reverse path feature extraction process is: the forward dimension increase result Flip along the time step dimension, then feed into the time step convolution block and feature convolution block of the multi-scale convolution fusion module, and pass the dynamic fusion weights in the multi-scale convolution fusion module on the reverse path The results of the two parallel time-step convolution blocks and feature convolution blocks are weighted fused, and the fusion results are flipped again to obtain the results of the multi-scale convolution fusion module on the reverse path. ;

[0012] Step 4: Results of the multi-scale convolutional fusion module on the forward path Input the state space model of the forward path for time series modeling and compare the modeling results with the dimension-upgrading results in step 2 Make residual connections to obtain the state space model branch modeling results of the forward path and the results of the multi-scale convolution fusion module on the reverse path. Input the state space model of the reverse path for time series modeling and compare the modeling results with the dimension-upgrading results in step 2 Residual connection is made to obtain the reverse path state space model branch modeling result, and the dynamic weight strategy is used to perform weighted fusion on the forward path state space model branch modeling result and the reverse path state space model branch modeling result to obtain the dynamic fusion weight Bidirectional modeling results after weighted fusion;

[0013] Step 5: Dynamically fusion weights The bidirectional modeling results after weighted fusion are gradually reduced in dimension through the linear layer and KAN network, and denormalized through the RevIN normalization module to restore to the same dimension as the input data sample. Consistent dimensions, final output reconstructed samples ;

[0014] Step 6: Use mean square error as loss function to train the rotating machinery fault detection model. During the test, use the threshold derived from the Youden index as the fault judgment criterion. When the reconstructed sample output in step 5 is The input data sample of step 2 When the mean square error loss is greater than the threshold, it is judged as a fault. When the reconstructed sample output in step 5 The data sample from step 2 When the mean square error loss is less than the threshold, it is considered normal.

[0015] A further improvement of the present application is that: Step 2 converts the data sample preprocessed in Step 1 into The dimension is increased through the KAN network, and then normalized through the RevIN normalization module. The normalized result is further increased through the linear layer to obtain the positive dimension increase result. And the result of dimension upgrading The calculation formula is:

[0016]

[0017] in, To use the linear layer dimension increase operation, is the RevIN normalization operation, This is the operation of dimension increase using the KAN network.

[0018] A further improvement of the present application is that the feature convolution block in the multi-scale convolution fusion module is used to convolve the feature dimension, perform a one-dimensional convolution operation along the feature dimension, and the convolution kernel of the feature convolution block scans different features to explore the interaction mechanism between multi-dimensional features, and outputs the feature dimension convolution result as :

[0019]

[0020] Among them, when performing feature extraction on the forward path, The input data of the multi-scale convolution fusion module is the result of the forward dimension increase. , when doing feature extraction on the reverse path, The result of positive dimension upgrading Dimensionality increase result by flipping along the time step dimension , is the flipping operation along the time step dimension, It is the feature convolution operation;

[0021] The time step convolution block in the multi-scale convolution fusion module acts on the time step dimension and needs to exchange the input data of the multi-scale convolution fusion module. The feature dimension and time step dimension of the convolution operation are taken into account, and then a one-dimensional convolution operation is performed along the time step dimension. The local pattern of the features on the time step convolution block changing over time is extracted by sliding the convolution kernel of the time step convolution block on the time dimension. The final output time step dimension convolution result is :

[0022]

[0023] in, represents the operation of exchanging the time step dimension and the feature dimension, For the time-step convolution operation, when extracting features on the forward path, The input data of the multi-scale convolution fusion module is the result of the forward dimension increase. , when doing feature extraction on the reverse path, The result of positive dimension upgrading Dimensionality increase result of flipping along the time step dimension ;

[0024] The dynamic weight strategy is used to convolve the extracted feature dimensions. And the time step dimension convolution result Perform weighted summation to obtain the result of the multi-scale convolution fusion module on the forward path:

[0025]

[0026] The result of the multi-scale convolution fusion module on the reverse path is:

[0027]

[0028] in, is the activation function, is the dynamic fusion weight in the multi-scale convolution fusion module on the forward path, is the dynamic fusion weight in the multi-scale convolution fusion module on the reverse path, Represents a flipping operation along the time step dimension.

[0029] A further improvement of the present application is that step 4 specifically includes the following steps:

[0030] Step 4.1: Fusion of the results of the multi-scale convolutional module of the forward path The state space model of the input forward path is used for time series modeling to capture the dependency of the current state on the historical state. The modeling result is consistent with the dimensionality increase result of step 2. Make a residual connection:

[0031]

[0032] in, Modeling results for the forward path state-space model branch, Modeling operations for state-space models, is the activation function;

[0033] Step 4.2: Combine the results of the multi-scale convolutional fusion module on the reverse path The state space model of the input reverse path is used for time series modeling and compared with the dimensionality increase result Make a residual connection:

[0034]

[0035] in, Modeling results for the reverse path state space model branch;

[0036] Step 4.3: Use dynamic fusion weights Modeling results of the forward path state space model branch and reverse path state space model branch modeling results Perform weighted addition:

[0037]

[0038] in, is the dynamic fusion weight The bidirectional modeling results after weighted fusion, is the activation function.

[0039] A further improvement of the present application is that: in step 5, the reconstructed sample is finally output for:

[0040]

[0041] in, Represented as the denormalization operation of the RevIN normalization module, To use the KAN network for dimensionality reduction, Indicates the dimensionality reduction operation using a linear layer.

[0042] A further improvement of the present application is that in step 6, the mean square error is used as the loss function to train the rotating machinery fault detection model, which is expressed as:

[0043] in, After preprocessing, Mechanical vibration signal samples, represents the final output reconstruction sample, is the total number of data samples in the training set.

[0044] The beneficial effects of this application are as follows:

[0045] Through the self-supervised fault detection model based on signal reconstruction, this application can maintain high fault detection performance while having a relatively low training cost, solving the problems of the existing technology's dependence on labels and high detection cost in fault detection solutions.

[0046] By combining the multi-scale convolution fusion method, this application solves the problem that the existing technology ignores the multi-scale feature information of time series signals, and integrates different-dimensional information by combining the dynamic weight strategy, effectively improving the feature extraction ability of the model.

[0047] This application proposes an improved bidirectional state space model architecture, which solves the problem that the existing technology ignores the complex time series characteristics of vibration signals itself, improves the model's analysis ability of the time series characteristics of vibration signals, improves the model's reconstruction ability of normal signals, and further improves the accuracy of fault detection.

[0048] This application is more suitable for actual industrial scenarios. Brief Description of the Drawings

[0049] Figure 1 is the flowchart of the detection method of this application.

[0050] Figure 2 is the schematic diagram of the rotating machinery fault detection model of this application.

[0051] Figure 3 is the improved bidirectional state space model architecture diagram of this application.

[0052] Figure 4 is the loss scatter plot on the visualization test set of this application. Detailed Embodiments

[0053] The following will disclose the embodiments of the present invention in the form of diagrams. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.

[0054] Without conflict, the embodiments in this application and the features in the embodiments can be combined with each other; and all model trainings in the comparative experiment table are programmed using the Python language and completed on an Intel Xeon Gold 6230R processor and an NVIDIA GeForce RTX 3090 graphics card.

[0055] The following will disclose the embodiments of the present invention with drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.

[0056] like Figure 1 As shown, the present application is a rotating machinery fault detection method based on self-supervised learning, and the rotating machinery fault detection method is implemented by a rotating machinery fault detection model, such as Figure 2 As shown, the rotating machinery fault detection model includes a data preprocessing module, a Kolmogorov–Arnold Networks network, namely a KAN network, a ReversibleInstance Normalization module, namely a RevIN normalization module, a linear layer and an improved bidirectional state space model architecture, wherein the improved bidirectional state space model architecture includes a multi-scale convolution fusion module, a state space model and a dynamic fusion weight The multi-scale convolution fusion module includes a time-step convolution block, a feature convolution block and a dynamic weight strategy, and performs feature extraction on the forward path and the reverse path. Specifically, the rotating machinery fault detection method includes the following steps:

[0057] Step 1: Collect the mechanical vibration signal of the equipment, and perform data preprocessing on the collected mechanical vibration signal through the data preprocessing module to obtain the preprocessed data sample , that is, using a sliding window to segment the collected mechanical vibration signal, using a fast Fourier transform operation on the mechanical vibration signal in the window to map the mechanical vibration signal in the window from the time domain to the frequency domain, and normalizing the mechanical vibration signal, dividing the data of the normalized mechanical vibration signal into a training set and a test set, wherein all the training sets are mechanical vibration signals working normally;

[0058] Step 2: The data samples preprocessed in step 1 The rotating machinery fault detection model is input, and the mechanical vibration signal in each data sample has a time step dimension and a feature dimension. The mechanical vibration signal in the data sample first passes through the KAN network to increase the feature dimension, and then passes through the RevIN normalization module for normalization processing. The normalized data samples are respectively sent to two identical linear layers to increase the feature dimension again, and the forward dimension increase result fed to the multi-scale convolution fusion module is obtained. And the dimension-raising result used for residual connection with the result of subsequent state space model time series modeling , the calculation formula is:

[0059]

[0060] in, To use the linear layer dimension increase operation, is the RevIN normalization operation, This is the operation of dimension increase using the KAN network.

[0061] Step 3: Upgrade the result of the forward dimension in step 2 The time step convolution block and feature convolution block of the multi-scale convolution fusion module are respectively input for feature extraction, and the results of the parallel time step convolution block and feature convolution block are weighted fused in combination with the dynamic fusion strategy to fuse the time step dimension and feature dimension to obtain the result of the multi-scale convolution fusion module. The forward path feature extraction process is as follows: the forward dimension upscaling result Directly input the time step convolution block and feature convolution block of the multi-scale convolution module for feature extraction, and pass the dynamic fusion weights in the multi-scale convolution fusion module on the forward path The results of the two parallel time-step convolution blocks and the feature convolution blocks are weighted fused to obtain the result of the multi-scale convolution fusion module on the forward path. The reverse path feature extraction process is: the forward dimension increase result Flip along the time step dimension, then feed into the time step convolution block and feature convolution block of the multi-scale convolution fusion module, and pass the dynamic fusion weights in the multi-scale convolution fusion module on the reverse path The results of the two parallel time-step convolution blocks and feature convolution blocks are weighted fused, and the fusion results are flipped again to obtain the results of the multi-scale convolution fusion module on the reverse path. Specifically,

[0062] The feature convolution block in the multi-scale convolution fusion module is used to convolve the feature dimension and perform a one-dimensional convolution operation along the feature dimension. The convolution kernel of the feature convolution block scans different features to explore the interaction mechanism between multi-dimensional features and outputs the feature dimension convolution result as :

[0063]

[0064] Among them, when performing feature extraction on the forward path, The input data of the multi-scale convolution fusion module is the result of the forward dimension increase. , when doing feature extraction on the reverse path, The result of positive dimension upgrading The result after flipping along the time step dimension , is the flipping operation along the time step dimension, It is the feature convolution operation;

[0065] The time step convolution block in the multi-scale convolution fusion module acts on the time step dimension and needs to exchange the input data of the multi-scale convolution fusion module. The feature dimension and time step dimension of the convolution operation are taken into account, and then a one-dimensional convolution operation is performed along the time step dimension. The local pattern of the features on the time step convolution block changing over time is extracted by sliding the convolution kernel of the time step convolution block on the time dimension. The final output time step dimension convolution result is :

[0066]

[0067] in, represents the operation of exchanging the time step dimension and the feature dimension, For the time-step convolution operation, when extracting features on the forward path, The input data of the multi-scale convolution fusion module is the forward dimension increase result of the linear layer. , when doing feature extraction on the reverse path, The result of positive dimension upgrading Dimensionality increase result of flipping along the time step dimension ;

[0068] In order to effectively integrate the output representation of time-step convolution and feature dimension convolution, the extracted feature dimension convolution results are extracted through a dynamic weight strategy. And the time step dimension convolution result Perform weighted summation to obtain the result of the multi-scale convolution fusion module on the forward path:

[0069]

[0070] The result of the multi-scale convolution fusion module on the reverse path is:

[0071]

[0072] in, is the activation function, is the dynamic fusion weight in the multi-scale convolution fusion module on the forward path, is the dynamic fusion weight in the multi-scale convolution fusion module on the reverse path, Represents the flipping operation along the time step dimension. Dynamic fusion weights in the multi-scale convolution fusion module on the forward path Dynamic fusion weights in multi-scale convolutional fusion modules on the reverse path As a trainable parameter obtained by model training, the goal of weighted fusion is to adaptively balance the information contribution of the time dimension and the feature dimension according to task requirements.

[0073] Step 4: Results of the multi-scale convolutional fusion module on the forward path Input the state space model of the forward path for time series modeling and compare the modeling results with the dimension-upgrading results in step 2 Make residual connections to obtain the state space model branch modeling results of the forward path and the results of the multi-scale convolution fusion module on the reverse path. Input the state space model of the reverse path for time series modeling and compare the modeling results with the dimension-upgrading results in step 2 Make residual connections to obtain the reverse path state space model branch modeling results, and use the dynamic weight strategy to adjust the forward path state space model branch modeling results and reverse path state space model branch modeling results Perform weighted fusion to obtain the bidirectional modeling result after dynamic weight strategy weighted fusion The specific steps include:

[0074] Step 4.1: Fusion of the results of the multi-scale convolutional module of the forward path The state space model of the input forward path is used for time series modeling to capture the dependency of the current state on the historical state. The modeling result is consistent with the dimensionality increase result of step 2. Make a residual connection:

[0075]

[0076] in, Modeling results for the forward path state-space model branch, Modeling operations for state-space models, is the activation function;

[0077] Step 4.2: Combine the results of the multi-scale convolutional fusion module on the reverse path The state space model of the input reverse path is used for time series modeling and compared with the dimensionality increase result Make a residual connection:

[0078]

[0079] in, Modeling results for the reverse path state space model branch;

[0080] Step 4.3: Use dynamic fusion weights Modeling results of the forward path state space model branch and reverse path state space model branch modeling results Perform weighted addition:

[0081]

[0082] in, It is the result of bidirectional modeling after weighted fusion with dynamic fusion weights and is the activation function. For the activation function.

[0083] Step 5: The result of bidirectional modeling after weighted fusion with dynamic fusion weights is gradually reduced in dimension through a linear layer and a KAN network, and is denormalized through a RevIN normalization module to restore to the same dimension as the input data sample and finally the reconstructed sample is output :

[0084]

[0085] Among them, represents the denormalization operation of the RevIN normalization module, is the dimensionality reduction operation using the KAN network, represents the dimensionality reduction operation using the linear layer.

[0086] Step 6: Use the mean squared error as the loss function to train the rotating machinery fault detection model. During testing, use the threshold derived from the Youden index as the fault determination criterion. When the mean squared error loss between the reconstructed sample output in Step 5 and the input data sample in Step 2 is greater than the threshold, it is determined as a fault. When the mean squared error loss between the reconstructed sample output in Step 5 and the data sample in Step 2 is less than the threshold, it is determined as normal. Specifically, when using the mean squared error as the loss function to train the rotating machinery fault detection model, it is expressed as:

[0087] Among them, represents the th mechanical vibration signal sample after preprocessing, represents the finally output reconstructed sample, is the total number of data samples in the training set.

[0088] This application uses the rolling bearing fault data set of Jiangnan University to verify the effectiveness of the proposed self-supervised fault detection method. The fault data is provided based on a fault diagnosis experiment of a three-phase induction motor (Mitsubishi SB-JR). The motor has a rated power of 3.7 kW, and the speed range is controlled by adjusting the input voltage to 400 to 800 r / min; the data set covers vibration signals in four states: normal state, outer ring fault, inner ring fault and rolling element fault, and the fault is artificially created by wire cutting technology. The vibration data at a speed of 600 r / min was used for the experiment. The data used were transformed from the time domain to the frequency domain using a sliding window method with a length of 1024. After normalization and standardization, the number of randomly sampled training set samples was 1246, the number of test set samples was 438, the length of each sample was 512, and the feature dimension was 1. Among them, all training sets are normal samples, and the ratio of normal samples to abnormal samples in the test set is 1:1. Each category of abnormal samples is randomly sampled from each abnormal class, and the training set and test set samples do not overlap.

[0089] In Table 1, the method of this application is compared with the advanced methods in recent years. The results of these benchmark methods are obtained by re-experimenting based on the open source code. Table 1 shows the accuracy, precision, recall and F1 score of each method on the Jiangnan University dataset. Compared with other methods in the table, the method proposed in this application achieves the best or suboptimal performance in multiple indicators, proving the effectiveness of the proposed method.

[0090] Table 1 Comparative experimental results on the Jiangnan University dataset

[0091]

[0092] In addition, a visual loss scatter plot is drawn. Figure 4 In the figure, green points are normal signals, blue points are fault signals, and the horizontal dotted line represents the fault threshold of classification. Sample loss points above the threshold are fault samples, and those below the threshold are normal samples. Figure 4 It can be seen that the present application can reconstruct normal signal samples very well, the reconstruction loss of fault samples is relatively large, and the fault detection performance is good.

[0093] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A rotating machinery fault detection method based on self-supervised learning, characterized in that: The rotating machinery fault detection method is implemented by a rotating machinery fault detection model, wherein the rotating machinery fault detection model includes an improved bidirectional state space model architecture, wherein the improved bidirectional state space model architecture includes a multi-scale convolution fusion module, a state space model and a dynamic weight strategy. Specifically, the rotating machinery fault detection method includes the following steps: Step 1: Collect the mechanical vibration signal of the equipment and pre-process it to obtain data samples , the data sample Dividing into a training set and a test set, wherein the training set is all mechanical vibration signals of normal category; Step 2: Data samples The rotating machinery fault detection model is input, the mechanical vibration signal in the data sample is first increased in feature dimension, and then normalized. The normalized data samples are respectively sent to two identical linear layers to increase the feature dimension again, and the forward dimension increase result fed to the multi-scale convolution fusion module is obtained. And the dimension-raising result used for subsequent residual connection ; Step 3: The multi-scale convolution fusion module performs the forward dimension upgrade results in step 2 from the forward and reverse directions respectively. Perform feature extraction to obtain the results of the multi-scale convolution fusion module on the forward path and the results of the multi-scale convolution fusion module on the reverse path; Step 4: Perform time series modeling on the results of the multi-scale convolution fusion module on the forward path and the results of the multi-scale convolution fusion module on the reverse path, respectively, to obtain the forward path state space model branch modeling results and the reverse path state space model branch modeling results, and use a dynamic weight strategy to perform weighted fusion on the forward path state space model branch modeling results and the reverse path state space model branch modeling results to obtain a bidirectional modeling result; Step 5: Reduce the dimension of the bidirectional modeling results obtained in step 4, perform denormalization, and finally output the reconstructed samples. ; Step 6: Use mean square error as loss function to train the rotating machinery fault detection model. During testing, fault detection is achieved by comparing the mean square error loss of data samples with the fault threshold.

2. A rotating machinery fault detection method based on self-supervised learning according to claim 1, characterized in that: The rotating machinery fault detection model includes a data preprocessing module, a Kolmogorov–Arnold Networks network, namely a KAN network, a Reversible Instance Normalization module, namely a RevIN normalization module, a linear layer and an improved bidirectional state space model architecture, wherein the improved bidirectional state space model architecture includes a multi-scale convolution fusion module, a state space model and a dynamic fusion weight ,The multi-scale convolution fusion module includes a time-step convolution block and a feature convolution block and a dynamic weight strategy, and performs feature extraction on the forward path and the reverse path.

3. A rotating machinery fault detection method based on self-supervised learning according to claim 2, characterized in that: The step 2 is specifically: the data sample preprocessed in step 1 is The rotating machinery fault detection model is input, and the mechanical vibration signal in each data sample has a time step dimension and a feature dimension. The mechanical vibration signal in the data sample first passes through the KAN network to increase the feature dimension, and then passes through the RevIN normalization module for normalization processing. The normalized data samples are respectively sent to two identical linear layers to increase the feature dimension again, and the forward dimension increase result fed to the multi-scale convolution fusion module is obtained. And used to make residual connection dimensionality increase results with the results of subsequent state space model time series modeling :The calculation formula is: , in, To use the linear layer dimension increase operation, is the RevIN normalization operation, This is the operation of dimension increase using the KAN network.

4. A rotating machinery fault detection method based on self-supervised learning according to claim 2, characterized in that: Step 3: Upgrade the result of the forward dimension in step 2 The time step convolution block and feature convolution block of the multi-scale convolution fusion module are respectively input for feature extraction, and the results of the parallel time step convolution block and feature convolution block are weighted fused in combination with the dynamic weight strategy to fuse the time step dimension and feature dimension to obtain the result of the multi-scale convolution fusion module. The forward path feature extraction process is as follows: the forward dimension upscaling result Directly input the time step convolution block and feature convolution block of the multi-scale convolution module for feature extraction, and pass the dynamic fusion weights in the multi-scale convolution fusion module on the forward path The results of the two parallel time-step convolution blocks and the feature convolution blocks are weighted fused to obtain the result of the multi-scale convolution fusion module on the forward path. The reverse path feature extraction process is: the forward dimension increase result Flip along the time step dimension, then feed into the time step convolution block and feature convolution block of the multi-scale convolution fusion module, and pass the dynamic fusion weights in the multi-scale convolution fusion module on the reverse path The results of the two parallel time-step convolution blocks and feature convolution blocks are weighted fused, and the fusion results are flipped again to obtain the results of the multi-scale convolution fusion module on the reverse path. , specifically including the following steps: Step 3.1: The feature convolution block in the multi-scale convolution fusion module is used to convolve the feature dimension, perform a one-dimensional convolution operation along the feature dimension, and the convolution kernel of the feature convolution block scans different features to output the feature dimension convolution result. : , Among them, when performing feature extraction on the forward path, The input data of the multi-scale convolution fusion module is the result of the forward dimension increase. , when doing feature extraction on the reverse path, The result of positive dimension upgrading Dimensionality increase result by flipping along the time step dimension , is the flipping operation along the time step dimension, It is the feature convolution operation; Step 3.2: The time step convolution block in the multi-scale convolution fusion module acts on the time step dimension, and the input data of the multi-scale convolution fusion module needs to be exchanged. The feature dimension and time step dimension of the convolution operation are taken into account, and then a one-dimensional convolution operation is performed along the time step dimension. The local pattern of the features on the time step convolution block changing over time is extracted by sliding the convolution kernel of the time step convolution block on the time dimension. The final output time step dimension convolution result is : , in, represents the operation of exchanging the time step dimension and the feature dimension, For the time-step convolution operation, when extracting features on the forward path, The input data of the multi-scale convolution fusion module is the result of the forward dimension increase. , when doing feature extraction on the reverse path, The result of positive dimension upgrading Dimensionality increase result by flipping along the time step dimension ; Step 3.3: Use the dynamic weight strategy to convolve the extracted feature dimensions And the time step dimension convolution result Perform weighted summation to obtain the result of the multi-scale convolution fusion module on the forward path: , The result of the multi-scale convolution fusion module on the reverse path is: , in, is the activation function, is the dynamic fusion weight in the multi-scale convolution fusion module on the forward path, is the dynamic fusion weight in the multi-scale convolution fusion module on the reverse path, Represents a flipping operation along the time step dimension.

5. A rotating machinery fault detection method based on self-supervised learning according to claim 4, characterized in that: In step 4, the result of the multi-scale convolutional fusion module on the forward path is Input the state space model of the forward path for time series modeling and compare the modeling results with the dimension-upgrading results in step 2 Make residual connections to obtain the state space model branch modeling results of the forward path and the results of the multi-scale convolution fusion module on the reverse path. Input the state space model of the reverse path for time series modeling and compare the modeling results with the dimension-upgrading results in step 2 Make residual connections to obtain the reverse path state space model branch modeling results, use the dynamic weight strategy to weightedly fuse the forward path state space model branch modeling results and the reverse path state space model branch modeling results to obtain the dynamic fusion weight The bidirectional modeling result after weighted fusion specifically includes the following steps: Step 4.1: Fusion of the results of the multi-scale convolutional module in the forward path The state space model of the input forward path is used for time series modeling to capture the dependency of the current state on the historical state. The modeling result is consistent with the dimension-upgrading result of step 2. Make a residual connection: , in, Modeling results for the forward path state-space model branch, Modeling operations for state-space models, is the activation function; Step 4.2: Combine the results of the multi-scale convolutional fusion module on the reverse path The state space model of the input reverse path is used for time series modeling and compared with the dimensionality increase result Make a residual connection: , in, Modeling results for the reverse path state space model branch; Step 4.3: Use dynamic fusion weights Modeling results of the forward path state space model branch and reverse path state space model branch modeling results Perform weighted addition: , in, is the dynamic fusion weight The bidirectional modeling results after weighted fusion, is the activation function.

6. A rotating machinery fault detection method based on self-supervised learning according to claim 5, characterized in that: In step 5, the dynamic fusion weights The bidirectional modeling results after weighted fusion are gradually reduced in dimension through the linear layer and KAN network, and denormalized through the RevIN normalization module to restore to the same dimension as the input data sample. Consistent dimensions, final output reconstructed samples : , in, Represented as the denormalization operation of the RevIN normalization module, To use the KAN network for dimensionality reduction, Indicates the dimensionality reduction operation using a linear layer.

7. A rotating machinery fault detection method based on self-supervised learning according to claim 1, characterized in that: In step 6, the mean square error is used as the loss function to train the rotating machinery fault detection model. During the test, the threshold derived from the Youden index is used as the fault judgment criterion. When the reconstructed sample output in step 5 is The input data sample of step 2 When the mean square error loss is greater than the threshold, it is judged as a fault. When the reconstructed sample output in step 5 The data sample from step 2 When the mean square error loss is less than the threshold, it is judged to be normal. The mean square error is used as the loss function to train the rotating machinery fault detection model, which is expressed as: , in, After preprocessing, Mechanical vibration signal samples, represents the final output reconstruction sample, is the total number of data samples in the training set.

Citation Information

Patent Citations

  • Multi-scale feature fusion gearbox fault diagnosis method based on self-attention mechanism

    CN116010900A

  • Walking rail insulation fault positioning method based on multi-scale convolution fusion one-dimensional Transform

    CN118501629A

  • Leukemia diagnosis system based on multi-scale feature fusion

    CN118692654A

  • Rotary machine fault diagnosis method based on residual contraction convolution and attention mechanism

    CN119475191A

  • Deep learning based operational domain verification using camera-based inputs for autonomous systems and applications

    US20230177839A1