Rotating machine fault diagnosis method based on denoising probability diffusion model
By using the denoising probability diffusion model and one-dimensional ResUnet network in rotary machinery fault diagnosis, the problem of insufficient data samples is solved, and efficient fault diagnosis and accurate fault identification are achieved.
Patent Information
- Application Number
- CN202510140347.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has the problem of insufficient data samples in rotary machinery fault diagnosis, resulting in limited diagnostic accuracy and availability.
Using a method based on the denoising probability diffusion model, feature extraction and Gaussian noise prediction are performed through a one-dimensional ResUnet network, trained models are generated, and abnormal scores are calculated by Manhattan distance for fault diagnosis.
The early and accurate fault diagnosis is achieved without the need for a large number of complete data sets, improving diagnostic efficiency and performance.
Smart Images

Figure CN120123638A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rotating machinery fault diagnosis, and particularly relates to a rotating machinery fault diagnosis method based on a denoising probability diffusion model. Background Art
[0002] Currently, with the development of global manufacturing, intelligent manufacturing is the core of current industrial technology, and rotating machinery is an essential equipment among them. However, the operating conditions of rotating machinery directly affect product quality and production efficiency. In industrial production, if a fault occurs, it may cause serious consequences. Therefore, diagnosing the operating state of rotating machinery and timely discovering equipment operation problems are of great significance for industrial production.
[0003] Traditional rotating machinery fault diagnosis methods are mainly based on expert experience, requiring practical operation experience and a large amount of professional knowledge. At the same time, these methods also require sufficient samples to support the diagnosis process. If there is a lack of sufficient data samples, the accuracy and usability of these methods will be limited.
[0004] With the development of machine learning and deep learning technologies, data-driven unsupervised learning methods have become one of the effective ways to solve this problem. Machine learning methods such as neural networks and support vector machines have been widely used in fault diagnosis and achieved good results. However, in real engineering applications, there is often a lack of sufficient healthy state and fault samples to support the application of these methods, making their diagnostic effects difficult to meet actual needs. In this regard, domain adaptation (DA) can transfer the information of source domain data to the target domain to solve the problem of lacking labeled data in the target domain. However, in actual production, due to the lack of target domain fault samples, the model established by DA has a low accuracy for fault diagnosis in the target domain. As a generative model that has emerged in recent years, GAN has a strong ability to fit the real data distribution. It can perform unsupervised learning on existing data to generate data similar to real data. This method has further improvements in expanding the dataset and assisting in interpretation for fault diagnosis, and can improve the prediction accuracy in the case of insufficient sample size. However, currently, the samples generated by GAN lack diversity, have high requirements for the selected hyperparameters and regularizers, and have problems such as unstable training processes, which have certain limitations for practical applications.
[0005] Therefore, a rotating machinery fault diagnosis method based on a denoising probability diffusion model is urgently needed to be proposed. Summary of the Invention
[0006] To solve the defects existing in the prior art, the present invention provides a rotating machinery fault diagnosis method based on a denoising probability diffusion model.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] The present invention provides a method for diagnosing faults in rotating machinery based on a denoising probability diffusion model, comprising the following steps:
[0009] S1. After data preprocessing, Gaussian noise is added to obtain a one-dimensional vibration signal submerged in noise. Feature extraction is performed through a one-dimensional ResUnet network, the added Gaussian noise is predicted, the distribution column closest to the input data is found, and a trained model is generated;
[0010] S2. An unknown sample is input, denoising is performed through the trained model, the corresponding generated data is inferred, the two are compared, the anomaly score is calculated through the Manhattan distance, a threshold is set to discriminate the fault data, and fault diagnosis is performed.
[0011] Preferably, in step S1, the data preprocessing method is to collect vibration signals from the rotating machinery as the original data, perform normalization processing to ensure that all data has the same ratio and scale, and cut it into segments of a specific length during processing. The sample data within the cutting length contains fault features.
[0012] Preferably, the one-dimensional ResUnet network structure in step S1 includes multiple encoder and decoder blocks. Each block includes a convolutional layer, a batch normalization layer, and a Swish activation function. The formula of the activation function is:
[0013] y = x * sigmoid(x)
[0014] where x is the input data of the convolutional layer and y is the output data;
[0015] The decoder block further includes an upsampling layer for increasing the dimension of the feature map, and a skip connection is added between the matching encoder and decoder blocks.
[0016] Preferably, in step S1, the weights of the one-dimensional ResUnet are optimized using maximum likelihood. Given a set of training samples, the log-likelihood of the observed noise given the input data is maximized, that is:
[0017]
[0018] where N is the number of training samples, is the data observed in the training samples, is the corresponding Gaussian noise at the last time step.
[0019] Preferably, in the denoising process of step S2, Gaussian noise is added to the data in the time domain range, and the features and waveforms of the original data are restored in this way, so as to accurately identify and locate the fault data. The reverse inference noise accumulation is expressed as:
[0020]
[0021] where E is the abnormal deviation of the fault sample relative to the normal data sample, is the noise prediction value at each step, and T is the number of times Gaussian noise is added.
[0022] Preferably, in step S2, the Manhattan distance between the generated data and the input data is measured:
[0023]
[0024] where A = (a 1 , a 2 , …, a n ) and B = (b 1 , b 2 , …, b n ) are two N-dimensional vectors, and M_D(A, B) is the Manhattan distance between these two vectors;
[0025] In addition, calculate the distance after the corresponding output f of a certain feature layer of the one-dimensional ResUnet network, and weight the sum of the two as the anomaly score;
[0026] Ano_score = λ·M_D(G(X t ), X t ) + (1 - λ)M_D(f(G(X t ), f(X t ))
[0027] where λ is the weight; M_D(G(X t ), X t ) is the Manhattan distance between the test sample and the corresponding generated data; M_D(f(G(X t ), f(X t )) is the Manhattan distance between the output of the test sample feature layer and the output of the corresponding generated data feature layer; G(X t ) is the corresponding generated data; f(X t ) is the output of the corresponding feature layer; X t is the test sample.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The present invention learns the distribution of normal samples of mechanical vibration signals in the latent space in an unsupervised manner. By this method, the general distribution of the data is captured, and fault diagnosis can be achieved without a large amount of complete data sets. Through efficient data modeling, the early detection and accuracy of fault detection are improved. During the diagnosis process, the fault is identified by locating the generated data closest to the test sample in this latent space, and the overall diagnosis efficiency and performance are improved by optimizing the data processing and analysis process. Description of the Drawings
[0030] Figure 1 is the flowchart of a method for diagnosing faults in rotating machinery based on a denoising probability diffusion model according to the present invention;
[0031] Figure 2 is the structural diagram of a one-dimensional ResUnet network in a method for diagnosing faults in rotating machinery based on a denoising probability diffusion model according to the present invention. Specific Embodiments
[0032] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0033] As Figures 1 to 2 shown, this embodiment provides a method for diagnosing faults in rotating machinery based on a denoising probability diffusion model, including the following steps:
[0034] S1. Training the model: After data preprocessing, Gaussian noise is added to obtain one-dimensional vibration signals submerged by noise. Feature extraction is performed through a one-dimensional ResUnet network, the added Gaussian noise is predicted, the distribution column closest to the input data is found, and a trained model is generated.
[0035] S2. Input the sample to be tested, denoise it through the trained model, infer the corresponding generated data, compare the two, calculate the anomaly score through the Manhattan distance, set a threshold to discriminate the fault data, and perform fault diagnosis.
[0036] In this embodiment, the data preprocessing method in step S1 is to collect vibration signals from the rotating machinery as the original data, perform normalization processing to ensure that all data have the same ratio and scale, and cut it into segments of a specific length during processing so that unified comparison and processing can be carried out in subsequent analysis. The sample data within the cut length contains fault features to ensure the accuracy and reliability of the experiment.
[0037] During the training process, the diffusion model uses noise addition to submerge the training data in random noise. This serves two main purposes: one is to increase the diversity of the data, enabling the model to better learn the patterns and features in the data; the other is to improve the robustness and reliability of the model, enabling the model to better adapt to different actual situations.
[0038] In this embodiment, the one-dimensional ResUnet network structure in step S1 includes multiple encoder and decoder blocks. Each block includes a convolutional layer, a batch normalization layer, and a Swish activation function. The formula for the activation function is:
[0039] y = x * sigmoid(x)
[0040] where x is the input data of the convolutional layer and y is the output data;
[0041] The decoder block also includes an upsampling layer for increasing the dimension of the feature map, and a skip connection is added between the matching encoder and decoder blocks. The structure diagram of the one-dimensional ResUnet is as Figure 2 shown.
[0042] Among them, the input and output data are one-dimensional vibration signals. Through the one-dimensional ResUnet network, multi-layer convolution is performed to extract data features, and multi-layer transposed convolution and upsampling are performed to restore the size of the original data and achieve the output of the data.
[0043] In this embodiment, the one-dimensional ResUnett network is used to learn the parametric mapping f θ : from the observed data X T to the corresponding Gaussian noise at the last time step
[0044]
[0045] where f θ is a one-dimensional ResUnet with parameters θ.
[0046] During training, the weights of the one-dimensional ResUnet are optimized using maximum likelihood in step S1. Given a set of training samples, the log-likelihood of the observed noise given the input data is maximized, i.e.:
[0047]
[0048] where N is the number of training samples, is the data observed in the training samples, is the corresponding Gaussian noise at the last time step.
[0049] Input X during the training process 0 is a normal one-dimensional vibration signal. Through T times of noise addition, time series data containing noise is obtained to implement the diffusion model. X is obtained by adding T times of Gaussian noise 0 . In the ideal result, when T approaches infinity, X 0 is pure Gaussian noise. However, considering the actual program, T can only be set large enough and can be approximately regarded as Gaussian noise
[0050] In this embodiment, test samples with different health conditions are used as the input of the model in step S2, and the possible Gaussian noise in the inference is gradually denoised to obtain the corresponding generated samples. The denoising process is performed on the data within the time domain, and Gaussian noise is added to it. In this way, the characteristics and waveforms of the original data are restored, so as to accurately identify and locate the fault data
[0051] Input X 0 , X t is the result obtained by adding t times of Gaussian noise. In the reverse inference, the parameters of the neural network are also trained to obtain p θ (x t-1 |x t ), and the added Gaussian noise is inferred to restore the normal time series data. The cumulative sum of the reverse inference noise is expressed as:
[0052]
[0053] where E is the abnormal deviation of the fault sample relative to the normal data sample, is the noise prediction value at each step, and T is the number of times of Gaussian noise addition. In step S2, the Manhattan distance between the generated data and the input data is measured:
[0054]
[0055] where A = (a 1 , a 2 , …, a n ) and B = (b 1 , b 2 , …, b n ) are two N-dimensional vectors, and M_D(A, B) is the Manhattan distance between these two vectors
[0056] In addition, the distance after calculating the corresponding output f of a certain feature layer of the one-dimensional ResUnet network is calculated, and the weighted sum of the two is used as the anomaly score
[0057] Ano_score = λ · M_D(G(X t ), X t)+(1-λ)M_D(f(G(X t )),f(X t ))
[0058] where λ is the weight; M_D(G(X t ),X t ) is the Manhattan distance between the test sample and the corresponding generated data; M_D(f(G(X t )),f(X t )) is the Manhattan distance between the output of the feature layer of the test sample and the output of the feature layer of the corresponding generated data; G(X t ) is the corresponding generated data; f(X t ) is the output of a corresponding feature layer; X t is the test sample.
[0059] The present invention is applicable to the field of bearing fault diagnosis. During the reverse denoising process, the test data containing positive and negative samples is used to replace the randomly sampled Gaussian noise X t . E can be regarded as the abnormal deviation of the fault sample relative to the normal data sample, and what is generated is the data in the latent space that is closest to the input sample of the test sample X t , thus realizing fault diagnosis and sample generation. Since the diffusion model is obtained by adding noise to X 0 to obtain X T , the dimensions of the two are the same.
[0060] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A rotating machinery fault diagnosis method based on a denoising probability diffusion model, characterized in that: The following steps are involved: S1. After data preprocessing, Gaussian noise is added to obtain the one-dimensional vibration signal submerged by noise. The one-dimensional ResUnet network is used to extract features, predict the added Gaussian noise, find the distribution column closest to the input data, and generate a trained model. S2. Input the sample to be tested, denoise it through the trained model, infer the corresponding generated data, compare the two, calculate the anomaly score through Manhattan distance, set the threshold to identify the fault data, and perform fault diagnosis.
2. The rotating machinery fault diagnosis method based on the denoising probability diffusion model according to claim 1 is characterized in that: The data preprocessing method in step S1 is to collect vibration signals from the rotating machinery as raw data, perform normalization processing to ensure that all data have the same proportion and scale, and cut them into segments of specific lengths during processing, and the sample data within the cut length contains fault characteristics.
3. The rotating machinery fault diagnosis method based on the denoising probability diffusion model according to claim 1 is characterized in that: The one-dimensional ResUnet network structure in step S1 includes multiple encoder and decoder blocks, each block includes a convolution layer, a batch normalization layer and a Swish activation function, and the activation function formula is: y=x*sigmoid(x) Among them, x is the input data of the convolution layer, and t is the output data; The decoder block also includes an upsampling layer to increase the dimensionality of the feature map, and skip connections are added between the matching encoder and decoder blocks.
4. The rotating machinery fault diagnosis method based on the denoising probability diffusion model according to claim 1 is characterized in that: In step S1, the weights of the one-dimensional ResUnet are optimized using maximum likelihood, given a set of training samples, and the logarithmic likelihood of the noise observed under given input data conditions is maximized, that is: Where N is the number of training samples, is the data observed in the training sample, is the corresponding Gaussian noise of the last time step.
5. The rotating machinery fault diagnosis method based on denoising probability diffusion model according to claim 1, characterized in that: The denoising process in step S2 adds Gaussian noise to the data in the time domain, thereby restoring the characteristics and waveform of the original data, thereby achieving accurate identification and positioning of the fault data.
6. The rotating machinery fault diagnosis method based on denoising probability diffusion model according to claim 1 is characterized in that: In step S2, the Manhattan distance between the generated data and the input data is measured: Where A=(a1,a2,…,a n ) and B=(b1,b2,…,b n ) are two N-dimensional vectors, M_D(A,B) is the Manhattan distance between the two vectors; In addition, the distance after the corresponding output f of a feature layer of the one-dimensional ResUnet network is calculated, and the weighted sum of the two is taken as the anomaly score; Ano_score=λ·M_D(G(X t ),X t )+1-λM_D(f(G(X t )),f(X t )) Where λ is the weight; M_D(G(X t ),X t ) is the Manhattan distance between the test sample and the corresponding generated data; M_D(f(G(X t )),f(X t )) is the Manhattan distance between the feature layer output of the test sample and the feature layer output of the corresponding generated data; G(X t ) is the corresponding generated data; f(X t ) is the corresponding feature layer output; X t For test samples.