Vibration signal adaptive denoising method based on global local feature mask denoising

By using the global local feature mask denoising method in vibration signal denoising, the convolution module of prime number cavitation coefficient sequence and the mixed Transformer extract feature, the problem of insufficient adaptability in the prior art is solved, and more efficient vibration signal denoising and diagnostic effects are achieved.

CN120011718APending Publication Date: 2025-05-16NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510487689.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing deep intelligent denoising networks are not adaptable enough in vibrating signal denoising, making it difficult to effectively remove common noise in all frequency bands and narrowband non-Gaussian interference, and have poor signal fidelity generalization capabilities.

Method used

The global local feature mask denoising method is adopted to extract multi-scale features through the convolution module of prime number hole coefficient sequence, and combined with a hybrid Transformer to extract local and global timing features to form a new mask for denoising.

Benefits of technology

It significantly improves the denoising accuracy and diagnostic effect of vibration signals, especially in low signal-to-noise ratio environments, and the diagnostic accuracy is increased by 10 to 30%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011718A_ABST
    Figure CN120011718A_ABST
Patent Text Reader

Abstract

The invention discloses a vibration signal self-adaptive denoising method for global local feature mask denoising, which comprises the following steps of: 1, segmenting a signal into frames through preprocessing, raising dimensions by using stacking, and converting a one-dimensional signal into a two-dimensional time domain signal; secondly, a convolution module with a prime number void coefficient sequence is built, and fault pulse scale changes are processed; and a third step of adaptively learning high and low frequency noise through a mixed Transform and designing a novel mask for denoising. And 4, constructing interpretable time-frequency domain joint constraints. And 5, introducing an overlapping addition strategy, recovering the processed signal into a one-dimensional signal, and outputting the one-dimensional signal to obtain a de-noised signal. According to the time domain two-dimensional denoising and recovery method for the one-dimensional signal, denoising parameters do not need to be adjusted independently, vibration signal fidelity denoising under the low signal-to-noise ratio can be achieved, and spectrum analysis and intelligent fault diagnosis can be effectively assisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of mechanical fault diagnosis, and in particular to a vibration signal adaptive denoising method for global and local feature mask denoising. Background Art

[0002] Vibration signals are one of the most widely used and direct data sources for mechanical health status analysis. However, in actual measurements, various types of equipment have electromagnetic interference and various types of noise, which makes it easy for the vibration characteristics related to the health status of the equipment in the collected vibration signals to be masked. Therefore, vibration signal denoising has always been an indispensable key part of vibration signal analysis.

[0003] Vibration signal denoising methods based on signal processing mainly include filtering, wavelet denoising and time-frequency analysis. Filtering mainly removes specific interference frequency bands; wavelet transform methods can better handle non-stationary vibration signals and can effectively distinguish the useful components of the signal from the noise; time-frequency analysis such as EMD, VMD and its derivative methods are suitable for situations where the signal is nonlinear and non-stationary, and can provide richer signal feature information. The optimal denoising method varies for different equipment vibration signals, different noise characteristics, and different noise levels, and expert experience is required to select. Researchers have also improved the adaptability of signal processing-based denoising methods by integrating various indicators such as kurtosis value indicators and hyperparameter optimization methods such as PSO.

[0004] Denoising methods based on signal processing can effectively remove noise, but they still require a lot of expert experience when dealing with complex noise, and their adaptability to different types of noise in complex environments is still insufficient. Deep learning can provide an end-to-end learning method, directly learning the final denoised signal from the original signal, and can also model environmental uncertainty by introducing noise characteristics, forms and other factors to improve the robustness and reliability of the model. Therefore, embedding the denoising strategy into the deep learning framework, integrating the advantages of the deep learning framework and the denoising strategy, has better accuracy and adaptability. For example, some studies have adopted a denoising model based on CNN, extracting features from the signal by constructing multiple convolutional layers, and reconstructing the denoised signal using deconvolution layers or upsampling layers. This method has achieved good results in processing one-dimensional vibration signals, and is generally based on generative network architectures such as VAE, DAE, CAE, etc. In addition, some studies have applied RNN or long short-term memory networks (LSTM) to vibration signal denoising. These network structures can capture the timing information in the signal and have good effects on processing continuous vibration signals.

[0005] Although deep networks have made some progress in the field of vibration signal denoising, there are still some shortcomings. (1) The existing deep intelligent denoising network does not have the ability to learn large-scale features, making it difficult to ensure that the denoising process fully learns the fault features of each frequency band, and it is difficult to effectively remove common noise and narrow-band non-Gaussian interference in the entire frequency band, so the signal fidelity generalization ability is poor. (2) The vibration signal is a time series signal. The generation mechanism determines that its short-range and long-range time series characteristics coexist. It requires a more effective time series relationship extraction method other than RNN and LSTM. There are very few studies on the existing Transformer-based denoising network, and the applicability of the existing Transformer in the periodic impact characteristics of vibration signals needs further study. (3) The poor solvability of the denoising process in the existing generation method is also one of the reasons why the intelligent denoising network is rarely used in practice. If the relevant analysis experience in the time domain and frequency domain analysis of vibration signals can be incorporated as an optimization constraint or objective function in the denoising process, the interpretability of the intelligent denoising process and the signal fidelity effect will be further improved. Summary of the invention

[0006] Purpose of the invention: In view of the poor adaptive denoising ability of existing denoising methods, the present invention proposes a vibration signal adaptive denoising method with global and local feature mask denoising to improve the denoising accuracy.

[0007] Technical solution: To achieve the purpose of the present invention, the technical solution adopted by the present invention is: a vibration signal adaptive denoising method of global local feature mask denoising, comprising the following steps:

[0008] Step 1: Data preprocessing: Noise the original data to obtain a data set containing noise-clean paired samples, split the data into frames, and then use stacking to process the data into a tensor of dimension three as the input of the encoder;

[0009] Step 2: Data encoding: Considering the multi-scale characteristics of the vibration signal, a convolution module with a prime number hole coefficient sequence is built as an encoder to process the scale changes of the fault pulse under variable speed and fault degree, and output the feature matrix;

[0010] Step 3: Noise masking: construct a hybrid Transformer to extract local and global features of two-dimensional time domain signals, adaptively learn high and low frequency noise and form a new mask for denoising;

[0011] Step 4: Signal reconstruction: The decoder has the same structure as the encoder, and the decoder is used to reconstruct the original signal sample from the feature matrix obtained after mask denoising;

[0012] Step 5, overlap-addition: Use overlap-addition to generate the denoised vibration signal waveform, and perform reverse restoration on the reconstructed signal according to the position coding of the signal sample slice into the fragment frame process in step 1;

[0013] Step 6: Model training: Using the vibration signal as input and the full-band denoising result as output, the adaptive denoising network constructed in steps 1 to 5 is trained using a data set containing noise-clean paired samples;

[0014] Step 7: Signal denoising: Input the vibration signal to be denoised into the trained adaptive denoising network to remove noise in the entire frequency band.

[0015] Furthermore, the data preprocessing in step 1 includes:

[0016] For length The original one-dimensional signal Cut to length , the jump size is Frames;

[0017] The processed frames are then stacked and output in the shape of The tensor is used as the input of the encoder, where is the number of channels, is the number of frames, is the length of the frame, The calculation formula is:

[0018] ;

[0019] In the data preprocessing stage, the original signal is divided into several time segments through overlapping blocks equipped with an overlapping frame strategy. A preset proportion of overlapping areas is retained between each segment. The pure data blocks extracted from the original uncontaminated signal serve as the reference target of the denoising system.

[0020] Further, the encoder in step 2 is composed of convolution blocks;

[0021] First, the encoder input data passes through a 1×1 convolution block and is normalized. Then, a dense dilated convolution block is constructed to form receptive fields of various scales. Four parallel dilated convolutions use different prime dilation rate sequences as dilation coefficients to obtain multi-scale features.

[0022] Finally, the single frame dimension is reduced through the convolution block, the layer is normalized and activated by the PReLU function. The PReLU formula is as follows:

[0023] ,

[0024] in, It is the result of applying nonlinear activation to feature data through PReLU. is the feature data processed by the encoder, It is an independent learnable parameter for each feature channel, which is used to adapt to the different nonlinear requirements of each channel.

[0025] Furthermore, step three constructs a hybrid Transformer for feature extraction, including:

[0026] First, two hybrid Transformer modules containing local Transformer and global Transformer are cascaded;

[0027] The Transformer adopts a simplified Transformer including multi-head attention blocks, layer normalization, GRU layers, residual connections and activation functions;

[0028] Among them, the local Transformer module is applied to a single data frame of the input, from the dimension Parallel processing of short-term context information within a fragment; the global Transformer module is applied to all fragments within the signal sample, that is, from the dimension Execute on to learn global dependencies;

[0029] Then, a noise masking strategy that integrates global and local temporal relationships is constructed through a convolution block with PReLU as the activation function. In this process, the number of feature channels C remains unchanged to form a mask matrix.

[0030] Finally, the Hadamard product is performed on the feature matrix output in step 2 and the mask matrix extracted above to implement the mask operation, that is, to obtain the denoised feature matrix.

[0031] Further, the signal reorganization in step 4 includes:

[0032] First, information is extracted through a dense hole convolution block of the same shape as the encoder, and then a deconvolution layer is used to enhance its resolution;

[0033] Next, the outputs of the first two layers are normalized and PReLU nonlinear operations are performed, and a 1×1 two-dimensional convolution is used to adjust the number of channels to 1, so that the output of each sample return .

[0034] Furthermore, in the optimization process of step 5, the overall loss function is measured in the time domain by the mean square error between the denoised reconstructed signal and the corresponding original signal, and in the frequency domain by the spectrogram difference. The loss function formula is as follows:

[0035] ,

[0036] ,

[0037] ,

[0038] in, and are the loss functions of the time dimension and the time-frequency domain dimension respectively, is the total loss function, is the original input signal, is the reconstructed signal after denoising; is the number of signal samples, is the representation of the real original signal in the time domain, To represent the denoised reconstructed signal in the time domain, and are the spectrum of the original signal and the spectrum of the denoised reconstructed signal, respectively. and denote the real and imaginary parts of the complex variable, respectively. Represents the original signal spectrum in the time frame ,frequency The real part of Represents the original signal spectrum in the time frame ,frequency The imaginary part of represents the real part of the denoised reconstructed signal spectrum, Represents the imaginary part of the denoised and reconstructed signal spectrum; Indicates the number of time frames, Indicates the frequency number, is the weight parameter;

[0039] Add additional feature loss to the loss function ,in is the input matrix obtained after data preprocessing for each sample, is the reconstruction matrix of the signal recombinant output;

[0040] The final loss function is ,in is the weight parameter.

[0041] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:

[0042] In view of the problem that vibration signals naturally have multi-scale key features, the present invention builds a prime core sequence dense hole convolution module to handle the change in the receptive field required for fault pulses under different speeds and fault levels, and constructs a hybrid transformer that fits the vibration signal to effectively extract local and global timing features, forming a new mask denoising method that effectively ensures high and low frequency noise learning. Experiments were conducted using open source Case Western Reserve University test data. Under low signal-to-noise conditions such as SNR=-10, -5, the denoising network can significantly improve the diagnostic effect, with diagnostic accuracy improved by 10~30%. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0044] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0045] The vibration signal adaptive denoising method of the present invention using global and local feature mask denoising is as follows: Figure 1 As shown, the following steps are included:

[0046] Step 1: Data preprocessing. Noise the input raw data to obtain a data set containing noise-clean paired samples. Segment the signal samples by overlapping segments to obtain a three-dimensional tensor as the encoder input. Suppose the input one-dimensional raw signal Length is , which is split into lengths in this step , the jump size is Frames, and transform the input data by stacking, resulting in a three-dimensional tensor As the input of the encoder, the number of channels Set to 1, is the number of frames, is the frame length. The calculation formula is:

[0047] (1)

[0048] At the same time, in the data preprocessing stage, the original signal is divided into several time segments through overlapping blocks equipped with overlapping frame strategies. A preset ratio of overlapping areas is retained between each segment to reduce the boundary effect generated when the signal is processed in blocks and ensure the continuity of the subsequent reconstructed signal. The pure data blocks extracted from the original uncontaminated signal serve as the ideal reference target for the denoising system.

[0049] Step 2: Data encoding. The encoder is mainly composed of convolution blocks. First, the convolution block expands the number of channels through a 1×1 convolution block and normalizes it through layer normalization; then, a dense hole convolution block is built to form receptive fields of various scales to adapt to the extraction of fault features of different cycles. The four parallel hole convolutions use prime number sequences 1, 2, 3, and 5 as expansion coefficients to reduce the chessboard effect; finally, the convolution block is used to reduce the single frame dimension, layer normalization is performed, and the PReLU function is used for activation. The PReLU formula is as follows:

[0050] (2)

[0051] in, It is the result of applying nonlinear activation to feature data through PReLU. is the feature data processed by the encoder, For input, It is an independent learnable parameter for each feature channel, which is used to better adapt to the different nonlinear requirements of each channel.

[0052] Step 3: Noise mask. The hybrid Transformer module includes a local Transformer and a global Transformer module. The present invention introduces a simplified Transformer model. The key fault shock in the vibration signal is periodic, and the position encoding part has a poor effect in the experiment. Therefore, the simplified Transformer used in the present invention only includes the key parts of multi-head attention block, layer normalization, GRU layer, residual connection and activation function, which are used to deal with the timing characteristics of the vibration signal. The core formula of the self-attention mechanism is as follows.

[0053] (3)

[0054] Among them, Q, K, and V represent query, key, and value matrices respectively. is the dimension of the key vector, which is used for scaling to stabilize the softmax function.

[0055] In multi-head attention, equation (3) is performed h times (h is the number of heads), and each head has an independent weight matrix, thereby mapping the input vector to h different subspaces. For each subspace, the self-attention mechanism is applied to calculate the attention weight, and the weighted sum value matrix V is added accordingly to obtain the output of each head , as shown in formula (4). Among them, , , are the Q, K, and V transformation matrices of the i-th head respectively.

[0056] (4)

[0057] Next, the outputs of all heads are concatenated, and then linearly transformed and layer normalized to obtain the output of multi-head attention, as shown in formula (5). This step integrates the information learned from different subspaces to enhance the model's expressiveness. is the output transformation matrix, and Concat represents the concatenation operation along the dimension.

[0058] (5)

[0059] In the present invention, the local Transformer module is applied to the single data frame obtained by segmenting the features output by the encoder, from the dimension Parallel processing of short-term context information within a fragment; the global Transformer module is further applied to all fragments in the sample, that is, from the dimension The global dependency is learned by performing the above operations, similar to learning long-range correlation features such as fault ridges on the time-frequency graph. In this step, two hybrid Transformer modules are cascaded first, and then passed through a convolutional block with PReLU as the activation function. During this process, the number of feature channels C remains unchanged, thus forming a mask matrix.

[0060] Finally, the feature matrix output in the second step is subjected to Hadamard product with the mask matrix extracted above to implement the mask operation, and the denoised feature matrix can be obtained.

[0061] Step 4: Signal reconstruction. The decoder uses a similar structure to the encoder to reconstruct the original noisy signal samples from the feature matrix output after mask denoising. First, information is extracted through a dense hole convolution block of the same shape as the encoder, and then the deconvolution layer is used to enhance its resolution. Next, the outputs of the first two layers are layer normalized and PReLU nonlinear operations are performed, and a 1×1 two-dimensional convolution is used to adjust the number of channels to 1, so that the reconstructed sample of each output return .

[0062] Step 5: Overlap-addition. Overlap-addition is used to generate the denoised vibration signal waveform. The main idea is to reversely restore the reconstructed signal based on the position coding of the signal sample slice into fragment frames in the first step. , , (3 vectors) are taken from the same position of the original signal samples, and they are summed and averaged to represent the signal of the same area in the reconstructed signal, that is, the fragment in the reconstructed signal .

[0063] During the optimization process, the overall loss function is measured in the time domain by the mean square error between the denoised reconstructed signal and the corresponding original signal, and in the frequency domain by the spectrogram difference. The loss function is as follows:

[0064] (6)

[0065] (7)

[0066] (8)

[0067] in, and are the loss functions of the time dimension and the time-frequency domain dimension respectively, is the total loss function, is the original input signal, is the reconstructed signal after denoising; is the number of signal samples, is the representation of the real original signal in the time domain, To represent the denoised reconstructed signal in the time domain, and are the spectrum of the original signal and the spectrum of the denoised reconstructed signal, respectively. and denote the real and imaginary parts of the complex variable, respectively. Represents the original signal spectrum in the time frame ,frequency The real part of Represents the original signal spectrum in the time frame ,frequency The imaginary part of represents the real part of the denoised reconstructed signal spectrum, Represents the imaginary part of the denoised and reconstructed signal spectrum; Indicates the number of time frames, Indicates the frequency number, is the weight parameter.

[0068] In addition, considering the signal fragments for the above-mentioned similar constraints can roughly constrain the fidelity characteristics of the signal fragments from the spectrum profile and improve the sample diversity. Therefore, additional feature loss is added to the loss function. ,in For each sample, the input matrix is ​​obtained after the first step. is the corresponding reconstruction matrix output by the fourth step.

[0069] The final objective function is , where the weight parameter It is set to 0.1 in the experiments.

[0070] Step 6: Model training: Using the vibration signal as input and the full-band denoising result as output, use a data set containing noise-clean paired samples to train the adaptive denoising network constructed in steps 1 to 5.

[0071] Step 7: Signal denoising. The intelligent denoising method proposed in the present invention is used to process the signal to be denoised, and the noise of the entire frequency band can be removed. Among them, the effect of removing the noise of the entire frequency band is the most prominent.

Claims

1. A vibration signal adaptive denoising method based on global and local feature mask denoising, characterized in that: The following steps are involved: Step 1: Data preprocessing: Noise the original data to obtain a data set containing noise-clean paired samples, split the data into frames, and then use stacking to process the data into a tensor of dimension three as the input of the encoder; Step 2: Data encoding: Considering the multi-scale characteristics of the vibration signal, a convolution module with a prime number hole coefficient sequence is built as an encoder to process the scale changes of the fault pulse under variable speed and fault degree, and output the feature matrix; Step 3: Noise masking: construct a hybrid Transformer to extract local and global features of two-dimensional time domain signals, adaptively learn high and low frequency noise and form a new mask for denoising; Step 4: Signal reconstruction: The decoder has the same structure as the encoder, and the decoder is used to reconstruct the original signal sample from the feature matrix obtained after mask denoising; Step 5, overlap-addition: Use overlap-addition to generate the denoised vibration signal waveform, and perform reverse restoration on the reconstructed signal according to the position coding of the signal sample slice into the fragment frame process in step 1; Step 6: Model training: Using the vibration signal as input and the full-band denoising result as output, the adaptive denoising network constructed in steps 1 to 5 is trained using a data set containing noise-clean paired samples; Step 7: Signal denoising: Input the vibration signal to be denoised into the trained adaptive denoising network to remove noise in the entire frequency band.

2. The vibration signal adaptive denoising method of global and local feature mask denoising according to claim 1 is characterized in that: Data preprocessing in step 1 includes: For length The original one-dimensional signal Cut to length , the jump size is Frames; The processed frames are then stacked and output in the shape of The tensor is used as the input of the encoder, where is the number of channels, is the number of frames, is the length of the frame, The calculation formula is: ; In the data preprocessing stage, the original signal is divided into several time segments through overlapping blocks equipped with an overlapping frame strategy. A preset proportion of overlapping areas is retained between each segment. The pure data blocks extracted from the original uncontaminated signal serve as the reference target of the denoising system.

3. The vibration signal adaptive denoising method of global and local feature mask denoising according to claim 1 is characterized in that: The encoder described in step 2 is composed of convolution blocks; First, the encoder input data passes through a 1×1 convolution block and is normalized. Then, a dense dilated convolution block is constructed to form receptive fields of various scales. Four parallel dilated convolutions use different prime dilation rate sequences as dilation coefficients to obtain multi-scale features. Finally, the single frame dimension is reduced through the convolution block, the layer is normalized and activated by the PReLU function. The PReLU formula is as follows: , in, It is the result of applying nonlinear activation to feature data through PReLU. is the feature data processed by the encoder, It is an independent learnable parameter for each feature channel, which is used to adapt to the different nonlinear requirements of each channel.

4. The vibration signal adaptive denoising method of global and local feature mask denoising according to claim 2 is characterized in that: Step 3: Build a hybrid Transformer for feature extraction, including: First, two hybrid Transformer modules containing local Transformer and global Transformer are cascaded; The Transformer adopts a simplified Transformer including multi-head attention blocks, layer normalization, GRU layers, residual connections and activation functions; Among them, the local Transformer module is applied to a single data frame of the input, from the dimension Parallel processing of short-term context information within a fragment; the global Transformer module is applied to all fragments within the signal sample, that is, from the dimension Execute on to learn global dependencies; Then, a noise masking strategy that integrates global and local temporal relationships is constructed through a convolution block with PReLU as the activation function. In this process, the number of feature channels C remains unchanged to form a mask matrix. Finally, the Hadamard product is performed on the feature matrix output in step 2 and the mask matrix extracted above to implement the mask operation, that is, to obtain the denoised feature matrix.

5. The vibration signal adaptive denoising method of global and local feature mask denoising according to claim 2 or 4, characterized in that: The signal reorganization in step 4 includes: First, information is extracted through a dense hole convolution block of the same shape as the encoder, and then a deconvolution layer is used to enhance its resolution; Next, the outputs of the first two layers are normalized and PReLU nonlinear operations are performed, and a 1×1 two-dimensional convolution is used to adjust the number of channels to 1, so that the output of each sample return .

6. A vibration signal adaptive denoising method using global and local feature mask denoising according to any one of claims 1 to 4, characterized in that: Step 5 During the optimization process, the overall loss function is measured in the time domain by the mean square error between the denoised reconstructed signal and the corresponding original signal, and in the frequency domain by the spectrogram difference. The loss function formula is as follows: , , , in, and are the loss functions of the time dimension and the time-frequency domain dimension respectively, is the total loss function, is the original input signal, is the reconstructed signal after denoising; is the number of signal samples, is the representation of the real original signal in the time domain, To represent the denoised reconstructed signal in the time domain, and are the spectrum of the original signal and the spectrum of the denoised reconstructed signal, respectively. and denote the real and imaginary parts of the complex variable, respectively. Represents the original signal spectrum in the time frame ,frequency The real part of Represents the original signal spectrum in the time frame ,frequency The imaginary part of represents the real part of the denoised reconstructed signal spectrum, Represents the imaginary part of the denoised and reconstructed signal spectrum; Indicates the number of time frames, Indicates the frequency number, is the weight parameter; Add additional feature loss to the loss function ,in is the input matrix obtained after data preprocessing for each sample, is the reconstruction matrix of the signal recombinant output; The final loss function is ,in is the weight parameter.

Citation Information

Patent Citations

  • Real image self-supervision denoising method and system

    CN116862779A

  • Feature extraction method and apparatus based on time domain and frequency domain of speech signal, and echo cancellation method and apparatus

    WO2023044962A1

Cited By

  • Unmanned vehicle high-reliability communication system and signal guarantee method thereof

    CN120857165A