Deep Denoising Method for Vibration Signals Based on CNN Structure with STFT Time-Frequency Domain Feature Extraction
Through STFT time-frequency domain feature extraction and CNN network, the problem of signal noise reduction dependence on prior knowledge in traditional methods is solved, and the vibration signal in structural health monitoring is automated to improve the noise removal effect.
Patent Information
- Application Number
- CN202211056214.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-31
AI Technical Summary
Traditional signal noise reduction methods rely on signal prior knowledge and artificial parameter settings, making it difficult to effectively deal with multivariate heterogeneous vibration signals in structural health monitoring, and deep learning lacks robustness and reliability in non-stationary signal noise reduction.
STFT time-frequency domain feature extraction and CNN deep convolution network are adopted to embed noise label settings to build mapping relationships to realize automated denoising of vibration signals and get rid of the dependence on signal prior knowledge.
It realizes automated intelligent noise reduction for different types of vibration signals, improves the noise denoising effect, and is suitable for signal preprocessing and engineering applications in structural health monitoring.
Smart Images

Figure CN115524404B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vibration signal noise reduction in structural health monitoring, and particularly relates to a CNN vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction. Background Art
[0002] In structural health monitoring, data-driven analysis based on vibration signals is one of the important ways to obtain the key states of structures. Due to the interference of the environment and system errors, the structural vibration signals collected by sensors will inevitably be affected by noise, which brings great obstacles to subsequent related analysis applications based on vibration signals (such as damage detection, modal identification, etc.). As a basic prerequisite for data-driven analysis, the research on related methods has important engineering significance for carrying out data-driven analysis and realizing structural health monitoring.
[0003] Most traditional signal noise reduction methods based on numerical filtering rely on prior knowledge of signals or noise and the artificial selection of parameters during the noise reduction process, that is, analyze the time-frequency information of the noise in the vibration signal to determine the parameters of the noise reduction physical model, and it is difficult to process multivariate heterogeneous vibration data in the field of structural health monitoring where the signal noise distribution cannot be obtained.
[0004] Convolutional neural networks have achieved extensive research applications in the field of signal processing due to their powerful feature extraction capabilities. However, due to the non-stationarity of structural vibration signals and the uncertainty of noise distribution, it brings great difficulties and obstacles to the application of deep learning technology in the field of structural vibration signal noise reduction. On the one hand, since the true value of the vibration signal cannot be obtained, it is difficult to set the training data labels; on the other hand, for the feature extraction of one-dimensional vibration signals, there is often a possibility of destroying the original features of the signal, resulting in low robustness and reliability of the trained model. Therefore, the present invention proposes a CNN deep noise reduction network based on STFT time-frequency domain feature extraction to realize the automatic intelligent denoising of structural vibration signals. Summary of the Invention
[0005] The purpose of the present invention is to propose a CNN structural vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction. This method does not need to rely on prior knowledge of the target signal or artificial parameter setting during the noise reduction process, can realize automatic noise reduction for different types of vibration signals (such as acceleration, strain, displacement, etc.), and its noise reduction effect can be improved by optimizing the model training set, and can be used for signal preprocessing in the field of vibration signal analysis research and embedded in the structural health system in engineering practice.
[0006] To achieve the purpose of the present invention, the present invention is realized through the following technical solutions:
[0007] A deep noise reduction method for vibration signals based on CNN structure with STFT time-frequency domain feature extraction, characterized by the following steps:
[0008] S1. Use a sensor to collect the structural vibration response signal and divide it into time series {Signal original} of appropriate length according to the sampling frequency to obtain the original signal;
[0009] S2. Embed different types and levels of noise (white noise and pink noise) into the original signal in S1 to obtain different types of mixed signals {Signal synthetic} to obtain one-dimensional time series signals;
[0010] S3. Use STFT (Short-Time Fourier Transform) to transform the one-dimensional time series signal in S2 and the original signal in S1 into two-dimensional spectrograms Extract high-order features of vibration signals through data augmentation processing;
[0011] S4. Use the two-dimensional spectrogram of the mixed signal as the input signal, and the two-dimensional spectrogram of the original signal as the target signal, so as to set corresponding labels for the noise and provide it to S5;
[0012] S5. Build a CNN deep convolutional network. The building idea is to build a mapping relationship f between and and then generate a denoised spectrogram to realize the prediction of clean signals; Specifically, the CNN deep convolutional network is repeatedly composed of multiple groups of convolutional layers and pooling layers. The specific architecture includes an input layer - convolutional layer - pooling layer - convolutional layer - pooling layer - fully connected layer - output layer in sequence. Among them, the convolutional layer is ConV2D (two-dimensional convolutional layer), the pooling layer is MaxPooling (maximum pooling), the activation function during training is RELU, and the loss function is MSE (root mean square error);
[0013] S6. Obtain the trained CNN deep convolutional network model, convert the measured noisy vibration signal into a spectrogram and then use it as the input for network testing / prediction, and output the denoised spectrogram; Use the inverse STFT to reconstruct the denoised spectrum into a one-dimensional time series signal, that is, obtain the denoised vibration signal.
[0014] The deep noise reduction method for vibration signals based on CNN structure with STFT time-frequency domain feature extraction is further characterized in that in S5 and S6, for the results of network training and testing, the signal-to-noise ratio SNR and root mean square error RMSE are used for evaluation, that is
[0015]
[0016]
[0017] where n represents the number of observation data, and E i represents the test data value, and P i represents the network prediction value.
[0018] The deep noise reduction method for vibration signals based on the CNN structure with STFT time-frequency domain feature extraction is characterized in that in S2, white noise is embedded according to the following formula:
[0019] Signal synthetic = Signal original + Noise × N l (8)
[0020] where Noise is a normal distribution random vector with a mean of zero and a standard deviation the same as the original signal, and N l is the noise level.
[0021] Different from the uniform distribution of white noise in different frequency bands, the power spectral density PSD of pink noise is inversely proportional to the frequency f, and is defined as follows:
[0022]
[0023] The deep noise reduction method for vibration signals based on the CNN structure with STFT time-frequency domain feature extraction is characterized in that in S3, the one-dimensional time series signal is transformed into a two-dimensional spectrogram by using STFT, which can effectively extract the high-order features of the signal. At the same time, to reduce the influence of spectral leakage on signal reconstruction, the original signal is windowed (multiplied by a window function such as the Hamming window) and then DFT (discrete Fourier transform) is performed during the STFT process.
[0024] The deep noise reduction method for vibration signals based on the CNN structure with STFT time-frequency domain feature extraction is characterized in that in S4, considering that the noise distribution of the structural vibration signal lacks prior conditions and the true value of the clean signal cannot be obtained, the training data cannot be directly calibrated. By adding noise to the original signal, the mixed signal is used as the input and the original signal is used as the output to build the mapping relationship between the noisy signal and the original signal, and then the data labels of the network input and output are set.
[0025] The deep noise reduction method for vibration signals based on the CNN structure with STFT time-frequency domain feature extraction is characterized in that in S5, an autoencoder-decoder is used to extract the mapping relationship between the noisy spectrogram and the clean spectrogram, and the purpose is to build and the mapping relationship f, and then generate the denoised spectrogram to predict the clean signal.
[0026] The CNN - based deep noise reduction method for vibration signals with STFT time - frequency domain feature extraction is characterized in that in S6, the iterative optimization objective of the mapping relationship f is to reduce the error function
[0027]
[0028] The CNN - based deep noise reduction method for vibration signals with STFT time - frequency domain feature extraction is characterized in that in S3 or S6, the one - dimensional time - series signal is processed by STFT and then transformed into a two - dimensional spectrogram, realizing the effective extraction of high - order features of vibration signals. At the same time, the data signal is transformed into an image signal and input into the CNN network for training and testing, making full use of the superior performance of the CNN network in extracting image features and further improving the effectiveness of the denoising network.
[0029] The parts not involved in this invention are applicable to the prior art.
[0030] The principle and beneficial effects of this invention are as follows: Using STFT to increase the dimension of one - dimensional time - series signals to two - dimensional spectrograms, realizing the effective extraction of high - order features of vibration signals; adopting the method of noise embedding to successfully set labels for non - stationary vibration signals whose clean signal true values cannot be obtained, providing a basis for network training; using the convolutional neural network CNN to denoise vibration signals, getting rid of the dependence on the prior knowledge of signals and the need for artificial setting of parameters in traditional denoising methods, and realizing the automatic and intelligent denoising of vibration signals.
[0031] This invention does not require human intervention during the noise reduction process and can be used for the automatic noise reduction of large - volume multi - heterogeneous data in the field of structural health monitoring, having great engineering application prospects in the development and embedding of structural health monitoring systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is the overall flow chart of the method of this invention;
[0033] Figure 2 is the STFT schematic diagram of the embodiment;
[0034] Figure 3 is the data processing and label setting of the training and test sets of the embodiment of this invention;
[0035] Figure 4 is the schematic diagram of the CNN deep convolutional network proposed by this invention;
[0036] Figure 5 is the signal model test flow chart of the embodiment of this invention DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Taking the acceleration response signal of a certain high-rise building structure under environmental excitation as an example, the technical solution of the present invention will be described.
[0038] The design idea of the present invention is as follows: converting a one-dimensional vibration signal into a two-dimensional spectrogram through STFT to extract its high-order features; realizing the label setting of the input signal and the target signal by embedding noise in the original vibration signal; building a CNN deep convolutional network to extract the mapping relationship between the noisy spectrum and the clean spectrum, and the model can realize the automatic denoising of vibration signals after being trained and optimized.
[0039] A deep noise reduction method for vibration signals based on the CNN structure with STFT time-frequency domain feature extraction, as Figure 1 shown, includes the following steps:
[0040] S1. Use an acceleration sensor to collect the acceleration time history signal of a certain high-rise building structure under environmental excitation, and divide the original signal into one-dimensional time series {Signal original}.
[0041] S2. Add different types and levels of noise to the original signal. In the embodiments of the present invention, white noise and pink noise with levels of 20%, 40%, 60%, and 80% are added to obtain different types of mixed signals {Signal synthetic}, where the white noise is embedded according to the following formula:
[0042] Signal synthetic = Signal original + Noise × N l (11)
[0043] In the formula, Noise is a normal distribution random vector with a mean of zero and a standard deviation the same as that of the original signal, and N l is the noise level.
[0044] Different from the uniform distribution of white noise in different frequency bands, the power spectral density PSD of pink noise is inversely proportional to the frequency f, and is defined as follows:
[0045]
[0046] S3. Use STFT (Short-Time Fourier Transform) to convert the one-dimensional time series signal into a two-dimensional spectrogram, and extract the high-order features of the vibration signal through data dimensionality increase processing. In this process, to reduce the influence of spectral leakage on signal reconstruction, the original signal is windowed with a Hamming window and then DFT (Discrete Fourier Transform) is performed. The Hamming window length is 256 and the frame shift is 64, as Figure 2 shown.
[0047] S4. Use the spectrogram of the mixed signal as the input signal and the spectrogram of the original signal as the target signal, and set the corresponding labels, such as Figure 3 .
[0048] S5. Build a CNN deep convolutional network (such as Figure 4 ), where the encoder is composed of multiple groups of convolution, batch normalization, max pooling, and RELU activation layers repeated, and the decoder is composed of convolution, batch normalization, and upsampling layers repeated. The CNN network parameters are shown in the following table:
[0049] Table 1. CNN Deep Convolutional Network Structure
[0050]
[0051] Total parameters: 32,933
[0052] Training parameters: 32,933
[0053] S6. The idea of building the noise reduction network is to build the mapping relationship f with , and then generate the denoised spectrogram to achieve the prediction of the clean signal. The iterative optimization goal of the mapping relationship f is to reduce the prediction error
[0054]
[0055] S7. Use the signal-to-noise ratio SNR and root mean square error RMSE to evaluate the results of network training and testing, and optimize the network accordingly, that is
[0056]
[0057]
[0058] where n represents the number of observed data, E i represents the test data value, and P i represents the network prediction value.
[0059] S8. Obtain the optimized network model according to the above method, convert the measured noisy vibration signal into a spectrogram and use it as the input for network testing, and output the denoised spectrogram. Use the inverse STFT transformation to reconstruct the denoised spectrogram into a one-dimensional time series signal, that is, obtain the noise-reduced vibration signal.
[0060] In the embodiment, the structural acceleration response time history collected from a certain building structure under three different typhoons is selected, as shown in Table 2. Use the above technical solution as the input for model testing, and the testing process is as Figure 5As shown, save the test results of the corresponding signal, i.e., the signal after noise reduction.
[0061] Table 2. Vibration signals of building structures under typhoon excitation
[0062]
[0063]
[0064] Calculate the SNR and RMSE of the test signals under each working condition, and find their average values for result denoising evaluation, as shown in Table 3. The parameter analysis of the test results shows that the signal-to-noise ratio of the denoised signal is effectively improved, and the root mean square error between the predicted signal and the clean signal is also significantly reduced. This embodiment demonstrates the effectiveness of the denoising method proposed by the present invention.
[0065] Table 3. Parameter analysis of vibration signal denoising results
[0066]
[0067] The above embodiments describe the basic principles, main features and technical routes of the present invention, and also show the effectiveness and superiority of the present invention in realizing vibration signal denoising. It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive, and do not limit the patent scope of the present invention accordingly. All those directly or indirectly changed by those skilled in the art according to the technical solution of the present invention are equally included in the patent protection scope of the present invention. The protection scope claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A deep noise reduction method for vibration signals based on the CNN structure with STFT time-frequency domain feature extraction, characterized in that, Including the following steps: S1. Use sensors to collect the structural vibration response signals and divide them into time series of appropriate lengths according to the sampling frequency to obtain the original signals; S2. Embed noises of different types and levels into the original signal described in S1 to obtain mixed signals of different types, resulting in one-dimensional time series signals. , obtaining one-dimensional time series signals. S3. Use the Short-Time Fourier Transform (STFT) to transform the one-dimensional time series signal described in S2 and the original signal described in S1 into two-dimensional spectrograms respectively , , and extract the high-order features of the vibration signal through data dimensionality augmentation processing S4. Use the two-dimensional spectrogram of the mixed signal as the input signal, and use the two-dimensional spectrogram of the original signal as the target signal, so as to set corresponding labels for the noise and provide them to S5; S5. Build a CNN deep convolutional network. The idea is to build and The mapping relationship f , and then generate the denoised spectrum Realize the prediction of clean signals; in specific implementation, the CNN deep convolutional network is composed of multiple groups of convolutional layers and pooling layers, and the specific architecture includes input layer-convolutional layer-pooling layer-convolutional layer-pooling layer-fully connected layer-output layer in sequence, wherein the convolutional layer is ConV2D, the pooling layer is MaxPooling, the activation function is RELU during training, and the loss function is MSE; S6. Obtain the trained CNN deep convolutional network model. After converting the measured noisy vibration signal into a spectrogram, use it as the input for network testing / prediction, and output the denoised spectrogram. Use the inverse STFT to reconstruct the denoised spectrum into a one-dimensional time series signal, that is, obtain the vibration signal after noise reduction.
2. The CNN structure vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction according to claim 1, characterized in that Furthermore, for the results of network training and testing in S5 and S6, the signal-to-noise ratio SNR and root mean square error RMSE are used for evaluation, that is (1) (2) In the formula n represents the number of observed data E i represents the test data value P i represents the network prediction value 3. The CNN structure vibration signal deep denoising method based on STFT time-frequency domain feature extraction according to claim 1, characterized in that, In S2, white noise is embedded according to the following formula: (3) where is a Gaussian distributed random vector with zero mean and the same standard deviation as the original signal, is the noise level; Unlike white noise which is uniformly distributed across different frequency bands, the power spectral density (PSD) of pink noise is inversely proportional to the frequency f and is defined as follows: (4) 。 4. The CNN structure vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction according to claim 1, characterized in that In S3, use STFT to convert the one-dimensional time series signal into a two-dimensional spectrogram to effectively extract the high-order features of the signal. At the same time, to reduce the impact of spectral leakage on signal reconstruction, after windowing the original signal in the STFT process, perform the discrete Fourier transform DFT.
5. The CNN structure vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction according to claim 1, characterized in that In S4, by adding noise to the original signal, use the mixed signal as the input and the original signal as the output to establish the mapping relationship between the noisy signal and the original signal, and then realize the setting of data labels for network input and output.
6. The CNN structure vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction according to claim 1, characterized in that In S5, an autoencoder-decoder is used to extract the mapping relationship between the noisy spectrogram and the clean spectrogram, aiming to establish the and mapping relationship f , and then generate the denoised spectrogram to predict the clean signal.
7. The CNN structure vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction according to claim 1, characterized in that In S6, the iterative optimization objective of the mapping relationship f is to minimize the error function (5)。 8. The CNN structure vibration signal deep noise reduction method based on STFT time-frequency domain feature extraction according to claim 1, characterized in that, In S3 or S6, after STFT processing of the one-dimensional time series signal, it is converted into a two-dimensional spectrogram, which effectively extracts the high-order features of the vibration signal. At the same time, converting the data signal into an image signal and inputting it into the CNN network for training and testing makes full use of the superior performance of the CNN network in extracting image features, further improving the effectiveness of the denoising network.
Citation Information
Patent Citations
Single-channel real-time noise reduction method based on convolutional recurrent neural network
CN109841226A
Phase-dependent shared deep convolutional neural network speech enhancement method
CN111081268A