Concrete filled steel tube void knocking acoustic detection method suitable for low signal-to-noise ratio environment
Through the stacked convolutional noise reduction autoencoder model combined with CNN and LSTM networks, the multimodal characteristics of steel pipe concrete detection signals are extracted, solving the accuracy problem of the existing technology in the low signal-to-noise ratio environment, and achieving high-precision de-emphasis defect recognition.
Patent Information
- Application Number
- CN202510046918.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The existing knock acoustic methods are difficult to effectively remove non-stationary and time-varying noise in low signal-to-noise environments, resulting in distortion of detection data and inaccurate defect identification.
The stacked convolution noise reduction autoencoder model is used to combine CNN and LSTM networks to extract frequency and time domain information, fuse to form multimodal features, build damage indicators, and achieve high-precision identification of steel pipe concrete detachment defects.
In a low signal-to-noise ratio environment, the accuracy and reliability of the detection data are significantly improved, the ability to capture timing characteristics is enhanced, the signal processing process is simplified, and the application difficulty of actual engineering detection is increased.
Smart Images

Figure CN119988901A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of defect detection based on acoustics, and in particular to a steel tube concrete void knocking acoustic detection method suitable for low signal-to-noise ratio environment. Background Art
[0002] CFST void detection can timely detect and repair void defects that may lead to reduced bearing capacity, accelerated corrosion and reduced durability, thereby avoiding potential structural failure and increased maintenance costs, which is crucial to ensuring structural safety and extending service life. The percussion acoustic method is widely used in the detection of large CFST arch bridges, CFST columns and other structures due to its advantages such as simplicity, low cost and no damage to the structure. However, in actual engineering inspections, environmental noise such as wind noise and traffic noise often interferes with the inspection process, significantly reducing the accuracy and reliability of the inspection data.
[0003] When the existing percussion acoustic method is applied to the degassing detection of steel tube concrete structures, filter noise reduction is used as a common signal processing method to reduce the interference of environmental noise on the detection data. However, since the filter design is usually based on the assumption of stationary noise, it is difficult for them to effectively remove non-stationary and time-varying noise, which may not only cause distortion of the detection data, but also may mistakenly delete useful signal features, thereby affecting the accurate identification of defects. Domestic and foreign scholars have proposed a series of methods such as wavelet transform, empirical mode decomposition, and variational mode decomposition to address this problem. Although good results have been achieved, these methods often require a lot of signal processing knowledge to determine parameters when facing different signals, have poor adaptability, and cannot effectively achieve end-to-end noise signal removal, which makes it difficult to apply them in actual engineering detection. Therefore, it is necessary to design a noise reduction and steel tube concrete degassing detection method based on percussion acoustics for application in actual engineering. Summary of the invention
[0004] The purpose of the present invention is to provide a steel tube concrete hollowing percussion acoustic detection method suitable for low signal-to-noise ratio environments, to solve the technical problem that the existing concrete hollowing prediction method has poor adaptability, which makes it difficult to apply in actual engineering detection. This method performs targeted noise reduction for the environmental noise encountered in actual engineering detection. At the same time, CNN and LSTM are used to extract frequency domain and time domain information, which are fused to form multi-modal features, build damage indicators, and then identify the location of steel tube concrete hollowing defects with high precision.
[0005] The method includes: collecting noise signals from actual engineering sites, combining them with simulated noise to produce multiple noise data sets, collecting clean knocking signals in a quiet environment, using clean signals superimposed on noise signals as training sets, and training to obtain a stacked convolutional denoising autoencoder model; knocking and collecting audio signals on the surface of steel tube concrete in an actual engineering environment, using the trained stacked convolutional denoising autoencoder model to perform denoising on the signals to be identified; and using a CNN-LSTM model to extract the time-frequency characteristics of the denoised signals to complete the identification of hollow defects.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A method for reducing noise of a steel tube concrete void detection signal autoencoder suitable for a low signal-to-noise ratio environment, characterized in that the method comprises the following steps:
[0008] S1: In a quiet environment, the steel tube concrete structure is struck at a fixed frequency and intensity, and the generated audio data is collected and saved using a microphone;
[0009] S2: Automatically pick up the starting point of the knock signal, using the Akaike information criterion to find the starting point of the knock audio, and save the audio of each knock as a separate file;
[0010] S3: Noise signal production. In order to restore the noise at the actual engineering site, the actual noise at the site is recorded and combined with the noise at the simulated site to expand the data set as the noise source;
[0011] S4: Dataset preparation, randomly adding noise sources to the clean audio files extracted from S2 to obtain signal data as a training set;
[0012] S5: Network structure construction, using a stacked denoising autoencoder model. By stacking several denoising autoencoder layers, the stack gradually learns higher-level feature representations and recovers clean signals from noisy inputs.
[0013] S6: Adjust hyperparameters and conduct training; hyperparameter settings include: learning rate, number of training iterations, number of batches, and optimizer parameters;
[0014] S7: Check whether the loss function of the validation set is gradually converging, and use the signal-to-noise ratio and root mean square error after denoising to evaluate the performance of the model;
[0015] S8: If the performance of the two indicators in S7 meets the requirements, the training ends; if not, return to S7 to continue training, and if it meets the requirements, the training ends.
[0016] Furthermore, in S3, the noise sources include wind noise, car noise, and bird noise, which are mixed in different intensity ratios to produce 8 types of noise data, namely noise 1 to noise 8.
[0017] Further, in S4, one or more of noises 1 to 8 are randomly added to the clean audio file, and at the same time, a random scaling factor is applied to change the intensity of noises 1 to 8 and randomly added into the sample to obtain signal data, the signal data is used as a training set, and the denoised audio data is used as a test set;
[0018] Furthermore, in S5, the detailed architecture of the convolutional denoising autoencoder is as follows: input->convolutional layer 1->BN+PReLU->max pooling layer 1->convolutional layer 2->BN+PReLU->max pooling layer 2->BN+PReLU->encoder LSTM layer->latent connection layer->decoder LSTM layer->deconvolutional layer 1->BN+PReLU->deconvolutional layer 2->BN+PReLU->deconvolutional layer 2->Tanh->output; the stacked convolutional denoising autoencoder takes the output of the encoder in a trained convolutional denoising autoencoder as the input of the next autoencoder.
[0019] Furthermore, in step 5, an encoder LSTM layer and a decoder LSTM layer are respectively added to the convolutional denoising autoencoder. The encoder LSTM layer in the convolutional denoising autoencoder improves the model's ability to capture temporal features and enhances the denoising effect by modeling the time dependency of the audio.
[0020] Furthermore, the two index evaluation model formulas in S7 are as follows:
[0021]
[0022] Among them: x is the original signal, is the noise reduction signal, E[x 2 ] is the energy of the original signal, is the energy of the noise signal, that is, the difference between the original signal and the noise-reduced signal.
[0023] Furthermore, the signal-to-noise ratio of the signal added with noise in S4 is maintained at 0-25 dB.
[0024] A steel tube concrete hollow percussion acoustic detection method suitable for a low signal-to-noise ratio environment, the method comprising the following steps:
[0025] SS1: The audio signal after noise reduction in claim 1 is labeled with "noise removed" and "noise not removed" as the training set, and a part of it is divided as the validation set to verify the model capability;
[0026] SS2: PSD signal is calculated using the Welch method. First, the signal is divided into multiple segments and windowed. The FFT windowed signal segments are transformed and the power is calculated. Then, the results of all segments are averaged to obtain the PSD estimation image of the entire signal.
[0027] SS3: Feature weighted splicing: Use convolutional neural networks to extract the frequency domain features expressed by each audio PSD signal, use long short-term memory networks to directly process each audio signal, extract the time domain features, and then perform weighted splicing;
[0028] SS4: After the time-frequency features are concatenated, the synthesized feature vector is used for classification tasks. The data with and without gaps are mixed as a test set, and the classification task is completed using the classification layer at the end of the model.
[0029] Furthermore, the PSD estimation image in SS2 is specifically expressed as:
[0030]
[0031] Among them, K is the total number of data segments, x i [n] is the i-th segment after windowing, w[n] is the window function, N w is the total energy of the window function.
[0032] Furthermore, the performance of weighted splicing in SS3 is:
[0033] x fused =[α·x CNN , β·x LSTM ] (4)
[0034] x CNN represents the frequency domain feature vector extracted by CNN, x LSTM represents the time domain feature vector extracted by LSTM, α and β are the weights of CNN and LSTM features respectively, and are the adjusted hyperparameters.
[0035] The present invention has the following beneficial effects due to the adoption of the above technical solution:
[0036] The present invention fully considers the noise conditions in the actual detection of existing projects, uses the noise collected by itself to make a noise data set, and randomly adds it. Compared with adding Gaussian white noise for training, the model has better robustness. At the same time, compared with the noise reduction technology of variational mode decomposition, the present invention does not require complex signal processing knowledge, uses stacked convolution noise reduction autoencoders for adaptive noise reduction, and adds LSTM to improve the model's ability to capture the temporal characteristics of audio signals, which has a broader application prospect for actual engineering detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flow chart summarizing the whole method;
[0038] Figure 2 Schematic diagram for finding the starting point;
[0039] Figure 3 A flowchart of a denoising method based on a stacked denoising autoencoder;
[0040] Figure 4 This is a flowchart of the feature extraction and classification method based on CNN and LSTM;
[0041] Figure 5 The denoising principle and layer-by-layer training diagram of the stacked denoising autoencoder;
[0042] Figure 6 This is the model architecture of a convolutional denoising autoencoder;
[0043] Figure 7 For cast steel tube concrete specimens;
[0044] Figure 8 This is the measurement point layout diagram of the steel tube concrete specimen with void defects;
[0045] Fig. 9 This is a comparison chart of the time domain curve and PSD curve before and after noise reduction;
[0046] Fig.10 The audio before and after noise reduction is Figure 4 Comparison of prediction accuracy on the steel pipe shown. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and preferred embodiments. However, it should be noted that many details listed in the specification are only for the purpose of enabling the reader to have a thorough understanding of one or more aspects of the present invention, and these aspects of the present invention can be implemented even without these specific details.
[0048] Combination Figure 1-Figure 10 As shown, the present invention provides a steel tube concrete void percussion acoustic detection method suitable for low signal-to-noise ratio environment, which is divided into an acoustic noise reduction method, including eight steps S1-S8. A steel tube concrete void prediction method, including four steps SS1-SS4. In order to demonstrate the embodiment of our proposed method, we use low-density plastic foam to simulate the void defect and set Figure 8 The four void defects with different thicknesses at different locations are shown and cured at room temperature for 28 days.
[0049] In S2, after the starting point is automatically picked up, the time for intercepting each audio segment is guaranteed to be 0.2s, and the sampling rate is 48000hz, that is, 9600 sampling points.
[0050] In S3, the recorded wind noise and vehicle noise are from actual bridge field recordings, and the animal noise is from a public data set. After considering the actual situation, the distribution of the eight noises is as follows:
[0051]
[0052] In S4, the denoising principle of the denoising autoencoder is as follows: Figure 5 As shown in Figure 2. In our method, the input audio data is added with our random noise, and the autoencoder training goal is to minimize the difference between the input noisy data and the autoencoder output. The loss function is expressed as
[0053]
[0054] where y i is the original clean data, is the output of the autoencoder, and n is the total number of samples.
[0055] In S4, the convolutional denoising autoencoder model architecture is as follows Figure 6 As shown in the figure, an additional LSTM layer is added to the model for audio signals, and the time-dependent modeling capability is introduced, so that the model can more effectively capture the temporal characteristics of audio signals, and thus better distinguish and remove noise. This structure enables the denoising autoencoder to extract not only spatial features but also temporal features, which is very helpful in improving the audio denoising effect.
[0056] In S4, a total of 4 encoders and 4 decoders are stacked. Greedy training is performed layer by layer by stacking four denoising autoencoders. This method trains the network layer by layer, training only one layer at a time, and using the output of the previous layer as the input of the next layer. After the training is completed, the training data is passed through the encoder part of this layer to generate a new compressed representation, which will serve as the input of the next layer. Each subsequent new layer uses the output of the previous layer as input and is trained independently until all layers have been trained. Finally, the backpropagation algorithm is used to adjust the weights of all layers to optimize the performance of the entire network.
[0057] The PReLU activation function is used at the end of each encoder and decoder. Since PReLU has learnable parameters, it can better adapt to different data characteristics, thereby improving the model's fitting ability and performance. Batch normalization (BN) technology is used to improve the speed and stability of neural network training.
[0058] The hyperparameter settings during this training process are: initial learning rate 0.005, maximum number of training rounds 100, the ratio of training set to validation set 8:2, and batch processing number 4.
[0059] After training is completed, check whether the loss function of the validation set is gradually converging, and use the signal-to-noise ratio and root mean square error after denoising to evaluate the model performance. The average SNR after denoising should reach more than 20dB, and the RMSE should remain at a low level, which is considered the end of training. If not, adjust the hyperparameters and continue training.
[0060] Reference for time domain effect and frequency domain effect after noise reduction Fig. 9 As shown, it can be seen that the noise reduction effect of this method is better.
[0061] According to the technical solution provided by the present invention, after noise reduction is completed, a method for predicting steel tube concrete voids based on CNN and LSTM is provided. The main purpose of using convolutional neural network (CNN) to process the spectrogram is to use its powerful spatial feature extraction ability to identify frequency patterns. The LSTM model directly processes the original audio signal, extracts features, and then performs weighted splicing. The weight is regarded as a hyperparameter. Since voids are extremely sensitive to frequency, they are initially set to 8:2. Other hyperparameters include training rounds and initial learning rate, which are set to 50 and 0.001 respectively.
[0062] In the SS2, a PSD image is used instead of a traditional spectrum diagram, which is more suitable for non-periodic signals such as audio.
[0063] In the SS4, the Softmax function is used in the last layer of the network to complete the empty prediction.
[0064] Reference Fig.10 As shown, in order to verify the effect of the present invention, in an actual noisy environment (the signal-to-noise ratio is less than 10 decibels), the noise reduction method and prediction method proposed in the present invention are compared with the non-noise-reduced data and the Bessel filter noise reduction method. It can be seen that the detection method proposed in the present invention has a high recognition accuracy, and the accuracy is significantly improved after noise reduction, which proves that the present invention still has strong applicability under low signal-to-noise ratio.
[0065] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for reducing noise of a steel tube concrete void detection signal autoencoder suitable for low signal-to-noise ratio environments, characterized by: The method comprises the following steps: S1: In a quiet environment, the steel tube concrete structure is struck at a fixed frequency and intensity, and the generated audio data is collected and saved using a microphone; S2: Automatically pick up the starting point of the knock signal, using the Akaike information criterion to find the starting point of the knock audio, and save the audio of each knock as a separate file; S3: Noise signal production. In order to restore the noise at the actual engineering site, the actual noise at the site is recorded and combined with the noise at the simulated site to expand the data set as the noise source; S4: Dataset preparation, randomly adding noise sources to the clean audio files extracted from S2 to obtain signal data as a training set; S5: Network structure construction, using a stacked denoising autoencoder model. By stacking several denoising autoencoder layers, the stack gradually learns higher-level feature representations and recovers clean signals from noisy inputs. S6: Adjust hyperparameters and conduct training; hyperparameter settings include: learning rate, number of training iterations, number of batches, and optimizer parameters; S7: Check whether the loss function of the validation set is gradually converging, and use the signal-to-noise ratio and root mean square error after denoising to evaluate the performance of the model; S8: If the performance of the two indicators in S7 meets the requirements, the training ends; if not, return to S7 to continue training, and if it meets the requirements, the training ends.
2. The method for reducing noise of a steel tube concrete void detection signal autoencoder suitable for a low signal-to-noise ratio environment according to claim 1, characterized in that: In S3, the noise sources include wind noise, car noise, and bird noise, which are mixed in different intensity ratios to produce 8 types of noise data, namely noise 1 to noise 8.
3. The method for reducing noise of a steel tube concrete void detection signal autoencoder suitable for a low signal-to-noise ratio environment according to claim 1, characterized in that: In S4, one or more of noises 1 to 8 are randomly added to the clean audio file. At the same time, a random scaling factor is applied to change the intensity of noises 1 to 8 and randomly added into the sample to obtain signal data. The signal data is used as a training set, and the denoised audio data is used as a test set.
4. The method for reducing noise of a steel tube concrete void detection signal autoencoder suitable for a low signal-to-noise ratio environment according to claim 1, characterized in that: In S5, the detailed architecture of the convolutional denoising autoencoder is as follows: input->convolutional layer 1->BN+PReLU->max pooling layer 1->convolutional layer 2->BN+PReLU->max pooling layer 2->BN+PReLU->encoder LSTM layer->latent connection layer->decoder LSTM layer->deconvolutional layer 1->BN+PReLU->deconvolutional layer 2->BN+PReLU->deconvolutional layer 2->Tanh->output; the stacked convolutional denoising autoencoder takes the output of the encoder in a trained convolutional denoising autoencoder as the input of the next autoencoder.
5. The method for reducing noise of steel tube concrete void detection signal autoencoder suitable for low signal-to-noise ratio environment according to claim 1, characterized in that: In step 5, an encoder LSTM layer and a decoder LSTM layer are added to the convolutional denoising autoencoder respectively. The encoder LSTM layer in the convolutional denoising autoencoder improves the model's ability to capture temporal features and enhances the denoising effect by modeling the time dependency of the audio.
6. The method for reducing noise of steel tube concrete void detection signal autoencoder applicable to low signal-to-noise ratio environment according to claim 1, characterized in that: The two index evaluation model formulas in S7 are as follows: Among them: x is the original signal, is the noise reduction signal, E[x 2 ] is the energy of the original signal, is the energy of the noise signal, that is, the difference between the original signal and the noise-reduced signal.
7. The method for reducing noise of steel tube concrete void detection signal autoencoder applicable to low signal-to-noise ratio environment according to claim 1, characterized in that: The signal-to-noise ratio of the signal with added noise in S4 is maintained at 0-25dB.
8. A method for detecting hollowing of concrete-filled steel tubes by percussion acoustics suitable for low signal-to-noise ratio environments, characterized in that: The method comprises the following steps: SS1: The audio signal after noise reduction in claim 1 is labeled with "noise removed" and "noise not removed" as the training set, and a part of it is divided as the validation set to verify the model capability; SS2: Welch method is used to calculate the PSD signal. First, the signal is divided into multiple segments and windowed. The FFT windowed signal segments are transformed and the power is calculated. Then, the results of all segments are averaged to obtain the PSD estimation image of the entire signal. SS3: Feature weighted splicing: Use convolutional neural networks to extract the frequency domain features expressed by each audio PSD signal, use long short-term memory networks to directly process each audio signal, extract the time domain features, and then perform weighted splicing; SS4: After the time-frequency features are concatenated, the synthesized feature vector is used for classification tasks. The data with and without gaps are mixed as a test set, and the classification task is completed using the classification layer at the end of the model.
9. According to claim 8, a method for detecting hollowing of concrete-filled steel tubes by percussion acoustics suitable for low signal-to-noise ratio environments, characterized in that: The PSD estimation image in SS2 is specifically expressed as: Among them, K is the total number of data segments, x i [n] is the i-th segment after windowing, w[n] is the window function, N w is the total energy of the window function.
10. The method for detecting hollowing of concrete-filled steel tubes by percussion acoustics in a low signal-to-noise ratio environment according to claim 8, characterized in that: The performance of weighted splicing in SS3 is: x fused =[α·x CNN , β·x LSTM ] (4) x CNN represents the frequency domain feature vector extracted by CNN, x LSTM represents the time domain feature vector extracted by LSTM, α and β are the weights of CNN and LSTM features respectively, and are the adjusted hyperparameters.
Citation Information
Patent Citations
Audio-based personalized recommendation method and device and mobile terminal
CN109558512A
Concrete pavement void intelligent detection device and detection method thereof
CN112557510A
Signal noise reduction method based on signal noise reduction auto-encoder SDE
CN114169368A
CNN + LSTM-based transformer iron core part looseness identification method and device
CN114283847A
Method and device for diagnosing void defect of concrete filled steel tube structure based on sound signal
CN115356397A
Cited By
Method and device for identifying void defect of steel plate concrete structure
CN120948609A
Porous asphalt concrete gap blockage identification method based on pavement noise signals
CN121049135A