Tunnel lining void detection method based on combination of improved auto-encoder and single-classification support vector machine and related equipment
Through the improved method of combining autoencoder with single-class support vector machine, the characteristics of tunnel lining sound and vibration signal data are extracted and reconstructed, and the problems of high probability of false alarms and poor data characterization in the prior art are solved, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510133613.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has a high probability of false alarm in tunnel lining detection, and it is not effective enough to characterize the sound and vibration data of dense tunnel lining.
Using an improved method of combining an autoencoder with a single-class support vector machine, the tunnel lining sound and vibration signal data is extracted and data reconstruction is carried out through the autoencoder, and the feature vector matrix is converted into a symmetric positive definite matrix and input it to a single-class support vector machine for detection.
It improves the accuracy of tunnel lining removal detection, reduces the probability of false alarms, and can better characterize the sound and vibration data of dense tunnel lining, enhancing the robustness of the model.
Smart Images

Figure CN120030424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tunnel safety detection, and in particular to a tunnel lining void detection method and related equipment based on the combination of an improved autoencoder and a single-classification support vector machine. Background Art
[0002] In recent years, thanks to the rapid development of deep learning, many methods in the field of computer vision have been gradually applied to the field of signal processing. However, for the problem of abnormal sound detection, since most problems only have a data set of one type of normal samples, many methods are mostly based on data reconstruction methods for anomaly detection. However, in the currently commonly used data reconstruction schemes, the optimization of the reconstruction of normal data is based on the absolute distance in Euclidean space. Such distance is too sensitive to abnormal points and has a high probability of false alarms. Summary of the invention
[0003] In order to improve the accuracy of tunnel lining void detection, the present invention proposes a tunnel lining void detection method and related equipment combining an improved autoencoder with a single-classification support vector machine.
[0004] In a first aspect, the present invention provides a tunnel lining void detection method based on a combination of an improved autoencoder and a single-classification support vector machine, comprising:
[0005] Obtain the acoustic vibration signal data of the tunnel lining to be tested;
[0006] Preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum;
[0007] Input the original Mel spectrum into an encoder in a trained autoencoder to obtain a feature vector matrix of the original Mel spectrum, and convert the feature vector matrix into a symmetric positive definite matrix;
[0008] The symmetric positive definite matrix is input into a trained single-class support vector machine to obtain a tunnel lining void detection result.
[0009] Furthermore, the training process of the autoencoder includes:
[0010] Construct a normal tunnel lining acoustic vibration signal dataset;
[0011] Using the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set as the input of the autoencoder to generate a reconstructed Mel spectrum set;
[0012] Minimize the reconstruction loss between the original Mel spectrum set and the reconstructed Mel spectrum set to iteratively optimize the parameters of the autoencoder and save the autoencoder with the optimal parameters.
[0013] Furthermore, it also includes: performing data enhancement on the normal tunnel lining acoustic vibration signal data set, and using the data enhanced normal tunnel lining acoustic vibration signal data set to train an autoencoder.
[0014] Furthermore, the data enhancement method includes one or more of audio translation, audio scaling and audio noise addition.
[0015] Furthermore, the training process of the single-classification support vector machine includes:
[0016] Inputting the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set into the encoder of the trained autoencoder to obtain all feature vector matrices of the original Mel spectrum set;
[0017] Each eigenvector matrix of the original Mel spectrum set is converted into a symmetric positive definite matrix, and a hyperplane is learned based on all symmetric positive definite matrices.
[0018] Furthermore, the reconstruction loss adopts a mean square error loss function or an Itakura-Saito spectral distance loss function.
[0019] Furthermore, the eigenvector matrix of the original Mel spectrum is converted into a symmetric positive definite matrix according to the following formula:
[0020]
[0021] Among them, Z T is the transpose of the eigenvector matrix Z, C SPD is the final symmetric positive definite matrix, and N is the dimension of the eigenvector matrix Z.
[0022] In a second aspect, the present invention provides a tunnel lining void detection device based on a combination of an improved autoencoder and a single-classification support vector machine, comprising:
[0023] An acquisition module is used to acquire acoustic vibration signal data of the tunnel lining to be tested;
[0024] A preprocessing module, used for preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum;
[0025] A feature extraction module, used for inputting the original Mel spectrum into an encoder in a trained autoencoder to obtain a feature vector matrix of the original Mel spectrum, and converting the feature vector matrix into a symmetric positive definite matrix;
[0026] The detection module is used to input the symmetric positive definite matrix into a trained single-class support vector machine to obtain a tunnel lining void detection result.
[0027] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0028] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0029] The beneficial effects of the present invention are:
[0030] (1) The tunnel lining void detection method based on the improved autoencoder combined with the single-class support vector machine proposed in the present invention converts the audio data in the data set into Mel-spectrogram data, and uses the autoencoder AE to reconstruct the data. After the AE model training is completed, only the output of the encoder in the model needs to be used as the feature representation of the original audio and transformed into a positive definite matrix. Then, the single-class support vector machine (One-class SVM, OSVM) is used to divide these positive definite matrices into hyperplanes, and finally the hyperplane (w, b) is used to calculate the abnormal scores of the test data, and finally these abnormal scores are used to determine whether the test data has a void phenomenon, and the detection accuracy is high.
[0031] (2) In the process of training the model, the present invention uses the IS (Itakura-Saito) spectral distance as the loss function to optimize the model. The loss function comprehensively considers the frequency, phase and other information of the signal, and can better characterize the acoustic vibration data of the dense tunnel lining, so that the distribution of the eigenvector matrix of the dense data in the data space can be better described by OSVM.
[0032] (3) Since the sample size of the tunnel lining acoustic vibration dataset is relatively small, this paper uses data enhancement methods to expand the tunnel lining acoustic vibration dataset and proposes a series of data enhancement methods to increase the diversity of samples and enhance the robustness of the model by simulating noise in the real environment. The feasibility and good performance of the AE-OSVM model are verified through experiments, and the generalization ability of the model is verified using the development dataset of DCASE. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 One of the flow diagrams of the tunnel lining void detection method based on the improved autoencoder combined with the single-classification support vector machine provided in the embodiment of the present invention;
[0034] Figure 2 A waveform diagram of part of the data in the tunnel lining acoustic vibration data set provided by an embodiment of the present invention;
[0035] Figure 3 Frequency information of part of the tunnel lining acoustic vibration data provided by the embodiment of the present invention;
[0036] Figure 4 Performing short-time Fourier transform on the tunnel lining acoustic vibration data provided by the embodiment of the present invention to obtain visualization of the time-frequency spectrum;
[0037] Figure 5 The Mel spectrum obtained by point multiplication of the Mel filter bank provided in the embodiment of the present invention;
[0038] Figure 6 Comparison with new training samples obtained by performing audio translation on the original audio provided in the embodiment of the present invention;
[0039] Figure 7 For comparison, white noise is added to the original audio signal to generate new training samples provided by the embodiment of the present invention;
[0040] Figure 8 A frequency spectrum of adding white noise to an original audio signal provided by an embodiment of the present invention;
[0041] Fig. 9 A second flow chart of a tunnel lining void detection method based on a combination of an improved autoencoder and a single-classification support vector machine provided in an embodiment of the present invention;
[0042] Fig.10 Model evaluation comparison on the Tunnel dataset provided by the embodiment of the present invention;
[0043] Fig.11 A schematic diagram of the structure of a tunnel lining void detection device based on a combination of an improved autoencoder and a single-classification support vector machine provided in an embodiment of the present invention;
[0044] Fig.12 A structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] like Figure 1 As shown, an embodiment of the present invention provides a tunnel lining void detection method based on a combination of an improved autoencoder and a single-classification support vector machine, comprising the following steps:
[0047] S101: Acquiring acoustic vibration signal data of the tunnel lining to be tested;
[0048] Specifically, after the tunnel wall is divided into grids, it is struck by an automatic mechanical device or manually by a hammer or other tool, and acoustic vibration signals of different grid areas are collected by using a microphone or a recorder built into the striking tool.
[0049] S102: preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum;
[0050] Specifically, raw audio data is an unprocessed audio signal, usually represented in the form of a waveform graph, where the horizontal axis represents time and the vertical axis represents the amplitude of the sound. This unprocessed audio has several disadvantages: large volume. Usually, audio data needs to retain all the information of the sound wave, and it needs to be sampled at a sampling frequency at least twice the maximum vibration frequency to ensure that the complete information of the sound wave is retained. Secondly, raw audio data is usually uncompressed, so the volume is very large. In addition, raw audio data is a continuous time series signal, which is characterized by continuous changes in the time domain, but human perception of sound is based on frequency domain information. Therefore, the raw audio data information is not easy to understand and analyze. For example Figure 2 As shown in the figure, it is the waveform representation of part of the data in the tunnel lining acoustic vibration data set. From these waveform data, only the start and end time of the audio signal and the maximum amplitude duration information can be observed, and the rest of the frequency and other information cannot be directly observed. In the tunnel lining void detection task of the present invention, more attention is paid to the characteristics of the signal in the frequency domain. Therefore, it is necessary to pre-process the original tunnel lining acoustic vibration signal data to be tested, mainly including audio data conversion.
[0051] Audio data transformation can convert the original time domain signal into a frequency domain signal, thereby extracting features more effectively and improving the interpretability of audio data. The original waveform data usually requires a high sampling frequency to save all the information, so the data volume is large and difficult to understand directly. Through audio data transformation, the data can be converted from the time domain to the frequency domain, making the representation of the data in the frequency domain easier to understand. This can improve the model's ability to understand and analyze audio signals, so that it can be better applied to tunnel lining degassing detection tasks.
[0052] Mel Spectrogram is a feature representation method for audio signal processing. It uses the Mel scale on the frequency axis to better simulate the human auditory system's perception of audio frequencies. Specifically, the Mel spectrum is a spectrum represented in the Mel scale, which is obtained by performing a dot multiplication of the tunnel lining acoustic vibration signal data with a set of Mel filters after time-frequency conversion.
[0053] Among them, time-frequency conversion aims to convert the signal from time domain representation to frequency domain representation, including but not limited to Fourier transform, short-time Fourier transform and fast Fourier transform.
[0054] For the tunnel lining acoustic vibration signal data f(t), its Fourier transform F(ω) is defined as (1):
[0055]
[0056] Where i is the imaginary unit and ω is the frequency. Fourier transform converts the time domain signal into frequency domain representation, revealing the distribution of the signal at different frequencies. By analyzing the frequency information, we can understand the frequency components of the signal and perform various operations such as filtering and noise reduction. Figure 3 The figure shows the frequency information obtained after Fourier transform of some data in the tunnel lining acoustic vibration data set. Figure 3 It can be seen that after Fourier transform, the frequency distribution of most dense audio is concentrated between 1kHz and 2kHz, which is a relatively low frequency range. In this area, the frequency distribution of dense audio is relatively concentrated, indicating that the dense acoustic vibration data has obvious signal activity in the low frequency band. However, there are also some high-frequency vibrations, which may be caused by a small amount of noise. Through Fourier transform, the frequency characteristics of audio data are more prominent, and the distinction between features is clearer, which helps to more accurately extract the characteristic information of acoustic vibration data.
[0057] Further, in Figure 3In the Fourier transform, only the frequency component of the signal can be observed, while the time information is lost in the original waveform. This may lead to a problem that the signal after Fourier transformation has the same spectrum and the same frequency component, but the phase of the original acoustic vibration signal is different. In order to solve this problem, the short-time Fourier transform (STFT) can be used to analyze the tunnel lining acoustic vibration signal. The short-time Fourier transform allows the characteristics of the signal to be observed simultaneously in two dimensions, time and frequency, so as to understand the nature of the signal more comprehensively. Its working principle is to divide the entire signal into short-time windows with the same length, and then perform Fourier transform on the signal in each window. In this way, the spectrum information of the signal at different times and frequencies can be obtained. In this embodiment, the process of performing short-time Fourier transform on the tunnel lining acoustic vibration signal is as follows: First, the tunnel lining acoustic vibration signal to be analyzed is divided into multiple time windows of fixed length, that is, the window function is used to weight the tunnel lining acoustic vibration signal to obtain a weighted signal. The window function that can be selected includes rectangular window, Hamming window, Hanning window, etc. In this embodiment, a Hamming window is selected for weighting, and its formula is as follows (2):
[0058]
[0059] Where n is the sampling point index in the window, and N represents the length of the window. By dividing the tunnel lining acoustic vibration signal into frames and adding windows, the frequency leakage in the Fourier transform process can be effectively reduced and the accuracy of spectrum analysis can be improved. Then, the Fourier transform is applied to each time window to obtain the spectrum information of several frame signals in the window. Since the short-time Fourier transform uses a sliding window, its transform result will also change with time. Its formula (3):
[0060]
[0061] Wherein, w(t) is the added window function (a Hamming window in this embodiment), and X(t,f) is the Fourier transform of w(t-τ)x(τ). As time t changes, the window function will slide on the time axis. After the window function w(t) is windowed, only the part of the signal intercepted by the window function is left for Fourier transform. It should be noted that in order to make the obtained spectrum smoother, there is usually a certain overlap when the signal is windowed, which can reduce the discontinuity of the spectrum estimation. Figure 4It shows the time-frequency spectrogram obtained after performing short-time Fourier transform on the acoustic vibration signal of the tunnel lining. From this time-frequency spectrum, not only can the frequency distribution of the acoustic vibration signal and the energy information of each frequency component be observed, but also the duration of each frequency component and the frequency attenuation in the time domain can be observed. This enables a more comprehensive time-frequency joint analysis of the acoustic vibration signal, thereby better understanding the characteristics of the acoustic vibration signal.
[0062] After obtaining the spectrum of the acoustic vibration signal data of the tunnel lining through time-frequency transformation, the Mel spectrum can be obtained according to the following transformation formula (4):
[0063]
[0064] Among them, f represents the original frequency of the signal, and m represents the converted Mel frequency. Specifically, the Mel filter is a set of triangular filters, with the starting points and ending points connected to each other to form a Mel filter bank. The frequency of each filter is linearly distributed on the Mel scale. And the Mel filter can also be inversely transformed into the spectrum under the normal frequency scale, and its formula is (5):
[0065] f = 700(10 m / 2595 - 1) (5)
[0066] Obviously, when the original frequency f is very large, after logarithmic operation, the change of the Mel frequency m tends to be gentle. As Figure 5 shown, it is the Mel spectrum obtained after transformation.
[0067] S103: Input the original Mel spectrum into the encoder of the trained autoencoder to obtain the feature vector matrix of the original Mel spectrum, and convert the feature vector matrix into a symmetric positive definite matrix;
[0068] Specifically, the detection method of the present invention is based on the autoencoder AE to train the model's reconstruction ability for the data of the dense tunnel lining. Although the AE method may have meaningless points in data reconstruction, only the encoder of the AE model is required to perform feature compression on the acoustic vibration data of the tunnel lining, and the decoder of the AE model is not required to decode the data. Therefore, it does not depend on the generation ability of the AE model.
[0069] S104: Input the symmetric positive definite matrix into the trained one-class support vector machine to obtain the tunnel lining void detection result.
[0070] Specifically, support vector machine is a classic binary classification method that separates training data samples with the largest classification interval by finding a hyperplane. However, in the tunnel anomaly detection task of the present invention, there is only one type of audio sample, so the traditional SVM is no longer applicable in this case. During the training process, the goal of the single-class support vector machine OSVM is to find a decision boundary so that all normal acoustic vibration data in the acoustic vibration data set can be correctly detected, while maximizing the distance between this boundary and the origin. The optimization problem of the single-class support vector machine is a convex optimization problem. By converting the eigenvector matrix of the Mel spectrum into a symmetric positive definite matrix and then inputting it into the single-class support vector machine, the convexity of the optimization problem can be ensured, and the computational efficiency and stability can be improved.
[0071] The tunnel lining void detection method provided in the embodiment of the present invention mainly determines whether there is a void in the tunnel by calculating the distance between the acoustic vibration sample data to be tested and the hyperplane. Therefore, the final inference result is that the tunnel is dense or there is a void, rather than being based on probabilistic reasoning. Therefore, the probability of false alarm of the present invention is lower and has better detection accuracy.
[0072] In one embodiment, the autoencoder is trained in the following manner, specifically comprising the following steps:
[0073] S201: construct a normal tunnel lining acoustic vibration signal dataset;
[0074] S202: using the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set as the input of the autoencoder to generate a reconstructed Mel spectrum set;
[0075] S203: Minimize the reconstruction loss between the original Mel spectrum set and the reconstructed Mel spectrum set to iteratively optimize the parameters of the autoencoder and save the autoencoder under the optimal parameters.
[0076] Specifically, AE is an unsupervised generative model that uses its own features as supervision to guide the neural network to try to learn a mapping relationship. First, it tries to map the original Mel spectrum to a lower-dimensional data space, and then remaps it from this low-dimensional data space back to the original data space. After the training is completed, only the encoder that maps to the low-dimensional data is needed.
[0077] A general AE model usually consists of two parts: an encoder and a decoder. The encoder is mainly used to map from the original data space to a low-dimensional data space, mainly for the purpose of data dimensionality reduction or feature compression, while the decoder is usually a mapping of the low-dimensional feature space to the original data space. The main purpose of AE is to find a set of optimal encoder / decoder pairs. In other words, for an encoder family F and a decoder family G, AE needs to find a set of optimal encoder / decoder pairs, the mathematical description of which is shown in (6):
[0078]
[0079] Among them, ε(·) represents the reconstruction error of AE for the input data x, θ and Corresponding to the parameters of the encoder and decoder, respectively, f and g correspond to a set of encoder / decoder pairs in the encoder family and decoder group, respectively.
[0080] In one embodiment, a simple mean square error loss function may be selected as the loss function for optimizing the training of AE.
[0081] Furthermore, considering that the mean square error loss function may focus too much on the absolute difference of the data itself, in one embodiment, the Itakura-Saito spectral distance (ISLoss) may be selected as the loss function of AE to optimize the model.
[0082] Specifically, IS spectral distance is a method for comparing two signals in the frequency domain, and is a measure of the difference between two signals. It is a nonlinear measure that takes into account the frequency, phase and other information of the signal, while the mean square error, as a linear measure, does not take this information into account. The calculation method of ISLoss is as follows (7):
[0083]
[0084] Where f is generally the frequency of the audio signal, but since the input is a Mel spectrum, f is the frequency band in each time window, F is the total number of frequencies in the frequency band in the Mel spectrum, and x(f) and is the power spectrum value of the two signals at frequency f, and ISLoss is the value of the calculated error. This loss value needs to be minimized to optimize the target. Using the above loss function for model training optimization can better characterize the dense tunnel lining acoustic vibration data, so that the distribution of the eigenvector matrix of dense data in the data space can be better described by OSVM.
[0085] Furthermore, in deep learning, a large amount of data is required to train the model. In practical applications, the normal tunnel lining acoustic vibration data set constructed may have the problem of small amount of data obtained and single data form. Therefore, in order to make the tunnel lining debonding detection model have higher accuracy and generalization, it is possible to choose to perform data enhancement on the constructed original data set, effectively expand the training data set, increase the diversity of data, and then use the enhanced data set to train the autoencoder to improve the robustness of the model to different environments and noises. In some embodiments, the data enhancement method includes audio translation, audio scaling, and audio noise addition; it is understandable that when enhancing the data, only one of the data enhancement methods can be used, or multiple methods can be superimposed.
[0086] (1) Audio panning
[0087] Audio panning is to generate several new audio samples by adjusting the position of the audio signal on the time axis, thereby increasing the diversity of the data. The audio enhancement method of audio panning is based on a simple assumption: the content of the audio signal is not affected by its position on the time axis, that is, all frequencies are the same and the phases are different. Therefore, based on this assumption, new training samples can be generated by adjusting the start time of the original audio signal and panning it to different start times. This enhancement method can be applied to situations where different recording positions and time delays need to be simulated. For example, in the tunnel lining debonding detection task, since the recording positions of different audios and different devices, or the timing of recording start may be different, this audio panning enhancement method can be used to generate training samples with different phases or different start times to improve the generalization ability of the model.
[0088] The mathematical formula of the audio panning enhancement method can be shown as (8):
[0089] y(t)=x(t+τ) (8)
[0090] Among them, x(t) is the original audio signal content, τ represents the time length required for translation, and y(t) is the new training sample obtained after the original audio signal is translated.
[0091] The audio panning method is simple in concept and easy to implement without complex algorithms. And because it does not actually modify the original audio signal, it does not introduce additional audio noise and can maintain the original content of the audio signal. It can enhance the diversity of data and improve the generalization ability of the model.
[0092] The data augmentation method of audio translation is a simple and effective audio data augmentation scheme. It can generate diverse training samples by adjusting the start time of the audio signal, thereby improving the generalization ability of the model. However, for tasks that are greatly affected by events, this method may change the semantic information, which may in turn affect the performance of the model. And this method requires sufficient free blank signals in the original audio signal. Otherwise, the edge part of the original audio signal may be truncated due to the translation operation, resulting in information loss. Therefore, when applying this method, it is necessary to reasonably select the translation time length according to the tunnel lining void detection task and data characteristics, and balance the relationship between data diversity and data consistency.
[0093] (2) Audio stretching
[0094] Audio stretching is based on the assumption that the content of the audio signal can change the duration of the audio signal by adjusting its playback speed without changing the frequency of the original audio signal. Therefore, the duration of the audio signal can be changed by accelerating or decelerating the playback speed of the original signal, thereby generating new training samples. This is a method of data augmentation in the time domain. Since the frequency is not changed, it will not affect the spectral components. This scheme can be used to simulate the recording speed or playback speed of different devices. For example, in the tunnel lining void detection task, the recording speeds of different devices may be different, and the playback speeds of different devices may also be different. Therefore, training samples with different durations can be generated by the stretching enhancement method. If the model requires that its samples have the same duration, they can be tiled or padded with blanks. The mathematical representation of the audio stretching data augmentation method is as in (9):
[0095] y(t) = x(α·t) (9)
[0096] Where, x(t) represents the original audio signal, α represents the stretching factor. When α is greater than 1, it means accelerating the audio, and vice versa, it means decelerating the audio. y(t) represents the audio signal obtained after accelerating or decelerating the process.
[0097] Since this scheme only changes the recording speed or playback speed of the audio signal and the energy of all frequency components remains unchanged, it will not introduce additional noise and can maintain the integrity of the content of the original audio signal. At the same time, it can simulate different recording speeds or playback speeds, and the robustness of the model can be enhanced through the training of these sample data. However, this scheme changes the duration of the audio signal. If the model has a requirement for the duration, the dataset may need to be adjusted. Such an operation may lead to a change in the learning speed of the model and affect the final prediction performance of the model. Therefore, an appropriate audio playback speed can be selected according to the needs.
[0098] (3) Audio noise
[0099] Audio noise is mainly generated by adding noise of appropriate intensity to the original audio signal to simulate the noise conditions in the real environment. This method can not only increase the diversity of data, but also simulate the real environment in reality, thereby increasing the robustness of the model. For example, in the tunnel lining void detection task, the real recording environment often has various noises and reverberations, such as environmental noise, electromagnetic interference, etc. Therefore, the audio can be noised to generate conditions with different degrees of environmental noise, thereby generating training samples, thereby increasing the generalization ability of the model. The more commonly used noise types include white noise, Gaussian noise, pink noise, etc.
[0100] White noise is a random signal with uniform power spectrum density, and its energy is constant at all frequencies. Therefore, the spectrum of white noise is very flat, without any prominent frequency. Its mathematical expression is as follows (10):
[0101] n(t)=A·rand(t) (10)
[0102] Among them, A is the amplitude of the white noise to be added, rand(t) is a random number that obeys a uniform distribution at each time point, and n(t) is the final white noise.
[0103] Gaussian noise is a random signal that obeys Gaussian distribution. Its energy will gradually decay at low and high frequencies, so its energy information presents the characteristics of a bell-shaped curve. Its mathematical expression is shown in (11):
[0104] n(t)=A·randn(t) (11)
[0105] Among them, randn(t) represents a random number that obeys the standard normal distribution at each time node.
[0106] The audio noise data enhancement method can simulate the noise conditions in the real environment, thereby improving the robustness of the model for the real environment data, and can increase the diversity of data and enhance the generalization ability of the model. Since there are many types of noise, the noise in the real situation is also diverse. Different intensities and types of noise can be added to the original data to meet the needs of different tasks. However, since noise is added to the energy and frequency of the original signal, the data of the original signal is changed, which may confuse the original signal and affect the learning effect of the model. Therefore, in practical applications, the appropriate noise type should be selected and too much noise should not be added to minimize the possibility of contaminating the original data set.
[0107] Taking the original Tunnel dataset as an example, in order to make the tunnel lining void detection model have higher accuracy and generalization, this paper uses the above method to enhance the data of the original Tunnel dataset, which has the problem of small data volume and serious imbalance of positive and negative samples. Figure 6 As shown in FIG. 1 , new samples are obtained after audio translation of part of the data in the Tunnel dataset.
[0108] from Figure 6 It can be seen that the audio after audio translation is different from the original audio except for the difference in the starting position of the audio. The energy between different frequencies and the frequency have not changed. The final frequency graph is the same as the frequency graph of the original audio. Different translation factors can be set to ensure the time span of the original data translation, so as to further expand the data set and improve the robustness of the model.
[0109] like Figure 7 As shown, in order to simulate the noise and other conditions in the real environment, white noise is added to part of the data in the Tunnel data set to obtain the audio with added noise.
[0110] from Figure 7 It can be seen from the figure that after adding white noise to the original audio signal, its waveform has low energy fluctuations at all times, and Figure 8 The figure shows the frequency change of the original audio signal after adding white noise.
[0111] from Figure 8From the numerical value of this frequency spectrum, we can observe that after adding noise, the energy of the audio signal has some small fluctuations compared to the original audio signal, and some energy fluctuations have appeared at frequencies where there was no energy originally, and there is energy at frequencies after 10kHz. The added white noise can generate white noise of different energies by adjusting the energy factor of the white noise, so as to simulate the real environment of different situations in reality, increase the diversity of the data set, and thus increase the robustness of the model.
[0112] In one embodiment, the training process of the single-class support vector machine OSVM mainly includes the following steps: After training the AE, the output of the AE is a low-dimensional latent space representation, which is learned from a certain frame of the input Mel spectrum through the encoder network. The latent space of the autoencoder is a 1×8 vector, which is a low-dimensional feature obtained by processing a spectrum frame through the encoder network. For an audio, by framing and windowing the original audio, an audio can be divided into several audio segments, and then each audio segment in the logarithmic Mel spectrum of these audio segments will be subjected to AE for feature extraction. After framing, an audio will obtain N audio segments, and finally after AE feature extraction, an N×8 feature vector matrix Z will be obtained.
[0113] The obtained N×8 eigenvector matrix Z is then converted into a symmetric positive definite matrix (SPD) to facilitate distance calculation in the subsequent data space. A symmetric positive definite matrix is a special type of square matrix with the properties of symmetry and positive definiteness.
[0114] In order to represent these eigenvector matrices as SPD matrices, it is necessary to perform operation (12) on the feature matrix extracted by AE:
[0115]
[0116] Where Z T is the transpose of the feature matrix Z, C SPD is the final SPD matrix, N is a dimension of each feature matrix, that is, the number of audio frames the original audio signal is divided into. This calculation is usually achieved by calculating the covariance matrix of the feature vector. For a certain audio segment, its feature vector, or the output of the encoder, can be directly used as the diagonal element of the SPD matrix.
[0117] Since there are only dense tunnel lining acoustic vibration data in the training set, a single-classification model is needed to detect tunnel lining voids. The present invention uses a single-classification support vector machine as a tunnel lining void detection model for void detection. The goal of OSVM is to find a decision boundary so that all normal acoustic vibration data in the acoustic vibration data set can be correctly detected, while maximizing the distance between this boundary and the origin. Among them, the hard-interval single-classification support vector machine requires that all data samples fall on the positive half of the hyperplane, while the soft-interval support vector machine does not strictly require that all acoustic vibration data samples fall on the positive half of the hyperplane, but the soft-interval classification method penalizes the dense data points that fall on the left side of the hyperplane. The penalty value is determined according to the distance from the acoustic vibration data to the hyperplane, and the corresponding hyperplane (w, b) needs to be solved by the problem in formula (13):
[0118]
[0119] Among them, w is the normal vector of the hyperplane, b is the intercept, ξ i is the slack variable and C is the regularization factor that balances the empirical error and the distance between the hyperplane and the origin.
[0120] By solving the above problem, a hyperplane (w, b) can be obtained to separate all dense data from the origin in the data space. By continuously reducing the penalty factor in the soft margin classifier, this hyperplane can learn the distribution of dense data in the acoustic vibration data set, and then calculate the abnormal score of the data point based on the distance between all data points in the training data set and the hyperplane (w, b). This abnormal score reflects the degree of deviation of the test sample from the normal data distribution, and can be used to determine whether the test data has lining delamination.
[0121] In order to verify the effectiveness of the scheme of the present invention, the present invention also provides the following experimental data. The scheme for detecting voids in tunnel linings based on an improved autoencoder and a single-classification support vector machine proposed in the present invention first reconstructs the acoustic vibration data of the tunnel lining by training an autoencoder, and then uses the encoder of the autoencoder as a feature extractor according to the trained model to obtain the audio features of the acoustic vibration data, and transforms the feature matrix into a positive definite matrix. Based on these positive definite matrices, a single-classification support vector machine is used to describe the data, and the data distribution of normal data samples is learned. Finally, the distance from the new test data to this hyperplane is calculated to detect whether there is a void in the test data. The process is mainly as follows: Fig. 9 As shown:
[0122] The various network parameters of the improved AE are shown in Table 1:
[0123] Table 1 Network structure and parameters of improved AE
[0124]
[0125] 1. Experiment
[0126] This experiment was run and tested on the Google Colab platform. Google Colab is a free cloud service developed by Google that aims to help users easily write, share, and train machine learning models. Colab provides an interactive programming environment based on Jupyter Notebook, where users can write and execute code in a web browser and view the running results of the code in real time. Colab also provides support for GPU and TPU accelerators, which users can use for free to speed up model training and reasoning, especially in some complex deep learning tasks. Google Colab also comes pre-installed with many commonly used Python libraries, such as Numpy, Pandas, Matplotlib, etc., and can use many deep learning frameworks such as TensorFlow and PyTorch at the same time. Some operations such as data processing and noise addition are performed using a local computer. The experimental configuration is shown in Table 2.
[0127] Table 2 Experimental environment configuration information
[0128]
[0129] The tunnel lining void detection method of the present invention mainly determines whether there is a void in the tunnel by calculating the distance, so the final inference result is whether the tunnel is dense or there is a void, rather than based on probabilistic reasoning. Therefore, the detection effect of the evaluation model is usually evaluated using precision, recall and AUC scores.
[0130] Precision (14) represents the probability of samples that are actually positive among all samples predicted to be positive:
[0131] precision=TP / (TP+FP) (14)
[0132] Where TP is the number of positive samples correctly predicted as positive by the model, and FP is the number of negative samples incorrectly predicted as positive.
[0133] Precision can reflect the classification effect of the classifier to a certain extent, but it cannot be used as a good indicator to measure the results when the data is unbalanced. The recall rate is introduced for evaluation. The recall rate (15) is the probability of being predicted as a positive sample among the actual positive samples:
[0134] recall=TP / (TP+FN) (15)
[0135] Where FN is the number of positive samples that are incorrectly predicted as negative.
[0136] The calculation method of AUC score is shown in (16)
[0137]
[0138] Among them, I(·) is an indicator function, which means that when the condition in the brackets is met, its value is 1, otherwise it is 0. m and n are the number of positive and negative samples in the data set, respectively, and p i represents the probability that the ith acoustic vibration data in the model pre-test data set detects a cavity in the tunnel lining, q j It indicates the probability that the model predicts that there is no cavity in the tunnel lining when the jth acoustic vibration data in the test data set is detected. The AUC score is calculated by summing and averaging the proportion of the probability that the model predicts that the positive sample is greater than the probability of the negative sample for all sample pairs. If the model's prediction results are completely random, the expected value of AUC is 0.5, indicating that the model has no ability to distinguish classification problems. Similarly, if AUC is greater than 0.5, it means that the model's prediction performance is better than random guessing.
[0139] In addition, during the model training process, some hyperparameters are needed to adjust the model. These hyperparameters are generally not obtained through the model's reasoning ability, but need to be adjusted repeatedly through experiments. The hyperparameter settings of this experiment are shown in Table 3:
[0140] Table 3 Experimental parameter settings
[0141]
[0142] (II) Results Analysis
[0143] The purpose of tunnel lining void detection is to accurately identify the acoustic vibration data with voids when only normal audio is used as the training set. Since the true labels of all dense sample data are set to 0 during the training process, the label of the void sample data should be 1. Therefore, because we should focus on the recognition rate of tunnel lining voids, the false alarm rate should be more concerned than the false alarm rate in the anomaly detection task. The F1 score can also be used to compare the recognition effect of the analysis model, and the AUC score is also an important indicator for analyzing the binary classification detection model. Table 4 shows the confusion matrix of the detection results of the AE-OSVM model on the Tunnel dataset. Since the Tunnel dataset has been enhanced, the amount of data divided into the test set has become three times the original amount, and the ratio of voids to dense data remains unchanged.
[0144] Table 4 Confusion matrix of Tunnel test dataset
[0145]
[0146] As shown in Table 5, the detection results of AE-OSVM and common methods on the tunnel lining acoustic vibration data set Tunnel are compared. From the results, it can be seen that AE-OSVM has a higher detection accuracy than the traditional model, which shows that the improved AE can better characterize dense data, so that the distribution of the feature matrix of dense data in the data space can be better described by OSVM, and then the feature matrix of the void data is farther away from the hyperplane described by OSVM, so the probability of false alarm is lower. Although the recall rate is slightly lower than the detection effect of the direct clustering algorithm, and there is a possibility of missing reports, considering these two parameters comprehensively, the AE-OSVM model can better detect the phenomenon of tunnel lining debonding.
[0147] Table 5 Model evaluation on the Tunnel dataset
[0148]
[0149] The main purpose of the data enhancement method used in this invention is to train the feature extractor AE to obtain better audio representation, thereby reducing the limitations of manual feature engineering in traditional machine learning. The detection results of the final model are as follows Fig.10 As shown, the model detection performance can be more intuitively seen. As can be seen from the figure, the improved AE-OSVM model proposed in the present invention has a higher precision rate, indicating that the model proposed in this section has a better data description of dense samples, more accurately describes the data distribution of dense samples, and can better reflect the accuracy of the model's detection of voids, avoiding false positives and reducing the manpower and material costs of subsequent detection. The higher F1 score shows that the proposed model performs better when considering both precision and recall.
[0150] In order to verify the generalization ability of the proposed model, the model proposed in this invention is applied to the Bearing and Valve datasets, which are the development datasets in the DCASE ChallengeTask2 task. The training sets of these two datasets are composed of audio recorded when the corresponding machine types are running normally. Tables 6 and 7 are the final results of anomaly detection using several models on these two datasets.
[0151] Table 6 Model evaluation on the Bearing dataset
[0152]
[0153] Table 7 Model evaluation on the Valve dataset
[0154]
[0155] It can be seen from the two tables that the model based on the optimized AE-OSVM can show good generalization performance and ensure high detection accuracy on both the Bearing and Valve datasets, indicating that AE-OSVM fits the probability distribution of normal operating data of bearings and valves well and can better describe the distribution of normal sample data in the data space. However, there is a possibility of underreporting, but the comprehensive F1 score can show a good model effect.
[0156] Compared with traditional AE, MobileNetV2 and other models on the Bearing dataset and Valve dataset, more accurate prediction performance can be achieved. The AE-SVM model has a better effect on fitting the data distribution of dense samples and has a higher detection accuracy than the clustering method.
[0157] like Fig.11 As shown, an embodiment of the present invention further provides a tunnel lining void detection device based on a combination of an improved autoencoder and a single-classification support vector machine, comprising: an acquisition module, a preprocessing module, a feature extraction module and a detection module.
[0158] Among them, the acquisition module is used to obtain the acoustic vibration signal data of the tunnel lining to be tested; the preprocessing module is used to preprocess the acoustic vibration signal data of the tunnel lining to be tested to obtain the original Mel spectrum; the feature extraction module is used to input the original Mel spectrum into the encoder in the trained autoencoder to obtain the eigenvector matrix of the original Mel spectrum, and convert the eigenvector matrix into a symmetric positive definite matrix; the detection module is used to input the symmetric positive definite matrix into the trained single-class support vector machine to obtain the tunnel lining void detection result.
[0159] It should be noted that the tunnel lining void detection device based on the combination of an improved autoencoder and a single-classification support vector machine provided in the embodiment of the present invention is intended to implement the above method. Its specific functions can be referred to the above method embodiments and will not be repeated here.
[0160] Fig.12 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.12As shown, the electronic device may include: a processor (Processor) 1201, a communication interface (Communications Interface) 1202, a memory (Memory) 1203 and a communication bus 1204, wherein the processor 1201, the communication interface 1202, and the memory 1203 communicate with each other through the communication bus 1204. The processor 1201 may call the logic instructions in the memory 1203 to execute the tunnel lining hollow detection method, which includes: obtaining the acoustic vibration signal data of the tunnel lining to be tested; preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain the original Mel spectrum; inputting the original Mel spectrum into the encoder in the trained autoencoder to obtain the eigenvector matrix of the original Mel spectrum, and converting the eigenvector matrix into a symmetric positive definite matrix; inputting the symmetric positive definite matrix into the trained single-class support vector machine to obtain the tunnel lining hollow detection result.
[0161] In addition, when the logic instructions in the above-mentioned memory 1203 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0162] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the tunnel lining void detection method provided by the above-mentioned method embodiments.
[0163] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the tunnel lining void detection method provided by the above-mentioned method embodiments is implemented.
[0164] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A tunnel lining void detection method based on an improved autoencoder combined with a single-classification support vector machine is characterized in that: include: Obtain the acoustic vibration signal data of the tunnel lining to be tested; Preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum; Input the original Mel spectrum into an encoder in a trained autoencoder to obtain a feature vector matrix of the original Mel spectrum, and convert the feature vector matrix into a symmetric positive definite matrix; The symmetric positive definite matrix is input into a trained single-class support vector machine to obtain a tunnel lining void detection result.
2. According to claim 1, the tunnel lining void detection method based on the combination of improved autoencoder and single-classification support vector machine is characterized in that: The training process of the autoencoder includes: Construct a normal tunnel lining acoustic vibration signal dataset; Using the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set as the input of the autoencoder to generate a reconstructed Mel spectrum set; Minimize the reconstruction loss between the original Mel spectrum set and the reconstructed Mel spectrum set to iteratively optimize the parameters of the autoencoder and save the autoencoder with the optimal parameters.
3. The tunnel lining void detection method based on the improved autoencoder combined with the single-classification support vector machine according to claim 2 is characterized in that: Also includes: The normal tunnel lining acoustic vibration signal dataset is data enhanced, and the normal tunnel lining acoustic vibration signal dataset after data enhancement is used to train an autoencoder.
4. The tunnel lining void detection method based on the improved autoencoder combined with the single-classification support vector machine according to claim 3 is characterized in that: The data enhancement method includes one or more of audio translation, audio scaling and audio noise addition.
5. The tunnel lining void detection method based on the combination of an improved autoencoder and a single-classification support vector machine according to any one of claims 2 to 4, characterized in that: The training process of the single-class support vector machine includes: Inputting the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set into the encoder of the trained autoencoder to obtain all feature vector matrices of the original Mel spectrum set; Each eigenvector matrix of the original Mel spectrum set is converted into a symmetric positive definite matrix, and a hyperplane is learned based on all symmetric positive definite matrices.
6. The tunnel lining void detection method based on the combination of an improved autoencoder and a single-classification support vector machine according to any one of claims 2 to 4, characterized in that: The reconstruction loss adopts a mean square error loss function or an Itakura-Saito spectral distance loss function.
7. The tunnel lining void detection method based on the combination of improved autoencoder and single-classification support vector machine according to claim 1 is characterized in that: The eigenvector matrix of the original Mel spectrum is converted into a symmetric positive definite matrix according to the following formula: Among them, Z T is the transpose of the eigenvector matrix Z, C SPD is the final symmetric positive definite matrix, and N is the dimension of the eigenvector matrix Z.
8. A tunnel lining void detection device based on an improved autoencoder combined with a single-classification support vector machine is characterized in that: include: An acquisition module is used to acquire acoustic vibration signal data of the tunnel lining to be tested; A preprocessing module, used for preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum; A feature extraction module, used for inputting the original Mel spectrum into an encoder in a trained autoencoder to obtain a feature vector matrix of the original Mel spectrum, and converting the feature vector matrix into a symmetric positive definite matrix; The detection module is used to input the symmetric positive definite matrix into a trained single-class support vector machine to obtain a tunnel lining void detection result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.