Tunnel lining void detection method based on data reconstruction and related equipment

Through a data reconstruction method, the unsupervised learning model is used to model the acoustic and vibrating signal of tunnel lining, which solves the problem of relying on scarce void data sets in the prior art, and achieves higher detection accuracy and data generalization capabilities.

CN119985696APending Publication Date: 2025-05-13KAIFENG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510133614.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing tunnel lining de-empty detection method relies on scarce void data sets, resulting in insufficient detection accuracy and data generalization capabilities.

Method used

Using a data reconstruction method, the tunnel lining sound and vibration signals are featured through an unsupervised learning model, and the model is trained using compact sound and vibration data to detect whether the tunnel lining has despair. The method includes obtaining the sound and vibration signal data, pre-processing it as the original Mel spectrum, inputting an unsupervised learning model for reconstruction, and calculating the reconstruction error to judge the gap.

Benefits of technology

Through data reconstruction, the detection accuracy of abnormal sound and vibration data of unknown types is improved, the problem of insufficient hollow samples is solved, and the detection performance is improved when facing unbalanced data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119985696A_ABST
    Figure CN119985696A_ABST
Patent Text Reader

Abstract

The invention provides a tunnel lining void detection method based on data reconstruction and related equipment. The method comprises the following steps: acquiring acoustic vibration signal data of a to-be-detected tunnel lining; preprocessing the to-be-detected tunnel lining acoustic vibration signal data to obtain an original Mel frequency spectrum; inputting the original Mel frequency spectrum into a preset unsupervised learning model to reconstruct the original Mel frequency spectrum to obtain a reconstructed Mel frequency spectrum; and calculating a reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum, and if the reconstruction error is greater than a preset error threshold, determining that the tunnel lining has void. According to the method, the problem of insufficient tunnel lining cavity samples is solved, and meanwhile, the detection accuracy of the model on abnormal sound and vibration data of unknown types is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tunnel safety detection, and in particular to a tunnel lining void detection method based on data reconstruction and related equipment. Background Art

[0002] Tunnel lining refers to a permanent support structure built with reinforced concrete and other materials along the perimeter of the tunnel to prevent deformation or collapse of the surrounding rock. Its function is to support and maintain the long-term stability and durability of the tunnel. Among the various diseases of tunnel structures, lining debonding is one of the most important factors. Lining debonding can cause problems such as uneven structural force and reduced bearing capacity. In addition, the presence of voids may indirectly lead to other diseases, such as tunnel water leakage. Therefore, timely and accurate detection of tunnel lining debonding has important engineering and social significance.

[0003] At present, the tunnel lining void detection methods mainly include core drilling method, geological radar detection method and acoustic vibration method. Among them, the acoustic vibration method is a method of exciting the detected object by certain means to make it produce mechanical vibration, and judging the quality of the detected object from the measurement results of this mechanical vibration. Using the acoustic vibration method for tunnel lining void detection is actually an abnormal sound detection and classification task. With the development of machine learning technology, some scholars have adopted the design of feature engineering combined with machine learning algorithms to enable computers to learn the pattern of data and then determine whether there is a void at the tapping point. With the rapid development of deep learning in the fields of image and speech, abnormal sound detection based on deep learning also provides a new detection direction for tunnel lining void detection.

[0004] In 2023, "Zhang X, Lin X, Zhang W, et al. Intelligent recognition of voids behind tunnel linings using deep learning and percussion sound [J]. Journal of Intelligent Construction, 2023, 1 (4)" proposed an intelligent detection scheme for voids behind tunnel linings based on deep learning and percussion sound. A large number of percussion signals were obtained through indoor percussion experiments, and then the Mel-frequency cepstral coefficients (MFCCs) were used to extract signal features. Based on this, a convolutional neural network (CNN) model was developed for automatic void defect diagnosis. However, this method requires the collection of a large number of void datasets, and in fact, tunnel lining void datasets are relatively scarce. Summary of the invention

[0005] Aiming at the problem that the existing tunnel lining void detection scheme relies on the tunnel lining void data set, while the tunnel lining void data is scarce, the present invention provides a tunnel lining void detection method based on data reconstruction and related equipment.

[0006] In a first aspect, the present invention provides a method for detecting tunnel lining voids based on data reconstruction, comprising:

[0007] Obtain the acoustic vibration signal data of the tunnel lining to be tested;

[0008] Preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum;

[0009] Inputting the original Mel spectrum into a preset unsupervised learning model to reconstruct the original Mel spectrum to obtain a reconstructed Mel spectrum; the unsupervised learning model is trained based on normal tunnel lining acoustic vibration data;

[0010] A reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum is calculated. If the reconstruction error is greater than a preset error threshold, it is considered that there is a gap in the tunnel lining.

[0011] Furthermore, the network structure of the unsupervised learning model adopts any one of an autoencoder, a variational autoencoder and a convolutional variational autoencoder; the convolutional variational autoencoder is obtained by replacing some fully connected layers in the variational autoencoder with convolutional layers.

[0012] Furthermore, the training process of the unsupervised learning model includes:

[0013] Construct a normal tunnel lining acoustic vibration signal dataset;

[0014] Using the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set as the input of the preset unsupervised learning network to generate a reconstructed Mel spectrum set;

[0015] The reconstruction loss between the original Mel-spectrogram set and the reconstructed Mel-spectrogram set is minimized to iteratively optimize the parameters of the unsupervised learning network, and the unsupervised learning network with the optimal parameters is used as the unsupervised learning model.

[0016] Furthermore, the training process of the unsupervised learning model further includes: obtaining a public sound data set; pre-training a preset unsupervised learning network using the public sound data set to obtain an initial unsupervised learning model;

[0017] Correspondingly, some network layer parameters in the initial unsupervised learning model are frozen, and the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set is used as the input of the initial unsupervised learning model to generate a reconstructed Mel spectrum set; the reconstruction loss between the original Mel spectrum set and the reconstructed Mel spectrum set is minimized to iteratively optimize the remaining network layer parameters, and the unsupervised learning network under the optimal parameters is used as the unsupervised learning model.

[0018] Furthermore, the reconstruction loss adopts a mean square error loss function or a consistency correlation coefficient loss function.

[0019] Furthermore, it also includes: determining the error threshold in the following manner, specifically including:

[0020] For each sample in the normal tunnel lining acoustic vibration signal data set, the mean square error of the original Mel spectrum and the reconstructed Mel spectrum of each sample is calculated, and the mean square error is regarded as a random variable to fit the mean square errors corresponding to all samples to the gamma distribution, so as to obtain the error threshold.

[0021] Furthermore, the preprocessing includes: performing time-frequency conversion on the acoustic vibration signal data of the tunnel lining to be tested to obtain a spectrum diagram, and passing the spectrum diagram through a group of Mel filters to obtain an original Mel spectrum; the time-frequency conversion includes one or more of Fourier transform, short-time Fourier transform and fast Fourier transform.

[0022] In a second aspect, the present invention provides a tunnel lining void detection device based on data reconstruction, comprising:

[0023] An acquisition module is used to acquire acoustic vibration signal data of the tunnel lining to be tested;

[0024] A preprocessing module, used for preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum;

[0025] A reconstruction module, used for inputting the original Mel spectrum into a preset unsupervised learning model to reconstruct the original Mel spectrum to obtain a reconstructed Mel spectrum; the unsupervised learning model is trained based on normal tunnel lining acoustic vibration data;

[0026] The detection module is used to calculate the reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum. If the reconstruction error is greater than a preset error threshold, it is considered that there is a gap in the tunnel lining.

[0027] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.

[0028] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.

[0029] Beneficial effects of the present invention:

[0030] (1) The present invention uses dense tunnel lining acoustic vibration data to train a neural network to obtain a data reconstruction model. Since the data reconstruction model is trained based on normal tunnel lining acoustic vibration data, the assumption that the reconstruction error of the data reconstruction model for unseen abnormal patterns is large realizes the detection of whether the tunnel lining is hollow. While solving the problem of insufficient tunnel lining void samples, it also improves the model's detection accuracy for unknown types of abnormal acoustic vibration data.

[0031] (2) During the training process of the data reconstruction model, the consistency correlation coefficient is used to calculate the reconstruction loss, which improves the detection performance of the model when facing unbalanced data sets.

[0032] (3) A convolutional variational autoencoder is proposed as a data reconstruction model. By replacing some of the fully connected layers of the variational autoencoder VAE with convolutional layers, the number of parameters of the fully connected layers during variational inference is reduced; and during feature extraction, not only can the relationship between different times and frequencies be calculated, but also the degree of attenuation of frequencies at different times can be calculated, increasing the attention to local features and further improving the detection performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A schematic diagram of a flow chart of a tunnel lining void detection method based on data reconstruction provided by an embodiment of the present invention;

[0034] Figure 2 A waveform diagram of part of the data in the tunnel lining acoustic vibration data set provided by an embodiment of the present invention;

[0035] Figure 3 Frequency information of part of the tunnel lining acoustic vibration data provided by the embodiment of the present invention;

[0036] Figure 4 Performing short-time Fourier transform on the tunnel lining acoustic vibration data provided by the embodiment of the present invention to obtain visualization of the time-frequency spectrum;

[0037] Figure 5 The Mel spectrum obtained by point multiplication of the Mel filter bank provided in the embodiment of the present invention;

[0038] Figure 6 A schematic diagram of the model structure of the autoencoder provided in an embodiment of the present invention;

[0039] Figure 7 A structural diagram of a probability encoder provided by an embodiment of the present invention;

[0040] Figure 8 A structural diagram of a probability generator provided by an embodiment of the present invention;

[0041] Fig. 9 A structural diagram of the Conv-VAE model provided in an embodiment of the present invention;

[0042] Fig.10 The Mel spectrum after reconstructing the data in the tunnel acoustic vibration data set provided by the embodiment of the present invention;

[0043] Fig.11 Comparison of evaluation indicators on the Tunnel dataset provided by the embodiment of the present invention;

[0044] Fig.12 A schematic diagram of the structure of a tunnel lining void detection device based on data reconstruction provided by an embodiment of the present invention;

[0045] Fig.13 A structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0047] Since the lining voids in actual tunnel projects are rare and highly diverse, it is difficult to collect detailed abnormal sound patterns. Therefore, a tunnel lining void detection method that can detect unknown types of abnormal sounds is needed. Figure 1 As shown, the tunnel lining void detection method based on data reconstruction proposed in the embodiment of the present invention mainly includes the following steps:

[0048] S101: Acquiring acoustic vibration signal data of the tunnel lining to be tested;

[0049] Specifically, an acoustic vibration collector is installed inside the tunnel to collect the acoustic vibration signal data of the tunnel lining; or the tunnel wall is manually hit with a hammer or other tool, and a microphone is used to collect the sound signal generated when the tunnel wall is hit.

[0050] S102: preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum;

[0051] Specifically, raw audio data is an unprocessed audio signal, usually represented in the form of a waveform graph, where the horizontal axis represents time and the vertical axis represents the amplitude of the sound. This unprocessed audio has several disadvantages: large volume. Usually, audio data needs to retain all the information of the sound wave, and it needs to be sampled at a sampling frequency at least twice the maximum vibration frequency to ensure that the complete information of the sound wave is retained. Secondly, raw audio data is usually uncompressed, so the volume is very large. In addition, raw audio data is a continuous time series signal, which is characterized by continuous changes in the time domain, but human perception of sound is based on frequency domain information. Therefore, the raw audio data information is not easy to understand and analyze. For example Figure 2 As shown in the figure, it is the waveform representation of part of the data in the tunnel lining acoustic vibration data set. From these waveform data, only the start and end time of the audio signal and the maximum amplitude duration information can be observed, and the rest of the frequency and other information cannot be directly observed. In the tunnel lining void detection task of the present invention, more attention is paid to the characteristics of the signal in the frequency domain. Therefore, it is necessary to pre-process the original tunnel lining acoustic vibration signal data to be tested, mainly including audio data conversion.

[0052] Audio data transformation can convert the original time domain signal into a frequency domain signal, thereby extracting features more effectively and improving the interpretability of audio data. The original waveform data usually requires a high sampling frequency to save all the information, so the data volume is large and difficult to understand directly. Through audio data transformation, the data can be converted from the time domain to the frequency domain, making the representation of the data in the frequency domain easier to understand. This can improve the model's ability to understand and analyze audio signals, so that it can be better applied to tunnel lining degassing detection tasks.

[0053] Mel Spectrogram is a feature representation method for audio signal processing. It uses the Mel scale on the frequency axis to better simulate the human auditory system's perception of audio frequencies. Specifically, the Mel spectrum is a spectrum represented in the Mel scale, which is obtained by performing a dot multiplication of the tunnel lining acoustic vibration signal data with a set of Mel filters after time-frequency conversion.

[0054] Among them, time-frequency conversion aims to convert the signal from time domain representation to frequency domain representation, including but not limited to Fourier transform, short-time Fourier transform and fast Fourier transform.

[0055] For the tunnel lining acoustic vibration signal data f(t), its Fourier transform F(ω) is defined as (1):

[0056]

[0057] Where i is the imaginary unit and ω is the frequency. Fourier transform converts the time domain signal into frequency domain representation, revealing the distribution of the signal at different frequencies. By analyzing the frequency information, we can understand the frequency components of the signal and perform various operations such as filtering and noise reduction. Figure 3 The figure shows the frequency information obtained after Fourier transform of some data in the tunnel lining acoustic vibration data set. Figure 3 It can be seen that after Fourier transform, the frequency distribution of most dense audio is concentrated between 1kHz and 2kHz, which is a relatively low frequency range. In this area, the frequency distribution of dense audio is relatively concentrated, indicating that the dense acoustic vibration data has obvious signal activity in the low frequency band. However, there are also some high-frequency vibrations, which may be caused by a small amount of noise. Through Fourier transform, the frequency characteristics of audio data are more prominent, and the distinction between features is clearer, which helps to more accurately extract the characteristic information of acoustic vibration data.

[0058] Further, in Figure 3 In the Fourier transform, only the frequency component of the signal can be observed, while the time information is lost in the original waveform. This may lead to a problem that the signal after Fourier transformation has the same spectrum and the same frequency component, but the phase of the original acoustic vibration signal is different. In order to solve this problem, the short-time Fourier transform (STFT) can be used to analyze the tunnel lining acoustic vibration signal. The short-time Fourier transform allows the characteristics of the signal to be observed simultaneously in two dimensions, time and frequency, so as to understand the nature of the signal more comprehensively. Its working principle is to divide the entire signal into short-time windows with the same length, and then perform Fourier transform on the signal in each window. In this way, the spectrum information of the signal at different times and frequencies can be obtained. In this embodiment, the process of performing short-time Fourier transform on the tunnel lining acoustic vibration signal is as follows: First, the tunnel lining acoustic vibration signal to be analyzed is divided into multiple time windows of fixed length, that is, the window function is used to weight the tunnel lining acoustic vibration signal to obtain a weighted signal. The window function that can be selected includes rectangular window, Hamming window, Hanning window, etc. In this embodiment, a Hamming window is selected for weighting, and its formula is as follows (2):

[0059]

[0060] Where n is the sampling point index in the window, and N represents the length of the window. By dividing the tunnel lining acoustic vibration signal into frames and adding windows, the frequency leakage in the Fourier transform process can be effectively reduced and the accuracy of spectrum analysis can be improved. Then, the Fourier transform is applied to each time window to obtain the spectrum information of several frame signals in the window. Since the short-time Fourier transform uses a sliding window, its transform result will also change with time. Its formula (3):

[0061]

[0062] Wherein, w(t) is the added window function (a Hamming window in this embodiment), and X(t,f) is the Fourier transform of w(t-τ)x(τ). As time t changes, the window function will slide on the time axis. After the window function w(t) is windowed, only the part of the signal intercepted by the window function is left for Fourier transform. It should be noted that in order to make the obtained spectrum smoother, there is usually a certain overlap when the signal is windowed, which can reduce the discontinuity of the spectrum estimation. Figure 4 The time-frequency spectrum obtained after short-time Fourier transform of the tunnel lining acoustic vibration signal is shown. From this time-frequency spectrum, we can not only observe the frequency distribution of the acoustic vibration signal and the energy information of each frequency component, but also the duration of each frequency component in the time domain and frequency attenuation. This can perform a more comprehensive time-frequency joint analysis of the acoustic vibration signal, so as to better understand the characteristics of the acoustic vibration signal.

[0063] After obtaining the spectrum of the tunnel lining acoustic vibration signal data through time-frequency transformation, the Mel spectrum can be obtained according to the following transformation formula (4):

[0064]

[0065] Where f represents the original frequency of the signal, and m represents the converted Mel frequency. In specific implementation, the Mel filter is a set of triangular filters, the starting points and end points of which are connected to each other to form a Mel filter bank. The frequency of each filter is linearly distributed on the Mel scale. The Mel filter can also be inversely transformed to a spectrum under a normal frequency scale, and the formula is (5):

[0066] f=700(10 m / 2595 -1) (5)

[0067] Obviously, when the original frequency f is very large, after logarithmic operation, the change of Mel frequency m tends to be flat. Figure 5 As shown, it is the Mel spectrum obtained after transformation.

[0068] S103: inputting the original Mel spectrum into a preset unsupervised learning model to reconstruct the original Mel spectrum to obtain a reconstructed Mel spectrum;

[0069] Specifically, since in actual tunnel engineering, tunnel lining voids rarely occur, and tunnel conditions in the real world are highly diverse, it is difficult to collect detailed patterns of abnormal sounds. Moreover, there is currently no such public data set, and the cost of acquiring void acoustic vibration data is also very high, so tunnel lining acoustic vibration data with voids is relatively scarce. However, dense acoustic vibration data is easy to obtain. Theoretically, if the model is good enough, modeling of dense acoustic vibration data will be more effective than modeling of void acoustic vibration data. Therefore, the method of the present invention uses an unsupervised learning model to extract features and model dense acoustic vibration data through data reconstruction, thereby assisting in the detection of tunnel lining acoustic vibration data with voids, solving the problem of insufficient tunnel lining void samples.

[0070] Unsupervised learning models do not rely on pre-labeled training data, but discover patterns and structures in the data by analyzing the intrinsic characteristics of the data. Therefore, the present invention can solve the problem of high data annotation costs by using unsupervised learning models for data reconstruction. A mel spectrum of the tunnel lining acoustic vibration data obtained after preprocessing the tunnel lining acoustic vibration data can be regarded as a single-channel two-dimensional image, and then the unsupervised learning model is used to reconstruct the obtained mel spectrum map features.

[0071] S104: Calculate a reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum. If the reconstruction error is greater than a preset error threshold, it is considered that there is a gap in the tunnel lining.

[0072] Specifically, since the acoustic vibration signal data of normal tunnel lining is modeled, the reconstruction error of the model will tend to be concentrated. When the model parameters are good enough, the errors of all normal data after reconstruction should be concentrated in a smaller numerical range (that is, the preset error threshold). Therefore, once the reconstruction error exceeds the preset error threshold, it can be considered that there is a void in the tunnel lining.

[0073] In one embodiment, the unsupervised learning model uses an auto-encoder. An auto-encoder is a generative learning model, and its structure is as follows: Figure 6. This method uses the input data itself as supervision to guide the neural network to try to learn a mapping relationship. It first maps the Mel spectrum to a relatively low-dimensional data space, and then uses the data in this low-dimensional space to reconstruct the original Mel spectrum. The autoencoder consists of two parts: an encoder and a decoder. The encoder mainly fits a mapping from the original data to the intermediate low-dimensional latent variable. It maps the original input data from the original high-dimensional space to a low-dimensional data space, thus achieving the purpose of data compression or dimensionality reduction; the decoder is the opposite of the encoder. It is mainly a mapping from the data space of the latent variable to the original data space, and from the low-dimensional latent variable space to the original data space. The core goal of the autoencoder is to find the optimal combination of encoders and decoders in a given data set. In other words, its task is to find those combinations from a set of possible encoders and decoders that can retain information to the greatest extent during encoding and minimize reconstruction errors during decoding. If G and F are used to represent the encoder family and decoder family under consideration, respectively, the autoencoder can be described as (6):

[0074]

[0075] where ε(·) represents the reconstruction error between the encoder and decoder, φ * and θ * They correspond to the parameters of the optimal encoder and the optimal decoder respectively. The original input data x and the reconstructed data f θ (g φ The closer (x)) is, the better the reconstruction effect of the autoencoder is. Therefore, the training loss of the general autoencoder can be calculated using the MSE loss (7):

[0076]

[0077] where g φ is the mapping function of the encoder, f θ is the mapping function of the decoder. Then for the original input x, the encoder can obtain the low-dimensional latent variable z=g φ (x), and then the decoder can map the low-dimensional implicit features to the original data space to achieve data reconstruction:

[0078] After the encoder and decoder are trained by the dataset, they can be generalized in different tasks. For example, if the encoder is used alone, the input data can be reduced in dimension. If the decoder is used alone, the noise randomly generated in the latent variable space can be used as the input of the decoder, which can be used as a generative model to generate data in the original data space. However, the feature space of the latent variables obtained after the autoencoder is trained is very limited and has no certain regularity. The latent variable space depends entirely on the data distribution of the original data space, which will cause some points in the latent variable space to give a series of meaningless data points after passing through the decoder. The result is that the data generated by directly inputting noise is far from the expected data.

[0079] Furthermore, in order to solve the problem of bias in the generated data of the autoencoder, in one embodiment, the unsupervised learning model uses a variational autoencoder based on a convolutional neural network to reconstruct the tunnel lining acoustic vibration signal data. Based on the variational reasoning autoencoder, although the variational autoencoder also has the word "autoencoder" in its name, this is mainly because the variational autoencoder and the autoencoder have similar structures, both of which are composed of structures such as encoders and decoders. As for the irregularity of the autoencoder in the latent variable space, the variational autoencoder introduces explicit regularization to address this problem. Unlike the autoencoder that maps the data in the original data space to a single point in the latent variable space, the variational autoencoder maps the data in the original data space to the probability distribution in the latent variable space, thus avoiding the appearance of meaningless points when the probability decoder generates data, and sampling is performed from the probability distribution of the latent variable data space during decoding, and data is generated through the decoder.

[0080] As a probabilistic generative model, the variational autoencoder first defines a probabilistic generative graph model to describe the process of reconstructing the Mel spectrum. To represent the N-dimensional variable composed of the Mel spectrum after the audio is transformed, and assume that It is generated by the latent variable z that cannot be directly observed. Therefore, for each data point, it is generated by the following two steps: first, the latent variable z is sampled from a simple prior distribution p(z), and then the Mel spectrum is reconstructed It is sampled from the conditional likelihood distribution p(x|z).

[0081] Considering such a probability model, considering the probability encoder and probability decoder, the probability decoder is naturally defined by the conditional likelihood distribution p(x|z) of the generated data x, while the probability encoder is the opposite of this process, it is defined by p(z|x). According to the famous Bayes theorem, the above-mentioned prior distribution p(z), conditional likelihood p(x|z), and posterior distribution p(z|x) are linked by this theorem, and the relationship is shown in formula (8):

[0082]

[0083] Now assume that the prior distribution of the latent variable p(z) is a standard Gaussian distribution, and the conditional likelihood p(x|z) is a mixed Gaussian distribution, whose mean and variance are defined by the transformation function f of the latent variable z. This transformation function is defined by the decoder, and it is a function belonging to the F family. Theoretically speaking, knowing the prior distribution p(z) and the likelihood p(x|z), the posterior distribution p(z|x) can be calculated by Bayes' theorem. However, this calculation is tricky because the true distribution p(x) of the Mel spectrum after preprocessing is unknown. At this time, variational inference is introduced to approximate the posterior distribution of the latent variable z. Variational inference is a technique in statistics for approximating complex distributions. It sets a parameterized distribution family and finds the best approximation of the target distribution in this family set. The general structures of the probability encoder and probability decoder are as follows: Figure 7 and Figure 8 shown.

[0084] Since the standard Gaussian distribution is used as the distribution of the latent variable z in the generative model, the standard Gaussian distribution q(z|x) will be used here to approximate the posterior distribution p(z|x) of the latent space, where q(z|x) can be defined by the encoder, and its mean and variance are obtained by two transformations of the input Mel spectrum. The transformation functions are defined as g and h, which belong to the families of functions G and H respectively. Therefore, the candidate family of variational inference is now defined, and now it is necessary to minimize the distance between the approximate distribution and the target posterior distribution p(z|x) by optimizing the parameters of the function. That is, it is necessary to find the best transformation g * and h * , thereby minimizing the distance between the standard Gaussian distribution and the target posterior distribution. The distance between two distributions is usually measured using relative entropy, that is, KL divergence. Information entropy gives a quantitative measure of the degree of difference between the two distributions, indicating the distance between the two distributions. Then the goal of this model optimization is (9):

[0085]

[0086] where g * and h *They represent the best transformation functions for the input x in the G and H distribution families respectively.

[0087] According to the definition of relative entropy, the above formula can be expanded to obtain (10):

[0088]

[0089] Then expand the logarithm, as shown in (11):

[0090]

[0091] Finally, the goal of model optimization is to maximize the logarithmic expectation of the conditional likelihood p(x|z) and minimize the distance between the standard normal distribution q(z|x) and the prior distribution p(z) of the latent variable z. In other words, maximizing the conditional likelihood is to obtain the generated Mel spectrum under the condition of the given latent variable z. The maximum probability can also be seen as the input spectrum x and the generated spectrum The most similar, the similarity can be measured by formula (12):

[0092] log p(x|z)=-||xf(z)|| 2 (12)

[0093] Therefore, based on the definition of the original decoder, we can get the best approximate distribution q(z|x) of a posterior distribution p(z|x) of a latent variable z through any f in the F distribution family. Therefore, for function f, when the latent variable z is sampled from q(z|x), it can maximize the generation In other words, for a given input Mel spectrum x, it is an N-dimensional random variable, when the latent variable z is sampled from the approximate distribution q(z|x) of the latent variable, and then sampled from the distribution p(x|z) Can maximize The probability of this is the Mel spectrum generated by data reconstruction. It is closest to the Mel spectrum x after preprocessing.

[0094] In machine learning and deep learning, the loss function is an important indicator to measure the difference between the model prediction value and the true value. In one embodiment, the mean square error (MSE) loss function is used to calculate the reconstruction loss. The MSE loss function mainly focuses on the difference between the predicted value and the true value, that is, calculating the square of the difference between them.

[0095] However, the MSE loss function is highly sensitive to outliers, which means that it may not be able to handle reconstructed outliers well. This will cause the model to over-focus on outliers during training, thereby affecting the overall performance. Therefore, further, in one embodiment, the consistency correlation coefficient (Concordance Correlation Coefficient, CCC) is used to measure the difference between the model prediction value and the true value, and this is used as a loss function to improve the performance of the tunnel lining debonding detection method. The core idea of ​​the CCC loss function is to measure the consistency between the predicted value of the Mel spectrum and the true Mel spectrum. This consistency not only includes the correlation between the two, but also includes an exact difference in their values. The CCC loss function provides a more comprehensive metric by combining the Pearson correlation coefficient, mean difference, and standard deviation difference. This measurement method makes the CCC loss function show higher robustness when dealing with outliers, because it not only focuses on the size of the absolute error of the predicted value, but also focuses on the relative relationship between the predicted value and the true value. The calculation formula of the CCC loss function is shown in (13):

[0096]

[0097] in is the Pearson correlation coefficient between the predicted value and the true value, σ x and are the standard deviations of the true value and the predicted value, μ x and are the means of the true value and the predicted value, respectively. The advantage of the CCC loss function lies in its robustness to outliers and its comprehensive evaluation of prediction performance. In the abnormal sound detection task, the CCC loss function can better guide model learning and improve the performance of the model when facing an unbalanced data set. In addition, the intuitiveness of the CCC loss function also makes the evaluation and comparison of models simpler and clearer.

[0098] In one embodiment, after reconstructing the data, the reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum can be calculated using the mean square error (14) of the two Mel spectrums:

[0099]

[0100] Where E represents the original Mel spectrum x and the reconstructed Mel spectrum The mean square error between them, i represents the components of the Mel spectrum.

[0101] Furthermore, normal tunnel lining acoustic vibration signal data is used in the training set of the model. When the model parameters are good enough, the errors of all normal data after reconstruction should be concentrated in a smaller numerical range. Correspondingly, the reconstructed mean square error is regarded as a random variable, which should obey the gamma distribution. By using the mean square error after all Mel spectrum reconstruction as a random variable and fitting it to the gamma distribution, the threshold of the reconstruction error can be determined. Then, the threshold can be used to detect whether there is a lining gap at the tunnel knocking point.

[0102] In one embodiment, Fig. 9 This is a structural diagram of an unsupervised learning model for data reconstruction. The model is mainly composed of convolutional layers, deconvolutional layers, pooling layers, fully connected layers, and resampling layers. The collected audio data needs to be transformed into Mel spectrum features in the preprocessing stage, and then the two-dimensional Mel spectrum is used as a single-channel image input into the model. In order to calculate the features of the time domain and frequency domain at the same time, some fully connected layers of the variational autoencoder VAE are replaced with convolutional layers. The purpose of using convolutional layers is to reduce the number of parameters of the fully connected layers during variational inference. In addition, due to the addition of convolutional layers, when extracting features, not only the relationship between different times and frequencies can be calculated, but also the degree of attenuation of frequencies at different times can be calculated, and the attention of local features is increased. The posterior distribution of the latent variable z is obtained through the probability encoder, and the sample is sampled from the posterior distribution of z. Then, the reconstructed Mel spectrum is obtained according to the maximum likelihood according to the probability decoder. The final judgment of whether the lining is empty is made based on the reconstruction effect of the model. The loss of the model is calculated by the mean square error between the reconstructed Mel spectrum and the original Mel spectrum and the distance between the prior distribution of the latent variable and the posterior distribution obtained by the probability encoder. The specific details of the model are shown in Table 1:

[0103] Table 1 Conv-VAE model structure

[0104]

[0105]

[0106] Since the latent variable z in this model is obtained by sampling from the approximate posterior distribution q(z|x), the back-propagation algorithm will not be able to proceed to this point. The latent variable z cannot be differentiated with respect to the mean μ and variance σ, so a resampling technique is required for the latent variable z. As described above, the approximate distribution q(z|x) of the posterior distribution of the latent variable obeys a Gaussian distribution. The mean μ and variance σ of the distribution can be obtained through data training, and then an additional normal distribution p(ξ) is used to sample ξ from p(ξ). Then z is linked to μ and σ through z=μ+ξ·σ. In this way, when the model loss is back-propagated to the z layer, the mean and variance can be differentiated separately, and back-propagation can proceed normally. As Fig.10 As shown, a sound vibration data randomly selected from the self-built tunnel sound vibration data set is compared with the original Mel spectrum and the reconstructed Mel spectrum after the data is reconstructed by the VAE model built in this experiment.

[0107] In order to verify the effectiveness of the solution of the present invention, the embodiments of the present invention also provide the following experimental data.

[0108] In this experiment, the data used for model training comes from the real tunnel knocking data collected by the tunnel bureau. Since there are few existing tunnel knocking samples, a transfer learning strategy was used for model training in the tunnel lining void detection experiment based on data reconstruction. Specifically, the model was pre-trained using the development dataset of DCASE 2023ChallengeTask2, and then a small number of fully connected layers were frozen. The pre-trained model was fine-tuned using tunnel lining acoustic vibration data, and finally applied to the tunnel lining void detection task. This strategy can make full use of the knowledge of the pre-trained model and adapt it to the needs of the target task through fine-tuning, thereby improving detection performance and generalization ability.

[0109] The goal of DCASE Challenge Task 2 is similar to this study. It is based on different machine types in seven different application scenarios, and uses audio recorded during normal use to detect abnormal conditions of these seven types of machines in these application scenarios. The existing tunnel knocking audio mainly includes 400 normal tunnel knocking point audios without lining debonding and 100 knocking audios of tunnels with lining debonding. Among them, 350 normal tunnel knocking sounds in the dataset are used to train the model, and the remaining 50 normal tunnel data and 100 abnormal tunnel data are used to test the model detection effect. The development dataset of DCASE Challenge Task 2 is composed of audios of seven different machine types, including Fan, Gearbox, Bearing, Sliderail, Toy Car, Toy Train, and Valve. Each set of data includes a training set consisting of 1,000 normal samples and a test set consisting of 100 normal samples and 100 abnormal samples.

[0110] (I) Experimental parameter setting

[0111] In the experiment, for the real tunnel percussion acoustic vibration data set, the length of each sample is processed as a 2-second slice. In order to perform time-frequency joint analysis, the librosa library is used to perform short-time Fourier transform on the time domain data of the audio. The librosa library is a powerful Python toolkit dedicated to audio analysis and processing. With the help of this library, functions such as time-frequency processing, feature extraction, and sound graphics can be easily implemented. The data is sampled at a frequency of 16000Hz. This is because when sampling audio, a sampling frequency of at least twice the highest frequency is required to retain effective information to the maximum extent. If a higher sampling frequency is used, the model parameters will increase, making it more difficult to train. Then the data is framed, and a window of 1024 is added to the data, that is, the frame length. The Fourier transform is performed on each frame of the data through the sliding window. Since there needs to be a certain connection between adjacent frames, a frame shift of 256 samples is selected here. The experiment runs on the Google Colab cloud platform and uses TeslaT4 for training. The framework used is the Google TensorFlow deep learning framework.

[0112] 2. Evaluation indicators

[0113] In order to accurately evaluate the defect detection performance of the proposed model, this paper conducts a detailed quantitative analysis of the experimental results. Including the area under the ROC curve (AUC), precision, recall and F1 score. The ROC curve is drawn based on the confusion matrix, where the horizontal axis is the false positive rate (FPR) and the vertical axis is the true positive rate (TPR). Therefore, the ideal situation of the model is that the lower the FPR, the higher the TPR. The AUC score is an overall evaluation of the classifier performance, ranging from 0 to 1, and the higher the score, the better the performance. pAUC refers to the AUC value of the partial area under the ROC curve within a specific range of interest. Precision, recall and F1 score are used to comprehensively evaluate the performance of the classifier. Precision (15) represents the probability that a sample is actually positive among all samples predicted to be positive:

[0114] precision=TP / (TP+FP) (15)

[0115] TP is the number of positive samples correctly predicted as positive by the model, and FP is the number of negative samples incorrectly predicted as positive. Precision can reflect the classification effect of the classifier to a certain extent, but it cannot be used as a good indicator to measure the results in the case of data imbalance. The recall rate is introduced for evaluation. The recall rate (16) is the probability of being predicted as a positive sample among the samples that are actually positive:

[0116] recall=TP / (TP+FN) (16)

[0117] FN is the number of positive samples that are incorrectly predicted as negative samples. The two parameters of the evaluation model, precision and recall, are basically a dilemma in practical applications. Therefore, in order to combine the performance of the two, it is necessary to find a balance between the two. An F1 score (17) is introduced to balance precision and recall:

[0118] F1=(2×precision×recall) / (precision+recall)(17)

[0119] The proposed convolutional variational autoencoder is trained using the self-built Tunnel dataset. After training, the mean square error between the preprocessed Mel spectrum and the Mel spectrum generated by the model is calculated as the anomaly score. The anomaly score is used as a random variable, and the gamma distribution is used to detect whether the tapping point of the input audio is a missed tapping point. After testing with the test set, the confusion matrix can be obtained, as shown in Table 2:

[0120] Table 2 Tunnel test set confusion matrix

[0121]

[0122] According to the TP, TN, FP, FN of the confusion matrix in the above table, and the anomaly score calculated by the model, the model can be evaluated accordingly, and the AUC, Precision, Recall and F1 scores can be calculated based on the above features. Then record this feature, and use AE, Mobile-Net, VAE and other models to detect tunnel lining voids on the Tunnel dataset, and obtain the evaluation data of each model on the self-built Tunnel dataset, and compare the evaluation results of Conv-VAE and other models. All the results are shown in Table 3:

[0123] Table 3 Model evaluation on the Tunnel dataset

[0124]

[0125] From the data in Table 3, it can be seen that after the Conv-VAE based on CCC optimized variational inference is trained with tunnel knocking audio data, its accuracy in detecting tunnel lining voids has been improved to varying degrees compared with the autoencoder and MobileNetV2, and slightly reduced compared with VAE. This is because the model sets a small threshold when focusing on the recognition rate of void samples, but the recall rate of the Conv-VAE model has been greatly improved, indicating that the model proposed in the present invention has more concentrated errors in the data reconstruction process. The loss function based on CCC can optimize the model with a better metric for the reconstruction of dense data samples, thereby reducing the absolute error of the reconstructed data. In addition, the F1 score of this model is higher than that of other models, and it can have better comprehensive detection results while considering both the false alarm rate and the false negative rate.

[0126] like Fig.11As shown in the figure, it is a more intuitive comparison between the model proposed in the present invention and AE, MobileNetV2 and VAE. The AUC score is the area under the ROC curve, which reflects the overall performance of the model between the true positive rate and the false positive rate at different thresholds. For the void detection task in this article, the high or low AUC shows the model's ability to distinguish between normal and void samples. It can be seen from the figure that the model proposed in the present invention has a higher AUC score and has a higher detection ability. The recall rate can measure the proportion of samples successfully predicted as voids in the actual void samples by the model. In the tunnel lining void detection task in this article, the void samples can be better identified. The recall rate shows the model's ability to identify void samples. It can be seen from the figure that the model proposed in the present invention has a higher recall rate, indicating that the reconstruction error of the Conv-VAE model for dense void samples is relatively more concentrated, making the calculated anomaly score more concentrated, and the reconstruction error of the void samples is relatively larger, which can better detect void samples, reduce the false alarm rate, and improve the safety of the tunnel. The F1 score is the harmonic mean of precision and recall. It combines the performance of precision and recall and is a comprehensive evaluation indicator used to measure the balance between positive and negative classes of the classifier. In the task of this paper, the model proposed by the present invention has a higher F1 score, which means that the model of the present invention performs better when considering both precision and recall, and can maintain a low false positive rate and false negative rate.

[0127] Then, for the development dataset of the DCASE challenge, this model is applied to its development dataset to verify the generalization performance of Conv-VAE. Since the development dataset of DCASE is composed of normal or abnormal audio recorded by seven different types of machines in different application scenarios, this model uses two of the datasets, Bearing and Valve, to verify the generalization performance of Conv-VAE. Conv-VAE is compared with AE, Mobile-Net and VAE models respectively, and the results are shown in Tables 4 and 5:

[0128] Table 4 Model evaluation on the Bearing dataset

[0129]

[0130] Table 5 Model evaluation on the Valve dataset

[0131]

[0132] From the data in Table 4 and Table 5, we can see that the Conv-VAE model can detect abnormal events well in the anomaly detection task. Although the detection accuracy is slightly reduced and there is a risk of false positives, it is more accurate in detecting abnormal sounds, has a higher detection recall rate, and a lower false negative rate. The harmonic mean F1 score is also higher.

[0133] Through the self-built dataset Tunnel in this experiment and the public datasets Bearing and Valve, it can be verified that the Conv-VAE proposed in this experiment has better performance in the abnormal sound detection task than models such as AE and MobileNetV2, and can have a better reconstruction effect on the samples in the dataset. Its recognition accuracy and the performance of the model in different datasets have been improved to varying degrees.

[0134] like Fig.12 As shown, an embodiment of the present invention further provides a tunnel lining void detection device based on data reconstruction, comprising: an acquisition module, a preprocessing module, a reconstruction module and a detection module.

[0135] Among them, the acquisition module is used to obtain the acoustic vibration signal data of the tunnel lining to be tested; the preprocessing module is used to preprocess the acoustic vibration signal data of the tunnel lining to be tested to obtain the original Mel spectrum; the reconstruction module is used to input the original Mel spectrum into a preset unsupervised learning model to reconstruct the original Mel spectrum to obtain a reconstructed Mel spectrum, and the unsupervised learning model is trained based on normal tunnel lining acoustic vibration data; the detection module is used to calculate the reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum. If the reconstruction error is greater than the preset error threshold, it is considered that there is a void in the tunnel lining.

[0136] It should be noted that the tunnel lining void detection device provided in the embodiment of the present invention is for realizing the above method. Its specific functions can be referred to the above method embodiments and will not be described in detail here.

[0137] Fig.13 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.13As shown, the electronic device may include: a processor (Processor) 1301, a communication interface (Communications Interface) 1302, a memory (Memory) 1303 and a communication bus 1304, wherein the processor 1301, the communication interface 1302, and the memory 1303 communicate with each other through the communication bus 1304. The processor 1301 may call the logic instructions in the memory 1303 to execute the tunnel lining hollow detection method, which includes: obtaining the acoustic vibration signal data of the tunnel lining to be tested; preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum; inputting the original Mel spectrum into a preset unsupervised learning model to reconstruct the original Mel spectrum to obtain a reconstructed Mel spectrum; calculating the reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum, and if the reconstruction error is greater than a preset error threshold, it is considered that the tunnel lining is hollow.

[0138] In addition, when the logic instructions in the above-mentioned memory 1303 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0139] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the tunnel lining void detection method provided by the above-mentioned method embodiments.

[0140] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the tunnel lining void detection method provided by the above-mentioned method embodiments is implemented.

[0141] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A tunnel lining void detection method based on data reconstruction, characterized in that: include: Obtain the acoustic vibration signal data of the tunnel lining to be tested; Preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum; Inputting the original Mel spectrum into a preset unsupervised learning model to reconstruct the original Mel spectrum to obtain a reconstructed Mel spectrum; the unsupervised learning model is trained based on normal tunnel lining acoustic vibration data; A reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum is calculated. If the reconstruction error is greater than a preset error threshold, it is considered that there is a gap in the tunnel lining.

2. The tunnel lining void detection method based on data reconstruction according to claim 1 is characterized in that: The network structure of the unsupervised learning model adopts any one of an autoencoder, a variational autoencoder and a convolutional variational autoencoder; the convolutional variational autoencoder is obtained by replacing some fully connected layers in the variational autoencoder with convolutional layers.

3. The tunnel lining void detection method based on data reconstruction according to claim 1 or 2 is characterized in that: The training process of the unsupervised learning model includes: Construct a normal tunnel lining acoustic vibration signal dataset; Using the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set as the input of the preset unsupervised learning network to generate a reconstructed Mel spectrum set; The reconstruction loss between the original Mel-spectrogram set and the reconstructed Mel-spectrogram set is minimized to iteratively optimize the parameters of the unsupervised learning network, and the unsupervised learning network with the optimal parameters is used as the unsupervised learning model.

4. The tunnel lining void detection method based on data reconstruction according to claim 3 is characterized in that: The training process of the unsupervised learning model further includes: obtaining a public sound data set; using the public sound data set to pre-train a preset unsupervised learning network to obtain an initial unsupervised learning model; Correspondingly, some network layer parameters in the initial unsupervised learning model are frozen, and the original Mel spectrum set corresponding to the normal tunnel lining acoustic vibration signal data set is used as the input of the initial unsupervised learning model to generate a reconstructed Mel spectrum set; the reconstruction loss between the original Mel spectrum set and the reconstructed Mel spectrum set is minimized to iteratively optimize the remaining network layer parameters, and the unsupervised learning network under the optimal parameters is used as the unsupervised learning model.

5. The method for detecting tunnel lining voids based on data reconstruction according to claim 3 is characterized in that: The reconstruction loss adopts a mean square error loss function or a consistency correlation coefficient loss function.

6. The tunnel lining void detection method based on data reconstruction according to claim 3 is characterized in that: Also includes: The error threshold is determined in the following manner, specifically including: For each sample in the normal tunnel lining acoustic vibration signal data set, the mean square error of the original Mel spectrum and the reconstructed Mel spectrum of each sample is calculated, and the mean square error is regarded as a random variable to fit the mean square errors corresponding to all samples to the gamma distribution, so as to obtain the error threshold.

7. The method for detecting tunnel lining voids based on data reconstruction according to claim 1 is characterized in that: The preprocessing includes: performing time-frequency conversion on the acoustic vibration signal data of the tunnel lining to be tested to obtain a spectrum diagram, and passing the spectrum diagram through a group of Mel filters to obtain an original Mel spectrum; the time-frequency conversion includes one or more of Fourier transform, short-time Fourier transform and fast Fourier transform.

8. A tunnel lining void detection device based on data reconstruction, characterized in that: include: An acquisition module is used to acquire acoustic vibration signal data of the tunnel lining to be tested; A preprocessing module, used for preprocessing the acoustic vibration signal data of the tunnel lining to be tested to obtain an original Mel spectrum; A reconstruction module, used for inputting the original Mel spectrum into a preset unsupervised learning model to reconstruct the original Mel spectrum to obtain a reconstructed Mel spectrum; the unsupervised learning model is trained based on normal tunnel lining acoustic vibration data; The detection module is used to calculate the reconstruction error between the original Mel spectrum and the reconstructed Mel spectrum. If the reconstruction error is greater than a preset error threshold, it is considered that there is a gap in the tunnel lining.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Transformer fault diagnosis method and system based on generative acoustic large model enhancement

    CN121708954A