Nondestructive pipeline risk level detection method based on audio processing

Through the lossless pipeline risk level detection method based on audio processing, the problems of single detection purposes and single feature extraction methods in the prior art are solved, effective processing of different environmental noises and accurate classification of pipeline risk levels are realized, and detection accuracy and environmental adaptability are improved.

CN120089160AActive Publication Date: 2025-06-03SUZHOU UNIV
View PDF 19 Cites 0 Cited by

Patent Information

Application Number
CN202510538735.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-06-03
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In the prior art, the detection purpose is single, the extraction feature method is single, the complexity and diversity of noise in different environments are not considered, and the imbalance of pipeline risk level samples is not considered, resulting in low classification accuracy.

Method used

The lossless pipeline risk level detection method based on audio processing is adopted, and the original audio signal is processed through isometric frame-by-frame processing and smooth processing, the industrial noise characteristic band and the residential noise characteristic band are determined, and the filter is dynamically selected for denoising. Combined with the adaptive filter and the wavelet threshold method, the time domain, frequency domain and time frequency domain characteristics are extracted, and the trained classifier is input to obtain the risk level classification results.

Benefits of technology

It improves the denoising efficiency, retains more effective audio signals, reduces the computational complexity, enhances environmental adaptability, improves risk detection accuracy, solves the problem of sample imbalance, and improves the generalization performance and classification accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089160A_ABST
    Figure CN120089160A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pipeline detection, in particular to a lossless pipeline risk level detection method based on audio processing, which comprises the following steps: if an initial noise mode of an audio sub-frame is a resident noise mode and the proportion of an industrial noise characteristic band in the audio sub-frame is greater than the proportion of a resident noise characteristic band with a preset multiple, judging that the audio sub-frame is a non-destructive pipeline risk level; the target noise mode of the audio framing is an industrial noise mode; if the initial noise mode of the audio sub-frame is an industrial noise mode and the proportion of the resident noise characteristic band in the audio sub-frame is greater than the proportion of the industrial noise characteristic band with a preset multiple, the target noise mode of the audio sub-frame is the resident noise mode; a filter corresponding to the target noise mode frequency of each audio sub-frame is selected for denoising, frame combination and feature processing are carried out after each target audio sub-frame is obtained, a feature set of a target audio signal is obtained and input into the trained classifier, and a risk level classification result of the to-be-detected pipeline is obtained. According to the invention, the pipeline risk grade classification result precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pipeline detection, and in particular to a non-destructive pipeline risk level detection method based on audio processing. Background Art

[0002] The research and application of pipeline defect detection technology play an important role in optimizing the design of urban water supply and drainage systems, improving management levels, and preventing flood risks. With the development of technology, the detection efficiency and accuracy have been continuously improved, which is beneficial to enhancing the management level of urban drainage systems and has reference significance for the detection technologies of other underground pipelines such as water supply and drainage pipelines and natural gas pipelines.

[0003] A water supply pipeline leakage detection and positioning method with the application number 2017101397727 uses audio enhancement and time delay estimation methods to judge whether there is a leakage in the pipeline. Although the noise in the signal is reduced, irreversible effects may be caused to the collected pipeline transportation sound during the enhancement process, reducing the final detection accuracy.

[0004] A signal processing method for water supply pipeline leakage detection and positioning with the application number 2009101039188 uses audio noise reduction, autocorrelation prediction comparison and other methods to judge whether there is a leakage in the pipeline. However, this method only relies on time domain correlation peak and autocorrelation entropy value analysis and does not utilize frequency domain features, resulting in insufficient feature discrimination and affecting the final detection accuracy.

[0005] A pipeline leakage detection method and system based on audio processing with the application number 2022107321722 uses frame-by-frame windowing preprocessing and multi-band spectral energy analysis to avoid directly using audio enhancement technology that may cause signal distortion and reduce the damage to the original signal; extracts spectral features through Fourier transform and calculates the energy ratio by dividing frequency bands to make up for the lack of frequency domain features; however, the detection purpose of this method is single, only judging whether the pipeline leaks, without considering multi-level risk classification, unable to guide differential maintenance strategies, and not considering the imbalance of pipeline risk level samples, which may lead to missed detection problems; only performs time domain frame-by-frame windowing preprocessing, without fully considering the complexity and diversity of noise in different environments, resulting in limited suppression effect on complex environmental noise, thus affecting the accuracy of interference signal extraction and reducing the risk detection accuracy; the audio feature extraction method is single, only obtaining the single-frame feature distribution by processing frame signals through Fourier transform and calculating the energy ratio of different frequency bands, without fully exploring the time domain, frequency domain and time-frequency domain features of audio signals, with limited feature expression ability, unable to accurately reflect the true condition of the pipeline, resulting in poor model generalization ability and reduced classification accuracy. Summary of the Invention

[0006] To this end, the technical problem to be solved by the present invention is to overcome the problems in the prior art, including single detection purpose, single feature extraction method, failure to consider the complexity and diversity of noise in different environments, and failure to consider the imbalance of pipeline risk level samples, resulting in low classification accuracy.

[0007] To solve the above technical problems, the present invention provides a non-destructive pipeline risk level detection method based on audio processing, including: Perform equal-length framing and smoothing processing on the original audio signals of the pipeline to be detected collected at each sampling moment to obtain each audio frame of the pipeline to be detected; Based on multi-window spectral estimation, determine the industrial noise characteristic band and the residential noise characteristic band in each audio frame; based on the energy ratio expression, calculate the proportion of the industrial noise characteristic band and the proportion of the residential noise characteristic band in each audio frame; Preset the initial noise mode of each audio frame; if the initial noise mode of the current audio frame is the residential noise mode, and the proportion of the industrial noise characteristic band in this audio frame is greater than the preset multiple of the proportion of the residential noise characteristic band, then the target noise mode of the current audio frame is the industrial noise mode, otherwise the target noise mode of the current audio frame is the residential noise mode; if the initial noise mode of the current audio frame is the industrial noise mode, and the proportion of the residential noise characteristic band in this audio frame is greater than the preset multiple of the proportion of the industrial noise characteristic band, then the target noise mode of the current audio frame is the residential noise mode, otherwise the target noise mode of the current audio frame is the industrial noise mode; Select a filter corresponding to the frequency of the target noise mode of each audio frame, filter out the noise signal corresponding to the target noise mode of each audio frame, and obtain each filtered audio frame; Use an adaptive filter to perform denoising processing on each filtered audio frame to obtain each target audio frame; Perform frame combination processing on all target audio frames to obtain the target audio signal of the pipeline to be detected; perform feature extraction on the target audio signal to obtain the feature set of the target audio signal; Input the feature set of the target audio signal into the trained classifier to obtain the risk level classification result of the pipeline to be detected.

[0008] Preferably, the performing equal-length framing and smoothing processing on the original audio signals of the pipeline to be detected collected at each sampling moment to obtain each audio frame of the pipeline to be detected includes: According to the preset frame length and the preset overlap rate, perform framing processing on the original audio signals of the pipeline to be detected collected at each sampling moment, and perform zero-padding processing at the end of the frames that do not meet the preset frame length to obtain each equal-length frame of the pipeline to be detected; Based on the Hanning window function, perform smoothing processing on each equal-length frame to obtain each audio frame of the pipeline to be detected.

[0009] Preferably, the initial noise pattern for each preset audio sub-frame includes: Set the initial noise pattern of the first audio sub-frame to industrial noise pattern or residential noise pattern; set the initial noise pattern of the second to the th audio sub-frame to the target noise pattern of its corresponding previous audio sub-frame.

[0010] Preferably, the selection of the filter corresponding to the target noise pattern frequency of each audio sub-frame includes: If the target noise pattern of the current audio sub-frame is industrial noise pattern, select a filter with a frequency of 50 - 100 Hz; If the target noise pattern of the current audio sub-frame is residential noise pattern, select a filter with a frequency of 1000 - 3000 Hz.

[0011] Preferably, after using the adaptive filter to denoise each filtered audio sub-frame to obtain each target audio sub-frame, it further includes: Use the wavelet threshold method to denoise each target audio signal again to obtain each denoised target audio sub-frame.

[0012] Preferably, after extracting the features of the target audio signal to obtain the feature set of the target audio signal, it further includes: Use the feature screening algorithm to obtain the feature subset of the target audio signal based on the feature set of the target audio signal.

[0013] Preferably, the training process of the classifier includes: Perform equal-length sub-framing and smoothing processing on the original audio signals of different pipelines obtained, to obtain each audio sub-frame of different pipelines, and calculate the proportion of the industrial noise feature band and the proportion of the residential noise feature band in each audio sub-frame of different pipelines; Based on the initial noise pattern of each audio sub-frame of each preset pipeline, determine the target noise pattern of each audio sub-frame of each pipeline; Select the filter corresponding to the target noise pattern frequency of each audio sub-frame of each pipeline, and combine with the adaptive filter to denoise each audio sub-frame of each pipeline to obtain each target audio sub-frame of each pipeline; Perform frame combination processing on each target audio sub-frame of each pipeline first and then perform feature extraction to obtain the feature set of the target audio signal of each pipeline; Based on the feature set of the target audio signal of different pipelines and their corresponding risk level labels, construct a data set, and divide the data set into a training set and a test set; Using the Synthetic Minority Over-sampling Technique (SMOTE) algorithm, new samples are generated between the minority class samples and their minority class neighbors in the training set through linear interpolation, and then a new training set is obtained. Using the new training set, a classifier is trained to obtain a trained classifier.

[0014] Preferably, the step of using the Synthetic Minority Over-sampling Technique (SMOTE) algorithm to generate new samples between the minority class samples and their minority class neighbors in the training set through linear interpolation and then obtaining a new training set includes: Obtain the number of samples for each class in the training set and the corresponding number of samples, and based on the number of samples of all classes, find the maximum value of the number of samples as the maximum number of samples of a class. If the number of samples of the current class is less than the maximum number of samples of a class, then use the K-d tree method to perform a nearest neighbor search for each sample of the current class to find each nearest neighbor sample of each sample of the current class, and construct a set of nearest neighbor samples for each sample of the current class. Randomly select any nearest neighbor sample from the set of nearest neighbor samples of the current sample of the current class as the target nearest neighbor sample of the current sample of the current class; perform linear interpolation between the current sample of the current class and its target nearest neighbor sample to obtain a new sample corresponding to the current sample of the current class. If the new sample corresponding to the current sample of the current class is between the minimum and maximum values of the features of the current sample of the current class, then retain the current new sample; otherwise, eliminate it. Add the current sample of the current class and its corresponding new sample to the result dataset. If the number of samples of the current class is not less than the maximum number of samples of a class, then add each sample of the current class to the result dataset. Shuffle the order of all samples in the result dataset to obtain a new training set.

[0015] Preferably, the classifier is any one of a support vector machine, a BP neural network, a convolutional neural network, a K-nearest neighbor, and a random forest.

[0016] Preferably, the risk level classification results include no risk, low risk, medium risk, and high risk.

[0017] The above technical solutions of the present invention have the following beneficial effects compared with the prior art: (1) A lossless pipeline risk level detection method based on audio processing according to the present invention performs equal-length framing processing and smoothing processing on the original audio signals of the pipeline to be detected collected at each sampling moment, obtaining each audio frame of the pipeline to be detected, which helps to respectively suppress the dominant noise in each frame during subsequent denoising processing, improve the denoising efficiency, retain more effective audio signals, and compared with processing the complete long audio signal, the data volume after framing processing is reduced, reducing the computational complexity of the denoising algorithm, and is conducive to more comprehensively and accurately analyzing the characteristics of the pipeline audio signal, improving the accuracy of judging the pipeline risk status; according to the initial noise pattern of each audio frame and the proportion of the industrial noise characteristic band and the proportion of the residential noise characteristic band in this audio frame, combined with a preset multiple, the target noise pattern of each audio frame is determined, which helps to perform targeted processing during subsequent denoising processing, accurately remove the main noise under the corresponding noise pattern, improve the noise reduction effect, make the pipeline audio signal purer, and provide a reliable basis for subsequent analysis; adjusting the processing strategy in a timely manner according to the change of the framing mode can ensure a stable and efficient state in different scenarios, enhancing its environmental adaptability; dynamically selecting filters can accurately match the noise pattern, and can filter specific frequency band noise. In an industrial noise environment, a filter targeting 50 - 100 Hz is selected to effectively attenuate the industrial equipment noise in this frequency band. In a residential noise environment, a filter adapted to 1000 - 3000 Hz is used to mainly suppress the living noise. Compared with a fixed filter, dynamic selection can remove noise more efficiently, retain the effective signals of the pipeline audio, make the audio signal purer, and improve the accuracy and effectiveness of noise reduction; through an adaptive filter, denoising processing is performed on each filtered audio frame, which can flexibly adjust the filtering method according to the characteristics of the audio frame, avoid over-damaging the key features of the audio, help subsequent analysis and understanding of the audio content, better retain the feature information related to the pipeline state, and provide support for accurately judging the risk level; wavelet threshold method is used for denoising after the adaptive filter to further eliminate the residual noise, improve the processing accuracy of each audio frame in the pipeline to be detected, and provide accurate data support for subsequent pipeline level detection.

[0018] (2) A non-destructive pipeline risk level detection method based on audio processing according to the present invention divides pipeline audio signals into four risk levels: risk-free, low-risk, medium-risk, and high-risk, and performs acquisition and subsequent classification. This is beneficial for adopting different coping strategies and has better application value. The time-domain, frequency-domain, and time-frequency domain features of the target audio signal are used as the input of the classifier. By introducing classifiers of different models, an optimal classification model is trained, and the optimal classification model is used to realize the monitoring and analysis of pipeline risks, improving the prediction accuracy of the pipeline risk level classification results. In order to further improve the correlation between the input data and the risk level detection, through the feature screening algorithm, irrelevant features that contribute less to the classification or prediction task are identified and removed, which helps to reduce the data dimension, reduce the complexity of data processing, and improve the training speed and efficiency of the subsequent model. In addition, before training the classifier, the Synthetic Minority Over-sampling Technique (SMOTE) algorithm is used, combined with the boundary protection mechanism, to increase the samples of the minority class, effectively solving the problem of unbalanced samples of the pipeline risk level of audio signals. Secondly, the K-dimensional tree data structure is used for nearest neighbor search, improving the efficiency of nearest neighbor search, especially having obvious advantages when dealing with high-dimensional data, reducing the running time of the SMOTE algorithm, and after completing the oversampling of all classes, randomly shuffling the order of each sample in the final result dataset to ensure that the data order is random, improving the universality of the samples and the generalization performance of the model, laying a foundation for further improving the prediction accuracy of the pipeline risk level classification results. Description of the Drawings

[0019] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention and in combination with the drawings, where: Figure 1 is a flowchart of a non-destructive pipeline risk level detection method based on audio processing provided by the present invention; Figure 2 is a frequency-domain comparison diagram of the original audio signal, the audio signal after denoising by dynamically selecting the corresponding filter and adaptive filter based on the noise pattern, and the audio signal after denoising by dynamically selecting the corresponding filter, adaptive filter, and wavelet threshold based on the noise pattern; where, Figure 2 (a) in represents the amplitude spectrum diagram of the original audio signal; Figure 2 (b) in represents the amplitude spectrum diagram of the audio signal after denoising by dynamically selecting the corresponding filter and adaptive filter based on the noise pattern; Figure 2 (c) in represents the amplitude spectrum diagram of the audio signal after denoising by dynamically selecting the corresponding filter, adaptive filter, and wavelet threshold based on the noise pattern; Figure 3 is a flowchart of the preprocessing of the pipeline audio signal; Figure 4 is a flowchart of the Synthetic Minority Over-sampling Technique (SMOTE) algorithm; Figure 5 is the amplitude spectrogram of pipeline audio signals with different risk levels; Figure 5 In (a) of represents the amplitude spectrogram of the pipeline audio signal with no risk level; Figure 5 In (b) of represents the amplitude spectrogram of the pipeline audio signal with a low risk level; Figure 5 In (c) of represents the amplitude spectrogram of the pipeline audio signal with a medium risk level; Figure 5 In (d) of represents the amplitude spectrogram of the pipeline audio signal with a high risk level; Figure 6 is the time-domain diagram of pipeline audio signals with different risk levels; Figure 6 In (a) of represents the time-domain diagram of the pipeline audio signal with no risk level; Figure 6 In (b) of represents the time-domain diagram of the pipeline audio signal with a low risk level; Figure 6 In (c) of represents the time-domain diagram of the pipeline audio signal with a medium risk level; Figure 6 In (d) of represents the time-domain diagram of the pipeline audio signal with a high risk level. Detailed implementation manners

[0020] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited are not intended to limit the present invention. Embodiment 1

[0021] Referring to Figure 1 as shown, Figure 1 is a flowchart of a non-destructive pipeline risk level detection method based on audio processing provided by the present invention; specifically including: S1: Perform equal-length framing and smoothing processing on the original audio signals of the pipeline to be detected collected at each sampling moment to obtain each audio frame of the pipeline to be detected, including: Perform framing processing on the original audio signals of the pipeline to be detected collected at each sampling moment according to a preset frame length and a preset overlap rate, and perform zero-padding processing on the frames that do not meet the preset frame length to obtain each equal-length frame of the pipeline to be detected; in a specific embodiment of the present invention, the preset frame length is 1024-point frame length; the preset overlap rate is 75%; Based on the Hanning window function, perform smoothing processing on each equal-length frame to obtain each audio frame of the pipeline to be detected; wherein, the expression of the Hanning window function is: ; wherein, represents the Hanning window function at the The value at a sampling moment; Indicates the time index of the Hann window function; Indicates the length of the Hann window function, i.e., the preset frame length; S2: Based on multi-taper spectral estimation, determine the industrial noise characteristic band and the residential noise characteristic band in each audio sub-frame; Based on the energy ratio expression, calculate the proportion of the industrial noise characteristic band and the proportion of the residential noise characteristic band in each audio sub-frame; Among them, in a specific embodiment of the present invention, the multi-taper spectral estimation is 4th-order Thomson multi-taper spectral estimation; The expressions for the proportion of the industrial noise characteristic band and the proportion of the residential noise characteristic band in each audio sub-frame are: ; ; Among them, Indicates the proportion of the industrial noise characteristic band in the th audio sub-frame; Indicates the industrial noise characteristic band in the th audio sub-frame; Indicates the proportion of the residential noise characteristic band in the th audio sub-frame; Indicates the residential noise characteristic band in the th audio sub-frame; Indicates the th audio sub-frame; S3: Preset the initial noise mode for each audio sub-frame; Among them, in a specific embodiment of the present invention, set the initial noise mode of the first audio sub-frame to the industrial noise mode or the residential noise mode; Set the initial noise mode of the second to the th audio sub-frames to the target noise mode of its corresponding previous audio sub-frame, and complete the preset of the initial noise mode for each audio sub-frame; If the initial noise mode of the current audio sub-frame is the residential noise mode, and the proportion of the industrial noise characteristic band in this audio sub-frame is greater than a preset multiple of the proportion of the residential noise characteristic band, then the target noise mode of the current audio sub-frame is the industrial noise mode, otherwise the target noise mode of the current audio sub-frame is the residential noise mode; If the initial noise mode of the current audio sub-frame is the industrial noise mode, and the proportion of the residential noise characteristic band in this audio sub-frame is greater than a preset multiple of the proportion of the industrial noise characteristic band, then the target noise mode of the current audio sub-frame is the residential noise mode, otherwise the target noise mode of the current audio sub-frame is the industrial noise mode; Among them, in a specific embodiment of the present invention, the preset multiple is 1.2; S4: Select a filter corresponding to the target noise pattern frequency of each audio sub-frame, filter out the noise signal corresponding to the target noise pattern of each audio sub-frame, and obtain each filtered audio sub-frame; Among them, the selection of the filter corresponding to the target noise pattern frequency of each audio sub-frame includes: if the target noise pattern of the current audio sub-frame is an industrial noise pattern, select a filter with a frequency of 50 - 100 Hz; if the target noise pattern of the current audio sub-frame is a residential noise pattern, select a filter with a frequency of 1000 - 3000 Hz; in a specific embodiment of the present invention, select a 12th-order Butterworth band-pass IIR filter corresponding to the target noise pattern frequency of each audio sub-frame; Among them, after being processed by the 12th-order Butterworth band-pass IIR filter, the signal at the th sampling moment in the filtered current audio sub-frame has the following expression: ; Among them, represents the th order feedforward coefficient, which determines the zero position of the filter; represents the th order feedback coefficient, which determines the pole position of the filter; represents the signal at the th sampling moment in the current audio sub-frame input to the Butterworth band-pass IIR filter; represents the filter order index; Among them, the transfer function of the 12th-order Butterworth band-pass IIR filter has the following expression: ; Among them, represents a variable; S5: Use an adaptive filter to perform denoising processing on each filtered audio sub-frame to obtain each target audio sub-frame; among them, in a specific embodiment of the present invention, the adaptive filter is a high-order normalized LMS (NLMS), set its order to 64, and set the step size to 0.01. Its weight update formula is: ; Among them, represents the filter weight vector corresponding to the signal at the th sampling moment in the current audio component; represents the filter weight vector corresponding to the signal at the th sampling moment in the current audio component; represents the step size factor, that is, the convergence coefficient; represents the error signal corresponding to the th sampling moment in the current audio component, ; represents the desired signal corresponding to the th sampling moment of the current audio component; represents the output signal corresponding to the th sampling moment of the current audio component, ; represents the regularization term; The above adaptive filter weight update formula realizes the adaptive update of the filter weights. Through the product term of the error signal and the input signal, the weight vector is automatically adjusted to make the output signal approximate the effective components in the desired signal; the step size factor controls the balance between the convergence speed and the steady-state error; the denominator term is the input vector energy. By normalizing it, the stability problem caused by the fixed step size of the traditional LMS is solved; when the amplitude of the input signal suddenly changes (such as in the pipeline noise scenario), the traditional LMS will diverge, while the NLMS dynamically adjusts the step size to decouple the convergence rate from the signal energy; the regularization term is to prevent the denominator from exploding when the input energy approaches zero (such as in the silent segment). The code implicitly adds a minimum value (the MATLAB default is ), ensuring numerical stability, which is a safety mechanism not available in the traditional LMS; To further eliminate the residual noise, in a specific embodiment of the present invention, the wavelet threshold method is also used to perform denoising processing on each target audio signal again to obtain each denoised target audio frame. Subsequently, each denoised target audio frame is processed, that is: cascading wavelet threshold denoising after the adaptive filter, using the multi-resolution characteristics of the wavelet to eliminate the residual noise. According to the time-frequency distribution characteristics of the pipeline audio signal, the wavelet basis function Sym8 is selected, and the decomposition layer number is automatically adjusted according to the input signal. The signal is decomposed into different scales and frequency sub-bands. The automatic adjustment formula for the wavelet decomposition layer number is: ; where, represents the length of the input audio frame, that is, the preset frame length; represents the floor operation, avoiding non-integer layer numbers to ensure strict decomposability. For short signals, the layer number is automatically reduced to avoid invalid decomposition; for long signals, the layer number is limited to 8 to balance the effect and the calculation amount; In each sub-band, the soft threshold method is used to remove the noise components. The wavelet coefficients are processed using the heuristic wavelet threshold (heursure) function method at each decomposed scale, and then the original signal is reconstructed using the estimated wavelet coefficients at each scale, retaining the effective audio features and improving the signal quality to highlight the audio details related to the pipeline risk; In a specific embodiment of the present invention, the frequency-domain comparison diagrams of the original audio signal, the audio signal after denoising by dynamically selecting the corresponding filter and the adaptive filter based on the noise pattern, and the audio signal after denoising by dynamically selecting the corresponding filter, the adaptive filter, and the wavelet threshold are as follows Figure 2 shown; it can be seen that the method adopted in the present invention not only highlights the important details of the audio signal, but also improves the filtering effect on high-frequency noise; In summary, S1-S5 are the preprocessing process of the pipeline audio signal, and its corresponding flowchart is as follows Figure 3 shown; S6: Perform frame combination processing on all target audio frames to obtain the target audio signal of the pipeline to be detected; extract the features of the target audio signal to obtain the feature set of the target audio signal; in a specific embodiment of the present invention, use the Opensmile speech feature extraction tool to extract the time-domain, frequency-domain, and time-frequency domain features of the target audio signal, and construct the feature set of the target audio signal; among them, the number of extracted audio features is 384; Since the number of extracted features is large, if a classifier needs to be trained for each feature subset, the computational overhead of the algorithm is large, the computational complexity is high, and it is difficult to implement. In a specific embodiment of the present invention, a feature screening algorithm is also used to obtain the feature subset of the target audio signal based on the feature set of the target audio signal, and then process the feature subset of the target audio signal, that is, input the feature subset of the target audio signal into the classifier; among them, the feature screening algorithm is the Relief algorithm; S7: Input the feature set of the target audio signal into the trained classifier to obtain the risk level classification result of the pipeline to be detected; among them, the classifier is any one of a support vector machine, a BP neural network, a convolutional neural network, a K-nearest neighbor, and a random forest; the risk level classification result includes no risk, low risk, medium risk, and high risk; among them, the amplitude spectrogram of the pipeline audio signal with different risk levels is as follows Figure 5 shown, and the time-domain diagram of the pipeline audio signal with different risk levels is as follows Figure 6 shown; the pipeline risk level classification standard and countermeasures are shown in Table 1.

[0022] Table 1 Pipeline risk level classification standard and countermeasures Risk level Classification criteria Countermeasures No risk Fully meet the safety requirements of the current national standards and specifications. The pipeline network is in good operating condition, and very few pipeline sections have structural safety hazards, and it is basically in a state of system safety and overall reliability. Carry out daily regular maintenance inspections, and pay attention to the impacts of external environments, adjacent surrounding activities, and weather conditions, etc. Low risk Meet the safety requirements of the current national standards and specifications. Very few pipeline sections have structural safety hazards, and it is basically in a state of system safety and overall reliability. Organize regular inspection and maintenance, and for pipelines in key areas, inspection or monitoring can be organized and implemented. Medium risk Basically meet the safety requirements of the current national standards and specifications. There are signs of deterioration and aggravation of diseases in the pipeline network, and risk events occur occasionally, which may affect system safety and overall functions. Maintenance measures should be taken for some pipelines, and regular inspection or monitoring should be strengthened. High risk Do not meet the safety requirements of the current national standards and specifications. Serious deterioration or diseases have occurred in the pipeline network, and risk events may occur, affecting system safety and overall functions. Maintenance or renovation and renewal measures should be taken for key areas, and pipelines with a diameter of DN800 and above should preferably be inspected and monitored. The training process of the classifier includes: Perform equal-length framing and smoothing processing on the original audio signals of different pipelines to obtain each audio frame of different pipelines, and calculate the proportion of the industrial noise feature band and the proportion of the residential noise feature band in each audio frame of different pipelines; Determine the target noise pattern for each audio sub-frame of each pipeline based on the initial noise pattern of each audio sub-frame of each pipeline preset. Select a filter corresponding to the frequency of the target noise pattern of each audio sub-frame of each pipeline, and combine it with an adaptive filter to denoise each audio sub-frame of each pipeline to obtain each target audio sub-frame of each pipeline. Perform frame combination processing on each target audio sub-frame of each pipeline and then perform feature extraction to obtain the target audio signal feature set of each pipeline. Based on the feature sets of the target audio signals of different pipelines and their corresponding risk level labels, construct a data set and divide the data set into a training set and a test set. Using the Synthetic Minority Over-sampling Technique (SMOTE) algorithm, through linear interpolation method, generate new samples between the minority class samples and their minority class neighbors in the training set, and then obtain a new training set. The specific process is as Figure 4 shown, including: Obtain the number of samples of each category in the training set and based on the number of samples of all categories, find the maximum value of the number of samples as the maximum category sample number. If the number of samples of the current category is less than the maximum category sample number, then use the K-d tree method to perform a nearest neighbor search on each sample of the current category to find each nearest neighbor sample of each sample of the current category, and construct a nearest neighbor sample set of each sample of the current category. Randomly select any nearest neighbor sample in the nearest neighbor sample set of the current sample of the current category as the target nearest neighbor sample of the current sample of the current category; perform linear interpolation between the current sample of the current category and its target nearest neighbor sample to obtain a new sample corresponding to the current sample of the current category. If the new sample corresponding to the current sample of the current category is between the minimum and maximum values of the features of the current sample of the current category, then retain the current new sample, otherwise eliminate it. Add the current sample of the current category and its corresponding new sample to the result data set. If the number of samples of the current category is not less than the maximum category sample number, then add each sample of the current category to the result data set. Shuffle the order of all samples in the result data set to obtain a new training set. Use the new training set to train a classifier to obtain a trained classifier. Among them, during the classifier training process, use the ten-fold cross-validation method, 90% of the data as the training set, 10% of the data as the test set, and take the average accuracy of the ten validations as the risk level detection accuracy of the classifier. It is found that the accuracy of random forest classification is the highest and the calculation time is short, which is suitable for classifying the pipeline risk level.

[0023] During the classifier training process, the Synthetic Minority Over-sampling Technique (SMOTE) algorithm is used to oversample the number of minority class samples to balance the multi-class dataset, solve the problem of unbalanced samples in the risk level of audio signal pipelines, and further improve the classification accuracy. The basic idea of the SMOTE algorithm is to analyze the minority class samples. For each minority class sample, new samples are generated by linear interpolation between it and its minority class neighbors to balance the number of samples in each class, that is, the number of samples in each class is equal to the number of samples in the largest class in the dataset. In a random forest, more minority class samples can enable the decision tree to consider the minority class more when constructing branches, thereby improving the classification performance for the minority class. The present invention improves the traditional SMOTE algorithm, specifically reflected in: (1) It has the ability to automatically process multi-class data, oversample all minority classes, balance the number of samples in each class, and significantly improve the practicality of the algorithm in multi-class scenarios; (2) Use the K-dimensional tree data structure for nearest neighbor search; the K-dimensional tree is an efficient spatial index structure that can greatly improve the efficiency of nearest neighbor search, especially has obvious advantages when dealing with high-dimensional data, and reduces the running time of the algorithm; (3) Introduce a boundary protection mechanism; after generating each synthetic sample, its feature values will be restricted between the minimum and maximum values of the original data features to ensure that the generated samples are within a reasonable range, improve the quality of the synthetic samples, and thus help to improve the performance of subsequent machine learning models; (4) After oversampling all classes, the final dataset will be randomly shuffled to ensure that the data order is random, avoid the adverse effects on model training caused by data order problems, and improve the generalization performance of the model. Embodiment 2

[0024] In order to verify a lossless pipeline risk level detection method based on audio processing provided by the present invention, more than 90 devices were installed in multiple places in Dongcheng District and Haidian District of Beijing for collection. Finally, a total of 5096 audio were collected. After denoising, the number of effective audio was 4910, and the time of each audio was 11 seconds; the effective audio signals were labeled, namely 1 - no risk, 2 - low risk, 3 - medium risk, 4 - high risk; Use the Opensmile language feature extraction tool to extract the features of the effective audio signals, and the feature configuration selects IS09_emotion; the number of features extracted from each audio is 384, and the size of the feature dataset is 4910 * 384; use the Relief algorithm to screen 50-dimensional features to obtain the dataset after feature screening; among them, the dataset after feature screening contains the feature subset of the target audio signal of each pipeline and its corresponding risk level label; The dataset after feature screening is divided into a training sample set and a test sample set. The classification process uses the ten-fold cross-validation method, with 90% of the data as the training set and 10% of the data as the test set. Before inputting the training sample set, the SMOTE algorithm is first used to balance the dataset, and then it is input into the classifier for training. When the training reaches the required accuracy or the maximum number of training times, the training stops. Among them, taking the use of a random forest as the classifier as an example, the parameter settings are as follows: the number of decision trees is 300, the minimum number of samples contained in the leaf nodes of the decision tree is 1, and the "Out-of-Bag" (OOB) samples are used to evaluate the impact of each input feature on the model prediction accuracy and the prediction process. The test set is input into the trained classifier for prediction, and the average accuracy of the ten-fold cross-validation is used as the risk level detection accuracy of the classifier. When using the improved random forest + SMOTE algorithm classifier, the accuracy can reach up to 92.46%.

[0025] In summary, the advantages of the present invention are as follows: 1. The pipeline risk level detection method based on audio processing proposed by the present invention collects and classifies audio signals according to four risk levels of risk-free, low-risk, medium-risk, and high-risk in accordance with current national standards and specifications. This is conducive to adopting different coping strategies and has better application value.

[0026] 2. The pipeline risk level detection method based on audio processing proposed by the present invention determines the target noise pattern of each audio frame, which helps to perform targeted processing in subsequent denoising, accurately remove the main noise under the corresponding noise pattern, improve the denoising effect, make the pipeline audio signal purer, and provide a reliable basis for subsequent analysis; adjusting the processing strategy in a timely manner according to the change of the frame mode can ensure a stable and efficient state in different scenarios and enhance its environmental adaptability; dynamically selecting filters can accurately match the noise pattern and filter specific frequency band noise. In an industrial noise environment, a filter targeting 50 - 100 Hz can be selected to effectively attenuate the industrial equipment noise in this frequency band. In a residential noise environment, a filter adapted to 1000 - 3000 Hz is used to mainly suppress the life noise. Compared with a fixed filter, dynamic selection can remove noise more efficiently, retain the effective pipeline audio signal, make the audio signal purer, and improve the accuracy and effectiveness of denoising.

[0027] 3. The pipeline risk level detection method based on audio processing proposed by the present invention designs an adaptive filter to filter the original audio signal, tracks the changes in the noise frequency and amplitude in real time, and effectively removes the environmental noise and the inherent vibration noise of the pipeline; the use of NLMS, combined with the variable-order filter structure and the regularization step size control, significantly improves the performance of the adaptive filtering system in a dynamic noise environment.

[0028] 4. The non-destructive pipeline risk level detection method based on audio processing proposed by the present invention uses wavelet transform to remove interference, adopts a soft threshold function, and can adaptively adjust the number of wavelet decomposition layers, effectively filtering the interference signal.

[0029] 5. The pipeline risk level detection method based on audio processing proposed by the present invention first uses an improved SMOTE algorithm to increase the minority class samples before training the classifier, effectively solving the problem of unbalanced samples of the pipeline risk level of audio signals.

[0030] 6. The pipeline risk level detection method based on audio processing proposed by the present invention realizes the monitoring and analysis of pipeline risks by introducing a deep learning network. The trained network model has a high recognition accuracy, solving the problems of missed detection and false detection caused by the current manual detection based on empirical knowledge.

[0031] Obviously, the above embodiments are only examples clearly described and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A non-destructive pipeline risk level detection method based on audio processing, characterized in that: include: The original audio signal of the pipeline to be detected collected at each sampling time is divided into equal-length frames and smoothed to obtain each audio frame of the pipeline to be detected; Based on multi-window spectrum estimation, determine the industrial noise characteristic band and the residential noise characteristic band in each audio sub-frame; Based on the energy ratio expression, calculate the proportion of industrial noise characteristic bands and the proportion of residential noise characteristic bands in each audio frame; Preset the initial noise mode of each audio subframe; if the initial noise mode of the current audio subframe is the residential noise mode, and the proportion of the industrial noise characteristic band in the audio subframe is greater than the preset multiple of the proportion of the residential noise characteristic band, then the target noise mode of the current audio subframe is the industrial noise mode, otherwise the target noise mode of the current audio subframe is the residential noise mode; if the initial noise mode of the current audio subframe is the industrial noise mode, and the proportion of the residential noise characteristic band in the audio subframe is greater than the preset multiple of the proportion of the industrial noise characteristic band, then the target noise mode of the current audio subframe is the residential noise mode, otherwise the target noise mode of the current audio subframe is the industrial noise mode; Selecting a filter corresponding to the target noise pattern frequency of each audio subframe, filtering out the noise signal corresponding to the target noise pattern of each audio subframe, and obtaining each audio subframe after filtering; Adopting an adaptive filter to perform denoising on each filtered audio sub-frame to obtain each target audio sub-frame; All target audio frames are combined to obtain the target audio signal of the pipeline to be detected; features are extracted from the target audio signal to obtain a feature set of the target audio signal; The feature set of the target audio signal is input into the trained classifier to obtain the risk level classification result of the pipeline to be detected.

2. According to the audio processing-based non-destructive pipeline risk level detection method of claim 1, it is characterized in that: The step of performing equal-length frame division and smoothing processing on the original audio signal of the pipeline to be detected collected at each sampling moment to obtain each audio frame of the pipeline to be detected includes: According to the preset frame length and the preset overlap rate, the original audio signal of the pipeline to be detected collected at each sampling time is framed, and the end of the frame that does not meet the preset frame length is zero-filled to obtain each frame of equal length of the pipeline to be detected; Based on the Hanning window function, each equal-length sub-frame is smoothed to obtain each audio sub-frame of the pipeline to be detected.

3. According to the non-destructive pipeline risk level detection method based on audio processing according to claim 1, it is characterized in that: The preset initial noise mode of each audio frame includes: The initial noise mode of the first audio frame is set to industrial noise mode or residential noise mode; the initial noise mode of the second to third audio frames is set to industrial noise mode or residential noise mode. The initial noise pattern of an audio sub-frame is set to the target noise pattern of its corresponding previous audio sub-frame.

4. According to the audio processing-based non-destructive pipeline risk level detection method of claim 1, it is characterized in that: The selecting of a filter corresponding to a target noise pattern frequency for each audio frame comprises: If the target noise mode of the current audio frame is the industrial noise mode, a filter with a frequency of 50-100 Hz is selected; If the target noise mode of the current audio frame is the residential noise mode, a filter with a frequency of 1000-3000 Hz is selected.

5. According to the audio processing-based non-destructive pipeline risk level detection method of claim 1, it is characterized in that: The method further comprises: using an adaptive filter to perform denoising on each filtered audio subframe to obtain each target audio subframe; The wavelet threshold method is used to perform denoising on each target audio signal again to obtain each target audio frame after denoising.

6. The non-destructive pipeline risk level detection method based on audio processing according to claim 1 is characterized in that: After extracting the features of the target audio signal to obtain the feature set of the target audio signal, the method further includes: A feature subset of the target audio signal is obtained based on the feature set of the target audio signal using a feature screening algorithm.

7. The non-destructive pipeline risk level detection method based on audio processing according to claim 1 is characterized in that: The classifier training process includes: The original audio signals of different pipelines are divided into equal-length frames and smoothed to obtain the audio frames of different pipelines, and the proportion of industrial noise characteristic bands and the proportion of residential noise characteristic bands in each audio frame of different pipelines are calculated; Determine a target noise pattern for each audio frame of each pipeline based on a preset initial noise pattern for each audio frame of each pipeline; Select a filter corresponding to the target noise pattern frequency of each audio frame of each pipeline, and combine it with an adaptive filter to denoise each audio frame of each pipeline to obtain each target audio frame of each pipeline; Each target audio frame of each pipeline is first combined and then feature extracted to obtain the target audio signal feature set of each pipeline; Based on the feature sets of target audio signals from different pipelines and their corresponding risk level labels, a dataset is constructed and divided into a training set and a test set. Using the synthetic minority oversampling algorithm, new samples are generated between the minority samples in the training set and their minority neighbors through linear interpolation method, thus obtaining a new training set. Use the new training set to train the classifier and obtain a trained classifier.

8. The non-destructive pipeline risk level detection method based on audio processing according to claim 7 is characterized in that: The synthetic minority oversampling technique algorithm is used to generate new samples between the minority class samples and their minority class neighbors in the training set by linear interpolation method, and then the new training set is obtained, including: Get each category and its corresponding number of samples in the training set, and find the maximum number of samples based on the number of samples of all categories as the maximum number of category samples; If the number of samples in the current category is less than the maximum number of samples in the category, the K-dimensional tree method is used to perform a neighbor search on each sample in the current category, find each neighbor sample of each sample in the current category, and construct a neighbor sample set of each sample in the current category; Randomly select any neighboring sample in the neighboring sample set of the current sample of the current category as the target neighboring sample of the current sample of the current category; perform linear interpolation between the current sample of the current category and its target neighboring sample to obtain a new sample corresponding to the current sample of the current category; If the new sample corresponding to the current sample of the current category is between the minimum and maximum values ​​of the current sample feature of the current category, the current new sample is retained, otherwise it is discarded; Add the current sample of the current category and its corresponding new sample to the result data set; If the number of samples in the current category is not less than the maximum number of samples in the category, then each sample in the current category is added to the result data set; The order of all samples in the result data set is disrupted to obtain a new training set.

9. The non-destructive pipeline risk level detection method based on audio processing according to claim 1 is characterized in that: The classifier is any one of support vector machine, BP neural network, convolutional neural network, K-nearest neighbor, and random forest.

10. The non-destructive pipeline risk level detection method based on audio processing according to claim 1 is characterized in that: The risk level classification results include no risk, low risk, medium risk and high risk.

Citation Information

Patent Citations

  • Expressway audio vehicle detection device and method thereof

    CN102682765A

  • 3D audio encoding acceleration method based on assembly line

    CN105206278A

  • Variable rate vocoder

    CN1071036A

  • Detecting method for sewer pipeline failures

    CN107355687A

  • Risk early warning system based on big data

    CN109636585A