Natural gas tiny leakage MFCC detection system and method based on voiceprint coupling
By combining multimodal fusion detection method with voiceprint and infrared image data in the natural gas detection system, using short-time Fourier transform and neural network model, the problems of low detection efficiency of natural gas micro leakage and high detection rate in the prior art are solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510353346.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-05-30
AI Technical Summary
The existing natural gas micro leak detection technology has the problems of low detection efficiency, poor detection effect, and prone to missed detection and missed detection, especially in complex environments, it is difficult to effectively detect micro leaks.
The MFCC detection system and method for micro-leakage of natural gas based on voiceprint coupling is adopted. By arranging a microphone array and infrared image acquisition device in the area to be detected, voiceprint data and infrared image data are collected, and combined with short-time Fourier transform and neural network model, multimodal fusion detection and risk assessment of abnormal voiceprints and infrared features are realized.
It improves the accuracy and reliability of natural gas micro leakage detection, overcomes the limitations of single soundprint detection being susceptible to environmental noise interference and missed detection, and improves the sensitivity and stability of micro leakage detection in complex environments.
Smart Images

Figure CN120062565A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pipeline detection, and in particular, to a MFCC detection system and method for natural gas micro-leakage based on acoustic coupling. Background Art
[0002] As a clean and efficient energy source, natural gas is widely used in industrial production and residential life. However, due to the flammable and explosive characteristics of natural gas, its safe transportation, storage and use pose high requirements for detection technologies. Especially during the operation of natural gas pipelines, valves and equipment, micro-leakages are often difficult to be detected in time by traditional detection technologies due to weak sound signals and small leakage amounts. Although the leakage amount of such micro-leakages is small, if they are not detected for a long time, cumulative effects may occur and lead to serious safety accidents.
[0003] At present, natural gas leakage monitoring technologies based on acoustic detection have been preliminarily applied. They utilize acoustic signals to capture specific spectral characteristics of leaks, and are a non-contact and rapid-response monitoring means. However, such technologies still have certain limitations: in complex industrial or outdoor environments, it is necessary for personnel to carry detection equipment for manual detection, and the acoustic characteristics of micro-leakages will be masked in complex environments, resulting in high false detection rates and missed detection rates. And only relying on acoustic signal analysis, lacking collaborative verification with other information, it is difficult to achieve high-sensitivity detection in diverse leakage scenarios.
[0004] Therefore, it is necessary to design a MFCC detection system and method for natural gas micro-leakage based on acoustic coupling to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes a MFCC detection system and method for natural gas micro-leakage based on acoustic coupling, aiming to solve the problems of low detection efficiency, poor detection effect, and easy occurrence of false detection and missed detection in the current detection of natural gas micro-leakage.
[0006] On the one hand, the present invention proposes a MFCC detection method for natural gas micro-leakage based on acoustic coupling, including:
[0007] Deploy a microphone array and an infrared image acquisition device in the area to be detected, collect the acoustic data of the microphone array within a preset time period, process the acoustic data according to the microphone frequency to determine the acoustic feature data at each moment, and obtain an acoustic feature data set;
[0008] Compare the acoustic feature data set with the abnormal range data, and determine a primary judgment flag according to the comparison result. The primary judgment flag includes a judgment abnormal flag, a suspected abnormal flag, and a judgment non-abnormal flag;
[0009] When the initial judgment identifier is a suspected abnormal identifier, perform a short-time Fourier transform on the voiceprint characterization data set to obtain the amplitude information of each frame, normalize the amplitude information, and map the normalized amplitude information to an RGB color table to generate a two-dimensional image; extract the infrared image data of the area to be detected at the same moment, and input the two-dimensional image and the infrared image data into a pre-trained neural network model to determine the risk coefficient;
[0010] Compare the risk coefficient with a risk threshold, and determine whether to issue a risk warning according to the comparison result.
[0011] Further, when processing the voiceprint data according to the microphone frequency to determine the voiceprint characterization data at each moment, it includes:
[0012] Use a Gaussian mixture distribution to determine the voiceprint characterization data at each moment:
[0013] The expression of the kernel density function is:
[0014]
[0015] where n represents the number of voiceprint data collected at each moment, h represents the smoothing bandwidth, represents the i-th voiceprint data collected at each moment, and x represents the value of a certain frequency point to be estimated;
[0016] Take the voiceprint data with the highest kernel density frequency as the voiceprint characterization data.
[0017] Further, when processing the voiceprint data according to the microphone frequency to determine the voiceprint characterization data at each moment, it also includes:
[0018] where the smoothing bandwidth is calculated by the following formula:
[0019]
[0020] where h represents the smoothing bandwidth, σ represents the standard deviation of the voiceprint data collected at each moment, β1 represents the data skewness, β2 represents the data kurtosis, n represents the number of voiceprint data collected at each moment, and k represents an adjustment coefficient;
[0021] where the data skewness is calculated by the following formula:
[0022]
[0023] The data kurtosis is calculated by the following formula:
[0024]
[0025] Wherein, n represents the number of voiceprint data collected at each moment, represents the i-th voiceprint data collected at each moment, represents the mean value of the voiceprint data collected at each moment, and σ represents the standard deviation of the voiceprint data collected at each moment.
[0026] Furthermore, the abnormal range data includes:
[0027] Collect the normal operation records corresponding to the normal operation of the pipeline in the area to be detected, and extract the corresponding normal voiceprint amplitude range from the normal operation records;
[0028] Compare the standard voiceprint amplitude of the pipeline with the normal voiceprint amplitude range, and generate the abnormal range data according to the numerical size relationship, wherein the abnormal range data includes a left boundary value and a right boundary value.
[0029] Furthermore, when determining the initial judgment flag according to the comparison result, it includes:
[0030] When the voiceprint amplitudes of all voiceprint characterization data in the voiceprint characterization dataset are within the abnormal range data, generate a non-abnormal judgment flag for the area to be detected;
[0031] When the voiceprint amplitudes of all voiceprint characterization data in the voiceprint characterization dataset are greater than the right boundary, generate an abnormal judgment flag for the area to be detected;
[0032] When there are voiceprint amplitudes of voiceprint characterization data within the abnormal range data in the voiceprint characterization dataset, and there are voiceprint amplitudes of voiceprint characterization data greater than the right boundary, generate a suspected abnormal flag for the area to be detected.
[0033] Furthermore, when the initial judgment flag is an abnormal judgment flag, it includes:
[0034] Generate an amplitude excess value according to the maximum voiceprint amplitude of the voiceprint characterization data in the voiceprint characterization dataset and the right boundary, the amplitude excess value is the difference between the maximum voiceprint amplitude and the right boundary, and determine the risk coefficient according to the amplitude excess value, and the risk coefficient is in a proportional relationship with the amplitude excess value.
[0035] Furthermore, when inputting the two-dimensional image and the infrared image data into a pre-trained neural network model to determine the risk coefficient, it includes:
[0036] Collect historical abnormal data to construct a historical dataset, and the historical abnormal data includes historical abnormal two-dimensional images and historical abnormal infrared image data;
[0037] Sample the historical dataset according to a preset ratio to obtain a training subset and a test subset;
[0038] Obtain a pre-selected neural network model, iteratively train the neural network model according to the training subset, evaluate the iteratively trained neural network model according to the test subset, and determine whether to stop the iterative training according to the evaluation value;
[0039] If the evaluation value of the neural network model after the current iterative training is less than the evaluation value of the neural network model after the previous iterative training, then reduce the amplitude of the change of the neural network model in the gradient direction, and continue the iterative training until the preset number of iterations is reached; if the evaluation value of the neural network model after the current iterative training is greater than or equal to the evaluation value of the neural network model after the previous iterative training, then stop the iterative training to obtain the pre-trained neural network model;
[0040] Input the two-dimensional image and the infrared image data into the pre-trained neural network model to obtain the risk coefficient.
[0041] Further, when comparing the risk coefficient with a risk threshold and determining whether to issue a risk warning according to the comparison result, it includes:
[0042] When the risk coefficient is greater than the risk threshold, it is determined that a risk warning is issued;
[0043] When the risk coefficient is less than or equal to the risk threshold, it is determined that no risk warning is issued.
[0044] Further, when it is determined that a risk warning is issued, it includes:
[0045] Obtain a risk coefficient difference according to the risk coefficient and the risk threshold, where the risk coefficient difference is the difference between the risk coefficient and the risk threshold, compare the risk coefficient difference with a first risk coefficient and a second risk coefficient respectively, and determine the risk level according to the comparison result; the first risk coefficient is less than the second risk coefficient;
[0046] When the risk coefficient difference is less than or equal to the first risk coefficient, determine that the risk level is the first risk level; when the risk coefficient difference is greater than the first risk coefficient and less than or equal to the second risk coefficient, determine that the risk level is the second risk level; when the risk coefficient difference is greater than the second risk coefficient, determine that the risk level is the third risk level; the first risk level is less than the second risk level, and the second risk level is less than the third risk level.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows: By arranging a microphone array and an infrared image acquisition device in the area to be detected, a multi-modal fusion detection system for voiceprint data and infrared image data is formed, which improves the detection accuracy and reliability of natural gas micro-leakage. The voiceprint data is collected by the microphone array and the voiceprint characterization data at each moment is extracted to realize the screening of abnormal voiceprints and identify the characteristic voiceprints. In case of suspected abnormality, the voiceprint amplitude information is further obtained by using the short-time Fourier transform, and a two-dimensional image is generated through normalization processing and color mapping, making the voiceprint features more intuitive and easier for subsequent analysis. At the same time, combined with the infrared image data at the same time, the voiceprint features and infrared features are input into a pre-trained neural network model to fully explore the spatial and temporal correlation of leakage features, calculate the risk coefficient and compare it with the risk threshold to realize risk assessment and early warning. It overcomes the limitations of single voiceprint detection being easily interfered by environmental noise and missed detection, and at the same time improves the sensitivity and stability of micro-leakage detection in complex environments through multi-modal fusion.
[0048] On the other hand, the present application also provides a natural gas micro-leakage MFCC detection system based on voiceprint coupling for applying the above-mentioned natural gas micro-leakage MFCC detection method based on voiceprint coupling, including:
[0049] A sensor module, including a microphone array and an infrared image acquisition device;
[0050] An acquisition unit, configured to collect the voiceprint data of the microphone array within a preset time period, process the voiceprint data according to the microphone frequency to determine the voiceprint characterization data at each moment, and obtain a voiceprint characterization data set;
[0051] A judgment unit, configured to compare the voiceprint characterization data set with the abnormal range data, and determine a primary judgment identifier according to the comparison result, where the primary judgment identifier includes a judgment abnormal identifier, a suspected abnormal identifier, and a judgment non-abnormal identifier;
[0052] A processing unit, configured to perform a short-time Fourier transform on the voiceprint characterization data set to obtain the amplitude information of each frame when the primary judgment identifier is a suspected abnormal identifier, perform normalization processing on the amplitude information, and map the normalized amplitude information to the RGB color table to generate a two-dimensional image; extract the infrared image data of the area to be detected at the same moment, and input the two-dimensional image and the infrared image data into a pre-trained neural network model to determine the risk coefficient;
[0053] An early warning unit, configured to compare the risk coefficient with the risk threshold, and determine whether to perform a risk early warning according to the comparison result.
[0054] It is understandable that the above-mentioned MFCC detection system and method for natural gas micro-leakage based on voiceprint coupling have the same beneficial effects, which will not be elaborated here. Description of the Drawings
[0055] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered as a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0056] Figure 1 is a flowchart of the MFCC detection method for natural gas micro-leakage based on voiceprint coupling provided by an embodiment of the present invention;
[0057] Figure 2 is a structural block diagram of the MFCC detection system for natural gas micro-leakage based on voiceprint coupling provided by an embodiment of the present invention. Detailed Embodiments
[0058] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0059] In some embodiments of the present application, referring to Figure 1 as shown, a MFCC detection method for natural gas micro-leakage based on voiceprint coupling includes:
[0060] S100: Arrange a microphone array and an infrared image acquisition device in the area to be detected, collect the voiceprint data of the microphone array within a preset time period, process the voiceprint data according to the microphone frequency to determine the voiceprint characterization data at each moment, and obtain a voiceprint characterization data set.
[0061] S200: Compare the voiceprint characterization data set with the abnormal range data, and determine the initial judgment identifier according to the comparison result. The initial judgment identifier includes a judgment abnormal identifier, a suspected abnormal identifier, and a judgment non-abnormal identifier.
[0062] S300: When the initial judgment flag is a suspected abnormal flag, perform a short-time Fourier transform on the voiceprint characterization data set to obtain the amplitude information of each frame, normalize the amplitude information, and map the normalized amplitude information to the RGB color table to generate a two-dimensional image. Extract the infrared image data of the area to be detected at the same moment, and input the two-dimensional image and the infrared image data into a pre-trained neural network model to determine the risk coefficient.
[0063] S400: Compare the risk coefficient with the risk threshold, and determine whether to issue a risk warning according to the comparison result.
[0064] Specifically, in S100, a microphone array and an infrared image acquisition device are arranged in the area to be detected, and the microphone array collects voiceprint data within a preset time period through multiple channels. Utilize the frequency response characteristics of the microphone to process the collected voiceprint data, including denoising, spectrum analysis, and feature extraction, so as to determine the voiceprint characterization data at each moment and form a voiceprint characterization data set with time and frequency information. This step aims to capture the specific voiceprint characteristics generated during natural gas leakage. In S200, compare the extracted voiceprint characterization data set with the preset abnormal range data. The abnormal range data is a standard library set according to the voiceprint characteristic range of natural gas leakage in the actual scenario. The comparison result can generate three initial judgment flags: Judgment abnormal flag: The voiceprint characteristics significantly conform to the leakage characteristics, and it is directly judged as leakage. Suspected abnormal flag: The voiceprint characteristics partially conform to the leakage characteristics, but it is not completely certain. Judgment non-abnormal flag: The voiceprint characteristics completely do not conform to the leakage characteristics, excluding the possibility of abnormality. Through the initial judgment flag, quickly screen out the suspected abnormal situations, avoid repeated processing of normal environments or obvious leaks, and improve the detection efficiency. In S300, for the situation where the initial judgment is a "suspected abnormal flag", further process the voiceprint characterization data set. The specific operations include: Short-time Fourier transform: Decompose the voiceprint signal into time-frequency domain information, obtain the amplitude information of each frame, and strengthen the description of frequency characteristics. Normalize the amplitude information to eliminate the influence of the amplitude range difference on the result. Subsequently, map the normalized data to the RGB color table to generate a visual two-dimensional image. Infrared image extraction: Synchronously obtain the infrared image data of the area to be detected at the same moment, and capture the infrared leakage signal characteristics, such as temperature changes and thermal disturbances. Input the generated two-dimensional image and the infrared image data into a pre-trained neural network model together. The model performs deep feature extraction and correlation analysis based on the multimodal data and outputs a risk coefficient to quantify the possibility and severity of leakage. In S400, compare the risk coefficient output by the neural network model with the preset risk threshold: If the risk coefficient is higher than the threshold, trigger a risk warning, indicating a high possibility of natural gas leakage. If the risk coefficient is lower than the threshold, continue monitoring without triggering a warning.
[0065] It can be understood that for MFCC (Mel Frequency Cepstral Coefficients), in the process of generating a two-dimensional image in this embodiment, the voiceprint data first undergoes the same feature processing as MFCC. Especially in the preprocessing stage, the signal is framed, subjected to FFT transformation, and spectral feature extraction, thereby constructing two-dimensional spectral data. Traditional MFCC only outputs as a set of feature vectors, mainly used as feature input for machine learning models. In this solution, these spectral features are further expanded and visualized to generate a two-dimensional image. The short-time Fourier transform is used to replace the Mel filter in MFCC, directly extracting the dynamic amplitude changes of time and frequency, retaining more high-dimensional information. And amplitude normalization and color information are added to the spectral data of MFCC, converting the features originally represented in numerical form into an intuitive RGB image.
[0066] It can be understood that through the multi-modal fusion detection of voiceprint and infrared images, the high sensitivity of the microphone array is used to capture the voiceprint features generated by natural gas leakage, and combined with the spatial information and thermal characteristics of the infrared image, the detection of micro natural gas leakage is realized. The short-time Fourier transform is used to convert the voiceprint data into a two-dimensional image, and the visualization of the leakage features is enhanced through color mapping, and effectively coupled with the infrared image in the time and space dimensions. The neural network model improves the recognition ability of micro leakage in complex environments through deep learning and analysis of multi-modal features. It solves the problem of high false detection rate in traditional voiceprint detection and also provides risk warning ability through quantitative evaluation of risk coefficients.
[0067] In some embodiments of the present application, when processing the voiceprint data according to the microphone frequency to determine the voiceprint characterization data at each moment, it includes:
[0068] Using Gaussian mixture distribution to determine the voiceprint characterization data at each moment:
[0069] The expression of the kernel density function is:
[0070]
[0071] where n represents the number of voiceprint data collected at each moment, h represents the smoothing bandwidth, represents the i-th voiceprint data collected at each moment, and x represents the value of a certain frequency point to be estimated.
[0072] The voiceprint data with the highest kernel density frequency is used as the voiceprint characterization data.
[0073] In some embodiments of the present application, when processing the voiceprint data according to the microphone frequency to determine the voiceprint characterization data at each moment, it further includes:
[0074] The smoothing bandwidth is calculated by the following formula:
[0075]
[0076] where h represents the smoothing bandwidth, σ represents the standard deviation of the voiceprint data collected at each moment, β1 represents the data skewness, β2 represents the data kurtosis, n represents the number of voiceprint data collected at each moment, and k represents the adjustment coefficient.
[0077] The data skewness is calculated by the following formula:
[0078]
[0079] The data kurtosis is calculated by the following formula:
[0080]
[0081] where n represents the number of voiceprint data collected at each moment, represents the i-th voiceprint data collected at each moment, represents the mean value of the voiceprint data collected at each moment, and σ represents the standard deviation of the voiceprint data collected at each moment.
[0082] It can be understood that the frequency point corresponding to the highest density value is found through kernel density estimation, and the voiceprint data at this frequency point is used as the voiceprint characterization data at the current moment. The main acoustic features are screened out, and at the same time, the interference of noise or outliers is eliminated. Through the kernel density estimation of Gaussian mixture distribution and the calculation of adaptive smoothing bandwidth, more accurate feature extraction of voiceprint data is achieved. Compared with the fixed bandwidth method, the introduction of adaptive bandwidth can dynamically adjust the estimation parameters according to the distribution of voiceprint data, so that the density curve can capture important features without being interfered by noise or outliers. At the same time, using skewness and kurtosis to quantify the data distribution form can effectively adapt to the changes of voiceprint characteristics in different scenarios. It can extract representative voiceprint characterization data from a large amount of voiceprint data in a complex environment, improving the accuracy of subsequent analysis and recognition.
[0083] In some embodiments of the present application, the abnormal range data includes: collecting the normal operation records corresponding to the normal operation of the pipeline in the area to be detected, and extracting the corresponding normal voiceprint amplitude range from the normal operation records. Comparing the standard voiceprint amplitude of the pipeline with the normal voiceprint amplitude range, and generating abnormal range data according to the numerical size relationship, where the abnormal range data includes the left boundary value and the right boundary value.
[0084] In some embodiments of the present application, when determining the initial judgment identifier according to the comparison result, it includes: when the voiceprint amplitude values of all voiceprint characterization data in the voiceprint characterization dataset are within the abnormal range data, a non-abnormal judgment identifier is generated for the area to be detected. When the voiceprint amplitude values of all voiceprint characterization data in the voiceprint characterization dataset are greater than the right boundary, an abnormal judgment identifier is generated for the area to be detected. When there are voiceprint amplitude values of voiceprint characterization data in the voiceprint characterization dataset within the abnormal range data and there are voiceprint amplitude values of voiceprint characterization data greater than the right boundary, a suspected abnormal identifier is generated for the area to be detected.
[0085] It can be understood that under the normal operating state of the pipeline, through long-term monitoring by a microphone array, a large amount of voiceprint data is collected, and the voiceprint amplitude range at different times and under different working conditions is recorded. Statistical analysis is performed on the voiceprint data in the normal operation record to calculate its amplitude range. For example, the "normal voiceprint amplitude range" is defined by the maximum value and the minimum value. The normal voiceprint amplitude range is compared with the standard voiceprint amplitude of the pipeline to further optimize the boundary value. By introducing the abnormal range data and the initial judgment identifier, the problems of high false positive rate and high false negative rate in traditional natural gas leakage detection are effectively solved. Compared with the simple single-threshold detection method, using the statistical characteristics of the normal operation record to construct the abnormal range data makes the determination of the voiceprint amplitude more scientific and more robust. At the same time, by independently marking the suspected abnormal state, potential small leakage risks can be discovered in advance, providing a reliable input basis for subsequent precise analysis and recognition by the deep learning model. The sensitivity and reliability of detecting small natural gas leaks are improved.
[0086] In some embodiments of the present application, when the initial judgment identifier is an abnormal judgment identifier, it includes: generating an amplitude exceedance value according to the maximum voiceprint amplitude of the voiceprint characterization data in the voiceprint characterization dataset and the right boundary, where the amplitude exceedance value is the difference between the maximum voiceprint amplitude and the right boundary, and determining the risk coefficient according to the amplitude exceedance value. The risk coefficient is in a direct proportional relationship with the amplitude exceedance value.
[0087] It can be understood that by introducing the calculation of the amplitude exceedance value and the risk coefficient, the voiceprint amplitude information in the abnormal judgment identifier is further quantified into a specific risk level. Compared with the traditional single-threshold early warning mode, it can accurately describe the size of the leakage scale and provide an accurate risk assessment basis. Based on the quantization process of the amplitude difference, the sensitivity and accuracy of the detection are enhanced, realizing a seamless connection from detection to assessment, and providing numerical support for the implementation of multi-level early warning and emergency response.
[0088] In some embodiments of the present application, when inputting the two-dimensional image and the infrared image data into a pre-trained neural network model to determine the risk coefficient, it includes:
[0089] Collect historical abnormal data to construct a historical dataset, where the historical abnormal data includes historical abnormal two-dimensional images and historical abnormal infrared image data.
[0090] Sample the historical dataset according to a preset ratio to obtain a training subset and a test subset.
[0091] Obtain a pre-selected neural network model, perform iterative training on the neural network model according to the training subset, evaluate the iteratively trained neural network model according to the test subset, and determine whether to stop the iterative training based on the evaluation value.
[0092] If the evaluation value of the neural network model after the current iterative training is less than the evaluation value of the neural network model after the previous iterative training, then reduce the amplitude of the change of the neural network model in the gradient direction and continue the iterative training until the preset number of iterations is reached. If the evaluation value of the neural network model after the current iterative training is greater than or equal to the evaluation value of the neural network model after the previous iterative training, then stop the iterative training to obtain a pre-trained neural network model.
[0093] Input the two-dimensional image and the infrared image data into the pre-trained neural network model to obtain a risk coefficient.
[0094] It is understandable that historical abnormal data (including historical abnormal two-dimensional images and historical abnormal infrared images) are collected to construct a historical dataset with rich features. These historical data truly reflect the characteristics of natural gas leakage of different types and scales, providing a basis for the training of neural networks. The historical dataset is divided into a training subset and a test subset according to a preset ratio (for example, 80% for training and 20% for testing). The training subset is used for the learning of the model, and the test subset is used to verify the model performance. A pre-designed or mature neural network model (such as a convolutional neural network) is selected, which is suitable for feature extraction and pattern recognition tasks of two-dimensional images and infrared images. The training subset is used to optimize the parameters of the neural network model, and the model parameters are adjusted through the gradient descent algorithm to minimize the loss function. After each iterative training, the performance of the model is evaluated using the test subset, usually with the loss value or accuracy as the evaluation index. If the evaluation value of the current model is less than the evaluation value of the previous iterative training, it indicates that the model performance has not reached the optimum, and the change amplitude of the model in the gradient direction (i.e., the learning rate is adjusted) is reduced to avoid excessive jumping. If the evaluation value of the current model is greater than or equal to the previous evaluation value, it indicates that the model performance has converged, and the iteration is stopped to obtain the final pre-trained model. Through the construction of the historical dataset and the refined neural network training process, the accurate quantification of natural gas leakage risk is achieved. Compared with the traditional single-signal analysis method, the use of the fusion of multi-modal data (two-dimensional acoustic fingerprint images and infrared images) improves the detection accuracy and robustness of the model. By dynamically adjusting the training strategy (such as reducing the change amplitude in the gradient direction), the convergence speed and stability of the model are optimized, ensuring the reliability and efficiency of risk assessment.
[0095] In some embodiments of the present application, when comparing the risk coefficient with the risk threshold and determining whether to issue a risk warning according to the comparison result, it includes: when the risk coefficient is greater than the risk threshold, it is determined to issue a risk warning. When the risk coefficient is less than or equal to the risk threshold, it is determined not to issue a risk warning.
[0096] In some embodiments of the present application, when it is determined to issue a risk warning, it includes: obtaining a risk coefficient difference according to the risk coefficient and the risk threshold, where the risk coefficient difference is the difference between the risk coefficient and the risk threshold, and comparing the risk coefficient difference with the first risk coefficient and the second risk coefficient respectively, and determining the risk level according to the comparison result. The first risk coefficient is less than the second risk coefficient.
[0097] Specifically, when the risk coefficient difference is less than or equal to the first risk coefficient, the risk level is determined to be the first risk level. When the risk coefficient difference is greater than the first risk coefficient and less than or equal to the second risk coefficient, the risk level is determined to be the second risk level. When the risk coefficient difference is greater than the second risk coefficient, the risk level is determined to be the third risk level. The first risk level is less than the second risk level, and the second risk level is less than the third risk level.
[0098] It can be understood that by introducing an evaluation mechanism of multiple risk levels, compared with the traditional single-threshold early warning, the risk degree of natural gas leakage can be judged more precisely. When the risk coefficient exceeds the threshold, multiple risk levels are divided according to the size of the risk coefficient difference, so as to provide more accurate guidance for subsequent emergency responses. The refined early warning method helps to timely identify and respond to potential safety hazards in different leakage scenarios, avoiding false alarms or missed alarms that may occur in traditional methods, and improving safety and reliability.
[0099] In the above embodiments, by arranging a microphone array and an infrared image acquisition device in the area to be detected, a multi-modal fusion detection system of voiceprint data and infrared image data is formed, which improves the detection accuracy and reliability of natural gas micro-leakage. The voiceprint data is collected by the microphone array and the voiceprint characterization data at each moment is extracted to realize the screening of abnormal voiceprints and identify the characteristic voiceprints. In case of suspected abnormality, the voiceprint amplitude information is further obtained by using the short-time Fourier transform, and a two-dimensional image is generated through normalization processing and color mapping, making the voiceprint features more intuitive and easy for subsequent analysis. At the same time, combined with the infrared image data at the same time, the voiceprint features and infrared features are input into a pre-trained neural network model to fully explore the spatial and temporal correlation of leakage features, calculate the risk coefficient and compare it with the risk threshold to realize risk assessment and early warning. It overcomes the limitations of single voiceprint detection being easily interfered by environmental noise and missed detection, and at the same time improves the sensitivity and stability of micro-leakage detection in complex environments through multi-modal fusion.
[0100] In another preferred manner based on the above embodiments, refer to Figure 2 As shown, this embodiment provides a MFCC detection system for natural gas micro-leakage based on voiceprint coupling, which is used to apply the above-mentioned MFCC detection method for natural gas micro-leakage based on voiceprint coupling, and includes:
[0101] A sensor module, including a microphone array and an infrared image acquisition device;
[0102] An acquisition unit, configured to collect the voiceprint data of the microphone array within a preset time period, process the voiceprint data according to the microphone frequency to determine the voiceprint characterization data at each moment, and obtain a voiceprint characterization data set;
[0103] A judgment unit, configured to compare the voiceprint characterization data set with the abnormal range data, and determine a primary judgment identifier according to the comparison result, where the primary judgment identifier includes a judgment abnormal identifier, a suspected abnormal identifier, and a judgment non-abnormal identifier;
[0104] A processing unit, configured to, when the primary judgment identifier is a suspected abnormal identifier, perform a short-time Fourier transform on the voiceprint characterization data set to obtain the amplitude information of each frame, perform a normalization process on the amplitude information, and map the normalized amplitude information to an RGB color table to generate a two-dimensional image; extract the infrared image data of the area to be detected at the same moment, and input the two-dimensional image and the infrared image data into a pre-trained neural network model to determine the risk coefficient;
[0105] An early warning unit, configured to compare the risk coefficient with a risk threshold, and determine whether to issue a risk warning according to the comparison result.
[0106] It can be understood that by arranging a microphone array and an infrared image acquisition device in the area to be detected, a multi-modal fusion detection system of voiceprint data and infrared image data is formed, which improves the detection accuracy and reliability of natural gas micro-leakage. The voiceprint data is collected by the microphone array and the voiceprint characterization data at each moment is extracted to realize the screening of abnormal voiceprints and identify the characteristic voiceprints. In the case of suspected abnormality, the short-time Fourier transform is further used to obtain the voiceprint amplitude information, and a two-dimensional image is generated through normalization processing and color mapping, making the voiceprint features more intuitive and easier for subsequent analysis. At the same time, combined with the infrared image data at the same time, the voiceprint features and infrared features are input into a pre-trained neural network model to fully explore the spatial and temporal correlation of leakage features, calculate the risk coefficient and compare it with the risk threshold to realize risk assessment and early warning. It overcomes the limitations of single voiceprint detection being easily interfered by environmental noise and missed detection, and at the same time improves the sensitivity and stability of micro-leakage detection in complex environments through multi-modal fusion.
[0107] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0108] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A natural gas micro-leakage MFCC detection method based on voiceprint coupling, characterized in that: include: A microphone array and an infrared image acquisition device are arranged in the area to be detected, voiceprint data within a preset time period of the microphone array are collected, and the voiceprint data are processed according to the microphone frequency to determine the voiceprint characterization data at each moment, so as to obtain a voiceprint characterization data set; Compare the voiceprint characterization data set with the abnormal range data, and determine the initial judgment mark according to the comparison result, wherein the initial judgment mark includes a judgment abnormal mark, a suspected abnormal mark, and a judgment non-abnormal mark; When the initial judgment mark is a suspected abnormal mark, the voiceprint characterization data set is subjected to short-time Fourier transform to obtain the amplitude information of each frame, the amplitude information is normalized, and the normalized amplitude information is mapped to the RGB color table to generate a two-dimensional image; the infrared image data of the area to be detected at the same time is extracted, and the two-dimensional image and the infrared image data are input into a pre-trained neural network model to determine the risk coefficient; The risk coefficient is compared with the risk threshold, and a determination is made based on the comparison result whether to issue a risk warning.
2. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 1 is characterized in that: When the voiceprint data is processed according to the microphone frequency to determine the voiceprint representation data at each moment, it includes: The voiceprint characterization data at each moment is determined using Gaussian mixture distribution: The kernel density function expression is: ; Where n represents the number of voiceprint data collected at each moment, h represents the smoothing bandwidth, represents the i-th voiceprint data collected at each moment, and x represents the value of a certain frequency point to be estimated; The voiceprint data with the highest frequency of kernel density is used as the voiceprint characterization data.
3. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 2 is characterized in that: When the voiceprint data is processed according to the microphone frequency to determine the voiceprint representation data at each moment, it also includes: The smoothing bandwidth is calculated by the following formula: ; Among them, h represents the smoothing bandwidth, σ represents the standard deviation of the voiceprint data collected at each moment, β1 represents the data skewness, β2 represents the data kurtosis, n represents the number of voiceprint data collected at each moment, and k represents the adjustment coefficient; The data skewness is calculated by the following formula: ; The data kurtosis is calculated by the following formula: ; Among them, n represents the number of voiceprint data collected at each moment, represents the i-th voiceprint data collected at each moment, represents the mean of the voiceprint data collected at each moment, and σ represents the standard deviation of the voiceprint data collected at each moment.
4. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 3 is characterized in that: The abnormal range data includes: Collecting normal operation records corresponding to normal operation of the pipeline in the area to be detected, and extracting the corresponding normal soundprint amplitude range from the normal operation records; The pipeline standard voiceprint amplitude is compared with the normal voiceprint amplitude range, and the abnormal range data is generated according to the numerical value relationship, wherein the abnormal range data includes a left boundary value and a right boundary value.
5. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 4 is characterized in that: When determining the initial identification based on the comparison results, it includes: When the voiceprint amplitudes of all the voiceprint representation data in the voiceprint representation data set are within the abnormal range data, a non-abnormal determination mark is generated for the area to be detected; When the voiceprint amplitudes of all the voiceprint representation data in the voiceprint representation data set are greater than the right boundary, generating an abnormality determination mark for the area to be detected; When the voiceprint amplitude of the voiceprint representation data in the voiceprint representation data set is within the abnormal range data, and the voiceprint amplitude of the voiceprint representation data is greater than the right boundary, a suspected abnormality mark is generated for the area to be detected.
6. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 5 is characterized in that: When the initial judgment mark is a judgment abnormality mark, it includes: An amplitude excess value is generated based on the maximum voiceprint amplitude of the voiceprint characterization data in the voiceprint characterization data set and the right boundary. The amplitude excess value is the difference between the maximum voiceprint amplitude and the right boundary. The risk coefficient is determined based on the amplitude excess value. The risk coefficient is proportional to the amplitude excess value.
7. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 6 is characterized in that: Inputting the two-dimensional image and infrared image data into a pre-trained neural network model to determine the risk factor includes: Collecting historical abnormal data to construct a historical data set, wherein the historical abnormal data includes historical abnormal two-dimensional images and historical abnormal infrared image data; Sampling the historical data set according to a preset ratio to obtain a training subset and a test subset; Acquire a pre-selected neural network model, and iteratively train the neural network model according to the training subset, evaluate the iteratively trained neural network model according to the test subset, and determine whether to stop the iterative training according to the evaluation value; If the evaluation value of the neural network model after the current iterative training is less than the evaluation value of the neural network model after the previous iterative training, the amplitude of the change of the neural network model in the gradient direction is reduced, and the iterative training is continued until the preset number of iterations is reached; if the evaluation value of the neural network model after the current iterative training is greater than or equal to the evaluation value of the neural network model after the previous iterative training, the iterative training is stopped to obtain the pre-trained neural network model; The two-dimensional image and infrared image data are input into the pre-trained neural network model to obtain the risk coefficient.
8. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 7 is characterized in that: The risk coefficient is compared with the risk threshold, and whether to issue a risk warning is determined according to the comparison result, including: When the risk coefficient is greater than the risk threshold, a risk warning is determined; When the risk coefficient is less than or equal to the risk threshold, it is determined that no risk warning is issued.
9. The MFCC detection method for natural gas micro-leakage based on voiceprint coupling according to claim 8 is characterized in that: When a risk warning is determined, it includes: A risk coefficient difference is obtained according to the risk coefficient and the risk threshold, the risk coefficient difference is the difference between the risk coefficient and the risk threshold, the risk coefficient difference is compared with the first risk coefficient and the second risk coefficient respectively, and the risk level is determined according to the comparison result; the first risk coefficient is less than the second risk coefficient; When the risk coefficient difference is less than or equal to the first risk coefficient, the risk level is determined to be the first risk level; when the risk coefficient difference is greater than the first risk coefficient and less than or equal to the second risk coefficient, the risk level is determined to be the second risk level; when the risk coefficient difference is greater than the second risk coefficient, the risk level is determined to be the third risk level; the first risk level is lower than the second risk level, and the second risk level is lower than the third risk level.
10. A natural gas micro-leakage MFCC detection system based on voiceprint coupling, used for applying the natural gas micro-leakage MFCC detection method based on voiceprint coupling as described in any one of claims 1 to 9, characterized in that: include: A sensor module, including a microphone array and an infrared image acquisition device; A collection unit is configured to collect voiceprint data within a preset time period of the microphone array, process the voiceprint data according to the microphone frequency to determine the voiceprint representation data at each moment, and obtain a voiceprint representation data set; A judgment unit is configured to compare the voiceprint characterization data set with the abnormal range data, and determine a preliminary judgment mark according to the comparison result, wherein the preliminary judgment mark includes a judgment abnormal mark, a suspected abnormal mark, and a judgment non-abnormal mark; The processing unit is configured to, when the initial judgment mark is a suspected abnormal mark, perform short-time Fourier transform on the voiceprint characterization data set to obtain amplitude information of each frame, normalize the amplitude information, and map the normalized amplitude information to an RGB color table to generate a two-dimensional image; extract infrared image data of the area to be detected at the same time, input the two-dimensional image and the infrared image data into a pre-trained neural network model to determine the risk coefficient; The early warning unit is configured to compare the risk coefficient with the risk threshold and determine whether to issue a risk early warning based on the comparison result.
Citation Information
Patent Citations
Rolling bearing sound signal fault diagnosis method based on short-time Fourier transform and sparse laminated automatic encoder
CN104819846A
Power equipment defect diagnosis method based on sound source information and thermal imaging feature fusion
CN112562698A
Voiceprint anomaly detection method based on multi-band self-supervision
CN114842870A
Ultrasonic signal data quality detection method
CN115144474A
Loudspeaker detection method and device, equipment and storage medium
CN115665628A
Cited By
Sealing detection method for ultrahigh vacuum exhaust device
CN120947933A