Sound feature signal recognition method and device, electronic equipment and storage medium
By combining wavelet denoising and empirical mode decomposition with discrete Fourier transform, the problem of low accuracy in identifying the operating sound of power grid equipment was solved, achieving more efficient noise removal and feature extraction, and improving the accuracy of power grid equipment status and fault identification.
Patent Information
- Application Number
- CN202411415383.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-10-11
AI Technical Summary
In existing technologies, the accuracy of power grid equipment operation sound recognition is too low, making it difficult to accurately determine equipment status and faults, mainly due to the influence of ambient noise in the equipment's operating environment.
By combining wavelet denoising, empirical mode decomposition, and discrete Fourier transform, various noises in the sound of equipment operation are removed, thereby improving the accuracy of recognition.
It effectively removes various noises from the equipment's operating sound, improves the accuracy of power grid equipment identification and feature extraction, and achieves more efficient signal recognition.
Smart Images

Figure CN119296593B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound processing technology, and in particular to methods, apparatus, electronic devices, and storage media for sound feature signal recognition. Background Technology
[0002] Power grid equipment such as substation main transformers, reactors, switchgear, overhead transmission lines, cables, distribution transformers, pole-mounted switches, switchgear, and ring main units generate different sounds during operation. These sounds contain important information such as the operating status of the equipment and potential faults. By collecting the operating sounds of the equipment and using sound recognition technology, the operating status of the power equipment can be diagnosed and monitored, and faults can be identified. However, when using sound recognition technology for identification, the collected operating sounds are affected by environmental factors, resulting in a variety of different noises. Excessive noise leads to a low sound recognition rate when identifying the operating sounds of the equipment, making it difficult to determine the operating status of the power grid equipment. Summary of the Invention
[0003] This invention provides a method, apparatus, electronic device, and storage medium for sound feature signal recognition, in order to solve the technical problem of low accuracy in the recognition of equipment operation sounds of power grid equipment in the prior art.
[0004] According to one aspect of the present invention, a method for recognizing sound feature signals is provided, comprising:
[0005] The audio signal to be processed is acquired through the communication unit;
[0006] Perform wavelet denoising on the audio signal to be processed to determine the first audio signal;
[0007] The first sound signal is denoised by empirical mode decomposition to determine the target sound signal;
[0008] The target sound signal is subjected to spectral analysis using discrete Fourier transform to determine the target time-frequency diagram.
[0009] The target time-frequency map is uploaded to the communication unit, and the communication unit performs signal recognition on the target time-frequency map to determine the sound recognition result.
[0010] According to another aspect of the present invention, a sound feature signal recognition device is provided, comprising:
[0011] The data communication module is used to acquire the audio signal to be processed through the communication unit;
[0012] The first noise reduction module is used to perform wavelet noise reduction on the audio signal to be processed and determine the first audio signal.
[0013] The second noise reduction module is used to denoise the first sound signal through empirical mode decomposition to determine the target sound signal;
[0014] The spectrum calculation module is used to perform spectrum calculation on the target sound signal through discrete Fourier transform to determine the target time-frequency diagram;
[0015] The data upload module is used to upload the target time-frequency map to the communication unit, and the communication unit performs signal recognition on the target time-frequency map to determine the sound recognition result.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the sound feature signal recognition method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the sound feature signal recognition method according to any embodiment of the present invention.
[0021] The technical solution of this invention involves acquiring a sound signal to be processed through a communication unit, establishing communication with the communication unit, acquiring the sound information to be processed through the communication unit, performing wavelet denoising on the sound signal to determine a first sound signal, and using wavelet denoising to preserve the key features of the sound signal, thereby achieving efficient denoising and improving the denoising effect; denoising the first sound signal through empirical mode decomposition to determine a target sound signal, and further decomposing and denoising the sound signal through empirical mode decomposition, enabling multi-scale and multi-resolution recognition and removal of noise in the sound, achieving a more refined denoising effect; calculating the spectrum of the target sound signal through discrete Fourier transform to determine the target time-frequency map, and displaying the sound in a visual form as an image, which can improve the accuracy of feature extraction and effectively improve the recognition accuracy of the sound signal; uploading the target time-frequency map to the communication unit, and using the communication unit to perform signal recognition on the target time-frequency map to determine the sound recognition result. This invention solves the technical problem of low accuracy in recognizing the operating sounds of power grid equipment in existing technologies. By performing two noise reduction steps, it can effectively remove various different noises present in the operating sounds of equipment, thereby improving the accuracy of power grid equipment recognition.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart of a sound feature signal recognition method is provided in this embodiment of the invention;
[0025] Figure 2 A flowchart of another sound feature signal recognition method provided in an embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of the structure of a sound feature signal recognition device provided in an embodiment of the present invention;
[0027] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Figure 1 This invention provides a flowchart of a method for recognizing sound feature signals. This embodiment is applicable to situations where sound signals from power grid equipment are identified and judged using an integrated sound sensing terminal. The method can be executed by a sound feature signal recognition device, which can be implemented in hardware and / or software and can be configured within the integrated sound sensing terminal. Figure 1 As shown, the method includes:
[0031] S110. Acquire the audio signal to be processed through the communication unit.
[0032] Optionally, the integrated voice sensing terminal includes an intelligent voiceprint perception function, which consists of an audio control processing unit and a communication interface. The audio control processing unit communicates with the communication unit through the communication interface. The communication unit has a three-layer architecture: a terminal layer, an edge layer, and a cloud layer. The terminal layer generates different types of data and transmits the data to the cloud layer for processing via wireless communication. The edge layer consists of edge nodes with limited computing power, including small base stations connected to edge servers. The edge servers are connected to the cloud layer via optical links. The cloud layer includes a data center layer and a service layer. The integrated voice sensing terminal connects to a vector sample library through the communication unit. The vector sample library is used to store the sound signals to be processed. It should be noted that the vector sample library is constructed through the following steps:
[0033] Collect various types of data from power grid equipment, including audio signals from the equipment.
[0034] The collected data undergoes data preprocessing, which includes data cleaning to remove duplicate, incomplete, or erroneous data; data correction and registration to eliminate biases, errors, or inconsistencies in the data; and data format conversion to unify the data into a suitable vector format.
[0035] The preprocessed data is then vectorized, converting it into vector data. Vectorization methods can include automated vectorization algorithms and point cloud processing; vector data can be points, lines, or surfaces.
[0036] To perform data modeling for vector data, select an appropriate data model, design a database, establish a data schema, and define data tables, fields, and constraints. For example, the data model for data modeling can be a geospatial data model; the database can be a Milvus vector database.
[0037] The vector database includes databases for raw sample storage, model information storage, label sample storage, raw data detection record storage, noise anomaly detection record storage, abnormal discharge detection record storage, and mechanical fault detection record storage. The raw sample storage stores the sound signals to be processed.
[0038] Specifically, the system connects to a vector database via a communication unit to obtain the sound signal to be processed from the original sample library.
[0039] S120. Perform wavelet denoising on the audio signal to be processed to determine the first audio signal.
[0040] The first sound signal can be obtained by wavelet transforming the sound signal to be processed. It should be noted that wavelet denoising involves performing a wavelet transform on the sound signal to be processed using wavelet basis functions, performing threshold denoising on the wavelet-transformed signal by selecting a threshold function, and then reconstructing the denoised signal using wavelets to obtain the first sound signal.
[0041] Specifically, wavelet denoising is performed on the audio signal to be processed to determine the first audio signal.
[0042] Optionally, in another optional embodiment of the present invention, the step of performing wavelet denoising on the sound signal to be processed to determine the first sound signal includes...
[0043] The discrete wavelet signal is obtained by convolving and downsampling the audio signal to be processed using a discrete wavelet filter; the discrete wavelet signal is then denoised using a preset threshold function; and finally, the discrete wavelet signal is upsampled and reconstructed using the wavelet coefficients corresponding to the discrete wavelet filter to obtain the first audio signal.
[0044] Discrete wavelet filters, in particular, can be filter banks used in wavelet denoising to decompose and reconstruct the audio signal being processed. A discrete wavelet filter consists of a low-pass filter and a high-pass filter. The low-pass filter extracts the low-frequency components of the signal, while the high-pass filter extracts the high-frequency components. Discrete wavelet filters can decompose the audio signal multiple times to obtain a finer-grained decomposed signal. The discrete wavelet signal can be obtained by performing multiple decompositions using discrete wavelet filters.
[0045] The preset threshold function can be a pre-set threshold function used to denoise the decomposed signal. For example, the preset threshold function can be a soft threshold or a hard threshold. The denoised wavelet signal can be obtained by processing a discrete wavelet signal using the preset threshold function to obtain the decomposed signal. It should be noted that the denoised wavelet signal is obtained by directly performing a function operation between the discrete wavelet signal and the preset threshold function. For example, when the preset threshold function is set to a hard threshold function, the decomposed signals in the discrete wavelet signal that are less than the threshold are set to 0, while the decomposed signals that are greater than the threshold are retained.
[0046] Optionally, the wavelet denoising process is as follows: The sound signal to be processed is filtered and calculated by the low-pass and high-pass filters in the discrete wavelet filter, and downsampled to obtain the low-frequency and high-frequency parts of the first decomposition. The low-frequency part of the first decomposition is then filtered and calculated by the low-pass and high-pass filters again, and downsampled. Through multiple decomposition cycles, a discrete wavelet signal is obtained. The discrete wavelet signal is then thresholded and denoised by a preset threshold function to obtain a denoised wavelet signal. Then, the inverse discrete wavelet filter is used as wavelet coefficients, and the denoised wavelet signal is reconstructed and upsampled multiple times to obtain the first sound signal.
[0047] S130. The first sound signal is denoised by empirical mode decomposition to determine the target sound signal.
[0048] Empirical Mode Decomposition (EMD) is an adaptive time-frequency analysis method that can decompose complex signals into several mode functions and effectively separate different frequency components of the signal.
[0049] The target sound signal can be the sound signal obtained by empirical mode decomposition and denoising of the first sound signal. It should be noted that after empirical mode decomposition, the first sound signal yields multiple mode functions and residuals. By denoising the mode functions, the denoised mode functions and residuals constitute the target sound signal.
[0050] Specifically, the first sound signal is denoised using empirical mode decomposition to determine the target sound signal.
[0051] S140. Calculate the spectrum of the target sound signal using discrete Fourier transform to determine the target time-frequency diagram.
[0052] The Discrete Fourier Transform (DFT) can perform a discrete Fourier transform on a signal, converting the target audio signal into a time-frequency graph. The target time-frequency graph is obtained by converting the target audio signal into a time-frequency graph. The target video graph is a two-dimensional graph, with the horizontal axis representing time and the vertical axis representing frequency. Gray levels represent the energy of a frequency at a given time point. The target time-frequency graph can represent the timbre characteristics of the target audio signal.
[0053] Specifically, the target sound signal is subjected to spectral analysis using Discrete Fourier Transform to determine the target time-frequency diagram.
[0054] Optionally, in another optional embodiment of the present invention, the step of calculating the spectrum of the target sound signal through discrete Fourier transform to determine the target time-frequency diagram includes:
[0055] Perform a discrete Fourier transform on the target sound signal to determine the sound spectrum signal; draw a spectrum diagram based on the sound spectrum signal to obtain the target time-frequency diagram.
[0056] Among them, the sound spectrum signal can be the frequency domain representation of the target sound signal after Fourier transform.
[0057] Optionally, the specific process of performing Fourier transform on the target sound signal is as follows: The continuous target sound signal is divided into frames, which are then divided into multiple short frames. Each short frame is assigned r sampling points. The framed signal can be represented as e(s,r), where s represents the s-th frame and r represents the sampling points within the frame. A Discrete Fourier Transform (DFT) is performed on each frame e(s,r) to obtain the sound spectrum signal e(s,k), where k represents the frequency index. The periodogram U(s,k) is calculated based on the sound spectrum signal e(s,k), and the periodogram U(s,k) is converted to a logarithmic scale Z(s,k). The spectrum is then plotted based on the logarithmic scale Z(s,k) to obtain the target time-frequency map.
[0058] S150. Upload the target time-frequency map to the communication unit, and use the communication unit to perform signal recognition on the target time-frequency map to determine the sound recognition result.
[0059] The sound recognition result can be the status information of the power grid equipment corresponding to the sound signal to be processed. For example, when the power grid equipment is a transformer, the sound to be processed can be high-frequency noise, and the sound recognition result can be partial discharge of the transformer.
[0060] Specifically, the target time-frequency map is sent to the communication unit via a communication interface. The edge server in the middle edge layer of the communication unit performs signal recognition, and a neural network model is used to perform signal recognition on the target time-frequency map to determine the sound recognition result. The neural network model can be a convolutional neural network with a self-attention layer; the convolutional neural network includes an input layer, a convolutional layer, an attention layer, a multi-scale layer, a residual layer, and an output layer.
[0061] The technical solution of this invention involves acquiring a sound signal to be processed through a communication unit, establishing communication with the communication unit, acquiring the sound information to be processed through the communication unit, performing wavelet denoising on the sound signal to determine a first sound signal, and using wavelet denoising to preserve the key features of the sound signal, thereby achieving efficient denoising and improving the denoising effect; denoising the first sound signal through empirical mode decomposition to determine a target sound signal, and further decomposing and denoising the sound signal through empirical mode decomposition, enabling multi-scale and multi-resolution recognition and removal of noise in the sound, achieving a more refined denoising effect; calculating the spectrum of the target sound signal through discrete Fourier transform to determine the target time-frequency map, and displaying the sound in a visual form as an image, which can improve the accuracy of feature extraction and effectively improve the recognition accuracy of the sound signal; uploading the target time-frequency map to the communication unit, and using the communication unit to perform signal recognition on the target time-frequency map to determine the sound recognition result. This invention solves the technical problem of low accuracy in recognizing the operating sounds of power grid equipment in existing technologies. By performing two noise reduction steps, it can effectively remove various different noises present in the operating sounds of equipment, thereby improving the accuracy of power grid equipment recognition.
[0062] Figure 2 This is a flowchart of another sound feature signal recognition method provided by an embodiment of the present invention. The relationship between this embodiment and the above embodiments is that this embodiment is a specific process of further denoising through empirical mode decomposition. Figure 2 As shown, the method includes:
[0063] S210. Acquire the audio signal to be processed through the communication unit.
[0064] S220. Perform wavelet denoising on the audio signal to be processed to determine the first audio signal.
[0065] S230. Perform empirical mode decomposition on the first sound signal to determine the first empirical mode function.
[0066] The first empirical mode function can be a function composed of at least one mode function and a residual. It should be noted that after performing empirical mode decomposition on the first audio signal, at least one mode function and a residual are obtained; the function composed of the mode function and the residual is defined as the first empirical mode function.
[0067] Specifically, empirical mode decomposition is performed on the first sound signal to determine the first empirical mode function.
[0068] Optionally, in another optional embodiment of the present invention, according to claim 3, the method is characterized in that, performing empirical mode decomposition on the first sound signal to determine the first empirical mode function includes:
[0069] The first sound signal is subjected to multiple rounds of noise addition using preset random Gaussian noise to obtain a set of noise-added sound signals; wherein, the set of noise-added sound signals is the noise-added sound signal obtained by adding noise to the first sound signal each time.
[0070] Empirical mode decomposition is performed on each noisy audio signal to determine the multi-mode function and residual corresponding to each noisy audio signal; wherein, the multi-mode function is composed of the mode function components of each signal.
[0071] The modal function components corresponding to each order of all noisy audio signals are averaged to determine the average modal function component corresponding to each order.
[0072] The residual corresponding to each noise-added audio signal is averaged to determine the average residual corresponding to the first audio signal.
[0073] The first empirical mode function is constructed based on multiple average mode function components and average residuals.
[0074] Among them, random Gaussian noise can be randomly generated Gaussian noise used for denoising; the amplitude distribution of random Gaussian noise conforms to a normal distribution.
[0075] The noise-added audio signal set can be a collection of noise-added audio signals generated by adding random Gaussian noise to the first audio signal. It should be noted that when adding noise to the first audio signal with random Gaussian noise, the first audio signal will be noise-added multiple times. The random Gaussian noise will change randomly each time it is noise-added, and the audio signal obtained by each noise addition will be saved as a noise-added audio signal, thus obtaining multiple noise-added audio signals. Then, a noise-added audio signal set can be constructed based on the multiple noise-added audio signals.
[0076] Among them, the multi-order mode function can be obtained by empirical mode decomposition. Since empirical mode decomposition is performed sequentially, the mode function obtained by the first decomposition is called the first-order mode function component, the mode function obtained by the second decomposition is called the second-order mode function component, the mode function obtained by the j-th decomposition is called the j-th order mode function component, and all the mode function components obtained from the first to the nth decomposition are called multi-order mode functions.
[0077] Optionally, during empirical mode decomposition, since multiple noise additions are made to the first sound signal, multiple noisy sound signals are formed. After empirical mode decomposition of the multiple noisy sound signals, the multi-order mode functions and residuals corresponding to each noisy sound signal are obtained. The mode function components of each noisy sound signal are averaged with the mode function components corresponding to other noisy sound signals, and the average mode function component corresponding to each order mode function component is calculated. For example, suppose there are 3 noisy audio signals. Through empirical mode decomposition, we obtain 3 multi-mode functions and residuals corresponding to the 3 noisy audio signals. Each noisy audio signal is decomposed into 4 mode function categories. Then the multi-mode functions are 4th-order mode functions. The way to average the 4th-order mode functions of the 3 noisy audio signals is as follows: average the 1st-order mode function components of the 3 noisy audio signals to obtain the 1st-order average mode function component. Repeat the calculation to obtain the 2nd-order average mode function component, the 3rd-order average mode function component, and the 4th-order average mode function component.
[0078] The average residual can be the average of the residuals corresponding to all noisy audio signals.
[0079] Specifically, the first sound signal is subjected to random Gaussian noise multiple times, and each noise-added signal is used as the noise-added sound signal corresponding to the first sound signal, thus forming a set of noise-added sound signals. Empirical mode decomposition is performed on each noise-added sound signal to obtain the multi-order mode function and residual corresponding to each noise-added sound signal. The mode function components and residuals corresponding to each order are averaged to determine the average mode function component and average residual corresponding to each order mode function component. The first empirical mode function is constructed based on the average mode function component and average residual of each order.
[0080] For example, the first sound signal is transmitted through x(t), and the noise signal is transmitted through x. i Let (t) represent the i-th noise addition, and w be the denoting factor. i (t) represents the random Gaussian white noise added in the i-th iteration. After decomposing the noisy signal using EMD, the j-th order mode function component and residual are obtained. The j-th order mode function component is represented by c. ij Let (t) be the expression, where j is 1, 2, 3, ..., m, and the residual is represented by r. i (t) is used to represent the signal; the formula for adding noise to the first sound signal using random Gaussian white noise is shown below:
[0081] x i (t)=x(t)+w i (t), i = 1, 2, 3, ... N;
[0082] The formula for decomposing a noisy audio signal using EMD is shown below:
[0083]
[0084] Furthermore, the j-th order average mode function component is represented by c. j The mean residual is represented by r(t), and the specific calculation method is shown below:
[0085]
[0086]
[0087] S240. Perform frequency denoising and energy denoising on the first empirical mode function to determine the target sound signal.
[0088] Specifically, the first empirical mode function is first denoised by frequency, and then denoised by capability, to obtain the target sound signal.
[0089] Optionally, in another optional embodiment of the present invention, the step of performing frequency denoising and energy denoising on the first empirical mode function to determine the target sound signal includes: identifying noise in the average mode function component of the first empirical mode function according to a preset frequency threshold to determine the first noise mode function component; removing the first noise mode function component from the first empirical mode function to obtain the second empirical mode function; and performing energy denoising on the second empirical mode function to determine the target sound signal.
[0090] The preset frequency threshold can be a pre-set frequency threshold used for denoising the average modal function components. It should be noted that since the highest frequency of the mechanical response is a fixed value, this highest frequency is chosen as the preset frequency threshold. Average modal function components with frequencies higher than the preset frequency threshold are considered noise and are designated as the first noise modal function components. That is, the first noise modal function components can be any number of average modal function components, and all average modal function components with frequencies higher than the preset frequency threshold are designated as the first noise modal function components.
[0091] The second empirical mode function is the first empirical mode function after removing the components of the first noise mode function.
[0092] Specifically, the average mode function component of the first empirical mode function is compared with a preset frequency threshold. If the frequency of the average mode function component is greater than the preset frequency threshold, it is identified as the first noise mode function component. The first noise mode function component is subtracted from the first empirical mode function to obtain the second empirical mode function. Energy denoising is performed on the second empirical mode function to determine the target sound signal.
[0093] Optionally, in another optional embodiment of the present invention, the step of performing energy denoising on the second empirical mode function to determine the target sound signal includes:
[0094] The energy density is calculated based on the amplitude of each mode function component in the second empirical mode function, and the signal energy density corresponding to each mode function component is obtained.
[0095] The energy corresponding to each mode function component is determined based on the signal energy density corresponding to each mode function component;
[0096] The energy ratio of each modal function component is calculated sequentially to determine the signal energy ratio of each modal function component.
[0097] Based on the signal energy ratio of each mode function classification, noise identification is performed on the energy corresponding to each mode function component to determine the second noise mode function component;
[0098] The target sound signal is obtained by removing the second noise mode function component from the second empirical mode function.
[0099] Signal energy density describes the energy distribution of each mode function component within a unit frequency range. It should be noted that the signal energy density of each mode function component can be calculated based on its amplitude.
[0100] Optionally, after obtaining the signal energy density of each modal function component, since the modal function components are all continuous-time signals, time-domain calculations or frequency-domain calculations can be performed based on the signal energy density of each modal function component to obtain the energy of each modal function component.
[0101] Optionally, after obtaining the energy of each modal function component, for each modal function component, the average energy of all preceding modal function components is calculated to determine the average energy contrast value of each modal function component. The energy of each modal function component and its average energy contrast value are then used to calculate the energy ratio, yielding the signal-to-energy ratio of each modal function component. For example, the formula for calculating the signal-to-energy ratio of each modal function component is as follows: The energy of the j-th order modal function component is represented by p... j The average energy contrast value corresponding to the j-th order modal function component is the average energy of the first j-1 order modal function components, denoted as q. j-1 Signal energy ratio using R pj The specific calculation formula is as follows:
[0102]
[0103] Optionally, after obtaining the signal-to-energy ratio of each modal function component, noise identification is performed on the modal function components based on their signal-to-energy ratios. The specific identification process is as follows: if the signal-to-energy ratio of a modal function component is not less than 1, the modal function component is considered a normal sound signal; if the signal-to-energy ratio of a modal function component is less than 1, the modal function component is identified as a second noise modal function component. The second noise modal function component is a modal function component whose signal energy is noise.
[0104] Specifically, the energy density is calculated based on the amplitude of each modal function component in the second empirical modal function to obtain the signal energy density corresponding to each modal function component; the energy corresponding to each modal function component is determined based on the signal energy density corresponding to each modal function component; the energy ratio is calculated for the energy corresponding to each modal function component to determine the signal energy ratio corresponding to each modal function component; noise is identified for the energy corresponding to each modal function component based on the signal energy ratio classified by each modal function to determine the second noise modal function component; the second noise modal function component is removed from the second empirical modal function to obtain the target sound signal.
[0105] S250. Calculate the spectrum of the target sound signal using discrete Fourier transform to determine the target time-frequency diagram;
[0106] S260. Upload the target time-frequency map to the communication unit, and use the communication unit to perform signal recognition on the target time-frequency map to determine the sound recognition result.
[0107] The technical solution of this invention denoises the first sound signal using Empirical Mode Decomposition (EMD) to determine the target sound signal. It then decomposes the first sound signal multiple times using EMD and further denoises it using frequency and energy denoising. This multi-scale, multi-resolution noise removal achieves a more refined denoising effect. It solves the problem of low accuracy in recognizing the operating sounds of power grid equipment in existing technologies. By performing two denoising steps, it effectively removes various types of noise present in the operating sounds of equipment, improving the accuracy of power grid equipment identification.
[0108] Figure 3 This is a schematic diagram of a sound feature signal recognition device provided in an embodiment of the present invention. Figure 3 As shown, the device includes: a data communication module 310, a first denoising module 320, a second denoising module 330, a spectrum calculation module 340, and a data upload module 350; wherein,
[0109] Data communication module 310 is used to acquire the audio signal to be processed through the communication unit;
[0110] The first noise reduction module 320 is used to perform wavelet noise reduction on the audio signal to be processed and determine the first audio signal.
[0111] The second noise reduction module 330 is used to denoise the first sound signal through empirical mode decomposition to determine the target sound signal;
[0112] The spectrum calculation module 340 is used to perform spectrum calculation on the target sound signal through discrete Fourier transform to determine the target time-frequency diagram;
[0113] The data upload module 350 is used to upload the target time-frequency map to the communication unit, and the communication unit performs signal recognition on the target time-frequency map to determine the sound recognition result.
[0114] The technical solution of this invention involves acquiring a sound signal to be processed through a communication unit, establishing communication with the communication unit, acquiring the sound information to be processed through the communication unit, performing wavelet denoising on the sound signal to determine a first sound signal, and using wavelet denoising to preserve the key features of the sound signal, thereby achieving efficient denoising and improving the denoising effect; denoising the first sound signal through empirical mode decomposition to determine a target sound signal, and further decomposing and denoising the sound signal through empirical mode decomposition, enabling multi-scale and multi-resolution recognition and removal of noise in the sound, achieving a more refined denoising effect; calculating the spectrum of the target sound signal through discrete Fourier transform to determine the target time-frequency map, and displaying the sound in a visual form as an image, which can improve the accuracy of feature extraction and effectively improve the recognition accuracy of the sound signal; uploading the target time-frequency map to the communication unit, and using the communication unit to perform signal recognition on the target time-frequency map to determine the sound recognition result. This invention solves the technical problem of low accuracy in recognizing the operating sounds of power grid equipment in existing technologies. By performing two noise reduction steps, it can effectively remove various different noises present in the operating sounds of equipment, thereby improving the accuracy of power grid equipment recognition.
[0115] Optionally, the first noise reduction module is specifically used for:
[0116] Discrete wavelet signals are obtained by convolving and downsampling the audio signal to be processed using a discrete wavelet filter.
[0117] The discrete wavelet signal is denoised by thresholding the discrete wavelet signal using a preset threshold function.
[0118] The discrete wavelet signal is upsampled and reconstructed by convolution using the wavelet coefficients corresponding to the discrete wavelet filter to obtain the first sound signal.
[0119] Optionally, the second noise reduction module is specifically used for:
[0120] Empirical mode decomposition is performed on the first sound signal to determine the first empirical mode function;
[0121] Frequency and energy denoising are performed on the first empirical mode function to determine the target sound signal.
[0122] Optionally, the second noise reduction module is further used for:
[0123] The first sound signal is subjected to multiple rounds of noise addition using preset random Gaussian noise to obtain a set of noise-added sound signals; wherein, the set of noise-added sound signals is the noise-added sound signal obtained by adding noise to the first sound signal each time.
[0124] Empirical mode decomposition is performed on each noisy audio signal to determine the multi-mode function and residual corresponding to each noisy audio signal; wherein, the multi-mode function is composed of the mode function components of each signal.
[0125] The modal function components corresponding to each order of all noisy audio signals are averaged to determine the average modal function component corresponding to each order.
[0126] The residual corresponding to each noise-added audio signal is averaged to determine the average residual corresponding to the first audio signal.
[0127] The first empirical mode function is constructed based on multiple average mode function components and average residuals.
[0128] Optionally, the second noise reduction module is further used for:
[0129] Based on a preset frequency threshold, noise is identified in the average mode function component of the first empirical mode function to determine the first noise mode function component;
[0130] The first noise mode function component is removed from the first empirical mode function to obtain the second empirical mode function;
[0131] Energy denoising is performed on the second empirical mode function to determine the target sound signal.
[0132] Optionally, the second noise reduction module is further used for:
[0133] The energy density is calculated based on the amplitude of each mode function component in the second empirical mode function, and the signal energy density corresponding to each mode function component is obtained.
[0134] The energy corresponding to each mode function component is determined based on the signal energy density corresponding to each mode function component;
[0135] The energy ratio of each modal function component is calculated sequentially to determine the signal energy ratio of each modal function component.
[0136] Based on the signal energy ratio of each mode function classification, noise identification is performed on the energy corresponding to each mode function component to determine the second noise mode function component;
[0137] The target sound signal is obtained by removing the second noise mode function component from the second empirical mode function.
[0138] Optionally, the spectrum calculation module is specifically used for:
[0139] Perform a discrete Fourier transform on the target sound signal to determine the sound spectrum signal;
[0140] The target time-frequency diagram is obtained by plotting the spectrum of the sound signal.
[0141] The sound feature signal recognition device provided in the embodiments of the present invention can execute the sound feature signal recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0142] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their patterns are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0143] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0144] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer grids such as the Internet and / or various telecommunications grids.
[0145] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as voice feature signal recognition methods.
[0146] In some embodiments, the voice feature signal recognition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the voice feature signal recognition method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the voice feature signal recognition method by any other suitable means (e.g., by means of firmware).
[0147] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0148] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the patterns / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0149] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0150] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0151] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or grid browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication grid). Examples of communication grids include local area networks (LANs), wide area networks (WANs), blockchain grids, and the Internet.
[0152] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact through a communication mesh. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0153] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0154] This embodiment provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the program implements the steps of the sound feature signal recognition method provided in any embodiment of the present invention. The method includes:
[0155] The audio signal to be processed is acquired through the communication unit;
[0156] Perform wavelet denoising on the audio signal to be processed to determine the first audio signal;
[0157] The first sound signal is denoised by empirical mode decomposition to determine the target sound signal;
[0158] The target sound signal is subjected to spectral analysis using discrete Fourier transform to determine the target time-frequency diagram.
[0159] The target time-frequency map is uploaded to the communication unit, and the communication unit performs signal recognition on the target time-frequency map to determine the sound recognition result.
[0160] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0161] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0162] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0163] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of mesh, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0164] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a grid of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0165] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0166] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for recognizing sound feature signals, characterized in that, include: The audio signal to be processed is acquired through the communication unit; The sound signal to be processed is subjected to wavelet denoising to determine the first sound signal; The first sound signal is denoised by empirical mode decomposition to determine the target sound signal; The target sound signal is subjected to spectral analysis using discrete Fourier transform to determine the target time-frequency diagram; The target time-frequency map is uploaded to the communication unit, and the communication unit performs signal recognition on the target time-frequency map to determine the sound recognition result; The step of denoising the first sound signal through empirical mode decomposition to determine the target sound signal includes: Perform empirical mode decomposition on the first sound signal to determine the first empirical mode function; The first empirical mode function is subjected to frequency denoising and energy denoising to determine the target sound signal; The step of performing empirical mode decomposition on the first sound signal to determine the first empirical mode function includes: The first sound signal is subjected to multiple additions of preset random Gaussian noise to obtain a set of noise-added sound signals; wherein, the set of noise-added sound signals is the noise-added sound signal obtained by adding noise to the first sound signal each time; the random Gaussian noise is randomly generated Gaussian noise used for denoising, and the amplitude distribution of the random Gaussian noise conforms to a normal distribution; Empirical mode decomposition is performed on each of the noise-added audio signals to determine the multi-order mode function and residual corresponding to each noise-added audio signal; wherein, the multi-order mode function is composed of the mode function components of each signal; The modal function components corresponding to each order of all the noise-added audio signals are averaged to determine the average modal function component corresponding to each order. The residual corresponding to each of the noise-added audio signals is averaged to determine the average residual corresponding to the first audio signal. The first empirical mode function is constructed based on multiple average mode function components and the average residual; The step of performing frequency denoising and energy denoising on the first empirical mode function to determine the target sound signal includes: Based on a preset frequency threshold, noise is identified in the average mode function component of the first empirical mode function to determine the first noise mode function component; The first noise mode function component is removed from the first empirical mode function to obtain the second empirical mode function; The second empirical mode function is subjected to energy denoising to determine the target sound signal; The step of performing energy denoising on the second empirical mode function to determine the target sound signal includes: The energy density is calculated based on the amplitude of each modal function component in the second empirical modal function to obtain the signal energy density corresponding to each modal function component; The energy corresponding to each modal function component is determined based on the signal energy density corresponding to each modal function component; The energy ratio is calculated sequentially for the energy corresponding to each modal function component to determine the signal energy ratio corresponding to each modal function component; the formula for calculating the signal energy ratio is: ; Wherein, the energy of the j-th order modal function component is represented by p j The average energy contrast value corresponding to the j-th order modal function component is the average energy of the first j-1 order modal function components, denoted as q. j-1 Signal energy ratio using R pj To represent; Based on the signal energy ratio of each mode function classification, noise identification is performed on the energy corresponding to each mode function component to determine the second noise mode function component; The target sound signal is obtained by removing the second noise mode function component from the second empirical mode function.
2. The method according to claim 1, characterized in that, The step of performing wavelet denoising on the sound signal to be processed to determine the first sound signal includes: The discrete wavelet signal is obtained by convolving and downsampling the sound signal to be processed using a discrete wavelet filter. The discrete wavelet signal is denoised by performing threshold denoising on the discrete wavelet signal using a preset threshold function to determine the denoised wavelet signal. The discrete wavelet signal is upsampled and reconstructed by convolution using the wavelet coefficients corresponding to the discrete wavelet filter to obtain the first sound signal.
3. The method according to claim 1, characterized in that, The step of calculating the spectrum of the target sound signal using discrete Fourier transform to determine the target time-frequency spectrum includes: Perform a Fourier transform on the target sound signal to determine the sound spectrum signal; The target time-frequency diagram is obtained by plotting a spectrum based on the sound spectrum signal.
4. A sound feature signal recognition device, characterized in that, include: The data communication module is used to acquire the audio signal to be processed through the communication unit; The first denoising module is used to perform wavelet denoising on the sound signal to be processed to determine the first sound signal; The second noise reduction module is used to denoise the first sound signal through empirical mode decomposition to determine the target sound signal; The spectrum calculation module is used to perform spectrum calculation on the target sound signal through discrete Fourier transform to determine the target time-frequency diagram; The data upload module is used to upload the target time-frequency map to the communication unit, and the communication unit performs signal recognition on the target time-frequency map to determine the sound recognition result. The second noise reduction module is specifically used for: Empirical mode decomposition is performed on the first sound signal to determine the first empirical mode function; The first empirical mode function is denoised by frequency and energy to determine the target sound signal. The second noise reduction module is also specifically used for: The first sound signal is denoised multiple times by adding preset random Gaussian noise to obtain a set of denoised sound signals; wherein, the set of denoised sound signals is the denoised sound signal obtained by adding noise to the first sound signal each time; the random Gaussian noise is randomly generated Gaussian noise used for denoising, and the amplitude distribution of the random Gaussian noise conforms to a normal distribution; Empirical mode decomposition is performed on each noisy audio signal to determine the multi-mode function and residual corresponding to each noisy audio signal; wherein, the multi-mode function is composed of the mode function components of each signal. The modal function components corresponding to each order of all noisy audio signals are averaged to determine the average modal function component corresponding to each order. The residual corresponding to each noise-added audio signal is averaged to determine the average residual corresponding to the first audio signal. The first empirical mode function is constructed based on multiple average mode function components and average residuals; The second noise reduction module is also specifically used for: Based on a preset frequency threshold, noise is identified in the average mode function component of the first empirical mode function to determine the first noise mode function component; The first noise mode function component is removed from the first empirical mode function to obtain the second empirical mode function; Energy denoising is performed on the second empirical mode function to determine the target sound signal; The second noise reduction module is also specifically used for: The energy density is calculated based on the amplitude of each mode function component in the second empirical mode function, and the signal energy density corresponding to each mode function component is obtained. The energy corresponding to each mode function component is determined based on the signal energy density corresponding to each mode function component; The energy ratio is calculated sequentially for the energy corresponding to each mode function component to determine the signal energy ratio for each mode function component; the formula for calculating the signal energy ratio is: ; Wherein, the energy of the j-th order modal function component is represented by p. j The average energy contrast value corresponding to the j-th order modal function component is the average energy of the first j-1 order modal function components, denoted as q. j-1 Signal energy ratio using R pj To represent; Based on the signal energy ratio of each mode function classification, noise identification is performed on the energy corresponding to each mode function component to determine the second noise mode function component; The target sound signal is obtained by removing the second noise mode function component from the second empirical mode function.
5. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the sound feature signal recognition method according to any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the sound feature signal recognition method according to any one of claims 1-3.
Citation Information
Patent Citations
Waveform self-adaptive and self-balanced damage diagnosis method
CN115393640A
Noise reduction method and system for acoustic emission signal of robot
CN117316172A
Belt conveyor carrier roller fault diagnosis method and system based on multi-scale feature fusion and residual mask convolution attention algorithm
CN117421581A