Sound feature extraction method and device, electronic equipment and storage medium

By performing noise reduction, interpolation filtering, pre-emphasis and frame-based windowing processing on the original sound signal, the Mel spectral coefficients are extracted, which solves the problem of low sound feature quality in noise environments and improves the training effect of the speech recognition model.

CN120108416AInactive Publication Date: 2025-06-06SICHUAN HUIDEXUAN CULTURE & ART CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510578840.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing Mel spectral coefficient extraction method reduces the quality of sound characteristics in a noisy environment, affecting the training results of the speech recognition model.

Method used

By denoising the original sound signal based on the ambient noise signal, then performing double interpolation filtering, pre-improvement processing and frame-by-frame windowing processing, and finally extracting the Mel spectral coefficients for each signal frame.

Benefits of technology

The quality and extraction efficiency of sound features are improved, noise interference is reduced, and the training results of the speech recognition model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108416A_ABST
    Figure CN120108416A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, and provides a voice feature extraction method and device, electronic equipment and a storage medium. The method comprises the following steps: firstly, performing noise reduction processing based on an environment noise signal and an original sound signal collected in the same environment to obtain a noise-reduced sound signal; then double interpolation filtering processing is carried out on the noise reduction sound signal to obtain a filtering sound signal, the intensity of a signal in a frequency range in the filtering sound signal is higher than that of the noise reduction sound signal, and the intensity of a signal exceeding the frequency range in the filtering sound signal is lower than that of a signal exceeding the frequency range in the noise reduction sound signal; performing pre-emphasis processing and framing windowing processing on the filtered sound signal to obtain a plurality of signal frames; and finally, performing feature extraction on each signal frame to obtain the Mel-spectrum coefficient of each signal frame, namely, obtaining the sound feature of the original sound signal. By reducing the noise in the sound signal and shortening the signal length required for feature extraction, the quality of the sound feature and the extraction efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a sound feature extraction method, device, electronic device and storage medium. Background Art

[0002] In the process of deep network learning of sound, the quality of sound features is very important for the training results of the model. For example, the quality of sound features is very important for the recognition accuracy of training speech recognition models. At present, MFSC (Mel-Frequency Spectral Coefficients) that conform to the auditory characteristics of the human ear are generally used as sound features. However, the existing method of extracting Mel-Frequency Spectral Coefficients will reduce the quality of sound features and affect the training results of the model when the noise of the sound signal is large, that is, the signal-to-noise ratio is low. For example, when collecting human voice signals outdoors in a noisy environment, it is usually necessary to collect sound signals for a long time for feature extraction, which will not only reduce the efficiency of sound extraction, but also affect the quality of sound features due to the presence of noise interference, resulting in poor training results of the sound model. Summary of the invention

[0003] In view of this, an object of the present invention is to provide a sound feature extraction method, device, electronic device and storage medium.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a method for extracting sound features, the method comprising: Performing noise reduction processing on the original sound signal based on the environmental noise signal to obtain a noise-reduced sound signal; the environmental noise signal and the original sound signal are collected in the same environment according to a preset sampling frequency; Performing a two-fold interpolation filtering process on the noise reduction sound signal to obtain a filtered sound signal; wherein the signal strength within a preset frequency range in the filtered sound signal is higher than the signal strength within the frequency range in the noise reduction sound signal; and the signal strength exceeding the frequency range in the filtered sound signal is lower than the signal strength exceeding the frequency range in the noise reduction sound signal; Performing pre-emphasis processing and frame windowing processing on the filtered sound signal to obtain a plurality of signal frames; wherein the same data exists between two adjacent signal frames; Feature extraction is performed on each of the signal frames to obtain the Mel spectrum coefficients of each of the signal frames, and the Mel spectrum coefficients of all the signal frames are collectively used as the sound features of the original sound signal.

[0005] In an optional implementation, the step of performing noise reduction processing on the original sound signal based on the ambient noise signal to obtain the noise-reduced sound signal includes: Calculate a noise parameter according to each noise sampling point in the environmental noise signal; The original sound signal is subjected to noise reduction processing according to the noise parameter to obtain the noise-reduced sound signal.

[0006] In an optional implementation, the step of performing noise reduction processing on the original sound signal according to the noise parameter to obtain the noise-reduced sound signal includes: Performing a fast Fourier transform on the original sound signal to obtain the amplitude and phase of each frequency component of the original sound signal; For each frequency component of the original sound signal, an initial power spectrum value is calculated according to the amplitude of the frequency component, and a noise reduction power spectrum value is calculated according to the initial power spectrum value and the noise parameter to obtain a noise reduction power spectrum value of each frequency component of the original sound signal; Based on the phases and noise reduction power spectrum values ​​of all frequency components of the original sound signal, an inverse fast Fourier transform is performed and the real part is taken to obtain the noise reduction sound signal.

[0007] In an optional implementation, the step of performing two-fold interpolation filtering on the noise reduction sound signal to obtain a filtered sound signal comprises: Inserting a placeholder between any two adjacent signal sampling points of the noise-reduced sound signal to obtain a sound signal to be interpolated; wherein the length of the sound signal to be interpolated is twice that of the noise-reduced sound signal; Using a preset interpolation filter, interpolation processing is performed on each placeholder in the sound signal to be interpolated, an interpolation sound signal is obtained, and filtering processing is performed to obtain the filtered sound signal.

[0008] In an optional implementation, the step of extracting features from each of the signal frames to obtain a Mel spectrum coefficient of each of the signal frames includes: Taking any one of the signal frames as a frame to be processed; Performing a fast Fourier transform on the frame to be processed to obtain a Fourier frequency value and an amplitude of each positive frequency component of the frame to be processed, and converting each of the Fourier frequency values ​​into a Mel frequency value to obtain a Mel frequency value and an amplitude of each positive frequency component of the frame to be processed; Calculating the Mel spectrum coefficient of the frame to be processed according to the preset Mel filter bank and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed; Each of the signal frames is traversed to obtain the Mel spectrum coefficient of each of the signal frames.

[0009] In an optional implementation, the step of calculating the Mel spectrum coefficient of the frame to be processed according to the preset Mel filter bank and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed includes: Acquire a Mel frequency sequence corresponding to the Mel filter group; the Mel frequency sequence includes a Mel frequency lower limit value, a Mel frequency center value, and a Mel frequency upper limit value of each Mel filter; For each of the Mel filters, determining a response value of the Mel filter for each positive frequency component of the frame to be processed according to a Mel frequency lower limit value, a Mel frequency center value and a Mel frequency upper limit value of the Mel filter and a Mel frequency value of each positive frequency component of the frame to be processed; Calculating a spectral coefficient of the Mel filter for the frame to be processed according to a response value of the Mel filter for each positive frequency component of the frame to be processed and an amplitude of each positive frequency component of the frame to be processed; The spectrum coefficients of each mel filter in the mel filter group for the frame to be processed are collectively used as the mel spectrum coefficients of the frame to be processed.

[0010] In a second aspect, the present invention provides a sound feature extraction device, the sound feature extraction device comprising: A noise reduction module, used for performing noise reduction processing on the original sound signal based on the environmental noise signal to obtain a noise-reduced sound signal; the environmental noise signal and the original sound signal are collected in the same environment according to a preset sampling frequency; Performing a two-fold interpolation filtering process on the noise reduction sound signal to obtain a filtered sound signal; wherein the signal strength within a preset frequency range in the filtered sound signal is higher than the signal strength within the frequency range in the noise reduction sound signal; and the signal strength exceeding the frequency range in the filtered sound signal is lower than the signal strength exceeding the frequency range in the noise reduction sound signal; A processing module, used for performing pre-emphasis processing and frame windowing processing on the filtered sound signal to obtain a plurality of signal frames; wherein the same data exists between two adjacent signal frames; The extraction module is used to perform feature extraction on each of the signal frames, obtain the Mel spectrum coefficients of each of the signal frames, and use the Mel spectrum coefficients of all the signal frames as the sound features of the original sound signal.

[0011] In an optional implementation, the noise reduction module is further used to: Calculate a noise parameter according to each noise sampling point in the environmental noise signal; The original sound signal is subjected to noise reduction processing according to the noise parameter to obtain the noise-reduced sound signal.

[0012] In a third aspect, the present invention provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the sound feature extraction method described in any one of the aforementioned embodiments is implemented.

[0013] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the sound feature extraction method described in any one of the aforementioned embodiments.

[0014] The present invention provides a sound feature extraction method, device, electronic device and storage medium, the method comprising: firstly, performing noise reduction processing based on an environmental noise signal and an original sound signal collected in the same environment to obtain a noise-reduced sound signal; then performing a two-fold interpolation filtering process on the noise-reduced sound signal to obtain a filtered sound signal, wherein the signal strength within the frequency range of the filtered sound signal is higher than that of the noise-reduced sound signal and the signal strength exceeding the frequency range in the filtered sound signal is lower than that of the noise-reduced sound signal; then performing pre-emphasis processing and frame-by-frame windowing processing on the filtered sound signal to obtain multiple signal frames; and there is the same data between two adjacent signal frames; finally, performing feature extraction on each signal frame to obtain the Mel spectrum coefficient of each signal frame, that is, to obtain the sound feature of the original sound signal. By performing noise reduction, interpolation and filtering processing on the collected sound signal, the noise in the sound signal is reduced and the required sound signal length is shortened, so as to extract the sound feature based on the sound signal with a good signal-to-noise ratio, thereby improving the quality of the sound feature and the efficiency of feature extraction.

[0015] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A block diagram of an electronic device provided by an embodiment of the present invention is shown; Figure 2 A schematic diagram of a flow chart of a sound feature extraction method provided by an embodiment of the present invention is shown; Figure 3 A schematic diagram showing the application effect of the existing sound feature extraction method; Figure 4A schematic diagram showing the application effect of the sound feature extraction method provided by an embodiment of the present invention; Figure 5 A functional module diagram of a sound feature extraction device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0018] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0020] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0021] See also Figure 1 , is a block diagram of an electronic device provided by an embodiment of the present invention. The electronic device includes a processor, a memory, and a communication module, and each component is directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0022] The processor is used to read / write data or programs stored in the memory and execute corresponding functions. It can be a general-purpose processor, including CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be a DSP digital signal processor, ASIC application-specific integrated circuit, FPGA ready-made programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0023] Memory is used to store programs or data. Memory can be RAM (Random Access Memory), ROM (Read Only Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electric Erasable Programmable Read-Only Memory), etc.

[0024] The communication module is used to communicate with other devices for signaling or data.

[0025] Understandably, Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device. The electronic device may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0026] The above-mentioned electronic device will be used as an execution subject to execute each step in each method provided in the embodiment of the present invention and achieve corresponding technical effects.

[0027] See also Figure 2 , is a flow chart of a sound feature extraction method provided in an embodiment of the present invention.

[0028] Step S202, performing noise reduction processing on the original sound signal based on the environmental noise signal to obtain a noise-reduced sound signal; the environmental noise signal and the original sound signal are collected in the same environment according to a preset sampling frequency.

[0029] Step S204, performing a two-fold interpolation filtering process on the noise reduction sound signal to obtain a filtered sound signal; wherein the signal strength within a preset frequency range in the filtered sound signal is higher than the signal strength within the frequency range in the noise reduction sound signal; and the signal strength exceeding the frequency range in the filtered sound signal is lower than the signal strength exceeding the frequency range in the noise reduction sound signal.

[0030] For ease of understanding, the embodiment of the present invention is explained by taking the extraction of sound features of a human voice signal as an example. In this embodiment, the environmental noise signal is first collected at a preset sampling frequency, such as 44kHz; then, in the same environment, the sound signal is collected at the sampling frequency, i.e., 44kHz, to obtain the original sound signal containing environmental noise and human voice. Then, based on the environmental noise signal, the original sound signal is subjected to noise reduction processing to obtain a noise-reduced sound signal. It can be understood that the embodiment of the present invention performs noise reduction processing on the sound signal by pre-collecting the environmental noise, so as to reduce the noise interference in the sound signal and improve the quality of the extracted sound features.

[0031] It is understandable that, based on the fact that the human ear is more sensitive to low-frequency signals, a frequency range that conforms to the auditory characteristics of the human ear, such as 0 Hz to 20 kHz, can be set, and the noise reduction sound signal can be filtered according to the frequency range to enhance the signal strength within the frequency range and weaken the signal strength outside the frequency range. In addition, the existing method requires a long time of sound signal to extract sufficient feature data, which will reduce the efficiency of sound feature extraction. Therefore, the embodiment of the present invention interpolates the short-time sound signal to ensure that sufficient feature data can be extracted while increasing the speed of sound feature extraction.

[0032] Then, based on the obtained noise reduction sound signal, a two-fold interpolation filtering process is performed on it, and a filtered sound signal having a length greater than the noise reduction sound signal can be obtained. In addition, the signal strength in the frequency range of 0 Hz to 20 kHz in the filtered sound signal is higher than the signal strength in the frequency range of 0 Hz to 20 kHz in the noise reduction sound signal, and the signal strength in the filtered sound signal exceeding the frequency range, that is, greater than 20 kHz, is lower than the signal strength in the noise reduction sound signal exceeding the frequency range, that is, greater than 20 kHz.

[0033] It can be understood that the embodiment of the present invention increases the length of the sound signal by performing a double interpolation filter on the sound signal after preliminary noise reduction to ensure that sufficient feature data can be extracted, and filters the high-frequency signal to which the human ear is insensitive to further reduce the noise in the sound signal, thereby improving the signal-to-noise ratio and providing data support for obtaining high-quality sound features.

[0034] Step S206, performing pre-emphasis processing and frame windowing processing on the filtered sound signal to obtain a plurality of signal frames; wherein the same data exists between two adjacent signal frames.

[0035] Step S208 , extracting features from each signal frame to obtain the Mel-spectrogram coefficients of each signal frame, and taking the Mel-spectrogram coefficients of all signal frames as the sound features of the original sound signal.

[0036] In this embodiment, based on the obtained filtered sound signal, a preset high-pass filter can be used to perform pre-emphasis processing on the filtered sound signal to compensate for the energy attenuation of the high frequency band in the filtered sound signal, that is, to obtain an emphasized sound signal. And the transfer function of the high-pass filter is: ; in, represents the transfer function, represents the pre-emphasis coefficient and , z represents the filtered sound signal.

[0037] Then, the emphasized sound signal is divided into frames according to a preset frame length L, that is, the emphasized sound signal is divided into a plurality of initial signal frames. Each initial signal frame includes L signal sampling points, and there is a period between two adjacent initial signal frames. Repeated signal sampling points. For example, assuming that the frame length L is 1024, each initial signal frame includes 1024 signal sampling points, and there are 512 repeated signal sampling points between two adjacent initial signal frames. In this way, the same signal data exists between two adjacent frames, thereby avoiding sudden changes between frames and improving the smoothness between two adjacent frames.

[0038] Next, each initial signal frame is multiplied by a preset window function such as a Hamming window function to obtain each signal frame. And the Hamming window function can be expressed as: ; in, Indicates the first The Hamming window coefficient of the signal sampling points, L represents the preset frame length.

[0039] Finally, feature extraction is performed on each signal frame to obtain the Mel spectrum coefficients representing the sound features of each signal frame, and the Mel spectrum coefficients of all signal frames are taken together as the sound features of the original sound signal, thus completing the extraction of the sound features of the original sound signal.

[0040] It can be understood that, compared with the prior art method of directly performing pre-emphasis processing, windowing and framing processing, and feature extraction on the collected sound signal, the embodiment of the present invention reduces the noise in the sound signal and shortens the required sound signal length by performing noise reduction, interpolation, and filtering processing on the collected sound signal, so as to extract sound features based on the sound signal with a good signal-to-noise ratio, thereby improving the quality of the sound features and the efficiency of feature extraction.

[0041] It can be seen that based on the above steps, firstly, noise reduction processing is performed based on the environmental noise signal and the original sound signal collected in the same environment to obtain a noise-reduced sound signal; then, the noise-reduced sound signal is subjected to double interpolation filtering processing to obtain a filtered sound signal, in which the signal strength within the frequency range of the filtered sound signal is higher than that of the noise-reduced sound signal, and the signal strength exceeding the frequency range in the filtered sound signal is lower than that of the noise-reduced sound signal; then, the filtered sound signal is subjected to pre-emphasis processing and frame-by-frame windowing processing to obtain multiple signal frames; and there is the same data between two adjacent signal frames; finally, feature extraction is performed on each signal frame to obtain the Mel spectrum coefficient of each signal frame, that is, the sound feature of the original sound signal is obtained. By performing noise reduction, interpolation and filtering processing on the collected sound signal, the noise in the sound signal is reduced and the required signal length is shortened, so as to extract sound features based on a sound signal with a good signal-to-noise ratio, thereby improving the quality of sound features and the efficiency of feature extraction.

[0042] Optionally, for step S202, an embodiment of the present invention provides a possible implementation method.

[0043] Step S202-1, calculating noise parameters according to each noise sampling point in the environmental noise signal. Step S202-3, performing noise reduction processing on the original sound signal according to the noise parameter to obtain a noise-reduced sound signal.

[0044] In this embodiment, the environmental noise signal includes multiple noise sampling points. For example, 2000 noise sampling points can be collected at a sampling frequency of 44kHz to obtain an environmental noise signal with a duration of about 45ms, and then noise parameters are calculated based on all noise sampling points in the environmental noise signal. The original sound signal is then subjected to noise reduction processing according to the noise parameters to obtain a noise-reduced sound signal.

[0045] It can be understood that the process of calculating the noise parameter based on all noise sampling points in the ambient noise signal can be expressed by a formula, namely: ; in, represents the noise parameter, represents the d+1th noise sampling point in the ambient noise signal; D represents the total number of noise sampling points in the ambient noise signal.

[0046] Optionally, for step S202-3, an embodiment of the present invention provides a possible implementation method.

[0047] Step S202-3-1, perform fast Fourier transform on the original sound signal to obtain the amplitude and phase of each frequency component of the original sound signal.

[0048] Step S202-3-3, for each frequency component of the original sound signal, calculate the initial power spectrum value according to the amplitude of the frequency component, and calculate the noise reduction power spectrum value according to the initial power spectrum value and the noise parameter to obtain the noise reduction power spectrum value of each frequency component of the original sound signal.

[0049] Step S202-3-5, based on the phase and noise reduction power spectrum values ​​of all frequency components of the original sound signal, perform inverse fast Fourier transform and take the real part to obtain the noise reduction sound signal.

[0050] In this embodiment, first, the original sound signal is subjected to a fast Fourier transform, and then spectrum coefficients of multiple frequency components can be obtained. These frequency coefficients are all complex numbers, and one spectrum coefficient represents the amplitude and phase of one frequency component.

[0051] For example, the original sound signal can be expressed as , and the spectral coefficients of multiple frequency components obtained based on the original sound signal can be expressed as .in, represents the n+1th signal sampling point in the original sound signal, N represents the total number of signal sampling points in the original sound signal, represents the spectral coefficient of the k+1th frequency component, and K represents the total number of frequency components of the original sound signal. And, Indicates frequency coefficient The modulus is the amplitude of the k+1th frequency component.

[0052] Then, for each frequency component of the original sound signal, an initial power spectrum value can be calculated according to the amplitude of the frequency component. And the noise reduction power spectrum value is calculated according to the initial power spectrum value and the noise parameter to obtain the noise reduction power spectrum value of the frequency component. By processing each frequency component in a similar manner, the noise reduction power spectrum value of each frequency component of the original sound signal can be obtained.

[0053] It can be understood that the calculation process of the noise reduction power spectrum value of a frequency component in the original sound signal can be expressed by a formula, namely: ; ; in, represents the initial power spectrum value of the k+1th frequency component, represents the denoised power spectrum value of the k+1th frequency component, represents the noise parameter, Represents the noise power.

[0054] It can be understood that if the initial power spectrum value of a frequency component is greater than the noise power calculated based on the noise parameters, the frequency component is considered to be a human voice, and the difference between the initial power spectrum value of the frequency component and the noise power is used as its noise reduction power spectrum value to reduce noise interference. If the initial power spectrum value of a frequency component is less than or equal to the noise power calculated based on the noise parameters, the frequency component is considered to be noise, and the noise reduction power spectrum value of the frequency component is set to 0 to reduce noise interference.

[0055] Finally, based on the phase and noise reduction power spectrum values ​​of all frequency components of the original sound signal, an inverse fast Fourier transform is performed and the real part is taken to obtain the noise reduction sound signal. For example, the noise reduction sound signal can be expressed as ,in represents the n+1th signal sampling point in the noise-reduced audio signal, and N represents the total number of signal sampling points in the noise-reduced audio signal.

[0056] Optionally, for step S204, an embodiment of the present invention provides a possible implementation method.

[0057] Step S204-1, inserting a placeholder between any two adjacent signal sampling points of the noise-reduced sound signal to obtain a sound signal to be interpolated; wherein the length of the sound signal to be interpolated is twice that of the noise-reduced sound signal.

[0058] Step S204 - 3 , using a preset interpolation filter, interpolate each placeholder in the sound signal to be interpolated, obtain the interpolated sound signal, and perform filtering to obtain a filtered sound signal.

[0059] In this embodiment, based on the obtained noise reduction sound signal, a placeholder may be inserted between any two adjacent signal sampling points of the noise reduction sound signal to obtain the sound signal to be interpolated. For example, 0 may be used as a placeholder, and the relationship between the sound signal to be interpolated and the noise reduction sound signal may be expressed as: ; in, Indicates the 2n+1th signal sampling point in the sound signal to be interpolated. represents the n+1th signal sampling point in the noise reduction sound signal. And the length of the sound signal to be interpolated is 2N, that is, its length is twice that of the noise reduction sound signal.

[0060] Then, a preset interpolation filter is used to interpolate each placeholder in the interpolated sound signal to calculate the specific value of the signal sampling point at each placeholder, thereby obtaining the interpolated sound signal; and the interpolated sound signal is filtered, that is, the signal with a frequency between 0 Hz and 20 kHz in the interpolated sound signal is enhanced, and the signal with a frequency exceeding 20 kHz is attenuated, thereby obtaining the filtered sound signal.

[0061] It can be understood that in order to enable the interpolation filter to achieve double interpolation filtering, the signal within the preset frequency range is enhanced, and the signal exceeding the preset frequency range is attenuated. The embodiment of the present invention also provides the coefficient reference value of the interpolation filter after 16-bit quantization, that is, multiplying by 2 to the 15th power and rounding down, that is: {0,-321,0,345,0,-373,0,406,0,-444,0,488,0,-542,0,609,0,-693,0,802,0,-950,0,1164,0,-1500, 0,2103,0,-3508,0,10528,16539,10528,0,-3508,0,2103,0,-1500,0,1164,0,-950,0,802,0,-693,0,609,0,-542,0,488,0,-444,0,406,0,-373,0,345,0,-321,0}.

[0062] Optionally, for step S208, an embodiment of the present invention provides a possible implementation method.

[0063] Step S208-1, taking any signal frame as a frame to be processed.

[0064] Step S208-3, perform fast Fourier transform on the frame to be processed to obtain the Fourier frequency value and amplitude of each positive frequency component of the frame to be processed, and convert each Fourier frequency value into a Mel frequency value to obtain the Mel frequency value and amplitude of each positive frequency component of the frame to be processed.

[0065] Step S208-5: Calculate the Mel spectrum coefficient of the frame to be processed according to the preset Mel filter bank and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed.

[0066] Step S208-7, traverse each signal frame to obtain the Mel spectrum coefficients of each signal frame.

[0067] It is understandable that the processing method for each signal frame in the embodiment of the present invention is similar. For the sake of simplicity, a signal frame is used as an example for explanation below. For example, the frame to be processed can be represented as .in represents the lth signal sampling point in the frame to be processed, and L represents the preset frame length, that is, the total number of signal sampling points in the frame to be processed.

[0068] First, according to the preset number of transformation points J, such as 1024, the frame to be processed is subjected to fast Fourier transform, and the Fourier frequency value and amplitude of each positive frequency component are obtained, and then The Fourier frequency value and amplitude of the positive frequency components.

[0069] It can be understood that since the signal sampling points in the frame to be processed are all real numbers, after the frame to be processed is transformed into the Fourier domain, the frequency components in its Fourier spectrum are symmetrical, that is, each positive frequency component has a corresponding negative frequency component, so the sound features of the frame to be processed can be obtained by analyzing only the positive frequency components. This can not only reduce the complexity of calculation, but also improve the speed of extracting sound features.

[0070] Then, the Fourier frequency value of each positive frequency component of the frame to be processed is converted into a Mel frequency value, and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed are obtained. It can be understood that the process of converting the Fourier frequency value of a positive frequency component in the frame to be processed into a Mel frequency value can be expressed by a formula, namely: ; in, Represents the Mel frequency value of the j+1th positive frequency component in the frame to be processed, represents the Fourier frequency value of the j+1th positive frequency component in the frame to be processed, represents the total number of positive frequency components of the frame to be processed, and J represents the total number of frequency components of the frame to be processed.

[0071] Finally, the Mel spectrum coefficient of the frame to be processed is calculated according to the preset Mel filter bank and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed, that is, the Mel spectrum coefficient of a signal frame is obtained. By processing each signal frame in a similar manner, the Mel spectrum coefficient of each signal frame can be obtained.

[0072] Optionally, for step S208-5, an embodiment of the present invention provides a possible implementation method.

[0073] Step S208-5-1, obtaining the Mel frequency sequence corresponding to the Mel filter group; the Mel frequency sequence includes the Mel frequency lower limit value, the Mel frequency center value and the Mel frequency upper limit value of each Mel filter.

[0074] Step S208-5-3, for each Mel filter, determine the response value of the Mel filter for each positive frequency component of the frame to be processed based on the Mel frequency lower limit value, Mel frequency center value and Mel frequency upper limit value of the Mel filter, and the Mel frequency value of each positive frequency component of the frame to be processed.

[0075] Step S208-5-5, calculating the spectral coefficient of the Mel filter for the frame to be processed based on the response value of the Mel filter for each positive frequency component of the frame to be processed and the amplitude of each positive frequency component of the frame to be processed. Step S208-5-7, taking the spectrum coefficients of each mel filter in the mel filter group for the frame to be processed as the mel spectrum coefficients of the frame to be processed.

[0076] In this embodiment, first, a Mel frequency sequence corresponding to the Mel filter bank is obtained, where the Mel frequency sequence includes a Mel frequency lower limit value, a Mel frequency center value, and a Mel frequency upper limit value of each Mel filter.

[0077] For example, assuming that the Mel filter bank includes M Mel filters, the Mel frequency sequence includes M+2 elements, and the first element in the Mel frequency sequence represents the Mel frequency lower limit value of the first Mel filter, the last element represents the Mel frequency upper limit value of the Mth Mel filter, i.e., the last Mel filter, and the remaining M elements represent the Mel frequency center value of each Mel filter. Moreover, the Mel frequency upper limit value of the mth Mel filter is the Mel frequency center value of the m+1th Mel filter, and the Mel frequency center value of the m+1th Mel filter is also the Mel frequency lower limit value of the m+2th Mel filter.

[0078] Then, for each Mel filter, the response value of the Mel filter for each positive frequency component of the frame to be processed is determined according to the Mel frequency lower limit value, Mel frequency center value and Mel frequency upper limit value of the Mel filter and the Mel frequency value of each positive frequency component of the frame to be processed. It can be understood that the process of determining the response value of a Mel filter for a positive frequency component of the frame to be processed can be expressed by a formula, namely: ; in, represents the response value of the mth Mel filter to the j+1th positive frequency component of the frame to be processed, Indicates the Mel frequency lower limit of the mth Mel filter, represents the Mel frequency center value of the mth Mel filter, Represents the Mel frequency upper limit of the mth Mel filter, Represents the Mel frequency value of the j+1th positive frequency component of the frame to be processed, represents the total number of positive frequency components of the frame to be processed, J represents the total number of frequency components of the frame to be processed, and M represents the total number of Mel filters in the Mel filter bank.

[0079] Next, the spectral coefficient of the Mel filter for the frame to be processed is calculated based on the response value of the Mel filter for each positive frequency component of the frame to be processed and the amplitude of each positive frequency component of the frame to be processed. It can be understood that the calculation process of the spectral coefficient of a Mel filter for the frame to be processed can be expressed by a formula, namely: ; in, represents the spectral coefficient of the mth Mel filter for the frame to be processed, represents the amplitude of the j+1th positive frequency component of the frame to be processed, represents the response value of the mth Mel filter to the j+1th positive frequency component of the frame to be processed, represents the total number of positive frequency components of the frame to be processed, J represents the total number of frequency components of the frame to be processed, and M represents the total number of Mel filters in the Mel filter bank.

[0080] Finally, according to the M Mel filters in the Mel filter bank, the spectrum coefficient of each Mel filter for the frame to be processed is calculated, and M spectrum coefficients are obtained, and these M spectrum coefficients are collectively used as the Mel spectrum coefficients of the frame to be processed.

[0081] It can be understood that when M is 64, the Mel filter bank includes 64 Mel filters, then the Mel frequency sequence includes 66 elements, and these 66 elements can be set to: {0,60,120,180,241,301,361,422,482,542,603,663,723,784,844,904,965,1025,1085,1146,1206,1266,1327,1387,1447,1508,1568,1628,1688,1749,1809,1869,1930,1990,2050, 2111,2171,2231,2292,2352,2412,2473,2533,2593,2654,2714,2774,2835,2895,2955,3016,3076,3136,3197,3257,3317,3377,3438,3498,3558,3619,3679,3739,3800,3860,3920}. Among them, 0 represents the Mel frequency lower limit of the first Mel filter, 3920 represents the Mel frequency upper limit of the 64th filter, and the remaining 64 elements represent the Mel frequency center values ​​of the 64 filters respectively.

[0082] In order to better understand the technical effect of the present invention. The embodiment of the present invention provides a set of comparative examples. The existing sound feature extraction method is used to extract features from a 2.32 second sound signal collected at a sampling frequency of 44kHz, and the obtained sound features are used to train a speech recognition model. The variation trend of the recognition accuracy of the speech recognition model with the number of iterations is shown in the figure. Figure 3 shown. Figure 3 The light blue line in the figure represents the actual recognition accuracy of the speech recognition model during the training process, the dark blue line represents the recognition accuracy obtained after smoothing the actual recognition accuracy of the speech recognition model, the black dotted line represents the recognition accuracy of the language recognition model during the verification process, and the red dotted line indicates that after multiple verifications, the recognition accuracy of the speech recognition model is finally 58%.

[0083] The method provided in the embodiment of the present invention is used to extract features from a 1.16 second sound signal collected at a sampling frequency of 44kHz, and the obtained sound features are used to train a speech recognition model. The variation trend of the recognition accuracy of the speech recognition model with the number of iterations is shown in the figure below: Figure 4 shown. Figure 4 The light blue line in the figure represents the actual recognition accuracy of the speech recognition model during the training process, the dark blue line represents the recognition accuracy obtained after smoothing the actual recognition accuracy of the speech recognition model, the black dotted line represents the recognition accuracy of the language recognition model during the verification process, and the red dotted line indicates that after multiple verifications, the recognition accuracy of the speech recognition model is finally 59%.

[0084] It can be seen that compared with the prior art, the results of training the speech recognition model using the sound features obtained from shorter sound signals in the embodiment of the present invention are better than those in the prior art. That is, the embodiment of the present invention can improve the extraction efficiency of sound features while ensuring the quality of sound features.

[0085] In order to execute the corresponding steps in the above embodiments and various possible methods, a method for implementing a sound feature extraction device is given below. Figure 5 , is a functional module diagram of a sound feature extraction device provided in an embodiment of the present invention. It should be noted that the basic principle and technical effects of the sound feature extraction device provided in this embodiment are the same as those of the above embodiment. For the sake of brief description, for matters not mentioned in this embodiment, reference may be made to the corresponding contents in the above embodiment. The sound feature extraction device includes: The noise reduction module is used to perform noise reduction processing on the original sound signal based on the environmental noise signal to obtain the noise-reduced sound signal; the environmental noise signal and the original sound signal are collected in the same environment according to the preset sampling frequency. The noise-reduced sound signal is subjected to double interpolation filtering processing to obtain the filtered sound signal; wherein the signal strength within the preset frequency range in the filtered sound signal is higher than the signal strength within the frequency range in the noise-reduced sound signal; the signal strength exceeding the frequency range in the filtered sound signal is lower than the signal strength exceeding the frequency range in the noise-reduced sound signal.

[0086] The processing module is used to perform pre-emphasis processing and frame windowing processing on the filtered sound signal to obtain multiple signal frames; wherein the same data exists between two adjacent signal frames.

[0087] The extraction module is used to extract features from each signal frame, obtain the Mel spectrum coefficients of each signal frame, and use the Mel spectrum coefficients of all signal frames as the sound features of the original sound signal.

[0088] Optionally, the noise reduction module is further used to: calculate a noise parameter according to each noise sampling point in the ambient noise signal; and perform noise reduction processing on the original sound signal according to the noise parameter to obtain a noise-reduced sound signal.

[0089] Optionally, the noise reduction module is also used to: perform fast Fourier transform on the original sound signal to obtain the amplitude and phase of each frequency component of the original sound signal; for each frequency component of the original sound signal, calculate the initial power spectrum value according to the amplitude of the frequency component, and calculate the noise reduction power spectrum value according to the initial power spectrum value and the noise parameter to obtain the noise reduction power spectrum value of each frequency component of the original sound signal; based on the phase and noise reduction power spectrum values ​​of all frequency components of the original sound signal, perform inverse fast Fourier transform and take the real part to obtain the noise reduction sound signal.

[0090] Optionally, the noise reduction module is also used to: insert a placeholder between any two adjacent signal sampling points of the noise reduction sound signal to obtain a sound signal to be interpolated; wherein the length of the sound signal to be interpolated is twice that of the noise reduction sound signal; use a preset interpolation filter to interpolate each placeholder in the sound signal to be interpolated, obtain the interpolated sound signal and perform filtering to obtain a filtered sound signal.

[0091] Optionally, the extraction module is also used to: take any signal frame as a frame to be processed; perform fast Fourier transform on the frame to be processed to obtain the Fourier frequency value and amplitude of each positive frequency component of the frame to be processed, and convert each Fourier frequency value into a Mel frequency value to obtain the Mel frequency value and amplitude of each positive frequency component of the frame to be processed; calculate the Mel spectrum coefficient of the frame to be processed according to a preset Mel filter group and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed; traverse each signal frame to obtain the Mel spectrum coefficient of each signal frame.

[0092] Optionally, the extraction module is also used to: obtain a Mel frequency sequence corresponding to the Mel filter group; the Mel frequency sequence includes a Mel frequency lower limit value, a Mel frequency center value, and a Mel frequency upper limit value of each Mel filter; for each Mel filter, determine the response value of the Mel filter to each positive frequency component of the frame to be processed according to the Mel frequency lower limit value, the Mel frequency center value, and the Mel frequency upper limit value of the Mel filter, and the Mel frequency value of each positive frequency component of the frame to be processed; calculate the spectral coefficient of the Mel filter for the frame to be processed according to the response value of the Mel filter to each positive frequency component of the frame to be processed and the amplitude of each positive frequency component of the frame to be processed; and use the spectral coefficient of each Mel filter in the Mel filter group for the frame to be processed as the Mel spectral coefficient of the frame to be processed.

[0093] An embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the sound feature extraction method disclosed in the embodiment of the present invention is implemented.

[0094] The embodiment of the present invention further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the sound feature extraction method disclosed in the embodiment of the present invention is implemented.

[0095] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0096] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0097] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0098] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for extracting sound features, characterized in that: The sound feature extraction method comprises: Performing noise reduction processing on the original sound signal based on the environmental noise signal to obtain a noise-reduced sound signal; the environmental noise signal and the original sound signal are collected in the same environment according to a preset sampling frequency; Performing a two-fold interpolation filtering process on the noise reduction sound signal to obtain a filtered sound signal; wherein the signal strength within a preset frequency range in the filtered sound signal is higher than the signal strength within the frequency range in the noise reduction sound signal; and the signal strength exceeding the frequency range in the filtered sound signal is lower than the signal strength exceeding the frequency range in the noise reduction sound signal; Performing pre-emphasis processing and frame windowing processing on the filtered sound signal to obtain a plurality of signal frames; wherein the same data exists between two adjacent signal frames; Feature extraction is performed on each of the signal frames to obtain the Mel spectrum coefficients of each of the signal frames, and the Mel spectrum coefficients of all the signal frames are collectively used as the sound features of the original sound signal.

2. The sound feature extraction method according to claim 1, characterized in that: The step of performing noise reduction processing on the original sound signal based on the environmental noise signal to obtain the noise-reduced sound signal comprises: Calculate a noise parameter according to each noise sampling point in the environmental noise signal; The original sound signal is subjected to noise reduction processing according to the noise parameter to obtain the noise-reduced sound signal.

3. The sound feature extraction method according to claim 2, characterized in that: The step of performing noise reduction processing on the original sound signal according to the noise parameter to obtain the noise-reduced sound signal comprises: Performing a fast Fourier transform on the original sound signal to obtain the amplitude and phase of each frequency component of the original sound signal; For each frequency component of the original sound signal, an initial power spectrum value is calculated according to the amplitude of the frequency component, and a noise reduction power spectrum value is calculated according to the initial power spectrum value and the noise parameter to obtain a noise reduction power spectrum value of each frequency component of the original sound signal; Based on the phases and noise reduction power spectrum values ​​of all frequency components of the original sound signal, an inverse fast Fourier transform is performed and the real part is taken to obtain the noise reduction sound signal.

4. The sound feature extraction method according to claim 1, characterized in that: The step of performing two-fold interpolation filtering on the noise reduction sound signal to obtain a filtered sound signal comprises: Inserting a placeholder between any two adjacent signal sampling points of the noise-reduced sound signal to obtain a sound signal to be interpolated; wherein the length of the sound signal to be interpolated is twice that of the noise-reduced sound signal; Using a preset interpolation filter, interpolation processing is performed on each placeholder in the sound signal to be interpolated, an interpolation sound signal is obtained, and filtering processing is performed to obtain the filtered sound signal.

5. The sound feature extraction method according to claim 1, characterized in that: The step of extracting features from each of the signal frames to obtain the Mel spectrum coefficients of each of the signal frames comprises: Taking any one of the signal frames as a frame to be processed; Performing a fast Fourier transform on the frame to be processed to obtain a Fourier frequency value and an amplitude of each positive frequency component of the frame to be processed, and converting each of the Fourier frequency values ​​into a Mel frequency value to obtain a Mel frequency value and an amplitude of each positive frequency component of the frame to be processed; Calculating the Mel spectrum coefficient of the frame to be processed according to the preset Mel filter bank and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed; Each of the signal frames is traversed to obtain the Mel spectrum coefficient of each of the signal frames.

6. The sound feature extraction method according to claim 5, characterized in that: The step of calculating the Mel spectrum coefficient of the frame to be processed according to the preset Mel filter group and the Mel frequency value and amplitude of each positive frequency component of the frame to be processed comprises: Acquire a Mel frequency sequence corresponding to the Mel filter group; the Mel frequency sequence includes a Mel frequency lower limit value, a Mel frequency center value, and a Mel frequency upper limit value of each Mel filter; For each of the Mel filters, determining a response value of the Mel filter for each positive frequency component of the frame to be processed according to a Mel frequency lower limit value, a Mel frequency center value and a Mel frequency upper limit value of the Mel filter and a Mel frequency value of each positive frequency component of the frame to be processed; Calculating a spectral coefficient of the Mel filter for the frame to be processed according to a response value of the Mel filter for each positive frequency component of the frame to be processed and an amplitude of each positive frequency component of the frame to be processed; The spectrum coefficients of each mel filter in the mel filter group for the frame to be processed are collectively used as the mel spectrum coefficients of the frame to be processed.

7. A sound feature extraction device, characterized in that: The sound feature extraction device comprises: A noise reduction module, used for performing noise reduction processing on the original sound signal based on the environmental noise signal to obtain a noise-reduced sound signal; the environmental noise signal and the original sound signal are collected in the same environment according to a preset sampling frequency; Performing a two-fold interpolation filtering process on the noise reduction sound signal to obtain a filtered sound signal; wherein the signal strength within a preset frequency range in the filtered sound signal is higher than the signal strength within the frequency range in the noise reduction sound signal; and the signal strength exceeding the frequency range in the filtered sound signal is lower than the signal strength exceeding the frequency range in the noise reduction sound signal; A processing module, used for performing pre-emphasis processing and frame windowing processing on the filtered sound signal to obtain a plurality of signal frames; wherein the same data exists between two adjacent signal frames; The extraction module is used to perform feature extraction on each of the signal frames, obtain the Mel spectrum coefficients of each of the signal frames, and use the Mel spectrum coefficients of all the signal frames as the sound features of the original sound signal.

8. The sound feature extraction device according to claim 7, characterized in that: The noise reduction module is also used for: Calculate a noise parameter according to each noise sampling point in the environmental noise signal; The original sound signal is subjected to noise reduction processing according to the noise parameter to obtain the noise-reduced sound signal.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the method for extracting sound features according to any one of claims 1 to 6 is implemented.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the sound feature extraction method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Class-D power amplifier, compensation method and digital signal processing device thereof

    CN109217827A

  • Device voice noise reduction, electronic device and storage medium

    CN114121031A

  • Audio signal synchronization method and device, equipment and storage medium

    CN115223578A

  • Infant crying emotion recognition method based on feature map lightweight convolution transformation

    CN115565550A

  • Audio quality analysis method and device, electronic equipment and storage medium

    CN116013367A