Audio recognition method and device, electronic device, and storage medium

By using Fourier transform and power spectral density calculation, the speech signal is dynamically enhanced, solving the problem of insufficient noise suppression in complex environments and improving the accuracy and precision of audio recognition.

CN120913546BActive Publication Date: 2026-01-27BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511454056.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-01-27
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

Existing audio recognition technologies struggle to adapt to dynamically changing noise in complex environments, resulting in insufficient background noise suppression and impacting recognition accuracy.

Method used

The first power spectral density of the non-speech signal and the second power spectral density of the speech signal are obtained by Fourier transform. The power ratio is calculated and combined with the dynamic gain factor to enhance the signal, suppress background noise, and improve the accuracy of speech signal recognition.

Benefits of technology

It effectively reduces background noise interference, improves audio recognition accuracy in complex scenarios, and ensures the clarity and accuracy of speech signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913546B_ABST
    Figure CN120913546B_ABST
Patent Text Reader

Abstract

The application provides an audio recognition method and device, an electronic device and a storage medium, and belongs to the technical field of audio processing. The method comprises the following steps: performing Fourier transform on a non-speech signal to obtain a noise frequency domain signal, and calculating a first power spectral density of a background noise signal in the non-speech signal; performing Fourier transform on a speech signal to obtain a speech frequency domain signal, and calculating a second power spectral density corresponding to an initial speech signal based on the speech frequency domain signal and the first power spectral density; calculating a power ratio of the initial speech signal and the background noise signal in the speech signal; performing signal enhancement on the initial speech signal based on the power ratio, the speech frequency domain signal, the first power spectral density and the second power spectral density to obtain a target speech signal; and performing recognition on the target speech signal based on a target model. The audio recognition method and device, the electronic device and the storage medium provided by the application can improve the accuracy of audio recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio processing technology, and more specifically, relates to an audio recognition method and apparatus, electronic device, and storage medium. Background Technology

[0002] As wearable devices, smart glasses need to achieve accurate audio recognition in complex environments to support functions such as voice interaction. However, their application scenarios often involve various background noises, such as traffic noise and conversations among people, which cause the collected audio signals to be mixed with a lot of interference, seriously affecting the recognition accuracy.

[0003] Existing audio recognition technologies mostly employ fixed filtering or single-model recognition, which makes it difficult to adapt to dynamically changing noise environments. This results in insufficient suppression of background noise, easily causing speech distortion and thus affecting the accuracy of audio recognition. Summary of the Invention

[0004] The purpose of this application is to provide an audio recognition method, device, electronic device, and storage medium to improve the accuracy of audio recognition.

[0005] A first aspect of this application provides an audio recognition method, including:

[0006] Acquire speech and non-speech signals from the target audio signal; the target audio signal is the preprocessed audio signal of the original mixed audio signal; the speech signal includes the user's initial speech signal and background noise signal, and the non-speech signal includes background noise signal;

[0007] Perform a Fourier transform on the non-speech signal to obtain the noise frequency domain signal corresponding to the non-speech signal, and calculate the first power spectral density of the background noise signal in the non-speech signal.

[0008] Perform a Fourier transform on the speech signal to obtain the corresponding audio domain signal, and calculate the second power spectral density corresponding to the initial speech signal based on the audio domain signal and the first power spectral density; calculate the power ratio of the initial speech signal to the background noise signal in the speech signal.

[0009] The target speech signal is obtained by enhancing the initial speech signal based on the power ratio, the speech frequency domain signal, the first power spectral density and the second power spectral density.

[0010] The target speech signal is identified based on the target model.

[0011] A second aspect of this application provides an audio recognition device, comprising:

[0012] The audio acquisition module is used to acquire speech signals and non-speech signals from the target audio signal; the target audio signal is the audio signal after preprocessing the original mixed audio signal; the speech signal includes the user's initial speech signal and background noise signal, and the non-speech signal includes background noise signal;

[0013] The first calculation module is used to perform Fourier transform on non-speech signals to obtain the noise frequency domain signal corresponding to the non-speech signals, and to calculate the first power spectral density of the background noise signal in the non-speech signals.

[0014] The second calculation module is used to perform Fourier transform on the speech signal to obtain the corresponding audio domain signal, and calculate the second power spectral density corresponding to the initial speech signal based on the audio domain signal and the first power spectral density; and calculate the power ratio of the initial speech signal to the background noise signal in the speech signal.

[0015] The signal enhancement module is used to enhance the initial speech signal based on the power ratio, the speech frequency domain signal, the first power spectral density and the second power spectral density to obtain the target speech signal.

[0016] The signal recognition module is used to recognize target speech signals based on the target model.

[0017] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the audio recognition method described above.

[0018] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the audio recognition method described above.

[0019] The beneficial effects of the audio recognition method, apparatus, electronic device, and storage medium provided in this application are as follows: This application divides the preprocessed target audio signal into speech signals and non-speech signals. The first power spectral density obtained by Fourier transform of the non-speech signal provides a dynamic and accurate background noise reference for noise suppression, overcoming the limitation of fixed filtering lacking a targeted noise reference. Then, Fourier transform is performed on the speech signal, and the second power spectral density of the initial speech signal is calculated in conjunction with the first power spectral density, along with the power ratio between the two, thus distinguishing between speech and noise energy. This avoids the problem of existing technologies being unable to dynamically identify the proportion of speech and noise energy. Finally, the initial speech signal is enhanced based on the power ratio, the speech domain signal, and the first and second power spectral densities to obtain the target speech signal, effectively reducing background noise interference and suppressing speech distortion to a certain extent. The target speech signal is then recognized through a target model to improve the accuracy of audio recognition in complex scenarios. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A schematic flowchart of an audio recognition method provided in an embodiment of this application;

[0022] Figure 2 This is a structural block diagram of an audio recognition device provided in an embodiment of this application;

[0023] Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0026] Please refer to Figure 1 , Figure 1This is a flowchart illustrating an audio recognition method provided in an embodiment of this application. The method can be executed by an electronic device, which can be a wearable device, such as smart glasses, a terminal device, such as a mobile phone or tablet computer, or a server. The method may include:

[0027] S101: Acquire speech signals and non-speech signals from the target audio signal; the target audio signal is the audio signal after preprocessing the original mixed audio signal; the speech signals include the user's initial speech signal and background noise signal, and the non-speech signals include background noise signal.

[0028] In this embodiment, the original mixed audio signal refers to the unprocessed audio signal directly acquired by the smart glasses microphone, containing a mixture of various sound components such as user voice and environmental background noise (e.g., traffic sounds, conversations). The target audio signal refers to the audio signal obtained after preprocessing, which retains valid voice and background noise information while removing some interference components (e.g., DC components, extreme amplitude distortion). A voice-type signal refers to a segment of the target audio signal that contains user voice components (i.e., the user's initial voice signal), which is also mixed with background noise; that is, it consists of "the user's initial voice signal + background noise signal". A non-voice-type signal refers to a segment of the target audio signal that does not contain user voice components and consists only of background noise signal.

[0029] In this embodiment, the raw mixed audio signal, which includes user voice, ambient background noise, and other mixed sounds, can be acquired through the built-in microphone of the smart glasses. For example, the sampling frequency is set to 16kHz and the quantization bit depth is 16bit. The raw mixed audio signal is then preprocessed to obtain the target audio signal.

[0030] The preprocessing process may include: using a 10Hz high-pass filter to filter the original mixed audio signal to remove the DC component and low-frequency vibration interference (such as low-frequency vibration of subway tracks) from the signal.

[0031] Automatic gain control is performed to normalize the signal amplitude to the range of [-1, 1], avoiding signal distortion caused by excessively high or low input volume.

[0032] The target audio signal is obtained after filtering and normalizing the original mixed audio signal.

[0033] In this embodiment, a speech activity detection algorithm based on a Gaussian mixture model can be used to acquire the target audio signal. Signals with short-time energy below an energy threshold and a short-time zero-crossing rate within a preset range are classified as non-speech signals, which only contain background noise. Signals with short-time energy above the energy threshold are classified as speech signals, which include the user's initial speech signal and background noise.

[0034] S102: Perform Fourier transform on the non-speech signal to obtain the noise frequency domain signal corresponding to the non-speech signal, and calculate the first power spectral density of the background noise signal in the non-speech signal.

[0035] In this embodiment, the Fourier transform can convert a non-speech signal from a time-varying amplitude representation to a frequency-varying spectral representation, thereby analyzing the energy distribution of different frequency components in the signal. The noise frequency domain signal is the frequency domain representation of the non-speech signal after Fourier transform, reflecting the amplitude and phase information of the background noise signal at different frequency points. The first power spectral density is the power spectral density of the background noise signal in the non-speech signal, a physical quantity describing the power of the background noise signal as a function of frequency. It is calculated by squared-down processing of the noise frequency domain signal's amplitude and combining it with frequency resolution, and is used to quantify the energy intensity of the background noise at various frequencies.

[0036] In this embodiment, before performing a Fourier transform on the non-speech signal, a windowing process (e.g., setting the window length to 256ms) can be applied to reduce spectral leakage during the Fourier transform. A Fast Fourier Transform (FFT) is then performed on the windowed non-speech signal to convert it from the time domain to the frequency domain, resulting in a noise frequency domain signal. The FFT points are set to 256 to match the window length, achieving the time-domain to frequency-domain conversion and obtaining the signal's distribution characteristics at different frequency components.

[0037] The power spectral density of the noise frequency domain signal is calculated by taking the modulus of the complex frequency domain signal obtained by Fourier transform and squaring the modulus. The squared result is then divided by the window function correction coefficient (e.g., the correction coefficient for the Hanning window is 1.63) and the frequency resolution (frequency resolution = sampling frequency / number of fast Fourier transform points, i.e., 16kHz / 256 = 62.5Hz) to obtain the power spectral density of a single frame of non-speech signal. The power spectral density of multiple consecutive frames is then averaged to obtain the first power spectral density of the background noise signal in the non-speech signal, which is used for subsequent noise feature analysis and speech enhancement processing.

[0038] S103: Perform Fourier transform on the speech signal to obtain the corresponding audio domain signal, and calculate the second power spectral density corresponding to the initial speech signal based on the audio domain signal and the first power spectral density; calculate the power ratio of the initial speech signal to the background noise signal in the speech signal.

[0039] In this embodiment, the speech frequency domain signal is the frequency domain representation of the speech signal after Fourier transform, reflecting the amplitude and phase mixing information of the initial speech signal and the background noise signal at different frequency points. The second power spectral density refers to the power spectral density of the initial speech signal, which is the pure speech energy distribution estimated by combining the speech frequency domain signal with the first power spectral density, describing the power variation characteristics of the initial speech signal with frequency. The power ratio is the power ratio of the initial speech signal to the background noise signal in the speech signal, calculated by the first average energy of the non-speech signal and the second average energy of the speech signal, reflecting the overall energy strength relationship between the two.

[0040] In this embodiment, before performing a Fourier transform on the speech signal, a Hanning window, the same as for non-speech signals, is used for windowing to reduce the impact of spectral leakage on subsequent frequency domain analysis. A Fast Fourier Transform is then performed on the windowed speech signal to convert the time-domain speech signal into a frequency-domain representation, resulting in a speech audio domain signal. This signal contains the amplitude and phase information of the initial speech and background noise at each frequency component.

[0041] In one embodiment of this disclosure, calculating the second power spectral density corresponding to the initial speech signal based on the audio domain signal and the first power spectral density includes:

[0042] Calculate the third power spectral density corresponding to the audio domain signal, and obtain the estimated power spectral density of the initial speech signal based on the difference between the third power spectral density and the first power spectral density.

[0043] The posterior signal-to-noise ratio is calculated based on the ratio of the third power spectral density to the first power spectral density.

[0044] The dynamic gain factor of the initial speech signal is determined based on the posterior signal-to-noise ratio.

[0045] The second power spectral density corresponding to the initial speech signal is determined by the product of the dynamic gain factor and the third power spectral density.

[0046] In this embodiment, the third power spectral density is the power spectral density of the speech signal (including the initial speech signal and the background noise signal), which is calculated by squaring the complex modulus of the speech audio domain signal, reflecting the distribution of the total energy of the speech and noise superimposed on each frequency. The estimated power spectral density of the initial speech signal is a preliminary approximation of the power spectrum of the clean speech signal, and is calculated by the difference between the third power spectral density and the first power spectral density, i.e.:

[0047] Estimated power spectral density = third power spectral density - first power spectral density.

[0048] In practical applications, the interaction term between the initial speech signal and the background noise signal is ignored by default in the power spectral density estimation calculation. If the detected background noise is non-stationary or strongly correlated with the initial speech signal (e.g., signal-to-noise ratio <5dB or noise containing homologous speech components), the estimated power spectral density value is corrected using the minimum mean square error estimation model. The correction formula is as follows:

[0049]

[0050] in, The cross-term compensation coefficient (1 < ≤1.5), To prevent the power spectral density from reaching a negative minimum.

[0051] The posterior signal-to-noise ratio (SNR) is an indicator that measures the relative strength of the initial speech signal to noise energy at each frequency point. It is calculated based on the ratio of the estimated power spectral density of the initial speech signal to the first power spectral density, using the following formula:

[0052] =Estimated power spectral density / First power spectral density.

[0053] This ratio directly reflects the proportion of pure speech energy to noise energy, and is closer to the real speech quality than the traditional noisy speech signal-to-noise ratio, providing a precise basis for dynamic gain adjustment.

[0054] The dynamic gain factor is a weighting coefficient adaptively optimized based on the posterior signal-to-noise ratio, used to balance noise suppression and speech preservation. The formula for calculating the dynamic gain factor is:

[0055]

[0056] The adjustment coefficient k=1 is set in the formula for calculating the dynamic gain factor. Based on the calculated posterior signal-to-noise ratio (SNR), the gain factor is dynamically adjusted. When the posterior SNR... When the value is greater than or equal to a set threshold (in high signal-to-noise ratio scenarios), i.e., the dynamic gain factor... Represented as: = / ( +1) to enhance noise suppression. When the posterior signal-to-noise ratio When the signal-to-noise ratio is less than the set threshold (in low signal-to-noise ratio scenarios), set the adjustment coefficient k=0.5. dynamic gain factor Represented as: = / ( +0.5 To avoid excessive suppression that could lead to speech distortion.

[0057] Multiplying the dynamic gain factor by the third power spectral density yields the second power spectral density corresponding to the initial speech signal. The calculation formula is: Second power spectral density = Dynamic gain factor × Third power spectral density. By weighting and adjusting the power spectrum of noisy speech using the dynamic gain factor, an accurate estimation of the power spectrum of clean speech can be achieved.

[0058] In this embodiment, the background noise of non-speech signals and speech signals in the same scene has the same origin, and the first power spectral density can accurately characterize the noise characteristics. By initially separating speech and noise energy through the difference method, and combining it with the posterior signal-to-noise ratio to dynamically optimize the gain, the estimation accuracy of the initial speech power spectrum in low signal-to-noise ratio scenes is effectively improved.

[0059] This embodiment calculates the third power spectral density of the speech audio domain signal, combines it with the first power spectral density to obtain the posterior signal-to-noise ratio (SNR), and then dynamically adjusts the gain factor and multiplies it by the third power spectral density to obtain the second power spectral density. This embodiment leverages the similarity of noise characteristics within the same scene and adapts the dynamic gain to different SNR scenarios to improve the accuracy of the initial speech power spectral estimation, providing a reliable basis for subsequent signal enhancement.

[0060] In one embodiment of this disclosure, calculating the power ratio of the initial speech signal to the background noise signal in a speech-type signal includes:

[0061] Calculate the first average energy of the non-speech signal, where the first average energy is the average energy of the background noise signal;

[0062] Calculate the second average energy of the speech signal. The second average energy is the average energy of the initial speech signal and the background noise signal superimposed on each other.

[0063] The estimated average energy of the initial speech signal is obtained based on the difference between the second average energy and the first average energy.

[0064] The ratio of the estimated average energy to the first average energy is calculated to obtain the power ratio of the initial speech signal to the background noise signal. In this embodiment, the first average energy is the average energy of the background noise signal among non-speech signals, obtained by integrating the first power spectral density across the entire frequency band and dividing by the frequency bandwidth, reflecting the average energy level of the background noise signal across the entire frequency band. The second average energy is the average energy of the speech signal, obtained by integrating the third power spectral density (including the superimposed power spectral density of the initial speech and background noise) across the entire frequency band and dividing by the frequency bandwidth, reflecting the average energy level of the speech signal across the entire frequency band.

[0065] In this embodiment, the first power spectral density of the background noise signal in the non-speech signal is integrated across the entire frequency band to obtain the total energy of the non-speech signal; the total energy is divided by the frequency bandwidth of the signal (i.e., the frequency range covered by the signal) to obtain the first average energy.

[0066] The total energy of the speech signal is obtained by integrating the third power spectral density (including the superimposed power spectral density of the initial speech and background noise) across the entire frequency band; the total energy is then divided by the same frequency bandwidth to obtain the second average energy.

[0067] Since the second average energy of the speech signal is the sum of the initial speech signal energy and the background noise signal energy, the estimated average energy of the initial speech signal can be obtained by subtracting the first average energy from the second average energy. Finally, the ratio of this estimated average energy to the first average energy is calculated to obtain the power ratio of the initial speech signal to the background noise signal.

[0068] In this embodiment, the background noise sources in non-speech signals and speech signals in the same scene are consistent and their energy characteristics are highly similar. Therefore, by using this difference calculation method, it is not necessary to directly separate the initial speech and noise. It can accurately reflect the energy ratio relationship between the initial speech and background noise. The calculation is simple and robust, which can effectively improve the accuracy and efficiency of subsequent signal enhancement and other processing based on this power ratio.

[0069] In this embodiment, a Fast Fourier Transform (FFT) is first performed on the speech signal to obtain the speech audio domain signal. The third power spectral density is calculated, and the posterior signal-to-noise ratio (SNR) is obtained by combining it with the first power spectral density. The gain factor is dynamically adjusted and multiplied by the third power spectral density to obtain the second power spectral density. Simultaneously, the first average energy of the non-speech signal and the second average energy of the speech signal are calculated to obtain the power ratio of the initial speech signal to the background noise signal. The calculation is simplified by utilizing the noise characteristics of the same scene, thus improving the estimation accuracy.

[0070] S104: The initial speech signal is enhanced based on the power ratio, the audio domain signal, the first power spectral density, and the second power spectral density to obtain the target speech signal.

[0071] In this embodiment, the target speech signal is a speech signal obtained after signal enhancement processing, which has removed most of the background noise interference and retained clear initial speech information, and is the direct input object for model recognition.

[0072] In this embodiment, the specific process of obtaining the target speech signal includes:

[0073] The noise suppression strength coefficient is determined based on the power ratio;

[0074] Substituting the noise suppression intensity coefficient, the first power spectral density, and the second power spectral density into the first formula, we obtain the frequency domain filter gain; the first formula is:

[0075] in, Indicates the frequency domain filter gain. This represents the second power spectral density corresponding to the initial speech signal. Indicates the noise suppression strength coefficient. This represents the first power spectral density corresponding to the background noise signal in non-speech signals;

[0076] Multiply the frequency domain filter gain by the audio domain signal to obtain an estimate of the enhanced initial speech signal;

[0077] The estimated value of the enhanced initial speech signal is subjected to inverse Fourier transform to obtain the target speech signal.

[0078] In this embodiment, the noise suppression intensity coefficient The power ratio is a dynamically adjusted coefficient used to control the suppression of background noise during the filtering process. The larger the power ratio, the larger the coefficient value, and the stronger the noise suppression effect. The frequency domain filtering gain is a frequency-dependent weight calculated by the first formula, reflecting the proportion of speech signal retained relative to noise at each frequency point, and is used to enhance speech and suppress noise in the frequency domain.

[0079] In this embodiment, the noise suppression intensity coefficient is determined based on the calculated power ratio. The higher the power ratio, The larger the value (e.g., power ratio > 0.5), the better. =1.2, power ratio ≤0.5 =0.8), to adapt to different noise intensity scenarios.

[0080] The first formula is an adaptive frequency domain filtering method based on the power spectral density difference between speech and noise. In the formula, The power spectral density of the pure speech is estimated. For background noise power spectral density, This is the noise suppression strength coefficient. The filter gain is dynamically calculated based on the energy ratio of the two components, at the frequency point where speech energy is dominant (…). The gain is close to 1 to preserve speech; at frequency points with strong noise energy, the gain is reduced to suppress noise. By dynamically adjusting the suppression level through the power ratio, a balance between noise suppression and speech preservation can be achieved.

[0081] Noise suppression strength coefficient First power spectral density Second power spectral density Substitute into the first formula The frequency domain filtering gain at each frequency point is obtained.

[0082] Multiplying the frequency domain filter gain by the audio domain signal yields an estimate of the enhanced initial speech signal. Since the audio domain signal contains noise and... It is an estimate of the clean speech (not the true value). Therefore, the frequency domain filter gain multiplied by the audio domain signal is an estimate of the enhanced initial speech signal. It approximates the true speech through frequency selective weighting, rather than being completely equivalent to the noise-free initial speech.

[0083] An inverse Fourier transform is performed on the estimated value of the enhanced initial speech signal to convert the frequency domain signal back to the time domain, yielding the denoised target speech signal. This signal, after noise suppression, has improved clarity, providing a reliable foundation for subsequent speech recognition.

[0084] This embodiment outputs an estimated value of the enhanced initial speech signal through dynamic power ratio adjustment and frequency domain adaptive filtering. This not only conforms to the actual scenario where noise and speech audio domains overlap and cannot be completely separated, but also balances noise suppression and speech preservation through engineering strategies, ultimately improving speech quality and subsequent recognition robustness.

[0085] S105: Recognize the target speech signal based on the target model.

[0086] In this embodiment, the target model is a pre-trained model used to recognize target speech signals. It has the ability to extract semantic information from speech features and output recognition results. The appropriate network structure or algorithm model can be selected according to the application scenario.

[0087] In this embodiment, feature extraction can be performed on the target speech signal obtained after signal enhancement, extracting 13-dimensional Mel frequency cepstral coefficients, and normalizing the feature sequence to unify the feature scale.

[0088] A pre-trained target recognition model (such as a deep neural network or speech recognition model) is loaded, and the normalized Mel-frequency cepstral coefficient feature sequence is input into the target model. The model analyzes the features layer by layer through a multi-layer network structure to extract speech semantic features. The target model performs inference calculations on the normalized Mel-frequency cepstral coefficient feature sequence and outputs the corresponding recognition result (such as text or instruction category). Finally, the recognition result is fed back to the interaction system of the smart glasses.

[0089] For example, suppose that in a subway commuting scenario, smart glasses need to achieve navigation interaction through voice commands. First, the built-in microphone collects the raw mixed audio signal containing the user's voice, rail vibrations, and conversations among people at a sampling rate of 16kHz. After removing low-frequency vibrations with a 10Hz high-pass filter, the amplitude is normalized to [-1,1] by automatic gain control, and the target audio signal is obtained after frame segmentation.

[0090] A Gaussian mixture model-based speech activity detection algorithm is used to acquire the target audio signal. Signals with short-term energy below a threshold are classified as non-speech signals (containing only ambient noise), while those above the threshold are classified as speech signals (containing user speech plus noise). A 256ms Hanning window is applied to the non-speech signals, followed by a 256-point Fast Fourier Transform. The squared amplitude is divided by a 62.5Hz frequency resolution to obtain the first power spectral density of the background noise.

[0091] Speech signals are processed in the same way to obtain the audio domain signal. The ratio of its third power spectral density to the first power spectral density is calculated as the posterior signal-to-noise ratio (SNR). For high SNR, the dynamic gain factor is taken as... / ( +1), take the value when the signal-to-noise ratio is low. / ( +0.5 Multiplying these two values ​​yields the second power spectral density. The ratio of the first average energy of the non-speech class to the second average energy of the speech class is calculated as the power ratio (set to 0.6), corresponding to the noise suppression coefficient. =1.2.

[0092] Substitute into the formula The frequency domain filter gain is obtained, multiplied with the audio domain signal, and then subjected to inverse fast Fourier transform to obtain the target speech signal. The 13-dimensional Mel frequency cepstral coefficient features are extracted and input into a pre-trained deep residual network model. The deep residual network model outputs the recognition result of "Navigate to East Station", realizing accurate voice interaction in noisy environments.

[0093] As can be seen from the above, this embodiment divides the preprocessed target audio signal into speech signals and non-speech signals. The first power spectral density obtained by Fourier transform of the non-speech signal provides a dynamic and accurate background noise benchmark for noise suppression, overcoming the limitation of fixed filtering lacking a targeted noise reference. Then, Fourier transform is performed on the speech signal, and the second power spectral density of the initial speech signal is calculated in conjunction with the first power spectral density, along with the power ratio between the two. This achieves the distinction between speech and noise energy, avoiding the problem of existing technologies being unable to dynamically identify the proportion of speech and noise energy. Finally, the initial speech signal is enhanced based on the power ratio, the speech domain signal, and the first and second power spectral densities to obtain the target speech signal. This effectively reduces background noise interference and can effectively suppress speech distortion to a certain extent. Furthermore, the target speech signal is recognized through the target model, thereby improving the accuracy of audio recognition in complex scenarios.

[0094] In one embodiment of this application, before recognizing the target speech signal based on the target model, the method further includes:

[0095] In response to the fact that the feature parameters and / or first average energy of the background noise signal in the non-speech signal meet the target condition, the target model for recognizing the target speech signal is determined based on the target condition.

[0096] The feature parameters are calculated based on the first power spectral density and include the bandwidth and energy entropy value of the background noise signal. In this embodiment, the feature parameters are quantitative indicators extracted from the first power spectral density of the background noise. This step includes bandwidth and energy entropy value, which are used to describe the frequency distribution range and energy dispersion of the background noise, providing a basis for model selection. Bandwidth is the frequency coverage range of the background noise signal, determined by the frequency range where the main energy is concentrated in the first power spectral density, reflecting the breadth of the noise's frequency distribution. Energy entropy value is an indicator describing the uniformity of the background noise energy distribution across its frequency components. The lower the entropy value, the more concentrated the energy; the higher the entropy value, the more dispersed the energy distribution. The target condition is the judgment criterion used to trigger model selection. It consists of the feature parameters of the background noise (bandwidth, energy entropy value) and the comparison result of the first average energy with a preset threshold. Different conditions correspond to different noise scenarios.

[0097] In this embodiment, characteristic parameters are calculated based on the first power spectral density of background noise in non-speech signals. The bandwidth can be determined by the frequency range where 80% of the energy in the first power spectral density is present; the energy entropy value can be obtained by calculating the probability distribution entropy value of the first power spectral density at each frequency component. The formula for the energy entropy value is:

[0098]

[0099] in, The energy entropy value. This represents the proportion of the energy of the background noise signal at a certain frequency component f to the total energy.

[0100] The calculated first average energy, bandwidth, and energy entropy of the background noise are compared with preset thresholds. When one or more of these parameters meet the target conditions, the target model selection mechanism is activated.

[0101] Based on the target conditions (such as high energy and high bandwidth, medium energy and low entropy, low energy, etc.), the corresponding recognition model is called from the pre-trained model library for subsequent recognition and processing of the target speech signal.

[0102] In one embodiment of this application, the target condition includes a first condition, a second condition, or a third condition;

[0103] Based on the target conditions, a target model for recognizing the target speech signal is determined, including:

[0104] If the target condition is the first condition, the deep residual network model is determined as the target model for recognizing the target speech signal. The first condition is that the first average energy of the non-speech signal is greater than the first average threshold and the bandwidth is greater than the first bandwidth threshold.

[0105] If the target condition is the second condition, the bidirectional long short-term memory network model is determined as the target model for recognizing the target speech signal. The second condition is that the first average energy of the non-speech signal is less than or equal to the first average threshold and greater than the second average threshold, and the energy entropy value is less than the first energy threshold.

[0106] If the target condition is the third condition, then the lightweight convolutional neural network is determined as the target model for recognizing the target speech signal. The third condition is that the first average energy of the non-speech signal is less than or equal to the second average threshold.

[0107] Background noise exhibits significant differences in energy, bandwidth, and distribution characteristics across various noise scenarios (e.g., high-energy broadband noise in subways versus low-energy narrow-band noise in offices). A single model is insufficient for all scenarios; complex noise scenarios require models with strong anti-interference capabilities, while simple scenarios need to balance efficiency and accuracy. By using sub-model identification, the optimal model can be matched based on noise characteristic parameters (bandwidth, energy entropy value) and / or energy level, avoiding the problems of insufficient accuracy or computational redundancy in some scenarios with a single model.

[0108] In this embodiment, a first average threshold (e.g., -40dB), a second average threshold (e.g., -60dB), a first bandwidth threshold (e.g., 3kHz), and a first energy threshold (e.g., 2.5) are preset as the criteria for classifying noise scenes.

[0109] If the first average energy of a non-speech signal is greater than the first average threshold (-40dB) and the bandwidth is greater than the first bandwidth threshold (3kHz), the first condition is met, corresponding to a high-energy, wide-bandwidth noise scenario (such as a noisy subway environment).

[0110] If the second average threshold (-60dB) < the first average energy ≤ the first average threshold (-40dB) and the energy entropy value < the first energy threshold (2.5), the second condition is satisfied, corresponding to medium energy and concentrated energy noise scenarios (such as air conditioner operation sound).

[0111] If the first average energy is less than or equal to the second average threshold (-60dB), the third condition is met, corresponding to a low-energy noise scenario (such as a quiet office).

[0112] In this embodiment, the deep residual network model alleviates the degradation problem of deep networks through residual connections, possesses strong noise resistance, and is suitable for complex broadband noise scenarios. The bidirectional long short-term memory network model excels at capturing temporal features and has a good suppression effect on energy-concentrated periodic noise, making it suitable for medium-energy noise scenarios. The lightweight convolutional neural network reduces computation by simplifying the network structure, and can quickly output recognition results in low-noise scenarios, balancing accuracy and efficiency.

[0113] Based on the matching target conditions, the corresponding model is called from the model library: the first condition loads a deep residual network model, the second condition loads a bidirectional long short-term memory network model, and the third condition loads a lightweight convolutional neural network model for target speech signal recognition.

[0114] For example, in a subway scenario during morning rush hour, smart glasses need to recognize user voice commands. First, the device collects mixed audio including track vibrations and conversations among people. After preprocessing, a non-speech signal (containing only ambient noise) is obtained. Its first average energy is calculated to be -35dB (>-40dB first average threshold). Through first power spectral density analysis, the frequency range with 80% energy is 200Hz-4kHz (bandwidth 4kHz > 3kHz first bandwidth threshold), satisfying the first condition. At this point, a deep residual network model is automatically loaded. This model enhances feature extraction capabilities through residual connections and can effectively resist broadband noise interference.

[0115] In an office setting, if the first average energy of non-speech signals is -50dB (between -40dB and -60dB) and the energy entropy value is calculated to be 2.0 (<2.5 of the first energy threshold), satisfying the second condition, the system switches to a bidirectional long short-term memory network model to utilize its temporal modeling advantages to suppress the periodic noise of air conditioner operation.

[0116] In a quiet library scenario, the first average energy is -65dB (≤-60dB), triggering the third condition. A lightweight convolutional neural network is invoked to achieve fast and accurate recognition with low computational cost.

[0117] In this embodiment, if only feature parameters are considered, the target model for recognizing the target speech signal is determined based on the target conditions, in response to the feature parameters corresponding to the background noise signal in the non-speech signal satisfying the target conditions.

[0118] By pre-setting thresholds for feature parameters, the target conditions are also divided into three categories. Different conditions correspond to different noise scenarios, and each is matched with a suitable recognition model, as follows:

[0119]

[0120] The logic for model selection is as follows:

[0121] The first condition corresponds to wideband and dispersed noise, with interference covering the main frequency bands of speech. A deep residual network model is needed to enhance feature extraction capabilities through residual connections to resist wideband interference. The second condition corresponds to medium-bandwidth and moderately concentrated noise, with a moderate interference range and energy distribution. Lightweight convolutional neural networks can simplify computation while taking into account noise suppression and recognition efficiency. The third condition corresponds to narrowband and highly concentrated noise, which is mostly periodic signals. Bidirectional long short-term memory network models are good at capturing temporal features and can accurately distinguish the temporal differences between periodic noise and speech, reducing interference misjudgments.

[0122] For example, assuming an office air-conditioned environment, non-speech signals (containing only air conditioner operating sounds) are collected, and the first power spectral density is calculated. Feature parameters are calculated: the frequency range of 80% of the cumulative energy is 500Hz~1.2kHz (bandwidth = 0.7kHz ≤ 1kHz), and the energy entropy value H = 1.8 ≤ 2.0, satisfying the third condition. A bidirectional long short-term memory network model is loaded, utilizing its temporal modeling capabilities to distinguish between the periodic noise of the air conditioner and the temporal characteristics of the user's email voice commands, avoiding misidentification of commands caused by noise interference. From the above, it can be concluded that this embodiment adapts to different noise scenarios by defining the first, second, and third conditions, dynamically selecting a deep residual network, a bidirectional long short-term memory network, or a lightweight convolutional neural network. High-energy broadband noise uses a robust model to ensure accuracy, medium-energy concentrated noise uses a temporal modeling model to suppress interference, and low-energy noise uses a lightweight model to improve efficiency. This method achieves accurate matching between the model and the scenario, optimizing resource consumption while ensuring recognition performance in complex environments, and improving the adaptability and practicality of audio recognition.

[0123] In one embodiment of this application, after recognizing the target speech signal based on the target model, the method further includes:

[0124] Obtain the recognition confidence score corresponding to the recognition result output by the target model;

[0125] In response to a recognition confidence level greater than or equal to a preset confidence threshold, the recognition result output by the target model is used as the target speech signal;

[0126] In response to a recognition confidence level less than a preset confidence threshold, fuzzy segments in the target speech signal with a matching degree less than a preset matching degree with the recognition result are extracted. The speech feature frequency band of the segment is determined based on the second power spectral density corresponding to the fuzzy segment. The amplitude of the target speech signal corresponding to the speech feature frequency band is enhanced to generate a new recognition result. The new recognition result output by the target model is used as the target speech signal.

[0127] In this embodiment, the recognition confidence score is a quantitative evaluation value of the reliability of the recognition result by the target model. It can be calculated based on the output probability distribution and has a value range of [0,1]. The higher the value, the more reliable the result. The preset confidence threshold is a critical value used to determine whether the recognition result can be directly used. It is set according to the accuracy requirements of the actual application scenario to balance recognition efficiency and accuracy. Fuzzy segments are the parts of the target speech signal with low matching degree with the recognition result and need further optimization processing. The speech feature frequency band is the frequency range (usually 300-3400Hz) in the speech signal that carries the main semantic information. It is determined based on the energy distribution of the second power spectral density and is the key area for enhancement processing. For example, in a noisy vegetable market scenario, the preset confidence threshold is set to 0.75. When the target model recognizes the user saying "weigh two jin of tomatoes", the output recognition confidence score is 0.7, which is less than the threshold. At this time, the comparison shows that the speech segment corresponding to the word "tomato" has a matching degree of only 0.6 with the recognition result (less than the preset matching degree of 0.65), which belongs to the fuzzy segment. The second power spectral density of the segment was extracted, and analysis revealed that the energy was concentrated in the 800-2000Hz range (this is the speech feature frequency band of the blurred segment, carrying the main semantic information of "tomato"). The amplitude of the target speech signal corresponding to the 800-2000Hz frequency band was multiplied by a gain coefficient of 1.3 to enhance it, and then input into the model for recognition. The confidence score of the new result "tomato" was increased to 0.85, which was then used as the final recognition result.

[0128] In this embodiment, after the target model completes the recognition of the target speech signal, the recognition confidence score attached to the model output is extracted. This confidence score is a quantitative assessment of the reliability of the recognition result by the model, and is usually calculated based on the probability distribution of the model output, for example, determined by the maximum probability value output by the softmax function.

[0129] The acquired recognition confidence score is compared with a preset confidence threshold (e.g., 0.75). If the recognition confidence score is greater than or equal to the preset confidence threshold, it indicates that the model's recognition result is highly reliable, and this recognition result is directly output as the final target speech signal recognition result. If the recognition confidence score is less than the preset confidence threshold, it indicates that the recognition result is not reliable enough. In this case, by comparing the target speech signal with the recognition result output by the model, signal segments with a matching degree less than a preset matching degree (e.g., 0.65) are extracted. These segments are considered ambiguous segments, usually those that are difficult to recognize due to noise interference or unclear speech features.

[0130] For the extracted blurry segment, the corresponding second power spectral density data is retrieved. By analyzing the energy distribution of the second power spectral density, the frequency band where the energy is concentrated is located. This frequency band is the speech feature frequency band of the blurry segment (for example, the main feature frequency band of human speech is usually in the range of 300Hz-3400Hz).

[0131] The amplitude of the target speech signal corresponding to the determined speech feature frequency band is enhanced (e.g., multiplied by a gain coefficient of 1.2-1.5) to highlight the speech features. The enhanced target speech signal is then re-input into the target model for recognition, generating a new recognition result. Finally, this new recognition result is output as the recognition result of the target speech signal.

[0132] In this embodiment, despite signal enhancement and model recognition, the confidence level of the recognition result may still be insufficient due to residual noise or speech distortion in complex noise environments. This embodiment introduces a confidence verification mechanism to perform secondary optimization on low-confidence results, avoiding direct output of erroneous results; at the same time, it accurately enhances the speech feature frequency bands for ambiguous segments, further improving recognition accuracy and ensuring the reliability of audio recognition in dynamic noise scenarios.

[0133] As can be seen from the above, this embodiment introduces a recognition confidence verification mechanism. When the recognition result is reliable (confidence ≥ threshold), it outputs directly to ensure efficiency. When the result is unreliable (confidence < threshold), it accurately extracts ambiguous segments, locates the speech feature frequency bands based on their second power spectral density, enhances the amplitude, and then re-recognizes. This method compensates for recognition defects under complex noise through secondary optimization, reduces erroneous outputs, improves the accuracy of speech recognition, and ensures the reliability of audio interaction in dynamic scenes.

[0134] Corresponding to the audio recognition method in the above embodiments, Figure 2 This is a structural block diagram of an audio recognition device provided according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2The audio recognition device 20 includes: an audio acquisition module 21, a first calculation module 22, a second calculation module 23, a signal enhancement module 24, and a signal recognition module 25.

[0135] The audio acquisition module 21 is used to acquire speech signals and non-speech signals in the target audio signal; the target audio signal is the audio signal after preprocessing the original mixed audio signal; the speech signals include the user's initial speech signal and background noise signal, and the non-speech signals include background noise signal;

[0136] The first calculation module 22 is used to perform Fourier transform on the non-speech signal to obtain the noise frequency domain signal corresponding to the non-speech signal, and to calculate the first power spectral density of the background noise signal in the non-speech signal.

[0137] The second calculation module 23 is used to perform Fourier transform on the speech signal to obtain the corresponding audio domain signal, and calculate the second power spectral density corresponding to the initial speech signal based on the audio domain signal and the first power spectral density; and calculate the power ratio of the initial speech signal to the background noise signal in the speech signal.

[0138] Signal enhancement module 24 is used to enhance the initial speech signal based on the power ratio, speech domain signal, first power spectral density and second power spectral density to obtain the target speech signal;

[0139] The signal recognition module 25 is used to recognize the target speech signal based on the target model.

[0140] In one embodiment of this application, the second calculation module 23 is specifically used for:

[0141] Calculate the third power spectral density corresponding to the audio domain signal, and obtain the estimated power spectral density of the initial speech signal based on the difference between the third power spectral density and the first power spectral density.

[0142] The posterior signal-to-noise ratio is calculated based on the ratio of the estimated power spectral density to the first power spectral density of the initial speech signal.

[0143] The dynamic gain factor of the initial speech signal is determined based on the posterior signal-to-noise ratio.

[0144] The second power spectral density corresponding to the initial speech signal is determined by the product of the dynamic gain factor and the third power spectral density.

[0145] In one embodiment of this application, the second calculation module 23 is further configured to:

[0146] Calculate the first average energy of the non-speech signal, where the first average energy is the average energy of the background noise signal;

[0147] Calculate the second average energy of the speech signal. The second average energy is the average energy of the initial speech signal and the background noise signal superimposed on each other.

[0148] The estimated average energy of the initial speech signal is obtained based on the difference between the second average energy and the first average energy.

[0149] The ratio of the estimated average energy to the first average energy is calculated to obtain the power ratio of the initial speech signal to the background noise signal.

[0150] In one embodiment of this application, the audio recognition device 20 further includes: a model selection module, specifically used for:

[0151] In response to the fact that the feature parameters and / or first average energy of the background noise signal in the non-speech signal meet the target condition, the target model for recognizing the target speech signal is determined based on the target condition.

[0152] The characteristic parameters are calculated based on the first power spectral density and include the bandwidth and energy entropy value of the background noise signal.

[0153] In one embodiment of this application, the target condition includes a first condition, a second condition, or a third condition; the model selection module is further configured to:

[0154] If the target condition is the first condition, the deep residual network model is determined as the target model for recognizing the target speech signal. The first condition is that the first average energy of the non-speech signal is greater than the first average threshold and the bandwidth is greater than the first bandwidth threshold.

[0155] If the target condition is the second condition, the bidirectional long short-term memory network model is determined as the target model for recognizing the target speech signal. The second condition is that the first average energy of the non-speech signal is less than or equal to the first average threshold and greater than the second average threshold, and the energy entropy value is less than the first energy threshold.

[0156] If the target condition is the third condition, then the lightweight convolutional neural network is determined as the target model for recognizing the target speech signal. The third condition is that the first average energy of the non-speech signal is less than or equal to the second average threshold.

[0157] In one embodiment of this application, the signal enhancement module 24 is specifically used for:

[0158] The noise suppression strength coefficient is determined based on the power ratio;

[0159] Substituting the noise suppression intensity coefficient, the first power spectral density, and the second power spectral density into the first formula, we obtain the frequency domain filter gain; the first formula is:

[0160] in, Indicates the frequency domain filter gain. This represents the second power spectral density corresponding to the initial speech signal. Indicates the noise suppression strength coefficient. This represents the first power spectral density corresponding to the background noise signal in non-speech signals;

[0161] Multiply the frequency domain filter gain by the audio domain signal to obtain an estimate of the enhanced initial speech signal;

[0162] The estimated value of the enhanced initial speech signal is subjected to inverse Fourier transform to obtain the target speech signal.

[0163] In one embodiment of this application, the audio recognition device 20 further includes an optimization module, specifically used for:

[0164] Obtain the recognition confidence score corresponding to the recognition result output by the target model;

[0165] In response to a recognition confidence level greater than or equal to a preset confidence threshold, the recognition result output by the target model is used as the target speech signal;

[0166] In response to a recognition confidence level less than a preset confidence threshold, fuzzy segments in the target speech signal with a matching degree less than a preset matching degree with the recognition result are extracted. The speech feature frequency band of the segment is determined based on the second power spectral density corresponding to the fuzzy segment. The amplitude of the target speech signal corresponding to the speech feature frequency band is enhanced to generate a new recognition result. The new recognition result is used as the target speech signal.

[0167] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the audio acquisition module 21, the first calculation module 22, the second calculation module 23, the signal enhancement module 24, and the signal recognition module 25 are shown.

[0168] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0169] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0170] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store information such as target conditions and target models.

[0171] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the audio recognition method provided in the embodiments of this application, or they can execute the implementation method of the electronic device described in the embodiments of this application, which will not be repeated here.

[0172] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0173] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0174] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0175] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0176] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces or units, or they may be electrical, mechanical, or other forms of connection.

[0177] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0178] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0179] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An audio recognition method, characterized in that, include: Acquire speech signals and non-speech signals from a target audio signal; the target audio signal is an audio signal obtained after preprocessing the original mixed audio signal; the speech signals include the user's initial speech signal and background noise signal, and the non-speech signals include background noise signal; Perform a Fourier transform on the non-speech signal to obtain the noise frequency domain signal corresponding to the non-speech signal, and calculate the first power spectral density of the background noise signal in the non-speech signal. Perform a Fourier transform on the speech signal to obtain the corresponding audio domain signal, and calculate the second power spectral density corresponding to the initial speech signal based on the audio domain signal and the first power spectral density. Calculate the power ratio of the initial speech signal to the background noise signal in the speech signal class; Based on the power ratio, the audio domain signal, the first power spectral density, and the second power spectral density, the initial speech signal is enhanced to obtain the target speech signal; The step of enhancing the initial speech signal based on the power ratio, the audio domain signal, the first power spectral density, and the second power spectral density to obtain the target speech signal includes: The noise suppression strength coefficient is determined based on the power ratio; Substituting the noise suppression intensity coefficient, the first power spectral density, and the second power spectral density into the first formula, we obtain the frequency domain filtering gain; the first formula is: in, Indicates the frequency domain filter gain. This represents the second power spectral density corresponding to the initial speech signal. Indicates the noise suppression strength coefficient. This represents the first power spectral density corresponding to the background noise signal in the non-speech signal; Multiply the frequency domain filter gain by the audio domain signal to obtain an estimate of the enhanced initial speech signal; The estimated value of the enhanced initial speech signal is subjected to an inverse Fourier transform to obtain the target speech signal; The target speech signal is identified based on the target model.

2. The audio recognition method as described in claim 1, characterized in that, The step of calculating the second power spectral density corresponding to the initial speech signal based on the speech frequency domain signal and the first power spectral density includes: Calculate the third power spectral density corresponding to the audio domain signal, and obtain the estimated power spectral density of the initial speech signal based on the difference between the third power spectral density and the first power spectral density. The posterior signal-to-noise ratio is calculated based on the ratio of the estimated power spectral density of the initial speech signal to the first power spectral density. The dynamic gain factor of the initial speech signal is determined based on the posterior signal-to-noise ratio. The second power spectral density corresponding to the initial speech signal is determined based on the product of the dynamic gain factor and the third power spectral density.

3. The audio recognition method as described in claim 1, characterized in that, The calculation of the power ratio of the initial speech signal to the background noise signal in the speech signal class includes: Calculate the first average energy of the non-speech signal, where the first average energy is the average energy of the background noise signal; Calculate the second average energy of the speech signal, where the second average energy is the average energy of the initial speech signal and the background noise signal superimposed on each other. The estimated average energy of the initial speech signal is obtained based on the difference between the second average energy and the first average energy. The ratio of the estimated average energy to the first average energy is calculated to obtain the power ratio of the initial speech signal to the background noise signal.

4. The audio recognition method as described in claim 3, characterized in that, Before recognizing the target speech signal based on the target model, the method further includes: In response to the fact that the feature parameters corresponding to the background noise signal in the non-speech signal and / or the first average energy satisfy the target condition, a target model for recognizing the target speech signal is determined based on the target condition. The characteristic parameters are calculated based on the first power spectral density, and the characteristic parameters include the bandwidth and energy entropy value of the background noise signal.

5. The audio recognition method as described in claim 4, characterized in that, The target conditions include a first condition, a second condition, or a third condition; The step of determining the target model for recognizing the target speech signal based on the target conditions includes: If the target condition is the first condition, the deep residual network model is determined as the target model for recognizing the target speech signal. The first condition is that the first average energy of the non-speech signal is greater than the first average threshold and the bandwidth is greater than the first bandwidth threshold. If the target condition is the second condition, the bidirectional long short-term memory network model is determined as the target model for recognizing the target speech signal. The second condition is that the first average energy of the non-speech signal is less than or equal to the first average threshold and greater than the second average threshold, and the energy entropy value is less than the first energy threshold. If the target condition is the third condition, a lightweight convolutional neural network is determined as the target model for recognizing the target speech signal. The third condition is that the first average energy of the non-speech signal is less than or equal to the second average threshold.

6. The audio recognition method as described in claim 1, characterized in that, After recognizing the target speech signal based on the target model, the method further includes: Obtain the recognition confidence score corresponding to the recognition result output by the target model; In response to the recognition confidence level being greater than or equal to a preset confidence threshold, the recognition result output by the target model is taken as the target speech signal; In response to the recognition confidence level being less than the preset confidence threshold, fuzzy segments in the target speech signal whose matching degree with the recognition result is less than the preset matching degree are extracted. The speech feature frequency band of the segment is determined based on the second power spectral density corresponding to the fuzzy segment. The amplitude of the target speech signal corresponding to the speech feature frequency band is enhanced to generate a new recognition result. The new recognition result is used as the target speech signal.

7. An audio recognition device, characterized in that, include: An audio acquisition module is used to acquire speech signals and non-speech signals from a target audio signal; the target audio signal is an audio signal after preprocessing the original mixed audio signal; the speech signals include the user's initial speech signal and background noise signal, and the non-speech signals include background noise signal; The first calculation module is used to perform Fourier transform on the non-speech signal to obtain the noise frequency domain signal corresponding to the non-speech signal, and to calculate the first power spectral density of the background noise signal in the non-speech signal. The second calculation module is used to perform Fourier transform on the speech signal to obtain the speech domain signal corresponding to the speech signal, and calculate the second power spectral density corresponding to the initial speech signal based on the speech domain signal and the first power spectral density. Calculate the power ratio of the initial speech signal to the background noise signal in the speech signal class; The signal enhancement module is used to enhance the initial speech signal based on the power ratio, the speech domain signal, the first power spectral density, and the second power spectral density to obtain the target speech signal; The signal enhancement module is specifically used for: The step of enhancing the initial speech signal based on the power ratio, the speech domain signal, the first power spectral density, and the second power spectral density to obtain the target speech signal includes: The noise suppression strength coefficient is determined based on the power ratio; Substituting the noise suppression intensity coefficient, the first power spectral density, and the second power spectral density into the first formula, we obtain the frequency domain filtering gain; the first formula is: in, Indicates the frequency domain filter gain. This represents the second power spectral density corresponding to the initial speech signal. Indicates the noise suppression strength coefficient. This represents the first power spectral density corresponding to the background noise signal in the non-speech signal; Multiply the frequency domain filter gain by the audio domain signal to obtain an estimate of the enhanced initial speech signal; The estimated value of the enhanced initial speech signal is subjected to an inverse Fourier transform to obtain the target speech signal; The signal recognition module is used to recognize the target speech signal based on the target model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice activity detection method for earphone, earphone and storage medium

    CN112017696A

  • Speech enhancement method and device, electronic equipment and storage medium

    CN113299308A