Voice noise reduction method and device, storage medium and electronic equipment
Patent Information
- Application Number
- CN202610727707.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-28
AI Technical Summary
然而,此类深度学习模型通常参数量大(往往在数MB以上)、单帧推理计算量高达数百MIPS,难以直接部署在低算力、低功耗的嵌入式芯片上
[0062]The speech denoising method, apparatus, storage medium, and electronic device provided in this application include: acquiring a frequency-domain noisy frequency signal corresponding to an original audio signal acquired by an audio acquisition component; performing a first denoising process on the frequency-domain noisy frequency signal to obtain a first audio signal, wherein the residual noise contained in the first audio signal is dominated by non-steady-state noise; extracting time-frequency features from the first audio signal to obtain acoustic features corresponding to the first audio signal, and inputting the acoustic features into a pre-trained lightweight denoising model; inferring from the acoustic features through the lightweight denoising model to output a noise suppression factor; wherein the lightweight denoising model is trained based on a hybrid noisy speech library, the hybrid noisy speech library containing at least non-noisy speech and noisy speech after the first denoising process; performing a second denoising process on the first audio signal according to the noise suppression factor to obtain a second audio signal; and reconstructing a time-domain denoised speech signal corresponding to the original audio signal based on the second audio signal. After acquiring the frequency-domain noisy signal, this application first filters out most of the steady-state noise and some non-steady-state noise through a first-stage noise reduction process, making the residual noise in the first audio signal predominantly non-steady-state noise, thus significantly reducing the processing pressure on the subsequent AI model. Secondly, a second-stage noise reduction process further suppresses residual noise, especially specifically suppressing residual non-steady-state noise, eliminating the need for the AI model to participate in the entire noise reduction process and significantly reducing the overall computational load. Furthermore, the AI model's input uses low-dimensional acoustic features extracted from its features, further compressing the model's parameters to adapt to the limited computing resources of embedded devices. Simultaneously, training is performed using a specially constructed hybrid noisy speech library to ensure the AI model's accuracy in suppressing residual noise after the first-stage noise reduction. Through this two-stage cascaded noise reduction architecture, this application can achieve efficient speech noise reduction in complex noisy environments on resource-constrained, low-computing-power embedded devices, ensuring the clarity and naturalness of the output speech, effectively solving the technical challenge of balancing lightweight deployment and high-performance noise reduction in related technologies.
Smart Images

Figure CN122658337A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio signal processing technology, and in particular to a speech noise reduction method, apparatus, storage medium and electronic device. Background Technology
[0002] With the widespread adoption of smart terminal devices, users' demands for voice communication quality are increasing, especially for clear and natural voice interaction in noisy environments, which has become a core user requirement. Therefore, voice noise reduction technology has become a key element in enhancing the competitiveness of smart terminal products.
[0003] In related technologies, end-to-end speech denoising solutions based on deep learning are mainly adopted. Specifically, a deep learning model is constructed to learn the time-frequency feature mapping relationship between noisy speech signals and clean speech signals; then, the deep learning model is used to denoise the input noisy frequency signal and output a clean speech signal. However, such deep learning models typically have a large number of parameters (often exceeding several MB) and a single-frame inference computation of up to hundreds of MIPS, making it difficult to deploy directly on low-computing-power, low-power embedded chips. Although model pruning, quantization, and structural simplification can be used to compress the model to reduce resource consumption, excessive compression can lead to a significant decrease in denoising performance, thereby affecting the clarity and intelligibility of the speech and failing to meet users' needs for high-quality voice interaction.
[0004] Therefore, there is an urgent need for a lightweight voice noise reduction solution that can be adapted to resource-constrained embedded devices while also having noise reduction performance. Summary of the Invention
[0005] This application provides a speech noise reduction method, apparatus, storage medium, and electronic device to achieve lightweight deployment, adaptability to resource-constrained embedded devices, and good noise reduction performance under low computing power conditions, thereby obtaining clear and natural speech output effects.
[0006] In a first aspect, this application provides a speech noise reduction method, including:
[0007] Obtain the frequency domain noise frequency signal corresponding to the original audio signal acquired by the audio acquisition component;
[0008] The first noise reduction process is performed on the frequency domain band noise signal to obtain the first audio signal. The residual noise contained in the first audio signal is dominated by non-steady-state noise.
[0009] The time-frequency features of the first audio signal are extracted to obtain the acoustic features corresponding to the first audio signal. The acoustic features are then input into a pre-trained lightweight noise reduction model. The lightweight noise reduction model infers the acoustic features and outputs a noise suppression factor. The lightweight noise reduction model is trained based on a hybrid noisy speech library. The hybrid noisy speech library contains at least non-noisy speech and noisy speech after the first noise reduction process.
[0010] Based on the noise suppression factor, the first audio signal is subjected to a second noise reduction process to obtain a second audio signal;
[0011] Based on the second audio signal, the time-domain denoised speech signal corresponding to the original audio signal is reconstructed.
[0012] In one possible implementation, the first audio signal is subjected to a second noise reduction process based on a noise suppression factor to obtain a second audio signal, including:
[0013] Multiply the noise suppression factor by the amplitude spectrum of the first audio signal to obtain the noise-reduced amplitude spectrum;
[0014] The second audio signal is determined based on the amplitude spectrum after noise reduction.
[0015] In one possible implementation, determining the second audio signal based on the denoised amplitude spectrum includes:
[0016] Based on the amplitude spectrum after noise reduction, determine the corresponding energy entropy ratio, which represents the probability of speech existence;
[0017] If the entropy ratio is less than the entropy ratio threshold and the duration is greater than or equal to the preset duration, it is determined that there is no speech activity. The amplitude spectrum after noise reduction is multiplied by the non-speech segment noise suppression factor to obtain the second audio signal. The non-speech segment suppression factor is used to suppress the residual noise of the non-speech segment.
[0018] In one possible implementation, the speech noise reduction method further includes:
[0019] If the entropy ratio is greater than or equal to the entropy ratio threshold, or the duration is less than the preset duration, then the amplitude spectrum after noise reduction is determined as the second audio signal.
[0020] In one possible implementation, the speech noise reduction method is applied to an electronic device, which is provided with a first audio acquisition component and a second audio acquisition component.
[0021] Acquire the frequency domain noise frequency signal corresponding to the raw audio signal acquired by the audio acquisition component, including:
[0022] Acquire the first frequency domain signal corresponding to the original audio signal acquired by the first audio acquisition component, and the second frequency domain signal corresponding to the original audio signal acquired by the second audio acquisition component;
[0023] The first frequency domain signal and the second frequency domain signal are processed by a first-order differential array to obtain a differential array signal;
[0024] Based on the spatial difference characteristics between the first frequency domain signal and the second frequency domain signal, it is determined whether wind noise exists in the frequency domain band noise signal;
[0025] If wind noise is confirmed to exist, the corresponding wind noise suppression factor is determined based on the spatial difference characteristics.
[0026] Based on the wind noise suppression factor, the differential array signal is subjected to wind noise suppression processing to obtain the signal after wind noise suppression;
[0027] The signal after wind noise suppression is used as the frequency domain noise frequency signal.
[0028] In one possible implementation, the speech noise reduction method further includes:
[0029] If it is determined that there is no wind noise, then the differential array signal is used as the frequency domain band noise signal.
[0030] In one possible implementation, the hybrid noisy speech library further includes noisy speech obtained after wind noise suppression processing.
[0031] In one possible implementation, both the first frequency domain signal and the second frequency domain signal are determined in the following manner:
[0032] Obtain the raw audio signal acquired by the corresponding audio acquisition component;
[0033] The original audio signal is subjected to front-end noise reduction processing to obtain a third audio signal. The front-end noise reduction processing includes anti-aliasing filtering and / or oversampling processing.
[0034] The fourth audio signal is obtained by performing a short-time Fourier transform on the third audio signal;
[0035] Linear echo cancellation processing is performed on the fourth audio signal to obtain the frequency domain signal corresponding to the original audio signal acquired by the corresponding audio acquisition component. The linear echo cancellation processing is implemented using an adaptive filtering algorithm to estimate and cancel the linear echo components in the fourth audio signal.
[0036] In one possible implementation, based on the second audio signal, a time-domain denoised speech signal corresponding to the original audio signal is reconstructed, including:
[0037] The amplitude spectrum corresponding to the second audio signal is combined with the phase spectrum corresponding to the frequency domain noise frequency signal to obtain the frequency domain denoised speech signal corresponding to the original audio signal.
[0038] The frequency-domain denoised speech signal is transformed by time-frequency transformation to obtain the time-domain denoised speech signal.
[0039] In one possible implementation, a first noise reduction process is performed on the frequency-domain noisy signal to obtain a first audio signal, including:
[0040] An improved Minima-Controlled Recursive Averaging (MCRA) noise estimation method is used to estimate the noise power spectrum of the frequency domain band-noise signal. The improved MCRA noise estimation uses a recursive method to continuously update the minimum search window instead of resetting the window every frame, and uses a two-stage recursive averaging method to estimate the speech presence probability to support a smaller search window length.
[0041] Based on the noise power spectrum, the frequency domain band noise signal is denoised using the Optimal Modified Log-Spectral Amplitude (OM-LSA) estimator or Wiener filtering to obtain the first audio signal.
[0042] Secondly, this application provides a speech noise reduction device, comprising:
[0043] The acquisition module is used to acquire the frequency domain noise frequency signal corresponding to the original audio signal acquired by the audio acquisition component;
[0044] The first noise reduction module is used to perform a first noise reduction process on the frequency domain band noise signal to obtain a first audio signal. The residual noise contained in the first audio signal is dominated by non-steady-state noise.
[0045] The inference module is used to extract time-frequency features from the first audio signal to obtain the acoustic features corresponding to the first audio signal, and input the acoustic features into the pre-trained lightweight noise reduction model. The lightweight noise reduction model infers the acoustic features and outputs a noise suppression factor. The lightweight noise reduction model is trained based on a hybrid noisy speech library, which contains at least non-noisy speech and noisy speech after the first noise reduction process.
[0046] The second noise reduction module is used to perform a second noise reduction process on the first audio signal according to the noise suppression factor to obtain a second audio signal;
[0047] The reconstruction module is used to reconstruct the time-domain denoised speech signal corresponding to the original audio signal based on the second audio signal.
[0048] In one possible implementation, the second noise reduction module is specifically used to: multiply the noise suppression factor by the amplitude spectrum of the first audio signal to obtain the noise-reduced amplitude spectrum; and determine the second audio signal based on the noise-reduced amplitude spectrum.
[0049] In one possible implementation, the second noise reduction module is further configured to: determine the corresponding energy entropy ratio based on the noise-reduced amplitude spectrum, wherein the energy entropy ratio represents the probability of speech presence; if the energy entropy ratio is less than the energy entropy ratio threshold and the duration is greater than or equal to a preset duration, then it is determined that there is no speech activity, and the noise-reduced amplitude spectrum is multiplied by a non-speech segment noise suppression factor to obtain a second audio signal, wherein the non-speech segment suppression factor is used to suppress residual noise in the non-speech segment.
[0050] In one possible implementation, the second noise reduction module is further configured to: determine the noise-reduced amplitude spectrum as the second audio signal if the energy entropy ratio is greater than or equal to the energy entropy ratio threshold, or the duration is less than a preset duration.
[0051] In one possible implementation, the speech noise reduction method is applied to an electronic device, which includes a first audio acquisition component and a second audio acquisition component. The acquisition module is specifically configured to: acquire a first frequency domain signal corresponding to the original audio signal acquired by the first audio acquisition component, and a second frequency domain signal corresponding to the original audio signal acquired by the second audio acquisition component; perform first-order differential array processing on the first and second frequency domain signals to obtain a differential array signal; determine whether wind noise exists in the frequency domain band noise signal based on the spatial difference characteristics between the first and second frequency domain signals; if wind noise is determined to exist, determine the corresponding wind noise suppression factor based on the spatial difference characteristics; perform wind noise suppression processing on the differential array signal based on the wind noise suppression factor to obtain a wind noise-suppressed signal; and use the wind noise-suppressed signal as the frequency domain band noise signal.
[0052] In one possible implementation, the acquisition module is further configured to: if it is determined that there is no wind noise, treat the differential array signal as a frequency domain band-noise signal.
[0053] In one possible implementation, the hybrid noisy speech library further includes noisy speech obtained after wind noise suppression processing.
[0054] In one possible implementation, both the first frequency domain signal and the second frequency domain signal are determined by: acquiring the original audio signal acquired by the corresponding audio acquisition component; performing front-end noise reduction processing on the original audio signal to obtain the third audio signal, wherein the front-end noise reduction processing includes anti-aliasing filtering and / or oversampling processing; performing a short-time Fourier transform on the third audio signal to obtain the fourth audio signal; and performing linear echo cancellation processing on the fourth audio signal to obtain the frequency domain signal corresponding to the original audio signal acquired by the corresponding audio acquisition component, wherein the linear echo cancellation processing is implemented using an adaptive filtering algorithm to estimate and cancel the linear echo components in the fourth audio signal.
[0055] In one possible implementation, the reconstruction module is specifically used to: combine the amplitude spectrum corresponding to the second audio signal with the phase spectrum corresponding to the frequency domain noise frequency signal to obtain the frequency domain denoised speech signal corresponding to the original audio signal; and perform time-frequency transformation on the frequency domain denoised speech signal to obtain the time domain denoised speech signal.
[0056] In one possible implementation, the first noise reduction module is specifically used to: perform improved MCRA noise estimation on the frequency domain band noise signal to obtain a noise power spectrum; wherein, the improved MCRA noise estimation uses a recursive method to continuously update the minimum search window instead of resetting the window every frame, and uses a two-stage recursive average estimation of the speech presence probability to support a smaller search window length; based on the noise power spectrum, OM-LSA or Wiener filtering is used to perform noise reduction processing on the frequency domain band noise signal to obtain a first audio signal.
[0057] Thirdly, this application provides an electronic device, including: a memory and a processor;
[0058] The memory stores instructions that the computer executes;
[0059] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0060] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible embodiments of the first aspect.
[0061] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0062] The speech denoising method, apparatus, storage medium, and electronic device provided in this application include: acquiring a frequency-domain noisy frequency signal corresponding to an original audio signal acquired by an audio acquisition component; performing a first denoising process on the frequency-domain noisy frequency signal to obtain a first audio signal, wherein the residual noise contained in the first audio signal is dominated by non-steady-state noise; extracting time-frequency features from the first audio signal to obtain acoustic features corresponding to the first audio signal, and inputting the acoustic features into a pre-trained lightweight denoising model; inferring from the acoustic features through the lightweight denoising model to output a noise suppression factor; wherein the lightweight denoising model is trained based on a hybrid noisy speech library, the hybrid noisy speech library containing at least non-noisy speech and noisy speech after the first denoising process; performing a second denoising process on the first audio signal according to the noise suppression factor to obtain a second audio signal; and reconstructing a time-domain denoised speech signal corresponding to the original audio signal based on the second audio signal. After acquiring the frequency-domain noisy signal, this application first filters out most of the steady-state noise and some non-steady-state noise through a first-stage noise reduction process, making the residual noise in the first audio signal predominantly non-steady-state noise, thus significantly reducing the processing pressure on the subsequent AI model. Secondly, a second-stage noise reduction process further suppresses residual noise, especially specifically suppressing residual non-steady-state noise, eliminating the need for the AI model to participate in the entire noise reduction process and significantly reducing the overall computational load. Furthermore, the AI model's input uses low-dimensional acoustic features extracted from its features, further compressing the model's parameters to adapt to the limited computing resources of embedded devices. Simultaneously, training is performed using a specially constructed hybrid noisy speech library to ensure the AI model's accuracy in suppressing residual noise after the first-stage noise reduction. Through this two-stage cascaded noise reduction architecture, this application can achieve efficient speech noise reduction in complex noisy environments on resource-constrained, low-computing-power embedded devices, ensuring the clarity and naturalness of the output speech, effectively solving the technical challenge of balancing lightweight deployment and high-performance noise reduction in related technologies. Attached Figure Description
[0063] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0064] Figure 1 A schematic flowchart illustrating the speech noise reduction method provided in an embodiment of this application;
[0065] Figure 2 A schematic diagram illustrating the training process and model inference application process of the lightweight noise reduction model provided in the embodiments of this application;
[0066] Figure 3 A flowchart illustrating the speech noise reduction method in a single audio acquisition component scenario provided in this application embodiment;
[0067] Figure 4 Comparison of speech noise reduction effects in a single audio acquisition component scenario provided in this application embodiment;
[0068] Figure 5 A flowchart illustrating the speech noise reduction method in a dual-audio acquisition component scenario provided in this application embodiment;
[0069] Figure 6 Comparison of speech noise reduction effects in scenarios using dual-audio acquisition components provided in this application embodiment;
[0070] Figure 7 This is a schematic diagram of the structure of the speech noise reduction device provided in the embodiments of this application;
[0071] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0072] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0073] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application.
[0074] The terms “first,” “second,” etc., used in this application’s specification are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, products, or apparatus.
[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0076] Voice noise reduction technology is widely used in real-time voice communication scenarios of portable electronic devices, especially suitable for True Wireless Stereo (TWS) earphones, smartwatches, in-vehicle voice terminals, remote conferencing terminals, and other electronic devices with voice pickup and wireless transmission capabilities. In these scenarios, devices typically capture user voice through microphones and perform front-end audio processing, noise suppression, and voice output on a local chip to send the processed voice to the other end of the call. Due to the small size and limited power supply of the devices, they can usually only accommodate low-power processors, limited memory capacity, and simplified audio processing links. Therefore, the entire system architecture often needs to be highly compactly designed between acquisition, analysis, noise reduction, and reconstruction. At the same time, the actual usage environment is often not ideal. Users may be in subway cars, airport halls, bus stops, shopping mall atriums, while cycling, or walking outdoors. The environment contains both steady noise such as air conditioner noise, engine noise, and equipment background noise, as well as non-steady noise such as crowd noise, horns, collision sounds, door opening and closing sounds, and wind noise. These noises will superimpose with the target voice in the time and frequency domains, making it difficult for the other end of the call to clearly identify the content of the speech. For real-time communication devices, in addition to achieving good noise reduction, it is also necessary to take into account low processing latency, low power consumption and high voice naturalness. Otherwise, problems such as voice interruption, loss of details, communication delay or shortened battery life may occur. Therefore, this type of technology has become an important research and industrial application direction in the field of embedded audio processing.
[0077] In related technologies, to improve call quality in complex environments, traditional digital signal processing noise reduction schemes or deep learning-based end-to-end speech noise reduction schemes are commonly used. Traditional schemes typically first convert the acquired noisy speech to the frequency domain, then estimate the noise based on its statistical characteristics, and weaken the noise components through filtering or spectral suppression. These methods are effective in dealing with relatively stable interference such as air conditioning noise, fan noise, and road noise, and their algorithm structure is relatively clear with controllable resource consumption, making them suitable for embedded platform deployment. However, when the noise type changes rapidly, or when multiple people are talking simultaneously, or when there are sudden traffic noises, mechanical friction noises, and airflow impact noises in the environment, traditional methods often struggle to distinguish the target speech from noise components in a timely and accurate manner, easily resulting in noise residue, musical tones, speech blurring, and loss of high-frequency details. In contrast, deep learning-based end-to-end speech noise reduction schemes can utilize models to learn the nonlinear relationship between complex noise and speech, typically exhibiting stronger suppression capabilities in non-stationary noise scenarios. However, such solutions often rely on a large number of model parameters, high multiplication and addition operations, and high data access bandwidth. While they can achieve good results on the cloud or high-performance processors, when deployed to resource-constrained embedded devices such as headphones and watches, they often face problems such as insufficient storage, high power consumption, and difficulty in controlling inference latency. Even with lightweight compression of the model through quantization and pruning, the model's ability to discriminate weak speech and complex residual noise may decrease. In practical use, this manifests as unstable suppression in complex scenarios, significant fluctuations in performance under different noise conditions, and even oversuppression in speech segments, making the speech sound muffled, harsh, or unnatural. Furthermore, the training samples used by existing deep learning methods often favor general noisy speech, which is not entirely consistent with the noise distribution remaining after pre-processing in the actual processing chain of embedded devices, thus limiting the model's generalization ability in real-world deployment environments. Therefore, although related technologies have certain advantages in terms of resource consumption or noise reduction performance, they still struggle to simultaneously achieve noise reduction effect, speech fidelity, system stability, and deployment feasibility in real-time call scenarios with limited resources, low latency, and complex noise.
[0078] In view of this, how to improve the speech denoising effect in complex noisy environments under conditions of limited resources, low computing power and low latency, while taking into account speech clarity, naturalness and stability, has become an urgent technical problem to be solved.
[0079] To address the aforementioned issues, this application provides a speech denoising scheme. During real-time calls or other speech acquisition and processing, a frequency-domain noisy signal is first acquired and subjected to a first-stage denoising process to prioritize suppressing stationary noise, ensuring that the residual noise in the resulting first audio signal is dominated by non-stationary noise. Subsequently, time-frequency features are extracted from the first audio signal to obtain corresponding acoustic features, which are then input into a pre-trained lightweight denoising model for inference, outputting a noise suppression factor. This lightweight denoising model is trained based on a hybrid noisy speech library, which includes at least non-noisy speech and noisy speech processed by the first denoising step. Further denoising is then applied to the first audio signal based on the noise suppression factor to obtain a second audio signal. Finally, a time-frequency transformation is performed on the second audio signal to reconstruct the time-domain denoised speech signal. This technical approach, through a two-stage cascaded denoising architecture, reduces the inference burden on subsequent models and improves the model's adaptability to real residual noise, making it more suitable for achieving a balance between low resource consumption and superior denoising performance in portable electronic devices.
[0080] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0081] Figure 1 This is a flowchart illustrating the speech noise reduction method provided in this embodiment. This embodiment applies to electronic devices, which may be TWS earphones, smartwatches, in-vehicle voice terminals, remote conferencing terminals, or other portable devices with speech acquisition and processing capabilities. The executing entity may be a main control processor, an audio digital signal processor, a low-power neural network acceleration unit, or a processing system composed of the aforementioned hardware within the electronic device.
[0082] like Figure 1 As shown, the speech noise reduction method includes:
[0083] S101. Obtain the frequency domain noise signal corresponding to the original audio signal acquired by the audio acquisition component.
[0084] Audio acquisition components include microphones, sound cards, and other components that collect and process speech through vibration. The raw audio signal can be considered as the raw time-domain signal acquired by the audio acquisition components, containing speech components and various noise components (such as steady-state noise including air conditioner noise, fan noise, road noise, engine noise, etc., and non-steady-state noise such as crowd noise, collision noise, horn noise, door opening and closing noise, and wind noise). The frequency-domain noise-band signal is considered as the frequency domain representation of the audio signal containing speech and noise components. This representation is typically obtained by framing, windowing, and short-time Fourier transform of the time-domain sampled signal, and includes amplitude and phase spectra.
[0085] For example, a microphone in an electronic device first acquires analog speech signals from the environment, which are then converted into a digital time-domain audio stream by an analog-to-digital converter at a target sampling rate. For instance, in a real-time call scenario, the sampling rate can be set to 16kHz to balance speech bandwidth, clarity, and computational overhead; in other scenarios, 8kHz, 24kHz, or 32kHz can be used. Subsequently, the processor performs framing processing on the digital time-domain audio stream, setting the frame length and frame shift. For example, the frame length can be set to 10ms, 20ms, or 32ms, and the frame shift can be set to 1 / 2 or 1 / 3 of the frame length to ensure overlapping areas between adjacent frames and avoid spectral leakage during subsequent waveform reconstruction. After applying a Hamming window, Hanning window, or other window function to reduce spectral leakage to each frame, a Fast Fourier Transform is performed to obtain the frequency domain representation of the corresponding frame, thereby forming a continuously input frequency-domain signal sequence with noise.
[0086] In another possible embodiment, the frequency domain noisy frequency signal can also be obtained from the noisy recording data pre-stored in the device's storage unit through the same framing and frequency domain transformation process, for offline testing, model adaptation, or production debugging.
[0087] Based on the above analysis, this step completes the standardized frequency domain representation of the original noisy speech (i.e., the original audio signal), mapping the speech and noise components, which were originally continuously changing on the time axis, to a computable time-frequency grid, providing a unified input for the subsequent first noise reduction process. Since frequency domain representation can more directly distinguish the harmonic structure of speech and the distribution of noise energy in different frequency bands, adopting a frequency-domain representation followed by hierarchical noise reduction in resource-constrained devices is beneficial for achieving more targeted noise suppression with limited computational resources.
[0088] S102. Perform a first noise reduction process on the frequency domain band noise signal to obtain a first audio signal. The residual noise contained in the first audio signal is dominated by non-steady-state noise.
[0089] The first noise reduction process is essentially a noise reduction operation based on Digital Signal Processing (DSP). This noise reduction operation estimates and suppresses noise in the frequency domain band-noise signal (specifically, the amplitude spectrum of the frequency domain band-noise signal) based on preset signal processing rules, spectral statistical characteristics, and real-time updated parameters. Its core purpose is not to eliminate all noise at once, but to take advantage of the low computational cost of DSP algorithms to prioritize filtering out steady-state noise that is easier to model and changes slowly over time, while also suppressing some non-steady-state noise. Thus, with limited chip resources, the residual noise that the subsequent AI model needs to process is concentrated into a signal form dominated by non-steady-state noise.
[0090] The first audio signal refers to the frequency domain audio signal after this step. Its main speech structure is preserved, while the steady-state noise energy in the background has been greatly reduced, and the non-steady-state noise has been partially reduced.
[0091] It should be noted that the first noise reduction process in this embodiment is not limited to the combination of improved MCRA and Wiener filtering. Other excellent traditional noise reduction algorithms based on DSP, such as spectral subtraction and OM-LSA, can also be used, as long as they can achieve the technical effect of "filtering out most of the steady-state noise and some non-steady-state noise, and making the residual noise dominated by non-steady-state noise".
[0092] In practical embedded deployments, to avoid significant damage to the main components of speech, the first noise reduction process in this step can optionally include a gain lower limit, a reserved noise threshold, and a speech protection mechanism to prevent speech distortion and deterioration of sound quality caused by excessive noise compression. For example, the gain lower limit can be set between 0.1 and 0.4 to effectively protect weak speech components with spectral characteristics similar to the background noise (such as voiceless consonants / s / , / f / , / h / , and low-frequency overtones of voiced consonants), avoiding muffled speech, loss of detail, or even word breaks due to excessive compression. The reserved noise threshold is used to set a minimum threshold for noise suppression: background noise with an intensity below the threshold is not processed, and a small amount of natural residual noise is intentionally retained, which avoids the damage to weak speech and maintains the naturalness of the sound, i.e., "retaining a small amount of background noise to ensure sound quality". The speech protection mechanism can reduce the suppression intensity of frequency bands that may have strong speech harmonics based on simple speech activity detection results, thereby preserving the formant structure and high-frequency clarity of the speech. Through this processing, the steady-state background noise that was originally mixed in with the noisy frequency signal is preferentially weakened, while non-steady-state components that are difficult to be fully modeled by traditional noise reduction methods, such as human voice interference, sudden noise, and wind noise disturbance, are retained as the main residues in the first audio signal and are then handled by the subsequent lightweight AI model.
[0093] Based on the above analysis, this step, by processing noise in layers, first uses a DSP algorithm with low computational overhead and strong determinism to undertake most of the steady-state noise suppression task and part of the non-steady-state noise suppression task. This not only reduces the noise complexity that the subsequent lightweight noise reduction model needs to handle, but also makes the input distribution faced by the model closer to the real residual noise distribution after the previous stage processing. This improves the noise reduction stability and deployment feasibility of the overall system under resource-constrained and low-latency conditions.
[0094] S103. Extract time-frequency features from the first audio signal to obtain the acoustic features corresponding to the first audio signal, and input the acoustic features into the pre-trained lightweight noise reduction model. Infer the acoustic features through the lightweight noise reduction model and output the noise suppression factor. The lightweight noise reduction model is trained based on a hybrid noisy speech library. The hybrid noisy speech library contains at least non-noisy speech and noisy speech after the first noise reduction process.
[0095] In this context, time-frequency feature extraction refers to the process of extracting information representing the difference between speech and residual noise from the time-domain variation and frequency-domain energy distribution of the first audio signal. Acoustic features can be logarithmic amplitude spectrum, power spectrum envelope, cepstral features, adjacent frame difference features, context splicing features, or tensors formed by combining multiple basic features. In this embodiment, acoustic features specifically refer to low-dimensional parameters extracted from the audio signal to characterize speech properties. Preferably, the acoustic features use Mel-Frequency Cepstral Coefficients (MFCCs), which have a 13-dimensional dimension.
[0096] Lightweight noise reduction models can be viewed as neural network models suitable for deployment on embedded devices in terms of parameter size, storage footprint, and inference multiply-accumulate computation. They can employ shallow convolutional networks, gated recurrent networks, depthwise separable convolutional networks, temporal convolutional networks, attention-simplified networks, or compact models formed by combinations of these structures. For example, a lightweight noise reduction model can be a miniature deep neural network (DNN) containing only two layers of neurons, or a hybrid model composed of a DNN and a gated recurrent unit (GRU). Its input is a 13-dimensional MFCC, and the number of model parameters is compressed to less than 30KB. Through quantization and fixed-point computation, the inference computation per frame can be controlled to within 100 MIPS, enabling real-time processing of sudden non-steady-state noise.
[0097] The noise suppression factor is the gain value output by the lightweight noise reduction model for each frequency point or frequency band in the current frame. It can be represented as a continuous value between 0 and 1, reflecting the degree of speech preservation and noise suppression in the corresponding time-frequency unit. The closer the noise suppression factor is to 0, the stronger the suppression of noise in that time-frequency unit; the closer it is to 1, the stronger the preservation of speech in that time-frequency unit.
[0098] For example, a Mel filter bank is used to process the amplitude spectrum of the first audio signal to obtain the Mel spectrum. Taking the logarithm of the Mel spectrum and then performing a discrete cosine transform yields MFCC features with lower dimensionality but stronger perceptual correlation. The generated MFCC features are then input into a pre-trained lightweight noise reduction model. During forward inference, the model maps time-frequency correlations and outputs the noise suppression factor for each frequency point or band.
[0099] It should be understood that the construction method of the model training data directly affects the deployment effect. Therefore, this step explicitly adopts a lightweight denoising model trained based on a hybrid noisy speech library. The hybrid noisy speech library contains at least two types of key data: one is noisy speech (i.e., clean speech data), which provides an ideal reference for the target speech; the other is noisy speech data after the first denoising process. This type of data is not ordinary raw noisy speech, but rather simulates the input form that still has residual noise after the actual device's front-end DSP denoising process. In addition, the hybrid noisy speech library also includes noisy speech that has not undergone the first denoising process.
[0100] For example, when constructing a hybrid noisy speech library, it can be generated as follows: First, clean speech is mixed with various real-world scene noises (including but not limited to bar noise, road noise, intersection noise, train carriage noise, coffee shop noise, Mensa noise, call center noise, and wind noise) at different signal-to-noise ratios to form the original noisy speech; then, half of the original noisy speech is subjected to DSP noise reduction processing (i.e., the first noise reduction processing) that is consistent with or approximately consistent with the device side, to obtain the noisy speech data after the first noise reduction processing; finally, the other half of the original noisy speech and the noisy speech data after the first noise reduction processing are used as training inputs; using the ideal ratio mask or ideal gain of clean speech to noisy speech as supervision labels, the model is trained to learn the mapping relationship from "residual noise speech after DSP noise reduction" to "ideal clean speech".
[0101] In one possible embodiment, the hybrid noisy speech library may also include the following enhancement data to improve the model's adaptability to real-world deployment environments: data collected under different device microphone response characteristics; data under different reverberation conditions; and data generated under different wind noise intensities. Furthermore, to accommodate the limited resources of embedded devices, model quantization can employ 8-bit fixed-point quantization or 16-bit half-precision floating-point representation, and model weights can be stored in on-chip memory or low-latency external memory to effectively control power consumption and access latency.
[0102] Based on the above analysis, the key to this step lies not in directly using neural networks to process the original complex noisy speech, but in focusing the model on the non-steady-state noise patterns that remain after DSP denoising. This significantly reduces the noise space the model needs to learn, lowering parameter requirements and inference burden. Furthermore, because the training data explicitly includes noisy speech processed by DSP denoising, the model can learn the input-output correspondence in the actual processing chain of the device, mitigating the problem of inconsistency between training and deployment distributions. Therefore, it is more conducive to improving the accuracy, stability, and weak speech protection capabilities of noise suppression factors in complex noise scenarios.
[0103] S104. Based on the noise suppression factor, perform a second noise reduction process on the first audio signal to obtain a second audio signal.
[0104] The second noise reduction process essentially utilizes the noise suppression factor output by the lightweight noise reduction model to perform noise reduction on the corresponding frequency components of the first audio signal. This further suppresses residual non-steady-state noise and steady-state noise remaining in the first audio signal while preserving the target speech components. The second audio signal is a frequency-domain audio signal enhanced by the model. Compared to the first audio signal, its energy is further weakened by complex interferences such as sudden noise, background human voice interference, and wind noise pulsation, while the main spectral structure of the speech is maintained, improving the naturalness and clarity of the sound. In other words, the noise suppression factor can specifically suppress non-steady-state noise that is difficult to handle by the first-stage noise reduction (DSP noise reduction).
[0105] It should be noted that in this step, the second noise reduction process only adjusts the gain of the amplitude spectrum of the first audio signal, and its corresponding phase spectrum remains unchanged (that is, the phase spectra of the first audio signal and the second audio signal are both equal to the phase spectrum of the frequency domain noise signal). The gain application method will be further defined in subsequent embodiments.
[0106] Based on the above analysis, this step transforms the inference results output by the lightweight noise reduction model into executable controllable quantities for the actual audio spectrum. Since the noise suppression factor is generated for the first audio signal after DSP noise reduction, its adjustment target focuses on the remaining complex noise components, achieving a more targeted suppression effect on non-steady-state noise with lower model complexity. Simultaneously, by smoothing and constraining the noise suppression factor, the risk of speech artifacts caused by model output jitter can be reduced, thus balancing speech clarity, naturalness, and real-time processing stability.
[0107] S105. Based on the second audio signal, the time-domain denoised speech signal corresponding to the original audio signal is reconstructed.
[0108] Among them, time-domain noise-reduced audio signals are clean speech signals in the time domain that electronic devices ultimately use for playback, transmission, recording, or further encoding and decoding processing. They should be continuous in time, stable in amplitude, and retain the intelligibility and natural listening experience of the original speech content as much as possible.
[0109] For example, in some embodiments, S105 may specifically include: combining the amplitude spectrum corresponding to the second audio signal with the phase spectrum corresponding to the frequency domain noise signal to obtain the frequency domain denoised speech signal corresponding to the original audio signal; and performing time-frequency transformation on the frequency domain denoised speech signal to obtain the time domain denoised speech signal.
[0110] Specifically, the amplitude spectrum of the second audio signal is used as the modulus of the complex spectrum, and the phase spectrum of the frequency domain signal with noise is used as the argument of the complex spectrum, and they are combined according to the complex representation:
[0111]
[0112] In the formula, The amplitude spectrum represents the amplitude distribution of the second audio signal at various frequency points. The phase spectrum of the frequency domain signal with noise (or the phase spectrum of the first audio signal can be used, since the phase spectrum does not change substantially in the second noise reduction process). This represents the reconstructed frequency-domain denoised speech signal; This indicates the frequency index.
[0113] Then, the frequency-domain denoised speech signal is subjected to an inverse short-time Fourier transform (ISTFT) to convert it back from the frequency domain to the time domain, resulting in a time-domain speech frame sequence.
[0114] Finally, adjacent temporal speech frames are spliced together using the overlap-add method to eliminate the boundary effects introduced by frame splitting and form a continuous and smooth temporal denoised speech signal.
[0115] It should be noted that the phase spectrum of the frequency domain signal with noise is used for reconstruction in this step, based on the auditory characteristic that the human ear is not sensitive to phase changes. This method avoids the additional computational overhead of phase estimation while ensuring the naturalness of the output speech, thus reducing the overall computing power requirements. In other possible embodiments, the phase spectrum can also be smoothed or corrected to further optimize the listening experience, but a trade-off must be made between computational complexity and the improvement in speech quality.
[0116] In the streaming processing architecture, the ISTFT and overlap-add operations are performed frame by frame. After each frame is processed, the corresponding time-domain signal is output. There is no need to buffer the entire audio segment, thus ensuring that the latency of the entire system is controlled within one frame (e.g., if one frame is 30ms, the latency can be controlled within 30ms; if one frame is 15ms, the latency can be controlled within 15ms), which meets the requirements of real-time voice calls.
[0117] To improve the final output quality, this step can further process the reconstructed time-domain noise-reduced frequency signal by applying amplitude limiting, DC bias elimination, short-time energy smoothing, or comfort noise compensation. Amplitude limiting prevents excessively high peak values that could cause distortion after some frames are superimposed; DC bias elimination avoids long-term offset affecting downstream equipment; and short-time energy smoothing reduces loudness jitter caused by frame-level gain changes. If the system design requires improved subjective naturalness, controlled comfort noise can be injected at a low amplitude after the noise has been significantly reduced, ensuring the output signal is not too abrupt during silence intervals. After completing the above processing, the time-domain noise-reduced frequency signal can be output frame by frame according to a real-time call buffering strategy, thereby meeting the requirements of low-latency communication.
[0118] Based on the above analysis, this step enables the phased noise reduction results to be implemented as real-time audio output in electronic devices. It restores continuous waveforms while maintaining manageable computational complexity, and fully incorporates the synergistic results of DSP noise reduction and lightweight model noise reduction into the final audio, thereby achieving a speech output effect that balances clarity, naturalness, stability, and low latency even in complex noise environments.
[0119] In this embodiment, after acquiring the frequency-domain noisy signal, a first-stage noise reduction process filters out most of the steady-state noise and some of the non-steady-state noise, making the residual noise in the first audio signal predominantly non-steady-state noise, thus significantly reducing the processing pressure on the subsequent AI model. Secondly, a second-stage noise reduction process further suppresses residual noise, especially specifically suppressing residual non-steady-state noise, eliminating the need for the AI model to participate in the entire noise reduction process and significantly reducing the overall computational load. Furthermore, the AI model's input uses low-dimensional acoustic features extracted from its features, further compressing the model's parameters to adapt to the limited computing resources of embedded devices. Simultaneously, training is performed using a specially constructed hybrid noisy speech library to ensure the AI model's accuracy in suppressing residual noise after the first-stage noise reduction. Through this two-stage cascaded noise reduction architecture, this application can achieve efficient speech noise reduction in complex noise environments on resource-constrained, low-computing-power embedded devices, ensuring the clarity and naturalness of the output speech, effectively solving the technical challenge of balancing lightweight deployment and high-performance noise reduction in related technologies.
[0120] Furthermore, some related technologies first identify the noise type of the frequency domain noisy signal (e.g., distinguish between steady-state and non-steady-state noise), and then select the corresponding noise reduction strategy based on the identification results. However, in real-world complex acoustic environments, multiple types of noise often coexist and superimpose, and the boundaries between noise types are blurred, making accurate and real-time noise classification difficult. Under these circumstances, the feasibility and robustness of selecting a noise reduction strategy based on the noise type identification results are greatly reduced, easily leading to misjudgments and untimely strategy switching, which in turn affects the noise reduction effect and speech quality. In contrast, this application does not require pre-distinguishing noise types, but instead adopts a two-stage cascaded noise reduction architecture. The first stage (DSP noise reduction) uses traditional signal processing methods to uniformly filter out most of the steady-state noise and some non-steady-state noise without relying on noise classification, making the residual noise dominated by non-steady-state noise. The second stage (lightweight AI noise reduction) uniformly processes the residual non-steady-state noise and a small amount of steady-state noise without distinguishing the specific noise source. This architecture fundamentally avoids the uncertainty caused by noise type identification, has stronger robustness and adaptability in complex noisy environments, and reduces the dependence on hardware resources, making it more suitable for deployment in resource-constrained embedded devices.
[0121] In some embodiments, a second noise reduction process is performed on the first audio signal according to a noise suppression factor to obtain a second audio signal, including:
[0122] Step 1041: Multiply the noise suppression factor by the amplitude spectrum of the first audio signal to obtain the noise-reduced amplitude spectrum.
[0123] The amplitude spectrum of the first audio signal is the amplitude component of the frequency domain signal obtained after the first noise reduction process (DSP noise reduction process), which includes residual noise dominated by non-steady-state noise.
[0124] For example, in step S103, the noise suppression factor output by the AI model is multiplied frequency-by-frequency by the amplitude spectrum of the first audio signal obtained in step S102:
[0125]
[0126] In the formula, This represents the amplitude spectrum of the first audio signal; Indicates the noise suppression factor; This is the amplitude spectrum after noise reduction.
[0127] Step 1042: Determine the second audio signal based on the amplitude spectrum after noise reduction.
[0128] For example, in one implementation, the amplitude spectrum after noise reduction is directly determined as the second audio signal.
[0129] However, in practical speech noise reduction applications, there is a trade-off between noise reduction depth and speech fidelity. Applying excessive noise reduction to speech segments can lead to the loss of high-frequency details, blurring of voiceless consonants, and weakening of emotional information, resulting in a harsh and unnatural-sounding output. Conversely, incomplete noise reduction in non-speech segments (silent or purely noisy segments) leaves residual noise that is perceived by the user, affecting the overall listening experience, especially in quiet environments. Therefore, a mechanism is needed that can accurately identify speech activity and dynamically adjust the noise reduction intensity based on the identification results, achieving adaptive control of "light noise reduction to protect speech when speech is present, and strong noise reduction to eliminate residual noise when speech is absent."
[0130] Based on this, in another implementation, the second audio signal is determined based on the amplitude spectrum after noise reduction, which may specifically include:
[0131] Step 1042-1: Determine the corresponding energy entropy ratio based on the amplitude spectrum after noise reduction. The energy entropy ratio represents the probability of speech existence.
[0132] The energy-entropy ratio is a metric used to distinguish the relative strength of speech components and background noise. It can be calculated from the energy distribution of the amplitude spectrum in the corresponding frequency band after noise reduction and the spectral entropy characteristics. Spectral entropy is used to quantify the irregularity or uncertainty of the current signal spectrum, that is, to characterize the dispersion of the amplitude spectrum in the frequency domain. Its core function is to distinguish speech from noise and to help determine whether there is valid speech in the current frame and the complexity of the noise. When speech activity increases, the spectrum usually exhibits stronger structure, the spectral entropy decreases, and the energy-entropy ratio increases accordingly. When in a silent or non-speech segment, the spectrum distribution tends to be flatter, the spectral entropy increases, and the energy-entropy ratio decreases accordingly.
[0133] For example, the short-time energy of the current audio frame is calculated from the denoised amplitude spectrum. This short-time energy reflects the signal strength of the current frame. Typically, the energy of a speech frame is higher than that of a noise frame. The spectral entropy of the current audio frame is then calculated from the denoised amplitude spectrum. This spectral entropy reflects the uniformity of the spectral energy distribution. Speech spectral energy is concentrated in certain frequency bands, resulting in lower spectral entropy; noise spectral energy is more uniformly distributed, resulting in higher spectral entropy. It should be noted that the formulas for calculating short-time energy and spectral entropy can be found in existing technologies and will not be elaborated here.
[0134] It should be understood that speech segments have higher energy and lower spectral entropy, resulting in a higher energy entropy ratio; while noise segments have lower energy and higher spectral entropy, resulting in a lower energy entropy ratio. Therefore, the energy entropy ratio can more robustly distinguish between speech and non-speech segments, especially in low signal-to-noise ratio environments where relying solely on energy detection is prone to failure. The energy entropy ratio, which incorporates spectral distribution information, has stronger anti-interference capabilities.
[0135] Step 1042-2: If the energy entropy ratio is less than the energy entropy ratio threshold and the duration is greater than or equal to the preset duration, it is determined that there is no speech activity. The amplitude spectrum after noise reduction is multiplied by the non-speech segment noise suppression factor to obtain the second audio signal. The non-speech segment suppression factor is used to suppress the residual noise of the non-speech segment.
[0136] In this embodiment, the second audio signal should be understood as the final noise-reduced audio signal.
[0137] Among them, the non-speech segment noise suppression factor can be a suppression coefficient vector that corresponds one-to-one with the frequency point. Its value range is usually set between 0.2 and 0.4, for example, 0.316, which is used to additionally attenuate the residual noise in the non-speech segment.
[0138] For example, the calculated energy entropy ratio is compared with a preset energy entropy ratio threshold, and the continuous duration (or number of consecutive frames) when the energy entropy ratio is lower than the energy entropy ratio threshold is counted. When both of the above conditions are met, the current audio frame is determined to be in a non-speech segment (i.e., there is no speech activity).
[0139] This dual-judgment mechanism can effectively avoid misjudgments caused by transient noise interference and improve the robustness of speech activity detection.
[0140] When a segment is determined to be non-speech segment, an additional non-speech segment noise suppression factor is applied to the amplitude spectrum after noise reduction to obtain the final second audio signal. This avoids excessive noise residue in areas without speech when noise remains, which would affect the overall listening experience of the speech. In other words, it significantly reduces the user's perception of residual noise and improves the overall listening experience.
[0141] Correspondingly, in some embodiments, the speech denoising method further includes: if the energy entropy ratio is greater than or equal to the energy entropy ratio threshold, or the duration is less than a preset duration, then the amplitude spectrum after denoising is determined as the second audio signal.
[0142] It can be understood that if the above two conditions are not met at the same time (i.e., the energy entropy ratio is greater than or equal to the energy entropy ratio threshold, or the duration does not reach the preset duration), it is determined that there is speech activity or the current frame is in a transition state. At this time, no additional non-speech segment noise suppression factor is applied, and the amplitude spectrum after noise reduction is directly used as the second audio signal.
[0143] This method avoids excessive noise reduction in the speech segment, preserving high-frequency details and emotional information of the speech.
[0144] The aforementioned mechanism, based on energy entropy ratio-based speech activity detection and duration determination, achieves accurate identification of speech and non-speech segments. Building upon this, backend optimization further reduces noise: speech segments maintain their original noise reduction intensity to protect speech quality, while non-speech segments receive additional non-speech noise suppression to further eliminate residual noise. This technical solution effectively resolves the contradiction between "noise reduction depth and speech fidelity" in speech denoising, achieving optimal noise reduction for non-speech segments while ensuring speech clarity and naturalness, thus improving the overall user's auditory experience. Furthermore, this solution requires low computational resources and no additional hardware resources, making it suitable for resource-constrained embedded devices such as TWS earphones and hearing aids.
[0145] In some embodiments, a first noise reduction process is performed on the frequency domain noisy signal to obtain a first audio signal, including:
[0146] Step 1021: Perform improved MCRA noise estimation on the frequency domain band noise signal to obtain the noise power spectrum.
[0147] Among them, the improved MCRA noise estimation uses a recursive method to continuously update the minimum search window instead of resetting the window every frame, and uses a two-stage recursive average to estimate the speech presence probability to support a smaller search window length.
[0148] This embodiment uses a recursive approach to continuously update the minimum power value of each frequency point. The recursive approach continuously updates the minimum value estimate by comparing the power of the current frame with the historical minimum value, eliminating the need to rescan the entire window every frame, unlike the traditional MCRA algorithm which resets the search window independently for each frame. Furthermore, a two-stage recursive averaging method is used to estimate the probability of speech presence: the first stage smooths the power spectrum of the current frame in the temporal domain to obtain a smoothed power spectrum; the second stage calculates the probability of speech presence based on the ratio of the smoothed power spectrum to the recursive minimum value, and recursively averages the noise power spectrum based on the speech presence probability, improving the estimation accuracy of the speech presence probability. Because a recursive minimum tracking mechanism is used, there is no need to rely on a large window to ensure stability, thus allowing for a smaller search window length (e.g., 1 / 2 to 1 / 4 of that in traditional algorithms, such as 15-30 frames). After recursive minimum tracking and two-stage recursive averaging, the noise power spectrum of the current frame is output, which will serve as the input for OM-LSA or Wiener filtering in step S1022.
[0149] Step 1022: Based on the noise power spectrum, use OM-LSA or Wiener filtering to perform noise reduction processing on the frequency domain noise signal to obtain the first audio signal.
[0150] Wiener filtering is used to construct a frequency domain gain based on the noise power spectrum to suppress noisy speech spectra. During processing, the noise power spectrum is combined with the observed power spectrum at the current frequency point to calculate the filtering coefficients for each frequency. These coefficients are then applied to the complex spectrum or amplitude spectrum of the noisy frequency signal in the frequency domain, thereby reducing noise components and preserving the target speech components. The result of Wiener filtering, after phase preservation or phase reconstruction, can form a first audio signal for subsequent processing.
[0151] OM-LSA is a frequency domain gain calculation method based on the probability of speech presence, which can suppress residual noise more effectively than Wiener filtering. The OM-LSA estimator calculates the optimal logarithmic spectral amplitude gain based on the probability of speech presence, multiplies the calculated logarithmic spectral amplitude gain with the amplitude spectrum of the frequency domain noisy signal, and preserves the original phase spectrum (phase unchanged) to obtain the first audio signal.
[0152] This embodiment employs a recursive approach to continuously update the minimum search window instead of resetting the window every frame, significantly reducing processing latency. The recursive update mechanism also supports smaller search window lengths, accelerating the response to noise changes and enhancing the tracking capability for non-stationary noise. Furthermore, by estimating the probability of speech presence through a two-stage recursive averaging method, it effectively avoids speech contamination of the noise model, protecting speech components from excessive suppression. It also supports both OM-LSA and Wiener filtering for noise reduction, allowing for flexible selection based on hardware resources, balancing noise reduction effectiveness and computational efficiency, and adapting to the real-time operation requirements of resource-constrained embedded devices.
[0153] In some embodiments, the speech noise reduction method is applied to an electronic device, which includes a first audio acquisition component and a second audio acquisition component. For example, the first and second audio acquisition components can be two microphone units spaced apart from each other, with a predetermined distance between their mounting positions on the electronic device housing, in order to acquire acoustic signals with spatial differences.
[0154] For example, taking the first and second audio acquisition components as microphones, i.e., an electronic device integrating a dual-microphone speech noise reduction system, wind noise is usually present in such systems. Wind noise is near-field turbulence noise formed by airflow directly impacting the microphone pickup hole, and its acoustic characteristics differ from other environmental noises (such as air conditioner noise and human voice reverberation). If wind noise is not effectively suppressed, it will cause low-frequency speech distortion and a "whooshing" blowing sound after noise reduction, seriously affecting speech clarity and user listening experience.
[0155] Therefore, for this dual-audio-frequency acquisition component speech noise reduction system, the spatial information of the dual-audio-frequency acquisition component can be used to effectively suppress wind noise without damaging low-frequency speech.
[0156] Accordingly, the frequency domain noise frequency signal corresponding to the original audio signal acquired by the audio acquisition component is obtained, including:
[0157] Step 1011: Obtain the first frequency domain signal corresponding to the original audio signal acquired by the first audio acquisition component, and the second frequency domain signal corresponding to the original audio signal acquired by the second audio acquisition component.
[0158] For example, two audio acquisition components acquire data synchronously to ensure signal timing alignment. The first frequency domain signal is a complex frequency domain signal obtained by performing a short-time Fourier transform (STFT) on the original audio signal (time domain) acquired by the first audio acquisition component, containing both amplitude and phase spectra. The second frequency domain signal is a complex frequency domain signal obtained by performing a short-time Fourier transform (STFT) on the original audio signal (time domain) acquired by the second audio acquisition component, also containing both amplitude and phase spectra.
[0159] Step 1012: Perform first-order differential array processing on the first frequency domain signal and the second frequency domain signal to obtain the differential array signal.
[0160] First-order differential array processing is a spatial differential beamforming method based on dual-audio acquisition components (such as dual microphones). For example, by differentially processing the frequency domain signals of two microphones, spatial directivity is formed, such as enhancing the voice directly in front and suppressing noise from the sides and rear. In wind noise processing, the single-channel frequency domain signal after the first-order differential array (i.e., the differential array signal output after differential array processing) serves as the basis signal for wind noise suppression.
[0161] For example, the first frequency domain signal and the second frequency domain signal are input into a first-order differential array module. Through differential operations, the differential array forms zeros (signal cancellation points) in a specific direction, thereby achieving spatial filtering.
[0162] It should be noted that for small-sized devices such as stick-shaped TWS earbuds, the distance between two audio acquisition components (such as microphones) is limited (usually 10-20mm), and a microphone array can be formed. Theoretical analysis and experimental verification show that differential arrays can still maintain good directivity in the low-frequency range, while additive minimum variance distortionless response (MVDR) exhibits significant directivity degradation under small aperture conditions, and its performance is not as good as that of differential arrays. Therefore, this embodiment preferentially uses a differential array (specifically a first-order differential array) to achieve beamforming of dual audio acquisition components (such as dual microphones).
[0163] Step 1013: Based on the spatial difference characteristics between the first frequency domain signal and the second frequency domain signal, determine whether there is wind noise in the frequency domain band noise signal.
[0164] Spatial difference features are characteristic parameters used to characterize the spatial relationship between the first and second frequency domain signals, including but not limited to: one or more combinations of the amplitude ratio, phase difference, energy difference, sum-difference ratio, coherence coefficient, or centroid spectral offset of the two frequency domain signals. It should be understood that the presence of wind noise will cause specific changes in the spatial difference features.
[0165] When the spatial difference characteristics meet the preset wind noise conditions (such as the sum-difference ratio being lower than the corresponding threshold, the centroid spectrum shifting to the low frequency exceeding the threshold, and the coherence coefficient being lower than the corresponding threshold), it is determined that there is wind noise in these two frequency domain signals.
[0166] For example, in one implementation, the sum-difference ratio is calculated based on the first and second frequency domain signals (when wind noise is present, the correlation between the two microphone signals decreases, and the sum-difference ratio changes), for example:
[0167]
[0168] In the formula, Frequency index; Indicates the ratio of sum to difference; Indicates the first frequency domain signal; This represents the second frequency domain signal.
[0169] Based on the first and second frequency domain signals, the centroid spectrum is calculated (wind noise energy is concentrated in low frequencies, which causes the centroid spectrum to shift to lower frequencies), for example:
[0170]
[0171] In the formula, Indicates the center-of-mass spectrum; This represents the frequency domain signal obtained by weighted fusion of the first and second frequency domain signals.
[0172] It should be noted that the above formulas for calculating the sum-difference ratio and centroid spectrum are merely exemplary implementations and are not intended to limit the scope of protection of this application.
[0173] Furthermore, the calculated spatial difference features (including the sum-difference ratio and the centroid spectrum) are compared with the preset wind noise detection threshold: when the sum-difference ratio is lower than the sum-difference ratio threshold and the centroid spectrum shifts to the low frequency beyond the low frequency shift threshold, it is determined that wind noise exists; otherwise, it is determined that there is no wind noise or the wind noise is negligible.
[0174] Step 1014: If wind noise is determined to exist, determine the corresponding wind noise suppression factor based on the spatial difference characteristics.
[0175] For example, in one implementation, a wind noise suppression factor is calculated based on the degree of deviation of spatial difference features from the wind noise detection threshold. A greater deviation indicates more severe wind noise, and a smaller wind noise suppression factor (i.e., stronger suppression) is generated. The wind noise suppression factor can be obtained through table lookup, piecewise mapping, or linear mapping. In other words, there is a mapping relationship between spatial difference features and the wind noise suppression factor.
[0176] Optionally, in some embodiments, when wind noise is determined to exist, the Wiener filter gain can be determined based on the noise power spectrum and the spectrum of the acquired original audio signal; based on the Wiener filter gain, the differential array signal is subjected to wind noise suppression processing to obtain the signal after wind noise suppression.
[0177] The noise power spectrum is determined based on the sum-difference ratio and the power spectrum of the acquired original audio signal. The sum-difference ratio is calculated as follows: the spectral difference between the first and second frequency domain signals is used as the denominator, and the sum of the spectra of the first and second frequency domain signals is used as the numerator. The spectrum of the acquired original audio signal is determined based on the weighted fusion result of the first and second frequency domain signals. The power spectrum of the acquired original audio signal is determined based on its spectrum. Specific implementation details of this embodiment can be found in existing technologies and will not be elaborated here.
[0178] Step 1015: Based on the wind noise suppression factor, perform wind noise suppression processing on the differential array signal to obtain the signal after wind noise suppression.
[0179] For example, the wind noise suppression factor is multiplied frequency-by-frequency by the amplitude spectrum of the differential array signal to obtain the wind noise-suppressed signal. Wind noise suppression only modifies the amplitude spectrum, not the phase spectrum. The phase spectrum of the wind noise-suppressed signal is the same as the phase spectrum of the differential array signal.
[0180] Step 1016: Use the signal after wind noise suppression as the frequency domain band noise signal.
[0181] The wind noise-suppressed signal is then used as the input signal for the subsequent first noise reduction process (DSP noise reduction), enabling the subsequent process to continue performing steady-state noise suppression and residual noise elimination on the frequency band noise signal in the frequency domain that has already had its wind noise reduced. This method allows wind noise identification, wind noise reduction, and subsequent noise reduction to work together to reduce the pollution of frequency domain features by wind noise on resource-constrained electronic devices.
[0182] In the implementation method of this application, the spatial difference characteristics between the two acquired signals (i.e., the first frequency domain signal and the second frequency domain signal) are used for wind noise identification and suppression. This weakens the wind noise component before the first noise reduction, reduces the burden of subsequent processing, and lowers the misjudgment and residual noise caused by wind noise. Since the differential array signal is specifically suppressed before entering the first noise reduction processing link, it helps to improve the speech clarity, stability, and real-time processing performance in complex environments, while reducing the dependence on highly complex models, making it suitable for deployment in portable electronic devices.
[0183] In some embodiments, the speech noise reduction method further includes: if it is determined that there is no wind noise, then treating the differential array signal as a frequency domain band noise frequency signal.
[0184] Understandably, after determining that there is no wind noise, the differential array signal is no longer subjected to wind noise suppression transformation. Instead, the differential array signal is directly used as the frequency domain band noise signal, and then as the input signal for the subsequent first noise reduction process (DSP noise reduction).
[0185] By directly using the differential array signal as the frequency domain noise signal when wind noise is absent, repeated suppression processing on the wind-noise-free signal can be avoided, reducing unnecessary computational overhead and spectral distortion. Simultaneously, more effective speech details are preserved, enabling subsequent first-stage noise reduction processing to suppress steady-state noise and improve residual noise modeling based on a more stable input. Therefore, the latency and power consumption of the entire speech noise reduction chain can be controlled, and the clarity and naturalness of the output speech are more easily maintained.
[0186] In some embodiments, both the first frequency domain signal and the second frequency domain signal are determined by: acquiring the original audio signal acquired by the corresponding audio acquisition component; performing front-end noise reduction processing on the original audio signal to obtain the third audio signal, wherein the front-end noise reduction processing includes anti-aliasing filtering and / or oversampling processing; performing short-time Fourier transform on the third audio signal to obtain the fourth audio signal; and performing linear echo cancellation processing on the fourth audio signal to obtain the frequency domain signal corresponding to the original audio signal acquired by the corresponding audio acquisition component, wherein the linear echo cancellation processing is implemented using an adaptive filtering algorithm to estimate and cancel the linear echo components in the fourth audio signal.
[0187] It is understandable that the determination methods for the first frequency domain signal and the second frequency domain signal are the same. The following explanation will take the first frequency domain signal as an example.
[0188] In the above processing, the raw audio signal can be directly acquired by the audio acquisition components of the electronic device (such as a microphone acquisition unit) without any processing, and contains various components such as speech, ambient noise, and echo. Front-end noise reduction is used to suppress high-frequency aliasing components in the sampling link before entering frequency domain analysis, and to improve the frequency resolution and numerical stability of subsequent transforms through oversampling. The resulting third audio signal still retains the noisy speech content to be processed. Oversampling is optional and is selectively executed based on the actual application scenario and hardware capabilities.
[0189] For example, the original audio signal is input to an anti-aliasing filter (analog low-pass filter) to filter out high-frequency noise above the cutoff frequency, resulting in a filtered analog signal. The filtered analog signal is then converted from analog to digital. If an oversampling strategy is used, sampling is performed at a frequency much higher than the target sampling rate (e.g., target 48kHz, actual sampling rate 192kHz). After oversampling, quantization noise is dispersed across a wider frequency band. Subsequently, a digital decimation filter is used to reduce the sampling rate back to the target value, while simultaneously filtering out out-of-band quantization noise, ultimately yielding the third audio signal (time domain). Then, a short-time Fourier transform is performed on the third audio signal. The short-time Fourier transform can be implemented using a frame-by-frame windowing method, with window functions such as Hamming windows, Hanning windows, or rectangular windows, to convert the third audio signal into a fourth audio signal (frequency domain) containing amplitude information (amplitude spectrum) and phase information (phase spectrum). Next, linear echo cancellation processing is performed on the fourth audio signal. The linear echo cancellation processing can establish an adaptive filter based on the least mean square algorithm, normalized least mean square algorithm, or recursive least square algorithm, etc., to estimate the echo path in the fourth audio signal online, and update the filter coefficients in real time according to the error signal, thereby reducing the echo interference generated by the far-end speech through the speaker-microphone path.
[0190] By setting up a collaborative link between front-end noise reduction and frequency domain echo cancellation before frequency domain processing, the impact of invalid high-frequency disturbances on subsequent frequency domain analysis can be reduced, and the interference of echoes on noise characterization can be decreased. This provides high-quality frequency domain input signals for subsequent differential array processing, wind noise detection and suppression, DSP noise reduction, AI noise reduction, and other algorithms. This processing method helps improve the accuracy of residual noise estimation, reduces the suppression burden of subsequent algorithms, and improves the clarity and naturalness of the final output speech, making it well-suited for resource-constrained portable electronic devices. The parameters, window function types, and adaptive filtering algorithms mentioned above are for illustrative purposes only. In practical applications, other parameter configurations or algorithm implementations can be selected, and this application does not limit these options.
[0191] Similarly, if an electronic device has only a single audio acquisition component, such as a single-microphone speech noise reduction system, then acquiring the frequency domain noise signal corresponding to the original audio signal acquired by the audio acquisition component includes: acquiring the original audio signal acquired by the audio acquisition component; performing front-end noise reduction processing on the original audio signal to obtain a third audio signal, wherein the front-end noise reduction processing includes anti-aliasing filtering and / or oversampling processing; performing a short-time Fourier transform on the third audio signal to obtain a fourth audio signal; and performing linear echo cancellation processing on the fourth audio signal to obtain the frequency domain noise signal corresponding to the original audio signal acquired by the audio acquisition component.
[0192] Based on the above embodiments, in some embodiments, the hybrid noisy speech library further includes noisy speech obtained after wind noise suppression processing.
[0193] Understandably, for electronic devices with dual audio acquisition components, the training samples of the lightweight noise reduction model deployed in the electronic device contain noisy speech obtained after wind noise suppression processing.
[0194] In this embodiment, "noisy speech obtained after wind noise suppression processing" refers to training samples formed after the original noisy speech has undergone wind noise detection and suppression processing. While retaining the main content of the target speech, the training samples may still contain varying degrees of residual wind noise, airflow disturbance sound, and slight distortion components introduced by the suppression algorithm. The wind noise suppression processing here can be implemented as described in the aforementioned embodiments to weaken the masking effect of low-frequency airflow impact on speech, thereby obtaining noisy speech samples for library construction. To adapt the lightweight noise reduction model to the audio distribution after pre-stage wind noise processing in real terminals, the hybrid noisy speech library can organize these samples together with clean speech, noisy speech after the first noise reduction processing (DSP noise reduction), and noisy speech without the first noise reduction processing (DSP noise reduction) into a training dataset, and can cover different wind speed conditions in terms of sample ratio. In practical applications, such samples can be generated by a laboratory fan that outputs wind noise at different wind speeds according to the anemometer settings. Then, the noise is collected by a manually placed earphone microphone. The pure voice is mixed with the wind noise at different wind speeds to generate an audio signal with wind noise. Finally, the results are combined with the output of the wind noise suppression module and stored. This application does not limit this process.
[0195] In practical implementation, after the noisy speech obtained through wind noise suppression is incorporated into the hybrid noisy speech library, the lightweight noise reduction model can learn the spectral distribution characteristics of the residual noise after wind noise suppression, as well as the spectral changes caused by the previous processing, thereby generating noise suppression factors more accurately in subsequent inference. This approach enables the model to handle not only conventional steady-state and non-steady-state noise, but also the residual noise state in wind noise scenarios, reducing over-suppression and muffled speech, and improving the model's stability and generalization ability in mobile call scenarios.
[0196] Next, combined Figure 2 The training process and inference application process of the lightweight noise reduction model are illustrated by example.
[0197] Figure 2 This diagram illustrates the training and inference processes of the lightweight noise reduction model provided in this embodiment. Figure 2 As shown in the diagram, the upper dashed box represents the model training part, and the lower dashed box represents the application part (i.e., the model inference noise reduction process). These will be explained separately below.
[0198] The model training section includes the following steps:
[0199] Training sample construction: Clean speech (i.e., noisy speech) and noise signals are linearly superimposed according to multiple preset signal-to-noise ratios to obtain paired noisy speech signals; each training sample contains noisy speech (including noisy speech data after the first noise reduction process and noisy speech without the first noise reduction process) and the corresponding clean speech (as ground truth label), forming a training data pair for supervised learning.
[0200] Time-frequency decomposition: Perform short-time Fourier transform (STFT) on the noisy speech signal to obtain the complex spectrum of the noisy speech; extract the amplitude spectrum |Y| of the noisy speech; perform the same STFT processing on the clean speech signal to extract the amplitude spectrum |X| of the clean speech, which is used as the ground truth label for network training.
[0201] Feature extraction: The amplitude spectrum |Y| of the noisy speech is extracted, and the Mel frequency cepstral coefficients (MFCC) are preferred as acoustic features. The purpose of feature extraction is to reduce dimensionality, compressing high-dimensional spectral data into low-dimensional feature vectors to reduce the number of parameters in the model input layer.
[0202] Deep learning network training: The extracted acoustic features are input into the deep learning network. The network learns and predicts the time-frequency mask (TF Mask). The time-frequency mask is a gain matrix with the same dimension as the input spectrum. The time-frequency mask is multiplied by the amplitude spectrum |Y| of the noisy speech to obtain the estimated amplitude of the denoised speech.
[0203] Loss calculation and weight iteration update: Calculate the loss between the estimated amplitude of the denoised speech and the amplitude spectrum |X| of the real clean speech. Use the backpropagation algorithm to backpropagate the loss gradient to each layer of the network. Use an optimizer (such as Adam) to adjust the network weight parameters. Iterate repeatedly until the model loss converges, and finally obtain the trained lightweight denoising model.
[0204] This lightweight noise reduction model can be deployed to embedded devices for inference applications.
[0205] The trained lightweight noise reduction model can be applied to real-world speech noise reduction scenarios. Its inference process is as follows:
[0206] Input preprocessing: The raw audio signal (noisy speech signal) acquired by the audio acquisition component is obtained. After anti-aliasing filtering and / or oversampling processing of the raw audio signal, STFT is performed to extract its amplitude spectrum |Y| and phase spectrum ∠Y. The phase spectrum is temporarily stored for subsequent waveform reconstruction.
[0207] Linear echo cancellation: Linear echo cancellation is performed on the amplitude spectrum |Y| to obtain the frequency domain noise frequency signal corresponding to the original audio signal.
[0208] DSP noise reduction processing: The first noise reduction processing is performed on the frequency domain band noise signal to obtain the first audio signal, denoted as amplitude spectrum |Y|'.
[0209] Feature extraction: Perform the same feature extraction (e.g., MFCC) on the noisy amplitude spectrum |Y|' as in the training phase.
[0210] Model inference: The extracted features are input into the trained lightweight denoising model. The model performs forward calculation and outputs the predicted time-frequency mask. After quantization (i.e., noise suppression factor), it is used as a gain to apply to the noisy amplitude spectrum |Y|' to obtain the denoised amplitude spectrum.
[0211] Time-domain waveform reconstruction: The amplitude spectrum after noise reduction is combined with the previously stored phase spectrum ∠Y, and the noise-reduced clean speech in the time domain (i.e., the noise-reduced speech signal in the time domain) is reconstructed through inverse short-time Fourier transform.
[0212] Next, through Figures 3-4 This paper provides an exemplary description of speech noise reduction methods and their effects in scenarios with a single audio acquisition component.
[0213] Figure 3 This is a flowchart illustrating the speech denoising method for a single audio acquisition component scenario provided in this application embodiment. Taking a microphone as the audio acquisition component, the method employs a two-stage cascaded denoising architecture of DSP denoising and AI denoising. The input is the original speech signal (i.e., the original audio signal) acquired by a single microphone. The signal sequentially undergoes time-frequency transformation, linear echo cancellation, traditional DSP denoising processing (including MCRA noise estimation, combined with OM-LSA or Wiener filtering, etc.), feature extraction and AI model inference, post-processing of Voice Activity Detection (VAD), and inverse time-frequency transformation, finally outputting the denoised time-domain speech signal corresponding to the original speech signal.
[0214] Specifically, such as Figure 3 As shown, the speech noise reduction method includes the following steps:
[0215] 3.1 Voice input.
[0216] The raw audio signal is acquired using a single microphone. This signal contains the target speech, ambient noise (steady-state and non-steady-state noise), and possible echo components.
[0217] 3.2 FFT processing.
[0218] Before performing FFT processing on the original audio signal, anti-aliasing filtering and / or oversampling are first applied to filter out high-frequency noise and reduce the computational load of subsequent algorithms. Then, Fourier transform is performed on the pre-filtered audio signal (i.e., the third audio signal) to convert the time-domain signal to the frequency domain, resulting in a frequency-domain complex signal (i.e., the fourth audio signal, which includes amplitude and phase spectra).
[0219] 3.3 Linear AEC processing.
[0220] Linear echo cancellation (AEC) is performed by using an adaptive filtering algorithm (such as LMS or NLMS) to establish a linear model of the echo path, estimating and canceling the linear echo components in the fourth audio signal to reduce echo interference generated by far-end speech through the speaker-microphone path, and finally obtaining the frequency domain noise frequency signal corresponding to the original audio signal.
[0221] 3.4 MCRA noise estimation.
[0222] An improved MCRA noise estimation method is used to perform noise estimation on the frequency domain band noise signal to obtain the noise power spectrum, which can be used for OM-LSA or Wiener filtering.
[0223] 3.5. Combine with OM-LSA or Wiener filtering.
[0224] By combining the noise power spectrum estimated by MCRA, the Wiener filter gain is calculated, and noise reduction processing is performed on the frequency domain noisy signal to suppress steady-state noise and some non-steady-state noise, outputting the first audio signal after the first stage of noise reduction. Alternatively, based on the obtained noise power spectrum, OM-LSA can be used to perform noise reduction processing on the frequency domain noisy signal, outputting the first audio signal after the first stage of noise reduction.
[0225] 3.6 Feature extraction (e.g., MFCC parameters).
[0226] Time-frequency features are extracted from the first audio signal. Mel frequency cepstral coefficients (MFCC) are preferred as acoustic features to compress high-dimensional spectral data into low-dimensional feature vectors, thereby reducing the input dimension of subsequent AI models.
[0227] 3.7 AI Model Inference.
[0228] The extracted acoustic features are input into a pre-trained lightweight noise reduction model. The model adopts a micro neural network structure (such as 2-layer DNN, DNN+GRU, etc.), and the number of parameters is compressed to less than 30KB. The model outputs a noise suppression factor to further suppress residual steady-state noise and non-steady-state noise in the first audio signal.
[0229] 3.8 Second noise reduction process.
[0230] The noise suppression factor output by the AI model is multiplied by the amplitude spectrum of the first audio signal to obtain the noise-reduced amplitude spectrum. The phase spectrum remains unchanged.
[0231] 3.9 VAD Post-processing.
[0232] The amplitude spectrum after noise reduction is subjected to joint adaptive detection based on short-time energy and entropy to accurately identify speech activity: when a speech segment is detected, the noise suppression intensity is automatically reduced to preserve high-frequency details and emotional information of the speech; when a non-speech segment is detected, the suppression intensity is maintained or increased to reduce noise residue. See steps 1042-1 to 1042-2 for details.
[0233] Then, the second audio signal, optimized by VAD, is output.
[0234] 3.10 IFFT processing.
[0235] The amplitude spectrum corresponding to the second audio signal is combined with the phase spectrum corresponding to the frequency domain noise frequency signal to obtain the frequency domain denoised speech signal corresponding to the original audio signal; the frequency domain denoised speech signal is then subjected to time-frequency transformation to obtain the time domain denoised speech signal.
[0236] 3.11 Output.
[0237] The output time-domain noise-reduced speech signal has high clarity and good naturalness, which can meet the low latency requirements of real-time voice calls.
[0238] Furthermore, based on the algorithm design, the embodiments of this application further improve noise reduction performance and system efficiency through hardware co-optimization. Specifically, this is manifested in the following aspects:
[0239] Hardware acceleration unit: It adopts fixed-point computing to replace floating-point operations, which significantly improves computing efficiency and reduces power consumption, providing hardware-level acceleration support for real-time voice processing;
[0240] Streaming architecture: It adopts a "real-time input-real-time processing-real-time output" streaming architecture, which does not require caching the entire audio segment. The latency of the entire system is controlled within one frame, making it suitable for low-memory embedded devices.
[0241] Dynamic power consumption regulation: Through Dynamic Voltage and Frequency Scaling (DVFS) technology, the chip's main frequency is reduced during non-voice periods to control system power consumption and meet the battery life requirements of portable devices.
[0242] Figure 4The images show a comparison of speech noise reduction effects in a single audio acquisition component scenario provided in this application embodiment. Taking a microphone as the audio acquisition component as an example, as... Figure 4 As shown, the noise reduction effect of the single-microphone noise reduction algorithm of this application embodiment is compared under eight typical noise environments, and the changes in speech signal before and after noise reduction are presented in the form of spectrogram.
[0243] Specifically, Figure 4 The test includes eight noise scenarios, from left to right: bar noise, road noise, intersection noise, train noise, car noise, coffee shop noise, restaurant noise, and call center noise. For each noise scenario, two rows of spectrograms are displayed: before and after noise reduction. The horizontal axis of the spectrogram represents time, the vertical axis represents frequency, and the color intensity indicates energy level. Figure 4 As can be seen, before noise reduction, the noise energy distribution in the spectrogram is widespread, and the speech energy is masked by noise, making it difficult to clearly identify the speech content. After processing by the speech noise reduction method of this application embodiment, the background noise energy is significantly suppressed, the harmonic structure of the speech frequency band (especially the mid-low frequency band) is clearly visible, and the speech intelligibility and clarity are significantly improved. The speech noise reduction method of this application embodiment shows stable and effective noise reduction effects in the above eight typical noise scenarios, and can adapt to a variety of complex acoustic environments, from steady noise (such as roads, trains, and cars) to non-steady noise (such as bars, cafes, and call centers).
[0244] It should be noted that the above eight noise types are typical noise scenarios used in objective testing based on the ACQUA 3Quest testing platform, and are only used as examples to illustrate the noise reduction effect of the algorithm in this application. In practical applications, the noise reduction algorithm of this application embodiment is not limited to suppressing the above eight noise types. Its training data covers dozens of real-world scene noises (including but not limited to bars, roads, intersections, trains, cars, cafes, restaurants, call centers, wind noise, etc.), so it can effectively suppress a wider range of real-world environmental noises and has good generalization ability.
[0245] Next, through Figures 5-6 This paper provides an exemplary description of speech noise reduction methods and their effects in scenarios involving dual-audio acquisition components.
[0246] Figure 5This is a flowchart illustrating the speech noise reduction method for a dual-audio acquisition component scenario provided in this application embodiment. Taking a microphone as the audio acquisition component, the method employs a noise reduction architecture that cascades dual-microphone spatial filtering with single-channel environmental noise cancellation (ENC). The input is the original speech signal acquired by two microphones (Mic0 and Mic1), which undergoes time-frequency transformation and linear echo cancellation, then passes through a first-order differential array and wind noise suppression via a wind noise detection and processing module. Finally, it is sent to a single-channel ENC module to perform "DSP noise reduction + AI noise reduction," outputting a denoised time-domain speech signal corresponding to the original speech signal.
[0247] Specifically, such as Figure 5 As shown, the speech noise reduction method includes the following steps:
[0248] 5.1 Voice input.
[0249] The original audio signals were captured separately using dual microphones.
[0250] 5.2 FFT processing.
[0251] Before performing FFT processing on the original audio signals, anti-aliasing filtering and / or oversampling are performed on each original audio signal to filter out high-frequency noise and reduce the computational load of subsequent algorithms. Then, Fourier transform is performed on the pre-filtered audio signal (i.e., the third audio signal) to convert the time-domain signal to the frequency domain, resulting in a frequency-domain complex signal (i.e., the fourth audio signal, which includes amplitude and phase spectra).
[0252] 5.3 Linear AEC processing.
[0253] Linear echo cancellation (AEC) is performed, which involves using an adaptive filtering algorithm (such as LMS or NLMS) to establish a linear model of the echo path, estimating and canceling the linear echo components in the fourth audio signal, thereby reducing echo interference generated by far-end speech through the speaker-microphone path, and finally obtaining the frequency domain signal with noise corresponding to the original audio signal. That is, the first frequency domain signal corresponding to Mic0 and the second frequency domain signal corresponding to Mic2.
[0254] 5.4 First-order difference array.
[0255] The first frequency domain signal and the second frequency domain signal are input into the first-order differential array module for first-order differential array processing to obtain the differential array signal. That is, the two signals are selected for audio directionality using the directivity of the array, and finally a single signal is obtained.
[0256] 5.5 Wind noise detection.
[0257] Based on the spatial difference characteristics (such as sum-difference ratio, centroid spectrum, etc.) between the first frequency domain signal and the second frequency domain signal, determine whether there is wind noise in the frequency domain band noise signal (or the first frequency domain signal and the second frequency domain signal).
[0258] The branching process is performed based on the results of wind noise detection. If wind noise is present, step 5.6 is executed; otherwise, step 5.6 is skipped and step 5.7 is executed directly, which treats the differential array signal as a frequency domain band noise signal.
[0259] 5.6 Wind noise suppression treatment.
[0260] Wind noise suppression can be achieved using the Wiener filtering method. The wind noise suppression factor is determined based on wind noise energy estimation, and the differential array signal undergoes spectral post-processing. The signal after wind noise suppression is then treated as a frequency-domain band-noise signal.
[0261] 5.7 Single-channel ENC (MCRA + OM-LSA or Wiener filtering + feature extraction + AI model inference + second noise reduction processing).
[0262] Similar to steps 3.4 to 3.8, they will not be repeated here.
[0263] 5.8 VAD Post-processing.
[0264] The amplitude spectrum after noise reduction is subjected to joint adaptive detection based on short-time energy and entropy to accurately identify speech activity: when a speech segment is detected, the noise suppression intensity is automatically reduced to preserve high-frequency details and emotional information of the speech; when a non-speech segment is detected, the suppression intensity is maintained or increased to reduce noise residue. See steps 1042-1 to 1042-2 for details.
[0265] Then, the second audio signal, optimized by VAD, is output.
[0266] 5.9 IFFT.
[0267] The amplitude spectrum corresponding to the second audio signal is combined with the phase spectrum corresponding to the frequency domain noise frequency signal to obtain the frequency domain denoised speech signal corresponding to the original audio signal; the frequency domain denoised speech signal is then subjected to time-frequency transformation to obtain the time domain denoised speech signal.
[0268] 5.10 Output.
[0269] The output time-domain noise-reduced speech signal has high clarity and good naturalness, which can meet the low latency requirements of real-time voice calls.
[0270] Figure 6 The images show a comparison of speech noise reduction effects in a dual-audio acquisition component scenario provided in this application embodiment. Taking a microphone as the audio acquisition component as an example, as... Figure 6As shown, the noise reduction effect of the dual-microphone noise reduction algorithm of this application embodiment is compared in eight typical noise environments. The changes in speech signals before and after noise reduction by dual microphones (Mic0 and Mic1) are presented in the form of spectrograms.
[0271] Specifically, Figure 6 The test includes eight noise scenarios, from left to right: bar noise, road noise, intersection noise, train noise, car noise, coffee shop noise, restaurant noise, and call center noise. For each noise scenario, three rows of spectrograms are displayed: before noise reduction (Mic0), before noise reduction (Mic1), and after noise reduction. The horizontal axis of the spectrogram represents time, the vertical axis represents frequency, and the color intensity indicates energy intensity. Figure 6 As can be seen, before noise reduction, the noise energy distribution in the spectrogram is widespread, and the speech energy is masked by noise, making it difficult to clearly identify the speech content. After processing by the speech noise reduction method of this application embodiment, the background noise energy is significantly suppressed, and the harmonic structure of the speech frequency band (especially the mid-to-low frequency band) is clearly visible. The speech noise reduction method of this application embodiment shows stable and effective noise reduction effects in the above eight typical noise scenarios.
[0272] To further illustrate the technical effects of the speech noise reduction method provided in the embodiments of this application, the following description is provided in conjunction with Tables 1 and 2.
[0273] Table 1 shows the objective performance comparison results of the speech denoising methods provided in this application (i.e., the single-microphone speech denoising algorithm and the dual-microphone speech denoising algorithm in the table) with competing algorithms on the 3Quest testing platform. As shown in Table 1, the single-microphone speech denoising algorithm and the dual-microphone speech denoising algorithm proposed in this application were compared with competing algorithms in three dimensions: Speech Mean Opinion Score (SMOS), Noise Mean Opinion Score (NMOS), and Global Mean Opinion Score (GMOS). Among them, SMOS reflects the clarity, naturalness, and overall listening experience of the denoised speech; a higher score indicates better speech quality. NMOS reflects the algorithm's ability to suppress background noise and the listening experience of residual noise; a higher score indicates better noise reduction effect. GMOS comprehensively reflects the overall performance of speech quality and noise reduction effect; a higher score indicates better overall algorithm performance.
[0274] Table 1
[0275]
[0276] As can be seen from the objective test results in Table 1, the speech noise reduction method proposed in this application has reached or exceeded the level of competing products in terms of objective test indicators. In particular, the dual-microphone scheme has outstanding performance in noise suppression and can meet the needs of high-quality speech communication in complex acoustic environments.
[0277] Table 2 shows the resource usage comparison data of the speech denoising methods provided in the embodiments of this application on the embedded platform. As shown in Table 2, the single-microphone speech denoising algorithm occupies 107KB of RAM resources, 48KB of Flash resources, and 72M of MIPS resources; the dual-microphone speech denoising algorithm occupies 170KB of RAM resources, 56KB of Flash resources, and 124M of MIPS resources.
[0278] Table 2
[0279]
[0280] As shown in Table 2, the single-microphone and dual-microphone voice noise reduction algorithms proposed in this application exhibit low resource consumption in terms of RAM, Flash, and MIPS, effectively adapting to resource-constrained embedded devices. Specifically, the single-microphone voice noise reduction algorithm, with a resource footprint of 107KB RAM, 48KB Flash, and 72M MIPS, can run in real time on a low-cost MCU. Although the dual-microphone voice noise reduction algorithm experiences a slight increase in resource consumption due to the addition of modules such as wind noise detection and processing and a first-order differential array (RAM 170KB, Flash 56KB, MIPS 124M), it remains within the resource budget of typical embedded devices such as TWS earphones.
[0281] The aforementioned resource utilization level fully meets the requirements of real-time voice calls for low latency and low resource consumption, while leaving sufficient resource margin for other system tasks, demonstrating good feasibility for mass production.
[0282] Figure 7 This is a schematic diagram of the speech noise reduction device provided in the embodiments of this application, as shown below. Figure 7 As shown, the speech noise reduction device 70 includes: an acquisition module 71, a first noise reduction module 72, an inference module 73, a second noise reduction module 74, and a reconstruction module 75. Wherein:
[0283] The acquisition module 71 is used to acquire the frequency domain noise frequency signal corresponding to the original audio signal acquired by the audio acquisition component;
[0284] The first noise reduction module 72 is used to perform a first noise reduction process on the frequency domain band noise signal to obtain a first audio signal. The residual noise contained in the first audio signal is dominated by non-steady-state noise.
[0285] The inference module 73 is used to extract time-frequency features from the first audio signal to obtain the acoustic features corresponding to the first audio signal, and input the acoustic features into the pre-trained lightweight noise reduction model. The lightweight noise reduction model infers the acoustic features and outputs a noise suppression factor. The lightweight noise reduction model is trained based on a hybrid noisy speech library. The hybrid noisy speech library contains at least non-noisy speech and noisy speech after the first noise reduction process.
[0286] The second noise reduction module 74 is used to perform a second noise reduction process on the first audio signal according to the noise suppression factor to obtain a second audio signal;
[0287] The reconstruction module 75 is used to reconstruct the time-domain denoised speech signal corresponding to the original audio signal based on the second audio signal.
[0288] In one possible implementation, the second noise reduction module 74 is specifically used to: multiply the noise suppression factor by the amplitude spectrum of the first audio signal to obtain the noise-reduced amplitude spectrum; and determine the second audio signal based on the noise-reduced amplitude spectrum.
[0289] In one possible implementation, the second noise reduction module 74 is further configured to: determine the corresponding energy entropy ratio based on the noise-reduced amplitude spectrum, wherein the energy entropy ratio represents the probability of speech presence; if the energy entropy ratio is less than the energy entropy ratio threshold and the duration is greater than or equal to a preset duration, then it is determined that there is no speech activity, and the noise-reduced amplitude spectrum is multiplied by a non-speech segment noise suppression factor to obtain a second audio signal, wherein the non-speech segment suppression factor is used to suppress residual noise in the non-speech segment.
[0290] In one possible implementation, the second noise reduction module 74 is further configured to: determine the noise-reduced amplitude spectrum as the second audio signal if the energy entropy ratio is greater than or equal to the energy entropy ratio threshold, or the duration is less than a preset duration.
[0291] In one possible implementation, the speech noise reduction method is applied to an electronic device, which includes a first audio acquisition component and a second audio acquisition component. The acquisition module 71 is specifically used to: acquire a first frequency domain signal corresponding to the original audio signal acquired by the first audio acquisition component, and a second frequency domain signal corresponding to the original audio signal acquired by the second audio acquisition component; perform first-order differential array processing on the first and second frequency domain signals to obtain a differential array signal; determine whether wind noise exists in the frequency domain band noise signal based on the spatial difference characteristics between the first and second frequency domain signals; if wind noise is determined to exist, determine the corresponding wind noise suppression factor based on the spatial difference characteristics; perform wind noise suppression processing on the differential array signal based on the wind noise suppression factor to obtain a wind noise-suppressed signal; and use the wind noise-suppressed signal as the frequency domain band noise signal.
[0292] In one possible implementation, the acquisition module 71 is further configured to: if it is determined that there is no wind noise, then treat the differential array signal as a frequency domain band noise signal.
[0293] In one possible implementation, the hybrid noisy speech library further includes noisy speech obtained after wind noise suppression processing.
[0294] In one possible implementation, both the first frequency domain signal and the second frequency domain signal are determined by: acquiring the original audio signal acquired by the corresponding audio acquisition component; performing front-end noise reduction processing on the original audio signal to obtain the third audio signal, wherein the front-end noise reduction processing includes anti-aliasing filtering and / or oversampling processing; performing a short-time Fourier transform on the third audio signal to obtain the fourth audio signal; and performing linear echo cancellation processing on the fourth audio signal to obtain the frequency domain signal corresponding to the original audio signal acquired by the corresponding audio acquisition component, wherein the linear echo cancellation processing is implemented using an adaptive filtering algorithm to estimate and cancel the linear echo components in the fourth audio signal.
[0295] In one possible implementation, the reconstruction module 75 is specifically used to: combine the amplitude spectrum corresponding to the second audio signal with the phase spectrum corresponding to the frequency domain noise frequency signal to obtain the frequency domain denoised speech signal corresponding to the original audio signal; and perform time-frequency transformation on the frequency domain denoised speech signal to obtain the time domain denoised speech signal.
[0296] In one possible implementation, the first noise reduction module 72 is specifically used to: perform improved MCRA noise estimation on the frequency domain band noise signal to obtain a noise power spectrum; wherein, the improved MCRA noise estimation uses a recursive method to continuously update the minimum search window instead of resetting the window every frame, and uses a two-stage recursive average estimation of the speech presence probability to support a smaller search window length; based on the noise power spectrum, OM-LSA or Wiener filtering is used to perform noise reduction processing on the frequency domain band noise signal to obtain a first audio signal.
[0297] The speech noise reduction device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0298] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 80 provided in this embodiment includes at least one processor 801 and a memory 802. Optionally, the electronic device 80 further includes a communication component 803. The processor 801, memory 802, and communication component 803 are connected via a bus 804.
[0299] In a specific implementation, at least one processor 801 executes computer execution instructions stored in memory 802, causing at least one processor 801 to perform the above-described method.
[0300] The specific implementation process of processor 801 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0301] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0302] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0303] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0304] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0305] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0306] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0307] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0308] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0309] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0310] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0311] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0312] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0313] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A speech noise reduction method, characterized in that, include: Obtain the frequency domain noise frequency signal corresponding to the original audio signal acquired by the audio acquisition component; The frequency domain band noise signal is subjected to a first noise reduction process to obtain a first audio signal, wherein the residual noise contained in the first audio signal is dominated by non-steady-state noise. The first audio signal is subjected to time-frequency feature extraction to obtain the acoustic features corresponding to the first audio signal, and the acoustic features are input into a pre-trained lightweight noise reduction model. The lightweight noise reduction model is used to infer the acoustic features and output a noise suppression factor. The lightweight noise reduction model is trained based on a hybrid noisy speech library, which includes at least non-noisy speech and noisy speech after the first noise reduction process. According to the noise suppression factor, the first audio signal is subjected to a second noise reduction process to obtain a second audio signal; Based on the second audio signal, the time-domain denoised speech signal corresponding to the original audio signal is reconstructed.
2. The speech noise reduction method according to claim 1, characterized in that, The step of performing a second noise reduction process on the first audio signal according to the noise suppression factor to obtain a second audio signal includes: The noise suppression factor is multiplied by the amplitude spectrum of the first audio signal to obtain the noise-reduced amplitude spectrum; The second audio signal is determined based on the amplitude spectrum after noise reduction.
3. The speech noise reduction method according to claim 2, characterized in that, Determining the second audio signal based on the noise-reduced amplitude spectrum includes: Based on the amplitude spectrum after noise reduction, the corresponding energy entropy ratio is determined, whereby the energy entropy ratio represents the probability of the presence of speech. If the energy entropy ratio is less than the energy entropy ratio threshold and the duration is greater than or equal to the preset duration, it is determined that there is no speech activity, and the amplitude spectrum after noise reduction is multiplied by the non-speech segment noise suppression factor to obtain the second audio signal. The non-speech segment suppression factor is used to suppress residual noise in the non-speech segment.
4. The speech noise reduction method according to claim 3, characterized in that, Also includes: If the energy entropy ratio is greater than or equal to the energy entropy ratio threshold, or the duration is less than the preset duration, then the amplitude spectrum after noise reduction is determined as the second audio signal.
5. The speech noise reduction method according to any one of claims 1 to 4, characterized in that, Applied to electronic devices, wherein the electronic devices are provided with a first audio acquisition component and a second audio acquisition component; The acquisition of the frequency domain noise frequency signal corresponding to the original audio signal acquired by the audio acquisition component includes: Acquire the first frequency domain signal corresponding to the original audio signal acquired by the first audio acquisition component, and the second frequency domain signal corresponding to the original audio signal acquired by the second audio acquisition component; The first frequency domain signal and the second frequency domain signal are processed by a first-order differential array to obtain a differential array signal; Based on the spatial difference characteristics between the first frequency domain signal and the second frequency domain signal, it is determined whether the frequency domain band noise signal contains wind noise; If wind noise is determined to exist, the corresponding wind noise suppression factor is determined based on the spatial difference characteristics. Based on the wind noise suppression factor, the differential array signal is subjected to wind noise suppression processing to obtain the signal after wind noise suppression; The signal after wind noise suppression is used as the frequency domain band noise signal.
6. The speech noise reduction method according to claim 5, characterized in that, Also includes: If it is determined that there is no wind noise, then the differential array signal is used as the frequency domain band noise frequency signal.
7. The speech noise reduction method according to claim 5, characterized in that, The hybrid noisy speech library also includes noisy speech obtained after wind noise suppression processing.
8. The speech noise reduction method according to claim 5, characterized in that, Both the first frequency domain signal and the second frequency domain signal are determined in the following way: Obtain the raw audio signal acquired by the corresponding audio acquisition component; The original audio signal is subjected to front-end noise reduction processing to obtain a third audio signal. The front-end noise reduction processing includes anti-aliasing filtering and / or oversampling processing. The third audio signal is subjected to a short-time Fourier transform to obtain the fourth audio signal; The fourth audio signal is subjected to linear echo cancellation processing to obtain the frequency domain signal corresponding to the original audio signal acquired by the corresponding audio acquisition component. The linear echo cancellation processing is implemented using an adaptive filtering algorithm to estimate and cancel the linear echo component in the fourth audio signal.
9. The speech noise reduction method according to any one of claims 1 to 4, characterized in that, The process of reconstructing the temporal-domain denoised speech signal corresponding to the original audio signal based on the second audio signal includes: The amplitude spectrum corresponding to the second audio signal is combined with the phase spectrum corresponding to the frequency domain band noise frequency signal to obtain the frequency domain noise-reduced speech signal corresponding to the original audio signal; The frequency-domain denoised speech signal is subjected to time-frequency transformation to obtain a time-domain denoised speech signal.
10. The speech denoising method according to any one of claims 1 to 4, characterized in that, The first noise reduction process on the frequency domain noisy signal to obtain the first audio signal includes: An improved MCRA noise estimation is performed on the frequency domain band noise signal to obtain the noise power spectrum; wherein, the improved MCRA noise estimation adopts a recursive method to continuously update the minimum search window instead of resetting the window every frame, and adopts a two-stage recursive average estimation of the speech presence probability to support a smaller search window length. Based on the noise power spectrum, the frequency domain band noise signal is denoised using the Optimal Modified Logarithmic Spectral Amplitude Estimator (OM-LSA) or Wiener filtering to obtain the first audio signal.
11. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the speech noise reduction method as described in any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed, are used to implement the speech noise reduction method as described in any one of claims 1 to 10.