Communication voice noise reduction method based on AI
By collecting noise from multiple scenarios, constructing a database, and using AI technology to generate labeled samples, a shared encoder was designed to handle speech separation and noise suppression. Combined with multi-microphone arrays and beamforming technology, the problem of traditional communication voice noise reduction technology being unable to adapt in complex noise environments was solved, achieving efficient and low-latency voice noise reduction.
Patent Information
- Application Number
- CN202511032828.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional voice noise reduction technology cannot adapt to complex noise in real-world scenarios and cannot adaptively adjust the noise reduction intensity, resulting in poor noise reduction performance in various types of non-steady-state noise environments.
By collecting noise from multiple scenarios, a dialect speech database is constructed. Samples with three-dimensional labels are generated using a conditional adversarial network. A shared encoder is designed to process speech separation, noise suppression, and reverberation cancellation. Combined with multi-microphone arrays and beamforming technology, a dedicated AI acceleration chip is used to achieve low-latency noise reduction. Furthermore, a dynamic noise recognition module is used to adjust the noise reduction intensity, providing a customized noise reduction strategy.
It achieves adaptive adjustment of noise reduction intensity in complex noise environments, improves speech fidelity and model generalization ability, reduces latency, and adapts to the noise reduction needs of different scenarios.
Smart Images

Figure BDA0005517846220000021 
Figure BDA0005517846220000032 
Figure FDA0005517846210000011
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of intelligent communication, specifically to an AI-based method for reducing noise in communication voice. Background Technology
[0002] Voice communication refers to the exchange of information through sound. Common voice communication devices include mobile phones and walkie-talkies.
[0003] Traditional communication voice noise reduction techniques mainly rely on single-microphone signal processing (such as spectral subtraction and Wiener filtering) or dual-microphone sound source separation. However, their design is based on ideal signal assumptions (such as noise stationarity) and cannot adapt to complex noise in real-world scenarios. Traditional algorithms are based on manual rules or mathematical derivations and perform poorly on various types of non-steady-state noise in real-world scenarios (such as keyboard sounds in office scenarios and floor impact sounds in home scenarios). They cannot adaptively adjust the noise reduction intensity. Therefore, an AI-based communication voice noise reduction method is proposed. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] To address the shortcomings of existing technologies, this invention provides an AI-based method for noise reduction in communication voice, which has advantages such as adaptive adjustment of noise reduction intensity and solves the aforementioned problems.
[0006] (II) Technical Solution
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] An AI-based method for noise reduction in communication voice includes the following steps:
[0009] Step 1: By collecting noise from multiple scenes, the noisy speech is recorded through a microphone array, covering complex environments such as traffic, restaurants, and conferences;
[0010] Step 2: Construct a dialect speech database, add a noise mixing layer (such as sudden horn blasts, keyboard tapping) and a scene evolution layer to simulate a real environment;
[0011] Step 3: Randomly extract human voice, reverberation, and noise samples, and generate noisy speech that closely resembles the real scene through operations such as filtering and signal-to-noise ratio adjustment, thereby reducing the difficulty of supervised learning;
[0012] Step 4: Design a shared encoder to jointly process speech separation (task R1), noise suppression (task R2), and reverberation cancellation (task R3) to improve the model's generalization ability;
[0013] Step 5: Deploy the deep learning model on edge devices (such as the microphone array of smart communication devices) to achieve low-latency noise reduction through real-time processing. Use a dedicated AI acceleration chip (such as NR2049-P) to support local computing and complete noise suppression in milliseconds. This is suitable for scenarios such as TWS earphones and in-vehicle communication.
[0014] Step Six: The first-level noise reduction node targets high-frequency burst noise (such as keyboard sounds, glass breaking sounds), while the second-level noise reduction node processes residual low-frequency noise. Layered filtering improves voice fidelity.
[0015] Step 7: The dynamic noise recognition module distinguishes between mechanical noise, aerodynamic noise, and external environmental noise, and adjusts the noise reduction intensity to suit scenarios such as elevators, conference rooms, and vehicles. Combined with multi-microphone arrays and beamforming technology, it focuses on the direction of the target speaker and suppresses multi-source interference.
[0016] Step 8: Sampling point reconstruction: Convert the noise-reduced sampling points into waveforms to generate clean speech.
[0017] Preferably, step two further generates samples with three-dimensional labels in a conditional adversarial network (cGAN): noise type, signal-to-noise ratio, and speech features. A scene evolution layer, a sound change layer, and a noise mixing layer are added to simulate the speech features of different dialects in complex environments, expanding the training data. The sound change layer simulates speech distortion (such as echo and reverberation), improving the model's generalization ability. The frequency domain representation of noisy speech is: X(ω) = S(ω) + N(ω), where X(ω) is the Fourier transform of the noisy speech, S(ω) is clean speech, and N(ω) is noise. Noise power spectrum estimation (stationary noise assumption):
[0018] K represents the number of frames without audio.
[0019] Preferably, step four employs an end-to-end speech enhancement model, directly using sampling point information as input and output, thus avoiding distortion problems caused by spectrum manipulation.
[0020] Preferably, step seven further provides customized noise reduction strategies by learning user voiceprint characteristics. For example, business people may focus on suppressing keyboard noise, while meeting scenarios may enhance the clarity of multi-person conversations. This involves dynamically adjusting the noise reduction intensity and frequency band strategy based on the user's voiceprint, including dynamic parameter adjustment, identifying environmental noise types (such as white noise and human voice), and adjusting the noise reduction mode. The adaptive gain control formula is as follows:
[0021] (Ratio of noisy signal power to noise power) (Ratio of clean signal power to noise power).
[0022] Preferably, step eight further includes voice enhancement: enhancing the target speaker through voiceprint locking technology (such as eye tracking).
[0023] (III) Beneficial Effects
[0024] Compared with existing technologies, this invention provides an AI-based method for noise reduction in communication voice, which has the following beneficial effects:
[0025] 1. This invention collects noise from multiple scenarios, uses a microphone array to record noisy speech, covering complex environments such as traffic, restaurants, and conferences, constructs a dialect speech database, adds a noise mixing layer (e.g., sudden honking, keyboard tapping) and a scene evolution layer to simulate real-world environments, and generates samples with three-dimensional labels in a Conditional Generative Adversarial Network (cGAN): noise type, signal-to-noise ratio, and speech features. It then adds a scene evolution layer, a sound change layer, and a noise mixing layer to simulate the speech features of different dialects in complex environments, expands the training data, and uses the sound change layer to simulate speech distortion (e.g., echo, reverberation) to improve the model's generalization ability. Finally, it randomly extracts human voices, reverberation, and... Noise samples are filtered and signal-to-noise ratio adjusted to generate noisy speech that closely resembles real-world scenarios. A dedicated AI acceleration chip (such as NR2049-P) is used to support local computation. An end-to-end speech enhancement model is adopted, directly using sampling point information as input and output to avoid distortion caused by spectrum manipulation. The deep learning model is deployed on edge devices (such as microphone arrays of smart communication devices) to achieve low-latency noise reduction through real-time processing. A dynamic noise recognition module distinguishes between mechanical noise, aerodynamic noise, and external environmental noise, and adjusts the noise reduction intensity to adapt to scenarios such as elevators, conference rooms, and vehicles. Thus, this method can adaptively adjust the noise reduction intensity.
[0026] 2. This invention designs a shared encoder to jointly process speech separation (task R1), noise suppression (task R2), and reverberation cancellation (task R3), thereby improving the model's generalization ability. Furthermore, the first-level noise reduction node targets high-frequency burst noise (such as keyboard sounds and glass breaking sounds), while the second-level noise reduction node processes residual low-frequency noise. Layered filtering improves speech fidelity, and finally, sampling point reconstruction is performed: the noise-reduced sampling points are converted into waveforms to generate clean speech. Detailed Implementation
[0027] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] This invention relates to an AI-based method for noise reduction in communication voice, comprising the following steps:
[0029] Step 1: By collecting noise from multiple scenes, the noisy speech is recorded through a microphone array, covering complex environments such as traffic, restaurants, and conferences;
[0030] Step 2: Construct a dialect speech database, add a noise mixing layer (such as sudden horn honking, keyboard typing) and a scene evolution layer to simulate the real environment;
[0031] Step two further expands the training data by generating samples with three-dimensional labels in a Conditional Adversarial Network (cGAN): noise type, signal-to-noise ratio, and speech features. Scene evolution layer, sound change layer, and noise mixing layer are added to simulate the speech features of different dialects in complex environments, thus improving the model's generalization ability. The sound change layer is used to simulate speech distortion (such as echo and reverberation). The frequency domain representation of noisy speech is: X(ω) = S(ω) + N(ω), where X(ω) is the Fourier transform of the noisy speech, S(ω) is the clean speech, and N(ω) is the noise. Noise power spectrum estimation (based on the assumption of stationary noise):
[0032] K represents the number of frames without audio;
[0033] Step 3: Randomly extract human voice, reverberation, and noise samples, and generate noisy speech that closely resembles the real scene through operations such as filtering and signal-to-noise ratio adjustment, thereby reducing the difficulty of supervised learning;
[0034] Step 4: Design a shared encoder to jointly process speech separation (task R1), noise suppression (task R2), and reverberation cancellation (task R3) to improve the model's generalization ability. The shared encoder jointly processes tasks such as speech enhancement and noise suppression to improve computational efficiency. Combined with hardware optimization, it focuses on the target speech direction and suppresses background noise. An end-to-end speech enhancement model is adopted, which directly uses the sampling point information as input and output to avoid distortion problems caused by spectrum manipulation.
[0035] Step 5: Deploy the deep learning model on edge devices (such as the microphone array of smart communication devices) to achieve low-latency noise reduction through real-time processing. Use a dedicated AI acceleration chip (such as NR2049-P) to support local computing and complete noise suppression in milliseconds. This is suitable for scenarios such as TWS earphones and in-vehicle communication. Use lightweight models (such as RNNoise and CRN) to process locally on the terminal device with latency as low as milliseconds. The model only depends on current and historical frame data to avoid future information leakage and meet real-time communication requirements.
[0036] Step Six: The first-level noise reduction node targets high-frequency burst noise (such as keyboard sounds, glass breaking sounds), while the second-level noise reduction node processes residual low-frequency noise. Layered filtering improves voice fidelity.
[0037] Step 7: The dynamic noise recognition module distinguishes between mechanical noise, aerodynamic noise, and external environmental noise, adjusting the noise reduction intensity to suit scenarios such as elevators, conference rooms, and vehicles. Combining multi-microphone arrays and beamforming technology, it focuses on the direction of the target speaker, suppressing multi-source interference. Furthermore, by learning user voiceprint characteristics, it provides customized noise reduction strategies. For example, for business people, it focuses on suppressing keyboard noise, while in conference scenarios, it enhances the clarity of multi-person conversations. This includes dynamically adjusting the noise reduction intensity and frequency band strategy based on the user's voiceprint, including dynamic parameter adjustment, identifying environmental noise types (such as white noise and human voice), and adjusting the noise reduction mode. Its adaptive gain control formula is as follows: (Ratio of noisy signal power to noise power) (Ratio of pure signal power to noise power), dynamically identify noise types (such as burst noise and steady-state noise), adjust noise reduction strategies, distinguish human voice from background noise, preserve the naturalness of speech, and avoid the "robot voice" of traditional algorithms.
[0038] Step 8: Sampling point reconstruction: Convert the noise-reduced sampling points into waveforms to generate clean speech, which also includes voice enhancement: enhance the target speaker through voiceprint locking technology (such as eye tracking).
[0039] The beneficial effects of this invention are as follows: This invention collects noise from multiple scenarios, records noisy speech through a microphone array, covering complex environments such as traffic, restaurants, and conferences, constructs a dialect speech database, adds a noise mixing layer (such as sudden honking or keyboard tapping) and a scene evolution layer to simulate real-world environments, generates samples with three-dimensional labels in a Conditional Generative Adversarial Network (cGAN): noise type, signal-to-noise ratio, and speech features, adds a scene evolution layer, a sound change layer, and a noise mixing layer to simulate the speech features of different dialects in complex environments, expands the training data, uses the sound change layer to simulate speech distortion (such as echo and reverberation) to improve the model's generalization ability, then randomly extracts human voice, reverberation, and noise samples, and generates noisy speech that closely resembles real-world scenarios through filtering, signal-to-noise ratio adjustment, and other operations, using a dedicated AI acceleration chip (such as NR2049-P) to support local computation, and adopts an end-to-end speech processing approach. The audio enhancement model directly uses sampling point information as input and output, avoiding distortion caused by spectrum manipulation. The deep learning model is deployed on edge devices (such as microphone arrays in smart communication devices) to achieve low-latency noise reduction through real-time processing. A dynamic noise recognition module distinguishes between mechanical noise, aerodynamic noise, and external environmental noise, adjusting the noise reduction intensity to adapt to scenarios such as elevators, conference rooms, and vehicles. A shared encoder is designed to jointly process speech separation (task R1), noise suppression (task R2), and reverberation cancellation (task R3), improving the model's generalization ability. Furthermore, the first-level noise reduction node targets high-frequency burst noise (such as keyboard clicks and glass shattering sounds), while the second-level noise reduction node handles residual low-frequency noise. Layered filtering improves speech fidelity. Finally, sampling point reconstruction is performed: the denoised sampling points are converted into waveforms to generate clean speech, thus allowing the method to adaptively adjust the noise reduction intensity.
[0040] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An AI-based method for noise reduction in communication voice, characterized in that, Includes the following steps: Step 1: By collecting noise from multiple scenes, the noisy speech is recorded through a microphone array, covering complex environments such as traffic, restaurants, and conferences; Step 2: Construct a dialect speech database, add a noise mixing layer (such as sudden horn honking, keyboard typing) and a scene evolution layer to simulate the real environment; Step 3: Randomly extract human voice, reverberation, and noise samples, and generate noisy speech that closely resembles the real scene through operations such as filtering and signal-to-noise ratio adjustment, thereby reducing the difficulty of supervised learning; Step 4: Design a shared encoder to jointly process speech separation (task R1), noise suppression (task R2), and reverberation cancellation (task R3) to improve the model's generalization ability; Step 5: Deploy the deep learning model on edge devices (such as the microphone array of smart communication devices) to achieve low-latency noise reduction through real-time processing. Use a dedicated AI acceleration chip to support local computing and complete noise suppression in milliseconds. Suitable for scenarios such as TWS earphones and in-vehicle communication. Step Six: The first-level noise reduction node targets high-frequency burst noise (such as keyboard sounds, glass breaking sounds), while the second-level noise reduction node processes residual low-frequency noise. Layered filtering improves voice fidelity. Step 7: The dynamic noise recognition module distinguishes between mechanical noise, aerodynamic noise, and external environmental noise, and adjusts the noise reduction intensity to suit scenarios such as elevators, conference rooms, and vehicles. Combined with multi-microphone arrays and beamforming technology, it focuses on the direction of the target speaker and suppresses multi-source interference. Step 8: Sampling point reconstruction: Convert the noise-reduced sampling points into waveforms to generate clean speech.
2. The AI-based voice noise reduction method according to claim 1, characterized in that: Step two further expands the training data by generating samples with three-dimensional labels in a Conditional Adversarial Network (cGAN): noise type, signal-to-noise ratio, and speech features. It adds a scene evolution layer, a voice change layer, and a noise mixing layer to simulate the speech features of different dialects in complex environments. The frequency domain representation of noisy speech is: X(ω)=S(ω)+N(ω), where X(ω) is the Fourier transform of the noisy speech, S(ω) is the clean speech, and N(ω) is the noise. Noise power spectrum estimation (based on the stationary noise assumption) is also performed. K represents the number of frames without audio.
3. The AI-based voice noise reduction method according to claim 1, characterized in that: Step four employs an end-to-end speech enhancement model, directly using sampling point information as input and output, thus avoiding distortion problems caused by spectrum manipulation.
4. The AI-based voice noise reduction method according to claim 1, characterized in that: Step seven further provides customized noise reduction strategies by learning user voiceprint characteristics. For example, business people may focus on suppressing keyboard noise, while meeting scenarios may enhance the clarity of multi-person conversations. This involves dynamically adjusting the noise reduction intensity and frequency band strategy based on the user's voiceprint, including dynamic parameter adjustment, identifying the type of environmental noise (such as white noise or human voice), and adjusting the noise reduction mode. The adaptive gain control formula is as follows: (Ratio of noisy signal power to noise power) (Ratio of clean signal power to noise power).
5. The AI-based voice noise reduction method according to claim 1, characterized in that: Step eight also includes voice enhancement: enhancing the target speaker using voiceprint locking technology (such as eye tracking).
Citation Information
Cited By
AI intelligent sound equipment signal processing method and system, sound equipment device and storage medium
CN121789667A