A microphone far-field pickup tri-state DRC control method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN ASMAX INFINITE TECH CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]针对现有技术的不足,本发明提供了一种麦克风远场拾音三态DRC控制方法,解决了现有远场拾音技术中信噪比提升困难、噪声抑制不佳以及对人工智能大模型识别性能影响的问题
1、本发明通过引入三路并行动态范围控制模块以及人声检测模块的动态选择机制,能够根据音频信号的实时特性进行人声增强和噪声抑制,避免了传统动态范围控制模块统一参数处理的局限性,确保了在人声状态下有效提升人声信号,在噪声状态下强力压制噪声,从而大幅提高输出音频的信噪比。
Smart Images

Figure CN121547712B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, specifically to a three-state DRC control method for far-field microphone pickup. Background Technology
[0002] The increasing use of large-scale AI models in fields such as speech recognition and natural language processing has driven the rapid development of applications in smart homes, in-vehicle systems, robotics, and remote conferencing. In these applications, voice interaction has become a core function. To ensure that large-scale AI models can accurately understand user intent, acquiring high-quality audio signals is a primary prerequisite. This is especially crucial for scenarios where the sound source is far away, making far-field microphone pickup technology paramount.
[0003] In existing far-field audio pickup solutions, electret microphones or MEMS microphones are typically used to acquire audio signals. These microphones, limited by their size and inherent performance, often produce audio signals with a low signal-to-noise ratio in far-field environments. While simple linear amplification can improve signal amplitude to enhance usability, it also amplifies background noise, resulting in poor overall signal quality and making it difficult for large-scale AI models to accurately recognize speech information.
[0004] However, existing far-field microphone pickup schemes, when directly performing noise reduction under conditions of extremely low signal-to-noise ratio, can lead to the loss of some speech information. When the ambient noise level is high, it may be incorrectly amplified by DRC (Dual Voice Control), while non-human voice signals with lower energy may be excessively suppressed, thus affecting the final audio recognition results output to the large-scale artificial intelligence model. Therefore, this invention provides a three-state DRC control method for far-field microphone pickup to address the shortcomings of existing technologies. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a three-state DRC control method for microphone far-field sound pickup, which solves the problems of difficulty in improving signal-to-noise ratio, poor noise suppression, and impact on the recognition performance of large artificial intelligence models in existing far-field sound pickup technologies.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The first aspect of the present invention provides a microphone far-field pickup three-state DRC control method, comprising the following steps: The raw audio signal acquired from the far field is pre-processed to generate an optimized audio signal. This pre-processing specifically includes: pre-adjusting the gain of the raw audio signal to ensure that the signal amplitude is within a preset reasonable range; performing high-pass filtering on the pre-adjusted signal to eliminate low-frequency noise components in the signal; and performing acoustic echo cancellation on the high-pass filtered signal to eliminate the echo generated by the system's own speaker in the microphone.
[0007] The optimized audio signal is input in parallel to three independent Dynamic Range Control (DRC) modules: a first DRC module, a second DRC module, and a third DRC module, thereby constructing three parallel audio signals. These three DRC modules use the same algorithm but have different parameter configurations: the first DRC module is configured as a voice enhancement DRC, designed to boost voice signals above a certain amplitude; the second DRC module is configured as a noise suppression DRC, designed to suppress noise amplitude to a preset low level; and the third DRC module is configured as a transitional intermediate-state DRC, with parameter configurations between the voice enhancement DRC and the noise suppression DRC, used to handle the transitional state of the signal.
[0008] The audio signal output from the first dynamic range control module after parallel processing is used as the input to the human voice detection module. The human voice detection module extracts features from the received audio signal, including but not limited to short-time energy, zero-crossing rate, or spectral centroid. Based on the extracted features and a pre-trained model or algorithm, the human voice detection module determines the characteristics of the current audio frame, including human voice, noise, or an intermediate state (i.e., a transition from non-human voice to human voice or vice versa). Based on the determination result, the human voice detection module generates Boolean-type human voice state signals, noise state signals, and intermediate state signals; these signals are mutually exclusive. It is worth noting that the input to the human voice detection module is the output of the first dynamic range control module, not the original audio signal. This is because the first dynamic range control module enhances the human voice, making the human voice signal fuller, thereby effectively avoiding detection errors that may occur in far-field environments due to the low amplitude of the original audio signal, and improving the accuracy of human voice detection.
[0009] In the compressed region, the dynamic range control of the first dynamic range control module includes adjusting the dynamic range based on the input loudness. Determine the output loudness The output loudness With the input loudness The following relationship exists between them: ; In the formula, This represents the first threshold. This represents the first compression ratio.
[0010] Based on the voice state signal, noise state signal, and intermediate state signal generated by the voice detection module, the system dynamically selects one of the three parallel audio signals as the current output signal. Specifically, when the voice state signal is valid, the output of the first dynamic range control module is selected as the current output signal to preserve the boosted voice; when the noise state signal is valid, the output of the second dynamic range control module is selected as the current output signal to ensure that the noise is effectively suppressed, even close to zero; when the intermediate state signal is valid, the output of the third dynamic range control module is selected as the current output signal to ensure the continuity of the sound and avoid missing words when the signal-to-noise ratio is low or the speaker's onset is unclear, by appropriately suppressing noise and boosting the voice amplitude.
[0011] Post-denoising processing is performed on the current output signal to obtain a denoised audio signal. This denoising process can be based on the output signal and employ any one or a combination of spectral subtraction denoising, Wiener filtering denoising, or deep learning denoising to further purify the audio signal and reduce residual noise.
[0012] The denoised audio signal is output to an AI model to improve the signal-to-noise ratio of human voice and suppress noise, thereby optimizing the recognition performance of the AI model. This output process specifically includes: performing automatic gain control on the denoised audio signal to ensure the output loudness remains stable within the acceptable range for the AI model; encoding and formatting the audio signal to meet the input requirements of the AI model; and transmitting the processed audio stream to the AI model for tasks such as speech recognition or semantic understanding.
[0013] A second aspect of the present invention provides a microphone far-field pickup three-state DRC control system for implementing the above-described method. It includes: The pre-processing module is used to pre-process the raw audio signals acquired from the far field to generate optimized audio signals.
[0014] The first dynamic range control module, the second dynamic range control module, and the third dynamic range control module are all connected to the pre-processing module and are used to receive the optimized audio signal and construct three parallel audio signals. Specifically, the first dynamic range control module is configured for voice enhancement DRC, the second dynamic range control module is configured for noise suppression DRC, and the third dynamic range control module is configured for transitional intermediate state DRC.
[0015] The human voice detection module has its input end connected to the output end of the first dynamic range control module. It is used to determine the human voice state, noise state, or intermediate state characteristics of the current audio frame based on the output of the first dynamic range control module, and generate human voice state signal, noise state signal, and intermediate state signal.
[0016] The selection module has its input end connected to the output ends of the first dynamic range control module, the second dynamic range control module, and the third dynamic range control module, and receives the human voice state signal, noise state signal, and intermediate state signal generated by the human voice detection module. It is used to dynamically select one of the three parallel processed audio signals as the current output signal according to the signal from the human voice detection module.
[0017] A post-noise reduction processing module is connected to the output terminal of the selection module and is used to perform noise reduction processing on the output signal of the selection module to obtain a noise-reduced audio signal.
[0018] The output module is connected to the output end of the post-noise reduction processing module and is used to output the noise-reduced audio signal to the artificial intelligence big model to optimize the artificial intelligence big model's recognition of the audio signal.
[0019] This invention provides a three-state DRC control method for far-field microphone pickup. It has the following beneficial effects: 1. This invention introduces a dynamic selection mechanism of a three-way parallel dynamic range control module and a human voice detection module, which can enhance human voice and suppress noise according to the real-time characteristics of the audio signal. This avoids the limitations of the unified parameter processing of the traditional dynamic range control module, ensuring that the human voice signal is effectively enhanced in the human voice state and the noise is strongly suppressed in the noise state, thereby greatly improving the signal-to-noise ratio of the output audio.
[0020] 2. When the human voice detection module determines that the audio frame is in the transition state between human voice and non-human voice, the present invention selects the output of the transition intermediate state dynamic range control module. It takes into account the errors that may occur in human voice detection or the complexity of signal onset in practice. By appropriately suppressing noise and enhancing human voice, it effectively avoids the loss of key audio information due to misjudgment, thereby ensuring the continuity and integrity of the sound.
[0021] 3. Before outputting the denoised audio signal to the large-scale artificial intelligence model, this invention performs automatic gain control, encoding, and format conversion, which ensures that the output loudness is stable within the optimal input range that the large-scale artificial intelligence model can accept. This avoids the negative impact of excessively strong or weak signals on model recognition, enabling the large-scale artificial intelligence model to stably perform speech recognition or semantic understanding tasks. Attached Figure Description
[0022] Figure 1This is a schematic diagram of the microphone far-field pickup three-state DRC control system architecture of the present invention; Figure 2 This is a flowchart of the method of the present invention; Figure 3 This is a schematic diagram of the far-field sound pickup algorithm framework of the present invention; Figure 4 This is a schematic diagram of the DRC algorithm framework of the present invention; Figure 5 This is a schematic diagram illustrating the principle of the DRC algorithm of the present invention.
[0023] Among them, 100 is the pre-processing module; 200 is the first dynamic range control module; 300 is the second dynamic range control module; 400 is the third dynamic range control module; 500 is the human voice detection module; 600 is the selection module; 700 is the post-noise reduction processing module; and 800 is the output module. Detailed Implementation
[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] See attached document Figure 1 , Figure 1 This is a schematic diagram of a microphone far-field pickup three-state DRC control system architecture according to an embodiment of the present invention. The present invention provides a microphone far-field pickup three-state DRC control system, including a pre-processing module 100, a first dynamic range control module 200, a second dynamic range control module 300, a third dynamic range control module 400, a voice detection module 500, a selection module 600, a post-noise reduction processing module 700, and an output module 800.
[0026] See attached document Figure 2 -Appendix Figure 5 This invention provides a microphone far-field pickup three-state DRC control method, comprising the following steps: S1. The raw audio signal acquired from the far field is pre-processed to generate an optimized audio signal. Specifically, this is performed by the pre-processing module 100. The pre-processing module 100 receives the raw audio signal from the microphone and performs gain pre-adjustment, high-pass filtering, and acoustic echo cancellation on it to output a pre-optimized audio signal.
[0027] S2. The optimized audio signal is input in parallel to three independent dynamic range control modules to construct three parallel-processed audio signals. This is accomplished jointly by the first dynamic range control module 200, the second dynamic range control module 300, and the third dynamic range control module 400. The output signal of the pre-processing module 100 is simultaneously input to the first dynamic range control module 200, the second dynamic range control module 300, and the third dynamic range control module 400. These three modules independently perform dynamic range control processing on the input audio signal and output the processed audio signal in parallel.
[0028] S3. The audio signal output after parallel processing is used as the input to the human voice detection module. The human voice detection module determines the human voice state, noise state, or intermediate state characteristics of the current audio frame based on the output of the first dynamic range control module, and generates human voice state signals, noise state signals, and intermediate state signals. This is executed by the human voice detection module 500. The human voice detection module 500 receives the output signal of the first dynamic range control module 200 as its input, and performs feature extraction and analysis based on the input signal to determine whether the current audio frame belongs to human voice, noise, or intermediate state (i.e., the transition from non-human voice to human voice or from human voice to non-human voice). The determination result is output in the form of Boolean-type human voice state signals, noise state signals, and intermediate state signals.
[0029] S4. Based on the generated human voice state signal, noise state signal, and intermediate state signal, dynamically select one audio signal from the three parallel-processed audio signals as the current output signal. This is performed by the selection module 600. The selection module 600 receives the human voice state signal, noise state signal, and intermediate state signal output from the human voice detection module 500, and simultaneously receives the parallel-processed audio signals from the first dynamic range control module 200, the second dynamic range control module 300, and the third dynamic range control module 400. Based on the received state signal, the selection module 600 dynamically selects one audio signal as its output, which is the current output signal of this method.
[0030] S5. Perform post-denoising processing on the signal output by the selection module 600 to obtain a denoised audio signal. This is performed by the post-denoising processing module 700. The post-denoising processing module 700 receives the audio signal output by the selection module 600 and performs further denoising processing on it to purify the audio signal and generate a denoised audio signal.
[0031] S6. Output the denoised audio signal to the AI model to optimize the AI model's recognition of the audio signal. This is performed by the output module 800. The output module 800 receives the denoised audio signal output by the post-denoising processing module 700, and performs automatic gain control, encoding, and format conversion on it. Finally, it transmits the audio signal that meets the input requirements of the AI model to the AI model.
[0032] For step S1, firstly, the gain of the original audio signal is pre-adjusted to ensure that the amplitude of the original audio signal is adjusted to a preset reasonable range, avoiding further deterioration of the signal-to-noise ratio due to an excessively weak signal, or clipping distortion due to an excessively strong signal. Gain pre-adjustment can be achieved by multiplying by a gain factor. This is achieved for a single original audio signal frame. The signal after gain pre-adjustment It can be represented as: ; In the formula, This represents the sample value of the original audio signal. This represents the preset gain factor. This represents the sampled value of the audio signal after gain pre-adjustment. This gain factor... The determination is based on the analysis of the energy or peak level of the original audio signal to ensure that the adjusted signal amplitude meets the dynamic range requirements of subsequent processing.
[0033] High-pass filtering is applied to the gain-pre-adjusted signal to eliminate low-frequency noise components in the audio signal, such as ambient hum, air conditioner noise, or noise caused by ground vibrations. High-pass filters are typically designed with a specific cutoff frequency (e.g., below 80Hz or 100Hz) to effectively filter out most non-human low-frequency noise below the human voice frequency band. This processing helps improve signal clarity.
[0034] Acoustic echo cancellation is applied to the high-pass filtered signal to eliminate the echoes generated by the system's own speakers in the microphone. In conference or robot interaction scenarios, voice commands or prompts played by the system may be picked up again by the microphone through the acoustic path, creating echoes that interfere with the recognition of far-field human voices. Acoustic echo cancellation constructs an echo path model using an adaptive filtering algorithm (e.g., based on the Least Mean Square (LMS) or Normalized Least Mean Square (NLMS) algorithm) and predicts and cancels echo components from the signal picked up by the microphone, thereby outputting a clean audio signal without echoes.
[0035] In step S2, the first dynamic range control module 200 is used as an input by the voice detection module 500 and configured for voice enhancement DRC. The main goal of this module is to enhance the voice signal above a certain amplitude, making the voice fuller and more prominent. Therefore, the parameter configuration of the first dynamic range control module 200 tends to have a lower threshold. and a large compression ratio A lower threshold ensures that even relatively small vocal amplitudes can be effectively enhanced, while a higher compression ratio guarantees that above the threshold, the signal's dynamic range is effectively compressed, thereby improving the overall loudness of the vocals. Its input loudness... With output loudness The formula relating them is: ; In the formula, This indicates the threshold of the first dynamic range control module 200; This indicates the compression ratio of the first dynamic range control module 200; Input loudness; This configuration significantly enhances the loudness of the human voice signal after it passes through the module.
[0036] The second dynamic range control module 300 is configured for noise suppression (DRC). The primary goal of this module is to suppress noise to the maximum extent possible, reducing its loudness to a preset low level, or even close to zero. To this end, the parameters of the second dynamic range control module 300 are typically set with high thresholds. and extremely high compression ratio A high threshold ensures that only signals with extremely high amplitudes are processed, while a very high compression ratio (e.g., close to infinity, achieving a limiter effect) ensures that any noise signals exceeding the threshold are strongly compressed, thus effectively suppressing noise to a lower amplitude.
[0037] The third dynamic range control module 400 is configured as a transitional intermediate state DRC. The parameters of this module are configured between voice enhancement DRC and noise suppression DRC. Its purpose is to appropriately suppress noise and enhance voice amplitude during the transition between voice and noise, so as to avoid abruptness or loss of important information during signal switching. The threshold of the third dynamic range control module 400... and compression ratio It will be adjusted according to actual needs, usually It will be higher than But lower ,and Will be between and This intermediate-state strategy aims to provide a smooth transition, ensuring that the continuity and intelligibility of the sound are maintained even when the VAD module's judgment is ambiguous or the signal is in a mixed state, and that noise is properly suppressed.
[0038] Step S3, using the output of the first dynamic range control module 200 as the input of the voice detection module 500, has technical advantages. The first dynamic range control module 200, acting as a voice enhancement DRC, amplifies the voice signal, making it fuller and more prominent. In far-field pickup environments, the voice amplitude in the original audio signal may be extremely low, causing the voice signal to be submerged in noise or deviating significantly from the amplitude of the VAD algorithm training data. If the original audio signal is used directly for VAD, the accuracy of voice detection will decrease. By using the signal enhanced by the first dynamic range control module 200, the signal-to-noise ratio of the voice signal is effectively improved at the VAD input, thereby enhancing the recognition accuracy and robustness of the voice detection module 500.
[0039] The voice detection module 500 extracts features from the received audio signal. These features characterize the acoustic properties of the audio signal and serve as a basis for determining the state of the voice. Common features include: Short-time energy Short-time energy measures the amplitude and intensity of an audio signal over a short period of time, and is used to distinguish between human voice and noise. It can be calculated using the following formula: ; In the formula, Indicates the first The audio signal of the frame after being weighted by a window function. This indicates the number of sampling points per frame.
[0040] Zero crossing rate Zero-crossing rate: Measures the number of times an audio signal crosses the zero axis per unit time, typically reflecting the signal's frequency characteristics. For certain types of noise (such as white noise) and voiceless consonants, the zero-crossing rate is high, while for voiced sounds, it is low. Zero-crossing rate It can be calculated using the following formula: ; In the formula, Indicates the first The audio signal sample value of the frame, For symbolic functions, This indicates the number of sampling points per frame.
[0041] Spectral centroid: Reflects the brightness of the spectrum. The spectral centroid of human voice is usually lower than that of high-frequency noise.
[0042] The human voice detection module 500 determines the characteristics of the current audio frame based on extracted features and a pre-trained algorithm. The pre-trained algorithm can be based on various machine learning models, learning feature patterns from a large amount of labeled data (human voice, noise, silence) to build a recognition model. The human voice detection module 500 inputs the features of the current audio frame into the pre-trained algorithm to obtain a judgment result as to whether the current frame is human voice, noise, or an intermediate state.
[0043] Based on the judgment result, the human voice detection module 500 generates three mutually exclusive Boolean signals: human voice state signal VAD1, noise state signal VAD3, and intermediate state signal VAD2. When the detection result is human voice, the human voice state signal VAD1 is set to an active state (e.g., high level); when the detection result is noise, the noise state signal VAD3 is set to an active state; and when the detection result is a transition state from non-human voice to human voice or from human voice to non-human voice, the intermediate state signal VAD2 is set to an active state. These state signals are transmitted to the selection module 600 to control the dynamic selection of subsequent audio streams.
[0044] For step S4, based on the voice state signal VAD1, noise state signal VAD3, and intermediate state signal VAD2 output by the voice detection module 500, one audio signal is dynamically selected as the current output signal from the three parallel-processed audio signals output by the first dynamic range control module 200, the second dynamic range control module 300, and the third dynamic range control module 400. This dynamic selection mechanism ensures that the system can output an optimized audio stream under different audio scenarios.
[0045] When the voice status signal VAD1 is active, it indicates that the voice detection module 500 has determined that the current audio frame mainly contains human voice. In this case, the selection module 600 selects the output of the first dynamic range control module 200 as the current output signal of this method. The first dynamic range control module 200 is configured for voice enhancement DRC, the purpose of which is to increase the amplitude of human voice. Therefore, selecting its output can effectively preserve and highlight the enhanced human voice signal, avoiding the loss of human voice energy due to noise suppression.
[0046] When the noise status signal VAD3 is active, this indicates that the voice detection module 500 has determined that the current audio frame mainly contains noise. In this case, the selection module 600 will select the output of the second dynamic range control module 300 as the current output signal of this method. The second dynamic range control module 300 is configured for noise suppression (DRC), the purpose of which is to reduce noise. Selecting its output can suppress ambient noise to the maximum extent and provide a cleaner background.
[0047] When the intermediate state signal VAD2 is valid, it indicates that the human voice detection module 500 has determined that the current audio frame is in a transitional state between human voice and non-human voice, or that the signal-to-noise ratio is relatively complex. In this case, the selection module 600 will select the output of the third dynamic range control module 400 as the current output signal of this method. The third dynamic range control module 400 is configured as a transitional intermediate state DRC, designed to provide a smooth transition, appropriately suppress noise, and appropriately boost the amplitude of human voice. This selection can effectively avoid audio loss that may occur when VAD detection is uncertain or when the signal switches, ensuring the continuity and robustness of the sound, and even noise can be appropriately suppressed.
[0048] In step S5, the post-noise reduction processing module 700 receives the audio signal output by the selection module 600 and performs further noise reduction processing on it to obtain a purified audio signal. Although noise suppression has already been performed in the DRC processing stage, additional noise reduction operations are performed to further improve the signal-to-noise ratio and audio quality.
[0049] Spectral subtraction denoising is a denoising method based on short-time spectral analysis. Its basic principle is to assume that the noise is stationary. By estimating the noise's spectrum and subtracting the estimated noise spectrum from the spectrum of the noisy speech, a spectral estimate of the clean speech is obtained. Specifically, the noisy signal is framed and subjected to Fourier transform to obtain its amplitude spectrum. Simultaneously estimate the amplitude spectrum of the noise. The amplitude spectrum of the noise-reduced signal. The following formula can be used to estimate: ; In the formula, Represents the power spectrum of a noisy signal; Represents the estimated power spectrum of the noise. It is an over-subtraction factor greater than 0, used to control the strength of noise suppression; This represents the estimated power spectrum of the denoised signal. After obtaining the estimated amplitude spectrum of the clean signal, the denoised time-domain audio signal can be reconstructed by performing an inverse Fourier transform, combined with the phase information of the noisy signal.
[0050] Wiener filtering is another common noise reduction technique. It designs a filter based on the minimum mean square error criterion to process noisy signals in the frequency domain. The goal of the Wiener filter is to minimize the mean square error between the estimated signal and the true clean signal. The frequency domain transfer function of the Wiener filter... It is usually expressed as: ; In the formula, The power spectrum representing a pure signal. This represents the power spectrum of the noise. In practical applications, the power spectra of the clean signal and noise need to be estimated. Wiener filtering typically provides smoother noise reduction than spectral subtraction, especially in low signal-to-noise ratio environments.
[0051] Deep learning-based denoising utilizes deep neural networks to learn complex patterns in human voice and noise. Deep learning models can directly learn mappings from noisy signals to predict clean signal or noise components. This approach typically achieves superior denoising results, especially in non-stationary or complex noise environments. The model learns from a large amount of noisy and clean speech data during the training phase, and then applies it to real-world signals for denoising during the inference phase.
[0052] For step S6, automatic gain control (AGC) is performed on the denoised audio signal. The purpose of AGC is to ensure that the output loudness is stable within the optimal input range acceptable to the large artificial intelligence model, avoiding clipping distortion due to excessively high signal loudness, or difficulty in model recognition due to excessively low loudness. The implementation process of AGC includes: Periodically calculate the short-time average power of the input audio signal: The output module 800 continuously divides the input audio signal into frames and calculates the average power of each frame or time period. For example, for the first frame... Short-time average power of audio frames It can be calculated using the following formula: ; In the formula, This represents the sample value of the denoised audio signal in the current frame. This indicates the number of sampling points per frame.
[0053] The gain factor is dynamically calculated based on the ratio of the short-time average power to the preset target output power: the system presets a target output power. The output module 800 will output the currently calculated short-time average power. and A comparison is made. Based on the ratio of the two values, a gain factor is dynamically calculated. Usually, when Below hour, It will be greater than 1 to be amplified; when Higher than hour, It will be less than 1 for attenuation. To avoid drastic changes in the gain factor that could cause listening discomfort, smoothing or a limit on the rate of gain change is usually introduced.
[0054] The gain factor is applied to the current audio frame to adjust the signal amplitude: the calculated gain factor The sampled values applied in real time to the current audio frame. The audio signal after AGC processing. It can be represented as: ; In the formula, This indicates the current audio signal that has not been processed by AGC at the [number]th [time]. The value of each sampling point.
[0055] In this way, the power of the output signal approaches the preset target output power. This provides a stable loudness and a suitable dynamic range input to large artificial intelligence models.
[0056] The AGC-processed audio signal is encoded and format-converted. Large-scale artificial intelligence models typically have specific requirements for the format of the input audio data, such as sampling rate, bit depth, number of channels, and encoding format (e.g., PCM, FLAC, Opus). The output module 800 will perform necessary sampling rate conversion, quantization, encoding compression (if necessary), and encapsulation into a compliant digital audio stream format according to the specific interface specifications of the large-scale artificial intelligence model.
[0057] The processed audio stream is transmitted to the large-scale artificial intelligence model. This transmission can be achieved through network interfaces (such as TCP / IP, UDP), shared memory, or other data transmission mechanisms. After receiving the audio signal optimized by this system, the large-scale artificial intelligence model will perform tasks such as speech recognition, semantic understanding, or voiceprint recognition.
Claims
1. A three-state DRC control method for far-field microphone pickup, characterized in that, Includes the following steps: S1. Perform preprocessing on the raw audio signal acquired from the far field to generate an optimized audio signal; S2. Input the optimized audio signal in parallel to the first dynamic range control module, the second dynamic range control module and the third dynamic range control module to construct three parallel audio signals; S3. The audio signal output after parallel processing is used as the input of the human voice detection module. The human voice detection module determines the human voice state, noise state or intermediate state characteristics of the current audio frame according to the output of the first dynamic range control module, and generates human voice state signal, noise state signal and intermediate state signal. S4. Based on the generated human voice state signal, noise state signal and intermediate state signal, dynamically select one of the three parallel audio signals as the current output signal; S5. Perform post-denoising processing on the output signal to obtain the denoised audio signal; S6. Output the noise-reduced audio signal to the AI model to optimize the AI model's recognition of the audio signal; In step S2, the first dynamic range control module is a voice enhancement DRC, the second dynamic range control module is a noise suppression DRC, and the third dynamic range control module is a transition intermediate state DRC. In step S4, dynamically selecting one audio signal from the three parallel-processed signals as the current output signal includes the following steps: When the human voice detection module generates a valid human voice status signal, the output of the first dynamic range control module is selected as the current output signal of this method. When the noise status signal generated by the human voice detection module is valid, the output of the second dynamic range control module is selected as the current output signal of this method. When the intermediate state signal generated by the human voice detection module is valid, the output of the third dynamic range control module is selected as the current output signal of this method.
2. The microphone far-field pickup three-state DRC control method according to claim 1, characterized in that, In step S1, the preprocessing of the raw audio signal acquired in the far field includes the following steps: The gain of the original audio signal is pre-adjusted to ensure that the signal amplitude is within the preset range; The signal after gain pre-adjustment is subjected to high-pass filtering to eliminate low-frequency noise components in the signal. The high-pass filtered signal is subjected to acoustic echo cancellation processing to eliminate the echo generated by the sound played by the speaker in the microphone.
3. The microphone far-field pickup three-state DRC control method according to claim 1, characterized in that, In step S3, generating the human voice state signal, the noise state signal, and the intermediate state signal includes the following steps: The human voice detection module extracts features from the output of the first dynamic range control module. The features include short-time energy, zero-crossing rate, or spectral centroid. The human voice detection module determines the characteristics of the current audio frame, including human voice, noise, or intermediate states, based on extracted features and pre-trained algorithms. Based on the judgment result, the human voice detection module generates a Boolean-type human voice state signal, a noise state signal, and an intermediate state signal.
4. The microphone far-field pickup three-state DRC control method according to claim 1, characterized in that, In step S5, the post-denoising process for the output signal includes the following steps: Based on the output signal, denoising is performed using spectral subtraction, Wiener filtering, or deep learning to purify the output audio signal and generate a denoised audio signal.
5. The microphone far-field pickup three-state DRC control method according to claim 1, characterized in that, In step S6, the optimization of the large artificial intelligence model for audio signal recognition includes the following steps: Automatic gain control is applied to the noise-reduced audio signal to ensure that the output loudness remains stable within the range of the large artificial intelligence model; The denoised audio signal is encoded and converted to meet the input requirements of large artificial intelligence models; The encoded and format-converted audio is transmitted to a large artificial intelligence model for speech recognition or semantic understanding tasks.
6. The microphone far-field pickup three-state DRC control method according to claim 1, characterized in that, The dynamic range control of the first dynamic range control module includes controlling the dynamic range based on the input loudness. Determine the output loudness In the compression region, the output loudness With the input loudness The following relationship exists between them: ; In the formula, Indicates the first threshold. This indicates the first compression ratio.
7. A microphone far-field pickup three-state DRC control method according to claim 5, characterized in that, The automatic gain control of the noise-reduced audio signal includes the following steps: Periodically calculate the short-time average power of the input audio signal; The gain factor is dynamically calculated based on the ratio of short-time average power to the preset target output power. The gain factor is applied to the current audio frame to adjust the signal amplitude so that the power of the output signal approaches the target output power.
8. A microphone far-field pickup three-state DRC control system, applied to the microphone far-field pickup three-state DRC control method according to any one of claims 1-7, characterized in that, include: The pre-processing module is used to pre-process the raw audio signal acquired from the far field to generate an optimized audio signal; The first dynamic range control module, the second dynamic range control module, and the third dynamic range control module are used to receive the optimized audio signal and construct three parallel audio signals. The human voice detection module is used to determine the human voice state, noise state, or intermediate state characteristics of the current audio frame based on the output of the first dynamic range control module, and to generate human voice state signal, noise state signal, and intermediate state signal. The selection module is used to dynamically select one of the three parallel audio signals as the current output signal based on the generated human voice state signal, noise state signal and intermediate state signal. A post-noise reduction processing module is used to perform noise reduction processing on the output signal of the selection module to obtain a noise-reduced audio signal. The output module is used to output the noise-reduced audio signal to the large artificial intelligence model to optimize the model's recognition of the audio signal.
Citation Information
Patent Citations
Noise reduction optimization method for noise reduction type MEMS microphone for smart home
CN118314917A
Noise reduction method combining different noise reduction algorithms of motorcycle riding earphone
CN120881452A