A Microphone AI Audio Processing Method Based on Domestically Produced Laptops

CN122575404APending Publication Date: 2026-08-14联想开天科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本发明的目的是为了解决现有技术中存在的缺点,而提出的一种基于信创笔记本的麦克风AI音频处理方法,该装置能够有效的解决定向或共享拾音模式、语音追踪模式和远距离拾音模式场景适用的问题

Benefits of technology

[0015] This invention proposes an AI audio processing method for microphones in domestically developed laptops. The beneficial effects are as follows: To achieve the removal and suppression of various complex noises, including steady-state and non-steady-state noise, and specifically targeting different types of noise such as keyboard sounds, typing sounds, air conditioner sounds, household operation sounds, cat meows, and dog barks, a large amount of noise and speech data from various environments has been collected. This includes consideration of usage scenarios with extremely low signal-to-noise ratios. Deep AI analysis and network training are then performed to achieve effective noise removal. Comparative recording tests were conducted on the computer using both the AI ​​noise reduction module and the system's built-in noise reduction module in environments with noisy white noise backgrounds. The method also supports speech pickup at distances of 4 meters and above. As the distance to the sound source increases, the algorithm automatically increases the audio gain, thereby effectively picking up distant speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575404A_ABST
    Figure CN122575404A_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio noise processing technology, and in particular to an AI audio processing method for microphones in domestically developed laptops. This invention achieves the removal and suppression of various complex noises, including steady-state and non-steady-state noise suppression. Specifically targeting different types of noise, such as keyboard sounds, typing sounds, air conditioner sounds, household operation sounds, cat meows, dog barks, and other non-steady-state noises, it collects a large amount of noise and speech data from various environments, including considering usage scenarios with extremely low signal-to-noise ratios, and performs deep AI analysis and network training to achieve effective noise removal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio noise processing technology, and in particular to an AI audio processing method for microphones based on domestically developed laptops. Background Technology

[0002] Currently, domestically developed IT products use the Linux operating system. When users make voice calls and record audio using these products, ambient noise significantly impacts the experience, resulting in noisy and unclear sound, negatively affecting the user experience. Furthermore, domestically developed operating systems are completely inadequate for various computer audio applications, such as directional or shared pickup modes, voice tracking modes, and long-distance pickup modes.

[0003] Existing solutions only offer noise reduction (NS) and echo cancellation (AEC) for steady-state noise, but they struggle to handle non-steady-state noise (such as keyboard clicks, typing sounds, air conditioner noises, household appliance noises, cat meows, dog barks, etc.) and are difficult to improve through tuning and optimization. In many computer audio applications, such as directional or shared pickup modes, voice tracking modes, and long-distance pickup modes, existing solutions are completely inadequate to meet the needs of these scenarios. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing an AI audio processing method for microphones in domestically developed laptops. This device can effectively solve the problems of applicability in scenarios such as directional or shared sound pickup mode, voice tracking mode, and long-distance sound pickup mode.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: Design a microphone AI audio processing method for domestically developed laptops. Based on an AI audio processing module running on top of the high-level Linux sound architecture kernel layer, it provides audio processing services to upper-layer audio applications through a pulse audio client; wherein: The audio data is processed frame by frame in the pipeline of the AI ​​audio processing module in short time frames. It interacts in parallel through the microphone link and the playback link to complete audio acquisition, noise reduction, echo cancellation, far-field sound pickup and speech enhancement.

[0006] In one embodiment, the AI ​​audio processing module further includes a cross-link data channel, which includes an echo reference transmission channel and a high-frequency boost coefficient transmission channel; one end of the echo reference transmission channel is connected to the frequency domain signal output of the playback link, and the other end is connected to the echo reference input of the acoustic echo cancellation module of the microphone link; one end of the high-frequency boost coefficient transmission channel is connected to the coefficient output of the brightness enhancement estimation module of the microphone link, and the other end is connected to the coefficient input of the brightness enhancement module of the playback link.

[0007] In one embodiment, the microphone link starts from the input signals of the multi-microphone array and the main microphone, obtains the azimuth angle of the sound source by sound source arrival direction estimation, and then passes through pre-gain, preprocessing, dynamic range management, time-domain to frequency-domain conversion to time-spectrum graph, acoustic echo cancellation to eliminate echo components, beamforming to enhance the speech in the target direction, artificial intelligence noise suppression to filter non-steady-state noise, and then through howling suppression, residual removal, brightness enhancement estimation to generate high-frequency boost coefficient, frequency equalization to optimize the human voice frequency band, frequency-domain to time-domain conversion to restore the time-domain waveform, far-field processing and post-gain to compensate for sound attenuation through automatic gain control, and finally through dynamic range compression and high-pass filtering to output the processed linear audio.

[0008] In one embodiment, the sound source arrival direction estimation uses a generalized cross-correlation-phase transform algorithm or multi-channel phase difference analysis to calculate the sound source azimuth angle, and transmits the angle information to the beamforming module.

[0009] In one embodiment, the artificial intelligence noise suppression inputs audio time-frequency features into a pre-trained deep neural network, outputs an ideal proportional mask or complex mask, and multiplies it point-by-point with the spectrum to achieve non-steady-state noise suppression.

[0010] In one embodiment, the far-field processing and post-gain are compensated for by automatic gain control to reduce acoustic attenuation.

[0011] In one embodiment, the playback link starts from the remote audio input, sequentially undergoes preprocessing, harmonic management, time-domain to frequency-domain conversion to frequency domain signal, secondary noise purification by remote noise suppression, then feedback suppression, brightness enhancement receiving the high-frequency boost coefficient to compensate for high-frequency details, frequency-domain to time-domain conversion to restore time-domain waveform, brightness enhancement gain application, dynamic range compression, and high-pass filtering, before outputting the speaker signal; wherein, the playback link sends the frequency-domain signal as echo reference data to the acoustic echo cancellation module of the microphone link.

[0012] In one embodiment, the playback link sends echo reference data to the acoustic echo cancellation module of the microphone link to achieve acoustic echo cancellation.

[0013] In one embodiment, the playback link receives the high-frequency boost coefficient transmitted by the microphone link to complete high-frequency detail compensation for the far-end audio.

[0014] A microphone AI audio processing method based on a domestically developed laptop, wherein the sound source arrival direction estimation calculates the sound source azimuth angle based on the multi-microphone array signal and transmits it to a beamforming module, and the beamforming module forms a spatial auditory beam based on the sound source azimuth angle to control the sound pickup direction, thereby realizing directional sound pickup or voice tracking. The multi-microphone array filters out non-steady-state noise such as keyboard sounds, typing sounds, and air conditioning sounds through beamforming and artificial intelligence noise suppression. Far-field processing and post-gain are used to compensate for long-distance sound attenuation through automatic gain control, thereby achieving far-field sound pickup.

[0015] This invention proposes an AI audio processing method for microphones in domestically developed laptops. The beneficial effects are as follows: To achieve the removal and suppression of various complex noises, including steady-state and non-steady-state noise, and specifically targeting different types of noise such as keyboard sounds, typing sounds, air conditioner sounds, household operation sounds, cat meows, and dog barks, a large amount of noise and speech data from various environments has been collected. This includes consideration of usage scenarios with extremely low signal-to-noise ratios. Deep AI analysis and network training are then performed to achieve effective noise removal. Comparative recording tests were conducted on the computer using both the AI ​​noise reduction module and the system's built-in noise reduction module in environments with noisy white noise backgrounds. The method also supports speech pickup at distances of 4 meters and above. As the distance to the sound source increases, the algorithm automatically increases the audio gain, thereby effectively picking up distant speech. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the integrated architecture of the Pulse Audio architecture and AI audio architecture within the domestically developed system for a microphone AI audio processing method based on a domestically developed laptop, as proposed in this invention.

[0017] Figure 2 This is a flowchart illustrating the AI ​​audio algorithm workflow of a microphone AI audio processing method for a domestically developed laptop, as proposed in this invention.

[0018] Figure 3 This image shows a comparison of noise reduction with AI audio enabled and disabled on a domestically developed laptop, based on a microphone AI audio processing method proposed in this invention.

[0019] Figure 4 This is a comparison diagram showing the microphone AI audio processing method proposed in this invention based on a domestically developed laptop with AI far-field sound pickup turned off and on under the domestically developed system. Detailed Implementation

[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0021] like Figure 1 As shown, a microphone AI audio processing system based on a domestically developed laptop includes: integrating an AI audio processing module into the Pulse Audio architecture of the domestically developed system, and using a frame-level processing mechanism to achieve audio acquisition, noise reduction, echo cancellation, far-field pickup and speech enhancement. The system is set to sample at 16kHz, and the audio data is processed in pipeline in short frames with a frame length of 512 and a frame shift of 128. The system is divided into two parallel interactive links: a microphone link (near-end pickup) and a playback link (far-end playback).

[0022] In scenarios such as voice calls and recordings in domestically developed IT systems, audio signals are collected through the laptop microphone. The collected signals include steady-state noise as well as non-steady-state noise such as keyboard sounds and typing sounds.

[0023] In one embodiment, the AI ​​audio processing module includes deploying an AI audio processing software package in a Pulse Server architecture, running on top of the ALSA Kernel layer, providing support for upper-layer audio applications through PA Client, and loading a pre-trained AI audio processing model. This model is trained on network data and speech corpora in various environments and scenarios with extremely low signal-to-noise ratios, and can achieve audio feature extraction and noise suppression.

[0024] refer to Figure 2 The microphone link (near-end signal processing) is used to clean and enhance the local pickup signal. The processing is performed in the following order: the microphone link starts from the input of the multi-microphone array (N mic) and the main microphone signal (Mic), and goes through direction of arrival (DOA) estimation, pre-gain, pre-processing, dynamic range management (HD), time-domain to frequency-domain (T2F), acoustic echo cancellation (AEC), beamforming (BF), noise suppression (NS), howling suppression (HS), Safe Clear, brightness enhancement estimation (BVE Est.), frequency equalization (EQ), frequency-domain to time-domain (F2T), far-field processing / post-gain (FFP / Post-Gain), dynamic range compression (DRC), and high-pass filtering (HPF), and finally outputs a processed linear output.

[0025] Among them, Direction of Arrival (DOA) estimation: receiving the input signal from the multi-microphone array, using the generalized cross-correlation-phase transform (GCC-PHAT) algorithm or multi-channel phase difference analysis, accurately calculating the spatial azimuth angle of the target sound source (such as a speaker), and transmitting this angle information to the subsequent beamforming module; Pre-Gain: Performs an initial linear amplification on the original microphone signal to increase the overall signal level and prevent the loss of precision (underload) in subsequent fixed-point or floating-point operations due to the signal being too small. Pre-processing: A fourth-order Butterworth high-pass filter with a cutoff frequency of 50Hz is used to filter out DC offset introduced by hardware circuitry and extremely low-frequency structural noise or wind noise. Dynamic Range Management / Harmonic Distortion Correction (HD): Employs a hyperbolic tangent function soft limiting algorithm to perform nonlinear mapping on the signal amplitude, preventing hard clipping of the pre-gained signal in extreme cases, which would produce harsh harmonic distortion. Time-domain to frequency-domain (T2F): Short-time Fourier transform (STFT) combined with Hanning window is used to complete the FFT transformation, which maps the one-dimensional time-domain signal into a two-dimensional time-frequency complex spectrum, providing a foundation for subsequent complex spectrum processing. Acoustic Echo Cancellation (AEC): Receives echo reference frequency domain data from the playback link, uses the NLMS adaptive filtering algorithm to simulate the acoustic propagation path and subtract echo components, thus eliminating speaker echoes. Beamforming (BF): Based on DOA angle information, it uses delay summation or minimum variance distortionless response (MVDR) algorithms to form a spatial "auditory beam" in the direction of the target speaker, enhancing the target speech while spatially suppressing interference noise from non-target directions; AI Noise Suppression (NS): Integrates an AI deep learning audio processing model into the Pulse Audio architecture. It extracts the time-frequency features of the current audio frame and feeds them into a pre-trained deep neural network model. This model is trained with data containing extremely low signal-to-noise ratios and a large amount of non-stationary noise (such as keyboard sounds, clapping sounds, air conditioner sounds, pet barking, etc.). The model infers and outputs an ideal proportional mask (IRM) or complex mask (cRM), which is multiplied with the current spectrum point by point. This breaks through the limitation of traditional Wiener filtering, which can only handle stationary noise, and accurately filters out complex non-stationary environmental noise, preserving pure speech. Howling suppression (HS): Real-time detection of abnormal peak frequencies in the spectrum and dynamic application of deep notch filtering to avoid howling under shared pickup or high volume conditions; Residual removal (Safe Clear): Employs a frequency domain soft threshold algorithm to attenuate residual music noise and artifacts after processing, ensuring a clean background; Brightness Enhancement Estimation (BVE Est.): Analyzes energy in the 5kHz-10kHz high-frequency band, dynamically generates high-frequency boost coefficients for local equalization and far-end audio enhancement, and quantifies the degree of loss of high-frequency details in speech in real time. This not only guides local equalization adjustments but also extends the range of this parameter across different frequency ranges. The link is passed to the playback link for use; Frequency Equalization (EQ): Optimizes the 200Hz-5000Hz human voice frequency band through a parametric equalizer to improve the fullness and naturalness of the voice. Frequency domain to time domain (F2T): The inverse short-time Fourier transform (iSTFT) combined with the overlap addition method is used to restore the processed spectrum to the time domain waveform; Far-field processing / post-gain (FFP / Post-Gain): Automatic gain control (AGC) compensates for sound attenuation at long distances, enabling far-field sound pickup at distances of 4 meters and above, meeting the requirement of "far-field sound pickup at distances of 4 meters and above", and automatically increasing audio gain to compensate for sound attenuation; Dynamic Range Compression (DRC): Attenuates the signal logarithmic domain gain by setting a threshold, balances the volume, and prevents clipping of loud sounds and blurring of soft sounds; High-pass filtering (HPF): Applying an 80 Hz high-pass filter at the end of the time domain removes residual low-frequency DC components and outputs the final pure linear audio.

[0026] The playback link (remote signal processing) is used to optimize the quality of remote audio playback. It performs processing in the following order: The playback link starts from the remote audio signal (Playback) input, and goes through pre-processing (Pre-Proc), harmonic management (HD), time-domain to frequency-domain (T2F), remote noise suppression (FENS), howling suppression (HS), brightness enhancement (BVE), frequency-domain to time-domain (F2T), brightness enhancement gain application (Apply BVE Gain), dynamic range compression (DRC), and high-pass filtering (HPF), and finally outputs the speaker signal (SPK Out).

[0027] Preprocessing and harmonic management: Receive the remote Playback signal and perform the same time-domain high-pass filtering and soft-limiting anti-distortion processing as the microphone link; Time-to-frequency (T2F): Converts the remote signal into the frequency domain and sends the original frequency domain data as an echo reference to the microphone link AEC module; Far-end noise suppression (FENS): Reuses AI noise reduction logic to perform secondary purification on the noise inherent in the far-end audio. If the audio transmitted from the far end itself has background noise, this module can perform secondary purification on it. Howling suppression (HS): Performs spectral spike detection and notch filtering on the downlink signal as a system anti-howling security mechanism.

[0028] Brightness Enhancement (BVE): Uses the high-frequency boost factor transmitted through the microphone link to compensate for the loss of high-frequency details in far-end audio.

[0029] Frequency-to-time (F2T): Restores the processed spectrum to a time-domain signal.

[0030] Brightness enhancement gain application: The high-frequency boost factor is used again to fine-tune the overall level of the time-domain signal.

[0031] Dynamic range compression and high-pass filtering: Performs amplitude limiting compression and low-frequency cleanup to protect the speaker and output a high-quality speaker playback signal.

[0032] The playback link provides an echo reference signal to the microphone link to support acoustic echo cancellation. The microphone link synchronizes the brightness enhancement estimation coefficients to the playback link to achieve cross-link high-frequency voice enhancement.

[0033] It acquires noisy microphone input signals and performs real-time noise reduction and echo cancellation based on the AI ​​model integrated into the Pulse Audio architecture. It supports far-field pickup, directional pickup, shared pickup, and voice tracking modes, and outputs high-definition, clean audio.

[0034] refer to Figures 3 to 4 In noisy white noise environments, the recording clarity is significantly better than the system's built-in module; the signal is stable and clear during far-field sound pickup, solving the problem of weak sound from the module.

[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A microphone AI audio processing method based on a domestically developed laptop, characterized in that: An AI-powered audio processing module, running on top of the advanced Linux sound architecture kernel layer, provides audio processing services to upper-layer audio applications via a pulse audio client; where: The audio data is processed frame by frame in the pipeline of the AI ​​audio processing module in short time frames. It interacts in parallel through the microphone link and the playback link to complete audio acquisition, noise reduction, echo cancellation, far-field sound pickup and speech enhancement.

2. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 1, is characterized in that... The AI ​​audio processing module also includes a cross-link data channel, which includes an echo reference transmission channel and a high-frequency boost coefficient transmission channel. One end of the echo reference transmission channel is connected to the frequency domain signal output of the playback link, and the other end is connected to the echo reference input of the acoustic echo cancellation module of the microphone link. One end of the high-frequency boost coefficient transmission channel is connected to the coefficient output of the brightness enhancement estimation module of the microphone link, and the other end is connected to the coefficient input of the brightness enhancement module of the playback link.

3. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 1, is characterized in that... The microphone link starts with the signal input from the multi-microphone array and the main microphone. After obtaining the azimuth angle of the sound source through sound source arrival direction estimation, it sequentially passes through pre-gain, preprocessing, dynamic range management, and time-domain to frequency-domain conversion to a time-spectrum graph. Then, acoustic echo cancellation eliminates echo components, beamforming enhances the speech in the target direction, artificial intelligence noise suppression filters out non-steady-state noise, and then it passes through howling suppression, residual removal, brightness enhancement estimation to generate high-frequency boost coefficients, frequency equalization to optimize the human voice frequency band, and frequency-domain to time-domain conversion to restore the time-domain waveform. After far-field processing and post-gain, sound attenuation is compensated by automatic gain control. Finally, after dynamic range compression and high-pass filtering, the processed linear audio is output.

4. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 3, is characterized in that... The sound source arrival direction estimation uses a generalized cross-correlation-phase transformation algorithm or multi-channel phase difference analysis to calculate the sound source azimuth angle, and transmits the angle information to the beamforming module.

5. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 4, is characterized in that... The artificial intelligence noise suppression inputs the audio time-frequency features into a pre-trained deep neural network, outputs an ideal proportional mask or complex mask, and multiplies it point by point with the spectrum to achieve non-steady-state noise suppression.

6. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 3, is characterized in that... The far-field processing and post-gain are compensated for sound attenuation through automatic gain control.

7. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 1, is characterized in that... The playback link starts from the remote audio input, and sequentially passes through preprocessing, harmonic management, time-domain to frequency-domain conversion to frequency domain signal, secondary noise purification by remote noise suppression, then feedback suppression, high-frequency enhancement to compensate for high-frequency details, frequency-domain to time-domain conversion to restore time-domain waveform, high-frequency enhancement gain application, dynamic range compression, and high-pass filtering, before outputting the speaker signal; wherein, the playback link sends the frequency domain signal as echo reference data to the acoustic echo cancellation module of the microphone link.

8. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 7, is characterized in that... The playback link sends echo reference data to the acoustic echo cancellation module of the microphone link to achieve acoustic echo cancellation.

9. The microphone AI audio processing method for laptops based on domestically developed information technology, as described in claim 7, is characterized in that... The playback link receives the high-frequency boost coefficient transmitted by the microphone link to complete the high-frequency detail compensation of the far-end audio.

10. The microphone AI audio processing method for laptops based on domestically developed information technology, according to any one of claims 1 to 9, is characterized in that: The sound source arrival direction estimation calculates the sound source azimuth angle based on the multi-microphone array signal and transmits it to the beamforming module. The beamforming module forms a spatial auditory beam based on the sound source azimuth angle to control the sound pickup direction, thereby achieving directional sound pickup or voice tracking. The multi-microphone array filters out non-steady-state noise such as keyboard sounds, typing sounds, and air conditioning sounds through beamforming and artificial intelligence noise suppression. Far-field processing and post-gain are used to compensate for long-distance sound attenuation through automatic gain control, thereby achieving far-field sound pickup.