Audio signal processing method and system of earphone
By using a deep learning network model trained with psychoacoustic constraints and adaptive echo cancellation technology, the problem of insufficient headphone noise reduction caused by speaker saturation effect is solved, achieving high-fidelity audio signal processing and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN WORGO TECH LTD
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-21
AI Technical Summary
Existing active noise cancellation systems in headphones suffer from poor noise cancellation performance due to speaker saturation effect, resulting in a poor user experience, especially in noisy environments.
A deep learning network model trained with psychoacoustic constraints is used to generate cancellation signals, and combined with adaptive echo cancellation, automatic gain control and nonlinear speech enhancement processing to generate high-fidelity audio signals.
While effectively suppressing environmental noise, it preserves the original audio signal components, providing a high-fidelity, distortion-free listening experience and improving the communication quality for users in noisy environments.
Smart Images

Figure CN121908191A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of headphone audio technology, and specifically relates to an audio signal processing method and system for headphones. Background Technology
[0002] Over-ear or in-ear headphones are indispensable personal audio devices in modern life, and their active noise cancellation (ANC) function has become a key feature to enhance the user experience. The basic principle of ANC technology is to use a reference microphone to collect external ambient noise, generate an anti-noise signal with the same amplitude but opposite phase through a built-in digital signal processor, and play it through the headphone speaker. The noise entering the ear is canceled out by the destructive interference of sound waves.
[0003] However, the noise cancellation performance of existing headphones is often unsatisfactory. The core reason is that most commercial ANC systems are based on linear algorithms, but the miniature speakers in headphones have an inherent saturation nonlinearity in actual operation. When the amplitude of the noise cancellation signal or music signal is large, this nonlinearity will cause the linear algorithm to fail, resulting in insufficient noise cancellation depth. Especially in noisy environments where strong noise cancellation is required, the user experience is greatly reduced. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides an audio signal processing method and system for headphones, which solves the technical problems in the prior art.
[0005] On the one hand, the invention provides the following technical solution: an audio signal processing method for headphones, the method comprising: Acquire environmental noise signals and user voice signals; The environmental noise signal and the audio signal to be played are input into a deep learning network model trained with psychoacoustic constraints to generate cancellation signals for active noise control. The cancellation signal, the user's voice signal, and the audio signal are mixed to generate an output signal that drives the speaker; Collect the residual error signal that is perceived by the error microphone after being played through the speaker; Based on the output signal and the residual error signal, the user speech signal is sequentially subjected to adaptive echo cancellation, automatic gain control and nonlinear speech enhancement processing to obtain an enhanced speech signal.
[0006] Compared to existing technologies, the advantages of this application are as follows: by inputting both ambient noise and the audio signal to be played into a deep learning network model trained with psychoacoustic constraints, a more accurate cancellation signal can be generated. This method overcomes the audio distortion problem caused by loudspeaker saturation effect in traditional nonlinear active noise control systems. While effectively suppressing ambient noise, it can retain the components of the original audio signal to the maximum extent, thereby providing users with a high-fidelity, distortion-free listening experience.
[0007] Furthermore, the deep learning network model is a psychoacoustic gated convolutional recurrent network, and the loss function used for its training is calculated as follows:
[0008] in, , It is the power spectrum of the estimation error. The masking threshold power spectrum is calculated based on the audio signal. It is a compensation factor. These are the number of frames and the number of frequencies, respectively.
[0009] Furthermore, the calculation of the masking threshold power spectrum includes the following steps: The frequency domain representation of the audio signal is converted to the Bark domain for critical band decomposition to obtain the energy of each critical band. Based on the energy of each critical band, the spectral flatness of the audio signal is calculated, and then the pitch coefficient is determined based on the spectral flatness. Based on the pitch coefficient, the masking energy offset value of each critical band is calculated, and the initial masking threshold is adjusted in combination with the absolute hearing threshold of the human ear to finally generate the masking threshold power spectrum.
[0010] Furthermore, after the step of acquiring the user's voice signal, the method further includes: Beamforming is applied to the user's voice signal to enhance the voice in the direction of the target sound source and suppress interference. The user's voice signal that is subsequently mixed and processed is the signal after beamforming.
[0011] Furthermore, the gain factor in the automatic gain control processing is dynamically adjusted based on the Mel-frequency cepstral coefficient cosine similarity between the user's speech signal and the target signal; the dynamic adjustment method includes: When the cosine similarity is higher than a preset threshold, the gain factor is set to 1; When the cosine similarity is lower than the preset threshold, the gain factor is increased; When the gain factor exceeds the upper limit of the dynamic range, the gain factor is compressed.
[0012] Furthermore, the gain factor iteration algorithm used in the automatic gain control processing is any one of the following: An adaptive algorithm based on envelope estimation and NLMS; An adaptive algorithm based on power estimation and NLMS; The MFCC-XNLMS algorithm based on envelope estimation and XE-NLMS.
[0013] Furthermore, the nonlinear speech enhancement processing includes: The user's voice signal is subjected to half-wave rectification nonlinear transformation to generate a harmonic regenerated signal; The harmonic regenerated signal is weighted and fused with the user's voice signal to reconstruct the suppressed harmonic components in the user's voice signal.
[0014] Furthermore, the expression for the half-wave rectifier nonlinear transformation function is as follows:
[0015] in, The input is the speech signal, and s(t) is the generated harmonic regeneration signal.
[0016] Furthermore, the input to the deep learning network model includes complex spectral features extracted from the environmental noise signal and the audio signal to be played, respectively.
[0017] Secondly, the invention provides the following technical solution: an audio signal processing system for headphones, the system comprising: The acquisition module is used to acquire environmental noise signals and user voice signals; The noise reduction module is used to input the environmental noise signal and the audio signal to be played into a deep learning network model trained with psychoacoustic constraints to generate a cancellation signal for active noise control. A mixing module is used to mix the cancellation signal, the user voice signal, and the audio signal to generate an output signal that drives the speaker; An error module is used to collect residual error signals that are sensed by an error microphone after being played through the speaker; The processing module is used to perform adaptive echo cancellation, automatic gain control and nonlinear speech enhancement processing on the user speech signal in sequence based on the output signal and the residual error signal to obtain an enhanced speech signal. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of an audio signal processing method for headphones provided in the first embodiment of the present invention; Figure 2 This is a structural block diagram of the audio signal processing system for headphones provided in the second embodiment of the present invention.
[0020] The embodiments of the present invention will be further described below with reference to the accompanying drawings. Detailed Implementation
[0021] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the embodiments of the present invention, and should not be construed as limiting the present invention.
[0022] In the description of the embodiments of the present invention, it should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention.
[0023] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0024] Example 1 The headphone system of this invention mainly includes: First microphone array: used to collect ambient noise signals.
[0025] Second microphone array: used to collect user voice signals.
[0026] Digital signal processor: This is the core processing unit used to run the audio signal processing method of this invention.
[0027] Speaker: Used to play the mixed audio signal.
[0028] Error microphone: Usually located inside the ear canal or near the speaker, used to collect residual error signals.
[0029] Wireless communication module: such as Bluetooth or Wi-Fi, used to communicate with the cloud or other devices.
[0030] In the first embodiment of the present invention, please refer to Figure 1 As shown, an audio signal processing method for headphones includes the following steps S01 to S05: S01, acquire ambient noise signal and user voice signal ; In this embodiment, ambient noise signals are collected through a first microphone array. ; The user's voice signal is acquired through a second microphone array.
[0031] Optionally, after the step of acquiring the user's voice signal, the method further includes: The user's voice signal is subjected to beamforming processing to enhance the voice in the direction of the target sound source and suppress interference. The processed signal is denoted as... ; The user's voice signal that is subsequently mixed and processed is the signal after beamforming.
[0032] S02, the environmental noise signal and the audio signal to be played are input into a deep learning network model trained with psychoacoustic constraints to generate a cancellation signal for active noise control; Specifically, the deep learning network model is a psychoacoustic gated convolutional recurrent network, and the loss function used for its training is calculated as follows:
[0033] in, , It is the power spectrum of the estimation error. The masking threshold power spectrum is calculated based on the audio signal. It is a compensation factor. These are the frame number and the frequency number, respectively. The calculation of the masking threshold power spectrum includes the following steps: The frequency domain representation of the audio signal is converted to the Bark domain for critical band decomposition to obtain the energy of each critical band. Based on the energy of each critical band, the spectral flatness of the audio signal is calculated, and then the pitch coefficient is determined based on the spectral flatness. Based on the pitch coefficient, the masking energy offset value of each critical band is calculated, and the initial masking threshold is adjusted in combination with the absolute hearing threshold of the human ear to finally generate the masking threshold power spectrum.
[0034] More specifically, the input to the deep learning network model includes complex spectral features extracted from the environmental noise signal and the audio signal to be played, respectively.
[0035] In this embodiment, the ambient noise signal and the audio signal to be played The input is fed into a pre-trained psychoacoustic gated convolutional recurrent network model, and the network finally outputs the cancellation signal. complex spectrum S03, the cancellation signal, the user voice signal and the audio signal are mixed to generate an output signal that drives the speaker; In this embodiment, the generated cancellation signal Audio signal to be played And (beam-shaped) user voice signal Digital mixing is performed to obtain the final output signal that drives the speaker. The signal is amplified and then played through a speaker. S04, Collect the residual error signal sensed by the error microphone after it has been played through the speaker; In this embodiment, an error microphone located inside the ear canal or near the speaker is used to collect the final auditory signal after the entire acoustic system has been processed, i.e., the residual error signal. The signal includes ambient noise. Output signal played by the speaker via secondary path The residual components after the effects of (speaker nonlinearity and cavity acoustics) are considered. That is:
[0036] in, The saturation nonlinearity effect of the loudspeaker was simulated.
[0037] S05, based on the output signal and the residual error signal, the user speech signal is sequentially subjected to adaptive echo cancellation, automatic gain control and nonlinear speech enhancement processing to obtain an enhanced speech signal.
[0038] Specifically, the gain factor in the automatic gain control process is dynamically adjusted based on the Mel-frequency cepstral coefficient cosine similarity between the user's speech signal and the target signal; the dynamic adjustment method includes: When the cosine similarity is higher than a preset threshold, the gain factor is set to 1; When the cosine similarity is lower than the preset threshold, the gain factor is increased; When the gain factor exceeds the upper limit of the dynamic range, the gain factor is compressed.
[0039] Specifically, the gain factor iteration algorithm used in the automatic gain control processing is any one of the following: An adaptive algorithm based on envelope estimation and NLMS; An adaptive algorithm based on power estimation and NLMS; The MFCC-XNLMS algorithm based on envelope estimation and XE-NLMS.
[0040] Specifically, the nonlinear speech enhancement processing includes: The user's voice signal is subjected to half-wave rectification nonlinear transformation to generate a harmonic regenerated signal; The harmonic regenerated signal is weighted and fused with the user's voice signal to reconstruct the suppressed harmonic components in the user's voice signal.
[0041] Specifically, the expression for the half-wave rectifier nonlinear transformation function is as follows:
[0042] in, The input is the speech signal, and s(t) is the generated harmonic regeneration signal.
[0043] In this embodiment; Echo cancellation: using the output signal sent to the speaker Using the reference signal, the acquired residual error signal As the desired signal, the filter-X least mean square algorithm is used to process the user's speech signal. Echo cancellation is performed, and the output signal has significantly suppressed echoes. This algorithm effectively estimates and subtracts the complex echo path from the loudspeaker to the error microphone, outputting a signal with significantly suppressed echoes. Here, the residual error signal... It is crucial because it accurately reflects the sound conditions that ultimately reach the human ear, containing the most accurate echo information.
[0044] Automatic gain control: for the signal after echo cancellation. The purpose of automatic gain control is to ensure that the processed speech signal remains stable and clear in a varying noise environment. The core of this invention's gain control lies in its gain factor. The update mechanism relies on the perceptual driving parameter of the cosine similarity of the Mel-cepstrum coefficients between the user's speech signal and a target signal.
[0045] Specifically, calculate the signal With an ideal target signal (or The cosine similarity of MFCC features in the quiet environment version.
[0046] Gain Factor Updates should follow these guidelines: If the similarity is greater than 0.98, the signal-to-noise ratio is considered high enough, and the setting is... .
[0047] If the similarity is ≤0.98, then according to the formula... Increase the gain factor, where The error is based on similarity calculation. The envelope of the input signal.
[0048] If the amplified signal If the magnitude exceeds 1, then for To compress, for example, let .
[0049] In this invention, the gain factor The adaptive iterative process can be implemented using the MFCC-XNLMS algorithm based on envelope estimation and XE-NLMS.
[0050] Nonlinear speech enhancement: enhancement of the amplified signal Perform half-wave rectification to generate harmonic regeneration signal :
[0051] Will With the original signal Perform weighted fusion: ,in This operation can effectively restore the high-frequency harmonics of speech, improving clarity and intelligibility.
[0052] Output and transmission: The final enhanced speech signal CELT encoding and AES encryption are performed. The data stream is transmitted to a mobile phone or cloud server via Bluetooth using the ASBP protocol for voice recognition or calls. Simultaneously, the user can hear the audio content mixed with the enhanced voice signal through the headset. In summary, the audio signal processing method for headphones has the following effects: By employing a deep learning network model trained with psychoacoustic constraints to generate the cancellation signal, this invention overcomes the audio distortion problem caused by speaker saturation in traditional nonlinear active noise control (ANC) systems. This method can intelligently preserve components of the original audio signal while deeply suppressing ambient noise, thus providing users with a high-fidelity, distortion-free listening experience and resolving the fundamental contradiction of the incompatibility between "noise reduction" and "sound quality" in traditional methods.
[0053] This invention creatively integrates active noise control and user voice communication at the signal level. By performing adaptive echo cancellation based on the final mixed signal, it can accurately estimate and eliminate the complex echo path from the speaker to the microphone, completely avoiding interference from the played audio to the user's voice. Combined with subsequent automatic gain control and nonlinear speech enhancement processing, the final output is a stable, echo-free, detailed, and natural enhanced speech signal, greatly improving the communication quality in two-way interactive scenarios such as phone calls and voice assistants.
[0054] By introducing beamforming into the speech path, this invention can directionally enhance speech in the direction of the user's mouth and effectively suppress interference noise from the environment and other directions. This significantly improves the input signal-to-noise ratio for subsequent processing. Combined with nonlinear harmonic regeneration technology, it can actively recover high-frequency harmonic components of speech lost during noise reduction and transmission, thereby greatly improving speech intelligibility and clarity, especially in noisy environments.
[0055] The Automatic Gain Control (AGC) module of this invention dynamically adjusts the gain based on the cosine similarity of Mel-frequency cepstral coefficients (MFCCs), rather than simply the signal energy. This makes the gain control more consistent with human hearing characteristics, preventing speech masking at low signal-to-noise ratios and avoiding saturation distortion caused by excessive gain at high signal-to-noise ratios. The entire system, driven by deep learning and perception, can adaptively cope with varying environmental noise and different audio content, exhibiting excellent robustness and intelligence.
[0056] Example 2 like Figure 2 As shown, in a second embodiment of the present invention, an audio signal processing system for headphones is provided, the system comprising: Acquisition module 10 is used to acquire environmental noise signals and user voice signals; The noise reduction module 20 is used to input the environmental noise signal and the audio signal to be played into a deep learning network model trained with psychoacoustic constraints to generate a cancellation signal for active noise control. The mixing module 30 is used to mix the cancellation signal, the user voice signal and the audio signal to generate an output signal that drives the speaker; Error module 40 is used to collect residual error signals that are sensed by error microphone after being played through the speaker; The processing module 50 is used to perform adaptive echo cancellation, automatic gain control and nonlinear speech enhancement processing on the user speech signal in sequence based on the output signal and the residual error signal to obtain an enhanced speech signal.
[0057] The audio signal processing system for headphones provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0058] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0059] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An audio signal processing method for headphones, characterized in that, The method includes: Acquire environmental noise signals and user voice signals; The environmental noise signal and the audio signal to be played are input into a deep learning network model trained with psychoacoustic constraints to generate cancellation signals for active noise control. The cancellation signal, the user's voice signal, and the audio signal are mixed to generate an output signal that drives the speaker; Collect the residual error signal that is perceived by the error microphone after being played through the speaker; Based on the output signal and the residual error signal, the user speech signal is sequentially subjected to adaptive echo cancellation, automatic gain control and nonlinear speech enhancement processing to obtain an enhanced speech signal.
2. The audio signal processing method for headphones according to claim 1, characterized in that, The deep learning network model is a psychoacoustic gated convolutional recurrent network, and its loss function is calculated as follows: in, , It is the power spectrum of the estimation error. The masking threshold power spectrum is calculated based on the audio signal. It is a compensation factor. These are the number of frames and the number of frequencies, respectively.
3. The audio signal processing method for headphones according to claim 2, characterized in that, The calculation of the masking threshold power spectrum includes the following steps: The frequency domain representation of the audio signal is converted to the Bark domain for critical band decomposition to obtain the energy of each critical band. Based on the energy of each critical band, the spectral flatness of the audio signal is calculated, and then the pitch coefficient is determined based on the spectral flatness. Based on the pitch coefficient, the masking energy offset value of each critical band is calculated, and the initial masking threshold is adjusted in combination with the absolute hearing threshold of the human ear to finally generate the masking threshold power spectrum.
4. The audio signal processing method for headphones according to claim 1, characterized in that, After the step of acquiring the user's voice signal, the method further includes: Beamforming is applied to the user's voice signal to enhance the voice in the direction of the target sound source and suppress interference. The user's voice signal that is subsequently mixed and processed is the signal after beamforming.
5. The audio signal processing method for headphones according to claim 1, characterized in that, The gain factor in the automatic gain control process is dynamically adjusted based on the cosine similarity of the Mel-frequency cepstral coefficients between the user's speech signal and the target signal. The dynamic adjustment methods include: When the cosine similarity is higher than a preset threshold, the gain factor is set to 1; When the cosine similarity is lower than the preset threshold, the gain factor is increased; When the gain factor exceeds the upper limit of the dynamic range, the gain factor is compressed.
6. The audio signal processing method for headphones according to claim 1, characterized in that, The gain factor iteration algorithm used in the automatic gain control process is any one of the following: An adaptive algorithm based on envelope estimation and NLMS; An adaptive algorithm based on power estimation and NLMS; The MFCC-XNLMS algorithm based on envelope estimation and XE-NLMS.
7. The audio signal processing method for headphones according to claim 1, characterized in that, The nonlinear speech enhancement processing includes: The user's voice signal is subjected to half-wave rectification nonlinear transformation to generate a harmonic regenerated signal; The harmonic regenerated signal is weighted and fused with the user's voice signal to reconstruct the suppressed harmonic components in the user's voice signal.
8. The audio signal processing method for headphones according to claim 7, characterized in that, The expression for the half-wave rectifier nonlinear transformation function is as follows: in, The input is the speech signal, and s(t) is the generated harmonic regeneration signal.
9. The audio signal processing method for headphones according to claim 1, characterized in that, The input to the deep learning network model includes complex spectral features extracted from the environmental noise signal and the audio signal to be played, respectively.
10. An audio signal processing system for headphones, characterized in that, The system includes: The acquisition module is used to acquire environmental noise signals and user voice signals; The noise reduction module is used to input the environmental noise signal and the audio signal to be played into a deep learning network model trained with psychoacoustic constraints to generate a cancellation signal for active noise control. A mixing module is used to mix the cancellation signal, the user voice signal, and the audio signal to generate an output signal that drives the speaker; An error module is used to collect residual error signals that are sensed by an error microphone after being played through the speaker; The processing module is used to perform adaptive echo cancellation, automatic gain control and nonlinear speech enhancement processing on the user speech signal in sequence based on the output signal and the residual error signal to obtain an enhanced speech signal.