Headphone call noise reduction methods and headphones
By utilizing the headset's dual-microphone acquisition and processing technology, combined with adaptive filtering and signal fusion, the problems of noise and interference with human voices during headset calls are solved, thereby improving call quality and user experience.
Patent Information
- Application Number
- CN202411915737.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-24
AI Technical Summary
During headset calls, external noise and interfering human voices severely affect call quality, leading to a decline in user experience.
The system uses the first microphone of the headphones to collect sound signals from inside the human ear and the second microphone to collect external sound signals. After framing, windowing, and Fourier transform, the audio frames are adjusted using adaptive filtering and attribute information. Combined with signal fusion processing, the system generates the target sound signal to optimize the noise reduction effect.
It effectively suppresses noise and interference with human voices, improves call clarity and intelligibility, restores the natural tone of the target speech, and enhances the user experience.
Smart Images

Figure CN119729287B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a method for noise reduction during headphone calls and a headphone. Background Technology
[0002] With the increasing popularity of headphones on the market, people are using them more and more to answer calls, make voice calls, and participate in meetings in offices, noisy shopping malls, and while riding public transportation. However, in actual call scenarios, there is a lot of external noise, such as wind noise, sudden mechanical noise, and other noises that interfere with human voices, causing the voice signal to be drowned out by background noise and seriously affecting the user's call experience. Summary of the Invention
[0003] This application provides a method for noise reduction during headphone calls and a headphone, which adjusts the third audio frame and the second audio frame based on the attribute information of the third audio frame to optimize the noise reduction effect of the headphone.
[0004] In a first aspect, a method for noise reduction during headphone calls is provided, comprising: obtaining a first sound signal inside a human ear through a first microphone of the headphone, and obtaining a second sound signal outside a human ear through a second microphone of the headphone; performing target conversion processing on the first sound signal to obtain a first audio frame, and performing the target conversion processing on the second sound signal to obtain a second audio frame, wherein the target conversion processing includes framing, windowing, and Fourier transform; performing adaptive filtering processing on the first audio frame based on the second audio frame to obtain a third audio frame; adjusting the third audio frame according to attribute information of the third audio frame to obtain an in-ear result sound signal, wherein the attribute information includes noise presence information, interference voice presence information, and target voice presence information; adjusting the second audio frame according to the attribute information to obtain an out-of-ear result sound signal; performing signal fusion processing on the in-ear result sound signal and the out-of-ear result sound signal, and generating a target sound signal to be output based on the signal fusion processing result.
[0005] In a second aspect, an earphone is provided, comprising: a first microphone for acquiring a first sound signal inside a human ear; a second microphone for acquiring a second sound signal outside a human ear; and a processing unit coupled to the first microphone and the second microphone, configured to perform the earphone call noise reduction method as described in any one of the first aspects.
[0006] By applying the above technical solution, a first sound signal inside the ear is obtained through the first microphone of the earphone, and a second sound signal outside the ear is obtained through the second microphone of the earphone. The first sound signal undergoes target conversion processing to obtain a first audio frame, and the second sound signal undergoes target conversion processing to obtain a second audio frame. The target conversion processing includes framing, windowing, and Fourier transform. Based on the second audio frame, the first audio frame undergoes adaptive filtering processing to obtain a third audio frame. The third audio frame is adjusted according to its attribute information to obtain the in-ear result sound signal. The attribute information includes noise presence information, interference voice presence information, and target voice presence information. The second audio frame is adjusted according to the attribute information to obtain the out-of-ear result sound signal. The in-ear and out-of-ear result sound signals are fused, and the target sound signal to be output is generated based on the signal fusion processing result. Because the attribute information includes noise presence information, interference voice presence information, and target voice presence information, noise reduction processing is performed based on the actual presence of noise and interference voices, thereby optimizing the noise reduction effect, improving the earphone call quality, and ultimately enhancing the user experience. Attached Figure Description
[0007] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 The flowchart of the headphone call noise reduction method according to an embodiment of this application is shown. Figure 1 ;
[0009] Figure 2 A flowchart illustrating the determination of attribute information of a third audio frame according to an embodiment of this application is shown;
[0010] Figure 3 A flowchart illustrating the adjustment of a third audio frame according to an embodiment of this application is shown;
[0011] Figure 4 This illustrates the process of determining the external auditory result sound signal according to an embodiment of this application. Figure 1 ;
[0012] Figure 5 This illustrates the process of determining the external auditory result sound signal according to an embodiment of this application. Figure 2 ;
[0013] Figure 6 The flowchart of the headphone call noise reduction method according to an embodiment of this application is shown. Figure 2 ;
[0014] Figure 7 The flowchart of the headphone call noise reduction method according to an embodiment of this application is shown. Figure 3 ;
[0015] Figure 8 A structural block diagram of an earphone according to an embodiment of this application is shown. Detailed Implementation
[0016] Various embodiments and features of this application are described herein with reference to the accompanying drawings.
[0017] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.
[0018] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0019] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.
[0020] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application.
[0021] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.
[0022] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this application, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.
[0023] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.
[0024] This application discloses a headphone call noise reduction method. It obtains a first sound signal inside the ear through a first microphone of the headphone and a second sound signal outside the ear through a second microphone. Based on the second audio frame corresponding to the second sound signal, it performs adaptive filtering on the first audio frame corresponding to the first sound signal to improve the signal-to-noise ratio of the first audio frame, thus obtaining a third audio frame. Then, based on the attribute information of the third audio frame, it adjusts the third audio frame and the second audio frame respectively to obtain the in-ear and out-of-ear result sound signals, which are then fused. The result of the signal fusion is used to generate the target sound signal to be output. Since the attribute information includes noise presence information, interference voice presence information, and target voice presence information, noise reduction processing is performed based on the actual presence of noise and interference voices, thereby optimizing the noise reduction effect and improving the headphone call quality.
[0025] Figure 1 The flowchart of the headphone call noise reduction method according to an embodiment of this application is shown. Figure 1 ,like Figure 1 As shown, it includes the following steps:
[0026] Step S101: Obtain a first sound signal inside the human ear through the first microphone of the earphone, and obtain a second sound signal outside the human ear through the second microphone of the earphone.
[0027] In this embodiment, the earphone is equipped with at least one first microphone and at least one second microphone. The first microphone and the second microphone can be digital microphones or analog microphones.
[0028] When a user wears the headphones, the first microphone is located inside the ear canal. The signal collected by the first microphone is mainly the target speech signal, but it will also leak in varying degrees of ambient noise depending on how it is worn. The signal characteristics are high low-frequency signal-to-noise ratio and severe attenuation of high-frequency components. The second microphone is generally exposed to the external acoustic environment and can collect the voices of people around and complex ambient noise. The user's voice is transmitted to the second microphone through the air.
[0029] It can obtain the first sound signal inside the human ear through the first microphone and the second sound signal outside the human ear through the second microphone, either in real time or when the user is making a call through the headset.
[0030] In some embodiments of this application, the second microphone can be a single Talk microphone, a single FF microphone (i.e., a feedforward microphone), or a dual microphone. In the case of a dual microphone, the signals from the dual microphones are processed by beamforming to obtain the second sound signal.
[0031] Step S102: Perform target conversion processing on the first audio signal to obtain a first audio frame, and perform the same target conversion processing on the second audio signal to obtain a second audio frame. The target conversion processing includes framing, windowing, and Fourier transform.
[0032] In this embodiment, to facilitate subsequent noise reduction processing, the first audio signal and the second audio signal are framed, windowed, and Fourier transformed respectively, and the first audio signal and the second audio signal are converted to the frequency domain to obtain the first audio frame corresponding to the first audio signal and the second audio frame corresponding to the second audio signal.
[0033] Step S103: Perform adaptive filtering on the first audio frame based on the second audio frame to obtain the third audio frame.
[0034] In this embodiment, based on the second audio frame, the first audio frame is adaptively filtered using an adaptive filter to obtain the third audio frame. This allows for dynamic adjustment of the first audio frame according to the second audio frame, eliminating environmental noise leaking into the first microphone due to improper user fit (or other varying degrees of tightness), and improving the signal-to-noise ratio of the first audio frame.
[0035] Step S104: Adjust the third audio frame according to the attribute information of the third audio frame to obtain the in-ear result sound signal. The attribute information includes noise presence information, interference voice presence information, and target voice presence information.
[0036] In this embodiment, the noise is ambient noise, the interfering voice is any human voice other than the user's voice, and the target voice is the user's voice. The attribute information of the third audio frame is determined, including information on the presence of noise, the presence of interfering voice, and the presence of the target voice. The third audio frame is then adjusted based on this attribute information to reduce noise and interfering voice in the third audio frame. The adjusted third audio frame is then used as the resulting in-ear sound signal.
[0037] Step S105: Adjust the second audio frame according to the attribute information to obtain the external sound signal.
[0038] The second audio frame is adjusted according to the attribute information to reduce noise and interference with human voices in the second audio frame, and the adjusted second audio frame is used as the external sound signal.
[0039] It is understandable that the execution order of steps S104 and S105 is not important.
[0040] Step S106: Perform signal fusion processing on the in-ear result sound signal and the out-of-ear result sound signal, and generate the target sound signal to be output based on the signal fusion processing result.
[0041] In this embodiment, after obtaining the intra-auricular result sound signal and the extra-auricular result sound signal, signal fusion processing is performed on the intra-auricular result sound signal and the extra-auricular result sound signal. For example, the intra-auricular result sound signal and the extra-auricular result sound signal can be fused directly, or the intra-auricular result sound signal and the extra-auricular result sound signal can be weighted and then fused. Then, the target sound signal to be output is generated according to the signal fusion processing result, and the target sound signal can be output as an uplink speech signal.
[0042] In some embodiments of this application, generating a target sound signal to be output based on the signal fusion processing result includes: performing an inverse Fourier transform on the signal fusion processing result to obtain the target sound signal, thereby ensuring that the subsequent target sound signal is output in the form of a time-domain signal.
[0043] The headphone call noise reduction method of this application embodiment obtains a first sound signal inside the ear through a first microphone of the headphone and a second sound signal outside the ear through a second microphone of the headphone; performs target conversion processing on the first sound signal to obtain a first audio frame, and performs target conversion processing on the second sound signal to obtain a second audio frame. The target conversion processing includes framing, windowing, and Fourier transform; performs adaptive filtering processing on the first audio frame based on the second audio frame to obtain a third audio frame; adjusts the third audio frame according to the attribute information of the third audio frame to obtain an in-ear result sound signal. The attribute information includes noise presence information, interference voice presence information, and target voice presence information; adjusts the second audio frame according to the attribute information to obtain an out-of-ear result sound signal; performs signal fusion processing on the in-ear result sound signal and the out-of-ear result sound signal, and generates a target sound signal to be output based on the signal fusion processing result. Since the attribute information includes information on the presence of noise, the presence of interfering human voices, and the presence of the target human voice, noise reduction processing is performed based on the actual presence of noise and interfering human voices. This effectively suppresses loud noise and interfering human voices, improves the clarity and intelligibility of calls, and perfectly restores the natural tone of the target speech, thus enhancing the user experience.
[0044] In some embodiments of this application, the noise presence information includes whether it belongs to a pure noise frame, the interfering human voice presence information includes whether it belongs to a pure interfering human voice frame, and the target human voice presence information includes whether it belongs to a target human voice frame. Figure 2 A flowchart illustrating the determination of attribute information of a third audio frame according to an embodiment of this application is provided. Before adjusting the third audio frame based on the attribute information of the third audio frame, as follows... Figure 2 As shown, it also includes the following steps:
[0045] Step S107: Perform speech activity detection on the third audio frame to determine the probability of speech presence in the frame corresponding to the third audio frame.
[0046] In this embodiment, noise estimation can be performed on the third audio frame, and the probability of the presence of speech in the third audio frame can be determined based on the noise estimation result. Alternatively, a neural network can be used to process the third audio frame to output the result of the speech activity and determine the probability of the presence of speech in the frame.
[0047] In some embodiments of this application, voice activity detection of the third audio frame may specifically include voice activity detection of each frequency point in the third audio frame, determining the voice presence probability of each frequency point, and then determining the frame voice presence probability based on the voice presence probability of each frequency point. For example, the average value of the voice presence probability of each frequency point can be determined as the frame voice presence probability.
[0048] Step S108: Determine whether the current audio frame belongs to the pure noise frame based on the probability of the presence of the frame speech.
[0049] In this embodiment, the current audio frame is the audio frame corresponding to the third and second audio frames, meaning the third and second audio frames should have the same attribute information. If the current audio frame is a pure noise frame, it means the wearer is not speaking and only ambient noise is present. The probability of speech presence in the frame can be used to determine whether the current audio frame is a pure noise frame.
[0050] Step S109: Determine the first auto-power spectral density corresponding to the third audio frame and the second auto-power spectral density corresponding to the second audio frame.
[0051] In this embodiment, the presence of interfering human voices is determined by comparing the energy of the third audio frame and the second audio frame. Auto-power spectral density is a function reflecting the power distribution of a random signal in the frequency domain; it describes the power density of the signal at different frequencies, i.e., the power per unit frequency range. Various methods can be used to determine auto-power spectral density, such as the periodogram method, the average periodogram method, the Welch method, and the autocorrelation method.
[0052] The first autopower spectral density corresponding to the third audio frame and the second autopower spectral density corresponding to the second audio frame are determined, and then interference voices are detected using the first autopower spectral density and the second autopower spectral density.
[0053] Step S110: Determine whether the current audio frame belongs to the pure interference human voice frame based on the ratio between the first autopower spectral density and the second autopower spectral density.
[0054] In this embodiment, a pure interference voice frame means that the current audio frame contains only interference voices. The ratio between the first autopower spectral density and the second autopower spectral density is determined, and the current audio frame is determined to be a pure interference voice frame based on the ratio.
[0055] Step S111: Determine whether the current audio frame belongs to the target human voice frame based on the probability of the presence of the frame speech and the ratio.
[0056] In this embodiment, the target human voice frame means that the current audio frame is the target speech emitted by the user. The current audio frame is determined to be a target human voice frame based on the probability and ratio of the existence of the speech in the frame.
[0057] The probability of the presence of a speech frame is determined by speech activity detection, and interference voices are detected by the self-power spectral density corresponding to the third and second audio frames, thereby achieving efficient and accurate determination of the attribute information of the third audio.
[0058] In some embodiments of this application, the current audio frame belongs to the pure noise frame when the probability of the presence of the frame speech is less than a first probability, the current audio frame belongs to the pure interference voice frame when the ratio is less than a first ratio, and the current audio frame belongs to the target voice frame when the probability of the presence of the frame speech is greater than a second probability and the ratio is greater than a second ratio, wherein the first probability is less than the second probability and the first ratio is less than the second ratio.
[0059] The first self-power spectral density is the self-power spectral density of the third audio frame within the target frequency range, and the second self-power spectral density is the self-power spectral density of the second audio frame within the target frequency range.
[0060] In this embodiment, the probability of the presence of speech in a frame is compared with a first probability and a second probability. The ratio is then compared with both the first and second ratios. If the probability of the presence of speech in a frame is less than the first probability, it indicates that the probability of speech presence in the current frame is low, and the current audio frame is determined to be a pure noise frame. If the ratio is less than the first ratio, it indicates that the target human voice is not present, and the current audio frame is determined to be a pure interference human voice frame. If the probability of the presence of speech in a frame is greater than the second probability, and the ratio is greater than the second ratio, it indicates that the target human voice is present, and the current audio frame is determined to be a target human voice frame, thus achieving more efficient determination of attribute information.
[0061] Furthermore, the target frequency range can be determined based on the frequency range of the wearer's jaw vibration signal and the sensitivity of the first microphone. For example, if the frequency range of the jaw vibration signal during normal human speech is between 100Hz and 1.5kHz, then the target frequency range could be 100Hz to 1.5kHz. Self-power spectral density deviating from this target frequency range is not processed. Since the first self-power spectral density is the self-power spectral density of the third audio frame within the target frequency range, and the second self-power spectral density is the self-power spectral density of the second audio frame within the target frequency range, the calculation of ratios for self-power spectral densities outside the target frequency range is avoided. This reduces the computational load and avoids misjudgments caused by high-frequency signal attenuation and noise from the first microphone, significantly improving the accuracy of detecting interference-laden human voices.
[0062] In some embodiments of this application, the ratio is determined by Formula 1, which is specifically as follows:
[0063]
[0064] Where Energy_diff is the ratio, Φ FB (ω) is the first self-power spectral density, Φ OUT (ω) is the second self-power spectral density, ω is the angular frequency; δ is a small quantity greater than 0 to avoid division by zero; ω1 and ω2 are the upper and lower limits of the target frequency range, respectively, the minimum value of ω1 is 0, and the maximum value of ω2 is 1 / 2 of the number of FFT (Fast Fourier Transform) points.
[0065] The ratio is determined by Formula 1, which further improves the accuracy of detecting interfering human voices.
[0066] In some embodiments of this application, Figure 3 The flowchart illustrates an embodiment of this application showing the adjustment of a third audio frame. The third audio frame is adjusted based on its attribute information, such as... Figure 3 As shown, it includes the following steps:
[0067] Step S1041: If the target human voice existence information belongs to the target human voice frame, determine the target frequency point in each frequency point of the third audio frame, and the probability of speech existence corresponding to the target frequency point is greater than the third probability.
[0068] In this embodiment, when the target human voice presence information belongs to the target human voice frame, the probability of speech presence at each frequency point in the third audio frame is determined, and the frequency points with a speech presence probability greater than the third probability are determined as target frequency points. Subsequently, only the target frequency points are subjected to adaptive gain compensation processing.
[0069] Optionally, the third probability can be the same as or different from the first probability.
[0070] Step S1042: Perform adaptive gain compensation processing on the target frequency point in the third audio frame based on the second audio frame.
[0071] In this embodiment, adaptive gain compensation processing is performed on the target frequency point in the third audio frame based on the second audio frame, so as to automatically adjust the gain of the target frequency point according to the frequency response of each frequency point in the second audio frame, so that the gain of the target frequency point is kept at a suitable level.
[0072] By identifying the target frequency, noise frequencies in the third audio frame can be avoided from being amplified. Adaptive gain compensation is then applied only to the target frequency to perfectly restore the natural timbre of the target voice, thereby improving the user's call experience.
[0073] In some embodiments of this application, Figure 4 This illustrates the process of determining the external auditory result sound signal according to an embodiment of this application. Figure 1 The second audio frame is adjusted according to the attribute information to obtain the external auditory result sound signal, such as... Figure 4 As shown, it includes the following steps:
[0074] Step S1051: If the noise presence information indicates that the frame belongs to the pure noise frame, update the adaptive filtering coefficients of the second audio frame.
[0075] In this embodiment, if the noise information indicates that it belongs to a pure noise frame, it means that the current audio frame is a pure noise frame, and the adaptive filtering coefficients of the second audio frame are updated.
[0076] Step S1052: Based on the updated adaptive filtering coefficients, perform adaptive filtering on the second audio frame to obtain the fourth audio frame.
[0077] In this embodiment, the second audio frame is adaptively filtered using an adaptive filter based on the updated adaptive filtering coefficients, thereby accurately and effectively filtering out noise in the second audio frame and obtaining the fourth audio frame.
[0078] Step S1053: Determine the external sound signal based on the information about the presence of the interfering human voice and the fourth audio frame.
[0079] Based on the presence of interfering human voices, the fourth audio frame is further adjusted to determine the external sound signal.
[0080] By updating the adaptive filtering coefficients of the second audio frame when the noise presence information belongs to a pure noise frame, the updated adaptive filtering coefficients are matched with the actual noise presence, thereby filtering out the noise in the second audio frame more accurately. Furthermore, by interfering with the human voice presence information, the interfering human voice in the fourth audio frame can be effectively suppressed, thus improving the quality of the external sound signal.
[0081] In some embodiments of this application, Figure 5 This illustrates the process of determining the external auditory result sound signal according to an embodiment of this application. Figure 2 The external auditory result sound signal is determined based on the information about the presence of the interfering human voice and the fourth audio frame, such as... Figure 5 As shown, it includes the following steps:
[0082] Step S10531: If the information of the presence of interfering human voice is a pure interfering human voice frame, the fourth audio frame is subjected to interfering human voice suppression processing to obtain a first gain.
[0083] In this embodiment, if the information of the interfering human voice is a pure interfering human voice frame, then the current audio frame is determined to be a pure interfering human voice frame, and the fourth audio frame is subjected to interfering human voice suppression processing to obtain a first gain, which is less than 1. In some embodiments of this application, if the information of the interfering human voice is not a pure interfering human voice frame, then the first gain is 1.
[0084] In an environment with multiple speakers, the second microphone indiscriminately captures the voices of each person, but only the target voice of the wearer is the signal needed. When there are only interfering voices, they are suppressed, thereby correctly outputting the voice of the target person and ensuring the accuracy of the call.
[0085] Step S10532: Perform nonlinear echo cancellation processing on the fourth audio frame to obtain a second gain.
[0086] By performing nonlinear echo cancellation processing on the fourth audio frame, a second gain can be output.
[0087] Step S10533: Perform single-microphone noise suppression processing on the fourth audio frame to obtain the third gain.
[0088] By performing single-microphone noise suppression processing on the fourth audio frame, a third gain can be output.
[0089] Step S10534: Determine the target gain based on the minimum value among the first gain, the second gain, and the third gain.
[0090] The minimum value among the first, second, and third gains is determined and set as the target gain, which can effectively suppress residual echoes, noise, and interfering human voices in the fourth audio frame.
[0091] Step S10535: Adjust the fourth audio frame according to the target gain to determine the external sound signal.
[0092] In this embodiment, steps S10531-S10533 do not actually adjust the fourth audio frame, but only obtain the corresponding gain. After determining the target gain in step S10534, the gain of the fourth audio frame is adjusted according to the target gain. This can eliminate environmental noise and interference with human voice in the external microphone signal, and will not cause human voice distortion due to over-suppression, thereby restoring a clear target human voice and further improving the quality of the external sound signal.
[0093] In some embodiments of this application, it also includes:
[0094] If the noise presence information indicates that the noise does not belong to the pure noise frame, the current adaptive filtering coefficients of the second audio frame remain unchanged.
[0095] Based on the current adaptive filtering coefficients, the second audio frame is subjected to adaptive filtering to obtain the fourth audio frame.
[0096] In this embodiment, if the noise information indicates that it does not belong to a pure noise frame, it means that there is target speech in the current audio frame. In this case, the filter coefficients of the current adaptive filter are kept unchanged, and the second audio frame is adaptively filtered using the adaptive filter to obtain the fourth audio frame, thereby avoiding human voice distortion caused by misprocessing.
[0097] In some embodiments of this application, after obtaining the third audio frame, at least one of the following is also included:
[0098] The third audio frame is subjected to single-microphone noise suppression processing;
[0099] The third audio frame is subjected to residual nonlinear echo cancellation processing.
[0100] In this embodiment, residual noise in the third audio frame can be eliminated by performing single-microphone noise suppression processing on the third audio frame. Residual echo signal in the third audio frame can be eliminated by performing residual nonlinear echo cancellation processing on the third audio frame.
[0101] In some embodiments of this application, single-microphone noise suppression processing can be performed using either DSP (digital signal processing) noise reduction or neural network noise reduction. Since the signal-to-noise ratio of the third audio frame obtained after adaptive filtering is relatively high, DSP noise reduction is generally sufficient.
[0102] In some embodiments of this application, the first microphone is a feedback microphone or a bone conduction microphone. Before performing adaptive filtering on the first audio frame based on the second audio frame to obtain the third audio frame, the method further includes:
[0103] When the first microphone is the feedback microphone, the first audio frame is subjected to fixed filtering processing using a fixed filter according to the second audio frame, and linear echo cancellation processing is performed on the first audio frame after fixed filtering processing. The coefficients of the fixed filter are generated by pure noise signals collected by the first microphone and the second microphone under different active noise cancellation modes.
[0104] When the first microphone is the bone conduction microphone, linear echo cancellation processing is performed on the first audio frame.
[0105] In this embodiment, the feedback microphone has active noise cancellation functionality. If the first microphone is a feedback microphone, the coefficients of a fixed filter are pre-generated using the pure noise signals collected by the first and second microphones under different active noise cancellation modes (including transparency mode, noise cancellation mode, and off mode). After obtaining the first audio frame, the first audio frame is subjected to fixed filtering processing using the fixed filter based on the second audio frame. This eliminates external noise leaking into the first audio frame under different active noise cancellation modes, ensuring that the signal-to-noise ratio of the first microphone remains consistent under different active noise cancellation modes, especially in transparency mode, so that call quality and listening experience are not affected by switching active noise cancellation modes. Since the first microphone is very close to the speaker, the collected echo signal is large and mostly linear echoes. Performing linear echo cancellation processing on the first audio frame can eliminate most of its linear echoes.
[0106] A bone conduction microphone, also known as a VPU microphone, is a type of bone conduction microphone. If the first microphone is a bone conduction microphone, since bone conduction microphones are not affected by active noise cancellation, there is no need to perform fixed filtering on the first audio frame. Instead, linear echo cancellation is performed on the first audio frame to eliminate most of the linear echo.
[0107] To further illustrate the technical concept of this application, the technical solution will now be explained in conjunction with specific application scenarios.
[0108] This application provides a method for noise reduction during headphone calls, wherein the headphone includes a first microphone and a second microphone. Figure 6 The flowchart of the headphone call noise reduction method according to an embodiment of this application is shown. Figure 2 For the first microphone, such as Figure 6 As shown, it includes the following steps:
[0109] Step S11, fixed filtering process.
[0110] In this embodiment, the first microphone is a feedback microphone. Before step S11, the first sound signal inside the human ear is obtained through the first microphone, and the second sound signal outside the human ear is obtained through the second microphone. The first sound signal and the second sound signal are framed, windowed, and Fourier transformed respectively to convert the first sound signal and the second sound signal to the frequency domain, thereby obtaining the first audio frame corresponding to the first sound signal and the second audio frame corresponding to the second sound signal.
[0111] The coefficients of a fixed filter are generated in advance by collecting pure noise signals from the first and second microphones under different active noise cancellation modes (including transparency mode, noise cancellation mode, and off mode). After obtaining the first audio frame, the first audio frame is subjected to fixed filtering processing using the fixed filter according to the second audio frame to eliminate external noise leaking into the first audio frame under different active noise cancellation modes. This ensures that the signal-to-noise ratio of the first microphone remains consistent under different active noise cancellation modes, especially in transparency mode, so that the call quality and listening experience are not affected by the switching of active noise cancellation modes.
[0112] Step S12, linear echo cancellation processing.
[0113] Because the first microphone is very close to the speaker, the collected echo signal is large and mostly linear echoes. Performing linear echo cancellation processing on the first audio frame can eliminate the vast majority of linear echoes in the first audio frame.
[0114] Step S13, adaptive filtering processing.
[0115] In this embodiment, based on the second audio frame, the first audio frame is adaptively filtered using an adaptive filter to obtain the third audio frame. This allows for dynamic adjustment of the first audio frame according to the second audio frame, eliminating environmental noise leaking into the first microphone due to improper user fit (or other varying degrees of tightness), and improving the signal-to-noise ratio of the first audio frame.
[0116] Step S14, voice activity detection.
[0117] The probability of speech presence in the third audio frame can be determined by performing noise estimation on the third audio frame and analyzing the results. Alternatively, a neural network can be used to process the third audio frame to output the result of the speech activity and determine the probability of speech presence in the third audio frame. Based on the probability of speech presence, it can be determined whether the current audio frame belongs to the pure noise frame.
[0118] Step S15, interference voice detection.
[0119] The first self-power spectral density corresponding to the third audio frame and the second self-power spectral density corresponding to the second audio frame are determined. The ratio between the first and second self-power spectral densities is used to determine whether the current audio frame belongs to a purely interfering voice frame. Here, the first self-power spectral density is the self-power spectral density of the third audio frame within the target frequency range, and the second self-power spectral density is the self-power spectral density of the second audio frame within the target frequency range. This avoids calculating the ratio of self-power spectral densities outside the target frequency range, thus reducing the computational load and avoiding misjudgments caused by high-frequency signal attenuation and noise from the first microphone, significantly improving the accuracy of interfering voice detection. Specifically, the ratio is determined using Formula 1, which is as follows:
[0120]
[0121] Where Energy_diff is the ratio, Φ FB (ω) is the first self-power spectral density, Φ OUT (ω) is the second self-power spectral density, ω is the angular frequency; δ is a small quantity greater than 0 to avoid division by zero; ω1 and ω2 are the upper and lower limits of the target frequency range, respectively, the minimum value of ω1 is 0, and the maximum value of ω2 is 1 / 2 of the number of FFT points.
[0122] The probability of the presence of speech in a frame is compared with a first probability and a second probability. The ratio of this ratio is then compared with both the first and second ratios. If the probability of the presence of speech in a frame is less than the first probability, it indicates a low probability of speech presence in the current frame, and the current audio frame is determined to be a pure noise frame. If the ratio is less than the first ratio, it indicates the absence of the target human voice, and the current audio frame is determined to be a pure interference human voice frame. If the probability of the presence of speech in a frame is greater than the second probability, and the ratio is greater than the second ratio, it indicates the presence of the target human voice, and that noise and interference human voice are relatively few, and the current audio frame is determined to be a target human voice frame.
[0123] Step S16, single-microphone noise suppression processing.
[0124] By performing single-microphone noise suppression processing on the third audio frame, residual noise in the third audio frame can be eliminated.
[0125] Step S17, residual nonlinear echo elimination.
[0126] By performing residual nonlinear echo cancellation on the third audio frame, the residual echo signal in the third audio frame can be eliminated.
[0127] Step S18, adaptive gain compensation processing.
[0128] If the target voice presence information belongs to the target voice frame, the probability of speech presence at each frequency point in the third audio frame is determined. Frequency points with a speech presence probability greater than the third probability are identified as target frequency points. Adaptive gain compensation processing is performed on the target frequency points in the third audio frame based on the second audio frame to obtain the in-ear result sound signal. The gain of the target frequency points is then automatically adjusted according to the frequency response of each frequency point in the second audio frame to maintain the gain of the target frequency points at an appropriate level.
[0129] By identifying the target frequency, noise frequencies in the third audio frame can be avoided from being amplified. Adaptive gain compensation is then applied only to the target frequency to perfectly restore the natural timbre of the target voice, thereby improving the user's call experience.
[0130] Figure 7 The flowchart of the headphone call noise reduction method according to an embodiment of this application is shown. Figure 3 For the second microphone, such as Figure 7 As shown, it includes the following steps:
[0131] Step S21: Determine whether the current frame is a pure noise frame.
[0132] In this embodiment, the current frame is the current audio frame. If the probability of speech in the frame is less than the first probability, it means that the probability of speech in the current frame is small, and the current audio frame is determined to be a pure noise frame.
[0133] Step S22, adaptive filtering processing.
[0134] If the noise presence information indicates that the frame belongs to the pure noise frame, the adaptive filtering coefficients of the second audio frame are updated. Based on the updated adaptive filtering coefficients, the second audio frame is subjected to adaptive filtering processing using an adaptive filter, thereby accurately and effectively filtering out the noise in the second audio frame and obtaining the fourth audio frame.
[0135] Step S23: Determine whether the current frame is a pure interference human voice frame.
[0136] If the ratio is less than the first ratio, it means that there is no target human voice, and the current audio frame is determined to be a pure interference human voice frame.
[0137] Step S24, interference with human voice suppression processing.
[0138] If the current audio frame is a pure interference voice frame, interference voice suppression processing is performed on the fourth audio frame to obtain a first gain, which is less than 1. In some embodiments of this application, if the current audio frame is not a pure interference voice frame, the first gain is 1.
[0139] Step S25, nonlinear echo cancellation processing.
[0140] By performing nonlinear echo cancellation processing on the fourth audio frame, a second gain can be output.
[0141] Step S26, single-microphone noise suppression processing.
[0142] The fourth audio frame is subjected to single-microphone noise suppression processing, which can output a third gain. Here, single-microphone neural network noise reduction is generally used.
[0143] The minimum value among the first gain, second gain, and third gain is determined, and this minimum value is set as the target gain. The fourth audio frame is adjusted according to the target gain to determine the external sound signal, thereby effectively suppressing residual echoes, noise, and interfering human voices in the fourth audio frame.
[0144] Step S27, signal fusion processing.
[0145] The results of the in-ear and out-of-ear sound signals are fused together. The result of the signal fusion is then subjected to inverse Fourier transform to obtain the target sound signal, thereby ensuring that the target sound signal is output in the form of a time-domain signal.
[0146] The headphone call noise reduction method in this application embodiment can effectively solve problems such as noise leakage, timbre changes, and unstable listening experience caused by variations in the tightness of the user's wearing, different ANC (Active Noise Cancellation) modes, and changes in environmental noise, resulting in the use of the first microphone signal. This significantly improves the voice quality of the wearer (including noise reduction, voice clarity, smoothness, and naturalness). Specifically, it can include the following technical effects:
[0147] 1. The embodiments of this application utilize a first microphone to collect the speech signal received in the ear canal, isolate external environmental noise, and have a high signal-to-noise ratio, which can significantly improve speech clarity and intelligibility.
[0148] 2. In this embodiment of the application, the coefficients of a fixed filter are generated from the pure noise signals collected by the first microphone and the second microphone under different active noise cancellation modes. The first microphone data is then filtered by the fixed filter to eliminate external noise leaked into the first microphone signal under different ANC modes. This ensures that the signal-to-noise ratio of the first microphone remains consistent under different ANC modes (especially transparency mode), thereby ensuring that the call quality and listening experience are not affected by ANC mode switching.
[0149] 3. In this embodiment of the application, linear echo cancellation is performed on the signal of the first microphone to ensure that the judgment of the target speech signal is not affected when the first microphone signal contains a large echo signal.
[0150] 4. In this embodiment of the application, the signal of the first microphone is adaptively filtered by the noise signal collected by the second microphone to eliminate the environmental noise leaked into the first microphone due to poor user wearing (or other different wearing tightness), thereby further improving the signal-to-noise ratio of the target signal in the signal of the first microphone.
[0151] 5. In this embodiment, the signal from the first microphone after filtering and linear echo cancellation is used to detect speech activity in order to obtain the probability of speech presence in the current audio frame. This ensures accurate speech activity detection even in noisy environments, where the signal from the external microphone array is submerged in noise.
[0152] 6. In this embodiment of the application, the energy ratio of the first microphone and the second microphone in a specific frequency band is used as a threshold to determine whether there is a target human voice in the current frame or only interfering human voice. The signal is detected and the interfering human voice is suppressed, especially the interfering human voice directly in front can be eliminated.
[0153] 7. In this embodiment of the application, the signal from the second microphone is used to perform adaptive gain compensation on the signal from the first microphone to eliminate the changes in listening experience caused by the characteristics of the first microphone signal, such as the heavier low-frequency components and the severe attenuation of high-frequency components. This ensures that even if the ambient noise changes continuously, the target human voice output after signal fusion is stable and the timbre is natural.
[0154] 8. After the processed signal from the first microphone is combined with the signal from the second microphone for signal fusion, the clarity of the low-frequency range of speech can be improved under strong non-stationary noise conditions such as high noise and wind noise, which greatly improves the intelligibility of speech during call quality.
[0155] This application also provides an earphone. Figure 8 A structural block diagram of an earphone according to an embodiment of this application is shown, such as... Figure 8As shown, it includes: a first microphone for collecting a first sound signal inside a human ear; a second microphone for collecting a second sound signal outside a human ear; and a processing unit coupled to the first microphone and the second microphone, configured to perform the headphone call noise reduction method as described in various embodiments of this application.
[0156] It is understood that, in the embodiments of this application, the headphones may also include components such as a shell, a Bluetooth module, a memory, and a speaker, but this is not a limitation.
[0157] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0158] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A method for noise reduction during headphone calls, characterized in that, include: The first sound signal inside the human ear is obtained through the first microphone of the earphone, and the second sound signal outside the human ear is obtained through the second microphone of the earphone. The first audio signal is subjected to target conversion processing to obtain a first audio frame, and the second audio signal is subjected to the same target conversion processing to obtain a second audio frame. The target conversion processing includes framing, windowing, and Fourier transform. The first audio frame is adaptively filtered based on the second audio frame to obtain the third audio frame. The third audio frame is adjusted according to the attribute information of the third audio frame to obtain the in-ear result sound signal. The attribute information includes noise presence information, interference voice presence information and target voice presence information. The second audio frame is adjusted according to the attribute information to obtain the external sound signal; The in-ear and out-of-ear sound signals are fused together, and the target sound signal to be output is generated based on the fusion result.
2. The headphone call noise reduction method as described in claim 1, characterized in that, The noise presence information includes whether it belongs to a pure noise frame; the interfering human voice presence information includes whether it belongs to a pure interfering human voice frame; the target human voice presence information includes whether it belongs to a target human voice frame; before adjusting the third audio frame according to the attribute information of the third audio frame, the method further includes: Speech activity detection is performed on the third audio frame to determine the probability of the presence of the corresponding frame speech. Determine whether the current audio frame belongs to the pure noise frame based on the probability of the presence of the frame speech; Determine the first auto-power spectral density corresponding to the third audio frame and the second auto-power spectral density corresponding to the second audio frame; Whether the current audio frame belongs to the pure interference human voice frame is determined based on the ratio between the first autopower spectral density and the second autopower spectral density. Whether the current audio frame belongs to the target human voice frame is determined based on the probability of the presence of the frame speech and the ratio.
3. The headphone call noise reduction method as described in claim 2, characterized in that, The current audio frame belongs to the pure noise frame when the probability of the presence of the frame speech is less than the first probability; the current audio frame belongs to the pure interference voice frame when the ratio is less than the first ratio; the current audio frame belongs to the target voice frame when the probability of the presence of the frame speech is greater than the second probability and the ratio is greater than the second ratio, wherein the first probability is less than the second probability and the first ratio is less than the second ratio. The first self-power spectral density is the self-power spectral density of the third audio frame within the target frequency range, and the second self-power spectral density is the self-power spectral density of the second audio frame within the target frequency range.
4. The headphone call noise reduction method as described in claim 2, characterized in that, Adjusting the third audio frame according to its attribute information includes: If the target human voice exists and the information belongs to the target human voice frame, the target frequency point in each frequency point of the third audio frame is determined, and the probability of the speech corresponding to the target frequency point is greater than the third probability. Adaptive gain compensation processing is performed on the target frequency point in the third audio frame based on the second audio frame.
5. The headphone call noise reduction method as described in claim 2, characterized in that, The second audio frame is adjusted according to the attribute information to obtain an external sound signal, including: If the noise presence information indicates that the frame belongs to the pure noise frame, update the adaptive filtering coefficients of the second audio frame. Based on the updated adaptive filtering coefficients, the second audio frame is subjected to adaptive filtering to obtain the fourth audio frame; The external auditory result sound signal is determined based on the information about the presence of the interfering human voice and the fourth audio frame.
6. The headphone call noise reduction method as described in claim 5, characterized in that, Determining the external auditory result sound signal based on the presence information of the interfering human voice and the fourth audio frame includes: If the information about the presence of interfering human voices is a pure interfering human voice frame, the fourth audio frame is subjected to interfering human voice suppression processing to obtain a first gain; The fourth audio frame is subjected to nonlinear echo cancellation processing to obtain a second gain; The fourth audio frame is subjected to single-microphone noise suppression processing to obtain a third gain; The target gain is determined based on the minimum value among the first gain, the second gain, and the third gain; The fourth audio frame is adjusted according to the target gain to determine the external auditory result sound signal.
7. The headphone call noise reduction method as described in claim 5, characterized in that, Also includes: If the noise presence information indicates that the noise does not belong to the pure noise frame, the current adaptive filtering coefficients of the second audio frame remain unchanged. Based on the current adaptive filtering coefficients, the second audio frame is subjected to adaptive filtering to obtain the fourth audio frame.
8. The headphone call noise reduction method as described in claim 1, characterized in that, After obtaining the third audio frame, at least one of the following is also included: The third audio frame is subjected to single-microphone noise suppression processing; The third audio frame is subjected to residual nonlinear echo cancellation processing.
9. The headphone call noise reduction method as described in claim 1, characterized in that, The first microphone is a feedback microphone or a bone conduction microphone. Before performing adaptive filtering on the first audio frame based on the second audio frame to obtain the third audio frame, the method further includes: When the first microphone is the feedback microphone, the first audio frame is subjected to fixed filtering processing using a fixed filter according to the second audio frame, and linear echo cancellation processing is performed on the first audio frame after fixed filtering processing. The coefficients of the fixed filter are generated by pure noise signals collected by the first microphone and the second microphone under different active noise cancellation modes. When the first microphone is the bone conduction microphone, linear echo cancellation processing is performed on the first audio frame.
10. An earphone, characterized in that, include: The first microphone is used to collect the first sound signal inside the human ear; The second microphone is used to collect a second sound signal from outside the human ear; The processing unit, coupled to the first microphone and the second microphone, is configured to perform the headphone call noise reduction method as described in any one of claims 1-9.
Citation Information
Patent Citations
Smart call noise reduction method and system of feedback type earphone
CN115884032A
Audio signal processing method and device, earphone equipment and storage medium
CN116528099A