Call noise cancelling method and earphone

By employing multiple microphones to detect and adapt to wind noise conditions, the method enhances audio clarity in earphones by switching to the microphone with better signal-to-noise ratio and using AI processing to cancel wind noise, addressing the ineffectiveness of conventional methods in loud wind environments.

JP2025160901AActive Publication Date: 2025-10-23ANKER INNOVATIONS TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025063734
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2025-04-08
Publication Date
2025-10-23
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Conventional call noise cancellation methods using AI models are ineffective in environments with loud wind noise, leading to low audio signal clarity due to the inability to distinguish wind noise from audio signals effectively.

Method used

The method employs multiple microphones in earphones to detect wind noise by comparing energy levels and coherence data, switching to the microphone with better signal-to-noise ratio, and using an AI noise canceling model to process the target audio signal, which includes replacing wind noise frequency bands to improve clarity.

Benefits of technology

Enhances audio signal clarity by effectively canceling wind noise, improving voice quality during calls by utilizing multiple microphones and AI processing to adapt to wind noise conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160901000001_ABST
    Figure 2025160901000001_ABST
Patent Text Reader

Abstract

To provide a call noise cancelling method and an earphone that can improve clarity of voice signals.SOLUTION: The present application relates to a call noise cancelling method and an earphone. The method includes the steps of: when there is wind noise in an external environment, acquiring an energy of a first voice signal collected by a first microphone and an energy of a second voice signal collected by a second microphone; determining the second voice signal as a target voice signal when a difference between the energy of the first voice signal and the energy of the second voice signal is larger than a predetermined threshold value; determining the target voice signal based on the first voice signal when the difference between the energy of the first voice signal and the energy of the second voice signal is smaller than or equal to the predetermined threshold value; and subjecting the target voice signal to noise cancelling processing for the wind noise. The present method can improve clarity of the voice signals.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the technical field of earphones, and in particular to a call noise canceling method and earphones. [Background technology]

[0002] Call noise cancellation effectively suppresses noise during calls, reduces external noise interference, and better captures people's voices, thereby improving voice quality.Wind noise is a special type of noise, and currently, noise cancellation is achieved by using an AI (artificial intelligence) model to remove wind noise from the audio signal collected by the main microphone in earphones.

[0003] However, when the above noise canceling method is used, there is a problem that the clarity of the audio signal is low because the noise canceling effect is low in an environment where wind noise is loud. Summary of the Invention

[0004] Based on this, it is necessary to provide a call noise canceling method and earphone that can improve the clarity of audio signals in order to address the above technical problems.

[0005] In a first aspect, an embodiment of the present application provides a method for canceling call noise, the method being used in an earphone including a first earphone having a first microphone and a second earphone having a second microphone, When wind noise is present in the external environment, acquiring energy of a first audio signal collected by the first microphone and energy of a second audio signal collected by the second microphone; determining the second audio signal as a target audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold; determining the target audio signal based on the first audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is less than or equal to the predetermined threshold; and performing wind noise canceling processing on the target audio signal.

[0006] In one embodiment, the first earphone further includes a third microphone, and the method further comprises: The method further includes determining whether wind noise is present in the external environment based on the first audio signal and a third audio signal collected by the third microphone.

[0007] In one embodiment, the step of determining whether wind noise is present in the external environment based on the first audio signal and the third audio signal collected by the third microphone includes: calculating coherence data of the first audio signal and the third audio signal based on the first audio signal and the third audio signal; and determining whether wind noise is present in the external environment based on the coherence data.

[0008] In one embodiment, before calculating coherence data for the first audio signal and the third audio signal, The method further includes the step of performing a signal delay process on the first audio signal or the third audio signal to make the first audio signal and the third audio signal in phase.

[0009] In one embodiment, the coherence data comprises coherence values ​​for a plurality of different frequency bands, and determining whether wind noise is present in the external environment based on the coherence data comprises: Determining that wind noise is present in the external environment if the coherence value for at least one frequency band is less than a threshold value corresponding to the frequency band.

[0010] In one embodiment, the step of determining the target audio signal based on the first audio signal comprises: The method includes replacing a wind noise frequency band in the first audio signal to obtain the target audio signal.

[0011] In one embodiment, the earphone further includes a feedback microphone, and the step of replacing a wind noise frequency band in the first audio signal to obtain the target audio signal includes: determining wind noise start and stop frequency points based on each of said coherence values; extracting a wind noise signal matching the start and stop frequency points from the fourth audio signal collected by the feedback microphone; and using the extracted wind noise signal to replace a wind noise frequency band in the first audio signal to obtain the target audio signal.

[0012] In one embodiment, the earphone is a headphone and the feedback microphone is located in the first earphone.

[0013] In one embodiment, the step of performing wind noise canceling processing on the target audio signal includes: The method includes a step of inputting the target sound signal into an AI noise canceling model to perform noise canceling processing for wind noise.

[0014] In a second aspect, an embodiment of the present application provides earphones, the earphones including a memory storing a computer program and a processor, a first earphone provided with a first microphone, and a second earphone provided with a second microphone, and the processor, when executing the computer program, performs the steps of the method according to the first aspect.

[0015] In the above-mentioned call noise canceling method and earphones, the earphones include a first earphone equipped with a first microphone and a second earphone equipped with a second microphone. When wind noise is present in the external environment, the method acquires the energy of a first audio signal collected by the first microphone and the energy of a second audio signal collected by the second microphone, and then determines the magnitude relationship between the energy of the first audio signal and the energy of the second audio signal. Since the wind noise signal has corresponding energy, if the difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold, i.e., if the energy of the first audio signal is significantly greater than the energy of the second audio signal, it indicates that the first audio signal contains large wind noise. In this case, the second audio signal can be used as a target audio signal. When the difference between the energy of the first audio signal and the energy of the second audio signal is equal to or less than the predetermined threshold, i.e., if the energy of the first audio signal is significantly greater than the energy of the second audio signal, it indicates that the first audio signal contains large wind noise. In this case, the second audio signal can be used as a target audio signal. If the energy of the first audio signal is not significantly greater than the energy of the second audio signal (for example, if the energy of the first audio signal corresponds to the energy of the second audio signal, or if the energy of the first audio signal is smaller than the energy of the second audio signal), it indicates that the wind noise contained in the first audio signal corresponds to the wind noise contained in the second audio signal, or if the wind noise contained in the first audio signal is smaller than the wind noise contained in the second audio signal, there is no need to use the second audio signal as the target audio signal, and the target audio signal is determined based on the first audio signal, and for example, the first audio signal is directly used as the target audio signal, or the target audio signal is obtained by performing noise canceling (or replacement) on the wind noise frequency band in the first audio signal, and then performing wind noise noise canceling processing on the target audio signal, thereby improving the effect of wind noise canceling and improving the clarity of the audio signal after noise cancellation.

[0016] In order to more clearly describe the technical solutions in the embodiments of the present application or related art, the following will briefly describe the drawings necessary for describing the embodiments or related art. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative work. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a diagram illustrating an application environment of a call noise canceling method according to an embodiment. [Figure 2] 1 is a flowchart of a call noise canceling method according to an embodiment. [Figure 3] FIG. 1 is a schematic diagram of a network structure of an exemplary AI noise canceling model according to another embodiment. [Figure 4] 4 is a flowchart of a call noise canceling method according to another embodiment. [Figure 5] 10 is a flowchart of step 205 according to another embodiment. [Figure 6] 10 is a flowchart for replacing a wind noise frequency band in a first audio signal according to another embodiment. [Figure 7] 4 is a flowchart of a call noise canceling method according to another embodiment. [Figure 8] 1 is a structural block diagram of a call noise canceling device according to an embodiment; [Figure 9] 1 is a diagram illustrating the internal structure of an earphone according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0018] In order to make the objectives, technical means and advantages of the present application clearer, the present application will be described in more detail below with reference to the drawings and examples. It will be understood that the specific examples described herein are merely illustrative of the present application and are not intended to limit the present application.

[0019] Call noise cancellation can effectively suppress noise during the call process, reduce the interference of external noise, and better capture people's voices, thereby improving the voice quality.

[0020] Conventional call noise cancellation methods often use a beamforming + AI (Artificial Intelligence) model to remove noise. The beamforming method removes noise by introducing phase and correlation between multiple microphones, but wind noise is a special type of noise that has no correlation between each microphone, so using the beamforming method will damage the audio signal.

[0021] In view of this, in the prior art, after detecting the presence of wind noise, it is common to prohibit the use of the beamforming method and directly use an AI model to remove wind noise from the audio signal collected by the main microphone in the earphone, thereby achieving noise cancellation.

[0022] However, wind noise has the characteristic that energy attenuates from low to high frequencies, and low frequencies are also a frequency band with high audio energy, so the audio signal becomes buried in the wind noise and becomes indistinguishable. This makes it impossible to effectively distinguish between wind noise and audio signals using an AI model, and this method has the problem of low noise cancellation effectiveness in environments with loud wind noise, resulting in low audio signal clarity.

[0023] The call noise canceling method according to the embodiment of the present application may be applied to the application environment shown in Fig. 1. The earphones include a first earphone 102 provided with a first microphone (not shown in Fig. 1) and a second earphone 104 provided with a second microphone (not shown in Fig. 1), and may be, for example, headphones or may be called headphones.

[0024] The first microphone may be a talk mic or a feedforward mic (FF mic), and the second microphone may be a talk mic or a FF mic, etc.

[0025] As shown in FIG. 2, the call noise canceling method according to an exemplary embodiment is described as being applied to the earphone in FIG. 1 as an example, and includes the following steps 201 to 204.

[0026] In step 201, when wind noise is present in the external environment, the energy of a first audio signal collected by a first microphone and the energy of a second audio signal collected by a second microphone are obtained.

[0027] In the process of voice pickup, the earphone collects a first voice signal of the user through a first microphone and collects a second voice signal of the user through a second microphone.

[0028] The first audio signal may be obtained by performing a Short-Time Fourier Transform (STFT) on the original audio signal collected by the first microphone, and the second audio signal may be obtained by performing a Short-Time Fourier Transform on the original audio signal collected by the second microphone.

[0029] The earphones then determine whether wind noise is present in the external environment, which refers to the environment in which the earphones are located.

[0030] In a possible embodiment, the earphones can themselves detect whether wind noise is present in the external environment. Illustratively, the earphones can detect wind noise using multiple microphones in the same earphone.

[0031] Preferably, the earphone detects wind noise using multiple microphones in the first earphone, for example, the first earphone further includes a third microphone, the first microphone is for example a talk mic, and the third microphone is for example an FF mic, and the earphone detects whether wind noise exists in the external environment based on the first audio signal collected by the first microphone and the third audio signal collected by the third microphone, and the detection process is described in the following embodiments.

[0032] Preferably, similar to the above method of detecting wind noise, the earphone may detect wind noise using multiple microphones in the second earphone. In this way, the distance between the multiple microphones in the same earphone is short, which is advantageous for improving the accuracy of wind noise detection because the influence of the user's head is small.

[0033] In another possible embodiment, the earphone may be communicatively connected to another earphone, which may be, for example, a smartphone, a tablet computer, a smart watch, a smart bracelet, etc., and the distance between the earphone and the other earphone is less than a predetermined distance threshold, i.e., the distance between the earphone and the other earphone is close, so that the other earphone can detect whether wind noise exists in the surrounding environment, and if wind noise exists, send a notification message to the earphone indicating that wind noise exists in the external environment.

[0034] When the earphone determines that wind noise is present in the external environment, it obtains the energy of the first audio signal and the energy of the second audio signal.

[0035] In the embodiment of the present application, the earphone can use a signal energy calculation formula to calculate the energy of the first audio signal and the energy of the second audio signal respectively.

[0036] In step 202, if the difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold, the second audio signal is determined as the target audio signal.

[0037] As can be seen, when the first and second earphones are located on either side of the user's head and there is no wind noise, the energy of the first audio signal should correspond to the energy of the second audio signal.

[0038] The wind noise signal also has a corresponding energy, and when the wind has a certain direction and the difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold, i.e., when the energy of the first audio signal is significantly greater than the energy of the second audio signal, it indicates that the first audio signal contains large wind noise, i.e., when the first microphone is facing the direction of the wind, the second audio signal can be taken as the target audio signal, i.e., the audio signal between the first audio signal and the second audio signal that contains small wind noise is taken as the target audio signal.

[0039] In this way, in a scene where wind noise is present, assuming that the first microphone is facing the wind direction, the signal-to-noise ratio of the other microphone (second microphone) is significantly better than that of the first microphone due to the user's head occlusion, and thus by switching to the second audio signal collected by the second microphone, the clarity of the audio can be significantly improved and the amount of noise cancellation can be reduced. Preferably, to further improve the clarity of the audio signal, noise cancellation (or replacement) can be performed on the wind noise frequency band in the second audio signal to obtain the target audio signal.

[0040] In step 203, if the difference between the energy of the first audio signal and the energy of the second audio signal is less than or equal to a predetermined threshold, a target audio signal is determined based on the first audio signal.

[0041] If the difference between the energy of the first audio signal and the energy of the second audio signal is below a predetermined threshold, i.e., if the energy of the first audio signal is not significantly greater than the energy of the second audio signal, several possible situations can be distinguished:

[0042] 1) The energy of the first audio signal corresponds to the energy of the second audio signal. In this case, if the wind noise contained in the first audio signal corresponds to the wind noise contained in the second audio signal, there is no need to switch to the second audio signal, and the target audio signal is determined based on the first audio signal, and the target audio signal is obtained by, for example, performing noise cancellation (or replacement) on the wind noise frequency band in the first audio signal.

[0043] 2) The energy of the first audio signal is significantly less than the energy of the second audio signal. In this case, the second audio signal contains relatively loud wind noise, i.e., the second microphone is facing the direction of the wind, and in this case, the first audio signal is used as the target audio signal, i.e., of the first audio signal and the second audio signal, the audio signal containing relatively small wind noise may be used as the target audio signal.

[0044] In this way, in a scene where wind noise is present, assuming that the second microphone is facing the wind direction, the signal-to-noise ratio of the front microphone (first microphone) is significantly better than that of the second microphone due to the user's head occlusion, and thus, by using the first audio signal collected by the first microphone, the clarity of the audio can be significantly improved and the amount of noise cancellation can be reduced. Preferably, to further improve the clarity of the audio signal, noise cancellation (or replacement) can be performed on the wind noise frequency band in the first audio signal to obtain a target audio signal.

[0045] In step 204, wind noise canceling processing is performed on the target audio signal.

[0046] For example, the earphones can input a target audio signal into an AI noise canceling model to perform wind noise cancellation processing.

[0047] In one embodiment of the present application, the AI ​​noise canceling model may use a time-frequency domain coupled network structure, as shown in FIG. 3, which is a schematic diagram of an exemplary AI noise canceling model network structure.

[0048] The earphones input the target audio signal as input to the first-layer network, which uses a multi-layer convolutional neural network (CNN) to extract time-frequency domain features of the target audio signal. The time-frequency domain features extracted by the multi-layer CNN are then input to the second-layer network, which uses a multi-layer recurrent neural network (RNN) to extract signal timing features. The signal timing features extracted by the multi-layer RNN are then input to the third fully connected layer (FC) to obtain a mask. Finally, the target audio signal and the mask are multiplied and an inverse short-time Fourier transform (ISTFT) is performed to obtain the output signal, i.e., the audio signal after wind noise cancellation processing.

[0049] In the above embodiment, when wind noise is present in the external environment, the energy of the first audio signal collected by the first microphone and the energy of the second audio signal collected by the second microphone are acquired, and then the magnitude relationship between the energy of the first audio signal and the energy of the second audio signal is determined. Since the wind noise signal has corresponding energy, if the difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold, i.e., if the energy of the first audio signal is significantly greater than the energy of the second audio signal, it indicates that the first audio signal contains a large amount of wind noise. In this case, the second audio signal can be used as the target audio signal. If the difference between the energy of the first audio signal and the energy of the second audio signal is equal to or less than a predetermined threshold, i.e., if the energy of the first audio signal is not significantly greater than the energy of the second audio signal (for example, When the signal energy corresponds to the energy of the second audio signal, or when the energy of the first audio signal is smaller than the energy of the second audio signal, it indicates that the wind noise contained in the first audio signal corresponds to the wind noise contained in the second audio signal, or when the wind noise contained in the first audio signal is smaller than the wind noise contained in the second audio signal, there is no need to use the second audio signal as the target audio signal, and the target audio signal is determined based on the first audio signal, and for example, the first audio signal is directly used as the target audio signal, or the target audio signal is obtained by performing noise canceling (or replacement) on the wind noise frequency band in the first audio signal, and then performing wind noise noise canceling processing on the target audio signal, thereby improving the effect of wind noise canceling and improving the clarity of the audio signal after noise cancellation.

[0050] In one embodiment, based on the embodiment shown in Fig. 2, the embodiment relates to a process in which an earphone detects whether wind noise exists in the external environment, as shown in Fig. 4. In this embodiment, the first earphone further includes a third microphone, and as shown in Fig. 4, the call noise canceling method of this embodiment further includes step 205 shown in Fig. 4.

[0051] In step 205, it is determined whether wind noise is present in the external environment based on the first audio signal and the third audio signal collected by the third microphone.

[0052] In the embodiment of the present application, the first microphone provided in the first earphone may be a talk microphone, and the third microphone may be an FF microphone; naturally, the first microphone may be an FF microphone, and the third microphone may be a talk microphone.

[0053] For example, if the first microphone is a talk microphone and the third microphone is an FF microphone, during the user's call, the first microphone collects the first audio signal and the third microphone collects the third audio signal. Since wind noise is a special noise, when wind noise exists, and the wind noise is particularly loud, the correlation between each microphone is weak. Then, the earphone can detect whether wind noise exists in the external environment based on the correlation between the first audio signal and the third audio signal.

[0054] Of course, the second earphone may further include a fourth microphone, and the earphone may determine whether or not wind noise exists in the external environment based on the second audio signal and the audio signal collected by the fourth microphone. There is no specific limitation on which earphone's microphones the earphone uses to detect wind noise.

[0055] The following is an exemplary description of a process in which the earphone determines whether wind noise exists in the external environment based on the first audio signal and the third audio signal collected by the third microphone.

[0056] As shown in FIG. 5, step 205 includes step 501 and step 502 shown in FIG.

[0057] In step 501, coherence data of the first and third audio signals is calculated based on the first and third audio signals.

[0058] In step 502, it is determined whether wind noise is present in the external environment based on the coherence data.

[0059] In a possible embodiment, the earphone can directly calculate the coherence data of the first audio signal and the third audio signal based on Equation 1.

[0060]

number

[0061] Next, the earphone substitutes τ(k,1) calculated using Equation 1 into Equation 2, and the calculated MSC(k,1) is the coherence data of the first audio signal and the third audio signal.

[0062] MSC(k,1)=|τ(k,1)| 2 formula 2 Due to the high-frequency attenuation characteristics of wind noise and the non-correlation characteristics between the microphone, there is a clear difference in the MSC distribution between the audio signal and the wind noise signal. By setting an appropriate wind noise threshold, if the coherence data MSC(k,1) of the first audio signal and the third audio signal is smaller than the wind noise threshold, it can be determined that wind noise exists and the wind noise threshold is, for example, 0.1, 0.2, etc.

[0063] In another possible embodiment, for the first audio signal and the third audio signal, the earphone may divide the first audio signal into multiple audio signal segments and divide the third audio signal into multiple audio signal segments according to frequency bands, respectively.

[0064] For example, by dividing the first audio signal and the third audio signal into a plurality of audio signals each corresponding to one frequency band, such as 0 Hz to 500 Hz, 500 Hz to 1000 Hz, and 1000 Hz to 1500 Hz, the first audio signal and the third audio signal are divided into a plurality of audio signals each corresponding to one of the frequency bands.

[0065] Next, for each frequency band, the earphone can calculate the coherence value between the first audio signal and the audio signal corresponding to that frequency band, and the coherence value between the third audio signal and the audio signal corresponding to that frequency band (i.e., MSC(k,1)) using the above Equation 1 and Equation 2. In this way, after the calculation is completed, the coherence value of each frequency band can be obtained and coherence data can be constructed, that is, the coherence data includes coherence values ​​of multiple different frequency bands.

[0066] In this way, if the coherence value for at least one frequency band is less than the threshold value corresponding to that frequency band, the earphone can determine that wind noise is present in the external environment, and similarly implement a process for determining whether wind noise is present in the external environment based on the coherence data.

[0067] In one embodiment, before calculating the coherence data of the first audio signal and the third audio signal, the earphone may perform a signal delay process on the first audio signal or the third audio signal to make the first audio signal and the third audio signal in phase. Exemplarily, since the positions of the first microphone and the third microphone in the earphone are relatively fixed, the time delay between the audio signals received by both microphones is also determined during the audio pickup process, for example, the audio signal collected by the first microphone leads the audio signal collected by the third microphone by half a phase. Thus, the earphone may delay the first audio signal by half a phase to make the first audio signal and the third audio signal in phase. For example, the audio signal collected by the third microphone is half-phase ahead of the audio signal collected by the first microphone. In this way, the earphone can delay the third audio signal by half-phase so that the first audio signal and the third audio signal are in phase. By making the first audio signal and the third audio signal in phase and calculating the coherence data of the first audio signal and the third audio signal, the accuracy of the coherence data can be improved, and the accuracy of wind noise detection can be improved.

[0068] In this embodiment, the first microphone and the third microphone are multiple microphones in the earphone on the same side, and the distance between the multiple microphones in the earphone on the same side is short, which is advantageous for improving the accuracy of wind noise detection as it reduces the influence of the user's head.

[0069] In one embodiment, based on the embodiment shown in FIG. 4 and FIG. 5, this embodiment relates to a process in which the earphone determines a target audio signal based on a first audio signal.

[0070] In this embodiment, if the difference between the energy of the first audio signal and the energy of the second audio signal is equal to or less than a predetermined threshold, the earphone replaces the wind noise frequency band in the first audio signal to obtain the target audio signal.

[0071] When the difference between the energy of the first audio signal and the energy of the second audio signal is equal to or less than a predetermined threshold, when the wind noise volume contained in the first audio signal is equal to the wind noise volume contained in the second audio signal, or when the wind noise volume contained in the first audio signal is smaller than the wind noise volume contained in the second audio signal, the earphones do not need to switch to the second audio signal and determine the target audio signal based on the first audio signal. The earphones can obtain the target audio signal by performing pre-noise cancellation processing for wind noise on the first audio signal. For example, the earphones obtain the target audio signal by replacing the frequency band of wind noise in the first audio signal, i.e., by replacing the frequency band in which wind noise exists in the first audio signal, they achieve pre-noise cancellation processing for wind noise.

[0072] The process by which the earphone replaces the wind noise frequency band in the first audio signal will now be described.

[0073] In this embodiment, the earphone further includes a feedback microphone (FB mic), which may be provided in the first earphone (naturally, a feedback microphone may be provided in the second earphone, and in this embodiment, the feedback microphone on the same side as the first microphone, i.e., the feedback microphone provided in the first earphone, is used). As shown in Fig. 6, the earphone can realize a process of replacing the wind noise frequency band in the first audio signal to obtain a target audio signal through steps 601 to 603 shown in Fig. 6.

[0074] In step 601, the start and stop frequency points of wind noise are determined based on each coherence value.

[0075] The start frequency point of wind noise is 0 Hz. In the embodiment shown in Figures 4 and 5, for each frequency band, the earphone calculates the coherence value of each frequency band and compares the coherence value of each frequency band with the threshold value corresponding to that frequency band, thereby determining whether wind noise exists in each frequency band. In order from the lowest frequency to the highest frequency of the frequency band, the stop frequency point of the last frequency band in which wind noise exists is set as the stop frequency point of wind noise.

[0076] For example, if the earphone determines, according to the above embodiment, that wind noise exists in the 0 Hz to 500 Hz frequency band, in the 500 Hz to 1000 Hz frequency band, and that wind noise does not exist in the 1000 Hz to 1500 Hz frequency band, it determines that the wind noise stop frequency point is 1000 Hz.

[0077] In step 602, a wind noise signal matching the start and stop frequency points is extracted from the fourth audio signal collected by the feedback microphone.

[0078] Unlike a system in which the call microphone and the feedforward microphone are located outside the earphone, the feedback microphone is located inside the earphone, which reduces the impact of wind noise in the environment. When the user makes a call or speaks, the microphones located in the first earphone and the second earphone collect the user's voice signal, respectively. Here, the voice signal collected by the feedback microphone is referred to as a fourth voice signal.

[0079] After determining the start and stop frequency points of the wind noise, the earphones determine the frequency band in which the wind noise is located. For example, if the start frequency point of the wind noise is 0 Hz and the stop frequency point of the wind noise is 1000 Hz, the frequency band in which the wind noise is located is 0 Hz to 1000 Hz.

[0080] The earphone extracts a signal in a frequency band in which wind noise exists from the fourth audio signal collected by the feedback microphone, and acquires a wind noise signal.

[0081] In step 603, the extracted wind noise signal is used to replace the wind noise frequency band in the first audio signal to obtain a target audio signal.

[0082] Next, the earphone uses the extracted wind noise signal to replace the wind noise frequency band (e.g., 0 Hz to 1000 Hz) formed by the start and stop frequency points of the wind noise in the first audio signal, to obtain the target audio signal.

[0083] In this way, since the feedback microphone is located inside the earphone and the impact of wind noise in the environment is small, the wind volume in the wind noise signal extracted from the fourth audio signal is necessarily smaller than the wind volume in the wind noise frequency band in the first audio signal. By combining the wind noise signal in the wind noise frequency band from the feedback microphone with the first audio signal, the target audio signal can be obtained and processed using AI noise canceling, improving the clarity of the audio and the noise canceling effect.

[0084] Hereinafter, an embodiment of the call noise canceling method of the present application will be exemplarily described with reference to a practical scene.

[0085] Assuming that the earphones are headphones, the first microphone provided in the first earphone is a talk mic, the third microphone provided in the first earphone is an FF mic, and the first earphone is further provided with a feedback microphone FB mic; assuming that the first earphone is a front earphone, the second earphone is a counterpart earphone and is provided with a second microphone, which may be a talk mic or an FF mic; in order to distinguish it from the talk mic and FF mic in the front earphone, the second microphone will be collectively referred to as the counterpart mic here.

[0086] As shown in FIG. 7, in the call noise canceling method of this embodiment, 1) The earphone performs STFT transformation on the audio signals collected by the talk mic, FF mic, FB mic and the other party's mic to obtain a first audio signal corresponding to the talk mic, a second audio signal corresponding to the other party's mic, a third audio signal corresponding to the FF mic and a fourth audio signal corresponding to the FB mic.

[0087] 2) The earphone delays the third audio signal collected by the FF mic to make it in phase with the first audio signal collected by the Talk mic.

[0088] 3) The earphone determines whether wind noise is present in the external environment based on the first audio signal and the third audio signal.

[0089] For the embodiment of step 3), reference can be made to the second embodiment in the example shown in Figure 5 above. That is, the first audio signal and the third audio signal are divided into a plurality of audio signal segments according to frequency bands, and for each frequency band, the MSC(k,1) corresponding to that frequency band is calculated. Whether wind noise exists in each frequency band is determined based on the threshold value corresponding to each frequency band. If wind noise exists in at least one frequency band, it is determined whether wind noise exists in the external environment, and the start and stop frequency points of the wind noise can be determined.

[0090] 4) When wind noise is present in the external environment, the energy of the first audio signal corresponding to the talk mic is calculated, the energy of the second audio signal corresponding to the other party's mic is calculated, and the magnitude relationship between the energy of the first audio signal and the energy of the second audio signal is determined.

[0091] 5) If the magnitude relationship is such that the energy of the first audio signal is significantly greater than the energy of the second audio signal (i.e., the difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold), the second audio signal is switched to be used as the target audio signal, and the target audio signal is input into the AI ​​noise canceling model to perform wind noise canceling processing.

[0092] 6) If the magnitude relationship is such that the difference between the energy of the first audio signal and the energy of the second audio signal is equal to or less than the predetermined threshold, a wind noise signal is extracted from the fourth audio signal corresponding to the FB mic, and the extracted wind noise signal is used to replace the wind noise frequency band in the first audio signal corresponding to the Talk mic to obtain a target audio signal, which is input into an AI noise canceling model to perform wind noise cancellation processing.

[0093] This embodiment realizes wind noise detection and cancellation for headphones with multiple microphones. The beneficial effects of the embodiment of the present application in improving wind noise cancellation are shown below with reference to the effect diagrams.

[0094] As can be understood, although the steps in the flowcharts according to the above embodiments are indicated in order by arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated otherwise in this specification, the execution of these steps is not limited to a strict order, and these steps may be executed in other orders. Furthermore, at least some of the steps in the flowcharts according to the above embodiments may include multiple steps or multiple stages, and these steps or stages may not necessarily be executed at the same time but may be executed at different times. The execution order of these steps or stages is also not necessarily sequential, and they may be executed in order or alternately with other steps or at least some of the steps or stages of other steps.

[0095] Based on the same inventive idea, the embodiments of the present application further provide a call noise canceling device that realizes the above call noise canceling method. The means for solving the problems related to the device are similar to the means for realizing the above method, so that the specific limitations of one or more of the following call noise canceling device embodiments can refer to the limitations of the above call noise canceling method, and further description will be omitted here.

[0096] As shown in FIG. 8 , a call noise canceling device according to an exemplary embodiment is provided in an earphone including a first earphone provided with a first microphone and a second earphone provided with a second microphone, an acquisition module 801 for acquiring the energy of a first audio signal collected by the first microphone and the energy of a second audio signal collected by the second microphone when wind noise is present in the external environment; a determination module 802 for determining the second audio signal as a target audio signal when a difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold, and determining the target audio signal based on the first audio signal when a difference between the energy of the first audio signal and the energy of the second audio signal is equal to or less than the predetermined threshold; and a noise canceling module 803 that performs noise canceling processing of wind noise on the target audio signal.

[0097] In one embodiment, the first earphone further includes a third microphone, and the device further includes: The system further includes a detection module that determines whether wind noise exists in an external environment based on the first audio signal and a third audio signal collected by the third microphone.

[0098] In one embodiment, the detection module comprises: a calculation unit for calculating coherence data of the first audio signal and the third audio signal based on the first audio signal and the third audio signal; a detection unit for determining whether wind noise is present in the external environment based on the coherence data.

[0099] In one embodiment, the device comprises: The audio signal processing device further includes a delay module that performs a signal delay process on the first audio signal or the third audio signal to make the first audio signal and the third audio signal in phase.

[0100] In one embodiment, the coherence data includes coherence values ​​of a plurality of different frequency bands, and the detection unit specifically determines that wind noise exists in the external environment when the coherence value of at least one frequency band is less than a threshold value corresponding to the frequency band.

[0101] In one embodiment, the determination module 802 specifically replaces a wind noise frequency band in the first audio signal to obtain the target audio signal.

[0102] In one embodiment, the earphone further includes a feedback microphone, and the determination module 802 further specifically determines wind noise start and stop frequency points based on each of the coherence values, extracts a wind noise signal matching the start and stop frequency points from the fourth audio signal collected by the feedback microphone, and uses the extracted wind noise signal to replace the wind noise frequency band in the first audio signal to obtain the target audio signal.

[0103] In one embodiment, the earphone is a headphone, and the feedback microphone is located in the first earphone.

[0104] In one embodiment, the noise canceling module 803 specifically inputs the target audio signal into an AI noise canceling model to perform noise canceling processing for wind noise.

[0105] All or part of the modules in the call noise canceling device may be realized by software, hardware, or a combination thereof. The modules may be built into or independent of the processor in the earphones in hardware form, or may be stored in the memory in the earphones in software form, so as to facilitate the processor to perform the operations corresponding to the modules.

[0106] An exemplary embodiment of the earphones may be headphones, including a first earphone equipped with a first microphone and a second earphone equipped with a second microphone, and its internal structure is shown in FIG. 9 . The earphones include a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the earphones provides calculation and control capabilities. The memory of the earphones includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the execution of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the earphones exchanges information between the processor and an external device. The communication interface of the earphones communicates with an external terminal via a wired or wireless method, and the wireless method can be achieved by WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the processor executes the computer program, When wind noise is present in the external environment, acquiring energy of a first audio signal collected by the first microphone and energy of a second audio signal collected by the second microphone; determining the second audio signal as a target audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold; determining the target audio signal based on the first audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is less than or equal to the predetermined threshold; and performing a wind noise canceling process on the target audio signal.

[0107] In one embodiment, the first earphone further includes a third microphone, and when the processor executes the computer program, The step of determining whether wind noise exists in the external environment based on the first audio signal and the third audio signal collected by the third microphone is further implemented.

[0108] In one embodiment, when the processor executes the computer program, specifically: calculating coherence data of the first audio signal and the third audio signal based on the first audio signal and the third audio signal; and determining whether wind noise is present in the external environment based on the coherence data.

[0109] In one embodiment, when the processor executes the computer program, The method further includes the step of performing a signal delay process on the first audio signal or the third audio signal to make the first audio signal and the third audio signal in phase.

[0110] In one embodiment, the coherence data comprises coherence values ​​for a plurality of different frequency bands, and when the processor executes a computer program, the computer program specifically: A step of determining that wind noise is present in the external environment is implemented if the coherence value of at least one frequency band is less than a threshold value corresponding to the frequency band.

[0111] In one embodiment, when the processor executes the computer program, specifically: A step of replacing a wind noise frequency band in the first audio signal to obtain the target audio signal is realized.

[0112] In one embodiment, the earphone further includes a feedback microphone, and when the processor executes the computer program, specifically: determining wind noise start and stop frequency points based on each of the coherence values; extracting a wind noise signal matching the start and stop frequency points from the fourth audio signal collected by the feedback microphone; and replacing the wind noise frequency band in the first audio signal using the extracted wind noise signal to obtain the target audio signal.

[0113] In one embodiment, the earphone is a headphone, and the feedback microphone is located in the first earphone.

[0114] In one embodiment, when the processor executes the computer program, specifically: A step of inputting the target audio signal into an AI noise canceling model to perform noise canceling processing for wind noise is realized.

[0115] As will be understood by those skilled in the art, the structure shown in FIG. 9 is merely a block diagram of a partial structure related to the solution of the present application, and does not limit the earphones to which the technical solution of the present application is applied; a specific earphone may include more or fewer components than those shown in the figure, or may combine some components, or may have a different component configuration.

[0116] A computer-readable storage medium according to one embodiment stores a computer program, which, when executed by a processor, When wind noise is present in the external environment, acquiring energy of a first audio signal collected by the first microphone and energy of a second audio signal collected by the second microphone; determining the second audio signal as a target audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold; determining the target audio signal based on the first audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is less than or equal to the predetermined threshold; and performing a wind noise canceling process on the target audio signal.

[0117] In one embodiment, the computer program, when executed by a processor, The step of determining whether wind noise exists in the external environment based on the first audio signal and the third audio signal collected by the third microphone is further implemented.

[0118] In one embodiment, the computer program, when executed by a processor, specifically: calculating coherence data of the first audio signal and the third audio signal based on the first audio signal and the third audio signal; and determining whether wind noise is present in the external environment based on the coherence data.

[0119] In one embodiment, the computer program, when executed by a processor, The method further includes the step of performing a signal delay process on the first audio signal or the third audio signal to make the first audio signal and the third audio signal in phase.

[0120] In one embodiment, the computer program, when executed by a processor, specifically: A step of determining that wind noise is present in the external environment is implemented if the coherence value of at least one frequency band is less than a threshold value corresponding to the frequency band.

[0121] In one embodiment, the computer program, when executed by a processor, specifically: A step of replacing a wind noise frequency band in the first audio signal to obtain the target audio signal is realized.

[0122] In one embodiment, the computer program, when executed by a processor, specifically: determining wind noise start and stop frequency points based on each of the coherence values; extracting a wind noise signal matching the start and stop frequency points from the fourth audio signal collected by the feedback microphone; and replacing the wind noise frequency band in the first audio signal using the extracted wind noise signal to obtain the target audio signal.

[0123] In one embodiment, the earphone is a headphone, and the feedback microphone is located in the first earphone.

[0124] In one embodiment, the computer program, when executed by a processor, specifically: A step of inputting the target audio signal into an AI noise canceling model to perform noise canceling processing for wind noise is realized.

[0125] In one embodiment, the computer program product, when executed by a processor, When wind noise is present in the external environment, acquiring energy of a first audio signal collected by the first microphone and energy of a second audio signal collected by the second microphone; determining the second audio signal as a target audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold; determining the target audio signal based on the first audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is less than or equal to the predetermined threshold; and performing a noise canceling process for the wind noise on the target audio signal.

[0126] In one embodiment, the computer program, when executed by a processor, The method further includes a step of determining whether wind noise is present in the external environment based on the first audio signal and the third audio signal collected by the third microphone, and a step of making the first audio signal and the third audio signal in phase.

[0127] In one embodiment, the computer program, when executed by a processor, specifically: calculating coherence data of the first audio signal and the third audio signal based on the first audio signal and the third audio signal; and determining whether wind noise is present in the external environment based on the coherence data.

[0128] In one embodiment, the computer program, when executed by a processor, specifically: A step of determining that wind noise is present in the external environment is implemented if the coherence value of at least one frequency band is less than a threshold value corresponding to the frequency band.

[0129] In one embodiment, the computer program, when executed by a processor, The method further includes the step of performing a signal delay process on the first audio signal or the third audio signal to make the first audio signal and the third audio signal in phase.

[0130] In one embodiment, the computer program, when executed by the processor, specifically: A step of replacing a wind noise frequency band in the first audio signal to obtain the target audio signal is realized.

[0131] In one embodiment, the computer program, when executed by a processor, specifically: determining wind noise start and stop frequency points based on each of the coherence values; extracting a wind noise signal matching the start and stop frequency points from the fourth audio signal collected by the feedback microphone; and replacing the wind noise frequency band in the first audio signal using the extracted wind noise signal to obtain the target audio signal.

[0132] In one embodiment, the earphone is a headphone, and the feedback microphone is located in the first earphone.

[0133] In one embodiment, the computer program, when executed by a processor, specifically: A step of inputting the target audio signal into an AI noise canceling model to perform noise canceling processing for wind noise is realized.

[0134] In addition, all user information (including, but not limited to, user device information, user personal information, user voice information, etc.) and data (including, but not limited to, data for analysis, stored data, displayed data, etc.) related to this application are either approved by the user or fully approved by each party, and the collection, use and processing of related data must comply with relevant regulations.

[0135] As will be understood by those skilled in the art, all or part of the flow of the method in the above embodiments can be realized by a computer program instructing associated hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it may include the flow of each of the above method embodiments. Any reference to a memory, database, or other medium used in each embodiment of the present application may include at least one of a non-volatile memory and a volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM®), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, the RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database in the embodiments may include at least one of a relational database and a non-relational database. The non-relational database may include, but is not limited to, a distributed database based on blockchain. The processor in each embodiment may be, but is not limited to, a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, programmable logic, data processing logic based on quantum computing, etc.

[0136] The technical features of the above embodiments can be combined in any manner, and for the sake of convenience, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it is considered that it should fall within the scope described in this specification.

[0137] The above examples merely illustrate some embodiments of the present application, and although the descriptions are specific and detailed, they should not be construed as limiting the scope of the claims of the present application. Those skilled in the art may make modifications and improvements without departing from the concept of the present application, and these also fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be determined based on the scope of the accompanying claims.

Claims

1. A call noise canceling method used in an earphone including a first earphone provided with a first microphone and a second earphone provided with a second microphone, When wind noise is present in the external environment, acquiring energy of a first audio signal collected by the first microphone and energy of a second audio signal collected by the second microphone; determining the second audio signal as a target audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is greater than a predetermined threshold; determining the target audio signal based on the first audio signal if a difference between the energy of the first audio signal and the energy of the second audio signal is less than or equal to the predetermined threshold; and performing wind noise canceling processing on the target audio signal.

2. The first earphone further includes a third microphone; 10. The method of claim 1, further comprising determining whether wind noise is present in the external environment based on the first audio signal and a third audio signal collected by the third microphone.

3. The step of determining whether wind noise exists in the external environment based on the first audio signal and the third audio signal collected by the third microphone includes: calculating coherence data of the first audio signal and the third audio signal based on the first audio signal and the third audio signal; and determining whether wind noise is present in the external environment based on the coherence data.

4. before calculating the coherence data for the first audio signal and the third audio signal; 4. The method of claim 3, further comprising the step of performing a signal delay process on the first audio signal or the third audio signal to make the first audio signal and the third audio signal in phase.

5. The coherence data includes coherence values ​​of a plurality of different frequency bands, and the step of determining whether wind noise is present in the external environment based on the coherence data includes:

4. The method of claim 3, comprising determining that wind noise is present in the external environment if the coherence value for at least one frequency band is less than a threshold value corresponding to the frequency band.

6. determining the target audio signal based on the first audio signal, 6. The method of claim 5, further comprising the step of: substituting a wind noise frequency band in the first sound signal to obtain the target sound signal.

7. The earphone further includes a feedback microphone, and the step of replacing a wind noise frequency band in the first audio signal to obtain the target audio signal includes: determining wind noise start and stop frequency points based on each of said coherence values; extracting a wind noise signal matching the start and stop frequency points from the fourth audio signal collected by the feedback microphone; and using the extracted wind noise signal to replace a wind noise frequency band in the first audio signal to obtain the target audio signal.

8. 8. The method of claim 7, wherein the earphone is a headphone and the feedback microphone is located in the first earphone.

9. The step of performing a wind noise canceling process on the target sound signal includes: The method of claim 1, further comprising inputting the target audio signal into an AI noise canceling model to perform wind noise canceling processing.

10. An earphone including a memory in which a computer program is stored and a processor, further including a first earphone provided with a first microphone and a second earphone provided with a second microphone, wherein when the processor executes the computer program, the earphone implements the steps of the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Wind noise reduction device

    JP2009055583A

  • Method and apparatus for detecting wind noise

    JP2015505069A

  • Acoustic device and acoustic control method

    JP2022167681A

  • Sensor Management for Wireless Devices

    JP2023554650A

  • Audio signal processing for noise reduction

    US20180270565A1