A noise suppression method and a noise suppression device for wireless earphones

Through multi-microphone configuration and signal processing technology, non-stable noise is accurately identified and data fusion is carried out, which solves the problem of insufficient noise suppression capability in the prior art, and significantly improves call quality and user experience.

CN115802225BActive Publication Date: 2025-05-30BESTECHNIC SHANGHAI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211369657.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-05-30
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify uplink target voice and effectively suppress noise under non-stable noise conditions, resulting in poor call quality and poor user experience.

Method used

Using a multi-microphone configuration, voice signals are collected through the first microphone, the second microphone and the third microphone, and frame processing and Fourier transform are performed to calculate the frequency domain coherence coefficient and signal energy ratio of the voice signal, determine whether the current frame is a noise frame, determine the wind noise and noise levels, and finally data fusion is performed based on the frequency threshold to suppress noise.

Benefits of technology

Effectively suppress strong non-stable noise, improve the clarity and intelligibility of voice signals, and improve call quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115802225B_ABST
    Figure CN115802225B_ABST
Patent Text Reader

Abstract

The present application relates to a noise suppression method and a noise suppression device for wireless earphones. The wireless earphones include a first microphone, a second microphone, and a third microphone. The noise suppression method includes: respectively collecting a first voice signal, a second voice signal, and a third voice signal by using the three microphones; determining whether it is a noise frame based on a first frequency-domain coherence coefficient between the first voice signal and the third voice signal and a signal energy ratio between the second voice signal and the third voice signal; determining a wind noise level based on a second frequency-domain coherence coefficient between the first voice signal and the second voice signal and the self-power spectral density of the first voice signal; determining a noise level; determining a frequency threshold according to the wind noise level and the noise level, and performing data fusion on the first voice signal and the third voice signal to suppress noise. This noise suppression method can effectively suppress external interfering voices and noise, improve the clarity of the low-frequency band of the voice, and improve the intelligibility of the target voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of headphone noise reduction, and more specifically, to a noise suppression method and a noise suppression device for wireless headphones. Background Art

[0002] At present, true wireless noise-canceling headphones have become an essential daily necessity in people's lives. However, when making or answering calls through headphones in noisy scenarios such as on the street, in a restaurant, or on the subway, the voice signal is submerged in background noise, resulting in poor call quality and low voice intelligibility. With the rapid development of electronic technology, users have higher and higher requirements for the voice quality output by them.

[0003] In response to the above problems, the commonly used voice activity detection systems in the industry adopt methods such as voice energy detection and voice feature extraction to extract voice features from the signals picked up by the microphone to determine whether there is an uplink voice signal of the wearer.

[0004] The existing single-microphone noise reduction method based on traditional DSP (digital signal processing) obtains the voice presence probability of the current frame through noise estimation, but this method is only applicable to stationary noise and has limited noise reduction ability. For non-stationary noise, this method cannot track the rapid change of the noise spectrum, and there is a large error in voice activity detection in a noisy environment. The single-microphone noise reduction method based on neural network determines whether it is a voice frame by extracting voice features from the signals picked up by the microphone, but it cannot distinguish whether it is the target voice signal or the interference of the speech of the people next to it. When the surrounding environmental noise is large, the accuracy of voice activity detection will decrease. Therefore, the existing technology has not been able to solve the problems of accurate recognition of the uplink target voice under non-stationary noise conditions and effective suppression of noise in the voice signal. Summary of the Invention

[0005] The present application is provided to solve the above-mentioned defects existing in the prior art. There is a need for a noise suppression method and a noise suppression device for wireless headphones, which can effectively suppress the noise and wind noise in the voice during a call under strong non-stationary noise conditions, especially enhance the clarity and intelligibility of the voice signal in the low-frequency band, improve the call quality, and improve the user experience.

[0006] According to a first aspect of the present application, a noise suppression method for wireless earphones is provided. The wireless earphones include a first microphone disposed at the lower end of the wireless earphone cavity, a second microphone disposed at the upper end of the wireless earphone cavity, and a third microphone located inside the cavity and placed in the ear when worn. The noise suppression method includes: collecting a first voice signal using the first microphone, collecting a second voice signal using the second microphone, and collecting a signal using the third microphone and performing echo cancellation processing to obtain a third voice signal. Based on the first frequency domain coherence coefficient of the current frame of the first voice signal and the current frame of the third voice signal in a first frequency range and the signal energy ratio of the current frame of the second voice signal and the current frame of the third voice signal in a second frequency range, it is determined whether the current frame is a noise frame, wherein the first frequency range and the second frequency range are determined based on the frequency range of the jaw vibration signal of the wearer of the wireless earphones and the sensitivity of the third microphone. Based on the second frequency domain coherence coefficient of the current frame of the first voice signal and the current frame of the second voice signal in the first frequency range and the self-power spectral density of the current frame of the first voice signal in the first frequency range, the wind noise level of the current frame is determined. Estimate noise energy related parameters for the current frame of the first voice signal, and determine the noise level of the current frame according to the estimated noise energy related parameters. Determine a frequency threshold according to the wind noise level and the noise level, and perform data fusion on the first voice signal and the third voice signal based on the frequency threshold to suppress the noise in the first voice signal.

[0007] According to a second aspect of the present application, there is provided a noise suppression device for wireless earphones. The wireless earphones include a first microphone disposed at the lower end of the wireless earphone cavity, a second microphone disposed at the upper end of the wireless earphone cavity, and a third microphone located inside the cavity and placed in the ear when worn. Among them, the first microphone is used to collect a first voice signal; the second microphone is used to collect a second voice signal; the third microphone is used to collect signals and perform echo cancellation processing to obtain a third voice signal. The noise suppression device includes a system-on-chip, which is configured to determine whether the current frame is a noise frame based on the first frequency domain coherence coefficient of the current frame of the first voice signal and the current frame of the third voice signal in a first frequency range in combination with the signal energy ratio of the current frame of the second voice signal and the current frame of the third voice signal in a second frequency range, where the first frequency range and the second frequency range are determined based on the frequency range of the jaw vibration signal of the wearer of the wireless earphones and the sensitivity of the third microphone. The system-on-chip is further configured to determine the wind noise level of the current frame based on the second frequency domain coherence coefficient of the current frame of the first voice signal and the current frame of the second voice signal in the first frequency range in combination with the self-power spectral density of the current frame of the first voice signal in the first frequency range. The system-on-chip is further configured to estimate noise energy related parameters for the current frame of the first voice signal, and determine the noise level of the current frame according to the estimated noise energy related parameters. The system-on-chip is further configured to determine a frequency threshold according to the wind noise level and the noise level, and perform data fusion on the first voice signal and the third voice signal based on the frequency threshold to suppress the noise in the first voice signal.

[0008] The noise suppression method and device for wireless earphones provided in various embodiments of the present application perform frame processing on the voice signals collected by multiple microphones, and determine the wind noise level and noise level of the current frame of the voice signal through the energy of each frame of the voice signal and the correlation between them, and determine the frequency threshold considering the wind noise level and noise level of the voice signal, and perform fusion processing on the first voice signal collected by the first microphone at the lower end of the wireless earphone cavity and the third voice signal collected by the in-ear microphone based on the frequency threshold. By utilizing the advantage that the processed third voice signal has a higher signal-to-noise ratio, the noise in the fused voice signal, especially in the low-frequency band where the wind noise is located, is quickly and effectively suppressed, so that the voice under strong non-stationary noise conditions can be clearer and more intelligible, the call quality is higher, and the user experience is better. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A flowchart showing a noise suppression method for wireless earphones according to an embodiment of the present application;

[0010] Figure 2 A flowchart showing the fusion process of a first voice signal and a third voice signal according to an embodiment of the present application;

[0011] Figure 3 A flowchart showing a noise suppression method for wireless earphones according to another embodiment of the present application; and

[0012] Figure 4 A partial structural schematic diagram of a noise suppression device according to an embodiment of the present application. Detailed implementation manners

[0013] To enable those skilled in the art to better understand the technical solutions of the present application, the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation manners. The embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, but this is not a limitation to the present application.

[0014] The "first", "second" and similar terms used in the present application do not indicate any order, quantity or importance, but are only used for distinction. Words such as "including" or "comprising" mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements.

[0015] Figure 1 A flowchart showing a noise suppression method for wireless earphones according to an embodiment of the present application. The wireless earphones include a first microphone 101 provided at the lower end of the wireless earphone cavity, a second microphone 102 provided at the upper end of the wireless earphone cavity, and a third microphone 103 located inside the cavity and placed in the ear when worn. Among them, the first microphone 101 is located at the lower end of the wireless earphone cavity and correspondingly collects the target voice signal emitted by the mouth of the wireless earphone wearer. Moreover, when there is noise in the surrounding environment, the surrounding noise will also be collected. For example, in the environment of a subway station, the sound of the train will also be collected. The second microphone 102 is located at the upper end of the wireless earphone cavity, facing outward or backward, and correspondingly mainly collects the voices of people around and complex environmental noise, and will also collect the target voice signal of the wireless earphone wearer, but the signal-to-noise ratio is relatively low. The third microphone 103 is located inside the ear and collects the voice signal inside the ear, which may be the vibration signal in the wearer's ear canal or the voice signal propagated in the ear canal. The voice signal collected by this in-ear microphone is mainly the target voice signal, and some surrounding environmental noise will also leak in and be collected, but the low-frequency signal-to-noise ratio is relatively high. In addition, since the third microphone 103 is very close to the position of the speaker, the collected echo signal is relatively large, and echo cancellation needs to be performed on the voice signal of the third microphone 103 collected.

[0016] The noise suppression method according to an embodiment of the present application may include: collecting a first voice signal 104 using the first microphone 101, collecting a second voice signal 105 using the second microphone 102, and collecting a signal using the third microphone 103 and performing echo cancellation processing 107 to obtain a third voice signal 106. In some embodiments, in order to enable fast processing of the voice signals obtained by each microphone and reduce the amount of computation at the same time, the voice signals of each channel are usually uniformly framed and windowed, and then subsequent Fourier transform and other processing are performed on the framed and windowed voice signal frames.

[0017] As Figure 1 shown, based on the current frame of the first voice signal 104 and the current frame of the third voice signal 106, calculate a first frequency-domain coherence coefficient 108 within a first frequency range. Specifically, for example, the current frames of the first voice signal 104 and the third voice signal 106 can be respectively Fourier-transformed to convert them from time-domain signals to frequency-domain signals, and then the first frequency-domain coherence coefficient 108 of the two frequency-domain signals within the first frequency range is calculated. In addition, a signal energy ratio 111 between the current frame of the second voice signal 105 and the current frame of the third voice signal 106 within a second frequency range can also be calculated. The specific method is similar to that described above, that is, by Fourier-transforming the current frames of each voice signal to convert them from time-domain signals to frequency-domain signals, and calculating the signal energy ratio 111 of the two frequency-domain signals within the second frequency range.

[0018] In some embodiments, the first frequency range and the second frequency range are determined based on the frequency range of the jaw vibration signal of the wearer of the wireless earphone and the sensitivity of the third microphone 103. Only as an example, for example, usually the frequency range of the jaw vibration signal during human speech is between 100 Hz and 1.5 kHz. Therefore, the frequency range for calculating the first frequency-domain coherence coefficient 108 and the signal energy ratio 111 can be set to an interval that at least covers this frequency range, and signals deviating from the above frequency range are not processed. In this way, not only can the amount of computation be greatly reduced, but also misjudgment caused by calculating signals that may be unrelated to speech in other frequency bands can be avoided. Therefore, the accuracy of determining whether the current frame is a noise frame can be improved. In some embodiments, the first frequency range and the second frequency range can be set to the same value. In other embodiments, the two can also be set to different values as needed. The present application does not limit this.

[0019] Then, combining the first frequency-domain coherence coefficient 108 and the signal energy ratio 111, it is determined whether the current frame is a noise frame in step S11. If the determination result in step S11 is "yes", that is, the current frame is a noise frame, it indicates that there may be only noise or interference signals in the current frame, and no recognizable target speech signal of the wearer is included. For example, when the wearer does not emit speech and there is noise or interference in the surrounding environment. Correspondingly, if it is determined that the current frame is not a noise frame, in one case, there is a target speech signal of the wearer in the current frame, which is called a speech frame, or in other cases, the current frame contains neither a target speech signal nor a large amount of noise or interference signals, and can be called a quiet frame. According to the result of whether it is a noise frame determined in step 112, the first speech signal can be processed correspondingly in subsequent steps.

[0020] It should be noted that since there may also be noise in the current frame that is not determined to be a noise frame, therefore, in some embodiments, regardless of whether the current frame is determined to be a noise frame, it is necessary to further determine the wind noise level and the noise level of the current frame. Regarding the wind noise level, as Figure 1 shown, the second frequency-domain coherence coefficient 110 of the two in the first frequency range can be calculated first according to the current frame of the first speech signal 104 and the current frame of the second speech signal 105, and the auto-power spectral density 109 of the current frame of the first speech signal in the first frequency range can be calculated. Then, in step S12, the second frequency-domain coherence coefficient 110 and the auto-power spectral density 109 are combined to determine the wind noise level of the current frame.

[0021] Regarding the noise level of the current frame, for example, noise energy-related parameters of the current frame of the first speech signal 104 can be estimated, and in step S13, the noise level of the current frame is determined according to the estimated noise energy-related parameters.

[0022] Similar to the foregoing processing, the calculation of the second frequency-domain coherence coefficient 110, the auto-power spectral density 109, and the noise energy-related parameters of the current frame of the first speech signal 104 also needs to be performed after converting the framed and windowed signals of each speech signal to the frequency domain, which will not be elaborated here.

[0023] In some embodiments, after obtaining the first speech signal 104, the second speech signal 105, and the third speech signal 106, each path in each speech signal can be framed separately but synchronously, and then the framed signals are windowed, and the Fourier transform is completed. And the calculations of the above first frequency-domain coherence coefficient 108, auto-power spectral density 109, second frequency-domain coherence coefficient 110, and signal energy ratio 111 can be performed simultaneously.

[0024] Next, when the wind noise level of the current frame is determined in step S12 and the noise level of the current frame is determined in step S13, a frequency threshold can be determined in step S14 according to the wind noise level and the noise level, and data fusion is performed on the first voice signal 104 and the third voice signal 106 based on the frequency threshold to suppress the noise in the first voice signal.

[0025] According to the noise suppression method of the embodiments of the present application, through steps S11 - S14, the signal energy of the framed voice signals collected by each microphone after frame processing and the correlation therebetween are used to determine the wind noise level and the noise level of the current frame of the voice signal. Considering the wind noise level and the noise level of the voice signal comprehensively, a frequency threshold is determined, and fusion processing is performed on the first voice signal collected by the first microphone at the lower end of the wireless headphone cavity and the third voice signal collected by the in-ear microphone based on this frequency threshold. Utilizing the advantage that the third voice signal has a higher signal-to-noise ratio after processing, the noise in the fused voice signal in each frequency band, especially in the low-frequency band where the wind noise is located, is quickly and effectively suppressed, so that the voice under strong non-stationary noise conditions can be clearer and more intelligible, the call quality is higher, and the user experience is better.

[0026] In some embodiments, the above first frequency domain coherence coefficient 108 can be calculated, for example, by formula (1):

[0027]

[0028] where, Φ ii (ω), Φ kk (ω) are the auto-power spectral densities of the first voice signal 104 and the third voice signal 106 respectively; Φ ik (ω) is the cross-power spectral density of the first voice signal 104 and the third voice signal 106; ω is the angular frequency; δ 1 is a small quantity greater than 0, used to avoid division by zero operation; ω 1 , ω 2 are respectively the upper and lower limits of the selected first frequency range, the minimum value of ω 1 is 0, and the maximum value of ω 2 can be 1 / 2 of the FFT frame length. The auto-power spectral density reflects the signal energy of the signal, and the first frequency domain coherence coefficient 108 defined as above can reflect the frequency correlation between the first voice signal 104 and the third voice signal 106 in the first frequency range.

[0029] In some embodiments, the signal energy ratio 111 between the current frame of the second voice signal 105 and the current frame of the third voice signal 106 in the second frequency range can be calculated by formula (2):

[0030]

[0031] Among them, S jk is the calculated signal energy ratio 111, and Φ jj (ω) is the auto-power spectral density of the second voice signal 105, and Φ kk (ω) is the auto-power spectral density of the third voice signal 106, and δ 2 is a small quantity greater than 0, used to avoid division by zero operation.

[0032] According to the characteristic that when the third voice signal 106 obtained by the third microphone 103 contains the target voice signal emitted by the wearer, it has particularly strong low-frequency energy. Therefore, when the signal energy ratio 111 is lower than the preset first threshold, it can be considered that the target voice signal of the wearer exists in the signal of the current frame. On the contrary, it can be considered that the current frame does not contain the target voice signal. And, in the case where the current frame is the quiet frame as described above, the value of the first frequency-domain coherence coefficient 108 is smaller than when there is noise, for example, a very small value close to 0. Therefore, when determining whether the current frame is a noise frame based on the first frequency-domain coherence coefficient 108 combined with the signal energy ratio 111, it can further include: when the signal energy ratio 111 is greater than or equal to the preset first threshold and the first frequency-domain coherence coefficient 108 is greater than the preset second threshold, determining that the current frame is a noise frame, that is, there is only noise or interference signal. Among them, the first threshold and the second threshold can be determined by experiments and set to appropriate values before the wireless earphone leaves the factory. Through the reasonable setting of the first threshold and the second threshold, the noise frame containing only noise signals can be more accurately identified.

[0033] The specific method for determining the wind noise level of the current frame is as follows. First, the second frequency-domain coherence coefficient 110 can be calculated by formula (3):

[0034]

[0035] Among them, C ij is the second frequency-domain coherence coefficient 110, and Φ ii (ω), Φ jj (ω) are the auto-power spectral densities of the first voice signal 104 and the second voice signal 105 respectively; Φ ij (ω) is the cross-power spectral density of the first voice signal 104 and the second voice signal 105; ω is the angular frequency; δ 1 is a small quantity greater than 0, used to avoid division by zero operation; ω 3 , ω 4 are respectively the upper and lower limits of the selected frequency range, the minimum value of ω 3 is 0, and the maximum value of ω 4 can be 1 / 2 of the FFT frame length.

[0036] Considering that wind noise has a large randomness, the squared modulus of the complex coherence function between different microphones is close to 0, and usually the signal energy is larger in the low frequency band. Therefore, the second frequency domain coherence coefficient 110 that characterizes the correlation between the first voice signal 104 and the second voice signal 105 in the selected frequency range can be calculated, and combined with the auto-power spectral density 109 of the first voice signal 104 that characterizes the signal energy in the selected frequency range, to jointly determine whether there is wind in the current frame and the corresponding wind noise level. Specifically, for example, the second frequency domain coherence coefficient 110 and the auto-power spectral density 109 of the first voice signal 104 can be input into a first finite state machine (FSM) together, to determine whether there is wind and the corresponding wind noise level in a combined manner. Only as an example, when the auto-power spectral density 109 is greater than a preset wind noise threshold and the second frequency domain coherence coefficient 110 is less than a preset coherence coefficient threshold, it can usually be determined that there is wind. In the case of determining that there is wind, the greater the auto-power spectral density 109, the greater the wind noise level. For example, it can be further divided into wind noise level two corresponding to strong wind, wind noise level one corresponding to light wind, etc. The specific level division thresholds can be preset in advance and will not be elaborated here.

[0037] In step S13, the noise level of the current frame can be determined according to the noise energy related parameters estimated for the current frame of the first voice signal 104. Among them, the noise energy related parameters can, for example, include the auto-power spectral density of the noise or the signal-to-noise ratio, etc. The noise level can be divided into levels by setting thresholds corresponding to the noise energy related parameters. For example, when the noise energy is greater than 1000, it is set as noise level one; when the noise energy is greater than 2000, it is set as noise level two; when it is greater than 3500, it is set as noise level four, etc. and will not be listed one by one here.

[0038] In some other embodiments, the second frequency domain coherence coefficient 110, the auto-power spectral density 109 of the first voice signal 104, and the noise energy related parameters estimated for the current frame of the first voice signal 104 can also be input into a second state machine, and through the uniformly set logical operations, whether there is wind, the wind noise level, and the noise level are output in an associated manner, and the present application does not limit this.

[0039] The following Figure 2 introduces the specific manner of determining the frequency threshold according to the wind noise level and the noise level, and performing data fusion on the first voice signal 104 and the third voice signal 106 based on the frequency threshold to suppress the noise in the first voice signal 104.

[0040] Figure 2The flowchart showing the fusion process of the first voice signal and the third voice signal according to an embodiment of the present application is presented. In step S21, a corresponding first frequency is determined based on the wind noise level, a corresponding second frequency is determined based on the noise level, and the larger one of the first frequency and the second frequency is selected as the frequency threshold. Each level of wind noise level has a corresponding first frequency. Similarly, each level of noise level also has a corresponding second frequency. Specifically, the corresponding relationship can be determined in advance according to experimental or test results, and the present application does not make specific limitations.

[0041] In step S22, when performing frequency-domain data fusion on the first voice signal 104 and the third voice signal 106, the first voice signal 104 below the frequency threshold is replaced with the third voice signal 106 to suppress the noise in the first voice signal 104. When outputting the first voice signal after noise suppression, an inverse Fourier transform needs to be performed to convert it from the frequency domain to a time-domain signal and output it.

[0042] In some other embodiments, in step S23, time-domain data fusion can be performed on the first voice signal 104 and the third voice signal 106. For example, parameters of high-pass filtering and low-pass filtering can be set based on the frequency threshold, and the high-pass filtered first voice signal 104 and the low-pass filtered third voice signal 106 are fused in the time domain to suppress the noise in the first voice signal. That is, the first voice signal 104 filters out the low-frequency part signal below the frequency threshold after passing through the high-pass filter, and the third voice signal 106 filters out the high-frequency part signal above the frequency threshold after passing through the low-pass filter. In this way, below the frequency threshold, the third voice signal 106 is used to replace the original first voice signal 104, and above the frequency threshold, the original first voice signal 104 remains unchanged. Then, the first voice signal after noise suppression combined with the third voice signal can be used as the output signal.

[0043] By steps S21 - S23, by comprehensively considering the wind noise level and the noise level and selecting the larger frequency threshold determined according to the wind noise level and the noise level, it is possible to select and use the third voice signal 106 with a higher signal-to-noise ratio to replace the first voice signal 104 within a larger frequency range, so that the fused voice signal has a higher signal-to-noise ratio. In particular, it can effectively suppress external interfering voices and noise, improve the clarity of the low-frequency band of the voice, improve the intelligibility of the target voice, thereby enhancing the call quality and improving the user experience of the wireless earphone.

[0044] Figure 3 The flowchart showing a noise suppression method for a wireless earphone according to another embodiment of the present application is presented. Figure 3The processing process before the fusion of the first voice signal 104 and the third voice signal 106 collected by the first microphone is shown.

[0045] First, in step S31, residual non-linear echo cancellation is respectively performed on the first voice signal 104 to obtain a first gain G out1 , and residual non-linear echo cancellation is performed on the third voice signal 106 to obtain a third gain G in1 . In this way, through the above non-linear processing, the residual echo signals in the voice signals collected by the first microphone and the second microphone can be respectively removed, so that the first voice signal 104 will output the first gain G out1 , and the first voice signal 104 will output the third gain G in1 .

[0046] In step S32, adaptive filtering is respectively performed on the first voice signal and the third voice signal. If it is determined according to the Figure 1 illustrated embodiment that the current frame is a noise frame, the coefficients of the first adaptive filtering of the first voice signal 104 in the current frame are updated, and the adaptive filtering coefficients of the third voice signal 106 in the current frame are updated. That is, it is judged whether to update the adaptive filtering coefficients according to formula (4):

[0047]

[0048] where C ik represents the first frequency-domain coherence coefficient, S jk represents the signal energy ratio, b represents the second threshold, and c represents the first threshold. When C ik is greater than the second threshold and S jk is greater than or equal to the first threshold, the current frame is a noise frame. When update = 1, it means that the coefficients of the first adaptive filter 301 and the third adaptive filter 302 need to be updated. When update = 0, it means that the coefficients of the above two filters do not need to be updated.

[0049] Specifically, if it is determined that the current frame is a noise frame, that is, only noise or interference signals exist in the current frame, and at this time the value of update in formula (4) is output as 1, then the current frames of the first voice signal 104 and the third voice signal 106 are respectively used to update the coefficients of the corresponding adaptive filter. The adaptive filter with updated coefficients can more effectively filter out noise. On the contrary, if it is determined that the current frame is not a noise frame, and at this time the value of update in formula (4) is output as 0, then the corresponding adaptive filter coefficients are not updated, but the original adaptive filter coefficients are used for filtering. Because, for example, when the current frame contains the target voice signal, that is, contains human voices, if the adaptive filter coefficients are updated at this time, it may cause the human voice signal in the current frame to be misprocessed. As a result, after the current frame passes through the filter, an undesired distortion of the human voice will occur. Therefore, it is necessary to appropriately update the adaptive filter coefficients based on accurately detecting whether the current frame is a noise frame, so as to avoid the bad experience of distorting the human voice of the target voice signal when attempting to filter out noise.

[0050] Next, in step S33, single-microphone noise suppression is performed on the first voice signal 104 after the adaptive filtering process 119 to obtain the second gain G out2 ; single-microphone noise suppression is performed on the third voice signal 106 after the adaptive filtering process to obtain the fourth gain G in2 . Through step S33, the background noise of the microphone can be reduced, and the first voice signal and the third voice signal after adaptive filtering can respectively obtain a gain. In some embodiments, the single-microphone noise suppression method can select DSP (digital signal processing) noise reduction or neural network noise reduction, or these two methods can be used for noise reduction respectively, and then a selection is made from the respectively obtained gains. For example, the gain obtained by DSP noise reduction for the first voice signal after adaptive filtering is G out4 , the gain obtained by neural network noise reduction is G out3 , and the final gain is G out2 = min(G out3 , G out4 ). By making an optimal selection, a method with a better noise suppression effect can be selected to increase the suppression amount of external high-frequency environmental noise.

[0051] In some embodiments, when performing frequency-domain data fusion on the first voice signal and the third voice signal, replacing the first voice signal below the frequency threshold with the third voice signal to suppress the noise in the first voice signal specifically includes as shown in formulas (5) and (6):

[0052]

[0053] y(t) = IFFT(ftF 1 (ω)) (6)

[0054] wherein, ftF 1 (ω) is the short - time spectrum of the fused voice signal, ftF(ω) and ftFB(ω) are the short - time spectra of the first voice signal 104 and the third voice signal 106 respectively, G mix (ω) is the gain compensation coefficient for the first voice signal and the third voice signal due to frequency response differences in a quiet situation; ω 0 is the frequency threshold, and the minimum value of ω 0 is 0, and the maximum value is 1 / 2 of the FFT length when performing FFT transformation on the first voice signal and the third voice signal. IFFT(ftF 1 (ω)) represents performing the inverse Fourier transform on ftF 1 (ω), and y(t) is the uplink voice signal in the time domain output after noise suppression obtained through the inverse Fourier transform.

[0055] In some embodiments, when performing time - domain data fusion on the first voice signal and the third voice signal, the high - pass - filtered first voice signal and the low - pass - filtered third voice signal are fused in the time domain, and setting the parameters of the high - pass filter and the low - pass filter using the frequency threshold is specifically as shown in formulas (7) and (8):

[0056] x 1 (t) = IFFT(ftF(ω)), x 3 (t) = IFFT(ftFB(ω)) (7)

[0057] y(t) = hpf(x 1 (t), ω 0 ) + lpf(x 3 (t), ω 0 ) (8)

[0058] wherein, x 1 (t) is the first voice signal in the time domain obtained by performing the inverse Fourier transform on ftF(ω), x 3 (t) is the third voice signal in the time domain obtained by performing the inverse Fourier transform on ftFB(ω), hpf(x 1 (t), ω 0 ) represents performing a high - pass filter with a cut - off frequency of ω 1 on x 0 (t), and lpf(x 3 (t), ω 0 ) represents performing a low - pass filter with a cut - off frequency of ω 3 on x 0Low-pass filtering.

[0059] The time-domain fused speech signal is the superposition result of the first speech signal in the time domain after high-pass filtering and the third speech signal in the time domain after low-pass filtering. Taking the frequency threshold as the cut-off frequency, the high-pass filter filters out the first speech signal with a frequency higher than the frequency threshold, and the low-pass filter filters out the third speech signal with a frequency lower than the frequency threshold, and then fusion is performed. The noise in the low-frequency band of the fused speech signal is smaller, which can effectively suppress the external interfering speech and noise, improve the clarity of the low-frequency band of the speech, and improve the intelligibility of the target speech.

[0060] Through formulas (5)-(8), the fused speech signal can simultaneously have a higher high-frequency environmental noise suppression amount. At the same time, it can also have a noise suppression amount in the low-frequency band adapted to the wind noise and noise level. Therefore, under the condition of non-stationary noise such as wind noise or noise, the output uplink speech signal can have good noise reduction effects in both the low-frequency band and the high-frequency band, the speech signal is clearer, the speech intelligibility during a call is higher, and the user experience is better.

[0061] In some embodiments, based on the current frame of the third speech signal, a DNN neural network can be used to assist in determining whether the current frame is a noise frame. The deep neural network is used to identify and judge the third speech signal. The signal-to-noise ratio of the third speech signal is relatively high. If noise is identified from the third speech signal, combined with the correlation between the first speech signal and the third speech signal and the signal energy ratio between the second speech signal and the third speech signal, it can be assisted to judge that the current frame is a noise frame.

[0062] In some embodiments, the third microphone includes one of a microphone, a bone conduction microphone, or a vibration sensor. The microphone can collect the sound waves propagated into the ear, the bone conduction microphone collects the sound signal propagated by bone vibration, and the vibration sensor collects the sound waves vibrating in the ear.

[0063] In some embodiments, the wireless earphone is one of an in-ear wireless earphone and a semi-in-ear wireless earphone.

[0064] According to an embodiment of the present application, there is also provided a noise suppression device for a wireless earphone. The noise suppression device for a wireless earphone according to an embodiment of the present application will be specifically described below.

[0065] Figure 4A partial structural schematic diagram of a noise suppression device according to an embodiment of the present application is shown. The wireless earphone 400 includes a first microphone 401 disposed at the lower end of the cavity of the wireless earphone 400, a second microphone 402 disposed at the upper end of the cavity of the wireless earphone 400, and a third microphone 403 located inside the cavity and placed in the ear when worn. Among them, the first microphone 401 is used to collect a first voice signal; the second microphone 402 is used to collect a second voice signal; the third microphone 403 is used to collect signals and perform echo cancellation processing to obtain a third voice signal. The first microphone 401, the second microphone 402, and the third microphone 403 are respectively located at different positions of the wireless earphone 400, and the focuses of signal collection are different. The first microphone 401 mainly collects the target voice signal emitted by the wearer's mouth and also collects surrounding noise. For example, in the environment of a subway station, it will also collect the sound of the train, etc. The second microphone 402 is affected by wind noise during air flow, and may also include the target voice signal of the wireless earphone wearer and the noise of the surrounding environment. The third microphone 403 is located inside the ear and collects the voice signal inside the ear, including the vibration signal in the wearer's ear canal and possibly the voice signal propagated in the ear canal. The voice signal collected by this in-ear microphone mainly includes the target voice signal and may also include the noise of the surrounding environment, but the signal-to-noise ratio is relatively high.

[0066] The noise suppression device 404 includes a system-on-chip 4041, and the system-on-chip 4041 is configured to determine whether the current frame is a noise frame based on the first frequency domain coherence coefficient of the current frame of the first voice signal and the current frame of the third voice signal in a first frequency range in combination with the signal energy ratio of the current frame of the second voice signal and the current frame of the third voice signal in a second frequency range, where the first frequency range and the second frequency range are determined based on the frequency range of the jaw vibration signal of the wearer of the wireless earphone and the sensitivity of the third microphone. By the degree of correlation of the first voice signal and the energy ratio of the second voice signal and the third voice signal, it is determined whether the current frame is a noise frame. If the current frame is a noise frame, it means that there may be only noise or interference signals in the current frame and no recognizable target voice signal of the wearer. For example, when the wearer does not emit a voice and there is noise or interference in the surrounding environment. Correspondingly, if it is determined that the current frame is not a noise frame, in one case, there is a directly or processed recognizable target voice signal of the wearer in the current frame, which is called a voice frame, or in other cases, the current frame contains neither a target voice signal nor a large amount of noise or interference signals, which can be called a quiet frame. According to the result of whether it is a noise frame determined in step 112, corresponding processing can be performed on the first voice signal in subsequent steps.

[0067] The system - on - chip 4041 is further configured to determine the wind - noise level of the current frame based on the second frequency - domain coherence coefficient of the current frames of the first voice signal and the second voice signal within the first frequency range and the auto - power spectral density of the current frame of the first voice signal within the first frequency range; estimate the noise - energy - related parameters for the current frame of the first voice signal, and determine the noise level of the current frame according to the estimated noise - energy - related parameters. Specifically, for example, the second frequency - domain coherence coefficient and the auto - power spectral density of the first voice signal can be input into a first Finite State Machine (FSM) together to determine whether there is wind and the corresponding wind - noise level in a combined manner. Only as an example, when the auto - power spectral density is greater than a preset wind - noise threshold and the second frequency - domain coherence coefficient 110 is less than a preset coherence - coefficient threshold, it can generally be determined that there is wind. In the case of determining that there is wind, the greater the auto - power spectral density, the greater the wind - noise level. For example, it can be further divided into wind - noise level two corresponding to strong wind and wind - noise level one corresponding to light wind, etc. The specific level - division thresholds can be set in advance and will not be elaborated here. The noise - energy - related parameters can include, for example, the auto - power spectral density of the noise or the signal - to - noise ratio, etc.

[0068] The system - on - chip 4041 is further configured to determine a frequency threshold according to the wind - noise level and the noise level, and perform data fusion on the first voice signal and the third voice signal based on the frequency threshold to suppress the noise in the first voice signal.

[0069] In some embodiments, the system - on - chip 4041 is further configured to: determine that the current frame is a noise frame when the signal - energy ratio is greater than or equal to a first threshold and the first frequency - domain coherence coefficient is greater than a second threshold. According to the characteristic that when the third voice signal acquired by the third microphone 403 contains the target voice signal emitted by the wearer, it has particularly strong low - frequency energy. Therefore, when the signal - energy ratio is lower than the preset first threshold, it can be considered that the target voice signal of the wearer exists in the signal of the current frame. Conversely, it can be considered that the current frame does not contain the target voice signal. And, in the case where the current frame is a quiet frame as described above, the value of the first frequency - domain coherence coefficient is smaller than when there is noise, for example, a very small value close to 0. Therefore, when the signal - energy ratio is greater than or equal to the preset first threshold and the first frequency - domain coherence coefficient is greater than the preset second threshold, it is determined that the current frame is a noise frame.

[0070] In some embodiments, the system - on - chip 4041 is further configured to: determine a corresponding first frequency based on the wind noise level, determine a corresponding second frequency based on the noise level, and select the larger of the first frequency and the second frequency as the frequency threshold. It is also configured to, when performing frequency - domain data fusion on the first voice signal and the third voice signal, replace the first voice signal below the frequency threshold with the third voice signal to suppress the noise in the first voice signal. It is further configured to, when performing time - domain data fusion on the first voice signal and the third voice signal, set the parameters of high - pass filtering and low - pass filtering based on the frequency threshold, and perform fusion processing in the time domain on the high - pass - filtered first voice signal and the low - pass - filtered third voice signal to suppress the noise in the first voice signal. This enables the selection of the third voice signal with a higher signal - to - noise ratio to replace the first voice signal within a larger frequency range, so that the fused voice signal has a higher signal - to - noise ratio. In particular, it can effectively suppress external interfering voices and noise, improve the clarity of the low - frequency band of the voice, enhance the intelligibility of the target voice, thereby improving the call quality and the user experience of the wireless earphones.

[0071] In some embodiments, the system - on - chip 4041 is further configured to perform residual non - linear echo cancellation on the first voice signal to obtain a first gain G out1 ; perform residual non - linear echo cancellation on the third voice signal to obtain a third gain G in1 . In this way, through the above non - linear processing, the residual echo signals in the voice signals collected by the first microphone and the second microphone can be removed respectively. Thus, the first voice signal will output the first gain G out1 , and the first voice signal will output the third gain G in1 .

[0072] The system - on - chip 4041 is also configured to perform adaptive filtering on the first voice signal and the third voice signal respectively. When determining that the current frame is a noise frame, update the first adaptive filtering coefficient of the first voice signal in the current frame and, update the third adaptive filtering coefficient of the third voice signal in the current frame. Based on accurately detecting whether the current frame is a noise frame, appropriately update the adaptive filter coefficients to avoid the bad experience of voice distortion of the target voice signal when attempting to filter out noise.

[0073] The system - on - chip 4041 is also configured to perform single - microphone noise suppression on the adaptively filtered first voice signal to obtain a second gain G out2 ; perform single - microphone noise suppression on the adaptively filtered third voice signal to obtain a fourth gain G in2By single-microphone noise suppression, the background noise of the microphone can be reduced, and a gain can be obtained for the first voice signal and the third voice signal respectively.

[0074] The system-on-chip 4041 is further configured to replace the first voice signal below the frequency threshold with the third voice signal to suppress the noise in the first voice signal when performing frequency-domain data fusion on the first voice signal and the third voice signal, which specifically includes as shown in formulas (5) and (6):

[0075]

[0076] y(t) = IFFT(ftF 1 (ω)) (6)

[0077] where ftF 1 (ω) is the short-time spectrum of the fused voice signal, ftF(ω) and ftFB(ω) are the short-time spectra of the first voice signal and the third voice signal respectively, G mix (ω) is the coefficient for gain compensation due to frequency response differences between the first voice signal and the third voice signal in a quiet situation; ω 0 is the frequency threshold, and the minimum value of ω 0 is 0, and the maximum value is 1 / 2 of the FFT length when performing FFT transformation on the first voice signal and the third voice signal. IFFT(ftF 1 (ω)) represents performing the inverse Fourier transform on ftF 1 (ω), and y(t) is the uplink voice signal in the time domain output after noise suppression obtained through the inverse Fourier transform.

[0078] The system-on-chip 4041 is further configured to perform fusion processing in the time domain on the first voice signal after high-pass filtering and the third voice signal after low-pass filtering when performing time-domain data fusion on the first voice signal and the third voice signal, and set the parameters of the high-pass filtering and the low-pass filtering using the frequency threshold, which specifically includes as shown in formulas (7) and (8):

[0079] x 1 (t) = IFFT(ftF(ω)), x 3 (t) = IFFT(ftFB(ω)) (7)

[0080] y(t) = hpf(x 1 (t), ω 0 ) + lpf(x 3 (t), ω 0 ) (8)

[0081] where x 1(t) is the first speech signal in the time domain obtained by performing the inverse Fourier transform on ftF(ω), x 3 (t) is the third speech signal in the time domain obtained by performing the inverse Fourier transform on ftFB(ω), hpf(x 1 (t), ω 0 ) represents performing high-pass filtering on x 1 (t) with a cut-off frequency of ω 0 . represents performing low-pass filtering on x 3 (t) with a cut-off frequency of ω 0 . The speech signal after time-domain fusion is the superposition result of the first speech signal in the time domain after high-pass filtering and the third speech signal in the time domain after low-pass filtering. Taking the frequency threshold as the cut-off frequency, the high-pass filter filters out the first speech signal with a frequency higher than the frequency threshold, and the low-pass filter filters out the third speech signal with a frequency lower than the frequency threshold, and then fusion is performed. The noise in the low-frequency band of the fused speech signal is smaller, which can effectively suppress external interfering speech and noise, improve the clarity of the low-frequency band of the speech, and improve the intelligibility of the target speech.

[0082] Through formulas (5)-(8), the fused speech signal can simultaneously have a higher high-frequency environmental noise suppression amount. At the same time, it can also have a noise suppression amount in the low-frequency band adapted to the wind noise and noise level. Therefore, under the condition of non-stationary noise such as wind noise or noise, the output uplink speech signal can have a better noise reduction effect in both the low-frequency band and the high-frequency band, the speech signal is clearer, the intelligibility of the speech during a call is higher, and the user experience is better.

[0083] In some embodiments, the system-on-chip 4041 is further configured to: based on the current frame of the first speech signal, the current frame of the second speech signal, and the current frame of the third speech signal, a DNN neural network can be used to assist in determining whether the current frame is a noise frame. Using a deep neural network to identify and judge the third speech signal, the signal-to-noise ratio of the third speech signal is relatively high. If noise is identified from the third speech signal, combined with the correlation between the first speech signal and the third speech signal and the signal energy ratio between the second speech signal and the third speech signal, it can assist in judging that the current frame is a noise frame.

[0084] The noise suppression device for wireless earphones according to the embodiments of the present application determines the wind noise level and noise level of the current frame of the speech signal based on the signal energy of the framed speech signals collected by each microphone after frame processing and the correlation therebetween, determines a frequency threshold considering the wind noise level and noise level of the speech signal, and performs fusion processing on the first speech signal collected by the first microphone at the lower end of the wireless earphone cavity and the third speech signal collected by the in-ear microphone based on the frequency threshold. By taking advantage of the higher signal-to-noise ratio of the processed third speech signal, the noise in each frequency band of the fused speech signal, especially in the low-frequency band where the wind noise is located, is quickly and effectively suppressed, so that the speech under strong non-stationary noise conditions can be more clearly understood, the call quality is higher, and the user experience is better.

[0085] Moreover, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application having equivalent elements, modifications, omissions, combinations (e.g., schemes that cross various embodiments), adaptations or changes. The elements in the claims will be broadly interpreted based on the language employed in the claims and are not limited to the examples described in this specification or during the implementation of the present application, and the examples will be construed as non-exclusive. Thus, the specification and examples are intended to be considered only as examples, and the true scope and spirit are indicated by the full scope of the claims and their equivalents.

[0086] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above detailed description, various features can be grouped together to simplify the present application. This should not be construed as an intention that a feature not claimed is necessary for any claim. On the contrary, the subject matter of the present application can be less than all the features of a particular embodiment of the application. Thus, the claims are incorporated herein as examples or embodiments into the detailed description, where each claim stands alone as a separate embodiment, and these embodiments can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalents empowered by these claims.

[0087] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present invention. The protection scope of the present invention is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.

Claims

1. A noise suppression method for wireless earphones, characterized in that, the wireless earphones include a first microphone disposed at the lower end of the wireless earphone cavity, a second microphone disposed at the upper end of the wireless earphone cavity, and a third microphone located inside the cavity and placed in the ear when worn. The noise suppression method includes: collecting a first voice signal using the first microphone, collecting a second voice signal using the second microphone, and collecting a signal using the third microphone and performing echo cancellation processing to obtain a third voice signal; judging whether the current frame is a noise frame based on the first frequency-domain coherence coefficient of the current frame of the first voice signal and the current frame of the third voice signal within a first frequency range in combination with the signal energy ratio of the current frame of the second voice signal and the current frame of the third voice signal within a second frequency range, wherein the first frequency range and the second frequency range are determined based on the frequency range of the jaw vibration signal of the wearer of the wireless earphones and the sensitivity of the third microphone; determining the wind noise level of the current frame based on the second frequency-domain coherence coefficient of the current frame of the first voice signal and the current frame of the second voice signal within the first frequency range in combination with the self-power spectral density of the current frame of the first voice signal within the first frequency range; estimating noise energy-related parameters for the current frame of the first voice signal, and determining the noise level of the current frame according to the estimated noise energy-related parameters; determining a frequency threshold according to the wind noise level and the noise level, and performing data fusion on the first voice signal and the third voice signal based on the frequency threshold to suppress the noise in the first voice signal.

2. The noise suppression method according to claim 1, characterized in that, judging whether the current frame is a noise frame based on the first frequency-domain coherence coefficient of the current frame of the first voice signal and the current frame of the third voice signal within a first frequency range in combination with the signal energy ratio of the current frame of the second voice signal and the current frame of the third voice signal within the first frequency range further includes: judging that the current frame is a noise frame when the signal energy ratio is greater than or equal to a first threshold and the first frequency-domain coherence coefficient is greater than a second threshold.

3. The noise suppression method according to claim 1 or 2, characterized in that, determining a frequency threshold according to the wind noise level and the noise level, and performing data fusion on the first voice signal and the third voice signal based on the frequency threshold to suppress the noise in the first voice signal further includes: determining a corresponding first frequency based on the wind noise level, determining a corresponding second frequency based on the noise level, and selecting the larger of the first frequency and the second frequency as the frequency threshold; when performing frequency-domain data fusion on the first voice signal and the third voice signal, replacing the first voice signal below the frequency threshold with the third voice signal to suppress the noise in the first voice signal; When performing time-domain data fusion on the first voice signal and the third voice signal, set the parameters of high-pass filtering and low-pass filtering based on the frequency threshold, and perform fusion processing in the time domain on the first voice signal after high-pass filtering and the third voice signal after low-pass filtering to suppress the noise in the first voice signal.

4. The noise suppression method according to claim 3, wherein, the noise suppression method further includes: Perform residual non-linear echo cancellation on the first voice signal to obtain a first gain G out1 ; perform residual non-linear echo cancellation on the third voice signal to obtain a third gain G in1 ; Performing adaptive filtering on the first voice signal and the third voice signal respectively. When determining that the current frame is a noise frame, update the first adaptive filtering coefficient of the first voice signal in the current frame, and update the third adaptive filtering coefficient of the third voice signal in the current frame; Perform single-microphone noise suppression on the first speech signal after adaptive filtering to obtain the second gain G out2 Perform single-microphone noise suppression on the third speech signal after adaptive filtering to obtain the fourth gain G in2 ; When performing frequency-domain data fusion on the first voice signal and the third voice signal, replacing the first voice signal below the frequency threshold with the third voice signal to suppress the noise in the first voice signal is specifically shown in Formulas (5) and (6): y(t) = IFFT(ftF 1 (ω)) (6) Among them, ftF 1 (ω) is the short-time spectrum of the fused speech signal, and ftF(ω) and ftFB(ω) are the short-time spectra of the first speech signal and the third speech signal respectively. G mix (ω) is the gain compensation coefficient for the first speech signal and the third speech signal due to the frequency response difference in the quiet situation; ω 0 is the frequency threshold, and ω 0 has a minimum value of 0 and a maximum value of 1 / 2 of the FFT length when performing FFT transformation on the first speech signal and the third speech signal. IFFT(ftF 1 (ω)) represents performing the inverse Fourier transform on ftF 1 (ω), and y(t) is the uplink speech signal in the time domain output after noise suppression obtained through the inverse Fourier transform; When performing time-domain data fusion on the first voice signal and the third voice signal, performing fusion processing in the time domain on the first voice signal after high-pass filtering and the third voice signal after low-pass filtering, and setting the parameters of the high-pass filtering and the low-pass filtering using the frequency threshold is specifically shown in Formulas (7) and (8): x 1 x(t) = IFFT(ftF(ω)) 3 x(t) = IFFT(ftFB(ω)) (7) y(t) = hpf(x 1 (t), ω 0 ) + lpf(x 3 (t), ω 0 ) (8) where x 1 (t) is the first speech signal in the time domain obtained by performing the inverse Fourier transform on ftF(ω), x 3 (t) is the third speech signal in the time domain obtained by performing the inverse Fourier transform on ftFB(ω), hpf(x 1 (t), ω 0 ) represents high-pass filtering of x 1 (t) with a cut-off frequency of ω 0 , and lpf(x 3 (t), ω 0 ) represents low-pass filtering of x 3 (t) with a cut-off frequency of ω 0 .

5. The noise suppression method according to claim 1 or 2, wherein, the noise suppression method further includes: Based on the current frame of the third voice signal, using a DNN neural network to assist in determining whether the current frame is a noise frame.

6. The noise suppression method according to claim 1 or 2, wherein, the third microphone includes one of a microphone, a bone conduction microphone, or a vibration sensor.

7. The noise suppression method according to claim 1 or 2, wherein, the wireless earphone is one of an in-ear wireless earphone or a semi-in-ear wireless earphone.

8. A noise suppression device for a wireless earphone, wherein, the wireless earphone includes a first microphone disposed at the lower end of the wireless earphone cavity, a second microphone disposed at the upper end of the wireless earphone cavity, and a third microphone located inside the cavity and placed in the ear when worn. Among them, the first microphone is used to collect a first voice signal; the second microphone is used to collect a second voice signal; the third microphone is used to collect signals and perform echo cancellation processing to obtain a third voice signal; the noise suppression device includes a system-on-chip, and the system-on-chip is configured to: Based on the first frequency-domain coherence coefficient of the current frame of the first voice signal and the current frame of the third voice signal in a first frequency range, combined with the signal energy ratio of the current frame of the second voice signal and the current frame of the third voice signal in a second frequency range, determine whether the current frame is a noise frame, where the first frequency range and the second frequency range are determined based on the frequency range of the jaw vibration signal of the wearer of the wireless earphone and the sensitivity of the third microphone; Determine the wind noise level of the current frame based on the second frequency domain coherence coefficient of the current frame of the first voice signal and the current frame of the second voice signal within the first frequency range, in combination with the auto-power spectral density of the current frame of the first voice signal within the first frequency range; Estimate the noise energy related parameters for the current frame of the first voice signal, and determine the noise level of the current frame according to the estimated noise energy related parameters; Determine a frequency threshold based on the wind noise level and the noise level, and perform data fusion on the first voice signal and the third voice signal based on the frequency threshold to suppress the noise in the first voice signal.

9. The noise suppression device according to claim 8, wherein, the system-on-chip is further configured to: Determine that the current frame is a noise frame when the signal energy ratio is greater than or equal to a second threshold and the first frequency domain coherence coefficient is greater than a third threshold.

10. The noise suppression device according to claim 8 or 9, wherein, the system-on-chip is further configured to: Determine a corresponding first frequency based on the wind noise level, determine a corresponding second frequency based on the noise level, and select the larger of the first frequency and the second frequency as the frequency threshold; When performing frequency domain data fusion on the first voice signal and the third voice signal, replace the first voice signal below the frequency threshold with the third voice signal to suppress the noise in the first voice signal; When performing time domain data fusion on the first voice signal and the third voice signal, set the parameters of the high-pass filter and the low-pass filter based on the frequency threshold, and perform fusion processing in the time domain on the high-pass filtered first voice signal and the low-pass filtered third voice signal to suppress the noise in the first voice signal.

11. The noise suppression device according to claim 10, wherein, the system-on-chip is further configured to: Perform residual non-linear echo cancellation on the first speech signal to obtain a first gain G out1 ; perform residual non-linear echo cancellation on the third speech signal to obtain a third gain G in1 ; Perform adaptive filtering on the first voice signal and the third voice signal respectively. When it is determined that the current frame is a noise frame, update the first adaptive filtering coefficient of the first voice signal of the current frame and, update the third adaptive filtering coefficient of the third voice signal of the current frame; Perform single-microphone noise suppression on the first speech signal after adaptive filtering to obtain the second gain G out2 ; Perform single-microphone noise suppression on the third speech signal after adaptive filtering to obtain the fourth gain G in2 ; When performing frequency domain data fusion on the first voice signal and the third voice signal, replacing the first voice signal below the frequency threshold with the third voice signal to suppress the noise in the first voice signal specifically includes as shown in formulas (5) and (6): y(t) = IFFT(ftF 1 (ω)) (6) Among them, ftF 1 (ω) is the short-time spectrum of the fused voice signal, ftF(ω), ftFB(ω) are the short-time spectra of the first voice signal and the third voice signal respectively, and G mix (ω) is the coefficient for gain compensation due to frequency response differences between the first voice signal and the third voice signal in a quiet situation; ω 0 is the frequency threshold, and ω 0 has a minimum value of 0 and a maximum value of 1 / 2 of the FFT length when performing FFT transformation on the first voice signal and the third voice signal. IFFT(ftF(ω)) represents performing the inverse Fourier transform on ftF(ω), and y(t) is the uplink voice signal in the time domain output after noise suppression obtained through the inverse Fourier transform; When performing time domain data fusion on the first voice signal and the third voice signal, performing fusion processing in the time domain on the high-pass filtered first voice signal and the low-pass filtered third voice signal, and setting the parameters of the high-pass filter and the low-pass filter using the frequency threshold specifically includes as shown in formulas (7) and (8): x 1 (t) = IFFT(ftF(ω)), x 3 (t) = IFFT(ftFB(ω)) (7) y(t) = hpf(x 1 (t), ω 0 ) + lpf(x 3 (t), ω 0 ) (8) where, x 1 (t) is the first speech signal in the time domain obtained by performing the inverse Fourier transform on ftF(ω), x 3 (t) is the third speech signal in the time domain obtained by performing the inverse Fourier transform on ftFB(ω), hpf(x 1 (t), ω 0 ) represents performing a high-pass filter with a cut-off frequency of ω 1 on x 0 (t), and lpf(x 3 (t), ω 0 ) represents performing a low-pass filter with a cut-off frequency of ω 3 on x 0 .

12. The noise suppression device according to claim 8 or 9, wherein, the system-on-chip is further configured to: Based on the current frame of the third voice signal, use a DNN neural network to assist in determining whether the current frame is a noise frame.

13. The noise suppression device according to claim 8 or 9, characterized in that, the third microphone includes one of a microphone, a bone conduction microphone or a vibration sensor.

14. The noise suppression device according to claim 8 or 9, characterized in that, the wireless earphone is one of an in-ear wireless earphone and a semi-in-ear wireless earphone.

Citation Information

Patent Citations

  • Dual-mic denoising system and denoising method of earphone

    CN107371079A

  • Voice activity detection method, noise inhibition method and noise inhibition system

    CN109920451A