Reducing noise in headsets using speech accelerometer signals

By combining the adaptive filtering technology of voice accelerometer signal and microphone signal in the headset, especially attenuating noise in the low frequency range and combining signals, the problem of low signal-to-noise ratio of microphone signal during noise suppression in the headset is solved, and effective noise reduction and voice retention in different noise environments are achieved.

CN120472919APending Publication Date: 2025-08-12HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510308331.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2019-09-05
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, when noise suppression in headsets, the signal-to-noise ratio of the microphone signal is too low, resulting in the inability to effectively restore the user's voice, especially in noisy environments.

Method used

By adjusting the microphone signal using a voice accelerometer signal, especially attenuating noise in the low frequency range, and combining it with the adjusted microphone signal, the signal-to-noise ratio is improved, and the noise is further reduced in combination with the traditional noise reduction method.

Benefits of technology

Under low signal-to-noise ratio conditions, the noise in the headset is effectively reduced, while retaining user voice information, adapting to different noise environments to flexibility and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472919A_ABST
    Figure CN120472919A_ABST
Patent Text Reader

Abstract

The invention relates to the field of voice and audio signal processing, namely noise reduction in headsets. In particular, the present invention proposes an apparatus and method for reducing noise in a headset. The apparatus is configured to: obtain a microphone signal; obtaining a voice accelerometer signal; adjusting the voice accelerometer signal according to the microphone signal; in a first mode of operation, the device is further configured to: modify the microphone signal by attenuating a low frequency; combining the modified microphone signal with the adjusted voice accelerometer signal; and outputting the combined signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 201980099938.4, and the original application date is September 5, 2019. The entire content of the original application is incorporated into this application by reference. Technical Field

[0002] The present invention relates to the field of speech and audio signal processing, and more particularly to noise reduction in headsets. The present invention provides a device and method for reducing noise in headsets. The device and method address the problem of noise contamination in the speech captured by a microphone when using a headset, such as when making a phone call in a noisy environment. Background Art

[0003] The traditional approach to addressing this noise problem is to use the same noise reduction methods used in the telephone industry. For example, spectral subtraction (which dates back to the late 1970s, see, for example, S.F. Boll, "Suppression of acoustic noise in speech using spectral subtraction," IEEE Transactions on Acoustics and Signal Processing, Vol. ASSP-27, April 1979)). Voice activity detection (VAD) and beamformers have also been studied over the years.

[0004] A newer approach is to use a voice accelerometer to generate a voice accelerometer signal. The advantage of this voice accelerometer signal is that its signal-to-noise ratio (SNR) is inherently better suited to the user's voice, but only within a narrow frequency range (typically below 1kHz).

[0005] US20140093093 A1 (relating to Apple's "AirPods") describes how VAD is calculated from a speech accelerometer and used to steer a beamformer and update noise estimates in noise suppression. Figure 8 compares the spectrum of the output signal from the "AirPods" (a) with the spectrum of the output signal produced by a device based on the above-mentioned conventional approach (b).

[0006] US20180367882 A1 (see Figure 9 ) describes the use of speech accelerometer signals for estimating speech in order to achieve more accurate speech estimation in noise suppression (and other speech enhancement tasks).

[0007] US20170337933 A1 (see Figure 10) describes noise suppression in two stages: first, in a "first equalizer," and second, using gain after the filter bank (FB). The gain is adjusted based on the voice signature (VS), which is derived from the voice accelerometer signal. The filter bank is used to provide frequency separation.

[0008] With the above method, if the SNR of the microphone signal is too low, that is, in the presence of noise, the user voice in the microphone signal cannot be recovered. Summary of the Invention

[0009] In view of the above-mentioned problems and shortcomings, embodiments of the present invention are intended to improve upon conventional methods. The present invention is directed to providing a device and method for reducing noise in a headset, respectively, which can effectively reduce noise in a microphone signal without losing information about the speech component in the microphone signal. Specifically, it should be possible to recover the user's speech (i.e., the sound / speech picked up by the microphone) in the microphone signal in the presence of noise (i.e., even if the microphone signal-to-noise ratio (SNR) is low).

[0010] This object is achieved by the embodiments of the invention described in the appended independent claims. Advantageous implementations of the embodiments of the invention are further defined in the dependent claims.

[0011] Specifically, embodiments of the present invention are based on the recognition that if the SNR of a microphone signal is too low, it is impossible to recover the user's voice component in the microphone signal using only the microphone signal itself. Accordingly, embodiments of the present invention propose additionally using a voice accelerometer signal to improve the SNR in the microphone signal. Noise reduction can then be continued using conventional methods that provide sufficient noise reduction at moderate noise levels. Specifically, by utilizing the voice accelerometer signal, the SNR of the microphone signal can be improved in low frequencies.

[0012] A first aspect of the present invention provides a device for reducing noise in a headset, the device being configured to: obtain a microphone signal; obtain a speech accelerometer signal; adjust the speech accelerometer signal according to the microphone signal; in a first operating mode, the device being further configured to: modify the microphone signal by attenuating low frequencies; merge the modified microphone signal with the adjusted speech accelerometer signal; and output the merged signal.

[0013] Specifically, adjusting the voice accelerometer signal according to the microphone signal may mean modifying the voice accelerometer signal so that its amplitude and phase are the same or at least similar to the amplitude and phase of the microphone signal. This can be achieved by adaptive filtering using the microphone signal and the voice accelerometer signal as input.

[0014] The "low frequencies" of the microphone signal attenuated in the first mode may include frequencies below 1 kHz, frequencies below 2 kHz, and frequencies below 3 kHz. The attenuated frequency range may be selected based on the noise level in the environment of the device.

[0015] In the first mode, the device according to the first aspect can use the voice accelerometer signal to improve the SNR of the microphone signal, particularly at the low frequencies. Thus, the noise in the microphone signal can be reduced without losing any information, particularly information about the voice component in the microphone signal. Conventional noise reduction methods can then be applied to the combined signal by the first device or another device, for example.

[0016] In an implementation of the first aspect, the device includes: a filter bank including a plurality of filters; the filter bank is configured to combine the modified microphone signal with the adjusted voice accelerometer signal in the first mode.

[0017] For example, depending on the SNR of the microphone signal, different filters within a filter bank or different filter banks may be used, particularly at the low frequencies. This can be easily implemented, making the device according to the first aspect very flexible to different environmental (noise) conditions.

[0018] In one implementation of the first aspect, the filter group includes a first filter and a second filter, the first filter is a high-pass filter, and the second filter is a low-pass filter; in the first mode, the device is further used to: use the first filter to filter the microphone signal to obtain the modified microphone signal; and use the second filter to filter the adjusted speech accelerometer signal before merging the adjusted speech accelerometer signal with the modified microphone signal.

[0019] In other words, the high-pass (first) filter is used to attenuate the low frequencies of the microphone signal. The low-pass (second) filter is used to attenuate the high frequencies of the voice accelerometer signal. A typical voice accelerometer signal only has content in the lower frequencies (e.g., below 3 kHz, or even just below 1 kHz). Accordingly, the low-pass filter can be used to pass frequencies below 3 kHz, or even just below 1 kHz, of the voice accelerometer signal.

[0020] In an implementation of the first aspect, the voice accelerometer signal is adjusted according to the microphone signal at a sampling rate lower than the sampling rate of the microphone signal; and the second filter is used to interpolate the adjusted voice accelerometer signal to the sampling rate of the microphone signal.

[0021] In an implementation of the first aspect, in the second operating mode, the device is further configured to: delay the microphone signal; and output the delayed microphone signal.

[0022] Accordingly, the device is configured to generate a delay in the microphone signal. The delay may be caused by filtering the microphone signal. Specifically, when the delay is generated, the microphone signal does not change, i.e., the content of the microphone signal remains unchanged but is shifted in time. Delaying the microphone signal enables smooth switching between the first operating mode and the second operating mode, and vice versa.

[0023] If noise reduction is not required, the second operating mode is preferably used. In this case, the voice accelerometer signal may not be combined with the microphone signal, and the low frequencies of the microphone signal may not be attenuated. Different operating modes make the device very flexible and efficient in different noise conditions.

[0024] In an implementation of the first aspect, the device is further configured to: apply noise suppression to the combined signal in the first mode and / or apply noise suppression to the delayed microphone signal in the second mode.

[0025] That is, further noise suppression or reduction can be performed on the signal output by the device, in particular based on conventional methods.

[0026] In an implementation of the first aspect, the filter bank is configured to delay the microphone signal in the second mode.

[0027] In an implementation of the first aspect, in the second mode, the device is further configured to: filter the microphone signal using the first filter and the second filter to obtain the delayed microphone signal.

[0028] The first filter and the second filter may be configured such that when the microphone signal is filtered using both filters, the content of the microphone signal is not changed. The microphone signal is simply delayed by the filtering process.

[0029] In an implementation of the first aspect, the apparatus is configured to adjust the filters in the filter bank when changing between the first mode and the second mode.

[0030] That is, in the first mode and the second mode, the filter bank can use different filters respectively. For example, in the first mode, a high-pass filter is needed to suppress the low frequencies, and the suppression depends on the noise level. In the second mode, a filter is needed that delays but does not change the microphone signal. It is also possible to adjust the filters in the filter bank for different noise conditions in the first mode, that is, to use different filters. For example, under normal noise conditions, a moderate attenuation of the low frequencies of the microphone signal may be sufficient. Under strong / wind noise conditions, a stronger attenuation of the low frequencies may be required. Under strong wind noise conditions, the higher frequencies of the microphone signal (for example, above 1kHz, or even above 3kHz) may also be additionally attenuated.

[0031] Therefore, different first filters can be used according to the noise.

[0032] In an implementation of the first aspect, each of the filters in the filter bank includes less than 10 taps, specifically each includes 7 taps.

[0033] Accordingly, the filters in the filter bank are relatively short and can therefore be realized particularly efficiently.

[0034] In an implementation of the first aspect, the device is configured to: determine whether an environment is quiet or noisy; if the environment is noisy, select the first mode; if the environment is quiet, select the second mode.

[0035] If the environment is quiet, the device may further select a third mode in which the microphone signal is output without any delay.

[0036] A "quiet" or "noisy" environment (the environment of the device, i.e., the surroundings of the device) can be distinguished based on the noise level or some more complex method. That is, if the noise level is below a threshold, the environment can be considered "quiet", and if the noise level is above the threshold, the environment can be considered "noisy". Of course, different thresholds can be used to more accurately distinguish between "noisy", "very noisy", etc., or to distinguish between "ordinary noise" (e.g., cafeteria, car, street, station, campus, etc.) and "wind noise". The noise level can be considered (only) within a certain frequency range. Exemplary thresholds for distinguishing between "quiet" and "noisy" environments can be a noise sound pressure level of 30dB, 50dB, or 60dB.

[0037] In an implementation of the first aspect, the device is further configured to, when determining that the environment is noisy: determine whether the environment is windy; and if the environment is windy, enhance the attenuation of the low frequency of the microphone signal in the first mode.

[0038] A "windy" environment can be determined based on the presence of wind noise in the microphone signal. Wind noise can be detected and, in particular, distinguished from normal noise. For example, wind noise can be highly non-stationary and may vary, for example, including gusts. Normal noise can be more stationary and uniform.

[0039] In an implementation of the first aspect, the device is configured to gradually change between the first mode and the second mode.

[0040] This helps avoid noise pumping effects and other discontinuities. For example, the filters in the filter bank can be adjusted so that they are between the filters used in the first mode and the filters used in the second mode. For example, when gradually changing from the second mode to the first mode, the attenuation of the low frequencies can be gradually increased.

[0041] In an implementation of the first aspect, the device includes an adaptive filter for adjusting the voice accelerometer signal according to the microphone signal, specifically, adjusting the amplitude and phase of the voice accelerometer signal to the amplitude and phase of the microphone signal.

[0042] Specifically, with respect to the speech component, the amplitude and phase are adjusted to be the same or at least similar, that is, the speech component in the adjusted speech accelerometer signal can match the speech component in the microphone signal.

[0043] In an implementation of the first aspect, the device (specifically, a headset) includes: a microphone, configured to generate the microphone signal; and a voice accelerometer, configured to generate the voice accelerometer signal.

[0044] The device may further comprise the adaptive filter and / or the filter bank and / or the processing circuit.

[0045] A second aspect of the present invention provides a method for reducing noise in a headset, the method comprising: obtaining a microphone signal; obtaining a speech accelerometer signal; adjusting the speech accelerometer signal according to the microphone signal; in a first operating mode, the method further comprises: modifying the microphone signal by attenuating low frequencies; merging the modified microphone signal with the adjusted speech accelerometer signal; and outputting the merged signal.

[0046] In an implementation of the second aspect, the method includes combining the modified microphone signal with the adjusted speech accelerometer signal in the first mode using a filter bank, wherein the filter bank includes a plurality of filters.

[0047] In one implementation of the second aspect, the filter group includes a first filter and a second filter, the first filter is a high-pass filter, and the second filter is a low-pass filter; in the first mode, the method further includes: filtering the microphone signal using the first filter to obtain the modified microphone signal; and filtering the adjusted speech accelerometer signal using the second filter before merging the adjusted speech accelerometer signal with the modified microphone signal.

[0048] In an implementation of the second aspect, the voice accelerometer signal is adjusted according to the microphone signal at a sampling rate lower than the sampling rate of the microphone signal; and the second filter interpolates the adjusted voice accelerometer signal to the sampling rate of the microphone signal.

[0049] In an implementation manner of the second aspect, in the second operating mode, the method further includes: delaying the microphone signal; and outputting the delayed microphone signal.

[0050] In an implementation of the second aspect, the method further includes: applying noise suppression to the combined signal in the first mode and / or applying noise suppression to the delayed microphone signal in the second mode.

[0051] In an implementation of the second aspect, the filter bank delays the microphone signal in the second mode.

[0052] In an implementation manner of the second aspect, in the second mode, the method further includes: filtering the microphone signal using the first filter and the second filter to obtain the delayed microphone signal.

[0053] In an implementation of the second aspect, the method includes adjusting the filters in the filter bank when changing between the first mode and the second mode.

[0054] In an implementation of the second aspect, each of the filters in the filter bank includes less than 10 taps, specifically each includes 7 taps.

[0055] In an implementation manner of the second aspect, the method includes: determining whether an environment is quiet or noisy; if the environment is noisy, selecting the first mode; if the environment is quiet, selecting the second mode.

[0056] In an implementation of the second aspect, the method further includes, when determining that the environment is noisy: determining whether the environment is windy; and if the environment is windy, enhancing the attenuation of the low frequency of the microphone signal in the first mode.

[0057] In an implementation of the second aspect, the method includes gradually changing between the first mode and the second mode.

[0058] In an implementation of the second aspect, the method includes: using an adaptive filter to adjust the voice accelerometer signal according to the microphone signal, specifically, adjusting the amplitude and phase of the voice accelerometer signal to the amplitude and phase of the microphone signal.

[0059] In an implementation of the second aspect, the method (specifically performed in a headset) includes: generating the microphone signal using a microphone; and generating the voice accelerometer signal using a voice accelerometer.

[0060] The method according to the second aspect and its implementations achieves all the above advantages and effects of the device according to the first aspect and its implementations. The above definitions and explanations of the device according to the first aspect also apply to the method according to the second aspect.

[0061] It should be noted that all devices, elements, units, and modules described in this application may be implemented in software or hardware elements or any combination thereof. All steps performed by various entities described in this application and functions to be performed by various entities are intended to indicate that the corresponding entities are suitable for or used to perform the corresponding steps and functions.

[0062] Although in the description of the following specific embodiments, specific functions or steps performed by an external entity are not reflected in the description of the specific detailed elements of the entity that performs the specific steps or functions, it should be clear to those skilled in the art that these methods and functions can be implemented in corresponding hardware or software elements or any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The following description of specific embodiments, in conjunction with the accompanying drawings, illustrates various aspects and implementations of the present invention, wherein:

[0064] Figure 1 The device provided by the embodiment of the present invention is shown;

[0065] Figure 2 The device provided by an embodiment of the present invention is shown, particularly in a first operating mode;

[0066] Figure 3 The device provided by the embodiment of the present invention is shown, especially in the second operating mode;

[0067] Figure 4(a) to Figure 4(c) An example of a filter bank of a device provided by an embodiment of the present invention is shown;

[0068] Figure 5 The following schematically illustrates a mode selection process performed by a device according to an embodiment of the present invention;

[0069] FIG6( a ) and FIG6 ( b ) compare the spectrum of the output signal of the device provided by the embodiment of the present invention with the spectrum of the output signal obtained by a traditional method;

[0070] Figure 7 The method provided by the embodiment of the present invention is shown;

[0071] Figures 8(a) and 8(b) compare the spectrum of the output signal of Apple's "AirPods" with the spectrum of the output signal obtained by another traditional method;

[0072] Figure 9 Another conventional noise reduction method is schematically shown;

[0073] Figure 10 Another traditional noise reduction method is schematically shown. DETAILED DESCRIPTION

[0074] Figure 1 FIG1 illustrates a device 100 provided by an embodiment of the present invention. Device 100 is adapted to reduce noise in a headset, or at least to support noise reduction in a headset. Device 100 may be included in the headset. Device 100 may also be a headset, for example, a headset for a mobile phone or other mobile device. However, while device 100 is adapted for use with a headset, it may also be adapted for use with any other microphone-based device or application.

[0075] The device 100 may include a processing circuit (not shown) for performing, carrying out, or initiating the various operations of the device 100 described herein. The processing circuit may include hardware and software. The hardware may include analog circuits or digital circuits, or both. The digital circuit may include components such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a multi-purpose processor. In one embodiment, the processing circuit includes one or more processors and a non-transitory memory connected to the one or more processors. The non-transitory memory may carry executable program code that, when executed by one or more processors, causes the device 100 to perform, carry out, or initiate the operations or methods described herein.

[0076] Device 100 is configured to obtain a microphone signal 101, specifically provided by one or more microphones or a microphone array. That is, microphone signal 101 may be provided by a single microphone or may be a signal generated by combining multiple microphones. Device 100 may include the one or more microphones to generate microphone signal 101 itself. However, device 100 may also receive microphone signal 101 from the one or more (external) microphones or from another device.

[0077] The device 100 is also configured to obtain a voice accelerometer signal 102, specifically provided by one or more voice accelerometers or an array of voice accelerometers. That is, the voice accelerometer signal 102 can be provided by a single voice accelerometer or by combining multiple voice accelerometers. The device 100 can include the one or more voice accelerometers to generate the voice accelerometer signal 102 itself. However, the device 100 can also receive the voice accelerometer signal 102 from the one or more voice accelerometers or from another device. It is worth noting that a voice accelerometer can be used to measure vibrations (e.g., in the x, y, and z directions), such as vibrations induced at / in a user's head. Specifically, the vibrations can be induced by the user's voice, i.e., when the user speaks into the one or more microphones. Therefore, the voice accelerometer signal 102 can reflect the user's voice, particularly at low frequencies. The voice accelerometer can also be implemented using a microphone, i.e., the voice accelerometer signal 102 can be a microphone signal, wherein the microphone can be placed in the user's ear.

[0078] The device 100 is further configured to adjust the voice accelerometer signal 102 based on the microphone signal 101 to obtain an adjusted voice accelerometer signal 103, for example by using the microphone signal 101 as an input to the adjustment process. For example, the adjustment can be performed by adaptively filtering the voice accelerometer signal 102 and the microphone signal 101. Specifically, the adjustment can align the amplitude and / or phase of the microphone signal 101 with the amplitude and / or phase of the voice accelerometer signal 102 with respect to the voice signal component (but not the noise component caused by noise leakage into the accelerometer signal 102).

[0079] The device 100 can operate in at least a first operating mode (e.g. Figure 1 ), specifically it can also be operated in the second operating mode ( Figure 1 ), more specifically operating in an intermediate mode. Accordingly, the device 100 may be configured to gradually change between the first mode and the second mode.

[0080] In the first operating mode, device 100 is further configured to modify microphone signal 101 by attenuating low frequencies (of microphone signal 101) to obtain modified microphone signal 104. Thus, noise in microphone signal 101, primarily reflected at low frequencies, can be reduced. Furthermore, device 100 is further configured to combine modified microphone signal 104 with adjusted voice accelerometer signal 103 and then output combined signal 105. Thus, (adjusted) voice accelerometer signal 102 can be used to improve the signal-to-noise ratio (SNR) of (modified) microphone signal 101. This first operating mode is best suited for noisy environments.

[0081] In the second operating mode, the device 100 is further configured to delay the microphone signal 101 to obtain a delayed microphone signal 204, and then output the delayed microphone signal 204 ( Figure 1 Not shown in, but see e.g. Figure 3 ). In particular, apart from the delay, the microphone signal 101 may not be modified in the second mode. For example, the microphone signal 101 is not combined with the (adjusted) voice accelerometer signal 102. The second operating mode is best suited for quiet environments.

[0082] In the first mode, the device 100 generates the combined signal 105, and in the second mode, the device 100 generates the delayed microphone signal 204. In the intermediate mode, the generated signal can be defined as:

[0083] Resulting signal = f*combined signal + (1-f)*delayed microphone signal

[0084] Here, if the first mode and the second mode are gradually changed, f gradually changes from zero (0) to one (1), and vice versa.

[0085] Figure 2 FIG. 1 shows a device 100 provided by an embodiment of the present invention. Figure 2 The embodiment of the device 100 shown is based on Figure 1 The embodiment of the device 100 shown in FIG. Figure 2 is shown as having further optional features. Accordingly, Figure 2 The device 100 includes Figure 1 Therefore, like features are labeled with like reference numerals and have like functions.

[0086] Figure 2 In particular, a first operating mode of the device 100 is shown. Figure 3 The embodiment of the present invention provides Figure 2 The same device 100 is shown, but specifically showing the second mode of operation of the device 100.

[0087] like Figure 2 and Figure 3As shown separately, the device 100 may include an adaptive filter 203, which may be configured to adjust the voice accelerometer signal 102 based on the microphone signal 101. Specifically, the adaptive filter 203 may be configured to adjust the amplitude and / or phase of the voice accelerometer signal 102 to the amplitude and / or phase of the microphone signal 101, particularly with respect to the sound / speech component in the signal. To this end, the adaptive filter 203 may receive the voice accelerometer signal 102 as an input, and may also receive the adjusted voice accelerometer signal 103 and the microphone signal 101 as another input (specifically, the input may be the sum of the signal 103 and the signal 101).

[0088] like Figure 2 and Figure 3 As further shown, the device 100 may include a filter bank 200, the filter bank 200 including a plurality of filters (201, 202), specifically including a first filter 201 and a second filter 202. In the first mode ( Figure 2 ), the filter bank 200 may be used to perform said merging of the modified microphone signal 104 with the adjusted speech accelerometer signal 103. In said second mode ( Figure 3 ), the filter bank 200 can be used to cause the delay of the microphone signal 101. The filters (201, 202) in the filter bank 200 can be adjustable, i.e. can be changed. For example, their filtering behavior with respect to the passed amplitude and / or frequency can be changed.

[0089] The first filter 201 may be a high-pass filter, and the second filter 202 may be a low-pass filter. Figure 2 ), the microphone signal 101 may be filtered using a first filter 201 to obtain a modified microphone signal 104, and the adjusted speech accelerometer signal 103 may be filtered using a second filter 202 before being combined with the modified microphone signal 104. In the second mode ( Figure 3 ), the microphone signal 101 can be filtered using both the first filter 201 and the second filter 202 to obtain a delayed microphone signal 204. It is worth noting that passing any type of signal through an (analysis-synthesis) filter bank produces a delayed version of the signal.

[0090] exist Figure 2In the illustrated device 100, i.e., in the first operating mode of the device 100, the SNR of the microphone signal 101 can be improved, particularly in the low frequency band, by first modifying the voice accelerometer signal 102 so that it has an amplitude and / or phase similar to that of the microphone signal 101 of the user's voice. This can be achieved by means of an adaptive filter 203 (wherein adaptation can only be performed when the user is speaking, i.e., when the pitch is detected). Furthermore, by secondly using the filter bank 200, the signals 101 and 102 are combined, and noise reduction is then continued as in conventional methods (based on the combined signal 105).

[0091] The filter bank 200 used in the first mode or in the second mode may depend on the SNR of the microphone signal 101, especially at low frequencies. In a quiet environment, the SNR of the microphone signal 101 is generally better than the SNR of the speech accelerometer signal 102, since movement can induce low-level noise in the speech accelerometer signal 102. Then, preferably, the speech accelerometer signal 102 is not used, but the microphone signal 101 itself and the filter bank 200 are (only) used to generate the delay (i.e. to generate the delayed microphone signal 204). This is as Figure 3 The second mode is shown, so this mode is best suited for quiet environments. In noisy environments (e.g., cafeterias, cars, streets, stations, campuses, etc.) or in windy environments, it may be sufficient to reduce low-frequency noise by 10-20 dB (by attenuating the microphone signal 101 at low frequencies). However, wind noise may also be very strong at low frequencies, in which case additional attenuation may be required. This is as Figure 2 The first mode is shown, and is therefore best suited for noisy environments, even windy environments. The low-frequency attenuation of the microphone signal 101 can be achieved by a high-pass filter as the first filter 201 .

[0092] Typically, the voice accelerometer signal 102 has content only in its lower frequencies. Therefore, the adaptive filtering performed using the adaptive filter 203 can be performed at a lower sampling rate, and the filter bank 200 can include interpolating the adaptively filtered voice accelerometer signal 103 to 16 kHz, etc. (i.e., after inserting enough zeros between samples, the attenuation of the second filter 202 in the filter bank 200 must be deeper). Therefore, the voice accelerometer signal 102 can be adjusted according to the microphone signal 101 at a sampling rate lower than the sampling rate of the microphone signal 101. The second filter 202 can then be used to interpolate the adjusted voice accelerometer signal 103 to the sampling rate of the microphone signal 101.

[0093] For the filters (201, 202) in the filter bank 200, very short filters can be selected, in which case they can be implemented particularly efficiently. For example, in the exemplary filter bank 200, each of the filters (201, 202) can have fewer than 10 taps, for example, each can have only 7 taps. Due to the interpolation of the adjusted speech accelerometer signal 104, the length of the combined second filter 202 can be longer, for example 32.

[0094] Figure 4(a) to Figure 4(c) The device 100 provided by the embodiment of the present invention (specifically, Figure 2 or Figure 3 Examples of filters (201, 202) in a filter bank 200 of the device 100 shown.

[0095] In Figure 4(a), filters (201, 202) are best suited for general noise (not wind noise). The first filter 201 is a high-pass filter for filtering the microphone signal 101; the second filter 202 is a low-pass filter for filtering the voice accelerometer signal 102. At very low frequencies near 0 kHz, the first filter 101 attenuates the microphone signal 101 below 3 kHz by up to 20 dB. The second filter 202 passes the voice accelerometer signal up to approximately 3-4 kHz.

[0096] In FIG4( b ), filters ( 201 , 202 ) are optimally suited for wind noise. First filter 201 is a high-pass filter for filtering microphone signal 101 ; second filter 202 is a low-pass filter for filtering voice accelerometer signal 102 . At very low frequencies near 0 kHz, first filter 101 attenuates microphone signal 101 below 3 kHz by up to 80 dB (even attenuating infinitely at approximately 0 kHz). In other words, compared to FIG4( a ), the low-frequency attenuation of microphone signal 101 is significantly enhanced. Second filter 202 passes the voice accelerometer signal up to approximately 3-4 kHz.

[0097] In FIG4(c), filters (201, 202) are optimized for normal noise, as in FIG4(a), but without any interpolation. First filter 201 is a high-pass filter for filtering microphone signal 101; second filter 202 is a low-pass filter for filtering voice accelerometer signal 102. As in FIG4(a), first filter 101 attenuates microphone signal 101 below 3 kHz by up to 20 dB at very low frequencies near 0 kHz. Second filter 202 can be shorter, for example, having 7 taps.

[0098] Figure 5The device 100 provided by the embodiment of the present invention is schematically shown (specifically, Figure 2 or Figure 3 1. The mode selection process performed by the device 100 shown in FIG. 1. Typically, the device 100 can determine whether the environment (its surroundings) is quiet or noisy, and can select the first mode when the environment is noisy, and select the second mode when the environment is quiet. In addition, when the device 100 determines that the environment is noisy, the device 100 can also determine whether the environment is windy. If the environment is windy, the low-frequency attenuation of the microphone signal 101 in the first mode can be enhanced. The mode can be changed gradually to avoid noise pumping effects and other discontinuities. In a windy environment, the higher frequencies (upper frequency band) of the microphone signal 101 can also be additionally attenuated, for example, by 18 dB.

[0099] Specifically, if Figure 5 As shown, the device 100 may further include an analysis filter bank 500, a high-pass filter 501, a first mixer 503, and a second mixer 502. The first mixer 503 may receive the speech accelerometer signal 101 and the microphone signal 101 as the output of the analysis filter bank 500 and may determine whether the environment is quiet. The second mixer 502 may receive the microphone signal 101 as the output of the analysis filter bank 500 and the same microphone signal 101 after passing through the high-pass filter 501 and may determine whether the environment is windy. The filter bank 200 may be selected separately, i.e., its filters (201, 202) may be selected according to the determined noise conditions, and may operate on the microphone signal 101 or the microphone signal 101 and the speech accelerometer signal 102 separately.

[0100] Figures 6(a) and 6(b) compare the frequency spectrum of the output signal of the device 100 provided by an embodiment of the present invention with the frequency spectrum of the output signal of a device that performs a conventional method (e.g., the spectral subtraction described in the background section). Specifically, spectrogram 6(a) is the result of using the device 100 in a first operating mode (i.e., generating a combined signal 105) and then applying a conventional noise suppression method to the combined signal 105. Specifically, spectrogram 6(b) is the result of applying the same conventional noise suppression method directly to the microphone signal (without using the voice accelerometer signal). Compared with the conventional device, spectrogram 6(a) of the output signal of the device 100 has significantly less noise. It is worth noting that in strong wind conditions, the device 100 can use the conventional method.

[0101] Figure 7 1 shows a method 700 provided by an embodiment of the present invention. The method 700 is suitable for reducing noise in a headset. The method 700 can be executed by the device 100, specifically by Figure 1 、 Figure 2 or Figure 3 The device 100 shown executes the method 700. Details of the method 700 may be implemented as described above with respect to the device 100.

[0102] The method 700 comprises: step 701 , obtaining 701 a microphone signal 101 ; step 702 , obtaining a speech accelerometer signal 102 ; and step 703 , adjusting the speech accelerometer signal 102 according to the microphone signal 101 .

[0103] In the first operating mode, the method 700 further includes: step 704, modifying the microphone signal by attenuating low frequencies (of the microphone signal 101); step 705, combining the modified microphone signal 104 with the adjusted voice accelerometer signal 103; and step 706, outputting the combined signal 105.

[0104] In the second operating mode ( Figure 7 (not shown), the method 700 further includes the following steps: delaying the microphone signal 101; and outputting the delayed microphone signal 204.

[0105] The invention has been described in conjunction with various exemplary embodiments and implementations. However, a person skilled in the art will be able to understand and implement other variations when implementing the claimed invention, based on a study of the drawings, the present invention and the independent claims. In the claims and the specification, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single element or other unit may fulfil the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not mean that a combination of these measures cannot be used effectively.

Claims

1. A device for reducing noise in a headset, characterized in that The device includes: a microphone, a vibration sensor, an adjustment module and a filter bank; The microphone is used to obtain a first microphone signal; The vibration sensor is used to obtain a vibration signal; The adjustment module is configured to adjust the vibration signal according to the first microphone signal; The filter bank is configured to output a second microphone signal based on the first microphone signal and the adjusted vibration signal.

2. The device according to claim 1, characterized in that In a first operating mode, the filter bank is configured to modify the first microphone signal by attenuating low frequencies, combine the modified first microphone signal with the adjusted vibration signal, and output the second microphone signal.

3. The device according to claim 2, characterized in that In a second operation mode, the filter bank is configured to delay the first microphone signal and output the second microphone signal.

4. The device according to claim 3, characterized in that The filter bank is further configured to gradually change between the first operation mode and the second operation mode.

5. The device according to claim 3 or 4, characterized in that The filter bank includes a plurality of filters, and the plurality of filters includes a first filter and a second filter, wherein the first filter is a high-pass filter and the second filter is a low-pass filter.

6. The device according to claim 5, characterized in that Each filter of the plurality of filters in the filter bank includes fewer than 10 taps.

7. The device according to claim 5 or 6, characterized in that In the first operation mode, the first filter is configured to filter the first microphone signal to obtain the modified first microphone signal; The second filter is configured to filter the adjusted vibration signal before combining the modified first microphone signal with the adjusted vibration signal.

8. The device according to any one of claims 5 to 7, characterized in that In the second operation mode, the first filter and the second filter are used to filter the first microphone signal to obtain the second microphone signal.

9. The device according to any one of claims 5 to 8, characterized in that The adjustment module is configured to adjust the vibration signal according to the first microphone signal at a sampling rate lower than a sampling rate of the first microphone signal; The second filter is configured to interpolate the adjusted vibration signal to the sampling rate of the first microphone signal.

10. The device according to any one of claims 5 to 9, characterized in that The filter bank is configured to adjust the plurality of filters when changing between the first operating mode and the second operating mode.

11. The device according to any one of claims 3 to 10, characterized in that The device also includes a noise suppression module; The noise suppression module is configured to apply noise suppression to the second microphone signal in the first operation mode and / or to apply noise suppression to the second microphone signal in the second operation mode.

12. The device according to any one of claims 3 to 11, characterized in that The apparatus further includes a first mixer; The first mixer is used to determine whether the environment is quiet or noisy; The filter bank is configured to: select the first operation mode if the environment is noisy, and select the second operation mode if the environment is quiet.

13. The device according to claim 12, characterized in that The apparatus further includes a second mixer; The second mixer is used to determine whether the environment is windy when the environment is confirmed to be noisy, The filter bank is configured to enhance the low-frequency attenuation of the first microphone signal in the first operation mode if there is wind in the environment.

14. The device according to any one of claims 1 to 13, characterized in that The adjustment module includes an adaptive filter; The adaptive filter is used to adjust the amplitude and phase of the vibration signal to the amplitude and phase of the first microphone signal.

15. A method for reducing noise in a headset, characterized in that The method comprises: obtaining a first microphone signal; Obtaining vibration signals; adjusting the vibration signal according to the first microphone signal; A second microphone signal is output based on the first microphone signal and the adjusted vibration signal.

16. The method according to claim 15, characterized in that The step of outputting a second microphone signal based on the first microphone signal and the adjusted vibration signal includes: In a first operating mode, the first microphone signal is modified by attenuating low frequencies, the modified first microphone signal is combined with the adjusted vibration signal, and the second microphone signal is output.

17. The method according to claim 16, characterized in that The step of outputting a second microphone signal based on the first microphone signal and the adjusted vibration signal includes: In a second operating mode, the first microphone signal is delayed and the second microphone signal is output.

18. The method according to claim 17, characterized in that Gradually change between the first operating mode and the second operating mode.

19. The method according to any one of claims 15 to 18, characterized in that: filtering the first microphone signal in the first operating mode to obtain the modified first microphone signal; The adjusted vibration signal is filtered before being combined with the modified first microphone signal.

20. The method according to any one of claims 15 to 19, characterized in that: In the second operation mode, the first microphone signal is filtered to obtain the second microphone signal.

Citation Information

Patent Citations

  • System and method of detecting a user's voice activity using an accelerometer

    US20140093093A1

  • Noise reduction methodology for wearable devices employing multitude of sensors

    US20170337933A1

  • Earbud speech estimation

    US20180367882A1