Bone conduction headphone voice enhancement system and method

By using a multi-microphone system and filtering technology to mix bone conduction and air conduction voice signals, the problem of decreased speech recognition performance of personal listening devices in noisy environments has been solved, thus improving the voice quality of conversations between users and remote devices.

CN116569564BActive Publication Date: 2026-03-20GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing personal listening devices struggle to effectively enhance the user's own voice in noisy environments, leading to decreased speech recognition performance and negatively impacting the auditory experience of distant conversation partners due to noise.

Method used

A multi-microphone system, including external and internal microphones, is used, combined with low-frequency and high-frequency spatial filters and spectral filters, to enhance voice quality by mixing bone conduction and air conduction voice signals.

Benefits of technology

It improves speech recognition performance and remote conversation quality in noisy environments, reduces noise interference, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116569564B_ABST
    Figure CN116569564B_ABST
Patent Text Reader

Abstract

Systems and methods for enhancing a user's own voice in a headset include at least two external microphones (104, 106), an internal microphone (108), an audio input assembly operable to receive and process microphone signals, and a cross module configured to generate an enhanced voice signal. The audio processing assembly includes a low frequency branch including a low pass filter bank, a low frequency spatial filter (212), a low frequency spectral filter (214), and a high frequency branch including a high pass filter bank, a high frequency spatial filter (232), and a high frequency spectral filter (234).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application is a continuation of U.S. Patent Application No. 17 / 123,091, filed December 15, 2020, the disclosure of which is incorporated herein by reference. TECHNICAL FIELD

[0003] The present disclosure relates generally to audio signal processing, and more specifically, for example, to personal listening devices configured to enhance a user’s own voice. BACKGROUND

[0004] Personal listening devices (e.g., earphones, earbuds, etc.) typically include one or more loudspeakers that allow a user to listen to audio and one or more microphones that are used to pick up the user’s own voice. For example, a smartphone user wearing a Bluetooth headset can wish to engage in a telephone conversation with a remote user. In another application, a user can wish to provide voice commands to a connected device using the headset. Today’s headsets are generally reliable in noise-free environments. However, in noisy situations, the performance of applications such as automated speech recognizers can be significantly degraded. In such situations, the user can need to raise their voice substantially (with undesirable consequences of drawing attention), without guaranteeing optimal performance. Likewise, the aural experience of the remote conversation partner can be undesirably affected by the presence of background noise.

[0005] In view of the foregoing, there is a continuing need for improved systems and methods to provide efficient and effective voice processing and noise cancellation in headsets. SUMMARY

[0006] In accordance with the present disclosure, systems and methods for enhancing a user's own voice in a personal listening device, such as a headset or earpiece, are disclosed. The system for enhancing a headset user's own voice (e.g., a headset system) and method includes a plurality (at least two) of external microphones, an internal microphone, an audio processing component operable to receive and process microphone signals, and a cross module configured to generate an enhanced voice signal. The audio processing component includes a low frequency branch comprising a low pass filter bank, a low frequency spatial filter, and a low frequency spectral filter, and a high frequency branch comprising a high pass filter bank, a high frequency spatial filter, and a high frequency spectral filter. Based on the proposed scheme, the resulting voice signal is enhanced in terms of speech quality by mixing the bone-conducted voice of the low frequency band and the noise-suppressed air-conducted voice of the high frequency band. In one exemplary embodiment, the system and method for enhancing a headset user's own voice can further include a voice activity detector operable to detect the presence and absence of speech in the received and / or processed signals. The audio processing component can further include a (low frequency spectral) equalizer for compensating the low frequency spectral filter output.

[0007] In one exemplary embodiment, the external microphones and the internal microphone are part of a headset. The audio processing component can be arranged within the headset or within another device coupled to the headset (wirelessly or wired), such as a mobile device or a server.

[0008] The scope of the disclosure is defined by the claims, which are incorporated in this section by reference. Those skilled in the art will have a more complete understanding of the embodiments of the disclosure, and the advantages thereof, by referring to the following detailed description in conjunction with the attached drawings. The drawings will first be described briefly. BRIEF DESCRIPTION OF DRAWINGS

[0009] Various aspects of the disclosure, together with its advantages, will be more fully understood by reference to the following description, taken in conjunction with the accompanying drawings. It should be understood that the drawings are illustrative only and are not limiting of the disclosure. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the disclosure.

[0010] Figure 1 A personal listening device and use environment in accordance with one or more embodiments of the present disclosure are illustrated.

[0011] Figure 2 is a schematic diagram of an exemplary voice enhancement system in accordance with one or more embodiments of the present disclosure.

[0012] Figure 3is a schematic diagram of a low-frequency spatial filter according to one or more embodiments of the present disclosure.

[0013] Figure 4 illustrates an example of a low-frequency spectral filter according to one or more embodiments of the present disclosure.

[0014] Figure 5 is a flowchart of exemplary operations of a hybrid module and spectral filter module according to one or more embodiments of the present disclosure.

[0015] Figure 6 is an example diagram of an audio input processing component according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] The present disclosure presents various embodiments of improved systems and methods for enhancing a user's own voice in a personal listening device.

[0017] Many personal listening devices, such as earphones and earbuds, include one or more external microphones configured to sense external audio signals (e.g., microphones configured to capture a user's voice, reference microphones configured to sense ambient noise for active noise cancellation, etc.) and an internal microphone (e.g., an ANC error microphone positioned within or adjacent to a user's ear canal). The internal microphone can be positioned such that it senses a bone-conducted speech signal when the user is speaking. The sensed signal from the internal microphone can include low frequencies boosted from occlusion effects, and in some cases, leakage noise from outside the earphone.

[0018] In various embodiments, an improved multi-channel speech enhancement system is disclosed for processing a bone-conducted voice signal. The system includes at least two external microphones configured to pick up sound from outside a housing of the listening device, and at least one internal microphone within (or adjacent to) the housing. The external microphones are positioned at different locations of the housing and capture the user's voice through air conduction. The internal microphone is positioned to allow the internal microphone to receive the user's own voice through bone conduction.

[0019] In some embodiments, the speech enhancement system includes four processing stages. In a first stage, the speech enhancement system separates the input signal into high- frequency and low-frequency processing branches. In a second stage, a spatial filter is employed in each processing branch. In a third stage, the spatial filter output is post-filtered by a spectral filter stage. In a fourth stage, the low-frequency spectral filter output is compensated by an equalizer and mixed with the high-frequency processing branch output by a crossover module.

[0020] REFERENCE Figure 1Example operating environments will now be described in accordance with one or more embodiments of the present disclosure. In various environments and applications, a user 100 wearing a headset (or other personal listening device or "hearable") such as an earbud headset 102 can wish to control a device 110 (e.g., a smartphone, tablet, motor vehicle, etc.) through voice control or otherwise communicate voice communications in a noisy environment, such as through voice conversation with a user of a remote device. In many noise-free environments, voice recognition using an automatic speech recognizer (ASR) can be accurate enough to allow a reliable and convenient user experience, such as through voice commands received by an external microphone such as external microphone 104 and / or external microphone 106. However, in noisy situations, the performance of the ASR can degrade significantly. In such situations, the user 100 can compensate by speaking much louder, but optimal performance is not guaranteed. Similarly, the listening experience of a remote conversation partner is also greatly affected by the presence of background noise, which can interfere with the user's voice communications.

[0021] A common complaint about personal listening devices is that the voice intelligibility in the phone is poor when the user wears it in an environment with significant background noise and / or strong wind. The noise can significantly hinder the user's voice intelligibility and degrade the user experience. Typically, the external microphone 104 receives more noise than the internal microphone 108 due to the attenuation effect of the earphone housing. In addition, wind noise can also occur at the external microphone due to the local air turbulence at the microphone. Wind noise is typically non-stationary, with its power mostly limited in the low frequency band, e.g., < 1500 Hz.

[0022] Unlike the air-conducted external microphone, the internal microphone 108 is positioned to enable it to sense the user's voice through bone conduction. The bone conduction response is strong in the low frequency band (< 1500 Hz) but weak in the high frequency band. If the earphone is well-sealed, the internal microphone is isolated from the wind, allowing it to receive a clearer user voice in the low frequency band. The systems and methods disclosed herein include enhancing the voice quality by mixing the bone-conducted voice in the low frequency band and the noise-suppressed air-conducted voice in the high frequency band.

[0023] In the illustrated embodiment, the earbud headset 102 is an active noise cancellation (ANC) earbud that includes multiple external microphones (e.g., external microphones 104 and 106) to capture the user's own voice and generate a reference signal corresponding to the ambient noise for cancellation. An internal microphone (e.g., internal microphone 108) is installed in the housing of the earbud headset 102 and is configured to provide an error signal for feedback ANC processing. Thus, the proposed system can use the existing internal microphone as a bone conduction microphone without adding extra microphones to the system.

[0024] In the present disclosure, a robust and computationally efficient noise cancellation system and method is disclosed based on utilizing both a headset-external microphone, such as external microphones 104 and 106, and a headset- or ear canal-internal microphone, such as internal microphone 108. In various embodiments, user 100 can whisper voice communications or voice commands to device 110, even in very noisy situations. The systems and methods disclosed herein improve voice processing applications, such as speech recognition and voice communication quality with a far-end user. In various embodiments, internal microphone 108 is a component of a noise cancellation system of a personal listening device, which further includes a speaker 112 configured to output sound for user 100 and / or generate an anti-noise signal to cancel ambient noise, an audio processing component 114 including digital and analog circuitry and logic for processing audio, including active noise cancellation and voice enhancement, for input and output, and a communication component 116 for communicating (e.g., wired, wireless, etc.) with a host device, such as device 110. In various embodiments, audio processing component 114 can be disposed in earbuds / headset 102, device 110, or one or more other devices or components.

[0025] The systems and methods disclosed herein have many advantages over existing solutions. First, the embodiments disclosed herein use two spatial filters separately for high and low frequency processing. The high frequency spatial filter suppresses high frequency noise in the external microphone signal. In some embodiments, a conventional air-conducted microphone spatial filtering solution, such as a fixed beamformer (e.g., delay-and-sum, super-directional beamformer, etc.), an adaptive beamformer (e.g., multi-channel Wiener filter (MWF), spatial maximum SNR filter (SMF), minimum variance distortion response (MVDR), etc.), and e.g., blind source separation, etc., can be used.

[0026] The geometry / position of the external microphone on the personal listening device can be optimized to achieve acceptable noise reduction performance, which can depend on the type of personal listening device and the intended use environment. The low frequency spatial filter suppresses low frequency noise by exploiting the speech and noise transfer functions between the external and internal microphones. This information cannot be well determined by the positions of the external and internal microphones alone. The earphone design and the user’s physical characteristics (head shape, bone, hair, skin, etc.) have a large impact on the transfer functions. Typical air-conducted solutions would perform poorly in most cases. Therefore, the embodiments disclosed herein use separate spatial filters for speech enhancement in high and low frequency processing, respectively.

[0027] Second, unlike most conventional speech enhancement systems that use only air-conducted microphones, the proposed system achieves higher output SNR in the low frequency band by using the bone-conducted microphone signal, which has a higher input SNR than the external microphone.

[0028] Third, the present disclosure discloses the application of a post-filter spectral filter to further improve voice quality. The function of this stage is to reduce the noise residue of the spatial filter stage. Existing solutions typically assume that the bone conduction signal is noise-free. However, this is not always true. Depending on the noise type, noise level, and the seal of the earphone, wind and background noise can still penetrate the earphone housing. The spectral filter stage is configured to reduce noise not only for high frequencies but also for low frequencies, and a multi-channel spectral filter can be used.

[0029] Fourth, the solution disclosed herein can be applied to both acoustic background noise and wind noise. Conventional solutions typically employ different techniques to handle different types of noise.

[0030] Figure 2 One embodiment of a system 200 with two external microphones (external microphone 1 and external microphone 2) and one internal microphone (internal microphone) is illustrated. Embodiments of the present disclosure can be implemented in a system with two or more external microphones and at least one internal microphone. For example, if there are two external microphones, one can be positioned at the left ear side and the other at the right ear side. The external microphones can also be on the same side, for example, one in front of the personal listening device and the other in the back.

[0031] The two external microphone signals (e.g., which include sounds received via air conduction) are denoted as X e,1 (f, t) and X e,2 (f, t). The internal microphone signal (e.g., which can include bone conduction sounds) is denoted as X i (f, t), where f denotes frequency and t denotes time.

[0032] The signals X e,1 (f, t), X e,2 (f, t), and X i (f, t) are passed through a low-pass filter bank 210 and processed to generate X e,1,l (f, t), X e,2,l (f, t), and X i,t (f, t). The two external microphone signals X e,1 (f, t) and X e,2 (f, t) are also passed through a high-pass filter bank 230, which processes the received signals to generate X e,1,h (f, t) and X e,2,h (f, t). Note that the internal microphone signal X i(f, t) does not have much speech signal at high frequencies and it is not used for the high frequency processing branch 204. The cutoff frequencies of the low pass filter bank 210 and the high pass filter bank 230 can be fixed and predetermined. In some embodiments, the optimal values depend on the acoustic design of the earpiece. In some embodiments, 3000 Hz is used as a default value.

[0033] Second, the low frequency spatial filter 212 of the low frequency branch 202 processes the low frequency signal X e,1,l (f, t), X e,2,l (f, t), and X i,l (f, t) and obtains a low frequency speech and error estimate D l (f, t), and ε l (f, t). The high frequency spatial filter 232 processes the high frequency signal X e,1,h (f, t), and X e,2,h (f, t) and obtains a high frequency speech and error estimate D h (f, t), and ε h (f, t).

[0034] Referring to Figure 3 One exemplary embodiment of the low frequency spatial filter 212 will now be described in accordance with one or more embodiments. The low frequency spatial filter 212 includes a filter module 310 and a noise suppression engine 320. The filter module 310 applies a spatial filter gain on the input signal and obtains a speech and error estimate,

[0035]

[0036] ε l (f, t) = X i,l (f, t) - D l (f, t),

[0037] where h S (f, t) is a spatial filter gain vector, X l (f, t) = [X e,1,l (f, t) X e,2,l (f, t) X i,l (f, t) T The superscript H denotes the Hermitian transpose. Since X e,1,l (f, t), X e,2,l (f, t), and X i,l (f, t) are not stationary, the filter gain is adaptively computed by the noise suppression engine 320.

[0038] The noise suppression engine 320 derives h S(f, t). There are several spatial filtering algorithms that can be used by the noise suppression engine 320, such as Independent Component Analysis (ICA), Multi-Channel Wiener Filter (MWF), Spatial Maximum SNR Filter (SMF), and their derivatives. An example ICA algorithm is discussed in U.S. Patent Publication No. US20150117649A1, entitled "Selective Audio Source Enhancement," which is incorporated by reference herein in its entirety.

[0039] Without loss of generality, for example, the MWF finds a spatial filter vector h that minimizes S (f, t),

[0040]

[0041] where E() denotes the expectation computation. The above minimization problem has been extensively studied, and one solution is

[0042]

[0043] where I is the identity matrix, and xx (f, t) is the covariance matrix of X l (f, t), and vv (f, t) is the covariance matrix of the noise. The covariance matrix xx (f, t) is estimated via

[0044]

[0045] where a is a smoothing factor. The noise covariance matrix vv (f, t) can be estimated in a similar manner as when there is only noise. The presence of speech can be identified by a Voice Activity Detection (VAD) flag, which is generated by the VAD module 220, discussed in further detail below.

[0046] The SMF is another spatial filter that maximizes the SNR of the speech estimate l (f, t). It is equivalent to solving a generalized eigenvalue problem

[0047] Φ xx (f, t)h S (f, t) = l max Φ vv (f, t)h S (f, t),

[0048] where l max is the largest eigenvalue of .

[0049] Like the low-frequency spatial filter 212, the high-frequency spatial filter 232 has the same general structure when its spatial filtering algorithm is adaptive, such as ICA, MWF, and SMF. When the spatial filter is fixed, such as using delay-and-sum or hyperdirected beamformers, the high-frequency spatial filter 232 can be simplified to a filter module, where h S The values of (f, t) are fixed and predetermined.

[0050] For example, for a system using a delay-and-sum beamformer, the spatial filter gain is where is the time delay between the two external microphones.

[0051] For a hyperdirected beamformer, for example,

[0052]

[0053] where Γ(f) is a 2x2 pseudo-coherence matrix corresponding to the spherical isotropic noise In different embodiments, the fixed spatial gain depends on the speech time delay between the two external microphones, which can be measured during the earphone design.

[0054] Referring to Figure 4 One exemplary embodiment of the low-frequency spectral filter 214 will now be described in further detail. In some embodiments, the high-frequency spectral filter 234 has the same structure, which is omitted here for simplicity. The low-frequency spectral filter 214 includes a feature evaluation module 410, an adaptive classifier 420, and an adaptive mask computation module 430.

[0055] The adaptive mask computation module 430 is configured to generate time and frequency varying mask gains to reduce the residual noise within D l (f, t). To derive the mask gains, certain inputs are used for the mask computation. These inputs include the speech and error estimate outputs D l (f, t) and ε l (f, t) from the spatial filter, the VAD 220 output, and the adaptive classification results obtained from the adaptive classifier module 420. Thus, the signals D l (f, t) and ε l (f, t) are forwarded to the feature evaluation module 410, which converts the signals into features representing the SNR of D l (f, t). The feature selection in one embodiment includes:

[0056]

[0057] L l,2(f, t) = c(|D l (f, t) | |ε l (f, t) |)

[0058] L l,3 (f, t) = c|D l (f, t) |

[0059] where c is a constant to limit the feature value in the range of 0 to 1. The feature evaluation module 410 can compute and forward one or more features to the adaptive classifier module 420.

[0060] The adaptive classifier is configured to perform online training and classification of features. In various embodiments, it can apply hard-decision or soft-decision classification algorithms. For hard-decision algorithms such as K-means, decision trees, logistic regression, and neural networks, the adaptive classifier identifies D l (f, t) as speech or noise. For soft-decision algorithms, the adaptive classifier computes the probability that D l (f, t) belongs to speech. Typical soft-decision classifiers that can be used include Gaussian mixture models, hidden Markov models, and Bayesian algorithms based on importance sampling such as Markov chain Monte Carlo.

[0061] The adaptive mask computation module 430 is configured to adapt the gain based on D l (f, t), ε l (f, t), the VAD output (from the VAD 220), and the real-time classification result from the adaptive classifier 420 to minimize the residual noise in D l (f, t). More details on the implementation of the adaptive mask computation module can be found in U.S. Patent Publication No. US20150117649A1 entitled "Selective Audio Source Enhancement", the entirety of which is incorporated herein by reference.

[0062] Returning to Figure 2 , in the low-pass branch 202, the enhanced speech S l (f, t) after the spectral filter is compensated by the equalizer 216 to cancel the bone conduction distortion. The equalizer 216 can be fixed or adaptive. In the adaptive configuration, when speech is detected by the VAD 220, the equalizer 216 tracks the transfer function between S l (f, t) and the external microphone and applies that transfer function to S l (f, t). The equalizer 216 can compensate across the entire low frequency band or only in a portion of the low frequency band. The high frequency processing branch 204 does not use the internal microphone signal X i (f, t), so its spectral filter output Sh (f, t) is free of bone conduction distortion.

[0063] Figure 5 is a flowchart illustrating an example process 500 for operating the adaptive equalizer 216. At step 510, the equalizer receives signals S l (f, t), X e,1,l (f, t) and X e,2,l (f, t), and at step 512, checks the VAD flag. If the VAD detects speech, the equalizer updates the transfer function and There are many well-known methods to track H1(f, t) and H2(f, t). One method is and where and is the average of X e,1,l (f, t), X e,2,l (f, t) and S l (f, t) over time. Other methods include the Wiener filter, the subspace method, and the least mean square filter. Here, we use the H1(f, t) estimate as an example. In the Wiener filter method, H1(f, t) is tracked by

[0064]

[0065] where, and

[0066] For example, the subspace method estimates the covariance matrix where and finds the eigenvector β = [β1 β2] corresponding to the largest eigenvalue of T . Then,

[0067] In the least mean square filter, H1(f, t) is tracked by

[0068]

[0069] After the estimation of H1(f, t) and H2(f, t), the adaptive equalizer compares the amplitude of the spectral output |S l (f, t) | to a threshold value, which is used in step 540 to determine the level of bone conduction distortion. In various embodiments, the threshold value can be a fixed predetermined value or a variable depending on the external microphone signal strength.

[0070] If the spectral output exceeds the amplitude threshold, the adaptive equalizer performs distortion compensation (step 550), i.e.

[0071]

[0072] where ci and c2 are constants. For example, ci = 1 and c2 = 0 compensates with respect to external microphone 1. If the spectral output is below the threshold, no compensation is needed (step 560), and Note that the adaptive equalizer described above performs both amplitude and phase compensation. In various embodiments, only amplitude compensation is performed.

[0073] Referring back to Figure 2 The final stage is the cross module 236, which mixes the outputs of the low and high bands. VAD information is used extensively in the system, and any suitable voice activity detector can be used with the present disclosure. For example, the estimated voice DOA and a priori knowledge of mouth position can be used to determine whether the user is speaking. Another example is the inter-channel level difference (ILD) between the internal and external microphones. When the user is speaking, the ILD will exceed the low band voice detection threshold.

[0074] Embodiments of the present disclosure can be implemented in various devices with two or more external microphones and at least one internal microphone within the device housing, such as earphones, smart glasses, and VR devices. Embodiments of the present disclosure can apply fixed and adaptive spatial filters in the spatial filtering stage, the fixed spatial filters can be delay-and-sum and superdirective beamformers, and the adaptive spatial filters can be independent component analysis (ICA), multi-channel Wiener filter (MWF), spatial maximum SNR filter (SMF), and their derivatives.

[0075] In various embodiments, various adaptive classifiers can be used in the spectral filtering stage, such as K-means, decision tree, logistic regression, neural network, hidden Markov model, Gaussian mixture model, Bayesian statistics, and their derivatives.

[0076] In various embodiments, various algorithms can be used in the spectral filtering stage, such as Wiener filter, subspace method, maximum a posteriori spectral estimator, maximum likelihood amplitude estimator.

[0077] Figure 6 is a schematic diagram of an audio processing component 600 for processing audio input data according to one example embodiment. The audio processing component 600 generally corresponds to Figures 1-5The systems and methods disclosed herein, and can share any of the functionality previously described herein. The audio processing component 600 can be implemented in hardware, or as a combination of hardware and software, and can be configured to operate on a digital signal processor, a general purpose computer, or other suitable platform.

[0078] As shown, the audio processing component 600 includes a memory 620, which can be configured to store program logic, and a digital signal processor 640. In addition, the audio processing component 600 includes a high frequency spatial filtering module 622, a low frequency spatial filtering module 624, a voice activity detector 626, a high frequency spectral filtering module 628, a low frequency spectral filtering module 630, an equalizer 632, an ANC processing component 634, and an audio input / output processing module 636, some or all of which can be stored as executable program instructions in the memory 620. Figure 6

[0079] Figure 6 Also shown in the figure are earpiece microphones, including external microphones 602 and 603, and an internal microphone 604, which are communicatively coupled to the audio processing component 600 in a wired (e.g., hardwire) or wireless (e.g., Bluetooth) manner. An analog-to-digital converter component 606 is configured to receive analog audio inputs and generate corresponding digital audio signals to the digital signal processor 640 for processing as described herein.

[0080] In some embodiments, the digital signal processor 640 can execute machine-readable instructions (e.g., software, firmware, or other instructions) stored in the memory 620. In this regard, the processor 640 can perform any of the various operations, processes, and techniques described herein. In other embodiments, the processor 640 can be replaced and / or supplemented with specialized hardware components to perform any desired combination of the various techniques described herein. The memory 620 can be implemented as a machine-readable medium that stores various machine-readable instructions and data. For example, in some embodiments, the memory 620 can store an operating system and one or more applications as machine-readable instructions that can be read and executed by the processor 640 to perform the various techniques described herein. In some embodiments, the memory 620 can be implemented as a non-volatile memory (e.g., a flash memory, a hard disk, a solid state drive, or other non-transitory machine-readable medium), a volatile memory, or a combination thereof.

[0081] In various embodiments, the audio processing component 600 is implemented within an earpiece, or a device such as a smartphone, tablet, mobile computer, electrical user device, or other device that processes audio data through an earpiece. In operation, the audio processing component 600 produces an output signal, which can be stored in memory, used by other device applications or components, or transmitted to another device for use. ​

[0082] It should be apparent that the above disclosed has many advantages over the prior art. The approach disclosed herein has lower implementation cost than the traditional approach, and does not require precise prior training / calibration, nor the availability of a specific activity detection sensor. It also has the advantage of being compatible with and easy to integrate with existing headsets, provided there is space to accommodate a second internal microphone. The traditional approach requires prior training, is computationally complex, and the presented results are unacceptable for many human listening environments.

[0083] In one embodiment, a method for enhancing a headset user's own voice includes receiving a plurality of external microphone signals from a plurality of external microphones configured to sense external sounds through air conduction, receiving an internal microphone signal from an internal microphone configured to sense bone conduction sounds from the user during speech, processing the external microphone signals and the internal microphone signal through low pass processing, including low frequency spatial filtering and low frequency spectral filtering each signal, processing the external microphone signals through high pass processing, including high frequency spatial filtering and high frequency spectral filtering each signal, and mixing the low pass processed signals and the high pass processed signals to generate an enhanced voice signal. Based on the proposed approach, the resulting voice signal is enhanced in terms of speech quality by mixing the bone conduction voice in the low frequency band and the noise suppressed air conduction voice in the high frequency band.

[0084] In various embodiments, the low pass processing further includes low pass filtering of the external microphone signals and the internal microphone signal, and / or the high pass processing further includes high pass filtering of the external microphone signals. The low frequency spatial filtering can include generating a low frequency speech and an error estimate, and the low frequency spectral filtering can result in generating an enhanced speech signal, which is "enhanced" in view of the filter speech signal implemented. The method can further include applying an equalization filter to the enhanced speech signal to mitigate distortion from the bone conduction sounds, detecting voice activity in the external microphone signals and / or the internal microphone signal, and / or receiving the speech signal, the error signal, and the voice activity detection data, and updating the transfer function if voice activity is detected. To detect voice activity, an inter-channel level difference (ILD) between the internal microphone and the external microphones can be used. When the user speaks, the ILD will exceed a low frequency band voice detection threshold, thereby generating voice activity detection data indicating detected voice activity.

[0085] In some embodiments of the method, the low-frequency spatial filtering comprises applying a spatial filter gain on the signal and generating a voice and error estimate, wherein the spatial filter gain is adaptively computed based at least in part on a noise suppression process. The low-frequency spectral filtering can comprise evaluating features from the voice and error estimate, adaptively classifying the features and computing an adaptive mask. In one exemplary embodiment, computing the adaptive mask comprises computing a mask gain to reduce residual noise in the low-pass processed signal. For example, computing the mask gain comprises using the voice and error estimate output from the low-frequency spatial filter (used for the low-frequency spatial filtering), the output from the voice activity detection and the adaptive classification result from an adaptive classifier module, the result of which indicates whether the voice output from the low-frequency spatial filter comprises speech. The mask gain is adapted to minimize the residual noise based on the aforementioned parameters, for example disclosed in US 20150117649 Al. The method can further comprise comparing the amplitude of the spectral output to a threshold to determine a bone conduction distortion level and applying voice compensation based on the comparison.

[0086] In some embodiments, a system comprises: a plurality of external microphones configured to sense external sounds through air conduction and generate corresponding external microphone signals; an internal microphone configured to sense bone conduction of a user during speech and generate a corresponding internal microphone signal; a low-pass processing branch configured to receive the external microphone signals and the internal microphone signal and generate a low-pass output signal; a high-pass processing branch configured to receive the external microphone signals and generate a high-pass output signal; and a crossover module configured to mix the low-pass output signal and the high-pass output signal to generate an enhanced voice signal. Other features and modifications disclosed herein can also be included.

[0087] The above disclosure is not intended to limit the disclosure to the precise form or specific use disclosed. Thus, various alternate embodiments and / or modifications to the disclosed embodiments, whether explicitly described or implied, are within the scope of the disclosure. Having thus described embodiments of the disclosure, a person of ordinary skill in the art will recognize that changes can be made in form and detail without departing from the scope of the disclosure. Accordingly, the disclosure is limited only by the claims.

Claims

1. A method for enhancing the voice of a headset user, comprising: Receive signals from multiple external microphones, the multiple external microphones being configured to sense external sound via air conduction; Receives signals from an internal microphone, which is configured to sense bone conduction sound from the user during speech; The external microphone signal and the internal microphone signal are processed by low-pass processing, the low-pass processing including: The low-frequency speech estimation and error estimation are obtained at least in part based on filtering a first set of signals corresponding to the external microphone signal and the internal microphone signal using a low-frequency spatial filter. The output of the low-frequency spectrum filter is obtained at least in part based on filtering the low-frequency speech estimation and the error estimation using a low-frequency spectrum filter; and One or more low-pass processed signals are generated, at least in part, based on the output of the low-frequency spectrum filter; The external microphone signal is processed by high-pass processing, while the internal microphone signal is not processed, to generate one or more high-pass processed signals. The high-pass processing includes filtering a second set of signals corresponding to the external microphone signal using a high-frequency spatial filter and a high-frequency spectral filter; and At least one of the one or more low-pass processed signals and at least one of the one or more high-pass processed signals are mixed to generate an enhanced voice signal.

2. The method according to claim 1, wherein, The low-pass processing further includes low-pass filtering of the external microphone signal and the internal microphone signal.

3. The method according to claim 1, wherein, The high-pass processing further includes high-pass filtering of the external microphone signal.

4. The method according to claim 1, wherein, Filtering via the low-frequency spatial filter includes generating the low-frequency speech estimate and the error estimate, and filtering via the low-frequency spectral filter includes generating an enhanced speech signal corresponding to the output of the low-frequency spectral filter.

5. The method of claim 4, further comprising applying an equalization filter to the enhanced speech signal to reduce distortion from the bone conduction sound.

6. The method of claim 1, further comprising detecting voice activity in the external microphone signal and / or the internal microphone signal.

7. The method according to claim 1, wherein, Filtering via the low-frequency spatial filter includes applying a spatial filtering gain to the first set of signals and generating the low-frequency speech estimate and the error estimate, wherein the spatial filtering gain is adaptively calculated at least in part based on a noise suppression process.

8. The method according to claim 7, wherein, Filtering via the low-frequency spectrum filter includes evaluating features from the low-frequency speech estimate and the error estimate, adaptively classifying the features, and calculating an adaptive mask.

9. The method of claim 1, further comprising: Receives voice signals, error signals, and voice activity detection data; as well as If voice activity is detected, the transfer function is updated.

10. The method of claim 9, further comprising: The amplitude of the spectral output is compared to a threshold to determine the level of bone conduction distortion. Voice compensation is applied based on the comparison.

11. A system comprising: Multiple external microphones are configured to sense external sounds via air conduction and generate external microphone signals corresponding to the sensed external sounds; An internal microphone is configured to sense bone conduction sounds from the user during speech and generate an internal microphone signal corresponding to the sensed bone conduction sounds; A low-pass processing branch is configured to process the external microphone signal and the internal microphone signal through low-pass processing, the low-pass processing including: The low-frequency speech estimation and error estimation are obtained at least in part based on filtering a first set of signals corresponding to the external microphone signal and the internal microphone signal using a low-frequency spatial filter. The output of the low-frequency spectrum filter is obtained at least in part based on filtering the low-frequency speech estimation and the error estimation using a low-frequency spectrum filter; and One or more low-pass processed signals are generated, at least in part, based on the output of the low-frequency spectrum filter; and A high-pass processing branch is configured to process the external microphone signal via high-pass processing, but not the internal microphone signal, to generate one or more high-pass processed signals, the high-pass processing including filtering a second set of signals corresponding to the external microphone signal via a high-frequency spatial filter and a high-frequency spectral filter; and The crossover module is configured to mix at least one of the one or more low-pass processed signals and at least one of the one or more high-pass processed signals to generate an enhanced voice signal.

12. The system according to claim 11, wherein, The low-pass processing branch further includes a low-pass filter bank configured to filter the external microphone signal and the internal microphone signal.

13. The system according to claim 11, wherein, The high-pass processing branch further includes a high-pass filter bank configured to filter the external microphone signal.

14. The system according to claim 11, wherein, The low-pass processing branch further includes the low-frequency spatial filter configured to generate the low-frequency speech estimate and the error estimate, and the low-frequency spectral filter configured to generate an enhanced speech signal corresponding to the output of the low-frequency spectral filter.

15. The system of claim 14, further comprising an equalization filter configured to mitigate distortion from bone conduction in the enhanced speech signal.

16. The system of claim 11, further comprising a voice activity detector configured to detect voice activity in the external microphone signal and / or the internal microphone signal.

17. The system according to claim 11, wherein, The low-frequency spatial filter is configured to apply a spatial filtering gain to the first set of signals and generate the low-frequency speech estimate and the error estimate, wherein the spatial filtering gain is adaptively calculated at least in part based on a noise suppression process.

18. The system according to claim 17, wherein, The low-pass processing branch further includes the low-frequency spectrum filter, which is configured to evaluate features from the low-frequency speech estimation and the error estimation, adaptively classify the features, and compute an adaptive mask.

19. The system of claim 11, further comprising an equalizer, the equalizer being configured to: Receives voice signals, error signals, and voice activity detection data; and If voice activity is detected, the transfer function is updated.

20. The system according to claim 19, wherein, The equalizer is further configured as follows: The amplitude of the speech signal's spectral output is compared to a threshold to determine the level of bone conduction distortion. Distortion compensation is applied based on the comparison.

Citation Information

Patent Citations

  • Selective Audio Source Enhancement

    US20150117649A1

  • Audio headset

    CN106792305A