Hearing aid with own voice relief
By using beamforming filters and processing circuits in hearing aids to detect user speech, enhance sound within a selected angle range, and suppress user speech, the problem of poor directional hearing aid performance in noisy environments is solved, resulting in a more natural auditory experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NEWANTH HEARING CO LTD
- Filing Date
- 2024-09-15
- Publication Date
- 2026-05-05
AI Technical Summary
Existing hearing aids are not effective in providing directional hearing assistance in noisy environments, especially since amplifying the user's voice can mask ambient sounds.
A novel beamforming filter is employed, which enhances sound within a selected angle range through a combination of microphone array and speaker, while suppressing user speech. The filter mode is switched by detecting the user's speaking state using processing circuitry.
It improves the clarity of conversation partners' voices in noisy environments, reduces user voice interference, and provides a more natural auditory experience.
Smart Images

Figure CN120239973B_ABST
Abstract
Description
[0001] field
[0002] This invention relates generally to hearing aids, and more particularly to devices and methods for improving directional hearing.
[0003] background
[0004] Speech comprehension in noisy environments is a significant challenge for people with hearing impairments. In addition to gain loss, hearing impairment is often accompanied by a reduction in the temporal resolution of the sensory system. These characteristics further reduce the ability of people with hearing impairments to filter target sources from background noise, particularly their ability to understand speech in noisy environments.
[0005] Some newer hearing aids offer directional hearing modes to improve speech intelligibility in noisy environments. This mode utilizes an array of microphones and applies beamforming technology to combine multiple microphone inputs into a single directional audio output channel. The output channel has spatial characteristics that increase the contribution of sound waves from the target direction relative to sound waves from other directions.
[0006] For example, PCT International Publication WO 2017 / 158507 (the disclosure of which is incorporated herein by reference) describes a hearing aid device comprising a housing configured to be physically attached to a mobile phone. An array of microphones is spaced apart within the housing and configured to generate electrical signals in response to acoustic input from the microphones. An interface is fixed within the housing together with processing circuitry coupled to receive and process the electrical signals from the microphones to generate a combined signal for output via the interface.
[0007] As another example, PCT International Publication WO 2021 / 074818 (the disclosure of which is incorporated herein by reference) describes a device for hearing assistance, comprising an eyeglass frame including a front piece and temples, wherein one or more microphones are mounted at corresponding first positions on the front piece and configured to output electrical signals in response to a first sound wave incident on the microphones. A speaker mounted at a second position on one of the temples outputs a second sound wave. Processing circuitry generates a drive signal for the speaker by processing the electrical signals output from the microphones, such that the speaker reproduces a selected sound appearing in the first sound wave with a delay equal to 20% of the transmission time of the first sound wave from the first position to the second position, thereby causing constructive interference between the first and second sound waves.
[0008] Overview
[0009] The embodiments of the present invention described below provide improved apparatus and methods for hearing assistance.
[0010] The invention will be more fully understood from the following detailed description of embodiments thereof, taken in conjunction with the accompanying drawings, in which: Brief description of the attached diagram
[0012] Figure 1 This is a schematic diagram illustrating a hearing aid based on an eyeglass frame according to an embodiment of the present invention;
[0013] Figure 2 This is a block diagram schematically illustrating details of a hearing aid device according to an embodiment of the present invention; and
[0014] Figure 3 This is a flowchart illustrating a method for operating a hearing aid according to an embodiment of the present invention.
[0015] Detailed description
[0016] Overview
[0017] While directional hearing aids are necessary and microphone arrays offer theoretical benefits in this regard, in practice, the directional performance of hearing aids falls far short of that achieved through natural hearing. Typically, good directional hearing aids require a relatively large number of microphones that are well-spaced and inconspicuously designed, allowing the user to easily orient the hearing aid's response toward a point of interest, such as a conversation partner in a noisy environment. The processing circuitry applies beamforming filters to the signals output by the microphones in response to incident sound waves to generate an audio output that amplifies the sound striking the microphone array within an angular range around the direction of interest while suppressing background noise. The audio output should reproduce the natural auditory experience as closely as possible while minimizing bothersome artifacts.
[0018] One of these artifacts is the user's own voice. When the user's mouth is close to the microphone, the processing circuitry captures and amplifies the user's voice along with sounds received from the direction of interest in the environment. Users find the amplified sound of their own voice unnatural and disturbing. In some cases, the amplification of the user's voice may even mask sounds in the environment that the user wants to hear.
[0019] The embodiments of the invention described herein address this problem using a novel beamforming filter that amplifies sound impacting a microphone array within a selected angular range while suppressing speech from a user. In the disclosed embodiments, a microphone array mounted near the user's head outputs an electrical signal in response to incoming sound waves incident on the microphones. A speaker mounted near the user's ear outputs sound waves in response to a drive signal. Processing circuitry generates the drive signal by amplifying and filtering the electrical signal using the beamforming filter.
[0020] In some embodiments, the processing circuitry detects that a user is speaking and, in response to such detection, applies a beamforming filter to suppress the sound spoken by the user. For this purpose, the processing circuitry may, for example, apply two different beamforming filters: one that suppresses the sound spoken by the user, and another that does not. Both filters amplify the sound within a selected angular range, but the filter that does not suppress the sound spoken by the user generally provides better directionality in the beamforming mode and provides reduced white noise gain relative to the filter that suppresses the user's own speech. Therefore, the filter that does not suppress the sound spoken by the user is preferred as long as the user is not actually speaking. The processor applies the beamforming filter that suppresses the user's own speech when it detects that the user is speaking. The processing circuitry can determine which beamforming filter should be used at any given time by calculating and comparing the relative power levels output by the beamforming filters.
[0021] In some embodiments, the microphone and speaker are mounted on a frame that is mounted on the user's head. In the embodiment described below, the microphone and speaker are mounted on an eyeglass frame. To mitigate the impact of the user's own voice, the processing circuitry applies a beamforming filter that suppresses sound originating from a location below the eyeglass frame. Alternatively, the microphone and speaker may be mounted on other types of frames, such as virtual reality (VR) or augmented reality (AR) headsets, or in other types of mounting arrangements.
[0022] System Description
[0023] Figure 1 This is a schematic illustration of a hearing aid 20 integrated into an eyeglass frame 22 according to another embodiment of the present invention. An array of microphones 23, 24 is mounted at corresponding positions on the eyeglass frame 22 and outputs electrical signals in response to sound waves incident on the microphones. In the illustrated example, microphone 23 is mounted on the front part 30 of the frame 22, while microphone 24 is mounted on the temple 32, which is connected to the corresponding edge of the front part 30. Although Figure 1 The extensive array of microphones 23 and 24 shown is useful in some applications of the invention; however, the principles of signal processing and hearing assistance described herein can also be applied alternatively using a smaller number of microphones (with necessary modifications). For example, these principles can be applied using the array of microphones 23 on the front piece 30, as well as in devices using microphone mounting arrangements that are not necessarily based on eyeglasses.
[0024] Processing circuitry 26 is fixed within or otherwise connected to the eyeglasses frame 22 and coupled via wires 27 (such as traces on a flexible printed circuit) to receive electrical signals output from microphones 23 and 24. Although in Figure 1 The processing circuitry 26 is shown, but for simplicity, some or all of the processing circuitry may alternatively be located in the front element 30 or in a unit externally connected to the frame 22, at some point in the temple 32. The processing circuitry 26 mixes signals from the microphone to generate an audio output with a specific directional response, for example by applying beamforming to enhance sound originating from a selected angular range while suppressing background sound originating outside that range. Typically, although not strictly necessary, the directional response is aligned with the angular orientation of the frame 22.
[0025] Furthermore, processing circuit 26 detects when the user is speaking and applies beamforming to suppress spoken sounds while amplifying sounds hitting the microphone array within a selected angular range. These signal processing functions of circuit 26 will be described in more detail below. Alternatively, processing circuit 26 may apply a beamforming filter that consistently suppresses spoken sounds regardless of whether the user is speaking.
[0026] Processing circuitry 26 can deliver audio output to the user's ears via any suitable type of interface and speaker. In the illustrated embodiment, the audio output is generated by drive signals for driving one or more audio speakers 28 mounted on the temples 32, typically close to the user's ears. Although in Figure 1 The diagram shows only a single speaker 28 on each temple 32, but the device 20 may alternatively include only a single speaker on one of the temples 32, or the device 20 may include two or more speakers mounted on one or both of the temples 32. In the latter case, the processing circuitry 26 may apply beamforming functionality to the drive signal to direct sound waves from the speaker to the user's ear. Alternatively, the drive signal may be transmitted to a speaker inserted into the ear, or may be transmitted wirelessly (e.g., as a magnetic signal) to a telecoil in a hearing aid (not shown) worn by a user with eyeglasses.
[0027] Signal processing
[0028] Figure 2 This is a block diagram schematically illustrating details of the processing circuitry 26 in a hearing aid device 20 according to an embodiment of the present invention. The processing circuitry 26 can be implemented in a single integrated circuit chip, or alternatively, the functionality of the processing circuitry 26 can be distributed among multiple chips, which can be located inside or outside the eyeglass frame 22. Although in Figure 2A particular implementation is shown, but the processing circuit 26 may alternatively include any suitable combination of analog and digital hardware circuitry, as well as a suitable interface for receiving electrical signals output by the microphones 23, 24 and outputting drive signals to the speaker 28.
[0029] In this embodiment, microphones 23 and 24 include an integrated analog-to-digital converter that outputs digital audio signals to processing circuitry 26. Alternatively, processing circuitry 26 may include an analog-to-digital converter for converting the analog output of the microphones into digital form. Processing circuitry 26 typically includes suitable programmable logic unit 40, such as a digital signal processor (DSP) or gate array, which implements the necessary filtering and mixing functions to generate and output digital drive signals for speaker 28.
[0030] These filtering and mixing functions typically involve the application of two beamforming filters 42 and 43, the coefficients of which are selected to produce a desired directional response. Specifically, the coefficients of filter 42 are calculated to amplify the sound impacting frame 22 (and therefore the microphones 23 and 24) within a selected angular range; while the coefficients of filter 43 are calculated to amplify the sound impacting frame 22 within that selected angular range while suppressing sound emitted by the user of device 20, originating from a location below frame 22. Details of the filters available for these purposes are further described below.
[0031] Alternatively or additionally, processing circuitry 26 may include a neural network (not shown) trained to determine and apply coefficients to be used in filters 42 and 43. Further alternatively or additionally, processing circuitry 26 may include a microprocessor programmed in software or firmware to perform at least some of the functions described herein.
[0032] In implementing filters 42 and 43, processing circuitry 26 can apply any suitable beamforming function known in the art in the time or frequency domain. For example, beamforming algorithms that can be used in this context are described in the aforementioned PCT International Publication WO 2017 / 158507 (particularly pages 10-11) and U.S. Patent 10,567,888 (particularly column 9).
[0033] In one embodiment, processing circuitry 26 applies a minimum variance distortionless response (MVDR) beamforming algorithm when deriving the coefficients of beamforming filters 42 and 43. This algorithm is advantageous in achieving fine spatial resolution and distinguishing between sounds originating from the direction of interest and sounds originating from the user's own speech. The MVDR algorithm maximizes the signal-to-noise ratio (SNR) of the audio output by minimizing the average energy while maintaining a low target distortion. This algorithm can be implemented by calculating a vector F(ω) of complex weights of the output signal from each microphone at each frequency, as expressed by the following formula:
[0034]
[0035] In this formula, W(ω) is the propagation delay vector between microphones 23, representing the desired response of the beamforming filter as a function of angle and frequency; and S zz (ω) is the cross-spectral density matrix, representing the covariance of the acoustic signal in the time-frequency domain. To calculate the coefficients of filter 42, S is measured or calculated for isotropic far-field noise. zz (ω). To calculate the coefficients of filter 43, the cross-spectral density matrix of the user's own speech is added to the far-field noise.
[0036] In an alternative embodiment, processing circuitry 26 applies a Linear Constrained Minimum Variance (LCMV) algorithm when deriving the coefficients of beamforming filters 42 and 43. LCMV beamforming allows the filters to pass signals from the desired direction with a specified gain and phase delay, while minimizing the power of interfering signals and noise from all other directions. Additional constraints are imposed in the calculation of filter 43 specifically to eliminate output power from the direction of the user's mouth.
[0037] In some embodiments, processing circuitry 26 includes selection logic 44 that selects a beamforming filter to be applied at any given time. Selection logic 44 selects the beamforming filter based on whether the user is speaking. The selection decision may be based on the outputs of beamforming filters 42 and 43 themselves. For this purpose, processing circuitry 26 applies both beamforming filters 42 and 43 to the electrical signals output by microphones 23 and 24, and calculates the corresponding power levels P1 and P2 output by beamforming filters 42 and 43, respectively. Selection logic 44 compares the power levels, for example, by applying a threshold to the ratio of the power levels. Selection logic 44 selects filter 42 as long as P1 / P2 is less than the threshold. When P1 / P2 exceeds the threshold, selection logic 44 selects filter 43. The threshold may be preset, or it may alternatively be adaptively and / or adjusted according to user preferences.
[0038] Audio output circuitry 46, including, for example, a suitable codec and digital-to-analog converter, converts the digital drive signals output from filters 42 and 43 into analog form. Analog filter 48 performs further filtering and analog amplification to optimize the analog drive signal to speaker 28.
[0039] Control circuitry 50 (such as an embedded microcontroller) controls the programmable functions and parameters of processing circuitry 26, which may include selection logic 44. Communication interface 52 (e.g., ...) (Or other wireless interfaces) allow users and / or audiologists to set and adjust these parameters as needed. Power circuitry 54 (such as a battery inserted into the temple 32) provides power to other components of the processing circuitry.
[0040] Operating method
[0041] Figure 3 This is a schematic flowchart illustrating a method for operating a hearing aid device with self-speech suppression according to an embodiment of the present invention. For specificity and clarity, the method is described below with reference to the components of the device 20 as described above. Alternatively, the principle of this method can be applied to other devices and systems with suitable microphone arrays and signal processing capabilities.
[0042] To prepare for the real-time application of this method, in filter calculation step 60, beamforming filters 42 and 43 are calculated and loaded into device 20. The filters can be predefined and loaded based on standardized spatial response parameters during the fabrication of device 20. Alternatively or additionally, considering parameters such as the user's physiology, hearing impairment, and subjective preferences, the filter coefficients can be optimized for the user.
[0043] Initially, in standard beamforming step 62, when device 20 is turned on, selection logic 44 selects filter 42 to process the acoustic signals from microphones 23 and 24 without suppressing its own speech. Processing circuit 26 periodically applies beamforming filters 42 and 43 to the electrical signals output from microphones 23 and 24, and in power comparison step 64, compares the respective power levels P1 and P2 output from beamforming filters 42 and 43. As long as the ratio P1 / P2 remains below an applicable threshold, selection logic 44 continues to apply filter 42.
[0044] In step 64, detecting that P1 / P2 has exceeded the threshold is considered an indication that the user of device 20 has started speaking. In this case, in the self-speech suppression step 66, selection logic 44 switches to filter 43. Optionally, to avoid rapid switching between filters 42 and 43 that may interfere with the user, selection logic 44 applies a time lag when applying and removing beamforming filters after detecting that the user has started or stopped speaking. For this purpose, for example, the selection logic can apply a smoothing filter to the power value:
[0045]
[0046] In this expression, It is the measurement value of P1 or P2 at time t; and the value of the filter coefficient α is between 0 and 1. Select logic 44 using ratio. This determines when to change the filtering mode.
[0047] Additionally or alternatively, when it is detected that a user has started speaking (or subsequently stopped speaking), the processing circuit 26 can blend beamforming filters 42 and 43 during a transition period to avoid abrupt changes in the acoustic output to the speaker 28. For this purpose, the processing circuit uses a blending factor β to blend the filter coefficients, which varies gradually between 0 and 1 over time. During the transition period, the vector F of the filter coefficients applied by the processing circuit 26 is a linear combination of the corresponding coefficient vectors F1 and F2 of filters 42 and 43: F = (1-β)F1 + βF2. For example, the value of β can be calculated using the following algorithm with small values γ and ε between 0 and 1 in a series of incremental steps over time:
[0048] If P1 / P2 > threshold, then β = β + γ; otherwise, β = β - ε.
[0049] If β > 1, then β = 1.
[0050] If β < 0, then β = 0.
[0051] return Figure 3 When the processing circuit 26 applies self-speech suppression using filter 43 in step 66, in the further power comparison step 68, selection logic 44 continues to check the ratio P1 / P2. As long as this ratio remains above the threshold, the processing circuit 26 continues to apply self-speech suppression in step 66. However, when the ratio drops below the threshold, in the self-speech suppression deactivation step 70, selection logic 44 switches back to beamforming filter 42. The same time lag and mixing process described above (with necessary modifications) can also be applied in step 70. The operation of device 20 continues in step 62 until a further change is detected.
[0052] Example
[0053] Example 1: A system for hearing assistance includes an array of microphones configured to be mounted near the head of a user of the system and to output an electrical signal in response to a first sound wave incident on the microphones. A speaker is configured to be mounted near the user's ear and is configured to output a second sound wave in response to a drive signal applied to the speaker. Processing circuitry is configured to generate a drive signal by amplifying and filtering the electrical signal using a beamforming filter that enhances the first sound wave striking the array of microphones within a selected angular range while suppressing the second sound wave spoken by the user.
[0054] Example 2: According to the system described in Example 1, the processing circuit is configured to detect that a user is speaking and, in response to detecting that a user is speaking, apply a beamforming filter to suppress a second sound.
[0055] Example 3: According to the system described in Example 2, wherein the beamforming filter that suppresses the second sound is a first beamforming filter, and wherein the processing circuit is configured to apply the second beamforming filter whenever the user is not speaking, and to apply the first beamforming filter when it is detected that the user is speaking, wherein the second beamforming filter amplifies the first sound but does not suppress the second sound.
[0056] Example 4: According to the system described in Example 3, the processing circuit is configured to detect that the user is speaking by applying both a first beamforming filter and a second beamforming filter to an electrical signal and comparing the corresponding power levels output by the first beamforming filter and the second beamforming filter.
[0057] Example 5: According to the system described in Example 3 or 4, the processing circuit is configured to generate a drive signal by mixing a first beamforming filter and a second beamforming filter during a transition period when it detects that the user has started or stopped speaking.
[0058] Example 6: A system according to any of Examples 2-5, wherein the processing circuitry is configured to apply a time lag when applying and removing beamforming filters after detecting that the user has started or stopped speaking.
[0059] Example 7: A system according to any of the examples 1-6, wherein the beamforming filter includes a minimum variance distortionless response (MVDR) beamformer.
[0060] Example 8: A system according to any of the examples 1-6, wherein the beamforming filter includes a linearly constrained minimum variance (LCMV) beamformer.
[0061] Example 9: A system according to any of Examples 1-8, and including a frame configured for mounting on a head, wherein a microphone and a speaker are mounted at corresponding positions on the frame.
[0062] Example 10: The system according to Example 9, wherein the frame includes an eyeglass frame.
[0063] Example 11: A head-mounted device includes a frame configured for mounting on a user's head. An array of microphones is mounted at corresponding locations on the frame and configured to output an electrical signal in response to a first sound wave incident on the microphones. A speaker is mounted on the frame and configured to output a second sound wave in response to a drive signal applied to the speaker. Processing circuitry is configured to generate a drive signal by amplifying and filtering the electrical signal using a beamforming filter that enhances the first sound wave striking the array of microphones within a selected angular range while suppressing a second sound wave originating from a location below the eyeglass frame.
[0064] Example 12: The device according to Example 11, wherein the frame includes an eyeglass frame.
[0065] Example 13: The device according to Example 12, wherein the eyeglass frame includes a front and temples connected to respective edges of the front, wherein at least some of the microphones are mounted on the front and a speaker is mounted on one of the temples.
[0066] The features of any of the examples in Examples 2-7 can be similarly applied to any of the examples in Examples 11-13.
[0067] Example 14: A method for hearing assistance includes mounting an array of microphones near a user's head, the microphones outputting electrical signals in response to a first sound wave incident on the microphones. A speaker is mounted near the user's ear and outputs a second sound wave in response to a drive signal applied to the speaker. The drive signal is generated by amplifying and filtering the electrical signals using a beamforming filter that enhances the first sound wave striking the array of microphones within a selected angular range while suppressing the second sound wave spoken by the user.
[0068] Features of any of the examples in Examples 2-10, 12 and 13 can be similarly applied to Example 14.
[0069] The embodiments described above are exemplified by way of example, and the invention is not limited to what is specifically shown and described above. Rather, the scope of the invention includes combinations and sub-combinations of the various features described above, as well as variations and modifications of these features that would occur to those skilled in the art upon reading the foregoing description and that are not disclosed in the prior art.
Claims
1. A system (20) for hearing assistance, comprising: An array of microphones (23, 24), the microphones being configured to be mounted near the head of the user of the system and to output electrical signals in response to a first sound wave incident on the microphones; A loudspeaker (28) configured to be mounted near the user's ear and configured to output a second sound wave in response to a drive signal applied to the loudspeaker; and Processing circuitry (26) is configured to generate the drive signal by amplifying and filtering the electrical signal using a beamforming filter, the beamforming filter enhancing a first sound striking the array of microphones within a selected angle range while suppressing a second sound spoken by the user. The processing circuit is configured to detect that the user is speaking, and in response to detecting that the user is speaking, apply the beamforming filter to suppress the second sound. Wherein, the beamforming filter that suppresses the second sound is a first beamforming filter, and wherein the processing circuit is configured to apply a second beamforming filter whenever the user is not speaking, and to apply the first beamforming filter when the user is detected speaking, wherein the second beamforming filter amplifies the first sound but does not suppress the second sound, and The processing circuit is configured to detect that the user is speaking by applying both the first beamforming filter and the second beamforming filter to the electrical signal and comparing the corresponding power levels output by the first beamforming filter and the second beamforming filter.
2. The system according to claim 1, wherein, The processing circuit is configured to generate the drive signal by mixing the first beamforming filter and the second beamforming filter during a transition period when it detects that the user has started or stopped speaking.
3. The system according to claim 1, wherein, The processing circuitry is configured to apply a time lag when applying and removing the beamforming filter after detecting that the user has started or stopped speaking.
4. The system according to any one of claims 1-3, wherein, The beamforming filter includes a minimum variance distortionless response (MVDR) beamformer.
5. The system according to any one of claims 1-3, wherein, The beamforming filter includes a linearly constrained minimum variance (LCMV) beamformer.
6. The system according to any one of claims 1-3, the system comprising a frame configured for mounting on the head, wherein the microphone and the speaker are mounted at corresponding positions on the frame.
7. The system according to claim 6, wherein, The frame includes eyeglass frames.
8. A method for hearing assistance, comprising: An array of microphones (23, 24) is mounted near the user's head, and the microphones output electrical signals in response to a first sound wave incident on the microphones; The speaker (28) is mounted near the user's ear, and the speaker outputs a second sound wave in response to a drive signal applied to the speaker; and The drive signal is generated by amplifying and filtering the electrical signal using a beamforming filter, which enhances the first sound impacting the array of microphones within a selected angle range while suppressing the second sound spoken by the user. Generating the drive signal includes detecting that the user is speaking, and applying the beamforming filter to suppress the second sound in response to detecting that the user is speaking. Wherein, the beamforming filter that suppresses the second sound is a first beamforming filter, and wherein generating the drive signal includes filtering the electrical signal with a second beamforming filter as long as the user is not speaking, and applying the first beamforming filter when the user is detected speaking, the second beamforming filter amplifying the first sound but not suppressing the second sound, and Detecting that the user is speaking includes applying both the first beamforming filter and the second beamforming filter to the electrical signal and comparing the corresponding power levels output by the first beamforming filter and the second beamforming filter.
9. The method according to claim 8, wherein, Generating the drive signal includes applying a time lag in applying and removing the beamforming filter after detecting that the user has started or stopped speaking.
Citation Information
Patent Citations
Directional hearing aid
US10567888B2
Hearing aid
WO2017158507A1
Beamforming devices for hearing assistance
WO2021074818A1
Bilateral hearing aid system comprising temporal decorrelation beamformers
CN113940097A
Binaural hearing system providing beamformed and omnidirectional signal outputs
CN114631331A