Hearing aids that reduce the sound of your own voice
The hearing aid device with beamforming filters suppresses the user's voice and enhances environmental sounds, addressing the unnatural amplification issue in directional hearing aids, thereby improving speech understanding in noise.
Patent Information
- Application Number
- JP2025538610
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-28
- Filing Date
- 2024-09-15
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2044-09-15
AI Technical Summary
Hearing aids with directional modes fail to effectively suppress the user's own voice, leading to an unnatural and unpleasant amplification that can mask desired environmental sounds.
A hearing aid device with an array of microphones and speakers mounted on a frame, such as eyeglasses, uses beamforming filters to suppress the user's voice when speaking and emphasize environmental sounds within a selected angular range, employing algorithms like MVDR and LCMV to optimize sound directionality.
The device provides a natural hearing experience by minimizing the user's own voice amplification and enhancing environmental sounds, improving speech intelligibility in noisy environments.
Smart Images

Figure 2026501606000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to hearing aids, and more particularly to devices and methods for improving directional hearing. [Background technology]
[0002] Speech understanding in noisy environments is a significant problem for people with hearing impairments. Hearing impairments are typically accompanied by a loss of gain as well as a reduction in the temporal resolution of the sensory system. These characteristics further reduce the hearing-impaired person's ability to filter target sound sources from background noise, especially to understand speech in noisy environments.
[0003] Some newer hearing aids offer a directional hearing mode that improves speech intelligibility in noisy environments. This mode uses an array of microphones and applies beamforming techniques to combine multiple microphone inputs into a single directional audio output channel. The output channel has spatial characteristics that increase the contribution of acoustic waves arriving from a target direction relative to contributions from other directions.
[0004] For example, PCT International Patent Application Publication No. 2017 / 158507, the disclosure of which is incorporated herein by reference, describes a hearing aid device including a case configured to be physically secured to a mobile phone. An array of microphones are spaced apart within the case and configured to generate electrical signals in response to acoustic input to the microphones. An interface is secured within the case along with processing circuitry coupled to receive and process the electrical signals from the microphones and generate a combined signal for output via the interface.
[0005] As another example, PCT International Patent Application Publication No. 2021 / 074818, the disclosure of which is incorporated herein by reference, describes a hearing aid device including a spectacle frame including a front piece and temples, with one or more microphones mounted at respective first positions on the front piece and configured to output electrical signals in response to first acoustic waves incident on the microphones. A speaker mounted at a second position on one of the temples outputs a second acoustic wave. Processing circuitry processes the electrical signals output by the microphones to generate a drive signal for the speaker, causing the speaker to reproduce a selected sound generated by the first acoustic wave from the first position to the second position with a delay equal to or less than 20% of the transit time of the first acoustic wave, thereby creating constructive interference between the first and second acoustic waves. Summary of the Invention
[0006] The embodiments of the present invention described below provide improved devices and methods for assisting hearing.
[0007] The present invention will be more fully understood from the following detailed description of the embodiments, when read in conjunction with the drawings. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic diagram illustrating a hearing aid device based on eyeglass frames, according to one embodiment of the present invention; [Figure 2] 1 is a block diagram that schematically illustrates details of a hearing aid device, according to an embodiment of the present invention; [Figure 3] 3 is a flow chart that schematically illustrates a method for operation of a hearing aid device, according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0009] Overview Despite the need for directional hearing aids and the theoretical advantages of microphone arrays in this regard, in practice, the directional performance of hearing aids falls far short of that achieved by natural hearing. Generally, good directional hearing aids require a relatively large number of microphones spaced far apart and designed to be unobtrusive, while still allowing the user to easily aim the hearing aid's directional response toward a point of interest, such as a conversation partner in a noisy environment. Processing circuitry applies beamforming filters to the signals output by the microphones in response to incident acoustic waves to generate an audio output that emphasizes sounds striking the microphone array within an angular range surrounding the direction of interest while suppressing background noise. The audio output should reproduce as natural a hearing experience as possible while minimizing distracting artifacts.
[0010] One of these artifacts is the user's own voice. When the user's mouth is close to the microphone, processing circuitry captures and amplifies the user's voice along with sounds received from the direction of interest in the environment. The user perceives the amplified sound of their own voice as unnatural and unpleasant. In some cases, amplifying the user's voice can mask sounds from the environment that the user wants to hear.
[0011] The embodiments of the invention described herein address this problem using a novel beamforming filter that emphasizes sounds striking an array of microphones within a selected angular range while suppressing sounds spoken by the user. In the disclosed embodiment, an array of microphones mounted near the user's head outputs electrical signals in response to incoming acoustic waves incident on the microphones. Speakers mounted near the user's ears output acoustic waves in response to drive signals. Processing circuitry generates the drive signals by amplifying and filtering the electrical signals using the beamforming filter.
[0012] In some embodiments, the processing circuitry detects when a user is speaking and, in response to such detection, applies a beamforming filter to suppress sounds spoken by the user. To this end, for example, the processing circuitry may apply two different beamforming filters: one that suppresses sounds spoken by the user and one that does not. While both filters emphasize sounds within a selected angular range, the filter that does not suppress sounds spoken by the user typically provides not only better directionality in the beamforming pattern but also reduced white noise gain compared to a filter that suppresses the user's own voice. Therefore, use of the filter that does not suppress sounds spoken by the user is preferred unless the user is actually speaking. When the processor detects that the user is speaking, it applies a beamforming filter that suppresses the user's own voice. The processing circuitry may detect which beamforming filter to use at which time by calculating and comparing the relative power levels output by the beamforming filters.
[0013] In some embodiments, the microphone and speaker are mounted in a frame attached to the user's head. In the embodiment described below, the microphone and speaker are mounted in an eyeglass frame. To mitigate the influence of the user's own voice, the processing circuitry applies a beamforming filter that suppresses sounds emanating from positions below the eyeglass frame. Alternatively, the microphone and speaker can be mounted in other types of frames, such as virtual reality (VR) or augmented reality (AR) headsets, or other types of mounting configurations.
[0014] System Description FIG. 1 is a schematic diagram of a hearing aid device 20 integrated into an eyeglass frame 22, according to another embodiment of the present invention. An array of microphones 23, 24 is mounted at respective locations on the eyeglass frame 22 and outputs electrical signals in response to acoustic waves incident on the microphones. In the illustrated embodiment, microphone 23 is mounted on a front piece 30 of the frame 22, while microphone 24 is mounted on temples 32 connected to respective edges of the front piece 30. While the wide array of microphones 23, 24 illustrated in FIG. 1 is useful in some applications of the present invention, the signal processing and hearing aid principles described herein may alternatively be applied, mutatis mutandis, using fewer microphones. For example, these principles may be applied using an array of microphones 23 on the front piece 30, as well as in devices using other microphone mounting configurations that are not necessarily eyeglass-based.
[0015] Processing circuitry 26 is fixed within or connected to eyeglass frame 22 and coupled by electrical wiring 27, such as traces on a flexible printed circuit, to receive the electrical signals output from microphones 23, 24. While processing circuitry 26 is illustrated in FIG. 1 , for simplicity, some or all of the processing circuitry may alternatively be located in front piece 30 or in a unit externally connected to frame 22 at certain locations in temple 32. Processing circuitry 26 mixes the signals from the microphones and generates audio output with a particular directional response by, for example, applying a beamforming function to emphasize sounds occurring within a selected angular range and suppress background sounds occurring outside that range. Typically, though not necessarily, the directional response coincides with the angular orientation of frame 22.
[0016] Additionally, processing circuitry 26 detects when the user is speaking and applies a beamforming function that suppresses sounds spoken by the user while emphasizing sounds striking the microphone array within a selected angular range. These signal processing functions of processing circuitry 26 are described in further detail below. Alternatively, processing circuitry 26 may apply a beamforming filter that always suppresses sounds spoken by the user, regardless of whether the user is speaking or not.
[0017] The processing circuitry 26 may transmit audio output to the user's ears via any suitable type of interface and speaker. In the illustrated embodiment, the audio output is generated by a drive signal to drive one or more audio speakers 28 mounted in temples 32, typically near the user's ears. While only one speaker 28 is shown in each temple 32 in FIG. 1 , device 20 may instead include only a single speaker in one of the temples 32, or two or more speakers mounted in one or both temples 32. In the latter case, processing circuitry 26 may apply a beamforming function to the drive signal to direct acoustic waves from the speakers toward the user's ears. Alternatively, the drive signal may be transmitted to a speaker inserted in the ear, or may be transmitted via a wireless connection, for example as a magnetic signal, to a telecoil in a user's hearing aid (not shown) wearing an eyeglass frame.
[0018] Signal Processing Figure 2 is a block diagram that schematically illustrates details of the processing circuitry 26 in the hearing aid device 20 according to an embodiment of the present invention. The processing circuitry 26 can be implemented in a single integrated circuit chip, or alternatively, the functions of the processing circuitry 26 can be distributed across multiple chips that can be located inside or outside the eyeglass frame 22. Although one particular implementation is illustrated in Figure 2, the processing circuitry 26 may alternatively comprise any suitable combination of analog and digital hardware circuits, along with appropriate interfaces for receiving electrical signals output by the microphones 23, 24 and outputting drive signals to the speaker 28.
[0019] In this embodiment, the microphones 23, 24 include integrated analog-to-digital converters and output digital audio signals to the processing circuitry 26. Alternatively, the processing circuitry 26 may include an analog-to-digital converter that converts the analog outputs of the microphones to digital form. The processing circuitry 26 typically includes suitable programmable logic components 40, such as a digital signal processor (DSP) or gate array, to perform the necessary filtering and mixing functions to generate and output drive signals for the speaker 28 in digital form.
[0020] These filtering and mixing functions typically involve the application of two beamforming filters 42, 43 having coefficients selected to create the desired directional response. Specifically, the coefficients of filter 42 are calculated to emphasize sounds impinging on frame 22 (and thus microphones 23, 24) within a selected angular range, while the coefficients of filter 43 are calculated to emphasize sounds impinging on frame 22 within this selected angular range while suppressing sounds spoken by the user of device 20 that originate from positions below frame 22. Details of filters that may be used for these purposes are described further below.
[0021] Alternatively or additionally, processing circuitry 26 may comprise a neural network (not shown) trained to determine and apply coefficients used in filters 42 and 43. Further alternatively or additionally, processing circuitry 26 comprises a microprocessor programmed in software or firmware to perform at least some of the functions described herein.
[0022] Processing circuitry 26 may apply any suitable beamforming function known in the art, either in the time domain or the frequency domain, in implementing filters 42, 43. Beamforming algorithms that may be used in this context are described, for example, in the above-mentioned PCT International Patent Application Publication No. 2017 / 158507 (especially pages 10-11) and U.S. Pat. No. 10,567,888 (especially paragraph 9).
[0023] In one embodiment, processing circuitry 26 applies a minimum variance distortion-free response (MVDR) beamforming algorithm in deriving the coefficients of beamforming filters 42 and 43. This type of algorithm is advantageous in achieving fine spatial resolution and distinguishing between sounds originating from a direction of interest and sounds originating from the user's own voice. The MVDR algorithm maximizes the signal-to-noise ratio (SNR) of the audio output by minimizing the average energy (while keeping the target distortion small). This algorithm can be implemented in frequency space by calculating a complex weight vector F(ω) of the output signal from each microphone at each frequency, expressed by the following equation:
[0024]
number
[0025] In an alternative embodiment, processing circuitry 26 applies a linearly constrained minimum variance (LCMV) algorithm in deriving the coefficients of filters 42 and 43. LCMV beamforming forces the filters to pass signals from the desired direction with a specified gain and phase delay while minimizing power from interfering signals and noise from all other directions. An additional constraint is imposed on computational filter 43 to specifically null out output power coming from the direction of the user's mouth.
[0026] In some embodiments, processing circuitry 26 includes selection logic 44, which selects the beamforming filter to be applied at any given time. Selection logic 44 selects the beamforming filter based on whether the user is speaking. The selection decision may be based on the outputs of beamforming filters 42 and 43 themselves. To this end, processing circuitry 26 applies both beamforming filters 42 and 43 to the electrical signals output by microphones 23 and 24 and calculates the respective power levels P1 and P2 output by beamforming filters 42 and 43. Selection logic 44 compares the power levels, for example, by applying a threshold to the ratio of the power levels. As long as P1 / P2 is less than the threshold, selection logic 44 selects filter 42. When P1 / P2 exceeds the threshold, selection logic 44 selects filter 43. The threshold may be preset or may be adjusted adaptively and / or according to user preferences.
[0027] Audio output circuitry 46, including, for example, an appropriate codec and digital-to-analog converter, converts the digital drive signals output by filters 42 and 43 to analog form. Analog filter 48 further performs filtering and analog amplification functions to optimize the analog drive signal to speaker 28.
[0028] A control circuit 50, such as an embedded microcontroller, controls the programmable functions and parameters of the processing circuitry 26, which may include selection logic 44. A communications interface 52, such as a Bluetooth® or other wireless interface, allows a user and / or hearing professional to set and adjust these parameters as needed. A power circuit 54, such as a battery inserted in temple 32, provides power to the other components of the processing circuitry.
[0029] How it works 3 is a flow chart that schematically illustrates a method of operation of a hearing aid device for suppressing self-voice, according to an embodiment of the present invention. For concreteness and clarity, the method will be described below with reference to the components of device 20 described above. Alternatively, the principles of the method may be applied to other devices and systems equipped with appropriate microphone arrays and signal processing capabilities.
[0030] In preparation for applying the method in real time, beamforming filters 42 and 43 are calculated and loaded into device 20 in a filter calculation step 60. The filters may be predefined and loaded at the time of manufacture of device 20 based on standardized spatial response parameters. Alternatively or additionally, the filter coefficients may be optimized for the user, taking into account parameters such as the user's physiology, hearing impairments, and subjective preferences.
[0031] Initially, when device 20 is turned on, selection logic 44 selects filter 42 so that the acoustic signals from microphones 23, 24 are processed without auto-voice suppression in a standard beamforming step 62. Periodically, processing circuitry 26 applies both beamforming filters 42 and 43 to the electrical signals output by microphones 23, 24 and compares them to the respective power levels P1 and P2 output by beamforming filters 42 and 43 in a power comparison step 64. As long as the ratio P1 / P2 is less than an applicable threshold, selection logic 44 continues to apply filter 42.
[0032] In step 64, detecting that P1 / P2 exceeds a threshold indicates that the user of device 20 has started speaking. In this case, selection logic 44 switches to filter 43 in self-voice suppression step 66. Optionally, to avoid rapid switching between filter 42 and filter 43, which may be distracting to the user, selection logic 44 applies a time lag in applying and removing beamforming filters after detecting that the user has started or stopped speaking. To this end, for example, the selection logic may apply a smoothing filter to the power values.
[0033]
number
number
number
[0034] Additionally or alternatively, upon detecting that the user has started speaking (or subsequently stopped speaking), processing circuitry 26 may blend beamforming filters 42 and 43 during a transition period to avoid abrupt changes in the acoustic output to speaker 28. To this end, processing circuitry blends the filter coefficients using a blending factor β that varies gradually over time between 0 and 1. During the transition period, the vector F of filter coefficients applied by processing circuitry 26 is a linear combination of the coefficient vectors F and F of filters 42 and 43, respectively: F=(1−β)F+βF. The value of β may be calculated using the following algorithm, for example, in a series of incremental steps over time using small values γ and ε between 0 and 1: If P1 / P2>threshold, β=β+γ, otherwise β=β-ε If β>1, then β=1 If β<0, then β=0
[0035] 3, while processing circuitry 26 applies self-voice suppression using filter 43 in step 66, selection logic 44 continues to check the ratio P1 / P2 in a further power comparison step 68. As long as the ratio remains above the threshold, processing circuitry 26 continues to apply self-voice suppression in step 66. However, if the ratio falls below the threshold, selection logic 44 switches back to beamforming filter 42 in a self-voice suppression deactivation step 70. The same time lag and blending procedures described above can be applied to step 70, mutatis mutandis. Operation of device 20 continues in step 62 until a further change is detected.
[0036] Example Example 1. A system for hearing assistance includes an array of microphones mounted near the head of a user of the system and configured to output electrical signals in response to first acoustic waves incident on the microphones. A speaker is configured for mounting near the user's ear and configured to output second acoustic waves in response to a drive signal applied to the speaker. Processing circuitry is configured to generate the drive signal by amplifying and filtering the electrical signal using a beamforming filter that emphasizes a first sound that strikes the array of microphones within a selected angular range while suppressing a second sound spoken by the user.
[0037] Example 2. The system according to Example 1, wherein the processing circuitry is configured to detect that a user is speaking and, in response to detecting that the user is speaking, apply a beamforming filter to suppress the second sound.
[0038] Example 3. The system according to Example 2, wherein the beamforming filter that suppresses the second sound is a first beamforming filter, and the processing circuitry is configured to apply the second beamforming filter that emphasizes the first sound without suppressing the second sound unless the user is speaking, and to apply the first beamforming filter upon detecting that the user is speaking.
[0039] Example 4. The system according to example 3, wherein the processing circuitry is configured to detect when a user is speaking by applying both the first and second beamforming filters to the electrical signal and comparing the respective power levels output by the first and second beamforming filters.
[0040] Example 5. The system according to example 3 or 4, wherein the processing circuitry is configured to generate the drive signal by blending the first and second beamforming filters during a transition period upon detecting that a user has started or stopped speaking.
[0041] Example 6. The system according to any of Examples 2-5, wherein the processing circuitry is configured to apply a time lag in applying and removing the beamforming filter after detecting that a user has started or stopped speaking.
[0042] Example 7. The system according to any of Examples 1-6, wherein the beamforming filter comprises a minimum variance distortion-free response (MVDR) beamformer.
[0043] Example 8. The system according to any of Examples 1-6, wherein the beamforming filter comprises a linearly constrained minimum variance (LCMV) beamformer.
[0044] Example 9. A system according to any of Examples 1-8, comprising a frame configured for attachment to a head, the microphone and speaker being mounted at respective locations on the frame.
[0045] Example 10. The system according to example 9, wherein the frame comprises an eyeglass frame.
[0046] Example 11. A head-mounted device comprising a frame configured for mounting on a user's head. An array of microphones is mounted at respective locations on the frame and configured to output electrical signals in response to first acoustic waves incident on the microphones. A speaker is mounted on the frame and configured to output second acoustic waves in response to a drive signal applied to the speaker. Processing circuitry is configured to generate the drive signal by amplifying and filtering the electrical signal using a beamforming filter that emphasizes first sounds impinging on the array of microphones within a selected angular range while suppressing second sounds emanating from a location below the eyeglass frame.
[0047] Example 12. The device according to example 11, wherein the frame comprises an eyeglass frame.
[0048] Example 13. The device according to example 12, wherein the eyeglass frame comprises a front piece and temples connected to respective edges of the front piece, at least some of the microphones are attached to the front piece, and the speaker is attached to one of the temples.
[0049] The features of any of Examples 2 to 7 can be similarly applied to any of Examples 11 to 13.
[0050] Example 14. A method for hearing aid comprises mounting an array of microphones near a user's head, the microphones outputting electrical signals in response to first acoustic waves incident on the microphones. A speaker is mounted near the user's ear and outputs second acoustic waves in response to a drive signal applied to the speaker. The drive signal is generated by amplifying and filtering the electrical signals using a beamforming filter that emphasizes a first sound impinging on the array of microphones within a selected angular range while suppressing a second sound spoken by the user.
[0051] Any of the features of Examples 2 to 10, 12, and 13 can be applied to Example 14 as well.
[0052] The above-described embodiments are given by way of example, and the present invention is not limited to what has been particularly shown and described above. Rather, the scope of the present invention includes both combinations and subcombinations of the various features described above, as well as variations and modifications thereof not disclosed in the prior art that would occur to those skilled in the art upon reading the foregoing description.
Claims
1. 1. A system for a hearing aid, comprising: an array of microphones mounted near a head of a user of the system and configured to output electrical signals in response to a first acoustic wave incident on the microphones; a speaker configured for mounting near the user's ear and configured to output a second acoustic wave in response to a drive signal applied to the speaker; and processing circuitry configured to generate the drive signal by amplifying and filtering the electrical signal using a beamforming filter that emphasizes a first sound that strikes the microphone array within a selected angular range while suppressing a second sound spoken by the user.
2. 2. The system of claim 1, wherein the processing circuitry is configured to detect when the user is speaking and, in response to detecting when the user is speaking, apply the beamforming filter to suppress the second sound.
3. 3. The system of claim 2, wherein the beamforming filter that suppresses the second sound is a first beamforming filter, and the processing circuitry is configured to apply a second beamforming filter that emphasizes the first sound without suppressing the second sound unless the user is speaking, and to apply the first beamforming filter upon detecting that the user is speaking.
4. 4. The system of claim 3, wherein the processing circuitry is configured to detect when the user is speaking by applying both the first beamforming filter and the second beamforming filter to the electrical signal and comparing the respective power levels output by the first beamforming filter and the second beamforming filter.
5. 4. The system of claim 3, wherein the processing circuitry is configured to generate the drive signal by blending the first beamforming filter and the second beamforming filter during a transition period upon detecting that the user has started or stopped speaking.
6. 3. The system of claim 2, wherein the processing circuitry is configured to apply a time lag in applying and removing the beamforming filter after detecting that the user has started or stopped speaking.
7. The system of any one of claims 1 to 6, wherein the beamforming filter comprises a minimum variance distortion-free response (MVDR) beamformer.
8. The system of any preceding claim, wherein the beamforming filter comprises a linearly constrained minimum variance (LCMV) beamformer.
9. The system of any one of claims 1 to 6, further comprising a frame configured to be attached to the head, the microphone and the speaker being attached at respective positions on the frame.
10. The system of claim 9 , wherein the frame comprises an eyeglass frame.
11. A head-mounted device, a frame configured for attachment to a user's head; an array of microphones mounted at respective locations on the frame and configured to output electrical signals in response to first acoustic waves incident on the microphones; a speaker attached to the frame and configured to output a second acoustic wave in response to a drive signal applied to the speaker; and processing circuitry configured to generate the drive signal by amplifying and filtering the electrical signal using a beamforming filter that emphasizes a first sound impinging on the array of microphones within a selected angular range while suppressing a second sound originating from a position below an eyeglass frame.
12. The device of claim 11 , wherein the frame comprises an eyeglass frame.
13. 13. The device of claim 12, wherein the eyeglass frame comprises a front piece and temples connected to respective edges of the front piece, at least some of the microphones being mounted to the front piece and the speaker being mounted to one of the temples.
14. 14. The device of claim 11, wherein the processing circuitry is configured to detect a level of the second sound and to apply the beamforming filter to suppress the second sound in response to detecting that the user is speaking.
15. 15. The device of claim 14, wherein the beamforming filter that suppresses the second sound is a first beamforming filter, and the processing circuitry is configured to apply a second beamforming filter that emphasizes the first sound without suppressing the second sound unless the user is speaking, and to apply the first beamforming filter upon detecting that the user is speaking.
16. 1. A method for a hearing aid, comprising: mounting an array of microphones near a user's head, the array of microphones outputting electrical signals in response to first acoustic waves incident on the microphones; attaching a speaker near the user's ear, the speaker outputting a second acoustic wave in response to a drive signal applied to the speaker; generating the drive signal by amplifying and filtering the electrical signal using a beamforming filter that emphasizes a first sound that strikes the microphone array within a selected angular range while suppressing a second sound spoken by the user.
17. 17. The method of claim 16, wherein generating the drive signal includes detecting that the user is speaking and applying the beamforming filter to suppress the second sound in response to detecting that the user is speaking.
18. 18. The method of claim 17, wherein the beamforming filter that suppresses the second sound is a first beamforming filter, and generating the drive signal includes applying the first beamforming filter to filter the electrical signal upon detecting that the user is speaking using a second beamforming filter that emphasizes the first sound without suppressing the second sound unless the user is speaking.
19. 20. The method of claim 18, wherein detecting that the user is speaking includes applying both the first beamforming filter and the second beamforming filter to the electrical signal and comparing respective power levels output by the first beamforming filter and the second beamforming filter.
20. 20. The method of any one of claims 16 to 19, wherein generating the drive signal comprises applying a time lag in applying and removing the beamforming filter after detecting that the user has started or stopped speaking.