Method, device, headphones and computer program for actively suppressing the occlusion effect during the playback of audio signals

The method addresses the occlusion effect in audio devices by using external and internal microphones to estimate and filter voice signals, enhancing the natural perception of one's own voice and reducing ambient noise, thus improving user comfort and device acceptance.

EP4158901B1Active Publication Date: 2025-07-16RWTH AACHEN UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
EP2021729292
Authority / Receiving Office
EP · EP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-29
Filing Date
2021-05-27
Publication Date
2025-07-16
Estimated Expiration
2041-05-27

Smart Images

  • Figure IMGF0001
    Figure IMGF0001
  • Figure IMGF0002
    Figure IMGF0002
  • Figure IMGF0003
    Figure IMGF0003
Patent Text Reader

Abstract

In the method according to the invention for actively suppressing the occlusion effect during the playback of audio signals by means of headphones (10) or a hearing aid, a sound signal occurring from outside is captured (20) by means of at least one outer microphone (11) of the headphones or of the hearing aid. A voice signal is captured (21) by means of at least one additional microphone (12, 17). The dry component of the captured voice signal is estimated (22), the dry component of the captured voice signal being the component of the captured voice signal without reverberation caused by the surrounding space and without ambient noises. By means of a filter, a voice component is extracted from the outer sound captured using the at least one outer microphone. Filter coefficients of the filter are determined (23) on the basis of the estimated dry component of the captured voice signal, or the estimated dry component of the captured voice signal is filtered such that a voice component having spaciousness comparable to that of the voice component at the outer microphones is produced (23). The extracted or produced voice component is output (24) by means of a loudspeaker of the headphones or of the hearing aid.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method for actively suppressing the occlusion effect during the reproduction of audio signals with headphones or a hearing aid. The present invention further relates to a device for implementing the method. Furthermore, the invention relates to headphones configured to carry out a method according to the invention or having a device according to the invention, as well as to a computer program with instructions that cause a computer to execute the steps of the method.

[0002] The muffled and unnatural perception of one's own voice when wearing headphones, hearing aids, or headsets is perceived as disturbing by wearers of such devices. This effect, known as the occlusion effect, occurs when the ear canal of the wearer of such headphones or hearing aids is partially or completely closed by the device. The occlusion effect is therefore particularly pronounced with so-called in-ear devices, in which the headphones or hearing aid are inserted into the opening of the ear canal and rest against its inner wall. The muffled perception of one's own voice is due, on the one hand, to the fact that the high-frequency components of one's own voice, transmitted through airborne sound, are significantly attenuated due to the headphones or hearing aid closing the ear canal.On the other hand, the low-frequency components of one's own voice are also transmitted into the ear canal by structure-borne sound, in particular via sound transmission from the cartilage or bones of the head, and cannot escape from the ear canal or can only partially escape due to the closure, so that the low-frequency components are even amplified.

[0003] Methods for compensating for the occlusion effect by correcting the airborne and structure-borne sound components in quiet environments are known. These involve attenuating the structure-borne sound components via a feedback control loop based on a microphone signal that reflects sound signals from the ear canal and is recorded with an internal microphone. The airborne sound components are recorded by an external microphone, filtered, and reproduced via an internal loudspeaker to create an acoustically transparent perception of the sound signals arriving from outside.

[0004] However, the airborne sound component includes not only one's own voice but also ambient noise. Since current technical solutions have so far failed in environments with high levels of background noise, measures that enable the most natural perception of one's own voice even under such conditions are the subject of current research.

[0005] Furthermore, various in-ear headphones and headsets already feature a "sidetone" or "hear-through" function. With the "sidetone" technology, it is possible to hear your own voice, for example, during a telephone call conducted with such headphones or headsets. A microphone records a voice signal that allows clear speech reproduction, but spatial and binaural information is lost. The "hear-through" technology makes it possible to perceive your surroundings and, for example, to have a conversation without having to remove the headphones. One or more external microphones are used for each side of the headphones, which preserves spatial information from your own voice; however, the signal contains unwanted ambient noise.

[0006] EP 2 920 980 A1 discloses a system for improving the perception of one's own voice, comprising an ear-canal microphone and an external microphone. An estimate of the ambient sound and an estimate of one's own voice are obtained from the microphone signals, which are then added together with variable gain factors.

[0007] EP 3 213 527 B1 and US 2014 / 126 735 A1 specify systems for reducing the occlusion effect in ANC headphones and headsets.

[0008] A headset that initially operates in a "noise-canceling" mode and then switches to a "hear-through" mode as soon as a speech activity detector determines that the user is in a call is described in EP 3 188 495 A1. Similarly, EP 2 362 678 A1 describes a communication headset with a switching function between a transparency and a communication mode. Furthermore, US 10,034,092 B1 describes digital audio signal processing techniques used to provide an acoustic transparency function in a headset. This involves considering multiple acoustic paths for different users or artificial heads to determine a transparency filter that delivers good results for most users.

[0009] It is an object of the invention to provide a method and a device for actively suppressing the occlusion effect when reproducing audio signals with a headphone or hearing aid in environments with a high noise level, as well as a corresponding headphone and a computer program for carrying out the method.

[0010] This object is achieved by a method having the features of claim 1, a corresponding device according to claim 8, and a corresponding headset according to claim 10. Preferred embodiments of the invention are the subject of the dependent claims.

[0011] In the method according to the invention for actively suppressing the occlusion effect during the reproduction of audio signals with headphones or a hearing aid, external sound in the form of an externally occurring sound signal is detected using at least one external microphone of the headphones or hearing aid. A voice signal is detected using at least one additional microphone. The dry portion of the detected voice signal is estimated, with the dry portion of the detected voice signal being the portion of the detected voice signal without reverberation or ambient noise caused by the surrounding space.A voice component is extracted from the external sound captured by the at least one external microphone using a filter, with filter coefficients of the filter being determined based on the estimated dry component of the captured voice signal. Or the estimated dry component of the captured voice signal is filtered to generate a voice component that has a comparable spatiality to the voice component at the external microphones. The extracted or generated voice component is output via a speaker of the headphones or hearing aid.

[0012] This allows for a more natural and undisturbed perception of one's own voice. This leads to a significant increase in comfort, which not only leads to increased acceptance of such headphones and hearing aids but also opens up the possibility of novel user experiences when using these products.

[0013] According to one embodiment of the invention, the voice signal is captured with at least one microphone or microphone array directed at the user's mouth and / or an internal microphone of the headphones or hearing aid. Both such a mouth microphone and the internal microphones offer a very good signal-to-noise ratio, either due to their directional characteristics, their spatial proximity, or shielding. In particular, a monaural dry component is estimated from the captured voice signal, and based on this, binaural voice signals are extracted from the signals of at least two external microphones of a left and right headphone or left and right hearing aid. Alternatively, the estimated monaural dry voice component can also be filtered such that binaural voice signals are generated with a comparable spatiality to the voice component at the external microphones.

[0014] This combines the advantages of the "sidetone" and "hearthrough" techniques, so that spatial and binaural information is preserved during the reproduction of the sound signals while simultaneously suppressing unwanted ambient noise.

[0015] According to one embodiment of the invention, the binaural voice signals are filtered before being output via a loudspeaker for a left and right pair of headphones or a left and right hearing aid.

[0016] Advantageously, the dry voice component at the outer microphone is estimated by filtering with the respective relative impulse response between the mouth microphone or microphone array and the outer microphone and then averaging.

[0017] Furthermore, the filter for extracting or generating the voice component based on the detected external sound and the estimated dry voice is preferably a Wiener filter, an adaptive filter or a filter that simulates a room impulse response.

[0018] According to a further embodiment of the invention, the estimated dry portion of the detected voice signal and the extracted or generated voice portion are linearly weighted and then added.

[0019] Accordingly, a device according to the invention for actively suppressing the occlusion effect during the reproduction of audio signals via a loudspeaker of a headset or hearing aid provided with at least one external microphone comprises at least one additional microphone for detecting a user's voice signal; a digital signal processor configured to estimate the dry portion of a voice signal detected by the at least one additional microphone, wherein the dry portion of the detected voice signal is the portion of the detected voice signal without reverberation or ambient noise caused by the surrounding space; to extract the voice portion from the external sound detected by the at least one external microphone using a filter, wherein filter coefficients of the filter are determined based on the estimated dry portion of the detected voice signal, or to filter the estimated dry portion of the detected voice signal such that a voice portion is generated that has a comparable spatiality to the voice portion at the external microphones; and to output the extracted or generated voice portion via the loudspeaker.

[0020] According to one embodiment of the invention, a digital filter is additionally provided to which the extracted or generated voice component is fed before being output via the loudspeaker.

[0021] The invention also relates to a headset which is configured to carry out the method according to the invention or has a device according to the invention.

[0022] Further features of the present invention will become apparent from the following description and claims in conjunction with the figures. Fig. 1 schematically shows an in-ear headphone with closure of the user's ear canal; Fig. 2 shows a flowchart of the inventive method for actively suppressing the occlusion effect; Fig. 3 shows a block diagram of a first embodiment of a headphone according to the invention; Fig. 4 shows a block diagram of a second embodiment of a headphone according to the invention; and Fig. 5 schematically shows a communication headset for carrying out the inventive method.

[0023] To better understand the principles of the present invention, embodiments of the invention are explained in more detail below with reference to the figures. It is understood that the invention is not limited to these embodiments and that the described features can also be combined or modified without departing from the scope of the invention as defined in the claims.

[0024] The method according to the invention can be used, for example, to reduce the occlusion effect in in-ear headphones, as in Figure 1shown schematically. The in-ear headphones 10 are located on the ear of a user, with an ear insert 14 of the in-ear headphones being inserted into the external auditory canal 15 to hold it in place. The ear insert seals the auditory canal to a certain extent, depending on the individual fit in the auditory canal and the material used. This results in external noise being at least partially shielded, so that this noise only reaches the user's eardrum 16 at a reduced level. This means that music playback via the headphones or the reproduction of a caller's voice during a telephone call using the headphones is less disrupted. On the other hand, the ear insert also muffles the user's voice, thus leading to the occlusion effect already mentioned above.

[0025] A noise signal x(t) arriving at the headphones from the environment, which may include the user's voice in particular, but also ambient noise, is recorded by an external microphone 11 directed away from the ear canal toward the headphones' surroundings. Furthermore, the in-ear headphones 10 have an internal microphone 12 directed toward the ear canal 15 in the direction of the user's ear canal or eardrum, and a loudspeaker 13 located near the internal microphone 12. A compensation signal u(t) can be output by means of the loudspeaker 13, which suppresses the occlusion effect as comprehensively as possible, or at least reduces it, so that the user ideally has the impression that they are not wearing headphones.

[0026] With the help of the outer microphone 11, the airborne sound components of the noise signal are recorded and a compensation signal is generated for this purpose. In addition, the inner microphone 12 records a residual signal e(t) after superimposing the compensation signal u(t) filtered by the secondary path S(s) with the noise signal x(t) filtered by the primary path P(s). In particular, this makes it possible to also record a structure-borne sound component and take it into account in the compensation signal. The acoustic primary path P(s) describes the transfer function for the acoustic transmission from the outer microphone 11 to the inner microphone 12 and can be measured, for example, using an external loudspeaker structure. The acoustic secondary path S(s) describes the transfer function from the internal loudspeaker 13 to the inner microphone 12 and can be measured using this loudspeaker and inner microphone.

[0027] The in-ear headphones shown here have only one external microphone, but multiple microphones arranged in a microphone array can also be used. Furthermore, the occlusion effect can also occur with other headphones, such as over-the-ear headphones with circumaural ear cushions that seal off the ear canal through a closed design, or with hearing aids. It can be compensated for as described below.

[0028] Figure 2 shows schematically the basic concept of a method for actively suppressing the occlusion effect, as it occurs, for example, in the reproduction of audio signals with in-ear headphones made of Figure 1can be carried out. In a first step 20, the external sound is recorded using at least one external microphone 11 of the headphones or hearing aid. This recorded external sound also includes an acoustic voice component that originates from a voice output of the user wearing the headphones. In a subsequent step 21, a voice signal that corresponds to the voice output of the user is recorded using at least one additional microphone, for example, a microphone of a communication headset directed towards the user's mouth, hereinafter also referred to as a mouth microphone for short.

[0029] Subsequently, in step 22, the dry portion of the voice signal captured with the additional microphone is estimated. As is known to those skilled in the art, a dry-recorded audio signal is understood to be a pure sound signal as it originally existed at the time of generation, i.e., without any reverberation due to reflections of the generated sound waves in a closed room or in a naturally confined area, and free from ambient acoustic disturbances. In this step, the voice signal is estimated as it was directly generated by the user's vocal tract.

[0030] Based on the estimated dry portion of the captured voice signal, the binaural voice signal contained in the microphone signal of the respective external microphone is estimated in the subsequent step 23 and extracted using a filter. The filter coefficients of the filter are determined based on the estimated dry portion of the captured voice signal. Alternatively, the estimated dry voice signal can also be filtered so that it has a comparable spatiality to the voice portion at the external microphones. The extracted or generated binaural voice portion is then output in step 24 via the corresponding speaker of the headphones or hearing aid. The signal is first adjusted using a feedforward filter to ensure the most acoustically transparent reproduction of the voice signals possible.

[0031] Figure 3shows a block diagram of a device according to the invention, which can be implemented in particular in headphones, but also in a hearing aid. Although headphones or hearing aids usually have sound transducers for both ears of the user, in the figure only the conceptual structure with reference to one ear is shown for the sake of clarity. Likewise, digital signal processing requires analog-to-digital converters for digitizing the sound signals captured by the microphones and digital-to-analog converters for converting the processed signals for output via the loudspeaker, but for the sake of simplicity these are not shown in the figure. Due to the digital signal processing, the signals are considered below in the time domain with a discrete time index n; the index z stands for a frequency domain representation of the time-discrete signals and filters.

[0032] As already mentioned in connection with Figure 1 As mentioned above, in addition to the loudspeaker 13, an outer microphone 11 and an inner microphone 12 are provided, each of which can be arranged in an earphone or a headphone cup. The outer microphone 11, which delivers the signal x(n), is mounted on the outside of the headphones. The loudspeaker 13 and the inner microphone 12, on the other hand, are arranged inside the headphones and directed towards the eardrum.

[0033] Furthermore, a mouth microphone 17 is provided. This can, for example, be part of a communications headset and be attached to a pivoting bracket so that it can be arranged in front of the user's mouth and aimed at the mouth. However, a microphone array consisting of several microphones can also be provided, which is arranged on the outside of the headphones or hearing aid and is aimed at the mouth, for example, using a beam-forming method. In addition to the primary path P(z), which designates the acoustic transmission from the outer microphone to the inner microphone, and the secondary path S(z) for the transmission from the loudspeaker to the inner microphone, the transmission path B(z) between the mouth microphone and the external reference microphone is also noted here. This path, for example, in a communications headset is given by the predefined position of the swivel microphone in front of the mouth relative to the position of the outer microphone.The transmission paths also include the influence of other components, such as the analog-to-digital converters and digital-to-analog converters (not shown).

[0034] If the user of the headphones or hearing aid outputs a voice output, a voice signal xv (n) corresponding to this voice output is recorded by the external microphone 11. The recorded voice signal xv (n) includes the room impulse response, which contains all relevant information about the current acoustic room properties. In addition to this voice signal, however, the external microphone 11 also records an interference signal xa (n) caused by ambient noise, since the external microphone 11 is attached to the outside of the headphones. The audio signal x(n) consisting of these two signal components is then processed as described below based on an estimate of the dry voice signal in order to achieve acoustic transparency for the user's own voice by outputting the processed speech signals u(n) via the loudspeaker 13 of the headphones or hearing aid.Here, the voice signal entering the headphones from the outside is transmitted both via the primary path P(z) from the external to the internal microphone and via the secondary path S(z) in the form of the signal that is actively output via loudspeaker 13. In this way, the missing airborne sound component of the listener's own voice is added back in. Acoustic interference between the sound signals transmitted via these two paths then leads to acoustic transparency for the voice signal.

[0035] In the illustrated embodiment, both the voice signal v(n) measured by the mouth microphone 17 and the error signal e(n) of the inner microphone are fed to an estimation unit 30 in which the pure, dry voice signal ṽ ( n ) as it is produced in the vocal tract and would be present without reverberation caused by the surrounding space and free from ambient acoustic disturbances. Based on this monaural estimate v ( n ), a second estimation unit 31 extracts the binaural voice signal from the signal captured by the outer microphone of the left or right headphones. Alternatively, the estimated dry voice signal can also be filtered so that it has a comparable spatiality to the voice component at the outer microphones. The binaural voice signals xv(n) are then filtered by a digital filter unit 32 with a negated transfer function and finally fed as a loudspeaker signal u(n) to a sound transducer for output via the headphones. The digital filter unit 32 is designed in particular as a feed-forward filter.

[0036] For estimation of the dry voice signal ṽ ( n) in the estimation unit 30, the voice signal v(n) can be measured by a mouth microphone 17 and then used as a speech reference. The estimation of the dry voice component at the outer microphone can be carried out, for example, by filtering the additional signals with the respective relative impulse response between the additional microphone and the outer microphone and subsequently averaging them. For this purpose, the mouth microphone signal v(n) can be estimated, for example, by B ( n ) of the relative transmission path B(z) between the mouth microphone and the external microphones. The voice signal v(n) is treated as a monaural source, but is then used for both headphones or ears.

[0037] Likewise, an error signal e(n) can be detected by the inner microphone 12, which is also used for estimating the dry voice signal ṽ ( n) and can be fed to the estimation unit 30 for this purpose. Since the ear is closed by the headphones, the user's own voice couples strongly into the ear canal via the body, so that information about the user's own voice can also be obtained using the microphone signals of the internal microphone. The error signal e(n) comprises an error component ev(n) based on the voice signal and a further error component eb(n) which is based on further disturbances such as impact sound transmitted into the ear canal via the user's body. Separate error signals are generated for each of the two headphones or ears. These can differ, for example, if the fit of the headphones differs. However, the separate error signals can also be averaged if necessary in order to obtain a monaural signal again.

[0038] The signals from the mouth microphone and the inner microphones can be adjusted, for example, through digital filtering and then combined by averaging to further improve the signal-to-noise ratio. It is important to note that the signals played through the headphone speakers are each convolved with an estimate of the respective secondary path and subtracted from the respective inner microphone signal to prevent signal feedback.

[0039] Since the internal microphones mainly record the structure-borne sound component of one's own voice, which does not allow for the breakdown of fricatives, for example, a bandwidth extension of the signals from the internal microphones is also conceivable.

[0040] Since both the oral microphone and the inner microphones offer a good signal-to-noise ratio, it can also be planned to perform an estimation based solely on the signal measured with the oral microphone or the signal from the inner microphone, rather than an estimate based on a combination of signals from both microphones. After all, under particularly favorable conditions, these can already provide a dry reference of the voice without the need for additional estimation.

[0041] In the second estimation unit 31, the binaural voice signal is estimated by extracting the binaural voice from the external microphone signals, which are distorted by ambient noise, based on the estimation of the dry voice, or by generating a voice signal that exhibits a comparable spatiality to the voice component at the external microphones. It is important that the processing has a short and constant delay so that the delay can be taken into account for the calculation of the forward filter W(z).

[0042] For this purpose, a Wiener filter or other noise suppression algorithms can be used. With the Wiener filter, the magnitude spectra of the recorded signals are evaluated in order to calculate a filter with an estimate of the speech signal and an estimate of the existing noise signal, which can be used to optimally extract the speech signal. For example, the magnitude spectrum of the mouth microphone can be combined with the magnitude spectrum of the inner microphones to estimate the magnitude spectrum of the dry voice signal and then extract the speech component from the signals of the outer microphones. The transfer function B(z) can be used to estimate how the dry voice from the mouth microphone arrives at the outer microphone in order to then compensate for the propagation times of the direct sound.

[0043] Since the transfer function B(z) of a communication headset is very similar even for different people, the impulse response can be determined, for example, by a series of measurements for a specific headset and then subsequently used for applications of headsets of this design.

[0044] One possibility is a Wiener filter in a "filter bank equalizer" structure. This structure requires a prototype low-pass filter with a constant group delay. The spectral weights of the Wiener filter require an estimate of the wanted and the noise signal. The estimate of the dry voice can be used to estimate the wanted signal component.

[0045] Alternatively, an adaptive filter a(n) can be used to estimate the binaural voice. Assuming that the external microphone signal x(n)=xa (n)+xv (n) consists of ambient noise xa (n) and a voice component xv (n) that is coherent with the estimate v ( n ) of the dry voice, an adaptive filter can be used to convert the voice component xv (n) into x(n) based on v ( n ) to reproduce. With the output x v ^ n of the adaptive filter, a rule for adapting the adaptive filter can be found based on the following cost function: C v = E x n − x v ^ n ∧ 2 , mit x v ^ n = a n * v ^ n .

[0046] Furthermore, the estimation unit 31 can analyze the acoustic influence of the room on the own voice and, based thereon, select or design a filter that can be applied to the estimated dry voice signal to generate a voice signal that has a comparable spatiality to the voice component at the external microphones.

[0047] The forward filter W ( z ) can be determined, for example, by solving the Wiener-Hopf equation w = Ψ s ′ s ′ − 1 φ s ′ p − h For this purpose, one or more measurements of the primary path P ( z ) and the secondary path S( z) is required. These measurements can be performed, for example, on an artificial head or on volunteers. It is important that any delay caused by the processing in the branch between the respective external microphone and the headphone speaker is taken into account by the secondary path used for the calculation of the forward filter. For example, if the signal x(n) If the signal from the speaker or any derived signals subsequently played through the loudspeaker are delayed in estimating the binaural voice, this delay must be accounted for by the secondary path. This is indicated by an apostrophe in the Wiener-Hopf equation above.

[0048] The desired transmission behavior from the outer to the inner microphone, which is usually characterized by a flat magnitude response for the natural perception of one’s own voice, is achieved by H ( z) in the z-range or by the impulse response h(n) and is also required for the Wiener-Hopf equation.

[0049] Figure 4 shows a block diagram of another device according to the invention. In addition to the units of the device according to the invention from Figure 3 A control unit 40 is provided here to control two weighting units 41 and 42. Since in the case shown v ( n) and xv (n) are coherent, i.e., they are not shifted or at least not noticeably shifted against each other in the time domain, both signals can be weighted with linear weighting factors α and 1-α, with 0≤α≤1, and then added together. The weighting units 41 and 42 thus enable the user to personalize the mix of dry and binaural voice. The user can thus decide and adjust how they perceive their voice, for example, what ratio the volume of the reverberation should have to the volume of their own voice. The control can also be automatic.

[0050] As described above, one consequence of the occlusion effect is that the low-frequency components of one's own voice are amplified. To compensate for this, the internal microphone signal can be additionally filtered with a feedback control to reduce the low-frequency components of one's own voice. This makes the perception of one's own voice appear even more natural when wearing headphones.

[0051] The estimation units 30 and 31 as well as the control unit 40 can be part of a processor unit that has one or more digital signal processors, but can also include other types of processors or combinations thereof. Furthermore, the filter coefficients of the digital filter 32 can be adjusted by the digital signal processor. The filter can be implemented as a time-invariant filter that is calculated once, loaded into the headset's firmware, and used in this form without any changes being made during runtime. An adaptive filter that changes during runtime and adapts to the current circumstances can also be used.

[0052] The device according to the invention is preferably fully integrated into a headset, since latency is very low due to the transmission of one's own voice through structure-borne sound. The mouth microphone can also be part of the headset, for example, in a so-called communications headset, attached to a headband that is placed in front of the mouth, or as a microphone array with directional characteristics integrated into a headset shell. Likewise, a separate microphone can also serve as a mouth microphone. In principle, however, parts of the device can also be part of an external device, such as a smartphone.

[0053] Figure 5shows schematically the use of a communications headset in which the method according to the invention can be carried out and which has the device described above for this purpose. A headphone 10 is provided for each of the user's ears, each of which has an external microphone 11, an internal microphone 12 and a loudspeaker 13 integrated therein. Furthermore, a mouth microphone 17 is provided which is attached to a pivoting bracket. Furthermore, a processor unit 50 is arranged in one of the two headphones, through which the estimation units and, if applicable, the control unit 40 are implemented. The individual components are connected to the processor unit 50, but this is not shown in the figure for the sake of clarity.

[0054] The invention can be used to suppress the occlusion effect when playing audio signals with any headphones or hearing aids, such as telephony or communication with communication headsets / hearables, so-called in-ear monitoring for checking one's own voice during a live performance, augmented / virtual reality applications or use with hearing aids. List of reference symbols

[0055] 10Single earphone, single hearing aid 11External microphone 12Internal microphone 13Speaker 14Earpiece 15Ear canal, 16Eardrum 17Oral microphone 20 - 24Procedure steps 30First estimation unit 31Second estimation unit 32Digital forward filter 40Control unit 41, 42Weighting unit 50Processor unit

Claims

1. A method for actively suppressing the occlusion effect during the playback of audio signals by means of headphones (10) or a hearing aid, in which - with at least one outer microphone (11) of the headphones or the hearing aid, external sound is captured (20) in the form of a sound signal occurring from the outside; - a voice signal is captured (21) with at least one additional microphone (12, 17); - the dry component of the captured voice signal is estimated (22), wherein the dry component of the captured voice signal is the component of the captured voice signal which has neither a reverberation caused by the surrounding space nor ambient noises; - a voice component is extracted by a filter from the external sound captured with the at least one outer microphone, wherein filter coefficients of the filter are determined (23) based on the estimated dry component of the captured voice signal, or the acoustic influence of the space on the own voice is analyzed based on the dry component of the captured voice signal and the external sound captured with the at least one outer microphone and based thereon the estimated dry component of the captured voice signal is filtered such that a voice component is produced (23) which has a comparable spatiality to the voice component at the at least one outer microphone; and - the extracted or generated voice component is output (24) via a loudspeaker of the headphones or hearing aid.

2. The method according to claim 1, wherein the at least one additional microphone with which the voice signal is captured (21) comprises at least one microphone or microphone array (17) directed towards the mouth of the user and / or an inner microphone of the headphones or hearing aid.

3. The method of claim 2, wherein a monaural dry component is estimated from the captured voice signal and based thereon binaural voice signals are extracted from the signals of at least two outer microphones of left and right headphones or left and right hearing aids, or the estimated monaural dry voice component is filtered to generate binaural voice signals with a comparable spatiality to the voice component at the outer microphones.

4. The method according to claim 3, wherein the binaural voice signals are filtered for left and right headphones or left and right hearing aids prior to being respectively output via a loudspeaker (13).

5. The method according to any one of claims 2 to 4, wherein the dry component of the captured voice signal is estimated (22) by performing a filtering with the respective relative impulse response between the at least one mouth microphone or microphone array (17) and the outer microphone (11) and subsequent averaging.

6. The method according to any one of the preceding claims, wherein the filter for extracting or generating the voice component based on the captured external sound and the estimated dry voice is a Wiener filter, an adaptive filter or a filter which simulates a room impulse response.

7. The method according to any one of the preceding claims, wherein the estimated dry component of the captured voice signal and the extracted or generated voice component are linearly weighted and added and then output via a loudspeaker of the headphones or hearing aid.

8. Device for actively suppressing the occlusion effect during the playback of audio signals by means of a loudspeaker (13) in a headphone (10) or hearing aid provided with at least one outer microphone (11), comprising - at least one additional microphone (17) for capturing a voice signal from a user; - a digital signal processor (50) which is arranged to - estimate the dry component of a voice signal captured with the at least one additional microphone (17), wherein the dry component of the captured voice signal is the component of the captured voice signal which has neither a reverberation caused by the surrounding space nor ambient noises; - extract from the external sound captured by the at least one outer microphone (11) the voice component using a filter, wherein filter coefficients of the filter are determined based on the estimated dry component of the captured voice signal, or analyze the acoustic influence of the space on the own voice based on the dry component of the captured voice signal and the external sound captured with the at least one outer microphone and based thereon filter the estimated dry component of the captured voice signal to produce a voice component which has comparable spatiality to the voice component at the at least one outer microphone; and - output the extracted or generated voice component via the loudspeaker (13).

9. The device according to claim 8, wherein a digital filter (32) is additionally provided, to which the extracted or generated voice component is supplied before it is output via the loudspeaker (13).

10. Headphones (10) comprising a device according to claim 8 or 9.

Citation Information

Patent Citations

  • Own voice shaping in a hearing instrument

    EP2920980A1

  • Information-processing device, information processing method, and program

    EP3163902A1

  • Audio output device, audio output method, program, and audio system

    EP3413590A1

  • Hearing Aid

    US20080192971A1

  • Reducing Occlusion Effect in ANR Headphones

    US20140126735A1