Method, apparatus, earphone and computer program for actively suppressing occlusion effect when reproducing an audio signal
By combining external and internal microphones for signal processing, spatial speech signal output is estimated and generated, solving the occlusion effect problem when wearing headphones or hearing aids in high-noise environments, and achieving natural speech perception and improved comfort.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-27
- Publication Date
- 2026-03-27
AI Technical Summary
In high-noise environments, existing technologies struggle to effectively suppress the occlusion effect when wearing headphones or hearing aids, resulting in unnatural speech perception.
By using at least one external microphone to collect external sound, extracting speech components using filters, and combining signal processing from an internal microphone and a mouth microphone, the interference component is estimated to generate a spatial speech signal output while suppressing environmental noise.
It enables natural perception of one's own voice in high-noise environments, improving the comfort and user acceptance of headphones or hearing aids, while preserving spatial and binaural information.
Smart Images

Figure CN115398934B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a method for actively suppressing occlusion effect when reproducing an audio signal with a headphone or a hearing aid. The invention also relates to a device for carrying out the method. Furthermore, the invention relates to a headphone which is arranged for carrying out the method according to the invention or having the device according to the invention, and to a computer program having instructions which cause a computer to carry out the steps of the method. BACKGROUND
[0002] The dull and unnatural perception of one's own voice when wearing a headphone, a hearing aid or a headset is perceived as disturbing by the wearer of these devices. This effect occurs when the ear canal of the wearer of such a headphone or hearing aid is partially or completely closed by the device, which is referred to as occlusion effect or closed effect. The occlusion effect is therefore also particularly pronounced in so-called In-Ear devices, in which the headphone or hearing aid is inserted into the opening area of the ear canal and rests against the inner wall of the ear canal. The dull perception of one's own voice is based on the fact, on the one hand, that the high-frequency components of one's own voice, which are transmitted by air sound, are significantly attenuated by the headphone or hearing aid closing the ear canal. On the other hand, in particular the low-frequency components of one's own voice are also transmitted into the ear canal by body sound, in particular by the sound transmission of the cartilage or bone of the head, and cannot or only partially escape from the ear canal due to the closure, so that an amplification of the low-frequency components even occurs.
[0003] Methods are known for compensating the occlusion effect by correcting the air sound components and the body sound components in a quiet environment. These methods include the attenuation of the body sound components by a feedback control loop based on a microphone signal, which reflects the sound signal from the ear canal and is recorded with an internal microphone. The air sound components are recorded by an external microphone, filtered and reproduced by an internal loudspeaker in order to create the feeling of acoustic transparency from the outside.
[0004] However, the air sound components not only include one's own voice, but also disturbing sounds from the surrounding environment. Since the current technical solutions do not work in environments with high noise levels, measures to enable the perception of one's own voice as natural as possible even in such conditions are the subject of current research.
[0005] Furthermore, different in-ear headphones and headsets already have a "sidetone" or "hear-through" function. In the "sidetone" approach, one's own voice can be heard, for example, during a telephone call made with such a headphone or headset. For this purpose, a microphone is used to record the voice signal, which, although it enables clear voice reproduction, loses spatial and binaural information in the process. The "hear-through" approach enables perception of the surroundings, for example, to be able to talk without taking off the headphones. For this purpose, one or more external microphones are used on each headphone side, whereby the spatial information of one's own voice is maintained, however, in this case the signal contains unwanted ambient noise.
[0006] In EP 3 188 495 A1, an earphone is described which initially works in a "noise cancellation" mode and then switches to a "hear-through" mode once a voice activity recognition means determines that the user is in a call. Similarly, EP 2 362 678 A1 also describes a communication headset with a switching function between a transparent mode and a communication mode.
[0007] Furthermore, in US 10,034,092 B1, a digital audio signal processing technique is described for providing an acoustic transparency function in headphones. Here, a plurality of acoustic paths are considered for different users or artificial heads in order to determine a transparency filter which provides good results for most users. SUMMARY
[0008] The technical problem addressed by the present application is therefore to provide a method and a device for actively suppressing occlusion effects when reproducing audio signals with headphones or hearing aids in an environment with a high level of disturbing noise, and to provide a corresponding earphone and a computer program for carrying out the method.
[0009] The above technical problem is solved by the method having the features of claim 1, the corresponding device according to claim 8, the corresponding earphone according to claim 10 and the computer program according to claim 11. Preferred design solutions of the present application are the subject of the dependent claims.
[0010] In the method for actively suppressing occlusion effects when reproducing an audio signal with headphones or a hearing aid according to the present application, external sounds in the form of sound signals from external occurrences are captured with at least one external microphone of the headphones or the hearing aid. A speech signal is captured with at least one additional microphone. A dry component of the captured speech signal is estimated, wherein the dry component of the captured speech signal is the part of the captured speech signal without reverberation or ambient noise caused by the surrounding space. A speech component is extracted from the external sounds captured with the at least one external microphone by a filter, wherein filter coefficients of the filter are determined based on the estimated dry component of the captured speech signal, or the estimated dry component of the captured speech signal is filtered such that a speech component is generated which has a similar spatiality as the speech component at the external microphone. The extracted or generated speech component is output by a loudspeaker of the headphones or the hearing aid.
[0011] In this way, the own speech is perceived more naturally and unobtrusively. This leads to a significant increase in comfort, which not only leads to an increased acceptance of such headphones or hearing aids, but also opens up new possibilities for user experiences when using these products.
[0012] According to embodiments of the present application, the speech signal is captured with at least one microphone or microphone array directed at the mouth of the user and / or an internal microphone of the headphones or the hearing aid. Both this mouth microphone and the internal microphone provide a very good signal-to-noise ratio, either because of their directional characteristics or because of their spatial proximity, or because of the shielding.
[0013] In particular, a monaural dry component is estimated from the captured speech signal, wherein a binaural speech signal is extracted from the signals of at least two external microphones of the left and right headphones or the left and right hearing aid based thereon. Alternatively, the estimated monaural dry speech component can also be filtered such that a binaural speech signal is generated which has a similar spatiality as the speech component at the external microphones.
[0014] This combines the advantages of the "sidetone" and "in-ear" methods, so that spatial and binaural information is preserved when reproducing the sound signal, while at the same time undesired ambient noise is suppressed.
[0015] According to embodiments of the present application, the binaural speech signal is filtered by the loudspeakers of the left and right headphones or the left and right hearing aid before the respective output.
[0016] Advantageously, the dry speech component at the external microphone is estimated by filtering with the respective relative impulse response between the mouth microphone or microphone array and the external microphone and subsequent averaging.
[0017] Furthermore, the filter for extracting or generating a speech component on the basis of the captured external sound and the estimated dry speech is preferably a Wiener filter, an adaptive filter or a filter simulating a spatial impulse response.
[0018] According to a further embodiment of the present application, the estimated dry component of the captured speech signal and the extracted or generated speech component are linearly weighted and then added.
[0019] Accordingly, the device for actively suppressing occlusion effects when reproducing an audio signal through a loudspeaker of a headphone or a hearing aid provided with at least one external microphone according to the present application comprises,
[0020] - at least one additional microphone for capturing a speech signal of a user;
[0021] - a digital signal processor, which is set up for:
[0022] - estimating a dry component of the speech signal captured with the at least one additional microphone, wherein the dry component of the captured speech signal is the part of the captured speech signal which is free of reverberation or ambient noise caused by the surrounding space;
[0023] - extracting a speech component from the external sound captured with the at least one external microphone using a filter, wherein filter coefficients of the filter are determined on the basis of the estimated dry component of the captured speech signal or the estimated dry component of the captured speech signal is filtered such that a speech component is generated which has a similar spatiality as the speech component at the external microphone; and
[0024] - outputting the extracted or generated speech component through the loudspeaker.
[0025] According to an embodiment of the present application, a digital filter is additionally provided, to which the extracted or generated speech component is fed before being output through the loudspeaker.
[0026] The present application also relates to a headphone which is set up for carrying out the method according to the present application or having the device according to the present application and to a computer program having instructions which cause a computer to carry out the steps of the method according to the present application. BRIEF DESCRIPTION OF DRAWINGS
[0027] Further features of the present application will become apparent from the following specification and drawings.
[0028] Figure 1 An in-ear headphone with an occluded ear canal of a user is schematically shown;
[0029] Figure 2 A flow chart of a method for actively suppressing occlusion effects according to the present application is shown;
[0030] Figure 3 A block diagram of a first embodiment of a headphone according to the present application is shown;
[0031] Figure 4 A block diagram of a second embodiment of a headphone according to the present application is shown; and
[0032] Figure 5 A communication headset for performing a method according to the present application is schematically shown. DETAILED DESCRIPTION
[0033] For a better understanding of the principles of the present application, embodiments of the present application are explained in more detail below with reference to the accompanying drawings. It is understood that the present application is not limited to these embodiments and that features described can also be combined or modified without departing from the scope of the present application as defined in the claims.
[0034] The method according to the present application can for example be used to reduce the occlusion effect in an in-ear headphone, as schematically shown in Figure 1 In this case, the in-ear headphone 10 is located on the ear of a user, wherein the ear insert 14 of the in-ear headphone is introduced into the external ear canal 15 in order to hold it in place. Depending on the specific location and material in the ear canal, the ear canal is sealed to some extent by the ear insert. This leads to the fact that external disturbing noise is at least partially shielded, so that these disturbing noises subsequently only reach the eardrum 16 of the user at a reduced level. Thus, on the one hand the music reproduction by the headphone is less disturbed, or in the case of a telephone call made with the aid of the headphone, the reproduction of the voice of the caller is less disturbed. On the other hand, the voice of the user is also attenuated by the ear insert, resulting in the occlusion effect already mentioned above.
[0035] The disturbing sound signal x(t) reaching the headphone from the surroundings, which can in particular contain the voice of the user, but also environmental noise, is picked up with the aid of an external microphone 11 which is oriented away from the ear canal towards the surroundings of the headphone. In addition, the in-ear headphone 10 has an internal microphone 12 which is directed towards the ear canal or the eardrum of the user and a loudspeaker 13 which is located in the vicinity of the internal microphone 12. With the aid of the loudspeaker 13 a compensation signal u(t) can be output, with the aid of which the occlusion effect is ideally conveyed to the user that he does not have the impression of wearing a headphone.
[0036] Here, the air sound components of the disturbing sound signals are picked up by means of the external microphone 11 and a compensation signal is generated therefor. In addition, the internal microphone 12 picks up the residual signal e(t) after the superimposition of the compensation signal u(t) filtered through the secondary path S(s) and the disturbing sound signal x(t) filtered through the primary path P(s) and, inter alia, can also pick up the body sound components and take them into account in the compensation signal. The acoustic primary path P(s) here describes the transfer function for the acoustic transmission from the external microphone 11 to the internal microphone 12 and can be measured, for example, with an external loudspeaker structure. The acoustic secondary path S(s) describes the transfer function from the internal loudspeaker 13 to the internal microphone 12 and can be measured with the use of this loudspeaker and internal microphone.
[0037] The shown in-ear earphone has only one external microphone, but a plurality of microphones arranged in a microphone array can also be used. Furthermore, the occlusion effect can also occur in other earphones, for example in a circum-aural earphone with an ear pad that surrounds the ear and closes the ear canal by means of a closed construction, or in a hearing aid, and is compensated as described below.
[0038] Figure 2 The basic concept of a method for actively suppressing the occlusion effect is schematically shown, which can be carried out, for example, when reproducing an audio signal using an in-ear earphone in Figure 1 Here, in a first step 20, the external sound is picked up with at least one external microphone 11 of the earphone or hearing aid. The picked-up external sound also includes an acoustic speech component from the speech output of the user wearing the earphone. In a subsequent step 21, a speech signal corresponding to the speech output of the user is picked up with at least one additional microphone, for example with a microphone of a communication headset that is directed at the mouth of the user (also referred to below as mouth microphone).
[0039] Next, in step 22, the dry component of the speech signal picked up with the additional microphone is estimated. As is known to the person skilled in the art, the recorded dry audio signal is understood to be a pure sound signal, as it exists initially at the time of production, that is to say without reverberation produced by reflections of the produced sound waves in an enclosed space or in a naturally limited area and without acoustic disturbances of the environment. Thus, in this step, the speech signal is estimated as it is produced directly through the sound channel of the user.
[0040] On the basis of the estimated dry components of the captured speech signal, in a following step 23, for the microphone signal of the respective outer microphone, the contained binaural speech signal is estimated and extracted using a filter, wherein the filter coefficients of the filter are determined on the basis of the estimated dry components of the captured speech signal. Alternatively, the estimated dry speech signal can also be filtered such that it has a similar spatiality as the speech components at the outer microphone. The extracted or generated binaural speech components are then output in step 24 via the respective loudspeaker of the earphone or hearing aid, wherein the signal is previously adjusted by means of a forward ("feedforward") filter, such that an acoustically as transparent as possible reproduction of the speech signal is possible.
[0041] Figure 3 A block diagram of a device according to the application is shown, which can be implemented in particular in an earphone, but also in a hearing aid. Although usually sound transducers are provided for both ears of a user in an earphone or hearing aid, in the attached figures only the conceptual structure with respect to one ear is shown for the sake of improved clarity. Likewise, for the digital signal processing, although an analog-to-digital converter for digitizing the sound signal captured with the microphone and a digital-to-analog converter for converting the processed signal for output via the loudspeaker are required, these are not shown in the attached figures for the sake of simplicity. On the basis of the digital signal processing, in the following the signals are observed in the time domain with discrete time indices n, the indices z correspondingly representing the frequency domain representation of the signals and filters.
[0042] As already mentioned in connection with Figure 1 In addition to the loudspeaker 13, an outer microphone 11 and an inner microphone 12 are provided, which can be arranged in the earpiece or earphone housing, respectively. Here, the outer microphone 11 providing the signal x(n) is mounted on the outside of the earphone. On the other hand, the loudspeaker 13 and the inner microphone 12 are arranged inside the earphone and oriented towards the eardrum.
[0043] In addition, a mouth microphone 17 is provided. This can be part of a communication headset, for example, and is mounted on a pivotable arch in order to be arranged in front of the mouth of the user and aligned with the mouth. It is also possible, however, to provide a microphone array consisting of a plurality of microphones, which is arranged on the outside of the earphone or hearing aid and aligned with the mouth, for example, by means of a beamforming method. In addition to the primary path P(z) representing the acoustic transmission from the outer microphone to the inner microphone and the secondary path S(z) for the transmission from the loudspeaker to the inner microphone, here also the transmission path B(z) between the mouth microphone and the outer reference microphone is noted, which is given, for example, in a communication headset by the pivotable microphone in front of the mouth at a predefined position relative to the position of the outer microphone. The transmission path here also contains the influence of other components, such as not shown analog-to-digital converters and digital-to-analog converters.
[0044] If a user makes a voice output through headphones or a hearing aid, the external microphone 11 will capture the corresponding voice signal x. v (n). The collected speech signal x v (n) here includes the spatial impulse response, which contains all relevant information about the current acoustic spatial characteristics. However, in addition to this speech signal, since the external microphone 11 is mounted on the outside of the earphone, the external microphone 11 also collects interference signals x caused by ambient noise. a The audio signal x(n), composed of these two signal components, is then processed based on an estimation of the dry speech signal as described below, so as to achieve acoustic transparency for its own speech by outputting the processed speech signal u(n) via the speaker 13 of the headphones or hearing aid. Here, the speech signal arriving at the headphones from the outside is transmitted not only from the outside to the internal microphone via the primary path P(z), but also transmitted as a signal via the secondary path S(z), which is actively output through the speaker 13. In this way, the missing air sound component is added to the speech itself. The acoustic interference of the sound signals transmitted through these two paths thus results in acoustic transparency for the speech signal.
[0045] In the illustrated embodiment, not only the speech signal v(n) measured by the mouth microphone 17, but also the error signal e(n) from the internal microphone is fed to the estimation unit 30, where the pure dry speech signal is estimated. As it is produced in the acoustic tract and exists without reverberation caused by the surrounding space and without acoustic interference from the environment. Based on this monoaural estimate... The second estimation unit 31 extracts the binaural speech signal from the signal acquired using the external microphone of the left or right earphone. Alternatively, the estimated dry speech signal can be filtered to give it spatial characteristics similar to the speech components at the external microphone. Binaural speech signal x v (n) is then passed through a negative transfer function ( The signal is filtered by a digital filter 32 and finally fed as a speaker signal u(n) to the audio converter for output through headphones. The digital filter 32 is specifically designed here as a feedforward filter.
[0046] In order to estimate the dry speech signal in estimation unit 30 The speech signal V(n) can be measured via the mouth microphone 17 and then used as a speech reference. The estimation of the dry speech component at the external microphone can be performed, for example, by filtering the additional signal using the corresponding relative impulse response between the additional microphone and the external microphone, followed by averaging. For this purpose, the mouth microphone signal v(n) can be estimated, for example, by estimating the relative transmission path B(z) between the mouth microphone and the external microphone. To filter. Here, the speech signal v(n) is considered a single-ear source, however, this single-ear source is subsequently used with two headphones or ears.
[0047] Similarly, the error signal e(n) can be acquired through the internal microphone 12, and this error signal can also be used for dry speech signals. The estimate can be fed to the estimation unit 30. Since the ears are closed by the headphones, the speaker's own speech is strongly coupled into the ear canal through the body; therefore, information about the speaker's own speech can also be obtained using the microphone signal from the internal microphone. The error signal e(n) includes an error component e based on the speech signal. v (n) and the other error component e b (n), the other error component e b (n) Based on other interference, such as footsteps transmitted through the user's body into the ear canal. Here, a separate error signal is generated for each of the two headphones or ears. For example, the error signal may differ when the headphones are paired differently. However, the individual error signals can also be averaged if necessary to obtain a single-ear signal again.
[0048] The signals from the mouth microphone and the internal microphone can be adjusted, for example, by digital filtering, and then combined by subsequent averaging to further improve the signal-to-noise ratio. It is important to note that the signal played through the headphone speaker is convolved with the estimates of the corresponding secondary paths and subtracted from the corresponding internal microphone signals to disable signal feedback.
[0049] Since the internal microphone primarily records the body sound component of one's own speech, which does not allow for decomposition such as fricatives, one can also consider extending the bandwidth of the internal microphone signal.
[0050] Because both the mouth microphone and the internal microphone provide a good signal-to-noise ratio, it can be stipulated that estimation can be performed solely based on the signal measured using either the mouth microphone or the internal microphone, instead of relying on a combination of signals from both microphones. Finally, at particularly advantageous scales, these can already provide a direct reference to the speech without the need for additional estimation.
[0051] In the second estimation unit 31, the binaural speech signal is estimated by extracting the binaural speech from the external microphone signal, which is affected by ambient noise, based on the estimation of the dry speech, or by generating a speech signal with spatiality similar to the speech components at the external microphone. Importantly, the processing involves a short and constant delay, which can be taken into account when calculating the forward filter W(z).
[0052] For this purpose, a Wiener filter or other algorithms for suppressing interference noise can be used, for example. In a Wiener filter, the amplitude spectrum of the acquired signal is evaluated so that a filter can be calculated using an estimate of the speech signal and an estimate of the current interference signal, which can be used to optimally extract the speech signal. Therefore, for example, the amplitude spectrum of the mouth microphone can be combined with the amplitude spectrum of the internal microphone to estimate the amplitude spectrum of the dry speech signal, and then the speech component can be extracted from the signal from the external microphone. In this case, the transfer function B(z) can be used to estimate how the dry speech travels from the mouth microphone to the external microphone, thereby compensating for the direct sound propagation time.
[0053] Since the transfer function B(z) in communication headsets is very similar for different people, the impulse response can be determined, for example, by a series of measurements for a specific headset, and then applied to the application of headsets of that structural form.
[0054] One possibility is Wiener filtering in a "filter bank equalizer" structure. This structure is based on a prototype low-pass filter with a constant bank runtime. The spectral weights of the Wiener filter are predicated on estimates of the useful signal and the interfering signal. For the estimation of the useful signal component, an estimate of the dry speech can be used.
[0055] Alternatively, an adaptive filter a(n) can be used to estimate binaural speech. Assume the external microphone signal x(n) = x a (n)+x v (n) is determined by environmental noise x a (n) and estimation of dry speech coherent speech components x v Composed of (n), it can be used to base on adaptive filters. Reproducing the speech component x using x(n) v (n). Utilizing the output of the adaptive filter The adjustment rule for the adaptive filter can be found based on the following cost function:
[0056] , in .
[0057] Furthermore, the estimation unit 31 can analyze the acoustic effects of space on its own speech and select or design a filter based on this, which can be applied to the estimated dry speech signal to generate a speech signal with spatiality similar to the speech components at the external microphone.
[0058] Forward filter It can be determined, for example, by solving the Wiener-Hopf equation.
[0059]
[0060] Therefore, it is necessary to measure the primary path once or multiple times. and secondary paths These measurements can be performed, for example, on an artificial head or on a subject. Importantly, the secondary path used to calculate the feedforward filter takes into account any processing-induced delays in the branches between the respective external microphones and headphone speakers. Therefore, if, when estimating binaural speech, for example, a signal... If any resulting signal, or any signal subsequently played through a speaker, is delayed, then the secondary path must account for that delay. This is indicated by an apostrophe in the Wiener-Hopf equation above.
[0061] The desired transmission characteristics from the external microphone to the internal microphone are achieved through... In the domain Or through impulse response This is used to describe, and is also necessary for the Wiener-Hopf equation, the internal microphone used to naturally perceive its own speech, which is typically characterized by a flat amplitude response.
[0062] Figure 4 A block diagram of another device according to the invention is shown. Besides... Figure 3 In addition to the unit of the device according to the invention, a control unit 40 for controlling the two weighting units 41 and 42 is also provided here. Because in the illustrated case... and x v (n) are coherent, meaning they are not offset from each other in the time domain, or at least not significantly offset from each other, so the two signals can be weighted using linear weighting factors α and 1-α, where 0 ≤ α ≤ 1, and then added together. Weighting units 41 and 42 thus allow users to personalize the mixing of dry speech and binaural speech. Therefore, the user can decide and set how he / she perceives his / her speech, for example, what the ratio of reverberation volume to the volume of his / her own speech should be. However, control can also be performed automatically.
[0063] As mentioned above, a consequence of the occlusion effect is that the low-frequency components of the own voice are amplified. In order to compensate for this, the internal microphone signal can additionally be filtered with a feedback controller, thereby reducing the low-frequency components of the own voice. In this way, the perception of the own voice in the case of wearing earphones is made to appear more natural.
[0064] Here, the estimation units 30 and 31 and the control unit 40 can be part of a processor unit, which has one or more digital signal processors, but can also contain different types of processors or combinations thereof. Furthermore, the filter coefficients of the digital filter 32 can be adjusted by means of a digital signal processor. The filter can be implemented as a time-invariant filter, which is calculated once, loaded onto the firmware of the earphone and used in this form without any changes at runtime. An adaptive filter can also be used, which changes at runtime and adapts to the current situation.
[0065] The device according to the application is preferably fully integrated in the earphone, since the delay due to the transmission of the own voice through the body sound is very small. In this case, the mouth microphone can also be part of the earphone, for example in the case of so-called communication headsets, which are fixed on the bow mounted in front of the mouth, or integrated in the head cover as a microphone array with directional characteristics. However, a separate microphone can also be used as a mouth microphone. In principle, however, the components of the device can also be part of an external device, for example a smartphone.
[0066] Figure 5 The application of a communication headset is schematically shown, in which the method according to the application can be carried out and which has the device described above for this purpose. Here, an earphone 10 is provided for each ear of the user, in which an external microphone 11, an internal microphone 12 and a loudspeaker 13 are integrated, respectively. Furthermore, a mouth microphone 17 is provided, which is mounted on a pivotable bow. Furthermore, a processor unit 50 is arranged in one of the two earphones, by means of which the estimation unit and possibly the control unit 40 are implemented. The individual components are connected here to the processor unit 50, but in order to improve clarity, the processor unit is not shown in the drawing.
[0067] The application can be used to suppress the occlusion effect when reproducing audio signals with any earphone or hearing aid, for example when using a communication headset / hearable for telephone services or communication, in the case of so-called in-ear monitoring, for checking the own voice in the case of live performances, augmented / virtual reality applications, or when using in a hearing aid.
[0068] List of reference signs
[0069] 10 individual earpiece, individual hearing aid
[0070] 11 external microphone
[0071] 12 internal microphone
[0072] 13 loudspeaker
[0073] 14 ear insert
[0074] 15 ear canal
[0075] 16 eardrum
[0076] 17 mouth microphone
[0077] 20-24 method steps
[0078] 30 first estimation unit
[0079] 31 second estimation unit
[0080] 32 digital forward filter
[0081] 40 control unit
[0082] 41, 42 weighting unit
[0083] 50 processor unit
Claims
1. A method for actively suppressing occlusion effects when reproducing audio signals using headphones (10) or a hearing aid, wherein, - Use at least one external microphone (11) of the earphone or hearing aid to collect external sound in the form of sound signals from the outside; - Acquire speech signals using at least one microphone or microphone array pointed at the user's mouth and / or the internal microphone of the earphone or hearing aid; - Estimate the dry component of the acquired speech signal, wherein the dry component of the acquired speech signal is the part of the acquired speech signal that has neither reverberation caused by the surrounding space nor environmental noise, and based on this extract the binaural speech signal from the signals of at least two external microphones of the left and right earphones or the left and right hearing aids, or filter the estimated dry speech component to produce a binaural speech signal with spatiality similar to the speech component at the external microphones. - Extract speech components from external sounds acquired using the at least one external microphone by means of a filter, wherein the filter coefficients of the filter are determined based on the estimated interference components of the acquired speech signal, or the acoustic effects of space on the speech itself are analyzed based on the interference components of the acquired speech signal and the external sounds acquired using the at least one external microphone, and the estimated interference components of the acquired speech signal are filtered accordingly, so as to produce speech components having spatiality similar to the speech components at the at least one external microphone. and - The extracted or generated speech components are output through the speaker of the headphones or hearing aid.
2. The method according to claim 1, wherein, The binaural speech signals are filtered by the speakers (13) of the left and right earphones or left and right hearing aids before being output.
3. The method according to claim 1 or 2, wherein, The interference component of the acquired speech signal is estimated by filtering using the corresponding relative impulse response between at least one mouth microphone or microphone array and the external microphone (11) and the subsequent average formation.
4. The method according to claim 1 or 2, wherein, The filters used to extract or generate speech components based on the acquired external sound and the estimated dry speech are Wiener filters, adaptive filters, or filters that simulate spatial impulse responses.
5. The method according to claim 1 or 2, wherein, The estimated interference component and the extracted or generated speech component of the acquired speech signal are linearly weighted, summed, and then output through the speaker of the headphones or hearing aid.
6. An apparatus for actively suppressing occlusion effects when reproducing audio signals through a speaker (13) of an earphone (10) or a hearing aid equipped with at least one external microphone (11), the apparatus comprising: - At least one additional microphone for capturing the user's voice signal; - Digital signal processor (50), the digital signal processor being configured to: - Estimate the interference component of the speech signal acquired using the at least one additional microphone, wherein the interference component of the acquired speech signal is the portion of the acquired speech signal that has neither reverberation caused by the surrounding space nor environmental noise. - Extract speech components from external sounds acquired using the at least one external microphone (11) using a filter, wherein, The filter coefficients of the filter are determined based on the estimated interference component of the acquired speech signal, or the acoustic effects of space on the speech itself are analyzed based on the interference component of the acquired speech signal and external sounds acquired using at least one external microphone, and the estimated interference component of the acquired speech signal is filtered accordingly to produce a speech component with spatiality similar to the speech component at the external microphone. and - The extracted or generated speech components are output through the speaker (13). The at least one additional microphone used to acquire the speech signal includes at least one microphone or microphone array pointing towards the user's mouth and / or the internal microphone of the earphone or hearing aid. Specifically, a dry component is estimated based on the acquired speech signal, and based on this, a binaural speech signal is extracted from the signals of at least two external microphones of the left and right earphones or left and right hearing aids, or the estimated dry speech component is filtered to generate a binaural speech signal with spatiality similar to the speech component at the external microphone.
7. The apparatus according to claim 6, wherein, A digital filter (32) is additionally provided, and the extracted or generated speech components are fed to the digital filter before being output through the speaker (13).
8. An earphone (10) having the device according to claim 6 or 7.
Citation Information
Patent Citations
A headset with hear-through mode
EP3188495A1
Spatial headphone transparency
US10034092B1
Own voice shaping in a hearing instrument
EP2920980B1