Adaptive Hybrid Method for Uncorrelated or Correlated Noisy Signals and Hearing Device

By configuring a mixing processor in the listening device, balancing the noise components in multiple audio data streams, the problems of non-natural noise and level changes in the mixed signal are solved, and a more natural signal perception effect is achieved.

CN112533121BActive Publication Date: 2025-06-24OTICON
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010997603.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-19
Filing Date
2020-09-21
Publication Date
2025-06-24
Estimated Expiration
2040-09-21

AI Technical Summary

Technical Problem

When mixing multiple noisy audio data streams, it is difficult to maintain the balance of the target component and the equalization of the noise component, resulting in unnatural noise and level changes in the mixed signal.

Method used

By configuring a mixing processor in a listening device, multiple audio data streams are received and mixed, and noise components are balanced in the processed version, non-natural signals caused by mixing are reduced or avoided.

Benefits of technology

When mixing multiple audio data streams, the target signal and the noise signal are maintained, the appearance of unnatural signals is reduced, and the natural perception effect of the signal is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112533121B_ABST
    Figure CN112533121B_ABST
Patent Text Reader

Abstract

The present application discloses an adaptive hybrid method for uncorrelated or correlated noisy signals and a hearing device, wherein the hearing device comprises: an input unit providing at least two input audio data streams, each input audio data stream comprising a mixture of a target signal component from a target sound source and a noise component from one or more noise sources; a mixing processor configured to receive at least two input audio data streams, to mix at least two input audio data streams or processed versions thereof, and to provide a processed input signal based thereon; an output unit configured to provide an output stimulus perceptible as sound by a user based on the processed input signal or a processed version thereof; wherein the processor is configured to process the noise components of at least two input audio data streams or processed versions thereof to reduce or avoid unnatural signals caused by mixing in the processed input signal by balancing the noise components of at least two input audio data streams in the processed input signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to hearing devices such as hearing aids, headphones or speakerphones, and in particular to hearing devices configured to receive multiple (possibly) noisy audio data streams, for example via an input transducer or by a wireless or wired receiver. Background Art

[0002] When mixing two or more noisy audio data streams, it is desirable that the target component and / or the noise component of the mixed streams satisfy certain characteristics. Summary of the invention

[0003] The target components should preferably be well balanced, i.e. the target from one sound source should preferably not be significantly louder or quieter than the target component from another sound source. The noise component should preferably also be well balanced and preferably not affected by the mixing, in that respect it is perceived as an annoying sound. In this specification, "balanced" or "well balanced" means "equalized", e.g. the balanced components are forced to be substantially equal, e.g. forced to be within a certain gap from each other, e.g. their numerical difference relative to the minimum value of the two components is less than 10%. The noise component that is "balanced" or "well balanced" or "equalized" can, for example, be the noise variance of the corresponding audio data stream before mixing.

[0004] As an example, consider a fading between two microphone signals consisting of a target signal + internal (audible) microphone noise. If the two microphones are of different types, the microphone noise levels may be different. When fading from one microphone signal to the other (i.e. gradually attenuating the level of the (current) signal and increasing the level of the other signal), it may be desirable to maintain the level of the target sound. If the microphones have different signal-to-noise ratios (SNRs), if the target level remains constant during the fading, the noise level will change. Thereby, the fading becomes audible. Even if the SNR at the two microphones is the same, the fading is still audible because during the fading, the relevant target sound (such as speech) and the unrelated noise (such as microphone noise and / or wind noise) are superimposed in different ways.

[0005] The mixing of data streams may be the result of switching from one program to another. Hearing devices such as hearing aids are usually equipped with multiple (user-selectable or automatically controlled) dedicated combinations of processing parameters ("settings"), which are optimized for different acoustic situations such as telephone programs, music programs, listening programs, conversation programs, etc. The aforementioned dedicated combinations of processing parameters are usually called "programs". The concept of the present invention regarding the mixing of data streams also applies to switching from one program to another (e.g. from a music listening program to a telephone program), where a gradual change from one sound input to another sound input will occur.

[0006] Hearing device

[0007] In one aspect of the present application, there is provided a hearing device such as a hearing aid, which is adapted to be located at or in a user's ear or is adapted to be fully or partially implanted in the user's head. The hearing device comprises:

[0008] - an input unit for providing at least two input audio data streams, each input audio data stream comprising a mixture of a target signal component from a target sound source and a noise component from one or more noise sources;

[0009] - a mixing processor for receiving at least two input audio data streams, for mixing at least two input audio data streams or processed versions thereof, and for providing a processed input signal based thereon;

[0010] - an output unit for providing an output stimulus perceptible as sound by the user based on the processed input signal or a processed version thereof.

[0011] The processor may be configured to process the noise components of at least two input audio data streams or processed versions thereof to reduce or avoid unnatural signals caused by mixing in the processed input signal. This may be achieved by balancing the noise components of at least two input audio data streams in the processed input signal.

[0012] The processor may be configured to estimate the noise variances of at least two input audio data streams before mixing. The noise variance may be measured, for example, as the background noise floor variance. The processor may be configured to process the noise components according to the noise variances of at least two input audio data streams.

[0013] Thus, an improved hearing device can be provided.

[0014] Mixing can start automatically, e.g., based on sensor input or detector input, e.g., based on analysis of the corresponding input signals. Mixing can be determined and / or started by the user, e.g., via a user interface (e.g., an APP implemented as a smart phone). Mixing can include or consist of a fade from one signal to another signal over a certain period of time. The fade can be defined by a fade parameter (e.g., varying over time) or by a fade curve that (e.g., gradually) decreases the gain (or weight) of one input audio stream (or its processed version) while increasing the gain (or weight) of another input audio stream (or its processed version). The fade can be regarded as a form of temporary mixing. In an embodiment, the fade from one audio stream to another includes one audio stream being "selected" (presented to the user) before the fade starts, and another audio stream being "selected" (presented to the user) when the fade ends (i.e., the end weights switch from α1 = 1 and α2 = 0 to α1 = 0 and α2 = 1 respectively, and vice versa). In an embodiment, the corresponding weights before and after the fade are not 0 and 1, but relatively large (close to 1, e.g., ≥ 0.9) weights are applied to the corresponding main audio streams (before and after the fade), and relatively small (non-zero, e.g., ≤ 0.1) weights are applied to the non-main audio streams (before and after the fade).

[0015] The processor can be configured to identify at least some (partial or all) of the noise components that would otherwise access multiple input audio streams or their processed versions. The noise components (such as microphone noise) can be known or determined before the operation of the hearing device and made available to the hearing device, e.g., stored in a memory or made available to the processor. The microphone noise can be extracted from the microphone's specifications or measured before the microphone is used. The hearing device can be configured to (adaptively) determine noise components (such as ambient noise) by estimating the noise in the environment during speech pauses.

[0016] The term "noise source" can include one or more of the following: microphone noise (noise inherent in the microphone), ambient acoustic or mechanical noise, and electromagnetic induction noise. The term "noise source" can include one or more competing (non-target) speech sources that are currently perceived as noise by the user.

[0017] The noise from one or more noise sources of at least two input audio data streams may be uncorrelated (e.g., microphone noise or wind noise). The noise from one or more noise sources of at least two input audio data streams may be correlated (e.g., acoustic noise from non-target speech, or noise from a fan or other machine).

[0018] The target sound sources of at least two of the plurality of audio data streams may be different (i.e., at least two input audio data streams originate from two different target sound sources). The input unit may include an input transducer (such as a microphone) for converting local sounds from the user environment of the hearing device wearer into an electrical input signal (such as a first audio data stream) representing the local sounds. The input unit may include (an antenna and) a receiver circuit for receiving a second electrical input signal (such as a second audio data stream) representing sounds different from the local sounds from the user environment of the hearing device wearer from a (possibly distant) transmitter. The at least two audio data streams may include a first and a second audio data stream respectively from the input transducer and the (antenna and) receiver circuit.

[0019] In this specification, the term "originate from" means "come from" or "be provided by", or a similar expression indicating that "A is the source of B" (here, the target sound source is the source of the audio data stream).

[0020] In this specification, the term "the target sound sources of at least two of the plurality of audio data streams are different" means that at least two audio data streams originate from two different target sound sources, such as two different speakers, or a target signal wirelessly received from a distant communication partner and an electrical input signal representing a target speaker in the user environment. In other words, this term does not cover two audio streams that are just staggered in time.

[0021] The target sound sources of at least two of the plurality of audio data streams may be the same (i.e., at least two input audio data streams originate from the same target sound source). The input unit may include at least two (such as first and second) input transducers (such as microphones), each input transducer for converting local sound from the user environment of the hearing device worn into a corresponding electrical input signal (e.g., converted into first and second electrical input signals, such as first and second audio data streams), each electrical input signal representing local sound (possibly different sound pressure levels, including different amounts or types of noise, etc.). At least two input transducers may be located in different parts of the hearing device, one, for example, located in the BTE part adapted to be located behind the user's ear (pinna), and one, for example, located in the ITE part adapted to be located in or at the user's ear canal. In an embodiment, the input unit includes an input transducer (such as a microphone) for converting local sound from the user environment of the hearing device worn into an electrical input signal representing the local sound (such as a first audio data stream). The input unit may further include a (antenna and) receiver circuit for receiving (such as wirelessly receiving) a second electrical input signal (such as a second audio data stream) from a transmitter, which represents sound of local sound from the user environment of the hearing device worn (e.g., from a speaker in the user environment wearing a microphone including an audio transmitter). At least two audio data streams may include a first and a second audio data stream respectively from the input transducer and the (antenna and) receiver circuit. The antenna and receiver circuit may include a coil (such as a pick-up coil or other inductor) for receiving an audio signal from an inductive transmitter. The antenna and receiver circuit may include an RF antenna for receiving electromagnetic signals in the GHz range, such as Bluetooth or equivalent components.

[0022] The processor may be configured to apply at least one signal processing algorithm to the processed input signal and provide a processed output signal. The processor may be connected to the input unit. The processor may be connected to the output unit. The processed output signal may be fed to the output unit. At least one signal processing algorithm may, for example, include a noise reduction algorithm. The processor may, for example, include a post-filter for filtering the processed input signal to attenuate the noise component in the processed input signal. The processor may, for example, include a compression amplification algorithm for applying amplification or attenuation varying with frequency and level to the processed input signal (thus, for example, compensating for the user's hearing impairment).

[0023] The hearing device may include a filter bank enabling signal processing in the (time-)frequency domain. The input unit may include a corresponding analysis filter bank for providing the plurality of input audio data streams in sub-band representation. The input unit may include corresponding analog-to-digital converters to digitalize each electrical input signal into an audio data stream.

[0024] The output unit may include a synthesis filter bank for converting the sub-band signals into an audio data stream in the time domain for generating the output stimulus. The output unit may include a loudspeaker for providing the output stimulus as an acoustic signal (vibrations in air, e.g., towards the eardrum of the user). The output unit may include a vibrator for providing the output stimulus as mechanical vibrations in the user's skull. The output unit may include an electrode array for providing the output stimulus as electrical stimulation of the user's cochlear nerve. Each electrode of the electrode array may be configured to receive stimuli targeted at different sub-frequency ranges of the human auditory system (e.g., below 20 kHz, such as below 12 kHz or below 10 kHz or below 8 kHz). In the latter case, the synthesis filter bank may be omitted (in some designs).

[0025] The processor may be configured to estimate the level of the target component of a plurality of input audio data streams. The level may be estimated, for example, as a sound pressure level, e.g., denoted in dB SPL. The processor may be configured to estimate the signal-to-noise ratio of a plurality of input audio data streams.

[0026] Mixing may include fading. The processor may be configured to fade from one input audio data stream to another input audio data stream. Sometimes, it is desirable to fade from one microphone signal to another microphone signal, e.g., if one microphone has more feedback than the other, or if the target signal-to-noise ratio of one microphone signal is (significantly) better than that of the other microphone signal. The processor may be configured to maintain the level of the target signal component when fading from one input audio data stream to another input audio data stream. In this specification, "fading" between a first and a second audio-containing signal means moving from a situation where the first signal is presented to the user (and the second signal is attenuated or disabled) to a situation where the second signal is presented to the user (and the first signal is attenuated or disabled). The hearing device may be configured to fade from an audio stream from a microphone to an audio stream from a beamformer and vice versa.

[0027] Fading from a first input audio data stream to a second input audio data stream over a certain fade time period may include the mixing processor being configured to provide the first data stream as a processed signal at a first time point t1 and the second data stream as a processed signal at a second time point t2, where the second time t2 is greater than the first time t1. The fade process starts by presenting one signal ("the first audio data stream") to the user and ends by presenting the other signal ("the second audio data stream") to the user. The fade time Δt fad = t2 - t1 may be less than a predetermined time range, e.g., Δt fad <20 s, or <10 s, e.g., <5 s.

[0028] The fading from a first input audio data stream to a second input audio data stream across a certain fading time period may include determining corresponding fading parameters or a fading curve that gradually reduces the weight of the first input audio data stream or its processed version while increasing the weight of the second input audio data stream or its processed version, wherein the noise level of the (perceived) processed input signal is not substantially changed during the fading.

[0029] The hearing device may be configured to start fading based on detecting feedback of one of the input audio data streams. The hearing device may include a feedback detector. The feedback detector may be configured to provide a measure of the feedback level currently experienced, from the output transducer of the hearing device to the input transducer (such as more than two input transducers).

[0030] The hearing device may include a voice activity detector configured to provide a VAD control signal indicating whether or with what probability a given input audio stream includes human speech. The voice activity detector may be configured to identify speech. The voice activity detector may be configured to indicate whether or with what probability a given sub - band of the input audio stream includes human speech such as speech.

[0031] The input unit may include at least two input transducers, each input transducer providing an electrical input signal representing sound, and a beamformer filtering unit configured to perform spatial filtering on the electrical input signals and provide at least one spatially filtered signal based thereon, the spatially filtered signal constituting or forming part of at least one of the plurality of input audio data streams. The input unit may for example include two, three or more input transducers.

[0032] The beamformer filtering unit may include at least two beamformers configured to provide at least two spatially filtered signals, which may constitute or form part of the plurality of input audio data streams. The direction from the user to the target sound source determined by the at least two beamformers may represent two different target directions (two different target sound sources). The number of beamformers may be less than the number of input transducers. The number of beamformers may be greater than or equal to the number of input transducers.

[0033] The fading between the corresponding input audio streams from at least two beamformers may be controlled according to the detected or selected target direction. The direction of the target sound source that the user is currently interested in (target direction) may be automatically determined. The target direction may also be selected by the user, for example via a user interface, such as implemented on a smart phone or other portable device, such as including a graphical user interface.

[0034] The hearing device may be configured such that the fade time Δt is increased. When fading between spatially filtered (beamformed) signals (representing sounds from two different target sound sources), the duration of the fade (fade time Δt) may preferably be increased compared to fading between two microphone signals representing sounds from the same target sound source (but picked up at (slightly) different positions). Thereby, the switch from one target signal to another (different) target signal (which may have very different signal-to-noise ratios and levels) is more gentle, otherwise, the switch may be perceived by the user as sudden or annoying.

[0035] The hearing device may be configured to fade between at least two input audio data streams or processed versions thereof, while ensuring that the noise components in the processed input signals are equalized to the level of the noise signal component exhibiting the maximum noise level in the input audio data stream.

[0036] The hearing device may be configured to fade between at least two input audio data streams or processed versions thereof, while ensuring that the levels of the target signal components in the processed input signals are equalized. The hearing device may for example be configured to maintain the level of the target signal in the processed input signal during the fade from one input audio stream to another.

[0037] The hearing device may be configured to fade between at least two input audio data streams or processed versions thereof, the hearing device comprising a single-channel post-filter for attenuating the noise in the processed input signal, wherein the post-filter is configured to increase the attenuation of the noise component of the processed input signal.

[0038] The hearing device may consist of a hearing aid, headphones, a headset, an ear protection device or a combination thereof or comprise a hearing aid, headphones, a headset, an ear protection device or a combination thereof. The hearing device may include a speakerphone (for example adapted to be located on a table).

[0039] The hearing device may be adapted to provide frequency-dependent gain and / or level-dependent compression and / or frequency shifting from one or more frequency ranges to one or more other frequency ranges (with or without frequency compression) to compensate for the user's hearing impairment. In an embodiment, the processor is configured to enhance the input signal and provide a processed output signal.

[0040] The output unit may be configured to provide a stimulus perceived by the user as an acoustic signal based on the processed electrical signal. The output unit may include a plurality of electrodes of a cochlear implant (for a CI type hearing device). The output unit may include an output transducer. The output transducer includes a receiver (speaker) for providing the stimulus as an acoustic signal to the user (for example in an acoustic (air-conduction based) hearing device). The output transducer may include a vibrator for providing the stimulus as a mechanical vibration of the skull to the user (for example in a bone-attached or bone-anchored hearing device).

[0041] A hearing device may include an input unit for providing an electrical input signal representative of sound. The input unit may include an input transducer, such as a microphone, for converting an input sound into an electrical input signal. The input unit may include a wireless receiver for receiving a wireless signal comprising or representing sound and providing an electrical input signal representative of said sound.

[0042] The wireless receiver may be configured, for example, to receive electromagnetic signals in the radio frequency range (3 kHz to 300 GHz). The wireless receiver may be configured, for example, to receive electromagnetic signals in the optical frequency range (e.g., infrared light from 300 GHz to 430 THz, or visible light, e.g., from 430 THz to 770 THz).

[0043] The hearing device may include a directional microphone system adapted to perform spatial filtering of sounds from the environment so as to enhance a target sound source among a plurality of sound sources in the local environment of a user wearing the hearing device. The directional system may be adapted to detect (such as adaptively detect) from which direction a particular part of a microphone signal originates. This may be achieved in a number of different ways as described in the prior art. In a hearing device, a microphone array beamformer is typically used to spatially attenuate background noise sources. Many beamformer variants can be found in the literature. The minimum variance distortionless response (MVDR) beamformer is widely used in microphone array signal processing. Ideally, the MVDR beamformer leaves the signal from the target direction (also called the look direction) unchanged while attenuating sound signals from other directions to the maximum extent. The generalized sidelobe canceller (GSC) structure is an equivalent representation of the MVDR beamformer, which offers computational and digital representation advantages over a direct implementation of the original form.

[0044] The hearing device may include an antenna and transceiver circuitry (such as a wireless receiver) for receiving a direct electrical input signal from another device, such as from an entertainment device (e.g., a television set), a communication device, a wireless microphone, or another hearing device. The direct electrical input signal may represent or include an audio signal and / or a control signal and / or an information signal. The hearing device may include demodulation circuitry for demodulating the received direct electrical input signal, thereby providing a direct electrical input signal representing the audio signal and / or the control signal, e.g., for setting operating parameters (such as volume) and / or processing parameters of the hearing device. Generally, the wireless link established by the antenna and transceiver circuitry of the hearing device can be of any type. The wireless link may be established between two devices, e.g., between an entertainment device (such as a TV) and the hearing device, or between two hearing devices, e.g., via a third intermediate device (such as a processing device, e.g., a remote control device, a smart phone, etc.). The wireless link may be used under power-limited conditions, e.g., in the case where the hearing device is or includes a portable (usually battery-powered) device. The wireless link may be a near-field communication-based link, e.g., an inductive link based on inductive coupling between antenna coils of a transmitter part and a receiver part. The wireless link may be based on far-field electromagnetic radiation. In an embodiment, the communication via the wireless link is arranged according to a specific modulation scheme, e.g., an analog modulation scheme such as FM (frequency modulation) or AM (amplitude modulation) or PM (phase modulation), or a digital modulation scheme such as ASK (amplitude shift keying) such as on-off keying, FSK (frequency shift keying), PSK (phase shift keying) such as MSK (minimum shift keying) or QAM (quadrature amplitude modulation), etc.

[0045] The communication between the hearing device and another device may be in the baseband (audio frequency range, e.g., between 0 and 20 kHz). Preferably, the communication between hearing devices is based on a certain type of modulation at a frequency higher than 100 kHz. Preferably, the frequency used to establish a communication link between the hearing device and another device is lower than 70 GHz, e.g., in the range from 50 MHz to 70 GHz, e.g., higher than 300 MHz, e.g., in the ISM range above 300 MHz, e.g., in the 900 MHz range or in the 2.4 GHz range or in the 5.8 GHz range or in the 60 GHz range (ISM = industrial, scientific, and medical, such standardized ranges are defined, for example, by the International Telecommunication Union ITU). The wireless link may be based on a standardized or proprietary technology. The wireless link may be based on Bluetooth technology (such as Bluetooth Low Energy technology).

[0046] The hearing device may be a portable (i.e., configured to be wearable) device or form part thereof, such as a device including a native energy source such as a battery, e.g., a rechargeable battery. The hearing device may be a lightweight and easily wearable device, e.g., having a total weight of less than 100 g.

[0047] The hearing device may include a forward or signal path between an input unit (such as an input transducer, e.g., a microphone or a microphone system and / or a direct electrical input such as a wireless receiver) and an output unit such as an output transducer. A signal processor may be located in the forward path. The signal processor may be adapted to provide frequency-varying gain according to the specific needs of the user. The hearing device may include an analysis path having functions for analyzing the input signal (such as determining the level, modulation, signal type, acoustic feedback estimate, etc.). Part or all of the signal processing in the analysis path and / or the signal path may be performed in the frequency domain. Part or all of the signal processing in the analysis path and / or the signal path may be performed in the time domain.

[0048] An analog electrical signal representing an acoustic signal may be converted to a digital audio signal during an analog-to-digital (AD) conversion process, where the analog signal is sampled at a predetermined sampling frequency or sampling rate f s at, f s for example in the range from 8 kHz to 48 kHz (to suit the specific needs of the application) at discrete time points t n (or n) to provide digital samples x n (or x[n]), each audio sample being represented by a predetermined N b bits representing the value of the acoustic signal at t n when, N b for example in the range from 1 to 48 bits such as 24 bits. Each audio sample is thus quantized using N b bits (resulting in 2 Nb different possible values for the audio sample). The digital sample x has a time length of 1 / f s , such as 50 μs, for f s = 20 kHz. In an embodiment, a plurality of audio samples are arranged in time frames. A time frame may include 64 or 128 audio data samples. Other frame lengths may be used according to the actual application.

[0049] The hearing device may include an analog-to-digital (AD) converter to digitize an analog input (such as from an input transducer such as a microphone) at a predetermined sampling rate such as 20 kHz. The hearing device may include a digital-to-analog (DA) converter to convert a digital signal to an analog output signal, for example for presentation to the user via an output transducer.

[0050] A hearing device such as an input unit and / or an antenna and transceiver circuitry may include a TF conversion unit for providing a time-frequency representation of an input signal. The time-frequency representation may include an array or mapping of corresponding complex-valued or real-valued values of the signal involved in a specific time and frequency range. The TF conversion unit may include a filter bank for filtering the (time-varying) input signal and providing a plurality of (time-varying) output signals, each output signal including a distinct input signal frequency range. The TF conversion unit includes a Fourier transform unit for converting the time-varying input signal into a (time-)frequency domain (time-varying) signal. The frequency range considered by the hearing device, from a minimum frequency f min to a maximum frequency f max may include a part of the typical human audible frequency range from 20 Hz to 20 kHz, for example a part of the range from 20 Hz to 12 kHz. Generally, the sampling rate f s is greater than or equal to twice the maximum frequency f max , that is, f s ≥2f max . The signals of the forward path and / or the analysis path of the hearing device may be split into NI (e.g., uniformly wide) frequency bands, where NI is, for example, greater than 5, such as greater than 10, such as greater than 50, such as greater than 100, such as greater than 500, and at least some of them are processed individually. The hearing aid may be adapted to process the signals of the forward and / or analysis path in NP different channels (NP≤NI). The channels may have the same or different widths (e.g., the width increases with frequency), and may overlap or not overlap.

[0051] The hearing device may be configured to operate in different modes, such as a normal mode and one or more specific modes, which may be selected by the user or automatically selected, for example. The operating mode may be optimized for a specific acoustic situation or environment. The operating mode may include a low-power mode, in which the functions of the hearing device are reduced (e.g., for energy saving), for example, disabling wireless communication and / or disabling specific features of the hearing device.

[0052] The hearing device may include a plurality of detectors configured to provide status signals related to the current network environment of the hearing device (such as the current acoustic environment), and / or related to the current state of the user wearing the hearing device, and / or related to the current state or operating mode of the hearing device. As an alternative or in addition, one or more detectors may form part of an external device that communicates with the hearing device (such as wirelessly). The external device may include, for example, another hearing device, a remote control, an audio transmission device, a phone (such as a smart phone), an external sensor, etc.

[0053] One or more of the plurality of detectors may act on the full-band signal (time domain). One or more of the plurality of detectors may act on the frequency-band split signal ((time-)frequency domain), for example, in a limited number of frequency bands.

[0054] Multiple detectors may include a level detector for estimating the current level of the signal in the forward path. The detector may be configured to determine whether the current level of the signal in the forward path is above or below a given (L-) threshold. The level detector may act on the full-band signal (time domain). The level detector may act on the band-split signal ((time-) frequency domain).

[0055] The hearing device may include a voice activity detector (VAD) for estimating whether (or with what probability) the input signal (at a particular point in time) includes a voice signal. In this specification, a voice signal includes speech signals from humans. It may also include other forms of vocalization (such as singing) produced by the human speech system. The voice activity detector may be adapted to classify the user's current acoustic environment as a "voice" or "no voice" environment. This has the advantage that time periods of the electro-microphone signal that include human vocalizations (such as speech) in the user's environment can be identified and thus separated from time periods that include only (or mainly) other sound sources (such as artificially generated noise). The voice activity detector may be adapted to also detect the user's own voice as "voice". Alternatively, the voice detector may be adapted to exclude the user's own voice from the detection of "voice".

[0056] The hearing device may include a self-voice detector for estimating whether (or with what probability) a particular input sound (such as a voice, such as speech) originates from the voice of the user of the hearing system. The hearing device (such as the self-voice detector) may be adapted to be able to distinguish the user's own voice from the voice of another person and possibly from a no-voice sound. This is highly advantageous when implemented in combination with a voice control interface in the hearing device.

[0057] Multiple detectors may include a motion detector, such as an acceleration sensor. The motion detector may be configured to detect movements of the user's facial muscles and / or bones caused, for example, by speech or chewing (such as jaw movements) and provide a detector signal indicating the movement.

[0058] The hearing device may include a classification unit configured to classify the current situation based on the input signal from (at least part of) the detector and possibly other inputs. In this specification, the "current situation" is defined by one or more of the following:

[0059] a) The physical environment (such as including the current electromagnetic environment, for example, the occurrence of planned or unplanned electromagnetic signals received by the hearing device (including audio and / or control signals), or other properties of the current environment that are different from acoustic);

[0060] b) The current acoustic situation (input level, feedback, etc.);

[0061] c) The user's current mode or state (motion, temperature, cognitive load, etc.);

[0062] d) The current mode or state of the hearing device and / or another device communicating with the hearing device (selected program, time elapsed since the last user interaction, etc.).

[0063] The classification unit may be based on or include a neural network, such as a trained neural network.

[0064] The hearing device may include an acoustic (and / or mechanical) feedback control or echo cancellation system. Acoustic feedback occurs because the output speaker signal from an audio system that provides amplification of the signal picked up by the microphone returns to the microphone via acoustic coupling parts through air or other media. The part of this speaker signal that returns to the microphone is amplified again by the audio system before it reappears at the speaker and then returns to the microphone again. As this cycle continues, when the audio system becomes unstable, the acoustic feedback effect becomes audible, such as unnatural signals or even worse, whistling. This problem typically occurs when the microphone and the speaker are placed close together, for example, in a hearing aid or other audio systems. Some other typical situations with feedback problems include telephony, broadcast systems, headphones, audio conferencing systems, etc. Adaptive feedback cancellation has the ability to track changes in the feedback path over time. It estimates the feedback path based on a linear time-invariant filter, but the filter weights are updated over time. The filter update can be calculated using a stochastic gradient algorithm, including certain forms of the least mean square (LMS) or normalized LMS (NLMS) algorithms. They all have the property of minimizing the mean square of the error signal, and NLMS additionally normalizes the filter update with respect to the square of the Euclidean norm of some reference signal. The feedback cancellation system may include a feedback detection / estimation unit. The hearing device may be configured to switch (such as fade) between microphone signals based on the estimated amount of feedback (as described in the present invention).

[0065] The hearing device may also include other suitable functions for the applications involved, such as compression, noise reduction, etc.

[0066] The hearing device may include a listening device such as a hearing aid, a hearing instrument, for example, a hearing instrument adapted to be located at the user's ear or fully or partially located in the ear canal, such as headphones, earphones, ear protection devices, or a combination thereof. The hearing aid system may include a horn loudspeaker (including a plurality of input transducers and a plurality of output transducers, for example, used in an audio conferencing scenario), for example, including a beamforming filter unit, for example, providing multiple beamforming capabilities.

[0067] Application

[0068] On the one hand, there is provided an application of the hearing device as described above, detailed in the "Detailed Description" section and defined in the claims. An application in a system including audio distribution can be provided, such as a system including a microphone and a speaker that are close enough to each other to cause feedback from the speaker to the microphone during user use. Applications in a system including one or more hearing aids (such as hearing instruments), headsets, earphones, active ear protection systems, etc. can be provided, such as uses in hands-free telephone systems, teleconference systems (such as including public address systems), broadcast systems, karaoke systems, classroom amplification systems, etc.

[0069] Method

[0070] On the one hand, the present application further provides a method of operating a hearing device such as a hearing aid, the hearing device being adapted to be located at or in a user's ear or being adapted to be fully or partially implanted in the user's head. The method includes:

[0071] - Providing at least two input audio data streams, each input audio data stream including a mixture of a target signal component from a target sound source and a noise component from one or more noise sources;

[0072] - Receiving at least two input audio data streams;

[0073] - Mixing at least two input audio data streams or their processed versions;

[0074] - Providing a processed input signal based on the mixing;

[0075] - Providing an output stimulus perceptible by the user as sound based on the processed input signal or its processed version.

[0076] The method further includes:

[0077] - Processing the noise components of at least two input audio data streams or their processed versions to reduce or avoid unnatural signals caused by the mixing in the processed input signal.

[0078] The method further includes:

[0079] - Balancing the noise components of at least two input audio data streams in the processed input signal.

[0080] When appropriately replaced by corresponding processes, some or all of the structural features of the device described above, detailed in the "Detailed Description", or defined in the claims can be combined with the implementation of the method of the present invention, and vice versa. The implementation of the method has the same advantages as the corresponding device.

[0081] Computer-readable medium

[0082] The present invention further provides a tangible computer-readable medium storing a computer program including program code, which, when the computer program runs on a data processing system, causes the data processing system to perform at least some (such as most or all) of the steps of the methods described above, detailed in the "Detailed Description" and defined in the claims.

[0083] By way of example and not limitation, the foregoing tangible computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store or execute the required program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk includes compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy disk and Blu-ray disk, where these disks typically magnetically reproduce data while they can optically reproduce data using lasers. Other storage media include those stored in DNA (such as in synthetic DNA strands). Combinations of the above disks should also be included within the scope of computer-readable media. In addition to being stored on a tangible medium, a computer program may also be transmitted via a transmission medium such as a wired or wireless link or network such as the Internet and loaded into the data processing system to run at a location different from the tangible medium.

[0084] Computer program

[0085] In addition, the present application provides a computer program (product) including instructions, which, when the program runs on a computer, causes the computer to perform the methods (steps) described above, detailed in the "Detailed Description" and defined in the claims.

[0086] Data processing system

[0087] On the one hand, the present invention further provides a data processing system including a processor and program code, which causes the processor to perform at least some (such as most or all) of the steps of the methods described above, detailed in the "Detailed Description" and defined in the claims.

[0088] Hearing system

[0089] On the other hand, there is provided a hearing system including the hearing device and the auxiliary device described above, detailed in the "Detailed Description" and defined in the claims.

[0090] The hearing system may be adapted to establish a communication link between the hearing device and the auxiliary device so that information (such as control and status signals, possibly audio signals) can be exchanged or forwarded from one device to another.

[0091] The auxiliary device may include a remote control, a smart phone, or other portable or wearable electronic devices such as a smart watch, etc.

[0092] The auxiliary device may be or include a remote control for controlling the functions and operations of the hearing device. The functions of the remote control may be implemented in a smart phone, which may run an APP enabling the functions of controlling the audio processing device via the smart phone (the hearing device includes an appropriate wireless interface to the smart phone, e.g., based on Bluetooth or some other standardized or proprietary solution).

[0093] The auxiliary device may be or include an audio gateway device, which is adapted to receive multiple audio signals (e.g., from entertainment devices such as TVs or music players, from telephone devices such as mobile phones, or from computers such as PCs) and is adapted to select and / or combine appropriate signals (or signal combinations) from the received audio signals for transmission to the hearing device.

[0094] The auxiliary device may be or include a wireless microphone, such as a desktop microphone or an on-chip microphone.

[0095] The auxiliary device may be or include another hearing device. The hearing system may include two hearing devices adapted to implement a binaural hearing system such as a binaural hearing aid system.

[0096] On the other hand, there is provided a hearing system including first and second hearing devices as described above, detailed in the "Detailed Description" or defined in the claims. The first and second hearing devices may be adapted to be located at or in the user's left and right ears or to be fully or partially implanted in the head at the user's left and right ears, and the first and second hearing devices are configured to exchange information therebetween.

[0097] Speakerphone

[0098] On the one hand, the present application further provides a speakerphone. The speakerphone includes a plurality of microphones configured to pick up sounds from the environment of the speakerphone and includes a hybrid processor as described above, detailed in the "Detailed Description" or defined in the claims. The hybrid processor is adapted to provide a processed input signal, which is transmitted to another device or system for further processing and / or presented to one or more distant users. The speakerphone is also configured to play sounds received from a distant sound source for perception in the environment of the speakerphone.

[0099] The speakerphone may include:

[0100] - A sound input signal path, including

[0101] -- An input unit providing at least two input audio data streams, each input audio data stream including a mixture of a target signal component from a target sound source and a noise component from one or more noise sources;

[0102] -- A mixing processor for receiving at least two input audio data streams, for mixing at least two input audio data streams or processed versions thereof, and for providing a processed input signal based thereon;

[0103] -- An output unit including a transmitter for transmitting the processed input signal or a processed version thereof to another device or system; and

[0104] - A speaker signal path including

[0105] -- A receiver for receiving an audio data stream from another device or system;

[0106] -- A signal processor for processing the audio data stream and providing a processed signal; and

[0107] -- A speaker for providing an acoustic sound signal representing the processed signal.

[0108] Thus, a speakerphone incorporating an adaptive mixing scheme according to the present invention can be implemented.

[0109] When appropriately adjusted, some or all of the structural features of the hearing device described above, detailed in the "Detailed Description" or defined in the claims, can be combined with the implementation of the speakerphone. The implementation of the speakerphone has the same advantages as the corresponding hearing device.

[0110] The input unit of the speakerphone may include a plurality of microphones, such as more than two, such as more than three, each microphone providing an input data stream representing sound in the environment. Based on two microphones, different beamformers can be generated, which listen in different directions. The input unit of the speakerphone may include a beamformer filtering unit that receives a plurality of input data streams and is configured to provide at least two spatially filtered (beamformed) signals directed at at least two target sound sources in the speakerphone environment. The plurality of microphones can be divided into microphone subgroups. Each subgroup may include at least two microphones. A given subgroup may include at least one microphone that does not form another microphone subgroup. A reference microphone can be determined among the plurality of microphones. All microphone subgroups may include the reference microphone (designated among the microphones of the subgroup). All microphone subgroups may include the same reference microphone. The speakerphone may be configured to fade between at least two spatially filtered signals and transmit the (resulting) processed input signal (or a further processed version thereof, such as a post-filtered version) to another device or system. The mixing unit may be configured to fade between two spatially filtered signals without changing the background noise level of the speakerphone environment in the processed input signal (or a further processed version thereof) that is transmitted to another device or system (such as a remote receiving listener of a communication device).

[0111] APP

[0112] On the other hand, the present invention also provides a non-transitory application called APP. The APP includes executable instructions configured to run on an auxiliary device to implement a user interface for a hearing device or a hearing system as described above, detailed in the "Detailed Description" and defined in the claims. In an embodiment, the APP is configured to run on a mobile phone such as a smart phone or another portable device enabling communication with the hearing device or the hearing system.

[0113] Definition

[0114] In this specification, a "hearing device" refers to a device suitable for improving, enhancing, and / or protecting a user's auditory ability, such as a hearing aid, for example, a hearing instrument or an active ear protection device or other audio processing device, which is achieved by receiving an acoustic signal from the user's environment, generating a corresponding audio signal, possibly modifying the audio signal, and providing the possibly modified audio signal as an audible signal to at least one ear of the user. A "hearing device" also refers to a device suitable for electronically receiving an audio signal, possibly modifying the audio signal, and providing the possibly modified audio signal as an audible signal to at least one ear of the user, such as a headset or an earphone. The audible signal can be provided, for example, in the form of an acoustic signal radiated into the user's outer ear, an acoustic signal transmitted as a mechanical vibration through the bone structure of the user's head and / or through parts of the middle ear to the user's inner ear, and an electrical signal transmitted directly or indirectly to the user's cochlear nerve.

[0115] The hearing device can be configured to be worn in any known manner, such as as a unit worn behind the ear (with a tube for guiding the radiated acoustic signal into the ear canal or with an output transducer arranged close to or in the ear canal, such as a speaker), as a unit arranged entirely or partially in the auricle and / or the ear canal, as a unit connected to a fixed structure implanted in the skull, such as a vibrator, or as a connectable or entirely or partially implantable unit, etc. The hearing device can include a single unit or several units that communicate electronically with each other. The speaker can be provided in a housing together with other components of the hearing device, or it can itself be an external unit (possibly combined with a flexible guiding element such as a dome-shaped element).

[0116] More generally, a hearing device includes an input transducer for receiving an acoustic signal from the user's environment and providing a corresponding input audio signal and / or a receiver for receiving the input audio signal electronically (i.e., wired or wirelessly), a (usually configurable) signal processing circuit for processing the input audio signal (such as a signal processor, e.g., including a configurable (programmable) processor, e.g., a digital signal processor), and an output unit for providing an audible signal to the user based on the processed audio signal. The signal processor may be adapted to process the input signal in the time domain or in a plurality of frequency bands. In some hearing devices, an amplifier and / or a compressor may form part of the signal processing circuit. The signal processing circuit typically includes one or more (integrated or separate) storage elements for executing programs and / or for storing parameters used (or potentially used) in the processing and / or for storing information suitable for the function of the hearing device and / or for storing information used, for example, in connection with an interface to the user and / or to a programming device (such as processed information, e.g., provided by the signal processing circuit). In some hearing devices, the output unit may include an output transducer, such as a loudspeaker for providing an air-conducted acoustic signal or a vibrator for providing a structure-borne or fluid-borne acoustic signal. In some hearing devices, the output unit may include one or more output electrodes for providing an electrical signal (such as a multi-electrode array for electrically stimulating the cochlear nerve). The hearing device may include a horn loudspeaker (including a plurality of input transducers and a plurality of output transducers, e.g., used in an audio conferencing scenario).

[0117] In some hearing devices, the vibrator may be adapted to transmit a structure-borne acoustic signal to the skull transcutaneously or percutaneously. In some hearing devices, the vibrator may be implanted in the middle ear and / or the inner ear. In some hearing devices, the vibrator may be adapted to provide a structure-borne acoustic signal to the middle ear bones and / or the cochlea. In some hearing devices, the vibrator may be adapted to provide a fluid-borne acoustic signal to the cochlear fluid, for example, through the oval window. In some hearing devices, the output electrodes may be implanted in the cochlea or on the inner side of the skull and may be adapted to provide an electrical signal to the hair cells of the cochlea, one or more auditory nerves, the auditory brainstem, the auditory midbrain, the auditory cortex, and / or other parts of the cerebral cortex.

[0118] Hearing devices such as hearing aids can be adapted to the needs of a particular user such as hearing impairment. The configurable signal processing circuit of the hearing device may be adapted to apply compression amplification that varies with frequency and level to the input signal. Customized gain (amplification or compression) that varies with frequency and level can be determined during the fitting process by a fitting system based on the user's hearing data such as an audiogram using fitting principles (such as adapted to speech). The gain that varies with frequency and level can be embodied, for example, in processing parameters, e.g., uploaded to the hearing device via an interface to a programming device (fitting system) and used by a processing algorithm executed by the configurable signal processing circuit of the hearing device.

[0119] "Hearing system" means a system including one or two hearing devices. "Binaural hearing system" means a system including two hearing devices and adapted to provide audible signals to both ears of a user in a coordinated manner. A hearing system or binaural hearing system may also include one or more "auxiliary devices" which communicate with the hearing devices and affect and / or benefit from the functions of the hearing devices. The auxiliary device may be, for example, a remote control, an audio gateway device, a mobile phone (such as a smart phone) or a music player. A hearing device, hearing system or binaural hearing system may be used, for example, to compensate for the loss of auditory ability of a hearing-impaired person, enhance or protect the auditory ability of a person with normal hearing and / or transmit an electronic audio signal to a person. A hearing device or hearing system may form part of or interact with, for example, a broadcast system, an active ear protection system, a hands-free telephone system, an automotive audio system, an entertainment (such as karaoke) system, a teleconference system, a classroom amplification system, etc.

[0120] Embodiments of the present invention may be used, for example, in hearing aid applications, such as hearing aids configured to communicate with another device, such as a binaural hearing aid system. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Various aspects of the present invention will be best understood from the following detailed description taken in conjunction with the accompanying drawings. For clarity, these drawings are schematic and simplified, showing only the details necessary for understanding the present invention and omitting other details. Throughout the specification, the same reference numerals are used for the same or corresponding parts. Each feature of each aspect may be combined with any or all features of other aspects. These and other aspects, features and / or technical effects will be apparent from and elucidated in conjunction with the following drawings, in which:

[0122] Figure 1A and 1B shows a situation where an acoustic and / or wirelessly propagated audio data stream (or a mixture thereof) is received in a hearing device, Figure 1A shows a side view of a user wearing a hearing device including a BTE part and an ITE part at the right ear, and Figure 1B shows a front view of a user wearing hearing devices at the left and right ears;

[0123] Figure 2A shows a block diagram of a first embodiment of a hearing device according to the present invention;

[0124] Figure 2B shows a processor for mixing first and second source signals x1(n) and x2(n), which are modified by time-varying gains α1 and α2 to a processed input signal y(n);

[0125] Figure 2C shows a block diagram of a second embodiment of a hearing device according to the present invention;

[0126] Figure 2D Shows a block diagram of a third embodiment of a hearing device according to the present invention;

[0127] Figure 2E Shows a block diagram of a fourth embodiment of a hearing device according to the present invention;

[0128] Figure 3 Shows an input stage of a hearing device including an input unit and an adaptive mixing unit according to the present invention, which provides a fade between two microphone signals having different noise variances with a fade factor α, such that an output y(n) having an unchanged noise (and target) level (before and after the fade) is provided;

[0129] Figure 4 Shows Figure 3 the input stage shown in, but assuming similar noise variances at each microphone;

[0130] Figure 5A Shows an in-ear receiver hearing device according to an embodiment of the present invention;

[0131] Figure 5B Shows a completely in-ear hearing device according to an embodiment of the present invention;

[0132] Figure 6A Shows an embodiment of a hearing system such as a binaural hearing aid system according to the present invention;

[0133] Figure 6B Shows an auxiliary device configured to execute an APP for a user interface of a hearing device or system, from which a running mode and a currently appropriate sound input can be selected;

[0134] Figure 7 Schematically shows a speakerphone including a plurality of microphones and a plurality of beamformers, which is configured to focus on a plurality of different target speakers in the environment around the speakerphone and enable an adaptive fade between spatially filtered signals as described in the present invention;

[0135] Figure 8 Shows an estimator for estimating the noise variances of at least two input audio data streams before audio stream mixing; and

[0136] Figure 9 Schematically shows an exemplary fade process between two input audio data streams having different target and noise levels.

[0137] The further scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various other embodiments will be apparent to those skilled in the art from the following detailed description. Detailed Description

[0138] The following detailed description presented in conjunction with the accompanying drawings is used as a description of various different configurations. The detailed description includes specific details for providing a thorough understanding of the various different concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. Several aspects of the apparatus and method are described by means of various different blocks, functional units, modules, elements, circuits, steps, processes, algorithms, etc. (collectively referred to as "elements" or "units"). Depending on the particular application, design constraints, or other reasons, these elements (or units) may be implemented using electronic hardware, computer programs, or any combination thereof.

[0139] Electronic hardware may include a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various different functions described in this specification. A computer program should be construed broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, programs, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0140] This application relates to the field of hearing devices such as hearing aids.

[0141] Figure 1A and 1B illustrates a scenario in which an acoustic and / or wirelessly propagated audio data stream (or a mixture thereof) is received in a hearing device, Figure 1A illustrates a side view of a user wearing a hearing device including a BTE portion and an ITE portion in the right ear, and Figure 1B illustrates a front view of a user wearing hearing devices in both the left and right ears.

[0142] In the case where a hearing aid user wants to listen to a mixture of audio data streams, the hearing aid should facilitate a natural perception of sound without any unnatural signals imposed by time-varying sound source balancing and / or fading. Below, we consider the mixture of two noisy speech sound streams.

[0143] Figure 1AShows a hearing device HD1 located at the ear of user U (here the right ear). The hearing device includes a BTE part adapted to be located at or behind the user's ear (pinna) and an ITE part adapted to be located at or in the user's ear canal. The BTE part includes an input unit. The input unit includes two microphones (BTE1, BTE2) for picking up sound from the user's environment and two wireless audio receivers for wirelessly receiving audio from a corresponding audio transmitter (here a pick-up coil ("pick-up coil") (or other receiver based on near-field communication) and an RF receiver ("wireless") (e.g., based on Bluetooth or similar technology)). Figure 1A The hearing device can be a stand-alone hearing device or (as shown here) form part of a binaural hearing system such as a binaural hearing aid system (as Figure 1B shown in). Figure 1B Shows a binaural hearing system including first and second hearing devices (HD1, HD2) (such as hearing aids) adapted to be located at or in the user's right and left ears. Figure 1B Shows that the ITE part of each hearing device (HD1, HD2) includes a microphone (ITE) located at the end of the ITE part facing the environment and a speaker (receiver) located at the end of the ITE part facing the eardrum. The speaker is thus configured to play into the residual cavity between the ITE part and the eardrum. The first and second hearing devices (HD1, HD2) thus belong to the receiver-in-the-ear (RITE) type and include three microphones, namely a microphone (referred to as the ITE microphone) located at the ear canal opening when the hearing device is on the user and two microphones (referred to as BTE microphones) at or behind the pinna. Such a type of hearing aid can have the advantage of being able to utilize the benefits of the pinna (ITE microphone, maintaining spatial cues) while providing sound to the user based on the BTE microphones if necessary when there is a risk of howling at the ITE microphone. The hearing device can for example include one or more microphones located elsewhere on the user's head or body, for example in the pinna, for example in the cochlea or in the ear canal, for example near the eardrum (for example to pick up sound from the residual cavity near the eardrum).

[0144] Assume that the audio signal is provided as a sub-band signal (time-frequency domain), for example a time-domain signal that has been transformed to the time-frequency domain using an analysis filter bank and transformed back to the time domain using a synthesis filter bank before being presented to the user (see for example Figure 2C units A and S in). The input unit can for example include at least two audio inputs. The audio inputs can for example include two microphones and / or two direct (wireless or wired) audio receivers or a mixture of a microphone and a direct audio receiver. The input unit IU (see for example Figure 2A, 2C, 2D, 2E) may also include a plurality of analog-to-digital converters (e.g., one analog-to-digital converter for each analog audio input) for converting the analog audio inputs into sampled (digital) electrical input signals. The input unit may also include a plurality of analysis filter banks ( Figure 2C A in), for providing the electrical input signal as a plurality of sub-band signals in a time-frequency representation, each sub-band signal corresponding to a sub-range of a frequency range representing the audio frequency of the audio signal involved (e.g., up to 20 kHz or lower, e.g., between 0 and 10 kHz). The output unit OU (e.g., see Figure 2A , 2C, 2D, 2E) may for example include a synthesis filter bank ( Figure 2C S in), for converting the sub-band signals into a time-domain signal for presentation to the user via an output transducer of the output unit (such as a loudspeaker or a vibrator of a bone conduction hearing device). The output unit may also include a digital-to-analog converter (or other drive circuitry) as appropriate. The output unit may also include an antenna and transmitter circuitry for passing the audio output signal to another element or device (such as another hearing device, if suitable for the application involved).

[0145] In this specification, the noisy speech mixture signal y(n) = y s (n) + y v (n) is a mixture of the noisy speech signals x1(n) = s1(n) + v1(n) and x2(n) = s2(n) + v2(n), where s(n) and v(n) refer to the speech component and the noise component respectively, and n represents time. The speech and noise components are assumed to be uncorrelated.

[0146] The variance of the noise component y v (n) = v1(n) + v2(n) is given by:

[0147]

[0148] where Re{X} refers to the real part of the complex number X, and Cov(v1, v2) refers to the covariance between v1 and v2. This expression is valid for both correlated and uncorrelated noise. If the noise components (v1, v2) are uncorrelated, the last part (2Re{Cov(v1, v2)}) is zero. In a hearing aid application, the signals are typically modified before mixing, such as for signal equalization / tapering, noise reduction, etc. The mixture signal is given by:

[0149] y(n) = α1x1(n) + α2x2(n)

[0150] where α1 and α2 are gain factors. These gain factors may be time-varying factors.

[0151] The noise variance of the mixture signal, including the gain, is given by:

[0152]

[0153] Since the noise and the speech component (presumably) are uncorrelated, a similar relationship can be found for the speech component. In practical applications, the noise and speech variances are typically time-varying estimators that vary with frequency and are found using a level estimator controlled by a voice activity detector (VAD). Specifically, it can be detected whether the noise between microphones is uncorrelated (e.g., based on the elements of the covariance matrix between microphones).

[0154] Mixing or fading of the source signals can result in annoying and audible unnatural signals when the noise backgrounds of the signals are not equal. To overcome this problem, the noise component in the second source can be modified to avoid unnatural signals.

[0155] Figure 2A A block diagram of a hearing device, such as a hearing aid, according to a first embodiment of the present invention is shown. The hearing device is, for example, adapted to be located at or in the user's ear or is adapted to be fully or partially implanted in the user's head. The hearing device includes, for example, an input unit IU that provides at least two input audio data streams (x1, x2) represented in sub-bands. Each input audio data stream includes a mixture of a target signal component and a noise component. The hearing device further includes a mixing processor PRO that receives at least two input audio data streams (x1, x2) or a processed version thereof from the input unit IU, mixes at least two input audio data streams or a processed version thereof, and provides a processed input signal y based on the mixing. The hearing device further includes an output unit OU configured to provide an output stimulus that can be perceived as sound by the user based on the processed input signal or a (further) processed version thereof. The mixing processor PRO is thus connected to the input unit IU and the output unit OU. Thereby, the forward (audio signal processing) path of the hearing device is implemented.

[0156] Figure 2B A processor PRO for mixing first and second source signals x1(n) and x2(n) (first and second input audio streams) is shown, which is modified by time-varying gains α1 and α2 to a processed input signal y(n). The mixing can, for example, include a fade from the first signal (first input audio stream) to the second signal (second input audio stream). In the present application, the functional unit that processes the mixing is referred to as an adaptive mixing unit ADM. In Figure 2B the processor PRO consists of one functional unit, namely the adaptive mixing unit ADM. Although this is not generally necessary (as shown in Figure 2D 、 2E ).

[0157] In Figure 2BIn an embodiment, prior to mixing, the second source may be modified by a compensation gain β. The goal of doing this is to provide noise source balance (by equalizing the noise levels) during (and after) mixing (such as fading).

[0158] The compensation gain β(n) is applied to the second source x2(n), see Figure 2B .

[0159] y(n) = α1x1(n) + βα2x2(n)

[0160] This means that the noise variance of the mixed signal, including the gain, is given by:

[0161]

[0162] We now normalize with the first input noise variance :

[0163]

[0164] The corrected gain β can now be found by choosing the desired output variance. For example, the desired output noise variance is chosen to be equal to the first input noise variance, i.e., Substituting this into the previous equation gives:

[0165]

[0166] This gives two solutions, one of which is negative. A negative β means mixing by subtraction, which we do not allow. Thus, we only consider the solution where β is positive, i.e.,

[0167]

[0168] where

[0169] and

[0170] The corrected gain β can be applied to time-frequency units that have been classified as noise only by, for example, a used voice activity detector (VAD) (time-frequency units that are noise only are units for which the VAD has indicated the absence of voice (such as speech)). The corrected gain β can be determined iteratively (such as gradient-based minimization of a quadratic polynomial). This avoids the square root.

[0171] The same principle can be applied to the speech component for target source balance. However, it may not be desirable to modify the spectral shape of the second input speech component to match the first input speech component. In that case, cross-frequency constraints can be applied. An exemplary constraint can be to maintain the user's loudness perception during fading. This constraint can be used to determine β for a given input audio stream.

[0172] In a practical hearing aid application, it is desirable to avoid operations such as squaring, square rooting, and division (due to their computational complexity (power limitations)). Most operations can be performed in the logarithmic domain, such that multiplication and division can be implemented using addition and subtraction, respectively. Through a mapping function or a look-up table, any other operation can be efficiently approximated.

[0173] Figure 2C A block diagram of a hearing device according to a second embodiment of the present invention is shown. Figure 2C The hearing device of Figure 2A is shown in combination with 2B but where the input unit IU, the adaptive mixing unit ADM, and the output unit OU are described in further detail. The input unit IU includes first and second input transducers (IT1, IT2), each input transducer providing a (preferably digitized) input audio stream as a (full-band) time-domain signal. The input unit IU also includes a corresponding analysis filter bank A for converting the two input audio streams into first and second sub-band signals, respectively, thereby providing the first and second input audio streams (x1, x2) in a time-frequency representation. The adaptive mixing unit ADM receives the first and second input audio streams (x1, x2) and applies time-varying gains (weights) α1, α2, β to the first (α1) and second (α2, β) input audio streams (x1, x2) to provide modified first and second input audio streams (x1α1 and x2α2β, respectively) and adds the modified audio streams to provide a processed input signal y (y = x1α1 + x2α2β), as Figure 2B shown in Figure 3 and described above. The adaptive mixing unit ADM also includes a weighting unit WGT and a noise variance estimation unit NVE. The weighting unit WGT is configured to determine first and second time-varying weights (α1, α2) for application to the first and second input audio streams (x1, x2), respectively. The weights (α1, α2) can be determined, for example, from a (time-varying) mixing function (e.g., a ramp function, see 4 , e.g., stored in the memory of the hearing device) based on a triggering input signal MT (e.g., from a user interface or based on the output of one or more detectors (such as a voice activity detector)) and, for example, based on the nature of the first and second input audio streams (such as modulation or noise, such as SNR). The triggering input signal MT can indicate, for example, the start of a transition from one audio input stream (e.g., provided by a microphone or a beamformer or a direct audio input) to another audio input stream, see Figure 6B。The noise variance estimation unit NVE is configured to determine a compensation gain β applied to the second input audio stream x2. The compensation gain β can be determined based on the properties of the first and second input audio streams (such as modulation or noise, e.g., SNR) and the current values of the time-varying weights (α1, α2), and non-necessarily, also based on the mixing trigger input signal MT, as described above.

[0174] Figure 2D A block diagram of a hearing device according to a third embodiment of the present invention is shown. Figure 2D The hearing device may include units combined with Figure 2A 、 2B and those described in 2C. Additionally, the processor PRO includes a hearing aid processor HAG for applying additional processing algorithms to the signal y' provided by the adaptive mixing unit ADM. The hearing aid processor HAG may be adapted, for example, to compensate for the user's hearing impairment, for example, by applying a compression amplification algorithm to the signal in the forward path, such as the processed input signal y' (the mixed or faded signal) or a signal derived therefrom. The customized compression amplification algorithm may be configured to apply a gain that varies with frequency and level according to the specific needs of the user. Additionally or alternatively, other processing algorithms may also be applied to the signal y', such as a calibration signal (such as a post-filter). The hearing aid processor HAG thus provides a processed signal y, which is fed to the output unit OU. In Figure 2D the embodiment, the processor PRO includes an adaptive mixing unit ADM and a hearing aid processor HAG. Additional functional units may also be included in the processor PRO, such as feedback control, etc.

[0175] Figure 2E A block diagram of a fourth embodiment of a hearing device according to the present invention is shown. Figure 2E The hearing device, such as a hearing aid, includes functional units combined with Figure 2D those described. Additionally, the input unit includes a beamforming filter unit BF. The beamforming filter unit BF includes two beamformers configured to provide corresponding (different) beamformed signals (x BF1 , x BF2 ) based on the first and second input signals (x1, x2) from the first and second input transducers (IT1, IT2) of the input unit IU, such as microphones. The first and second beamformed signals (x BF1 , x BF2 ) are provided, for example, as (different) linear combinations of the first and second input signals (x1, x2), such as x BF1 = C 11 x1 + C 12 x2 and x BF2 = C 21 x1 + C 22x2, where the filter weights C of the first beamformer 11 , C 12 and the filter weights C of the second beamformer 21 , C 22 (generally) are complex (fixed or adaptively determined, usually frequency-varying) parameters. Figure 2E An embodiment of may for example correspond to a fade between two beamformed signals (from two different spatial locations), for example controlled by voice activity detection in the two signals (“select the signal including voice”). Such a scenario may for example include a fixed beamformer, such as in an automotive scenario with a possible sound source in a fixed position (e.g., to the side or behind or in front of a user wearing a hearing device). Alternatively, the scenario may be a multi-speaker scenario, where the direction of the main speaker is adaptively determined and may fade between them for example based on voice activity in the beamformed signals. Figure 2E An embodiment of is shown to include two input transducers (IT1, IT2), but may include more than two such as three, four or more input transducers. Adaptive mixing may for example be based on two beamformed signals generated from more than two electrical input signals or based on different (or partially overlapping) electrical input signals. In the example of three input transducers, one input transducer may for example be defined as a reference input transducer, whose electrical input signal is used as an input to two beamformers, and the other two electrical input signals are used in their respective beamformers.

[0176] Alternatively, we may add uncorrelated noise to the mixture of uncorrelated noise such that the fade from one signal to the other becomes inaudible. By adding uncorrelated noise, the same behavior of the uncorrelated noise source can be simulated at two microphones. The cost of having more similar behavior at two microphones is that more noise will be added to the microphone with the least noise.

[0177] The above-mentioned beamformed signals x BF1 and x BF2 's noise characteristics can be equalized by generating corresponding signals

[0178] Y1 = (1 - α1)x ref + α1x BF1

[0179] Y2 = (1 - α2)x ref + α2x BF2

[0180] where x refis an input data stream for a common reference microphone among a plurality of microphones. By adding the scaled reference microphone signal to each beamformed signal (where the beamformer is directed towards different target directions), similar noise characteristics can be obtained in two signals Y1 and Y2, thereby making the transition between the two signals less audible.

[0181] When fading between two "improved beamformed signals" Y1, Y2, the processed input signal can be expressed as

[0182] Y = λY1+(1 - λ)Y2

[0183] where λ is a fade parameter for fading between two "improved beamformed signals" Y1, Y2, and where α1 and α2 are determined such that the noise levels in Y1 and Y2 are comparable. More than two beamformed signals can also be used, such as Y1, …, Y N . In this case, α1, …, α N are chosen such that the background noise levels in each beamformed signal are approximately the same. The different α values can be adaptively determined over time and frequency.

[0184] The proposed solution is shown in Figure 3 which shows fading between two microphone signals in a hearing device with a fade factor α such that an output y(n) with an unchanged noise (and target) level is provided while fading from one microphone signal to the other microphone signal. A possible fade function α(t) (when fading from a microphone signal x1 from microphone M1 to a microphone signal x2 from microphone M2) is shown in the middle of Figure 3 (rectangular box). The fade function is shown as a piecewise linear function that varies (over time) from a maximum value (such as 1) to a minimum value (such as 0) across a time period Δt. Other monotonic progressions of this function can be foreseen, such as an S-shaped (or S-like) function, or a linear fade in the logarithmic domain, etc. The time period Δt across which the transition occurs can be changed according to the specific application or listening situation. The time period Δt can be in the range of 0.5 to 5 s, for example.

[0185] We assume a system with at least two input signals. It can be, for example, two microphone signals (as shown in Figure 3 ), two pick-up coil (or other wirelessly received) signals, one microphone signal and one pick-up coil (or other wirelessly received) signal such as an audio signal for TV streaming, or other signals. Each signal consists of two parts: the desired target signal (s1, s2), which is assumed to be correlated (preferably the same); and some additional uncorrelated noise (v1, v2). Referring to Figure 3 , each input x i (i = 1, 2) consists of a target component s i(n) and the noise component v i (n), where n is the time index. We assume that the target is similar at the two inputs, while the noise components are additive and uncorrelated, with noise variances and This is illustrated schematically by two time segments of two (time-varying) microphone signals x1 and x2. The two time segments are inserted between the microphones (M1, M2) and the corresponding combining unit X at Figure 3 The noise variance at each microphone may be different, for example due to adjustment to make the target signals (levels) similar at each microphone. The noise at each microphone may include, for example, uncorrelated noise (such as predominantly), for example microphone noise and / or wind noise. The microphone noise (for each individual microphone) may be given before system operation, for example by measurement or estimation (e.g., based on the microphone specifications) and, for example, stored in a memory.

[0186] Sometimes, it is desirable to fade from one microphone signal to another, for example when one microphone has more feedback than the other (e.g., in the Figure 5A or microphone configuration shown in 5B, where one microphone is more susceptible to feedback from the output transducer than the other). Since the uncorrelated and correlated parts of the input signals do not mix in a similar way, when fading (from the first microphone M1 to the second microphone M2), it is proposed to add uncorrelated noise (v3) to the system when fading from one microphone to another to obtain an "unaltered signal" (in terms of noise and / or target signal or total signal level).

[0187] In other words, in an embodiment, we aim to fade between two microphone signals x1 and x2 with a fade factor such that an output y(n) with an unaltered target signal level is obtained, i.e.,

[0188] y(n) = αx1+(1 - α)x2

[0189] Similarly, we aim to keep the noise level constant by adding some additional noise v3 equal to the maximum noise level of the two microphone levels, i.e.,

[0190]

[0191] To estimate the noise variance of the additional random variable, we isolate the additional noise in the above equation, i.e.,

[0192]

[0193] Assume that v1 is a Gaussian random variable with a known variance and v2 is a Gaussian random variable with a known variance For Gaussian random variables (such as microphone noise and wind noise), we can generate and add a third Gaussian random variable v3 with an adaptive variance that varies with and α. The result of the proposed method is that the output noise level corresponds to the microphone signal with the highest noise variance.

[0194] In the

[0195] scheme, the input stage of the hearing device includes an input unit IU and an adaptive mixing unit ADM that provides a gradual transition between two microphone signals with different noise variances. The adaptive mixing is performed with a gradual factor α such that an output y(n) with an unchanged noise level (before and after the transition) is provided. A hearing device processor for applying one or more processing algorithms (such as a noise reduction algorithm (e.g., including post-filtering (single-channel noise reduction)) and / or a compression amplification algorithm, etc.) can be included downstream of the input stage (e.g., see Figure 3 the hearing aid processor HAG in Figure 2D , 2E ). In addition, a voice activity detector can be used to qualify the microphone signals. The target signal component may or may not be equalized (equalization of the noise component is more important in both cases).

[0196] Figure 4 shows the Figure 3 input stage shown in Figure 4 , but where it is assumed that there is a similar noise variance at each microphone. However, even if the noise variances are the same, we must add noise during the transition to maintain a stable noise level. Figure 3 The input stage of 2 is similar to the input stage of Figure 2E , except that the two microphones M1, M2 exhibit the same noise variance σ

[0197] Figure 5A shows an in-the-ear receiver hearing device according to an embodiment of the present invention.

[0198] Figure 5B Shows a completely in-the-ear hearing device according to an embodiment of the present invention.

[0199] Figure 5A Shows a BTE / RITE hearing device (BTE = "behind-the-ear", RITE = "in-the-ear receiver") according to a first embodiment of the present invention. An exemplary hearing device HD such as a hearing aid belongs to a specific type (sometimes called in-the-ear receiver type or RITE type), including a BTE part (BTE) adapted to be located at or behind the user's ear and an ITE part (ITE) adapted to be located in or at the user's ear canal and including a receiver (speaker, SPK). The BTE part and the ITE part are connected (such as electrically connected) by a connecting element IC and internal wiring in the ITE and BTE parts (for example, see the wiring Wx in the BTE part). Alternatively, the connecting element may be constituted entirely or partially by a wireless link between the BTE part and the ITE part. Of course, other styles may also be used, such as including a custom ear mold adapted to the user's ear and / or ear canal (for example, see Figure 5B ).

[0200] In Figure 5A the hearing device embodiment, the BTE part includes an input unit having two input transducers (such as microphones) (M BTE1 , M BTE2 ), each input transducer for providing an electrical input audio signal representing an input sound signal (S BTE ) (derived from the sound field S around the hearing device). The input unit also includes two wireless receivers (WLR1, WLR2) for providing corresponding directly received auxiliary audio and / or control input signals (and / or enabling the transmission of audio and / or control signals to other devices such as a remote control or a processing device, or a telephone, or another hearing device). The hearing device HD includes a substrate SUB on which a plurality of electronic components are mounted, including a memory MEM, which, for example, stores different hearing aid programs (such as user-specific data such as data related to a hearing diagram, or parameter settings derived therefrom such as parameter settings defining the aforementioned (user-specific) program, or other parameters of an algorithm such as beamformer filter weights, and / or gradient parameters) and / or a hearing aid configuration such as an input source combination (M BTE1 , M BTE2 (M ITE ), WLR1, WLR2), for example, optimized for a plurality of different listening situations. In a specific operating mode, more than two electrical input signals from the microphones are combined to provide a beamformed signal provided by applying appropriate (for example, complex) weights to (at least part of) the corresponding signals.

[0201] The substrate SUB further includes a configurable signal processor DSP (such as a digital signal processor), for example including a processor for applying a gain that varies with frequency and level, for example providing beamforming, noise reduction, filter bank functions, and other digital functions of a hearing device, for example implementing the features (for example in combination with Figure 1A-1B as described in FIGS. 2A-2E). The configurable signal processor DSP is adapted to access the memory MEM to, for example, select appropriate parameters for the current configuration or operating mode and / or listening situation and / or write data to the memory (such as algorithm parameters, for example for recording user behavior). The configurable signal processor DSP is also configured to process one or more electrical input audio signals and / or one or more directly received auxiliary audio input signals based on the currently selected (activated) hearing aid program / parameter settings (for example, automatically selected, such as based on one or more sensors, or selected based on an input from a user interface). The mentioned functional units (and other elements) can be divided into circuits and elements according to the application involved (for example, for size, power consumption, analog-digital processing, acceptable latency, etc.), for example integrated in one or more integrated circuits, or as a combination of one or more integrated circuits and one or more separate electronic components (such as inductors, capacitors, etc.). The configurable signal processor DSP provides a processed audio signal that is intended to be presented to the user. The substrate also includes a front-end IC (FE) for interfacing the configurable signal processor DSP with input and output transducers, etc., and generally including an interface between analog and digital signals (such as an interface to a microphone and / or a speaker and possibly to a sensor / detector). The input and output transducers can be individually separate elements, or integrated with other electronic circuits (such as based on MEMS).

[0202] The hearing device HD also includes an output unit (such as an output transducer) for providing a stimulus that can be perceived as sound by the user based on the processed audio signal from the processor or a signal derived therefrom. In Figure 5A a hearing device embodiment, the ITE part includes an output unit (at least a part of which) in the form of a speaker (also called a receiver) SPK for converting an electrical signal into an acoustic (airborne) signal, which (when the hearing device is mounted at the user's ear) is directed towards the eardrum to provide a sound signal (S ED ). The ITE part also includes a guiding element such as a dome DO for guiding and positioning the ITE part in the user's ear canal. In Figure 5A an embodiment, the ITE part also includes another input transducer such as a microphone (M ITE ) for providing an electrical input audio signal representing the input sound signal (S ITE ) at the ear canal. The sound (S ITE)Propagation from the environment through a direct acoustic path through the semi-open dome piece DO to the residual cavity at the eardrum is indicated by the dashed arrow in Figure 5A , 5B (denoted as the direct path). The directly propagated sound (indicated by sound field S dir ) is mixed with the sound from the hearing device HD (indicated by sound field S HI ) to form a combined sound field (S ED ) at the eardrum. The ITE part may include (possibly customized) an ear mold for providing a fairly tight fit with the user's ear canal. The ear mold may include a ventilation channel to provide (controlled) leakage of sound from the residual cavity between the ear mold and the eardrum (thereby managing the occlusion effect), see Figure 5B .

[0203] (from input transducers M BTE1 , M BTE2 , M ITE ) The electrical input signal can be processed in the time domain or (time-) frequency domain (or partly in the time domain and partly in the frequency domain, if considered advantageous for the application in question).

[0204] In Figure 5A 's embodiment, the connector IC includes electrical conductors for connecting the electrical components of the BTE part and the ITE part. The connector IC may include an electrical connector CON to connect the cable IC to a matching connector in the BTE part. In another embodiment, the connector IC is an acoustic tube, and the speaker SPK is located in the BTE part. In yet another embodiment, the hearing device does not include a BTE part, but the entire hearing device is enclosed in an ear mold (ITE part), for example see Figure 5B .

[0205] Figure 5A And 5B The illustrated embodiments of the hearing device HD are portable devices, which include a battery BAT such as a rechargeable battery, which is, for example, based on lithium-ion battery technology, for example for powering the electronic components of the BTE part and possibly the ITE part. In an embodiment, the hearing device such as a hearing aid is adapted to provide gain that varies with frequency and / or compression that varies with level and / or frequency shifting (with or without frequency compression of one or more frequency ranges to one or more other frequency ranges), for example to compensate for the user's hearing impairment. The BTE part may, for example, include a connector (such as a DAI or USB connector) for connecting a "boot" with additional functions (such as an FM boot or a spare battery, etc.) or a programming device or a charger, etc. to the hearing device HD. As an alternative or in addition, the hearing device may include a wireless interface for programming and / or charging the hearing device.

[0206] Figure 5B Shows another embodiment of the hearing aid HD according to the present invention.Figure 5B An ITE type hearing aid according to an embodiment of the present invention is schematically shown. The hearing aid HD includes or consists of an ITE part, which includes a housing, which can be a standard housing aimed at fitting a group of users, or can be customized for the user's ear (for example, as an ear mold, for example, providing a proper fit with the outer ear and / or ear canal). Figure 5B The housing schematically shown therein has a symmetric shape, for example, about a longitudinal axis (when installed) from the environment towards the user's eardrum, but this is not necessary. It can be customized according to the shape of the ear canal of a specific user. The hearing aid can be configured to be located in the outer part of the ear canal, for example, visible from the outer part; or it can be configured to be completely located in the ear canal, possibly deep into the ear canal, for example, completely or partially in the bony part of the ear canal.

[0207] To minimize the sound leaking from the ear canal (played by the hearing aid towards the user's eardrum), good mechanical contact between the hearing aid housing and the skin / tissue of the ear canal is required. When attempting to minimize the aforementioned leakage, the housing of the ITE part can be customized for the specific user's ear.

[0208] The hearing aid HD includes Q microphones M q , i = 1, …, Q, here Q = 2. The two microphones (M1, M2) are located in the housing and have a predetermined distance d therebetween, such as 8 - 10 mm, for example, on a part of the surface of the housing facing the environment when the hearing aid is installed in or at the user's ear. The microphones (M1, M2) are located on the housing, for example, such that when the hearing aid is installed in or at the user's ear, their microphone axes (the axes passing through the centers of the two microphones) point in the forward direction relative to the user, for example, the user's line of sight (for example, determined by the user's nose, for example, substantially in the horizontal plane). Thus, the two microphones are well adapted to generate a directional signal towards the front (and / or rear) of the user. The microphones are configured to receive the sound (X 1,ac , X 2,ac) is converted into corresponding (analog) electrical signals (x1, x2) representing the sound. The microphone is connected to a corresponding analog-to-digital converter AD to provide the corresponding (analog) electrical signals (x1, x2) as digitized signals (x1, x2). The digitized signals can be further connected to a corresponding filter bank to provide each electrical input signal (time-domain signal) as a sub-band signal (frequency-domain signal). The (digitized) electrical input signals (x1, x2) are fed to a digital signal processor DSP for processing the audio signals (x1, x2), for example including one or more of the following: spatial filtering (beamforming), adaptive mixing (such as fading), (such as single-channel) noise reduction, compression (amplification / attenuation varying with frequency and level according to the user's needs such as hearing impairment), spatial cue retention / restoration, etc. The digital signal processor DSP can, for example, include an appropriate filter bank (such as analysis and synthesis filter banks) to enable processing in the frequency domain (individual processing of sub-band signals). The digital signal processor DSP is configured to provide a processed signal y, which includes a representation of the sound field S (for example including an estimate of the target signal therein). The processed signal y is fed to an output transducer (here a loudspeaker SPK), for example via a synthesis filter bank, and optionally, is fed to a digital-to-analog converter DA for converting the processed (digital electrical) signal y (or an analog version y) into a sound signal S out .

[0209] The hearing aid HD can, for example, include a venting channel (vent), configured to minimize the occlusion effect (when the user is speaking). In addition to enabling the establishment of an (unintended) acoustic propagation path from the residual cavity between the hearing aid housing and the earmold (see Figure 5B ), the venting channel also provides a direct acoustic propagation path for sound from the environment to the residual cavity. The directly propagated sound S reaching the residual cavity dir is mixed with the acoustic output of the hearing aid HD to produce a synthetic sound S at the earmold ED . In one operating mode, active noise suppression (ANS) is activated to attempt to cancel the directly propagated sound S dir .

[0210] The ventilation duct (vent) is asymmetrically located in the hearing aid housing. Such an asymmetric position may be the result of design limitations caused by hearing aid components such as the battery. Consequently, the first and second microphones (M1, M2) have different feedback paths from the speaker SPK. The first microphone M1 is closer to the ventilation duct than the second microphone M2. All other things being equal, the feedback metric FBM1 of the first microphone is greater than the feedback metric FBM2 of the second microphone. A scheme according to the present invention for controlling the use (such as switching, e.g., fading therebetween) of a beamforming signal or a signal from a single input transducer in the forward path of a hearing aid can be applied to ITE hearing aids. Thereby providing more complexity in the positioning of the input transducer and the ventilation duct relative to each other without compromising (reducing) the full-on gain value of the hearing aid. In a particular operating mode, the signal from the (single) microphone with the lowest feedback is used for amplification and presentation to the user. Thus providing a fade according to the present invention between the first and second microphone signals. Thereby minimizing the risk of feedback howling.

[0211] The hearing aid HD includes an energy source such as a battery BAT, e.g., a rechargeable battery, for powering the components of the hearing device.

[0212] Figure 6A An embodiment of a hearing system such as a binaural hearing aid system according to the present invention is shown. The hearing system includes left and right hearing devices that communicate with an auxiliary device, the auxiliary device being, for example, a remote control device, a communication device such as a mobile phone, or a similar device capable of establishing a communication link to one or both of the left and right hearing devices. Figure 6B An auxiliary device configured to run an application program (APP) is shown, the APP implementing a user interface for the hearing device or system, from which an operating mode for selecting a specific sound input can be selected, e.g., an input from a specific microphone or a specific input (such as a pick-up coil or RF input) of a wired or wireless direct reception of sound from another device or a specific beamforming signal.

[0213] Figure 6A 、 6B Together show an application scenario including an embodiment of a binaural hearing aid system according to the present invention, which includes first (left) and second (right) hearing devices (HD1, HD2) and an auxiliary device AD. The auxiliary device AD includes a mobile phone, e.g., a smart phone. In Figure 6A the embodiment, the hearing device and the auxiliary device are configured to establish a wireless link WL-RF therebetween, e.g., in the form of a digital transmission link conforming to a Bluetooth standard (such as Bluetooth Low Energy or an equivalent technology). Alternatively, these links can be implemented in any other convenient wireless and / or wired manner and according to any suitable modulation type or transmission standard (which may be different for different audio sources). Figure 6A 、 6BThe auxiliary device (such as a smart phone) includes a user interface UI that provides remote control functions for the hearing aid device or system, such as for changing the program or operating mode or operating parameters (such as volume) in the hearing device, etc. Figure 6B The user interface UI of shows an APP for selecting the operating mode of the hearing system or device (denoted as "Select Audio Input" (selecting audio input among a microphone, a beamformer, and a direct audio input)), where the user currently prefers a specific input among multiple audio inputs (and can be selected via the user interface). In Figure 6B In the example of, the currently preferred audio input can be selected among the following audio inputs:

[0214] BTE microphone 1;

[0215] ITE microphone;

[0216] Smart phone;

[0217] Forward beamformer;

[0218] Side-facing beamformer;

[0219] Backward beamformer;

[0220] Pickup coil;

[0221] Telephone;

[0222] Music player.

[0223] In Figure 6B In the screen of, "ITE microphone" has been selected, as indicated by the solid "tick box" on the left and the bold "ITE microphone". The screen also includes the instruction "Click on the preferred input. When ready, press 'Start'", referring to the start button at the bottom of the screen.

[0224] When the user changes the currently preferred audio input, for example, from a forward beamformer to a side-facing beamformer (such as in a car situation), the gradual change between the two inputs proposed in the present invention automatically starts.

[0225] In an embodiment of the APP, the user may be allowed to control the details of the gradual change function between the two audio input signals, such as the time period (Δt) of the transition and / or the possible residual weight of the previously preferred audio input (if relevant). In an embodiment, different gradual change parameter (function, time period, residual weight, etc.) configurations may be defined for different pairs of audio inputs.

[0226] The hearing devices (HD1, HD2) are shown in Figure 6A as devices installed at the user U's ear (behind the ear), for example, see Figure 5A. Other styles may be used, such as being entirely located in the ear (e.g., in the ear canal, e.g., see Figure 5B ), being entirely or partially implanted in the head, and so on. As Figure 6A shows, each hearing instrument may include a wireless transceiver to establish an inter-aural wireless link IA-WL between hearing devices, e.g., based on inductive communication or RF communication (e.g., Bluetooth technology). Each hearing device also includes a transceiver for establishing a wireless link WL-RF to an auxiliary device AD (e.g., based on a radiation field (RF)), at least for receiving and / or transmitting signals such as control signals such as information signals, e.g., including audio signals. The transceivers are indicated by RF-IA-Rx / Tx-1 and RF-IA-Rx / Tx-2 in the right (HD2) and left (HD1) hearing devices, respectively.

[0227] In an embodiment, the remote control APP is configured to interact with a single hearing device (rather than with a binaural hearing aid system).

[0228] In Figure 6A , 6B 's embodiment, the auxiliary device is described as a smart phone. However, the auxiliary device may be other portable electronic devices, such as an FM transmitter, a dedicated remote control device, a smart watch, a tablet computer, etc.

[0229] Figure 7 A speakerphone including an input unit is schematically shown. The input unit includes a plurality of microphones configured to pick up sounds from the environment of the speakerphone and a plurality of beamformers configured to focus on a plurality of different target speakers in the environment around the speakerphone and enable an adaptive fade between spatially filtered signals as described in the present invention. The input unit of the speakerphone SPKPHO includes a microphone array, which includes a plurality (here 8) of microphones MIC arranged in a predetermined pattern (here evenly distributed along a circumference). The speakerphone also includes a speaker SPK (here located at the center of the speakerphone). The speaker is configured to play sounds received from a distant sound source so as to be perceived in the environment of the speakerphone. The speakerphone includes a hybrid processor as described in the present invention. The hybrid processor is adapted to provide a processed input signal based on at least a portion of the signals from the plurality of microphones. The processed input signal (or its processed version) is passed to another device or system for further processing and / or presented to one or more users. The speakerphone is also configured to play sounds received from a distant sound source so as to be perceived in the environment of the speakerphone.

[0230] The input unit of a loudspeaker phone may include a beamforming filter unit that receives electrical input signals from a plurality of microphones MIC. The beamforming filter unit is configured to provide at least two spatially filtered (beamformed) signals (four are shown here: BF1, BF2, BF3, BF4) that are directed towards at least two target sound sources (four are shown here: S1, S2, S3, S4) in the loudspeaker phone environment. The plurality of microphones may be divided into microphone subgroups. Each beamformer may be based on a subgroup of the microphones or all of the microphones. The loudspeaker phone may be configured to fade between at least two spatially filtered signals and pass the (resulting) processed input signal (or a further processed version thereof, such as a post-filtered version) to another device or system. The currently active beamformer (BF2) is marked by a thick frame. Loudspeaker S2 is currently active. When (primary) speech activity is detected in another beamformer, the fading process according to the present invention is started. The start of the fading process may be determined by a corresponding voice activity detector (e.g., one voice activity detector per spatially filtered signal).

[0231] Figure 8 An estimator NEST for estimating the noise variance of at least two input audio data streams before audio stream mixing is shown. The input unit IU provides at least two input audio data streams (here x1, x2 from respective microphones M1, M2, and the (digitized) microphone signals x1, x2 are transformed into the frequency domain by respective analysis filter banks A). Each input audio data stream (x1, x2) includes a mixture of a target signal component (s1, s2) from a target sound source and a noise component (v1, v2) from one or more noise sources, as illustrated in Figure 3 , 4 . In Figure 8 , a process for estimating the respective noise variance and is illustrated.

[0232] The two input audio data streams (x1, x2) are multiplied in a combining unit "x" and low-pass filtered by a low-pass filter LP to provide an estimate of the correlation COR between the two data streams (x1, x2). The cross-correlation between x1 and x2 is determined as = <x1 x2*>, where * denotes complex conjugation (see * on the input to the multiplying unit "x" from x2), and <·> denotes smoothing over time, e.g., using a low-pass filter LP, assuming the processing is performed in the filter bank domain (i.e., the (time-)frequency domain), see the analysis filter bank FBA in Figure 8 . The correlation COR is fed to a controller CTR, which is used to control the estimation of the respective noise variance (using control signals U1, U2) and to determine the type of noise (signal NTP) present in the current input audio data streams (x1, x2).

[0233] Each of two input audio data streams (x1, x2) has a separate (identical) noise variance estimation path. Each noise estimation path includes an ABS square function for providing a magnitude squared representation of the two input audio data streams (x1, x2) (|x1| 2 , |x2| 2 ). In each noise estimation path (where m = 1, 2 corresponds to M1, M2), the squared magnitude values (|x1| 2 , |x2| 2 ) are low-pass filtered (LP, LPm, m = 1, 2) in two different parallel signal paths. The low-pass filter LP is configured to be continuously updated to provide an envelope of the squared magnitude values (<|x1| 2 , <|x2| 2 >). The levels (L1, L2) of the envelopes of the squared magnitude values (<|x1| 2 , <|x2| 2 ) are estimated in the respective level estimators LD, and the estimated levels (L1, L2) are fed to the controller CTR. The low-pass filter LP1 in microphone path 1 (and correspondingly LP2 in microphone path M2) is updated under the control of a signal U1 (U2) from the controller CTR. The control signal U1 (U2) is determined based on the correlation COR between the two input audio data streams (x1, x2) and the estimated levels (L1 (L2)) of the envelopes of the squared magnitude values of the respective input audio data streams (x1 (x2)) (<|x1| 2 > (<|x2| 2 >)) at a given time. When the correlation COR is low, the outputs of the low-pass filters LP1 and LP2 represent the noise variances of the first and second input audio data streams x1 and x2 and

[0234] The controller CTR is configured to provide control signals U1, U2, NTP according to the following criteria:

[0235] - If L1 is low (e.g., below a first level threshold L th1 ), update LP1 (U1 = 1);

[0236] - If L2 is low (e.g., below a second level threshold L th2 ), update LP2 (U2 = 1);

[0237] - If COR is low (e.g., below a first correlation threshold COR th1 ), and at the same time Lm (m = 1, 2) is low, signal type = microphone noise (= NTP);

[0238] - If the COR is low (e.g., below a first relevant threshold COR th1 ), and at the same time Lm (m = 1, 2) varies, the signal type = wind noise (= NTP);

[0239] - If the COR is high (e.g., above a second relevant threshold COR th2 ), and at the same time Lm (m = 1, 2) varies, the signal type = speech (= NTP).

[0240] The control signal NTP can be used, for example, to distinguish noise (including the distinction between wind noise and, for example, microphone noise) and noise - free (such as speech), thus implementing a voice activity detector. This control signal can be used, for example, elsewhere in the hearing aid.

[0241] Figure 8 The estimator NEST, for example, can form part of the processor PRO, for example, see Figure 2A , 2B, 2C, 2D, 2E. The estimator NEST can form part of the adaptive mixing unit ADM, for example, see Figure 2C , 2D, 2E. The estimator NEST can form part of the noise variance estimation unit NVE, for example, see Figure 2C .

[0242] Figure 9 Schematically shows an exemplary fade - in process between two input audio data streams with different target and noise levels. The two upper curves schematically show the first and second input audio data streams x1 and x2 (denoted as audio stream #1 and audio stream #2, respectively). The sequence of alternating speech and non - speech (level - time) is shown. The first and second data streams have different maximum and minimum input levels (for simplicity, both are assumed to be constant). Audio stream #1 exhibits a maximum level LS1 and a minimum level LN1. Audio stream #2 exhibits a maximum level LS2 and a minimum level LN2. The maximum levels (LS1, LS2) can be assumed to represent the average speech level (envelope, upper tracking line). The minimum levels (LN1, LN2) can be assumed to represent the average noise level (envelope, bottom tracking line). All four levels (LS1, LS2, LN1, LN2) are marked in the middle curve representing the waveform of the second input audio data stream (audio stream #2). Similarly, in the bottom curve showing the fade - in from audio stream #1 across the fade - in time Δt fad to audio stream #2. The fade - in time (Δt fad = T1 + T2) can be, for example, greater than the minimum time and less than the maximum time, for example, 1s ≤ Δt fad ≤ 5s. The bottom curve shows an example of the fade - in process, where the noise levels (LN1, LN2) represent the noise variance estimators and Instead of suddenly changing from audio stream #1 to audio stream #2, the noise level of the mixed signal remains at the level (LN1) of audio stream #1 for a short period of time (T1) before gradually changing (with time T2) to the level (L2) of audio stream #2. This avoids an obvious unnatural signal during the mixing process.

[0243] When appropriately substituted by corresponding processes, the structural features of the devices described above, detailed in the "Detailed Description" and defined in the claims, can be combined with the steps of the method of the present invention.

[0244] Unless explicitly stated otherwise, the singular forms "a", "the" as used herein are intended to include the plural forms (i.e., having the meaning of "at least one"). It should be further understood that the terms "having", "including", and / or "comprising" as used in the specification indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that unless explicitly stated otherwise, when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may be present. As used herein, the term "and / or" includes any and all combinations of one or more of the listed related items. Unless explicitly stated otherwise, the steps of any method disclosed herein do not have to be performed in the exact order disclosed.

[0245] It should be appreciated that references to "an embodiment" or "embodiments" or "aspect" or "may" include features described in connection with that embodiment are intended to mean that the particular features, structures, or characteristics so described are included in at least one embodiment of the invention. Moreover, the particular features, structures, or characteristics may be appropriately combined in one or more embodiments of the invention. The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.

[0246] The claims are not limited to the aspects shown herein, but encompass the full scope consistent with the claim language, where unless explicitly stated otherwise, an element recited in the singular does not mean "one and only one" but rather "one or more". Unless explicitly stated otherwise, the term "some" means one or more.

[0247] Accordingly, the scope of the present invention should be determined in accordance with the claims.

Claims

1. A hearing device adapted to be located at or in a user's ear or adapted to be fully or partially implanted in a user's head, the hearing device comprising: An input unit providing at least two input audio data streams (x1, x2), each input audio data stream comprising a mixture of a target signal component (s1, s2) from a target sound source and a noise component (v1, v2) from one or more noise sources; A processor for receiving at least two input audio data streams (x1, x2), for mixing at least two input audio data streams or their processed versions and for providing a processed input signal (y) based thereon; An output unit providing an output stimulus perceptible as sound by the user based on the processed input signal (y) or its processed version; wherein the processor is configured to - Estimate the noise variance of each of at least two input audio data streams (x1, x2) before mixing According to the noise variances of at least two input audio data streams (x1, x2) Process the noise components (v1, v2) so as to reduce or avoid unnatural signals caused by mixing in the processed input signal (y) by equalizing the noise components of at least two input audio data streams (x1, x2) in the processed input signal (y); and - Fade from a first input audio data stream (x1) among the at least two input audio data streams (x1, x2) to a second input audio data stream (x2) over a gradual time period, where the fade includes determining a fade parameter or fade curve that gradually decreases the weight (α) of the first input audio data stream (x1) or its processed version while increasing the weight (1-α) of the second input audio data stream (x2) or its processed version, adding additional noise (v3) to the processed input signal (y), and the variance of the additional noise (v3) is determined by the estimated noise variances of the at least two input audio data streams (x1, x2) and the weight (α) such that the noise components (v1, v2) in the processed input signal (y) are equal in level to the noise signal component exhibiting the maximum noise level among the at least two input audio data streams (x1, x2).

2. The hearing device according to claim 1, wherein at least two input audio data streams (x1, x2) originate from two different target sound sources.

3. The hearing device according to claim 1, wherein at least two input audio data streams (x1, x2) originate from the same target sound source.

4. The hearing device according to claim 1, wherein the processor is configured to estimate the levels of the target components of at least two input audio data streams (x1, x2).

5. The hearing device according to claim 1, wherein the noise variance is measured as the background noise floor variance.

6. The hearing device according to claim 1, wherein the processor is configured to identify or use some or all of the noise components of at least part of the at least two input audio data streams (x1, x2) or their processed versions.

7. The hearing device according to claim 1, configured to adaptively determine the noise components (v1, v2) by estimating the noise in the environment during speech pauses.

8. The hearing device according to claim 1, wherein the input unit comprises at least two input transducers, each input transducer providing an electrical input signal representative of sound, and a beamformer filtering unit for spatially filtering the electrical input signal and providing at least one spatially filtered signal based thereon, the spatially filtered signal constituting or forming part of at least one of the at least two input audio data streams.

9. The hearing device according to claim 8, wherein the beamformer filtering unit comprises at least two beamformers configured to provide at least two spatially filtered signals, which may constitute or form part of at least two input audio data streams (x1, x2).

10. The hearing device according to claim 9, wherein the gradation between the respective input audio data streams from at least two beamformers is controlled according to a detected or selected target direction.

11. The hearing device according to claim 1, consisting of or comprising a hearing aid, a headset, an earphone, an ear protection device or a combination thereof.

Citation Information

Patent Citations

  • Method and Device for Acoustic Management Control of Multiple Microphones

    US20090016542A1

  • Hearing device comprising a beamformer filtering unit

    US20170295437A1