Hearing device including a noise reduction system
Through self-voice detector and beamformer technology, the listening device can recognize and attenuate unwanted noise signals, solving the problem of difficult to distinguish user voice from noise in the prior art, and improving the intelligibility and communication quality of the voice signal.
Patent Information
- Application Number
- CN202010955909.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-11
- Filing Date
- 2020-09-11
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-09-11
AI Technical Summary
Existing hearing devices are difficult to effectively distinguish and attenuate the voice signals that users want and do not want, resulting in poor noise signal processing.
Using self-voice detector and beamformer technology, the user's voice is recognized through the self-voice detector and the beamformer is used to attenuate unwanted noise signals, combining the update of the noise covariance matrix and the voice activity detector to achieve accurate identification and attenuation of the noise signal.
It improves the intelligibility of the voice signal in the listening device, reduces unwanted noise interference, and provides a clearer voice communication experience.
Smart Images

Figure CN112492434B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to hearing devices such as hearing aids or headsets, and in particular to noise reduction in hearing devices. The present invention specifically relates to the application of a good (high-quality) estimate of the voice of a user wearing a hearing device, for example for transmission to another device such as a remote communication partner or listener and / or for transmission to a voice interface for voice control of the hearing device (or other device or system). Background Art
[0002] A hearing device can determine whether an audio signal includes voice (or speech), for example, by applying a voice activity detector. However, voices often originate from desired and undesired sound sources simultaneously, making it difficult to distinguish between desired and undesired voice signals and to attenuate the undesired voice signals. Therefore, it is desirable to be able to attenuate the voice from an undesired sound source while enhancing the voice from a desired sound source. Summary of the Invention
[0003] Hearing device
[0004] In one aspect of the present application, a hearing device is disclosed. The hearing device can be adapted to be located at or in the user's ear, or to be fully or partially implanted in the user's head.
[0005] The hearing device can include an input unit for providing at least one electrical input signal representing sound in the user's environment. The environment can refer to the free space around the user, which is fixed and / or dynamically depends on whether the user is standing still or walking around, and which contains audio (such as sound) reaching the user's location. For example, the environment can refer to a closed classroom in which the user is located, or, when the user is located outside a building, for example, it can refer to the open space around the user.
[0006] The electrical input signal can include a target voice signal from a target sound source and additional signal components (hereinafter referred to as noise signal components) from one or more other sound sources. The target sound source can refer to one or more sound sources such as one or more persons (for example, the user of the hearing device and / or other persons) or one or more electronic devices (such as a television, radio, etc.), which generate and / or emit voice signals that the user wants to hear. One or more other sound sources can refer to one or more persons, electronic devices or other sound sources (such as instruments, animals, etc.), which generate and / or emit additional signal components, namely noise signal components, which are regarded as signal components that the user does not want and should preferably be attenuated.
[0007] The hearing device can include a noise reduction system for providing an estimate of the target voice signal.
[0008] The noise signal components can be at least partially attenuated.
[0009] The hearing device may include an own voice detector for repeatedly estimating whether or with what probability at least one electrical input signal or a signal derived therefrom includes speech from the user's voice.
[0010] The hearing device may also be configured such that a noise signal component is identified during a time period.
[0011] The own voice detector may indicate that at least one electrical input signal or a signal derived therefrom is from the user's voice or is from the user's voice with a probability higher than an own voice presence probability (OVPP) threshold.
[0012] Thus, the noise signal component, which may also include speech from an unwanted sound source, may be detected during a time interval when the own voice detector estimates that the user is speaking, for example instead of (as common in the art) during a time interval without voice activity, or in addition to being detected during a time interval without voice activity, also during a time interval when the own voice detector estimates that the user is speaking. Accordingly, the noise signal component to be attenuated may also be updated while the user is speaking. For example, if a person is speaking during the same time period as the user, the sound from that person may be identified and marked as noise, which should be attenuated.
[0013] Furthermore, using own voice detection to identify the noise signal component no longer requires an additional detector (such as a camera) dedicated to, for example, identifying whether a particular person is an unwanted noise source when speaking during the same time period as the user through image analysis.
[0014] Thus, noise reduction can be improved.
[0015] The input unit may include a microphone. The input unit may include at least two microphones. The input unit may include more than three microphones.
[0016] Each microphone may provide an electrical input signal. The electrical input signal may include a target speech signal and a noise signal component.
[0017] The hearing device may include a voice activity detector for repeatedly estimating whether or with what probability at least one electrical input signal or a signal derived therefrom includes speech.
[0018] Thus, the speech included in at least one electrical input signal can be enhanced.
[0019] The hearing device may include one or more beamformers. For example, the beamformer filter may include more than two beamformers.
[0020] The input unit may be configured to provide at least two electrical input signals connected to one or more beamformers. The one or more beamformers may be configured to provide at least one beamformed signal.
[0021] The one or more beamformers may include one or more self-voice cancellation beamformers configured to attenuate signal components originating from the user's mouth while signal components from (e.g., all) other directions remain unchanged or are less attenuated.
[0022] The one or more may include one or more target beamformers for enhancing the voice of a target sound source (relative to sounds from other directions different from the direction of the target sound source).
[0023] The target signal may be assumed to be the user's self-voice.
[0024] The one or more beamformers may include a self-voice beamformer configured to retain signal components from the user's mouth while attenuating signal components from (e.g., all) other directions. The self-voice beamformer may be determined before operation of the hearing device (e.g., during fitting), and the corresponding filter weights may be stored, for example, in the memory of the hearing device. The acoustic transfer function from the user's mouth to each microphone of the hearing device may be determined, for example, before operation of the hearing device, or using a model (such as a head and torso model, e.g., the HATS, Head and Torso Simulator 4128C from Brüel& Sound&Vibration Measurement A / S), or by measuring one or more persons including the user. The absolute or relative acoustic transfer function may be represented by the vector d = (d1,…,d M )), where each element represents the (absolute or relative) transfer function from the mouth to a particular microphone among the M microphones. One of the microphones may be defined as a reference microphone, and the relative transfer function may be defined as the transfer function from the reference microphone to the remaining microphones of the hearing device (or hearing system). The self-voice filter weights W OV may be determined before or during operation of the hearing device. The self-voice filter weights are a function of the vector d OV (k) of the noisy microphone signals, the noise covariance matrix estimator and the inter-microphone covariance matrix C x (k,n), where k and n are the frequency index and the time index, respectively. For a given type of beamformer (such as an MVDR beamformer), the calculation of the filter weights is conventional in the art and is illustrated, for example, in the detailed description section of this specification.
[0025] The beamformer may include a minimum variance distortionless response (MVDR) beamformer.
[0026] The beamformer may include a multichannel Wiener filter (MWF) beamformer.
[0027] The beamformer may include an MVDR beamformer and an MWF beamformer.
[0028] The beamformer may include an MVDR filter and a single-channel post-filter following it.
[0029] For example, the beamformer may include an MVDR beamformer and a single-channel post-Wiener filter.
[0030] The advantage of using an MVDR filter is that it does not distort the target component.
[0031] The advantage of using an MWF filter is that it maximizes the broadband signal-to-noise ratio (SNR).
[0032] The noise signal component may be represented by a noise covariance matrix estimator.
[0033] The noise covariance matrix may be based on the cross power spectral densities (CPSDs) of the noise signal components.
[0034] Thereby providing a simple (mathematically tractable) description of the noise field.
[0035] The hearing device may include a beamformer filter comprising a plurality of beamformers.
[0036] The noise covariance matrix may be updated when the self-voice detector indicates that at least one electrical input signal or a signal derived therefrom is from the user's voice.
[0037] The noise covariance matrix may be updated when the self-voice detector indicates that at least one electrical input signal or a signal derived therefrom is from the user's voice with a probability higher than the OVPP threshold.
[0038] Thus, the voice (unwanted speech) from a (competitive) speaker that the user is (currently) not interested in and / or interferes with the user's speech can be attenuated.
[0039] The noise signal component may alternatively be identified during a time period when the voice activity detector indicates the absence of speech in at least one electrical input signal or a signal derived therefrom.
[0040] The noise signal component can be identified during periods when the voice activity detector indicates that there is no speech or that speech is present with a probability below the speech presence probability (SPP) threshold.
[0041] The hearing device can be configured to estimate the noise signal component using a maximum likelihood estimator.
[0042] Thereby, providing a noise covariance matrix estimator that best "explains" (has the maximum likelihood) the observed microphone signal.
[0043] The target speech signal from the target sound source can include (or constitute) the self-voice speech signal from the user of the hearing device.
[0044] The target sound source can include (or constitute) an external speaker in the environment of the user of the hearing device.
[0045] The hearing device can include a voice interface for voice control of the hearing device or other devices or systems.
[0046] The input of the voice interface can be based, for example, on an estimate of the user's self-voice provided by a self-voice beamformer, which is configured to retain the signal component from the user's mouth while attenuating the signal components from (e.g., all) other directions. The hearing device can include a wake word detector based on an estimate of the user's voice. The hearing device can be configured to activate the voice interface when (e.g., with a probability higher than the wake word detection threshold such as greater than 60%) a wake word is detected.
[0047] The voice interface can be included in the part of the hearing device that is placed at, behind, or in the user's ear. The hearing device can include one or more "auxiliary devices" that communicate with the hearing device and affect and / or benefit from the functions of the hearing device. The auxiliary device can be, for example, a remote control, an audio gateway device, a mobile phone (such as a smart phone), or a music player. In this case, one or more auxiliary devices can include a voice interface.
[0048] By providing a hearing device that includes a voice interface, seamless processing of the functions of the hearing device is provided.
[0049] The hearing device can be constituted by or include a hearing aid, a headset, an active ear protection device, or a combination thereof.
[0050] The hearing device can include a headset. The hearing device can include a hearing aid. The hearing device can, for example, include an antenna and transceiver circuitry configured to establish a communication link to another device or system. The hearing device can be used, for example, to implement a hands-free phone.
[0051] The hearing device may further include a timer configured to determine an overlapping time period between the self-voice speech signal and another speech signal.
[0052] The another speech signal may refer to a speech signal generated by a person, a radio, a television, etc.
[0053] The timer may be associated with the self-voice detector. When the target speech signal includes speech from the hearing device user, the timer may start when the another speech signal is detected during the time period when the self-voice detector detects the speech signal from the user. The timer may end when the self-voice detector does not detect the speech signal from the user. Thus, the unwanted speech signal may be identified and attenuated.
[0054] The hearing device may be configured to determine whether the time period exceeds a time limit, and if so, mark the another speech signal as part of a noise signal component.
[0055] For example, the time limit may be at least one-half second, at least one second, at least two seconds.
[0056] The another speech signal may be speech from a competing speaker, which may itself be regarded as noise relative to the target speech signal. Thus, the another speech signal may be marked as part of a noise signal component so that the another speech signal may be attenuated.
[0057] The hearing device may be configured to, for a predetermined time period, mark the another speech signal as part of a noise signal component. Thereafter, the another speech signal may not be marked as part of a noise signal component. For example, when a person is not part of a conversation with the hearing device user, the speech signal from that person may be attenuated, but at a subsequent time, when that person participates in a conversation with the hearing device user, it may not be attenuated.
[0058] The noise reduction system may be updated recursively. The noise signal component may be identified recursively. Thus, a recursive update of the noise covariance matrix may be provided. For example, a speech signal from a sound source that has previously been identified and marked as part of a noise signal component may be attenuated over time to a continuously decreasing extent. At a certain time, the sound source may be exempt from being attenuated unless the sound source is again identified and marked as part of a noise signal component.
[0059] The hearing device may be adapted to provide frequency-varying gain and / or level-varying compression and / or frequency shifting from one or more frequency ranges to one or more other frequency ranges (with or without frequency compression) to compensate for the user's hearing impairment. The hearing device may include a signal processor for enhancing the input signal and providing a processed output signal.
[0060] The hearing device may include an output unit for providing a stimulus that is perceived by the user as an acoustic signal based on the processed electrical signal. The output unit may include a plurality of electrodes of a cochlear implant (for CI-type hearing devices) or a vibrator of a bone conduction hearing device. The output unit may include an output transducer. The output transducer may include a receiver (loudspeaker) for providing the stimulus as an acoustic signal to the user (e.g., in an acoustic (air-conduction-based) hearing device). The output transducer may include a vibrator for providing the stimulus as a mechanical vibration of the skull to the user (e.g., in a bone-attached or bone-anchored hearing device). The output unit may include a wireless transmitter for transmitting a wireless signal including or representing sound to another device.
[0061] The hearing device includes an input unit for providing one or more electrical input signals representing sound. The input unit may include an input transducer such as a microphone for converting the input sound into an electrical input signal. The input unit may include a wireless receiver for receiving a wireless signal including or representing sound and providing an electrical input signal representing the sound.
[0062] The wireless receiver and / or transmitter (such as a transceiver) may be configured, for example, to receive and / or transmit electromagnetic signals in the radio frequency range (3 kHz to 300 GHz). The wireless receiver and / or transmitter may be configured, for example, to receive and / or transmit electromagnetic signals in the optical frequency range (e.g., infrared light 300 GHz to 430 THz, or visible light, e.g., 430 THz to 770 THz).
[0063] The hearing device may include an antenna and transceiver circuitry (such as a wireless receiver) for receiving signals from another device such as an entertainment device (e.g., a television), a communication device (e.g., a smart phone), a wireless microphone, a personal computer, or another hearing device and / or for transmitting signals to the aforementioned other device. The signals may represent or include audio signals and / or control signals and / or information signals. The hearing device may include appropriate modulation / demodulation circuitry for modulating / demodulating the transmitted / received signals. The signals may represent audio signals and / or control signals, such as for setting operating parameters (e.g., volume) and / or processing parameters and / or voice control commands, etc., of the hearing device. Generally speaking, the wireless link established by the antenna and transceiver circuitry of the hearing device can be of any type. The wireless link may be established between two devices, such as between an entertainment device (e.g., a TV) or a communication device (e.g., a smart phone) and the hearing device, or between two hearing devices, such as via a third intermediate device (e.g., a processing device, such as a remote control device, a smart phone, etc.). The wireless link may be a near-field communication-based link, such as an inductive link based on inductive coupling between the antenna coils of a transmitter part and a receiver part. The wireless link may be based on far-field electromagnetic radiation. Communication via the wireless link may be arranged according to a specific modulation scheme, such as an analog modulation scheme, such as FM (frequency modulation) or AM (amplitude modulation) or PM (phase modulation), or a digital modulation scheme, such as ASK (amplitude shift keying) such as on-off keying, FSK (frequency shift keying), PSK (phase shift keying) such as MSK (minimum shift keying) or QAM (quadrature amplitude modulation), etc.
[0064] Communication between the hearing device and another device may be in the baseband (audio frequency range, such as between 0 and 20 kHz). Communication between hearing devices may be based on a certain type of modulation at a frequency higher than 100 kHz. Preferably, the frequency used to establish a communication link between the hearing device and another device is lower than 70 GHz, such as in the range from 50 MHz to 70 GHz, such as higher than 300 MHz, such as in the ISM range higher than 300 MHz, such as in the 900 MHz range or in the 2.4 GHz range or in the 5.8 GHz range or in the 60 GHz range (ISM = industrial, scientific, and medical, such standardized ranges are defined, for example, by the International Telecommunication Union ITU). The wireless link may be based on a standardized or proprietary technology. The wireless link may be based on Bluetooth technology (such as Bluetooth Low Energy technology).
[0065] The hearing device may have a maximum outer dimension on the order of 0.08 m (such as for headphones). The hearing device may have a maximum outer dimension on the order of 0.04 m (such as for hearing instruments).
[0066] A hearing device may include a directional microphone system adapted to spatially filter sounds from the environment so as to enhance a target sound source among a plurality of sound sources in the local environment of a user wearing the hearing device. The directional system may be adapted to detect (such as adaptively detect) from which direction a particular part of the microphone signal originates. This can be achieved in a variety of different ways described in the prior art, for example. In a hearing device, a microphone array beamformer is typically used to spatially attenuate background noise sources. Many beamformer variants can be found in the literature. The minimum variance distortionless response (MVDR) beamformer is widely used in microphone array signal processing. Ideally, the MVDR beamformer leaves the signal from the target direction (also called the look direction) unchanged while maximally attenuating sound signals from other directions. The generalized sidelobe canceller (GSC) structure is an equivalent representation of the MVDR beamformer, which offers computational and digital representation advantages compared to a direct implementation of the original form.
[0067] The hearing device may be a portable (i.e., configured to be wearable) device or form part of one, such as a device including a native energy source such as a battery, for example a rechargeable battery. The hearing device may be a lightweight, easily wearable device, for example having a total weight of less than 100 g, such as less than 20 g, such as less than 10 g.
[0068] The hearing device may include a forward or signal path between an input unit (such as an input transducer, for example a microphone or a microphone system and / or a direct electrical input such as a wireless receiver) and an output unit such as an output transducer. A signal processor may be located in this forward path. The signal processor may be adapted to provide frequency-variable gain according to the specific needs of the user. The hearing device may include an analysis path having functional elements for analyzing the input signal (such as determining level, modulation, signal type, acoustic feedback estimate, etc.). Part or all of the signal processing in the analysis path and / or the signal path may be performed in the frequency domain. Part or all of the signal processing in the analysis path and / or the signal path may be performed in the time domain.
[0069] An analog electrical signal representing an acoustic signal may be converted to a digital audio signal during an analog-to-digital (AD) conversion process, where the analog signal is sampled at a predetermined sampling frequency or sampling rate f s at, f s for example in a range from 8 kHz to 48 kHz (adapting to the specific needs of the application) to provide digital samples x n (or x[n]) at discrete time points t n (or n), each audio sample representing the value of the acoustic signal at t b by a predetermined N n bits, N b for example in a range from 1 to 48 bits such as 24 bits. Each audio sample is thus quantized using N b bits (resulting in 2Nb different possible values). The digital sample x has a 1 / f s time length, such as 50 μs, for f s = 20 kHz. Multiple audio samples can be arranged in time frames. A time frame can include 64 or 128 audio data samples. Other frame lengths can be used according to the actual application.
[0070] The hearing device can include an analog-to-digital (AD) converter to digitize an analog input (e.g., from an input transducer such as a microphone) at a predetermined sampling rate such as 20 kHz. The hearing device can include a digital-to-analog (DA) converter to convert a digital signal into an analog output signal, e.g., for presentation to the user via an output transducer.
[0071] A hearing device such as an input unit and / or an antenna and transceiver circuit can include a TF conversion unit for providing a time-frequency representation of an input signal. The time-frequency representation can include an array or mapping of the corresponding complex or real values of the signal involved in a specific time and frequency range. The TF conversion unit can include a filter bank for filtering the (time-varying) input signal and providing a plurality of (time-varying) output signals, each output signal including a distinct input signal frequency range. The TF conversion unit can include a Fourier transform unit for converting the time-varying input signal into a (time-) frequency domain (time-varying) signal. The frequency range considered by the hearing device, from the minimum frequency f min to the maximum frequency f max can include a part of the typical human audible frequency range from 20 Hz to 20 kHz, e.g., a part of the range from 20 Hz to 12 kHz. Generally, the sampling rate f s is greater than or equal to twice the maximum frequency f max , i.e., f s ≥ 2f max . The signals in the forward path and / or the analysis path of the hearing device can be split into NI (e.g., uniformly wide) frequency bands, where NI is, for example, greater than 5, such as greater than 10, such as greater than 50, such as greater than 100, such as greater than 500, and at least some of them are processed individually. The hearing aid can be adapted to process the signals in the forward and / or analysis paths in NP different channels (NP ≤ NI). The channels can have the same width or different widths (e.g., the width increases with frequency), and can overlap or not overlap.
[0072] The hearing device may be configured to operate in different modes, such as a normal mode and one or more specific modes, which may be user-selectable or may be automatically selected. The operating mode may be optimized for a specific acoustic situation or environment. The operating mode may include a low-power mode, in which the functionality of the hearing device is reduced (e.g., in order to save energy), such as disabling wireless communication and / or disabling specific features of the hearing device. The operating mode may be a voice control mode, in which the voice interface is activated, for example, by a specific wake word such as "Hey Oticon". The operating mode may be a communication mode, in which the hearing device is configured to pick up the user's voice and transfer it to another device (and possibly receive audio from another device, such as to enable hands-free calling).
[0073] The hearing device may include a plurality of detectors configured to provide status signals related to the current network environment of the hearing device (such as the current acoustic environment), and / or related to the current state of the user wearing the hearing device, and / or related to the current state or operating mode of the hearing device. As an alternative or in addition, one or more detectors may form part of an external device communicating with the hearing device (such as wirelessly). The external device may include, for example, another hearing device, a remote control, an audio transmission device, a telephone (such as a smartphone), an external sensor, etc.
[0074] One or more of the plurality of detectors may act on the full-band signal (time domain). One or more of the plurality of detectors may act on the frequency-band split signal ((time-)frequency domain), for example, in a limited number of frequency bands.
[0075] The plurality of detectors may include a level detector for estimating the current level of the signal in the forward path. The predefined criterion includes whether the current level of the signal in the forward path is above or below a given (L-)threshold. The level detector may act on the full-frequency band signal (time domain). The level detector may act on the frequency-band split signal ((time-)frequency domain).
[0076] The hearing device may include a voice detector (VD) for estimating whether (or with what probability) the input signal (at a specific time point) includes a voice signal. In this specification, a voice signal includes speech signals from humans. It may also include other forms of vocalization produced by the human speech system (such as singing). The voice detector unit may be adapted to classify the user's current acoustic environment as a "voice" or "no voice" environment. This has the advantage that time periods of the microphone signal including human vocalizations (such as speech) in the user's environment can be identified and thus separated from time periods including only (or mainly) other sound sources (such as artificially generated noise). The voice detector may be adapted to also detect the user's own voice as "voice". As an alternative, the voice detector is adapted to exclude the user's own voice from the detection of "voice".
[0077] The hearing device may include a self-voice detector for estimating whether (or with what probability) a particular input sound (such as voice, e.g., speech) originates from the voice of the user of the hearing system. The microphone system of the hearing device may be adapted to be able to distinguish the user's own voice from the voice of another person and possibly from non-voice sounds.
[0078] The plurality of detectors may include a motion detector, such as an acceleration sensor. The motion detector may be configured to detect movements of the user's facial muscles and / or bones, such as due to speech or chewing (e.g., jaw movement), and provide a detector signal indicative of the movement.
[0079] The hearing device may include a classification unit configured to classify the current situation based on input signals from (at least part of) the detectors and possibly other inputs. In this specification, the "current situation" is defined by one or more of the following:
[0080] a) The physical environment (e.g., including the current electromagnetic environment, such as the presence of planned or unplanned electromagnetic signals received by the hearing device (including audio and / or control signals), or other properties of the current environment that are different from acoustic);
[0081] b) The current acoustic situation (input level, feedback, etc.);
[0082] c) The current mode or state of the user (movement, temperature, cognitive load, etc.);
[0083] d) The current mode or state of the hearing device and / or another device communicating with the hearing device (selected program, time elapsed since the last user interaction, etc.).
[0084] The classification unit may be based on or include a neural network, such as a trained neural network.
[0085] The hearing device may also include other suitable functions for the applications involved, such as compression, feedback control, etc.
[0086] The hearing device may include a listening device such as a hearing aid, a hearing instrument, e.g., a hearing instrument adapted to be located at the user's ear or fully or partially in the ear canal, such as a headset, an earphone, an ear protection device, or a combination thereof. The hearing system may include a loudspeaker (including a plurality of input transducers and a plurality of output transducers, e.g., used in an audio conferencing scenario), e.g., including a beamformer filtering unit, e.g., providing multiple beamforming capabilities.
[0087] In one aspect of the present application, a binaural hearing system including a first hearing device and an auxiliary device is disclosed. The binaural hearing system may be configured to enable data exchange between the first hearing device and the auxiliary device.
[0088] In one aspect of the present application, a binaural hearing system including first and second hearing devices is disclosed. The binaural hearing system can be configured to enable data exchange between the first and second hearing devices, for example, via an intermediate auxiliary device.
[0089] Application
[0090] In one aspect, there is provided an application of the hearing device as described above, detailed in the "Detailed Description" section and defined in the claims. Applications can be provided in systems including one or more hearing aids (such as hearing instruments), headsets, earphones, active ear protection systems, etc., for example, in hands-free telephone systems, remote conferencing systems (such as including loudspeakers), broadcast systems, karaoke systems, classroom amplification systems, etc.
[0091] Method
[0092] In one aspect, there is provided a method of operating a hearing device.
[0093] The hearing device can be adapted to be located at or in the user's ear, or to be fully or partially implanted in the user's head.
[0094] The method can include providing at least one electrical input signal representative of sound in the user's environment.
[0095] The electrical input signal can include a target speech signal from a target sound source and additional signal components (referred to as noise signal components) from one or more other sound sources.
[0096] The method can include providing an estimate of the target speech signal.
[0097] The noise signal components can be at least partially attenuated.
[0098] The method can include repeatedly estimating whether or with what probability at least one electrical input signal or a signal derived therefrom includes speech from the user's voice.
[0099] The method can further include identifying the noise signal components during a time period.
[0100] The self-voice detector can indicate that at least one electrical input signal or a signal derived therefrom is from the user's voice or with a probability higher than a self-voice presence probability (OVPP) threshold.
[0101] When appropriately replaced by corresponding processes, some or all of the structural features of the device described above, detailed in the "Detailed Description", or defined in the claims can be combined with the implementation of the method of the present invention, and vice versa. The implementation of the method has the same advantages as the corresponding device.
[0102] Computer-readable medium or data carrier
[0103] The present invention further provides a tangible computer-readable medium (data carrier) for storing a computer program including program code (instructions), which, when the computer program runs on a data processing system, causes the data processing system (computer) to execute (complete) at least part (such as most or all) of the steps of the method described above, detailed in the "Detailed Description" and defined in the claims.
[0104] By way of example and not limitation, the foregoing tangible computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to execute or store the required program code in the form of instructions or data structures and can be accessed by a computer. As used herein, a disk includes a compact disk (CD), a laser disk, an optical disk, a digital versatile disk (DVD), a floppy disk, and a Blu-ray disk, where these disks typically magnetically replicate data while these disks can optically replicate data using a laser. Other storage media include those stored in DNA (for example, in synthetic DNA strands). Combinations of the above disks should also be included within the scope of computer-readable media. In addition to being stored on a tangible medium, a computer program can also be transmitted via a transmission medium such as a wired or wireless link or a network such as the Internet and loaded into a data processing system to run at a location different from the tangible medium.
[0105] The method steps for providing an estimate of a target speech signal in which a noise signal component is at least partially attenuated can be implemented in software.
[0106] The noise signal component can be at least partially attenuated.
[0107] The method steps for repeatedly estimating whether or with what probability at least one electrical input signal or a signal derived therefrom includes speech from a user's voice can be implemented in software.
[0108] The method steps for identifying a noise signal component during a period when a self-voice detector indicates that at least one electrical input signal or a signal derived therefrom is from a user's voice or is from a user's voice with a probability higher than a self-voice presence probability (OVPP) threshold can be implemented in software.
[0109] Computer program
[0110] In addition, the present application provides a computer program (product) including instructions, which, when the program runs on a computer, causes the computer to execute the method (steps) described above, detailed in the "Detailed Description" and defined in the claims.
[0111] Data processing system
[0112] On the one hand, the present invention further provides a data processing system, including a processor and program code, the program code causing the processor to execute at least part (such as most or all) of the steps of the method described above, detailed in the "Detailed Description" and defined in the claims.
[0113] Hearing system
[0114] On the other hand, there is provided a hearing system including the hearing device and the auxiliary device described above, detailed in the "Detailed Description" and defined in the claims.
[0115] The hearing system may be adapted to establish a communication link between the hearing device and the auxiliary device so that information (such as control and status signals, possibly audio signals) can be exchanged or forwarded from one device to another.
[0116] The auxiliary device may include a remote control, a smart phone, or other portable or wearable electronic devices such as a smart watch, etc.
[0117] The auxiliary device may constitute or include a remote control for controlling the functions and operations of the hearing device. The functions of the remote control may be implemented in a smart phone, and the smart phone may run an APP that enables the functions of controlling the audio processing device via the smart phone (the hearing device includes an appropriate wireless interface to the smart phone, such as based on Bluetooth or some other standardized or proprietary solution).
[0118] The auxiliary device may be or include an audio gateway device, which is adapted to receive multiple audio signals (such as from an entertainment device such as a TV or a music player, from a telephone device such as a mobile phone, or from a computer such as a PC) and is adapted to select and / or combine appropriate signals (or signal combinations) among the received audio signals for transmission to the hearing device.
[0119] The auxiliary device may be constituted by or include another hearing device. The hearing system may include two hearing devices adapted to implement a binaural hearing system such as a binaural hearing aid system.
[0120] APP
[0121] On the other hand, the present invention also provides a non-transitory application called an APP. The APP includes executable instructions configured to run on the auxiliary device to implement a user interface for the hearing device or the hearing system described above, detailed in the "Detailed Description" and defined in the claims. The APP may be configured to run on a mobile phone such as a smart phone or another portable device enabling communication with the hearing device or the hearing system.
[0122] Definition
[0123] In this specification, a "hearing device" refers to a device suitable for improving, enhancing, and / or protecting a user's hearing ability, such as a hearing aid, for example, a hearing instrument or an active ear protection device or other audio processing device, which is achieved by receiving an acoustic signal from the user's environment, generating a corresponding audio signal, possibly modifying the audio signal, and providing the possibly modified audio signal as an audible signal to at least one ear of the user. A "hearing device" also refers to a device suitable for electronically receiving an audio signal, possibly modifying the audio signal, and providing the possibly modified audio signal as an audible signal to at least one ear of the user, such as a headset or an earphone. The audible signal can be provided, for example, in the following forms: an acoustic signal radiated into the user's outer ear, an acoustic signal transmitted as a mechanical vibration through the bone structure of the user's head and / or through parts of the middle ear to the user's inner ear, and an electrical signal transmitted directly or indirectly to the user's cochlear nerve.
[0124] The hearing device can be configured to be worn in any known manner, such as a unit worn behind the ear (with a tube for guiding the radiated acoustic signal into the ear canal or with an output transducer, such as a speaker, arranged to be close to or in the ear canal), a unit arranged wholly or partly in the auricle and / or the ear canal, a unit connected to a fixed structure implanted in the skull, such as a vibrator, or a connectable or wholly or partly implantable unit, etc. The hearing device can include a single unit or several units in electronic communication with each other. The speaker can be provided in a housing together with other components of the hearing device, or it can itself be an external unit (possibly combined with a flexible guiding element, such as a dome-shaped element). The hearing device can be implemented in a single unit (housing) or can be implemented in multiple units connected to each other.
[0125] More generally, a hearing device includes an input transducer for receiving an acoustic signal from the user's environment and providing a corresponding input audio signal and / or a receiver for receiving the input audio signal electronically (i.e., wired or wirelessly), a (usually configurable) signal processing circuit for processing the input audio signal (such as a signal processor, e.g., including a configurable (programmable) processor, e.g., a digital signal processor), and an output unit for providing an audible signal to the user based on the processed audio signal. The signal processor may be adapted to process the input signal in the time domain or in multiple frequency bands. In some hearing devices, an amplifier and / or a compressor may form part of the signal processing circuit. The signal processing circuit typically includes one or more (integrated or separate) storage elements for executing programs and / or for storing parameters used (or potentially used) in the processing and / or for storing information suitable for the function of the hearing device and / or for storing information used, for example, in connection with an interface to the user and / or to a programming device (such as processed information, e.g., provided by the signal processing circuit). In some hearing devices, the output unit may include an output transducer, such as a loudspeaker for providing an air-conducted acoustic signal or a vibrator for providing a structure- or fluid-conducted acoustic signal. In some hearing devices, the output unit may include one or more output electrodes for providing an electrical signal (such as a multi-electrode array for electrically stimulating the cochlear nerve). The hearing device may include a horn loudspeaker (including multiple input transducers and multiple output transducers, e.g., used in an audio conferencing scenario).
[0126] In some hearing devices, the vibrator may be adapted to transmit a structure-conducted acoustic signal to the skull transcutaneously or percutaneously. In some hearing devices, the vibrator may be implanted in the middle ear and / or the inner ear. In some hearing devices, the vibrator may be adapted to provide a structure-conducted acoustic signal to the middle ear bones and / or the cochlea. In some hearing devices, the vibrator may be adapted to provide a fluid-conducted acoustic signal to the cochlear fluid, for example, through the oval window. In some hearing devices, the output electrodes may be implanted in the cochlea or on the inner side of the skull and may be adapted to provide an electrical signal to the hair cells of the cochlea, one or more auditory nerves, the auditory brainstem, the auditory midbrain, the auditory cortex, and / or other parts of the cerebral cortex.
[0127] Hearing devices such as hearing aids can be adapted to the needs of a particular user such as hearing impairment. The configurable signal processing circuit of the hearing device may be adapted to apply compression amplification that varies with frequency and level to the input signal. Customized gain (amplification or compression) that varies with frequency and level may be determined during the fitting process by a fitting system based on the user's hearing data such as an audiogram using fitting principles (such as adapted to speech). The gain that varies with frequency and level may be embodied, for example, in processing parameters, e.g., uploaded to the hearing device via an interface to a programming device (fitting system) and used by a processing algorithm executed by the configurable signal processing circuit of the hearing device.
[0128] "Hearing system" refers to a system including one or two hearing devices. "Binaural hearing system" refers to a system including two hearing devices and adapted to provide audible signals to both ears of a user in a coordinated manner. A hearing system or binaural hearing system may also include one or more "auxiliary devices" which communicate with the hearing devices and affect and / or benefit from the functions of the hearing devices. Auxiliary devices may be, for example, a remote control, an audio gateway device, a mobile phone (such as a smart phone) or a music player. A hearing device, hearing system or binaural hearing system may be used, for example, to compensate for the loss of auditory ability of a hearing-impaired person, enhance or protect the auditory ability of a person with normal hearing and / or transmit an electronic audio signal to a person. A hearing device or hearing system may form part of or interact with, for example, a broadcast system, an active ear protection system, a hands-free telephone system, an automotive audio system, an entertainment (such as karaoke) system, a teleconference system, a classroom amplification system, etc.
[0129] Embodiments of the present invention can be used, for example, in applications that require a good (high-quality) estimate of the voice of a user wearing a hearing device. BRIEF DESCRIPTION OF THE DRAWINGS
[0130] Various aspects of the present invention will be best understood from the following detailed description taken in conjunction with the accompanying drawings. For clarity, these drawings are schematic and simplified, showing only the details necessary for understanding the present invention and omitting other details. Throughout the specification, the same reference numerals are used for the same or corresponding parts. Each feature of each aspect may be combined with any or all features of other aspects. These and other aspects, features and / or technical effects will be apparent from and elucidated in conjunction with the following illustrations, in which:
[0131] Figure 1A An exemplary application scenario of a hearing device system according to the present invention is shown;
[0132] Figures 1B - 1D The corresponding voice activity, voice activity detector (VAD) and noise update for the same time period according to the present invention are shown respectively;
[0133] Figure 2A An exemplary application scenario of a hearing device system according to the present invention is shown;
[0134] Figures 2B - 2D The corresponding voice activity, voice activity detector (VAD) and noise update for the same time period according to the present invention are shown respectively;
[0135] Figure 3A An exemplary application scenario of a hearing device system according to the present invention is shown;
[0136] Figures 3B - 3DRespectively shown are corresponding voice activities, voice activity detectors (VADs), and noise updates according to the present invention for the same time period.
[0137] Figure 4A An exemplary input unit is shown connected to an exemplary noise reduction system.
[0138] Figure 4B An exemplary input unit according to the present invention is shown connected to an exemplary noise reduction system.
[0139] Figure 5A An exemplary block diagram of a hearing aid including a noise reduction system according to an embodiment of the present invention is shown.
[0140] Figure 5B An exemplary block diagram of a hearing aid including a noise reduction system in a hands-free phone operation mode according to an embodiment of the present invention is shown.
[0141] Figure 5C An exemplary block diagram of a hearing aid including a noise reduction system and a voice control interface according to an embodiment of the present invention is shown.
[0142] Figure 6 An exemplary application scenario of a hearing device system according to the present invention is shown.
[0143] From the detailed description given below, the further scope of application of the present invention will become apparent. However, it should be understood that while the detailed description and specific examples indicate preferred embodiments of the present invention, they are given for illustrative purposes only. For those skilled in the art, other embodiments of the present invention will be apparent based on the following detailed description. Detailed Description of the Specific Embodiments
[0144] The following specific description presented in conjunction with the drawings is used as a description of various different configurations. The specific description includes specific details for providing a thorough understanding of multiple different concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. Several aspects of the apparatus and method are described by multiple different blocks, functional units, modules, elements, circuits, steps, processes, algorithms, etc. (collectively referred to as "elements"). Depending on a particular application, design constraints, or other reasons, these elements can be implemented using electronic hardware, computer programs, or any combination thereof.
[0145] The electronic hardware may include microelectromechanical systems (MEMS), integrated circuits (such as application specific integrated circuits), microprocessors, microcontrollers, digital signal processors (DSP), field programmable gate arrays (FPGA), programmable logic devices (PLD), gated logic, discrete hardware circuits, printed circuit boards (PCB) (such as flexible PCB), and other suitable hardware configured to perform a plurality of different functions described in this specification, such as sensors, for example, for sensing and / or recording physical properties of the environment, devices, users, etc. A computer program should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, programs, functions, etc., whether called software, firmware, middleware, microcode, hardware description language or other names.
[0146] This application relates to the field of hearing devices such as hearing aids.
[0147] In real audio applications, speech enhancement and noise reduction are usually required, where the noise from the acoustic environment masking the desired speech signal usually results in reduced speech intelligibility. Examples of audio applications where noise reduction is beneficial are hands-free wireless communication devices, such as headsets, automatic speech recognition systems, and hearing aids (HA). Specifically, in applications such as headset communication devices where a ("remote") human listener needs to understand the noisy self-voice picked up by the microphone of the headset, the noise can greatly reduce the sound quality and speech intelligibility, making the conversation more difficult.
[0148] In this specification, a "headset application" may include a normal headset application for communicating with a "remote speaker" via a network (such as an office or call center application), and may also include a hearing aid application in a specific "communication or telephone mode" where the hearing aid is adapted to pick up the user's voice and transmit it to another device (such as a remote communication partner) while possibly receiving audio from other devices (such as a remote communication partner).
[0149] The noise reduction algorithms implemented in a multi-microphone device may include a set of linear filters, such as spatial filters and temporal filters for shaping the sounds picked up by these microphones. The spatial filter can change the sound by enhancing or attenuating the sound as a function of direction, and the temporal filter can change the frequency response of the noisy signal to enhance or attenuate specific frequencies. To find the optimal filter coefficients, it is usually necessary to know the noise characteristics of the acoustic environment. Unfortunately, these noise characteristics are usually unknown and need to be estimated online. The characteristics that are usually required as inputs to multi-channel noise reduction algorithms are, for example, the cross-power spectral densities (CPSDs) of the noise. For example, both the minimum variance distortionless response (MVDR) beamformer and the multi-channel Wiener filter (MWF) beamformer require the noise CPSDs, and these two are common beamformers implemented in multi-microphone noise reduction systems.
[0150] To estimate the noise statistics, researchers have developed a large number of estimators for the noise statistics, such as [1–5]. In [1,4], a maximum likelihood (ML) estimator of the noise CPSD matrix during speech is proposed, assuming that the noise CPSD matrix remains the same up to a scalar multiplier. This estimator performs well when the underlying structure of the noise CPSD matrix does not change over time, such as for cabin noise and homogeneous noise fields, but may fail in other cases. In many real acoustic environments, the underlying structure of the noise CPSD matrix cannot be assumed to be fixed, such as when there are significant, non-stationary interfering noise sources in the acoustic scene. Specifically, when the interference is a competing speaker, many noise reduction systems fail to effectively suppress the competing speaker because it is difficult to determine whether the self-voice or the competing speaker is the desired speech.
[0151] In Figure 1A is shown the environment of a hearing device user 1. The environment is shown as including the hearing device user 1, a target sound source 2, and a noise signal component 3.
[0152] The hearing device user 1 may wear a hearing device that includes a first microphone 4 and a second microphone 5 on the user 1's left ear and a third microphone 6 and a fourth microphone 7 on the user 1's right ear.
[0153] The target sound source 2 may be located near the hearing device user 1 and may be configured to generate a target speech signal and transmit it into the environment of the user 1. The target sound source 2 can be a person, a radio, a television, etc., which generates the target speech signal. The target speech signal may be directed towards the direction of the user 1 or may be directed away from the direction of the user 1.
[0154] The noise signal component 3 is shown surrounding the hearing device user 1 and the target sound source 2, thus causing the target sound source signal to be received at the hearing device user 1. The noise signal component can include local noise sources (such as machines, fans, etc.) and / or distributed (diffuse, homogeneous) noise sound sources.
[0155] Each of the first microphone 4, the second microphone 5, the third microphone 6, and the fourth microphone 7 can provide an electrical input signal including the target speech signal and the noise signal component 3.
[0156] In Figure 1B , the voice activity (VA) is shown as a function of time periods. It is assumed that the target sound source 2 and the user 1 are speaking back-to-back, i.e., there is no pause or only a minimal pause between the voices in the conversation. The user 1 is shown speaking during the time periods between t1 and t2 and between t5 and t6 (denoted as "self-voice"), while the target sound source 2 is shown speaking during the time periods between t3 and t4 and between t7 and t8 (denoted as "target sound source"). During the entire time period of Figure 1B , there is a noise signal with a noise level having random fluctuations (the solid curve denoted as "noise").
[0157] Figure 1C Shows how Figure 1B the exemplary voice activity can be detected using self-voice VAD (such as self-voice detector (OVD)) and using VAD (i.e., classical VAD).
[0158] The self-voice VAD can detect that the user is speaking during the time periods between t1 and t2 and between t5 and t6. On the other hand, this VAD will detect that speech is being generated (from both the user 1 and the target sound source 2) during the entire time period from t1 to t8. However, depending on the resolution of the VAD used, there may be small interruptions in the detected voice activity during the time periods between t2 and t3, between t4 and t5, and between t6 and t7.
[0159] Figure 1D Shows that the hearing device is capable of updating the noise reduction system to provide an estimate of the target speech signal and at least partially attenuate the noise signal component 3.
[0160] In the classical method ( Figure 1D the upper part of), the VAD can be used to detect the presence of speech, and the noise reduction system of the hearing device will be updated only during the time when no speech (from both the user 1 and the target sound source 2) is being generated, because the VAD cannot distinguish between the speech from the user 1 and the speech from the target sound source 2. Thus, the noise reduction system will be updated only during the time when the VAD does not detect speech, i.e., from t0 to t1 and after t8.
[0161] Using self-voice VAD ( Figure 1Dat the lower part), the noise reduction system of the hearing device can be updated not only when no speech is detected, but also when the self - voice VAD detects speech from user 1, i.e., from t0 to t2, from t5 to t6, and after t8.
[0162] Thus, the noise signal component can be identified during a time period (time interval) when the self - voice detector indicates that at least one electrical input signal or a signal derived therefrom is from the speech of user 1 or is from the speech of user 1 with a probability higher than the Over - Voice - Presence - Probability (OVPP) threshold, such as 60% or 70%.
[0163] In the hearing device, by combining the self - voice VAD and the VAD, the noise reduction system can be configured to detect both when user 1 is speaking and when the target sound source 2 is speaking. Thus, the noise reduction system can be updated during a time period when no speech signal is generated and during a time period when user 1 is speaking, but is prevented from being updated during a time period when only the target sound source 2 generates a target speech signal (speaks).
[0164] In Figure 2A , the environment of the hearing device user 1 is shown. The environment is shown as including the hearing device user 1, the competing speaker 8, and the noise signal component 3.
[0165] As in the case of Figure 1A , the hearing device user 1 can wear a hearing device, which includes a first microphone 4 and a second microphone 5 on the left ear of user 1 and a third microphone 6 and a fourth microphone 7 on the right ear of user 1.
[0166] The competing speaker 8 can be located near the hearing device user 1 and can be configured to generate a competing speech signal (i.e., an unwanted speech signal) and transmit it into the environment of user 1. The competing speaker 8 can be a person, a radio, a television, etc., which generates a competing speech signal. The competing speech signal can be directed towards the direction of user 1 or can be directed away from the direction of user 1.
[0167] The noise signal component 3 is shown surrounding the hearing device user 1 and the competing speaker 8, thus causing an estimate of the self - voice of user 1, i.e., the desired speech signal, to be received at the microphones 4, 5, 6, 7 of the hearing device (for example, in the case where the hearing device includes or implements headphones).
[0168] In Figure 2B , the voice activity (VA) is shown as a function of time period (time). It is assumed that user 1 speaks during the time period from t1 to t3, and the competing speaker 8 speaks during the time period from t2 to t4, whereby the speech of the competing speaker 8 overlaps with the speech of user 1. During the entire time period of Figure 2B , there is a noise signal with a noise level having random fluctuations.
[0169] Figure 2C shows Figure 2B how the exemplary voice activity of Figure 2C can be detected using self-voice VAD and using (general) VAD.
[0170] The self-voice VAD ( Figure 2C lower part of Figure 2C ) can detect that the user 1 is speaking during the time period between t1 and t3. On the other hand, this VAD ( Figure 2C upper part of Figure 2C ) will detect that speech is being generated (from user 1 and competing speaker 8) during the entire time period from t1 to t4.
[0171] Figure 2D shows that the hearing device can update the noise reduction system to provide an estimate of the target speech signal and at least partially attenuate the noise signal component 3.
[0172] In the classical method ( Figure 2D upper part of
[0172] ), the VAD is used to detect the presence of speech, and the noise reduction system of the hearing device will only be updated when no speech (from user 1 and competing speaker 8) is being generated, because the general VAD cannot distinguish between speech from user 1 and speech from competing speaker 8. Thus, the noise reduction system can only be updated when the VAD does not detect speech, i.e., from t0 to t1 (and after t4).
[0173] Using the self-voice VAD ( Figure 2D lower part of
[0172] ), the noise reduction system of the hearing device can be configured to update not only when no speech is detected, i.e., from t0 to t1 (and after t4), but also when the self-voice VAD detects speech from user 1, i.e., (in total) during the time from t0 to t3.
[0174] Thus, the noise signal component (including that from competing speaker 8) can be identified during the time period when the self-voice detector indicates that at least one electrical input signal or a signal derived therefrom is from the voice of user 1 or has a probability higher than the own voice presence probability (OVPP) threshold of being from the voice of user 1.
[0175] By combining the self-voice VAD and the general VAD in the hearing device, the noise reduction system can be configured to detect both when user 1 is speaking and when competing speaker 8 is speaking alone. Thus, the noise reduction system can be updated during time intervals when no speech signal is being generated and during time intervals when user 1 is speaking, but is prevented from being updated during time intervals when competing speaker 8 is generating a speech signal.
[0176] In Figure 3A
[0176] , the environment of the hearing device user 1 is shown. The environment is shown as including the hearing device user 1, the target sound source 2, the competing speaker 8, and the noise signal component 3.
[0177] InFigure 1A and Figure 2A As in the case of Figure 1A , a user 1 of a hearing device can wear a hearing device that includes a first microphone 4 and a second microphone 5 on the left ear of the user 1 and a third microphone 6 and a fourth microphone 7 on the right ear of the user 1.
[0178] A target sound source 2 and a competing talker 8 can be located near the user 1 of the hearing device and can be configured to generate voice signals and transmit them into the environment of the user 1. The target voice signal and / or the competing talker voice signal can be directed towards the direction of the user 1 or can be directed away from the direction of the user 1.
[0179] A noise signal component 3 is shown surrounding the user 1 of the hearing device, the competing talker 8, and the target sound source 2, and thus can affect the target sound source signal received at the user 1 of the hearing device.
[0180] The first microphone 4, the second microphone 5, the third microphone 6, and the fourth microphone 7 can provide an electrical input signal including the target voice signal, the competing talker signal, and the noise signal component 3.
[0181] In Figure 3B , voice activity (VA) is shown as a function of time intervals (time). It is assumed that the target sound source 2 and the user are speaking back to back, and the competing talker 8 overlaps with the voices of the target sound source 2 and the user 1. The user 1 is shown speaking during time intervals between t1 and t2 and between t5 and t6 (“self-voice”), while the target sound source 2 is shown speaking during time intervals between t3 and t4 and between t7 and t8 (“target sound source”). The competing talker 8 is shown speaking during a time interval between t1* and t7* (“competing talker”). During the entire time interval of Figure 3B , there is a noise signal (a solid curve labeled “noise”) with a noise level having random fluctuations.
[0182] Figure 3C Shows Figure 3B how the exemplary voice activity of
[0183]
[0184] Figure 3D Shows the time intervals during which the hearing device can update the noise reduction system to provide an estimate of the target voice signal and at least partially attenuate the noise signal component 3 including the competing talker signal.
[0185] In classical methods, VAD can be used to detect the presence of speech, and the noise reduction system of the hearing device will only be updated when there is no speech (from user 1, competing speaker 8, and from target sound source 2) because VAD cannot distinguish between speech from user 1, speech from competing speaker 8, and speech from target sound source 2. Thus, the noise reduction system will only be updated when VAD does not detect speech, that is, from t0 to t1 and after t8.
[0186] Using a self-voice VAD, the noise reduction system of the hearing device can be configured to be updated not only when no speech is detected, but also when the self-voice VAD detects speech from user 1, that is, updated during the time from t0 to t2, from t5 to t6, and after t8.
[0187] Thus, the noise signal component can be identified during the time period when the self-voice detector indicates that at least one electrical input signal or a signal derived therefrom is from the voice of user 1 or is from the voice of user 1 with a probability higher than the Over Voice Presence Probability (OVPP) threshold.
[0188] By combining a self-voice VAD and a general VAD in the hearing device, the noise reduction system can be configured to detect both when user 1 is speaking and when target sound source 2 and competing speaker 8 are speaking. Thus, the noise reduction system can be updated during the time intervals when no speech signal is generated and during the time intervals when user 1 is speaking, but is prevented from being updated during the time intervals when target sound source 2 generates a target speech signal.
[0189] In Figure 4A and 4B the noise reduction system NRS is connected to an input unit IU including M input transducers (IT1,…,IT M ) such as microphones, where M is greater than or equal to 2. The M input transducers can be located in a single hearing device such as a hearing aid (e.g., in or at the user's ear). The M input transducers can be distributed across two (separate) hearing devices such as hearing aids (e.g., in two hearing devices located in the user's two ears or at the ears). The latter configuration can form part of a binaural hearing system such as a binaural hearing aid system or constitute a binaural hearing system such as a binaural hearing aid system. Each hearing device of the binaural hearing aid system can include one or more (at least one) e.g., more than two input transducers (such as microphones). Figure 6 Fig. shows the microphone configuration of a binaural hearing aid system, where each hearing aid includes two microphones. Figure 5A , 5B , 5C show several different embodiments of a hearing device (such as a hearing aid) including a noise reduction system according to the present invention.
[0190] Figure 4AAn exemplary input unit IU is shown connected to an exemplary noise reduction system. Each of the M input transducers (at its respective, different location) receives a sound signal (s1, …, s M ) from an input sound field that includes ambient sound. The input unit IU includes M input subunits (IU1, …, IU M ). Each input unit includes input transducers (IT1, …, IT M ) such as microphones for converting the input sound signals into electrical input signals (s’1, …, s’ M ). Each input transducer may include an analog-to-digital converter for converting the analog input signal into a digital signal (at a certain sampling rate, such as 20 kHz or higher). Each input unit also includes an analysis filter bank for converting the time-domain (digital) signal into K (e.g., ≥16, or ≥24 or ≥64) sub-band signals (S1(k,n), …, S M (k,n), where k and n are the frequency index and time index, respectively, and where k = 1, …, K. The corresponding electrical input signals (S1(k,n), …, S M (k,n)) of the time-frequency representation (k,n) are fed to the noise reduction system NRS.
[0191] The noise reduction system NRS is configured to provide an estimate of a target speech signal (such as the self-voice of a hearing aid user and / or the voice of a target speaker in the user's environment) where the noise signal component is at least partially attenuated. The noise reduction system NRS includes a plurality of beamformers. The noise reduction system NRS includes a beamformer BF such as an MVDR beamformer or an MWF beamformer, which is connected to the input unit IU and is configured to receive the electrical input signals (S1(k,n), …, S M (k,n)) of the time-frequency representation. The beamformer BF is configured to provide at least one beamformed (spatially filtered) signal, such as an estimate of the target speech signal
[0192] Achieving directivity / directionality through beamforming is an effective way to attenuate unwanted noise because the gain that varies with direction can cancel out the noise from one direction while retaining the sound of interest coming from another direction, thereby potentially improving the intelligibility of the target speech signal (and thus providing spatial filtering). Generally, beamformers in hearing devices such as hearing aids have a beam pattern that is continuously adjusted to minimize the noise component while the sound coming from the target direction is not altered. Generally, the acoustic properties of the noise signal vary over time. Therefore, the noise reduction system is implemented as an adaptive system that adjusts the directional beam pattern to minimize the noise while the target sound (direction) is not changed.
[0193] Figure 4AThe noise reduction system NRS also includes a voice activity detector VAD for repeatedly estimating whether or not, or with what probability, at least one (most or all) electrical input signal or a signal derived therefrom includes speech. The electrical input signal (S1(k,n),…,S M (k,n)) or at least one of them (or a processed version thereof, such as a beamformed version) is fed to the VAD, on the basis of which a voice activity signal VA is provided indicating whether or not, or with what probability, the electrical input signal or its processed version includes speech. The VA is fed to an update unit UPD-C noise for updating the noise covariance matrix C noise . The noise covariance matrix is determined from the (noisy) electrical input signal (S1(k,n),…,S M (k,n)) in the absence of speech (at a given time point) (assuming that only noise is present in the sound field at these moments). The updated noise covariance matrix C noise (k,n) is used by an update filter weight unit UPD-W, where the updated filter weights W(k,n) at a given moment when the noise covariance matrix is updated are determined based on the latest noise covariance matrix C noise (k,n) and an estimate of the current relative or absolute acoustic transfer function from the target sound source to the input transducer of the input unit IU of the hearing system (or device) (e.g., in the view vector d(k,m) in the setting). The calculation of the noise covariance matrix C noise (k,n) and the beamformer weights W(k,n) is known in the prior art, for example, as described in
[11] and / or EP2701145A1. The updated beamformer weights W(k,n) are applied in the beamformer BF to the electrical input signal (S1(k,n),…,S M (k,n)), whereby an estimate of the target signal is provided
[0194] Figure 4B FIG. shows an exemplary input unit IU connected to an exemplary noise reduction system NRS according to the present invention. Figure 4B The embodiment of Figure 4A is substantially equivalent to Figure 4A the embodiment of M) or whether the signal derived therefrom includes speech from the user's voice and with what probability. Some acoustic events have distinct directional beam patterns and these acoustic events can be distinguished from other acoustic events. The self-voice of a hearing device user is an example of such an event. This is exploited in the present invention. By simultaneously monitoring the (general) voice presence (indicated by the voice activity signal VA from the VAD) and the (specific) self-voice presence (indicated by the self-voice activity signal OVA from the OVAD), another scheme for identifying a time period suitable for updating the noise covariance matrix C noise (k,n) can be advantageously used (different from the general absence of voice). As Figure 1A 、 2D 、as shown in the 3D example, the noise reduction system according to the present invention is configured to update the noise covariance matrix C noise (k,n) during the self-voice speech activity (possibly and during the general absence of speech). The update unit UPD-C noise may for example include a self-voice cancellation beamformer configured to cancel (or attenuate) the sound from the user's mouth while leaving the sound from other directions unchanged (or less attenuated). The update filter weights unit UPD-W may include the function of a (single-channel) post-filter, where in addition to the spatial filtering of the target signal, the noise component is also attenuated by the self-voice cancellation beamformer of the update unit UPD-C noise . The update filter weights unit UPD-W may receive or calculate the self-voice transfer function (mouth to microphone), for example set in the look vector d (see input d). The look vector may be determined before or during the operation of the hearing device. The look vector may be used to determine the current filter weights. The look vector may represent the transfer function or relative transfer function to the user's self-voice or to an external target sound source such as a target speaker in the environment. The look vector of the user's self-voice and the look vector of the environmental target speaker can both be provided to the noise reduction system or adaptively determined by the noise reduction system. The noise reduction system NRS may include a mode selection input ("mode") configured to indicate the operating mode and / or update strategy of the system such as the beamformer, for example whether the target signal is the user's self-voice or a target signal from the user's environment (and possibly indicate the direction or location of such a target sound source). The mode control signal may for example be provided from a user interface such as from a remote control device (for example implemented as an APP of a smart phone or a similar device such as a smart watch, etc.). The user interface may include a voice control interface (for example see Figure 5C ). The mode control signal may for example be automatically generated, for example using one or more sensors, for example starting from the reception of a wireless signal such as from a phone. The output of the beamformer BF can be an estimate of the user's voice or an estimate of the target sound from the environment for example see Figure 5B .
[0195] Figure 5A shows an exemplary block diagram of a hearing device, such as a hearing aid HD, which includes a noise reduction system NRS according to the present invention. The hearing device includes an input unit IU for picking up sound s from the environment in and providing M electrical input signals (S1, …, S M ), and a noise reduction system NRS for estimating the target signal Figure 4A , 4B in the input sound s based on the electrical input signals and optionally based on additional information (such as a mode control signal (“mode”)) described in in the input sound s. The hearing aid further includes a processor PRO for applying one or more processing algorithms to the signals in the forward path from the input to the output transducer (e.g., here applied to the estimated quantity of the target signal provided in a time-frequency representation ). One or more processing algorithms may include, for example, a compression algorithm configured to amplify (or attenuate) the signal according to the user's needs, thereby compensating for the user's hearing impairment, for example. Other processing algorithms may include frequency shifting, feedback control, etc. The processor provides a processed output OUT, which is fed to an output unit OU, and the output signal out is thus converted into a stimulus s that can be perceived by the user as sound out (perceived output sound), such as an acoustic vibration (in air and / or the skull) or an electrical stimulation of the cochlear nerve. In non-hearing aid applications such as a headset, the processor may be configured to further enhance the signal from the noise reduction system or may be omitted (such that the estimated quantity of the target signal is fed directly to the output unit). The target signal may be the user's own voice and / or a target sound in the user's environment (e.g., a person (different from the user) speaking, e.g., communicating with the user).
[0196] Figure 5B shows an exemplary block diagram of a hearing device, such as a hearing aid HD, which includes a noise reduction system NRS according to an embodiment of the present invention and operating in a hands-free telephone mode. Figure 5B The embodiments of Figure 5A include functional modules described in combination with the embodiments of Figure 5B . However, in particular, MIC the embodiments of and transmitting the estimator (self-speech audio) via a synthesis filter bank FBS and appropriate transmitter Tx and antenna circuitry to another device, such as a telephone or similar device, or system. Additionally, the hearing aid HD includes an auxiliary audio input (audio input) configured to receive a direct audio input from another device or system, such as a telephone (or similar device) (e.g., via wired or wireless means). In Figure 5B an embodiment, a wirelessly received input, such as a verbal communication from a communication partner, is shown to be received by the hearing aid via an antenna and an input unit IU AUX The auxiliary input unit IU AUX includes appropriate receiver circuitry, an analog-to-digital converter (if appropriate), and an analysis filter bank to provide an audio signal S aux in the time-frequency representation as sub-band signals S aux (k,n). Figure 5B The forward path of the hearing aid in Figure 5A an embodiment includes the same elements as described in the embodiment in combination with Figure 5A and additionally includes a selector-mixer SEL-MIX such that the signal in the forward path, which is processed in the processor PRO and presented to the user as a perceptible sound stimulus, is configurable. Under the control of a mode control signal, the output S x (k,n) of the selector-mixer SEL-MIX can be a) an environmental signal S ENV (k,n) (e.g., an estimator of a target signal in the environment, or an omnidirectional signal, such as from one of the microphones); b) an auxiliary input signal S aux (k,n) from another device; or c) a mixture thereof (e.g., a weighted mixture (possibly configurable via a user interface)). Additionally, compared to Figure 5A the embodiment in Figure 5B the forward path of the embodiment in Figure 5B includes a synthesis filter bank FBS configured to convert a signal in the time-frequency domain represented by a plurality of sub-band signals (here, the signal OUT(k,n) from the processor PRO) into a time-domain signal out. The hearing aid (forward path) also includes an output transducer OT for converting the output signal out into a stimulus s out perceptible as sound (output sound) by the user, such as acoustic vibrations (in air and / or in the skull). The output transducer OT may include a digital-to-analog converter, if appropriate.
[0197] A first noise reduction system NRS1 is configured to provide an estimator of the user's self-speech The first noise reduction system NRS1 may include a self-speech retention beamformer and a self-speech cancellation beamformer. The self-speech cancellation beamformer includes a noise source when the user is speaking.
[0198] The second noise reduction system NRS2 is configured to provide an estimate of the target sound source (e.g., the voice of a speaker in the user's environment) ). The second noise reduction system NRS2 may include an ambient target sound source hold beamformer and an ambient target sound source cancellation beamformer and / or a self-voice cancellation beamformer. The target cancellation beamformer includes the noise source when the target speaker is speaking. The self-voice cancellation beamformer includes the noise source when the user is speaking.
[0199] Figure 5B May represent a general headphone application, e.g., by separating the microphone-to-transmitter path IU MIC -Tx from the direct audio input to the speaker path IU AUX -OT. This can be done in several ways, e.g., by removing the second noise reduction system NRS2 and the selector-mixer SEL-MIX, and possibly removing the synthesis filter bank FBS (if the auxiliary input signal S aux is processed in the time domain), so as to feed the auxiliary input signal S aux directly to the processor PRO, which may or (generally) may not be configured to compensate for the user's hearing impairment.
[0200] Figure 5C An exemplary block diagram of a hearing aid including a noise reduction system according to the present invention is shown, which includes a voice control interface. Figure 5C Embodiments of include the same forward path as Figure 5B embodiments of, except that in Figure 5C embodiments, the option of including (e.g., wirelessly received) auxiliary audio signals in the beamformed signal consisting of the electrical input signal from the input transducer is omitted. In another embodiment, Figure 5B and 5C embodiments may be mixed such that Figure 5C the hearing aid additionally includes an auxiliary input from another device and the option of transmitting the self-voice signal to another device (to implement a communication mode) may also be implemented. The start (or termination) of a communication mode (such as a telephone mode) may be provided, for example, via a voice interface such as a voice control signal Vctr. In Figure 5C embodiments, the estimate of the user's self-voice provided by the first noise reduction system NRS1 is used as an input to the voice control interface VCI. The voice control interface VCI may, for example, be based on a wake word (spoken by the user and estimated from the user's voice initiated by the detection of (extraction). When the voice control interface is initiated, one command word among a plurality of predetermined command words can be extracted, and a control signal (VCtr, xVCtr) can be generated according to it. The functions of the hearing aid (e.g., implemented by the processor PRO) can be controlled via the voice interface VCI, referring to the signal Vctr. The extracted wake word (e.g., "Hey Siri", "Hey Google" or "OK Google", "Alexa", "X Oticon", etc.) and / or command word can be passed to another device (e.g., a smart phone or other voice-controllable device), referring to the control signal xvctr passed to another device via (optionally, the synthesis filter bank FBS and) the antenna and transceiver circuit TX.
[0201] Example 1
[0202] In the present application, a maximum likelihood (ML) estimator of the noise CPSD matrix is disclosed, which overcomes the limitations of the methods proposed in [1,4] (e.g., when there are significant interferences in the acoustic environment). An extended noise CPSD matrix model is proposed. Below, the signal model of the noisy observations in the acoustic scenario is presented. Based on this signal model, the ML estimator of the interference + noise CPSD matrix is derived, and the proposed method is illustrated by applying it to self-voice retrieval.
[0203] The acoustic scenario consists of a user equipped with multiple hearing aids or a headset that can access at least M > 2 microphones. These microphones pick up sounds from the environment, and the noisy signals are sampled as discrete sequences For all m = 1, …, M microphones, As Figure 6 shown, the user is active in this acoustic scenario, and the desired clean speech signal (which we call self-voice) generated by the user is defined as the discrete sequence s o (t). The interference is modeled as a point source (called v c (t)), and the noise in the acoustic environment is v e,m (t). The noisy signal picked up by the microphone is then the sum of all three components, i.e.,
[0204] x m (t) = s o (t) * d o,m (t) + v c (t) * d m (t, θ c ) + v e,m (t), (1)
[0205] where * refers to convolution, d o,m (t) is the relative impulse response between the m-th microphone and the self-voice source, dm (t, θ c ) is the relative impulse response between the m-th microphone and the interference arriving from direction θ c ∈ Θ, where, without loss of generality, we assume that Θ is a discrete set of directions, Θ = {-180°, …, 180}, with I elements. The goal of the noise reduction system is then to retrieve s m (t) from the noisy observation x o (t).
[0206] We apply the short-time Fourier transform (STFT) to x m (t) to transform the noisy signal into the time-frequency (TF) domain, with frame length T, decimation factor D, and analysis window w A (t), such that
[0207]
[0208] is the TF-domain representation of the noisy signal, where k is the frequency bin index, and n is the frame index. The signal model of the noisy observation in the TF domain then becomes
[0209]
[0210] For convenience, the noisy observation is vectorized such that x(k, n) = [x1(k, n), …, x M (k, n)] T and
[0211]
[0212] We further assume that the relative transfer function (RTF) vectors (i.e., d o (k, n) and d(k, n, θ c )) remain the same over time, so that we can define and In practice, it is usually the case that s o (k, n), v c (k, n), and v e (k, n) are uncorrelated random processes, meaning that the CPSD matrix of the noisy observation, i.e., is given by
[0213]
[0214] where λ s (k, n), λ c (k, n), and λe (k,n) are the power spectral densities (PSDs) of the self-talk, interference, and noise, respectively. Γ e (k,n) is the normalized noise CPSD matrix, 1 is the reference microphone index, and we assume that Γ e (k,n) is a known matrix, but for an approximately homogeneous noise field, it can be modeled as
[0215]
[0216] We assume that the self-talk RTF vector d o (k) is known because it can be measured in advance before deployment. The remaining parameters to be estimated are λ c (k,n), λ e (k,n), and θ c , and the proposed ML estimators for these parameters will be presented in the following section.
[0217] To estimate the interference + noise PSDs, i.e., λ c (k,n) and λ e (k,n), and the interference direction θ c , we first apply a self-talk cancellation beamformer to obtain the signal of only interference + noise (e.g., signals from self-talk and competing speakers). The self-talk cancellation beamformer is implemented using the self-talk blocking matrix B o (k). A common method to find the self-talk blocking matrix is to first find the orthogonal projection matrix of d o (k), and then select the first M - 1 column vectors of this projection matrix. More clearly, let I M×M be the M×M identity matrix, then I M×M-1 is the first M - 1 column vectors of I M×M . The self-talk blocking matrix is then given by
[0218]
[0219] where B o (k) ∈ C M×M-1 . The self-talk blocked signal z(k,n) can be expressed as
[0220]
[0221] and the self-talk blocked CPSD matrix is
[0222]
[0223] In proposing λ c (k,n), λ e (k,n), and θ cBefore the ML estimator, we introduce the self - voice + interference blocking matrix
[0224] This step is necessary because the noise PSD λ e (k,n) of the ML estimator also requires the interference to be removed from the signal z(k,n) blocked by the self - voice. Forming the self - voice + interference blocking matrix follows a procedure similar to forming the self - voice blocking matrix. The self - voice + interference blocking matrix can be
[0225]
[0226] where Self - voice + interference blocking matrix is a function of direction because the direction of the interference is generally unknown. The self - voice + interference blocked signal is then
[0227]
[0228] and the blocked self - voice + interference CPSD matrix is
[0229]
[0230] Only when θ i = θ c at that time.
[0231] It is common to assume that the self - voice, interference, and noise are uncorrelated in time [6]. Under this assumption, the blocked self - voice + interference signal is distributed according to a circularly symmetric complex Gaussian distribution, that is which means the likelihood function of the N observations of z(k,n) is given by
[0232]
[0233] tr(·) refers to the trace operator, and is the sample estimator of the CPSD matrix blocked by the self - voice. The interference + noise PSD λ c (k,n) and λ e (k,n) of the ML estimators have been derived in [1,4]. The ML estimator of λ e (k,n) is given by
[0234]
[0235] is the sample covariance of the self - voice + interference blocked signal, and the ML estimator of the interference PSD is then as given in [7]
[0236]
[0237] wherein is the MVDR beamformer constructed from the blocked self-speech CPSD matrix, i.e.,
[0238]
[0239] Inserting the ML estimator and into the likelihood function, we obtain the concentrated likelihood function which we simplify to Commonly, the log-likelihood function is maximized by applying the natural logarithm function to the concentrated likelihood function. It can then be shown that the concentrated log-likelihood function is proportional to [8, 9].
[0240]
[0241] Under the assumption that there is only a single interference in the acoustic environment and the noisy observations across frequency windows are uncorrelated, the broadband concentrated log-likelihood function can then be obtained
[0242]
[0243] where K is the total number of frequency windows of the one-sided spectrum. To obtain the ML estimator of the interference direction, we maximize the following function
[0244]
[0245] Since θ i belongs to a discrete set of directions, the ML estimator of θ c is obtained by an exhaustive search across θ i Finally, to obtain the estimator of the interference + noise CPSD matrix, we insert the ML estimator into the interference + noise CPSD model, i.e.,
[0246]
[0247] For self-speech retrieval, we implement the MWF beamformer. As is well known, the MWF can be decomposed into an MVDR beamformer and a single-channel post-Zener filter. The MVDR beamformer is given by
[0248]
[0249] and the single-channel post-Zener filter is
[0250]
[0251] The MWF beamformer coefficients are then
[0252] w MWF (k, n) = w MVDR (k, n) · g(k, n). (23) Finally, the self-speech signal can be estimated as a linear combination of the noisy observations using a beamformer weight, i.e.,
[0253]
[0254] The enhanced TF-domain signal y(k, n) is then transformed back to the time domain using an inverse STFT such that y(t) is the retrieved self-speech time-domain signal.
[0255] When appropriately substituted by corresponding processes, the structural features of the apparatus described above, detailed in the "Detailed Description" and defined in the claims, can be combined with the steps of the method of the present invention.
[0256] Unless otherwise specified, the singular forms "a", "the" used herein are intended to include the plural forms (i.e., having the meaning of "at least one"). It should be further understood that the terms "having", "including" and / or "comprising" used in the specification indicate the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations. It should be understood that unless otherwise specified, when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may be present. As used herein, the term "and / or" includes any and all combinations of one or more of the listed related items. Unless otherwise specified, the steps of any method disclosed herein do not have to be performed in the exact order disclosed.
[0257] It should be appreciated that references to "an embodiment" or "embodiments" or "aspect" or "may" include features described in connection with that embodiment are intended to be included in at least one embodiment of the present invention. Moreover, the particular features, structures, or characteristics may be appropriately combined in one or more embodiments of the present invention. The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.
[0258] The claims are not limited to the aspects shown herein, but encompass the full scope consistent with the claim language, where unless otherwise specified, an element recited in the singular does not mean "one and only one" but rather "one or more". Unless otherwise specified, the term "some" means one or more.
[0259] Accordingly, the scope of the present invention should be determined in accordance with the claims.
[0260] References
[0261] [1] U.Kjems and J.Jensen, “Maximum likelihood based noise covariance matrix estimation for multimicrophone speech enhancement,” in 2012 Proceedings of the 20th European Signal Processing Conference (EUSIPCO), Aug 2012, pp.295–299.
[0262] [2] Yujie Gu and A.Leshem, “Robust Adaptive Beamforming Based on Interference Covariance Matrix Reconstruction and Steering Vector Estimation,” IEEE Transactions on Signal Processing, vol.60, no.7, pp.3881–3885, July 2012.
[0263] [3] Richard C.Hendriks and Timo Gerkmann, “Estimation of the noise correlation matrix,” in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Prague, Czech Republic, May 2011, pp.4740–4743, IEEE.
[0264] [4] Jesper Jensen and Michael Syskind Pedersen, “Analysis of beamformer directed single-channel noise reduction system for hearing aid applications,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Queensland, Australia, Apr. 2015, pp. 5728–5732, IEEE.
[0265] [5] Mehrez Souden, Jingdong Chen, Jacob Benesty, and Sofi`ene Affes, “An Integrated Solution for Online Multichannel Noise Tracking and Reduction,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 7, pp. 2159–2169, Sept. 2011.
[0266] [6] K.L. Bell, Y. Ephraim, and H.L. Van Trees, “A Bayesian approach to robust adaptive beamforming,” IEEE Transactions on Signal Processing, vol. 48, no. 2, pp. 386–398, Feb. 2000.
[0267] [7]Adam Kuklasinski, Simon Doclo, Timo Gerkmann, Soren Holdt Jensen, and Jesper Jensen, “Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparison,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Queensland, Australia, Apr. 2015, pp. 91–95, IEEE.
[0268] [8]Mehdi Zohourian, Gerald Enzner, and Rainer Martin, “Binaural Speaker Localization Integrated Into an Adaptive Beamformer for Hearing Aids,” IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 3, pp. 515–528, Mar. 2018.
[0269] [9]Hao Ye and D. DeGroat, “Maximum likelihood DOA estimation and asymptotic Cramer-Rao bounds for additive unknown colored noise,” IEEE Transactions on Signal Processing, vol. 43, no. 4, pp. 938–949, Apr. 1995.
[0270]
[10] Michael Brandstein and Darren Ward, Microphone Arrays: Signal Processing Techniques and Applications, 2001.
[0271]
[11] EP2701145A1 (Retune, Oticon) February 26, 2014.
Claims
1. A hearing device adapted to be located at or in a user's ear or adapted to be fully or partially implanted in a user's head, the hearing device comprising: An input unit for providing at least one electrical input signal representative of sounds in the user's environment, the electrical input signal comprising a target speech signal from a target sound source and additional signal components, i.e., noise signal components, from one or more other sound sources; And A noise reduction system for providing an estimate of the target speech signal, wherein the noise signal components are at least partially attenuated, the noise reduction system comprising a self-voice detector and a voice activity detector; Wherein the self-voice detector is adapted to repeatedly estimate whether or with what probability at least one electrical input signal or a signal derived therefrom comprises speech originating from the user's voice; Wherein the voice activity detector is adapted to repeatedly estimate whether or with what probability at least one electrical input signal or a signal derived therefrom comprises speech; Wherein the noise signal components are identified during a time period when the self-voice detector indicates that at least one electrical input signal or a signal derived therefrom originates from the user's voice or originates from the user's voice with a probability higher than a self-voice presence probability threshold; Wherein the noise reduction system is updated during a time interval when the voice activity detector estimates that there is no speech and during a time interval when the self-voice detector estimates that the speech originates from the user's voice.
2. The hearing device according to claim 1, wherein the input unit comprises a microphone, each microphone providing an electrical input signal comprising a target speech signal and a noise signal component.
3. The hearing device according to claim 1, the noise reduction system comprising one or more beamformers, wherein the input unit is configured to provide at least two electrical input signals connected to the one or more beamformers, and wherein the one or more beamformers are configured to provide at least one beamformed signal.
4. The hearing device according to claim 3, wherein the one or more beamformers comprise one or more self-voice cancellation beamformers configured to attenuate signal components originating from the user's mouth while signal components from all other directions remain unchanged or are attenuated less than the attenuation of the signal components originating from the user's mouth.
5. The hearing device according to claim 1, wherein the noise signal components are additionally identified during a time period when the voice activity detector indicates that there is no speech in at least one electrical input signal or a signal derived therefrom or there is speech with a probability lower than a speech presence probability threshold.
6. The hearing device according to claim 1, comprising a voice interface for voice control of the hearing device or other device or system.
7. The hearing device according to any one of claims 1-6, wherein the target speech signal from the target sound source comprises a self-voice speech signal of the hearing device user.
8. The hearing device according to any one of claims 1-6, wherein the target sound source comprises an external speaker in the hearing device user's environment.
9. The hearing device according to claim 1, which is constituted by a hearing aid, a head-mounted earphone, an active ear protection device or a combination thereof, or includes a hearing aid, a head-mounted earphone, an active ear protection device or a combination thereof.
10. The hearing device according to claim 1, wherein the hearing device further includes a timer configured to determine an overlapping time period between a self-voice speech signal and another speech signal.
11. The hearing device according to claim 10, wherein the hearing device is configured to determine whether the overlapping time period exceeds a time limit, and if so, mark the other speech signal as part of a noise signal component.
12. A method of operating a hearing device, the hearing device being adapted to be located at or in a user's ear or to be fully or partially implanted in the user's head, the method comprising: providing at least one electrical input signal representative of sounds in the user's environment, the electrical input signal including a target speech signal from a target sound source and additional signal components from one or more other sound sources, namely noise signal components; providing a noise reduction system for providing an estimate of the target speech signal, wherein the noise signal components are at least partially attenuated, the noise reduction system including a self-voice detector and a voice activity detector, the self-voice detector being used to repeatedly estimate whether or with what probability at least one electrical input signal or a signal derived therefrom includes speech originating from the user's voice, and the voice activity detector being used to repeatedly estimate whether or with what probability at least one electrical input signal or a signal derived therefrom includes speech; identifying the noise signal components during a time period when the self-voice detector indicates that at least one electrical input signal or a signal derived therefrom originates from the user's voice or originates from the user's voice with a probability higher than a self-voice presence probability threshold; updating the noise reduction system during a time interval when the voice activity detector estimates that there is no speech and during a time interval when the self-voice detector estimates that the speech originates from the user's voice.
13. A binaural hearing system, including first and second hearing devices according to any one of claims 1-11, the binaural hearing system being configured to enable data exchange between the first and second hearing devices.
14. A computer-readable medium having stored thereon a computer program including instructions that, when executed by a computer, cause the computer to perform the steps of the method according to claim 12.
Citation Information
Patent Citations
Noise estimation for use with noise reduction and echo cancellation in personal communication
EP2701145A1
System and method of detecting a user's voice activity using an accelerometer
US20140093091A1