Hearing system comprising a hearing instrument and method of operating the hearing instrument

WO2026176118A1PCT designated stage Publication Date: 2026-08-27WS AUDIOLOGY AS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/054960
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-24
Filing Date
2026-02-24
Publication Date
2026-08-27

Smart Images

  • Figure EP2026054960_27082026_PF_FP_ABST
    Figure EP2026054960_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a hearing system (2) with a hearing instrument (4) to be worn in or at an ear of a user and a method for operating a hearing instrument (4). A sound signal (I) is captured from an environment of the hearing instrument (4) and processed to derive a processed sound signal (O) which is output to a user of the hearing instrument (2). Said processing includes applying a noise suppression algorithm by which speech peaks (SP) in each of which the captured sound signal (I) contains sound of a spoken consonant or a spoken consonant of a predefined consonant type are detected, and by which a noise suppression gain (gs) is applied to the captured sound signal (I) or a signal derived there-from, such that the processed sound signal (O) is attenuated to a noise suppression level outside said speech peaks (SP). Said processing also includes applying a consonant extension algorithm by which, at the end of each speech peak (SP), the signal level of the processed sound signal (O) is maintained or gradually reduced to the noise suppression level within a predefined extension time (Ts).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] Hearing system comprising a hearing instrument and method of operating the hearing instrument

[0003] The invention relates to a method for operating a hearing instrument. The invention further relates to a hearing system comprising a hearing instrument.

[0004] In general, a hearing instrument is an electronic device being designed to support the hearing of person wearing it (which person is called the “user” or “wearer” of the hearing instrument). In particular, the invention relates to hearing instruments that are specifically configured to at least partially compensate a hearing impairment of a hearing-impaired user. Such hearing instruments are also called “hearing aids”. In addition to such hearing aids, there are hearing instruments that are designed to support the hearing of normal-hearing users (i.e. persons without a hearing impairment). Such hearing instruments, being sometimes referred to as “Personal Sound Amplification Products” (PSAP), may be provided, e.g., to enhance the hearing of the wearer in complex acoustic environments or to protect the hearing of the wearer from damage or overstress.

[0005] Hearing instruments, in particular hearing aids, are typically designed to be worn in or at an ear of the user, e.g. as a Behind-The-Ear (BTE) or In-The-Ear (ITE) device. With respect to its internal structure, a hearing instrument normally comprises at least one (acousto-electrical) input transducer, a signal processor, and an output transducer. During operation of the hearing instrument, the at least one input transducer captures a sound signal from an environment of the hearing instrument and converts it into an input audio signal (i.e. an electrical signal transporting a sound information) that is also referred to as the “captured sound signal”. In the signal processor, the input audio signal is processed, in particular amplified dependent on frequency, e.g., to compensate the hearingimpairment of the user. The signal processor outputs the processed sound signal, also referred to as the output audio signal, to the output transducer. Most often, the output transducer is an electro-acoustic transducer (also called “receiver”) that converts the output audio signal into a processed airborne sound which is emitted into the ear canal of the user. Alternatively, the output transducer may be an electro-mechanical transducerthat converts the output audio signal into a structure-borne sound (vibrations) that is transmitted, e.g., to the cranial bone of the user.

[0006] Furthermore, besides classical hearing instruments, there are implanted hearing aids such as cochlear implants, and hearing instruments the output transducers of which directly stimulate the auditory nerve of the user.

[0007] The term “hearing system” denotes one device or an assembly of devices and / or other structures providing functions required for the operation of a hearing instrument. A hearing system may consist of a single stand-alone hearing instrument. As an alternative, a hearing system may comprise a hearing instrument and at least one further electronic device which may, e.g., be one of another hearing instrument for the other ear of the user, a remote control, and a programming tool for the hearing instrument. Moreover, modem hearing systems often comprise a hearing instrument and a software application (hearing app) for controlling, programming and / or otherwise supporting the hearing instrument’s operation, which software application is or can be installed on a computer or a mobile communication device such as a mobile phone (smart phone). In the latter case, typically, the computer or the mobile communication device are not a part of the hearing system. In particular, most often, the computer or the mobile communication device will be manufactured and sold independently of the hearing system.

[0008] A problem with conventional hearing instruments is that many users, in particular hearing impaired users of hearing aids, have problems following speech in noise. Often, those users cannot filter out the consonants from the noise which causes problems to understand speech and follow the rhythm of the speech. A conventional way to address this problem and, thus, to allow for better speech understanding is to apply a noise suppression algorithm that specifically reduces non-speech sound in the captured sound signal. Such noise suppression algorithms try to keep the consonants in speech sound unattenuated and only attenuate noise. Conventional noise suppression algorithms normally analyze the captured sound signal for presence of speech, in particular spoken consonants. A low level gain (subsequently referred to as the “noise suppression gain”) attenuating the processed sound signal, is applied to the captured sound signal when and as long as no speech, in particular no spoken consonants are detected.However, conventional noise suppression algorithms often fail to improve speech intelligibility (speech understanding) in noise to a satisfying extent as they normally induce new artefacts and disturbing effects in the processed sound signal while reducing noise. In particular, most noise suppression algorithms produce the perception of chopping the speech signal, thus particularly hampering correct understanding of consonants.

[0009] It is therefore an object of the present invention to improve speech intelligibility of the processed sound created by a hearing instrument, especially in noisy sound environments.

[0010] According to a first aspect of the invention, the above object is met by a method as defined by claim 1 for operating a hearing instrument. Preferred embodiments of the method are defined in the dependent claims 2 to 6 and in the subsequent description.

[0011] In accordance with the method, a sound signal is captured from an environment of the hearing instrument. The captured sound signal is processed (by applying at least one signal processing step) to thereby derive, directly or indirectly, a processed sound signal. Said processed sound signal is output to a (human) user of the hearing instrument.

[0012] Said processing includes applying a noise suppression algorithm by which speech peaks are detected, and by which a noise suppression gain is applied to the captured sound signal or a signal derived therefrom, such that the processed sound signal is attenuated to a noise suppression level outside said peaks. The noise suppression algorithm, thus, attenuates the sound signal to a greater extend (i.e. more strongly) or amplifies the sound signal less strongly outside peaks than during speech peaks. In other words, application of said noise suppression algorithm reduces the level of an overall gain applied to the captured sound signal outside speech peaks, as compared to a level of said overall gain applied during said speech peaks.

[0013] Herein and subsequently, the term “speech peak” denotes a time sequence of the captured sound signal in which the latter contains sound of a spoken consonant. Herein, the noise suppression algorithm may be sensitive for all consonants (and, thus, not distinguish between different kinds of consonants). In this case, any spoken consonant is attributed to a speech peak. In alternative implementations of the invention, the noise suppressionalgorithm may be configured to be sensitive for at least one specific predefined consonant type (e.g. fricatives and / or affricates, etc.) only. In this case, only speech sound of the respective consonant type(s) involves a speech peak detected by the noise suppression algorithm.

[0014] In advantageous implementations of the invention, a speech peak is recognized by detecting an onset, i.e. a sudden increase of the signal level of the captured sound signal in a frequency range that corresponds to a characteristic frequency range of the respective consonant type. For instance, the predefined consonant type may include at least one of

[0015] - fricatives (“s”, “z”, “f’, “v”, etc.) having characteristic frequencies in a range between 4 kHz and 8 kH,

[0016] - affricates (“ds”, “pf ’, “tsh”, etc.) having characteristic frequencies in a range be-tween 3 kHz and 6 kH,

[0017] - plosives (“p”, “b”, “f ’, “d”, “k”, “g”, etc.) having characteristic frequencies in a range between 2 kHz and 3.5 kH,

[0018] - liquids (“1”, “r”, etc.) having characteristic frequencies in a range between 1.5 kHz and 3 kH,

[0019] - nasals (“m”, “n”, etc.) having characteristic frequencies in a range between 1 kHz and 2 kH,

[0020] - etc.

[0021] Herein, preferably, a spoken consonant and a corresponding speech peak are only detected if the onset, i.e. a sudden increase of the signal level of the captured sound signal, is limited to the characteristic frequencies of the respective phoneme and, thus, is not detected outside the characteristic frequencies.

[0022] Additionally or alternatively, at least one further consonant detection algorithm can be applied to the captured sound signal in order to identify a spoken consonant and a corresponding speech peak. Such consonant detection algorithms that can be used in the present method, within the scope of the invention, are known from- Pl: M. Yurt, A.N. Escalante B., V. I. Morgenshtern: “Fricative phoneme detection with zero delay”, ICLR 2020 Conference Blind Submission, 25 Sept 2019 (modified: 05 May 2023); to download (17.01.2025): https: / / open- review.net / attachment?id=BJxlmeBKwS&name=original_pdf

[0023] - P2: M. Yurt, et al., „Fricative phoneme detection using deep neural networks and its comparison to traditional methods“, Interspeech 2021, DOI:10.21437 / Inter- speech.2021-645; to download (17.01.2025): https: / / www.isca-archive.org / in- terspeech_202 l / yurt2 l_interspeech.pdf

[0024] - P3: C. Glackin, et al., „ Convolutional neural networks for phoneme recognition^ Int. Conf, on Pattern Recognition Applications and Methods 2018,

[0025] DOI: 10.5220 / 0006653001900195; to download (17.01.2025): https: / / pdfs.se- manticscholar.org / ec9b / 40428 lce4c3b6a5acde48ba0ec533da0a7e3.pdf.

[0026] In general, within the scope of the invention, said noise suppression gain may have a single scalar value that is applied uniformly to all frequencies of the captured sound signal. However, preferably, the noise suppression gain has a frequency dependence and, thus, has different values for different frequencies. Herein, the noise suppression gain may be provided, e.g., as a frequency-dependent mathematical function or as a vector having different values for different frequency bands.

[0027] In order to attenuate the processed sound signal, the noise suppression gain may be provided with a value between 0 and 1, if the noise suppression gain is multiplied to the captured sound signal or a signal derived therefrom. In alternative implementations, the noise suppression gain may be provided with a low positive value greater than 1 and may be applied such that it replaces a stronger gain applied during speech peaks.

[0028] Furthermore, according to the invention, said processing includes a consonant extension algorithm by which, at the end of each speech peak, the signal level of the processed sound signal is maintained or gradually reduced to the noise suppression level within a predefined extension time (that starts at the end of the speech peak). The consonant extension algorithm delays the effect of the noise suppression after a speech peak for theduration of the extension time and, thus, artificially extends (prolongs) the signature of the spoken consonant in the processed sound signal. This effect is limited to the duration of the predefined extension time. Thus, at the end of the extension time, the processed sound signal reaches the noise suppression level. The described prolonging of the consonant sound reduces artefacts and disturbing effects in the processed audio signal that would be induced otherwise by the noise suppression algorithm. In particular, the perception of chopping the speech signal, as is typically produced by conventional noise suppression algorithms, is effectively alleviated. In essence, the invention uses background noise as a vocoder, to prolong consonants in the captured speech sound to make them more audible to the user. As a consequence, the sound of spoken consonants in the processed sound signal can be understood and distinguished more easily by the user. Thus, speech intelligibility of the processed sound created by the hearing instrument, especially in noisy sound environments, is significantly improved.

[0029] In preferred embodiments of the invention, the consonant extension algorithm maintains or gradually reduces the signal level of the processed sound signal by applying a further gain (referred to as the “consonant extension gain”) to the captured sound signal or a signal derived therefrom. Alternatively or additionally, the consonant extension algorithm applies an additional noise to the captured sound signal or a signal derived therefrom. In both cases, the consonant extension algorithm thereby increases the level of the processed sound signal as compared to the noise suppression level. The consonant extension gain and / or the additional noise are applied for the duration of the extension time only. Thus, at the end of the extension time, the consonant extension gain and / or the additional noise are switched off or set to a neutral value in which they no longer influence the processed sound signal.

[0030] Preferably, the extension time is between 5 and 50 msec, in particular 30 msec.

[0031] In different embodiments of the invention, the consonant extension algorithm may be implemented as a separate algorithm or as a part of another algorithm, e.g., as a part of the noise suppression algorithm. In the latter case, e.g., the consonant extension gain and the noise suppression gain described above may be combined into one time-dependent gam.In general, within the scope of the invention, the consonant extension gain and / or the level of the additional noise may be set to zero or another neutral value abruptly at the end of the extension time. However, in more preferred embodiments of the invention, the consonant extension gain and / or the level of the additional noise are smoothly reduced during the extension time, such that they reach zero (or another neutral value) at the end of the latter. In other words, the effect of the consonant extension gain and / or the additional noise is faded out successively. This helps to alleviate artefacts or disturbing effects in the processed sound signal and further improves speech intelligibility in the processed sound.

[0032] In a particularly advantageous embodiment of the invention, the consonant extension gain and / or the additional noise are adjusted such that they overcompensate the attenuation effected by the noise suppression gain, at least in an initial phase of the extension time. The consonant extension gain and / or the additional noise, thus, increase an overall gain applied to the captured sound signal at an onset of the extension time, i.e. upon detection the end of a speech peak. In a particularly advantageous implementation of the invention, the consonant extension gain and / or the additional noise are adjusted, at the onset of the extension time, such that the level of the processed sound signal is kept constant, as compared to the level of the processed sound signal at the end of a preceding speech peak. Thus, a smooth transition of the signal level of the processed sound between a speech peak and the subsequent extension time is produced. This also helps to alleviate artefacts or disturbing effects in the processed sound signal and further improves speech intelligibility in the processed sound.

[0033] In preferred embodiments of the invention, the processing of the captured sound signal includes splitting the captured sound signal into a plurality of frequency band signals. Herein, preferably, the consonant extension algorithm is applied exclusively to frequency band signals within a predefined frequency sub-range of the frequencies included in the captured sound signal. Herein and subsequently, the term “frequency sub-range” denotes a range of frequencies being comprised in but smaller than the entire frequency range of the captured or processed sound signal. Preferably, said frequency sub-range corresponds to the range of characteristic frequencies of the consonants, i.e., depending on the specificimplementation, characteristic frequencies of consonants in general or one or more specific consonant types to be extended (e.g. to 4 kHz to 8 kH for extension of fricatives).

[0034] In simple implementations, within the scope of the invention, the method can be executed independently of whether or not the captured sound signal contains speech. In such implementations, the method will improve speech intelligibility whenever there is speech. In non-speech intervals (in which the captured sound signal does not include speech) the method will run idle as, ideally, no speech peaks will be recognized or extended. However, if the consonant extension algorithm is performed during non-speech intervals also, this may lead to artefacts caused by extension of falsely recognized speech peaks which may lead to suboptimal sound quality. Thus, in more refined implementations of the method, the consonant extension algorithm (and, preferably, the noise suppression algorithm) are performed in speech intervals only, i.e., in time intervals in which the captured sound signal contains speech. Herein, the presence of speech is detected by analyzing the captured sound signal, i.e., by applying a voice recognition algorithm to the captured sound signal. Preferably, the voice recognition algorithm is configured to distinguish the own voice of the user of the hearing instrument from foreign voice (i.e. voice of a speaker different from the user). In this case, preferably, the consonant extension algorithm (and, preferably, the noise suppression algorithm) are performed in foreign speech intervals only, i.e., in time intervals in which the captured sound signal contains speech of a speaker different from the user.

[0035] According to a second aspect of the invention, as specified in claim 7, a hearing system with a hearing instrument (as previously specified) is provided. Preferred embodiments of the hearing system are defined in the dependent claims 8 to 12 and in the subsequent description.

[0036] The hearing instrument comprises

[0037] - at least one input transducer configured to capture an (original) sound signal from an environment of the hearing instrument;

[0038] a signal processor configured to process the captured sound signal to thereby derive a processed sound signal; and- an output transducer configured to output the processed sound signal or a signal derived therefrom to a user of the hearing instrument.

[0039] In particular, the input transducer converts the original sound signal into an input audio signal (containing information on the captured sound) that is fed to the signal processor, and the signal processor outputs an output audio signal (containing information on the processed sound) to the output transducer which converts the output audio signal into airborne sound, structure-borne sound or into a signal directly stimulating the auditory nerve.

[0040] In general, the hearing system according to the second aspect of the invention is configured to automatically perform the method according to the first aspect of the invention. To this end, the hearing system comprises:

[0041] - a noise suppression module configured to execute the noise suppression algorithm described above, i.e. to detect speech peaks (as defined above), and to apply the suppression gain (as defined above) to the captured sound signal or a signal derived therefrom, such that the captured sound signal is attenuated to the noise suppression level outside said speech peaks; and

[0042] - a consonant extension module configured to execute the consonant extension algorithm described above, i.e. to maintain or gradually reduce the signal level of the processed sound signal to the noise suppression level, at the end of each speech peak within a predefined extension time.

[0043] For each embodiment or variant of the method according to the first aspect of the invention, there is a corresponding embodiment or variant of the hearing system according to the second aspect of the invention. Thus, disclosure related to the method also applies, mutatis mutandis, to the hearing system, and vice-versa.

[0044] In particular, in preferred embodiments of the hearing system,

[0045] - the consonant extension module is configured to apply the consonant extension gain and / or add additional noise to the captured sound signal or a signal derived therefrom,to thereby increase the level of the processed sound signal as compared to the noise suppression level, for the duration of the extension time;

[0046] - the extension time is between 5 and 50 msec, in particular 30 msec;

[0047] - the noise suppression module is further configured to reduce the consonant extension gain and / or a level of the additional noise smoothly, during the extension time;

[0048] - the consonant extension module is further configured to choose the consonant extension gain and / or a level of said additional noise such that they overcompensate the attenuation effected by the noise suppression gain, at least in an initial phase of the extension time; in particular, the level of the consonant extension gain and / or the level of the additional noise are adjusted, at the start of the extension time, such that the level of the processed sound signal is kept constant, as compared to the level of the processed sound signal at the end of the preceding speech peak;

[0049] - the hearing instrument further comprises a filter bank configured to split said captured sound signal into a plurality of frequency band signals, wherein the consonant extension module exclusively acts on (i.e. influences) frequency band signals within a predefined frequency sub-range of the frequencies included in the captured sound signal, which frequency sub-range corresponds to the range of characteristic frequencies of the consonants to be extended.

[0050] Within the scope of the invention, the noise suppression module and the consonant extension module can be implemented as separate (software or hardware) modules or as a combined module.

[0051] Optionally, the hearing system includes a voice recognition module configured to analyze the captured sound signal to detect speech-intervals and non-speech intervals as described above. Preferably, the voice recognition module is configured to distinguish the own voice of the user of the hearing instrument from foreign voice. Herein, preferably, the hearing system is configured to activate the consonant extension module (and, optionally, the noise suppression module) in speech intervals only, preferably in foreign-speech intervals only.Preferably, the signal processor is designed as a digital electronic device. It may be a single module or consist of a plurality of sub-processors. The signal processor or at least one of said sub-processors may be a programmable device (e.g., a microcontroller). In this case, the functionality mentioned above, or part of said functionality is implemented as software (firmware). Also, the signal processor or at least one of said sub-processors may be a non-programmable device (e.g., an ASIC). In this case, the functionality mentioned above, or part of said functionality may be implemented as hardware circuitry.

[0052] In a preferred embodiment of the invention, the noise suppression module and / or the consonant extension module (and / or, if provided, voice recognition module) are arranged in the hearing instrument. In particular, each of these modules may be designed as a hardware or software component of the signal processor or as separate electronic component. However, in other embodiments of the invention, the noise suppression module and / or the consonant extension module (and / or, if provided, the voice recognition module) or at least a functional part thereof may be located on an external electronic device such as a second hearing instrument of the hearing system. One or more of said modules can also be implemented as a functional part of a hearing app to be installed on a mobile electronic device (such as a smartphone) of the user.

[0053] Embodiments of the present invention will be described with reference to the accompanying drawings in which

[0054] Fig. 1 shows a schematic representation of a hearing system comprising a hearing instrument to be worn at the ear of a user, wherein the hearing instrument comprises two input transducers arranged to capture a sound signal from an environment of the hearing instrument, a signal processor arranged to process the captured sound signal, and an output transducer arranged to emit the processed sound signal to the user, wherein a noise suppression module and a consonant extension module are implemented in the signal processor;

[0055] Fig. 2 shows, in a schematic block diagram, a functional structure of a first embodiment of the hearing system of Fig. 1;

[0056] Fig. 3 shows, in a diagram over time, a schematic representation of a frequency band signal derived from the captured sound signal, said frequency bandsignal comprising speech peaks in each of which the captured sound signal contains sound of a spoken consonant, and intermediate sections in which the captured sound signal does not contain sound of a spoken consonant;

[0057] Fig. 4 shows, in three stacked synchronous diagrams over time a temporal section of a partially processed frequency band signal, and time averaged signals derived therefrom (diagram (a), top); a noise suppression gain provided by the noise suppression module (diagram (b), middle); and a consonant extension gain provided by the consonant extension module (diagram (c), bottom); and

[0058] Fig. 5 shows, in a schematic block diagram, a functional structure of a second embodiment of the hearing system of Fig. 1.

[0059] Like reference numerals indicate like parts, structures and elements unless otherwise indicated.

[0060] Fig. 1 shows a hearing system 2 comprising a hearing instrument 4 which, by way of example, may be designed as a Behind-The-Ear (BTE) device. In the preferred implementation, the hearing instrument 4 is a hearing aid that is configured to support the hearing of a hearing-impaired user. As shown in Fig. 1, the hearing system 2 may solely consist of the hearing instrument 4 and, thus, may not comprise any further device or structure. However, optionally, the hearing system 2 may comprise at least one further device such as a second hearing aid instrument (not shown) to be worn in or at the other ear of the user to provide binaural support to the user. Alternatively, or additionally, the hearing system 2 may comprise a structure such as an app for remote control and / or configuration of the hearing instrument 4 being installed, e.g., on a smartphone of the user. In this case, the smartphone is not a part of the hearing system 2. Instead, the hearing system 2 only uses the smartphone as an external resource for storage, computing power and communication functionalities.

[0061] The hearing instrument 4 comprises a housing 6 that is configured to be worn behind one of the ears of the user. Inside that housing 6, the hearing instrument 4 comprises two microphones 8 as input transducers and a receiver 10 as output transducer. The hearing instrument 4 further comprises a battery 12 and a signal processor 14. Preferably, thesignal processor 14 comprises both a programmable sub-unit (such as a microprocessor) and a non-programmable sub-unit (such as an ASIC). In the signal processor 14, i.a., a noise suppression modulel6 and a consonant extension module 18 are implemented. By preference, both modules 16 and 18 are designed as software components being installed in the signal processor 14. However, any one of these modules 16 and 18 can also be implemented as a hardware circuitry such an ASIC or a functional part thereof. The signal processor 14 is powered by battery 12, i.e., the battery 12 provides an electrical supply voltage U to the signal processor 14.

[0062] During normal operation of the hearing instrument 4, the microphones 8 capture a sound signal from an environment of the hearing instrument 4. The microphones 8 convert the sound into an input audio signal I, containing information on the captured sound. The input audio signal I is fed to the signal processor 14. The signal processor 14 processes the input audio signal I, i.a., to provide directed sound information (beamforming), to perform noise suppression and / or dynamic compression. Moreover, if the hearing instrument 4 is designed as a hearing aid, the signal processor 14 is configured to individually amplify different spectral portions of the input audio signal I based on audiogram data of the user to compensate the user-specific hearing loss.

[0063] The signal processor 14 emits an output audio signal O containing information on the processed sound to the receiver 10. The receiver 10 converts the output audio signal O into processed air-borne sound that is emitted into the ear canal of the user, via a sound channel 20 connecting the receiver 10 to a tip 22 of the housing 6 and a flexible sound tube (not shown) connecting the tip 22 to an ear piece inserted in the ear canal of the user.

[0064] Fig. 2 shows a functional structure of the signal processor 14 in more detail. As can be seen herein, in the signal processor 14, the input audio signal I of the two microphones 8 is fed to a beam former 24 that outputs a beamformed input audio signal I’ to an analysis filter bank 26. The analysis filter bank 26 converts the beamformed input audio signal I’ into a frequency domain representation of this signal in that it spectrally splits the beam-formed input audio signal I’ into a plurality of (e.g. 128) frequency band signals if each of which contains a spectral portion of the beamformed input audio signal I’corresponding to a small frequency range with a center frequency (which is denoted by the index letter “ ’).

[0065] At least one signal processing algorithm 28 (as previously described) is applied to the frequency band signals if to execute a signal processing step and, thus, modify the input audio signal I, to support the hearing of the user. E.g., the at least one signal processing algorithm 28 applies a gain gPto the frequency band signals if and outputs partially processed (or pre-processed) frequency band signals o’f (o’f = gP• if). In general, said gain gPis a frequency-dependent parameter and may, thus, have a different value for each center frequency f. In simple implementations, the gain gPmay be constant in time t. However, in other implementations, the gain gPmay also have a dependency of time t (gP= gp ).

[0066] In a multiplier 30, a further gain ge(also referred to as the “speech enhancement gain” ge) is applied to the frequency band signals o’f, thus producing a frequency representation of the output audio signal O comprising processed frequency band signals or The frequency band signals Of are fed to a (synthesis) filter bank 32 that combines the frequency band signals Of, thereby creating the output audio signal O.

[0067] Parallel hereto, the frequency band signals if are fed to an envelope detector 34 that extracts the time-dependent signal level (envelope) of each of the frequency band signals if. The envelope detector 34 outputs frequency band signals i’f containing said signal levels to the noise suppression module 16.

[0068] The captured sound signal I may contain speech sound, e.g. of a person speaking with the user of the hearing system 2. As is schematically illustrated in Fig. 3, showing a exemplary (simplified) variation of a frequency band signal i’f over time t, such speech sound typically results in speech peaks SP in some of the frequency band signals i’f. During said speech peaks SP, the respective frequency band signal i’f has a comparably high signal level, and intermediate sections IS, in which the respective frequency band signal i’f has a lower signal level, at least in a time average [i’f] of the frequency band signal i’f. Each of said speech peaks SP corresponds to a time section of the respective frequency band signal i’f, in which the latter contains speech sound of a spoken consonant. In contrast, the intermediate sections IS contain speech sound of spoken vowels orno speech sound at all. The speech peaks SP only occur in frequency band signals i’f having center frequencies within the characteristic frequency range of the respective consonant.

[0069] As is evident from Fig. 2, the time average [i’f] is a short-time average calculated for an averaging time that is short as compared to the typical temporal length of a speech peak SP (i.e. short as compared to the typical temporal length of a spoken consonant). For example, the time average [i’f] may be calculated for an averaging time of 10 msec.

[0070] In order to improve speech intelligibility in the output audio signal O, the noise suppression module 16 is configured to selectively attenuate the output audio signal O in the intermediate sections IS, while leaving speech peaks SP (and, thus, spoken consonants) unattenuated. To this end, the noise suppression module 16 comprises a consonant detector 36 and a gain setting unit 38.

[0071] The consonant detector 36 is configured to detect onsets (i.e. temporal start points) and offsets (i.e. temporal end points) of speech peaks SP in the frequency band signals i’f.

[0072] In order to detect onsets of speech peaks SP, in a simple implementation, the consonant detector 36 compares one or more frequency band signals i’f with center frequencies within the characteristic frequency range of a consonant type to be detected (e.g. between 4kH and 8 kHz for fricatives) with one or more frequency band signals i’f with center frequencies outside said characteristic frequency range. For each of said frequency band signals i’f, the consonant detector 36 determines the difference of a short-time average (such as the time average [i’f] of Fig. 2) and a long-time average of the respective frequency band signal i’f. Herein, the long-time average is calculated for an averaging time that corresponds to or is longer than a typical temporal length of a speech peak SP (e.g. for an averaging time of 125 msec). The consonant detector 36 compares said difference with a threshold and detects an onset of a speech peak SP if, for the frequency band signal(s) i’f within the characteristic frequency range of a consonant type to be detected, said difference exceeds the threshold, while said threshold is not exceeded outside said characteristic frequency range.In order to recognize offsets of speech peaks SP and, thus, sound of spoken consonants in the frequency band signals i’f, the consonant detector 36 uses a reverse function as compared to onset recognition. For example, the consonant detector 36 compares the short-time average and the long-time average (as defined for the onset recognition) for each of said one or more frequency band signals i’f with center frequencies within the characteristic frequency range of a consonant type in interest. It recognizes offsets whenever, for said frequency band signals i’f, the difference of said short-time and long-time averages (having a negative value in case of offsets) falls below a predefined threshold, i.e. if the short-time average undershoots the long-time average by more than the predefined threshold. As in the case of onset detection, the consonant detector 36 can also check one of more frequency band signals i’f with center frequencies outside the characteristic frequency range of a consonant type in interest, for not showing the offset. Alternatively or additionally, the consonant detector 36 uses at least one of the consonant detection methods known from Pl to P3 to detect spoken consonants and, thus, speech peaks SP.

[0073] Whenever, and as long as the consonant detector 36 detects a speech peak SP (i.e. the presence of the sound of a spoken consonant in the captured sound signal I) it outputs a consonant detection signal d to the gain setting unit 38. Whenever, and as long as the gain setting unit 38 receives said signal d, i.e. outside detected speech peaks SP, it outputs a noise suppression gain gs.

[0074] Moreover, the signal d is fed to the consonant extension module 18. Whenever said module 18 registers an end of a speech peak SP (by detecting that the consonant detection signal d is switched off) it outputs a consonant extension gain gcfor a predefined extension time Ts. In a preferred implementation, the extension time Ts is set to 30 msec. In the exemplary embodiment of Fig. 2, the noise suppression gain gsand the consonant extension gain gcare combined in adder 40 to provide the speech enhancement gain ge(ge=gs + gc).

[0075] The combined effect the noise suppression gain gsand the consonant extension gain g(to the processed sound signal O is illustrated in Fig. 4.The top diagram (diagram (a)) of this figure shows a temporal section of one of the frequency band signals o’f that includes an offset 42 of a speech peek SP (i.e. a transition between a speech peak SP and a subsequent intermediate section IS). For example, the temporal section of the frequency band signal o’f shown in diagram (a) of Fig. 4 may correspond, with respect to time t and center frequency f, to a temporal section At of the frequency band signal i’f shown in Fig. 3. The diagram (a) of Fig. 4 also shows, with a dashed line, a short-time average [o’f] of the frequency band signal o’f that may be calculated analogously to the short-time average [i’f] of Fig. 3.

[0076] As shown in diagram (b) of Fig. 4, the noise suppression module 16 provides the noise suppression gain gs, during intermediate sections IS, with a low value that attenuates the frequency band signal or If the noise suppression gain gsis multiplied to the frequency band signal o’f (as shown in Fig. 2), the noise suppression gain gsis set to a value between 0 and 1, during intermediate sections IS. This value may be maintained constant, as shown in Fig. 4, by way of example, or may be provided with a predefined time-dependency. Moreover, that value to which the noise suppression gain gsis set during intermediate sections IS, may vary with frequency. During speech peaks SP, the noise suppression gain gs is set to 1, such that it does not influence the frequency band signal or For sake of illustration, diagram (a) of Fig. 4 shows, with a dash-dotted line, a short-time average [of]* of a corresponding frequency band signal Of as it would be derived from the frequency band signal o’f, by the multiplier 30, in absence of the consonant extension gain gc(i.e. if the frequency band signal Of was solely influenced by the noise suppression gain gs: Of = gs • o’f). It can be seen from this illustration that the noise suppression gain gs enhances the offset 42 in that it reduces the level of the frequency band signal Of and its short-time average [of]* to a (low) noise suppression level.

[0077] The consonant extension module 18 provides the consonant extension gain gcwith a value that overcompensates the effect of the noise suppression gain gs, for almost the entire duration of the extension time Ts. In fact, in the example shown in Fig. 4, the consonant extension gain gcis chosen such that it compensates both the effect of the noise suppression gain gsand the natural decrease of the signal level of the frequency band signal o’f at the offset 42. Thus, the consonant extension module 18 artificially prolongs the signature of the spoken consonant in the captured sound by more stronglyamplifying noise for the duration of the extension time Ts. At the end of the extension time Ts, the consonant extension module 18 reduces the value of the consonant extension gain gc(and, thus, its influence to the frequency band signal o’f) to zero.

[0078] A short-time average [or] of the frequency band signal Of provided by the multiplier 30, as modified by the noise suppression gain gsand the consonant extension gain gc, is shown in diagram (a) of Fig. 4 with a solid line. It can be seen from this figure that the level of the frequency band signal Of and its short-time average [or] is maintained, at least approximately, at the offset 42 of the speech peak SP and reach the noise suppression level only at the end of the extension time Ts.

[0079] The noise suppression gain gsand the consonant extension gain gcare applied to frequency band signals o’f with center frequencies within the characteristic frequency range of a consonant type in interest only. In other words, frequency band signals o’f with center frequencies outside said the characteristic frequency range are not influenced by noise suppression gain gsand the consonant extension gain gc.

[0080] Fig. 5 shows a further embodiment of the hearing system 2. In this embodiment, different from the embodiment according to Fig. 2, the consonant extension module 18 is configured to output an (artificial) additional noise N, for the extension time Ts, whenever it registers an end of a speech peak SP (by detecting that the signal d is switched off). Said additional noise N is combined by an adder 44 with the frequency band signal o’f (after or before the latter is multiplied with the noise suppression gain gs). The additional noise N is chosen such that it artificially extends (prolongs) the sound of the spoken consonant for the duration of the extension time Ts. As a result, the short-time average [of] of the frequency band signal Of which is provided by the adder 44 and, thus, modified by the noise suppression gain gsand the additional noise N, resembles the short-time average [of] shown in diagram (a) of Fig. 4. At the end of the extension time Ts, the consonant extension module 18 reduces the value of the additional noise N (and, thus, its influence to the frequency band signal o’f) to zero.

[0081] Unless described otherwise, the hearing instrument 4 shown in Fig. 5 corresponds to the embodiment shown in Fig. 2.

[0082] In further variations of the embodiments of Fig. 2 or Fig. 5, the hearing instrument 4 also includes a voice recognition module (not shown) that is configured to detect speech-intervals and non-speech intervals in the input audio signal I and to distinguish the own voice of the user of the hearing instrument 4 from foreign voice. Herein, the voice recognition module activates the consonant extension module 18 and the noise suppression module 18 during foreign-speech intervals only, i.e. in time intervals in which the input audio signal contains speech of speaker different from the user.

[0083] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the invention as shown in the specific examples without departing from the spirit and scope of the invention as broadly described in the claims. The present examples are, therefore, to be considered in all aspects as illustrative and not restrictive. Unless specified otherwise or excluded for causal reasons, all elements of the embodiments described above can be used in different combinations, in accordance with the subject matter of the claims, without departing from the scope of the invention.LIST OF REFERENCE NUMERALS

[0084] 2 hearing system

[0085] 4 hearing instrument

[0086] 6 housing

[0087] 8 microphone

[0088] 10 receiver

[0089] 12 battery

[0090] 14 signal processor

[0091] 16 noise suppression module

[0092] 18 consonant extension module

[0093] 20 sound channel

[0094] 22 tip

[0095] 24 beam former

[0096] 26 (analysis) filter bank

[0097] 28 signal processing algorithm

[0098] 30 multiplier

[0099] 32 (synthesis) filter bank

[0100] 34 envelope detector

[0101] 36 consonant detector

[0102] 38 gain setting unit

[0103] 40 adder

[0104] 42 offset

[0105] 44 adder

[0106] [i’f] short-time average (of frequency band signal i’f) [Of] short-time average (of frequency band signal Of) [Of]* short-time average (of frequency band signal Of) [o’f] short-time average (of frequency band signal o’f) At (temporal) section

[0107] d (consonant detection) signal

[0108] if frequency band signal

[0109] i’f frequency band signalgc (consonant extension) gain

[0110] ge(speech enhancement) gain

[0111] gp gain

[0112] gs (noise suppression) gain

[0113] of (processed) frequency band signals

[0114] o’f (partially processed) frequency band signals t time

[0115] I input audio signal

[0116] I’ (beamformed) input audio signal

[0117] IS intermediate section

[0118] N noise

[0119] O output audio signal

[0120] SP speech peak

[0121] Ts extension time

[0122] U supply voltage

Claims

CLAIMS1. A method for operating a hearing instrument (4), the method comprising:capturing a sound signal (I) from an environment of the hearing instrument (4);processing the captured sound signal (I) to derive a processed sound signal (O); andoutputting the processed sound signal (O) to a user of the hearing instrument (4);wherein said processing includes applying a noise suppression algorithm by which speech peaks (SP) in each of which the captured sound signal (I) contains sound of a spoken consonant or a spoken consonant of a predefined consonant type are detected, and by which a noise suppression gain (gs) is applied to the captured sound signal (I) or a signal (o’f) derived therefrom, such that the processed sound signal (O) is attenuated to a noise suppression level outside said speech peaks (SP); and,wherein said processing includes applying a consonant extension algorithm by which, at the end of each speech peak (SP), the signal level of the processed sound signal (O) is maintained or gradually reduced to the noise suppression level within a predefined extension time (Ts).

2. The method according to claim 1,wherein, by the consonant extension algorithm, a consonant extension gain (gc) and / or an additional noise (N) is applied to the captured sound signal (I) or a signal derived therefrom (o’f), for the predefined extension time (Ts), to thereby increase the level of the processed sound signal (O) as compared to the noise suppression level.

3. The method according to claim 1,wherein the extension time (Ts) is between 5 and 50 msec, in particular 30 msec.

4. The method according to claim 2 or 3,wherein, during the extension time (Ts), the consonant extension gain (gc) and / or a level of said additional noise (N) are smoothly reduced.

5. The method according to one of claims 2 to 4,wherein the consonant extension gain (gc) and / or the additional noise (N) are adjusted such that they overcompensate the attenuation effected by the noise suppression gain (gs), at least in an initial phase of the extension time (Ts).

6. The method according to one of claims 1 to 5,wherein the processing of the captured sound signal (I) includes splitting the captured sound signal (I) into a plurality of frequency band signals (if), and wherein the consonant extension algorithm is applied exclusively to a sub-group of said frequency band signals (if) or signals (o’f) derived therefrom having frequencies within a predefined frequency sub-range of the frequencies included in the captured sound signal (I) which frequency sub-range corresponds to the characteristic frequencies of consonants or to the characteristic frequencies of consonants of said predefined consonant type.

7. A hearing system (2) with a hearing instrument (4) to be worn in or at an ear of a user, the hearing instrument (4) comprising:at least one input transducer (8) configured to capture a sound signal (I) from an environment of the hearing instrument (4);a signal processor (14) configured to process the captured sound signal (I) to thereby derive a processed sound signal (O),an output transducer (10) configured to output the processed sound signal (O) to a user of the hearing instrument (4);wherein the hearing system (2) includes a noise suppression module (16) configured to detect speech peaks (SP) in each of which the captured sound signal (I) contains sound of a spoken consonant or a spoken consonant of a predefined consonant type, and to apply a noise suppression gain (gs) to the captured sound signal(I) or a signal (o’f) derived therefrom, such that the processed sound signal (O) is attenuated to a noise suppression level outside said peaks (SP); andwherein the hearing system (2) includes a consonant extension module (18) configured to maintain or gradually reduce the signal level of the processed sound signal (O) to the noise suppression level, at the end of each speech peak (SP) within a predefined extension time (Ts).

8. The hearing system (2) according to claim 7,wherein the consonant extension module (18) is configured to apply a consonant extension gain (gc) and / or an additional noise (N) to the captured sound signal (I) or a signal (o’f) derived therefrom, to thereby increase the level of the processed sound signal (O) as compared to the noise suppression level.

9. The hearing system (2) according to claim 7 or 8,wherein the extension time (Ts) is between 5 and 50 msec, in particular 30 msec.

10. The hearing system (2) according to claim 8 or 9,wherein the consonant extension module (18) is configured to reduce the consonant extension gain (gc) and / or a level of said additional noise (N) smoothly.

11. The hearing system (2) according to one of claims 8 to 10,wherein the consonant extension module (18) is configured to choose the consonant extension gain (gc) and / or a level of said additional noise (N) such that they overcompensate the attenuation effected by the noise suppression gain (gs), at least in an initial phase of the extension time (Ts).

12. The hearing system (2) according to one of claims 7 to 11,wherein the hearing instrument (4) further comprises a filter bank (26) configured to split said captured sound signal (I) into a plurality of frequency band signals (if), wherein the consonant extension module (18) exclusively acts on a sub-group of said frequency band signals (if) or signals (o’f) derived therefrom havingfrequencies within a predefined frequency sub-range of the frequencies included in the captured sound signal (I) which frequency sub-range corresponds to the characteristic frequencies of consonants or to the characteristic frequencies of consonants of said predefined consonant type.