Hearing system and method of operating the same

The hearing system addresses dynamic adaptation challenges by employing a noise embedding mechanism in the SEM, ensuring effective speech enhancement with reduced computational demands.

WO2026068002A1PCT designated stage Publication Date: 2026-04-02SIVANTOS PTE LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing hearing instruments face challenges in adapting dynamically to diverse acoustic environments, leading to suboptimal speech enhancement performance and user satisfaction due to high computational complexity and limited processing capacity.

Method used

A hearing system that employs a speech enhancement model (SEM) with a noise embedding mechanism, utilizing a convolutional neural network (CNN) and an embedding generation model (EGM) to suppress noise components, where the embedding signal is generated externally or selected from pre-generated signals, reducing computational complexity without compromising performance.

Benefits of technology

The system achieves improved adaptability and reduced numerical power and energy consumption by using noise embeddings, enhancing speech clarity in diverse acoustic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024078986_02042026_PF_FP_ABST
    Figure EP2024078986_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A hearing system (2) including at least one hearing instrument (4) and a method for operating the hearing system (2) are provided. According to the invention, an ambient sound from an environment of the hearing instrument (4) is captured by at least one first input transducer (8) of the hearing instrument (4) to produce a first input audio signal (I), wherein the ambient sound includes a speech component (S) and a noise component (N). The first input audio signal (I) is modified, at least in part by a signal processor (14) of the hearing instrument (4), to produce a processed audio signal (O) which is output as perceivable sound to a user of the hearing instrument (4), by an output transducer (10) of the hearing instrument (4). A speech enhancement model (26) is applied to the first input audio signal (I), or an audio signal (X) derived therefrom to produce an enhanced audio signal (Y) in which the noise component (N) of the ambient sound is suppressed while the speech component is (S) retained. An embedding signal (E) that include infomation characterizing the noise component is supplied to the speech enhancement model (26) to configure the speech enhancement model (26) for suppressing the noise component (N).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] FDST Patentanwalte, Nurnberg Seite 1

[0002] P240455P-CT / JB

[0003] DESCRIPTION

[0004] Hearing system and method of operating the same

[0005] The present invention relates to a hearing system including at least one hearing instrument. The invention also relates to a method of operating a hearing system.

[0006] In general, a hearing instrument is an electronic device being designed to support the hearing of a person wearing it (which person is called the “user” or “wearer” of the hearing instrument). In particular, the present invention relates to hearing instruments that are specifically configured to at least partially compensate a hearing impairment of a hearing-impaired user. Such hearing instruments are also called “hearing aids”. In addition to said hearing aids, there are hearing instruments that are designed to support the hearing of normal-hearing users (i.e. persons without a hearing impairment). Such hearing instruments, being sometimes referred to as “Personal Sound Amplification Products” (PSAP), may be provided, e.g., to enhance the hearing of the wearer in complex acoustic environments or to protect the hearing of the wearer from damage or overstress.

[0007] Hearing instruments, in particular hearing aids, are typically designed to be worn in or at an ear of the user, e.g. as a Behind-The-Ear (BTE) or In-The-Ear (ITE) device. With respect to its internal structure, a hearing instrument normally comprises at least one (acousto-electrical) input transducer, a signal processor, and an output transducer. During operation of the hearing instrument, the at least one input transducer captures a sound from an environment of the hearing instrument and converts it into an input audio signal (i.e. an electrical signal transporting sound information). In the signal processor, the input audio signal is processed, in particular amplified dependent on frequency, e.g., to at least partially compensate the hearing-impairment of the user. The signal processor outputs the processed signal

[0008] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 2

[0009] (also called output audio signal) to the output transducer. Most often, the output transducer is an electro-acoustic transducer (also called “receiver”) that converts the output audio signal into a processed airborne sound which is emitted into the ear canal of the user. In BTE devices, said receiver may be located in a housing to be worn behind the ear of the user. In this case, the sound output by the receiver is guided to the ear canal by a hollow sound tube connecting a tip of the housing with an earpiece to be introduced in the ear canal. In a further kind of BTE devices, referred to as Receiver-In-Canal (RIC) devices, the receiver is located in the earpiece, thus, outside the housing. Alternatively, the output transducer may be an electro-mechanical transducer that converts the output audio signal into a structure-borne sound (vibrations) that is transmitted, e.g., to the cranial bone of the user.

[0010] Furthermore, besides classical hearing aids, there are implanted hearing instruments such as cochlear implants, and hearing instruments the output transducers of which directly stimulate the auditory nerve of the user.

[0011] The term “hearing system” denotes one device or an assembly of devices and / or other structures providing functions required for the operation of a hearing instrument. A hearing system may consist of a single stand-alone hearing instrument. As an alternative, a hearing system may comprise a hearing instrument and at least one further electronic device which may, e.g., be one of another hearing instruments for the other ear of the user, a remote control, and a programming tool for the hearing instrument. Moreover, modem hearing systems often comprise at least one hearing instrument and a software application for controlling and / or programming the at least one hearing instrument, which software application (hereinafter referred to as the “hearing app”) is or can be installed on a computer, in particular a mobile communication device such as a mobile phone (smartphone). In the latter case, typically, the computer is not a part of the hearing system but is only used by the hearing system as a resource of data storage, numerical power, and communication services. Most often, the computer (in particular, the mobile communication device) on which the hearing app is or may be installed will be manufactured and sold independently of the hearing system.

[0012] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 3

[0013] In general, hearing instruments enhance auditory perception and quality of life of the users, particularly in challenging, noisy environments. This is of particular importance for hearing-impaired users since the impairment not only impacts their ability to communicate effectively, but also isolates individuals from their social environments. I. a., modern hearing instruments can distinguish between speech and noise, enhancing the former while suppressing the latter, thereby allowing the users to participate in conversations even in busy settings like restaurants or marketplaces.

[0014] Speech enhancement (SE) is one of the core functionalities of modem hearing instruments. Real-time approaches for providing SE, as disclosed in references

[0001] and [2], traditionally depend on statistical models. However, the contemporary landscape has changed significantly, with data-driven, deep learning-based models trained on large available datasets showing great promise in SE tasks. In particular, convolutional neural networks (CNNs) use masking and direct feature mapping techniques to enhance speech. An advanced solution (referred to as “Deepfil- ternet”) known for its superior noise suppression and speech clarity was proposed in reference [3],

[0015] It was recognized that deeper CNN architectures yield improved results. However, their computational complexity can challenge the real-time processing capacities, in particular in small devices such as hearing instruments. Additionally, they often struggle to adapt dynamically to varied real-world acoustic environments. This leads to a significant gap in user satisfaction and usability. For example, users report a lack of clarity in noisy settings and delayed responses to sudden noise changes.

[0016] It is an object of the present invention to provide a solution for SE in a hearing instrument with improved adaptability and performance in diverse acoustic environments.

[0017] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 4

[0018] According to a first aspect of the invention, the above object is met by a method as defined by claim 1 for operating a hearing system. The above-mentioned object is also met, according to a second aspect of invention, by a hearing system as defined by claim 13. Preferred embodiments of the invention are described in the dependent claims and the subsequent description.

[0019] In general, the invention relates to a hearing system and a method for operating the latter, which hearing system comprises at least one hearing instrument. As mentioned above, the hearing system may consist of a single stand-alone hearing instrument or may include at least one further device or app in addition to said at least one hearing instrument. In particular, the hearing system may be a binaural hearing system comprising two hearing instruments for the two ears of a user.

[0020] The at least one hearing instrument of the system comprises at least one (first) input transducer configured to capture an ambient sound from an environment of the hearing instrument. Optionally, in addition to said at least one (first) input transducer located in the at least one hearing instrument, the hearing system may comprise or access at least one (second) input transducer of an external device, i.e. , located outside the at least one hearing instrument. The external device may be a part of the hearing system (e.g. a dedicated remote control for the hearing instrument). As an alternative, the second input transducer and the external device may not belong to the hearing system. For instance, the at least one (second) input transducer may be, e.g., a microphone of a smartphone of which the hearing app is installed.

[0021] The at least one hearing instrument of the system comprises a signal processor configured to process the captured sound signal to produce a processed sound signal; and an output transducer configured to output the processed sound signal to a user of the hearing system. By preference, the at least one hearing instrument is designed to be worn at or in the ear of the user, e.g. as a BTE (in particular RIC) device or as an ITE device or any other design mentioned above. By preference, the output transducer of the at least one hearing instrument is a receiver (i.e. electro-acoustic transducer). Alternatively, within the scope of the invention, the output

[0022] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 5 transducer may be an electro-mechanical transducer or an output transducer directly stimulating the auditory nerve of the user.

[0023] In accordance with the method for operating the hearing system, in a first sound capturing step, an ambient sound is captured from an environment of the at least one hearing instrument by the at least one (first) input transducer of the hearing instrument to produce a first input audio signal, wherein the ambient sound includes a speech component and a noise component. Herein, the speech component is not necessarily contained in the ambient sound at all times. Instead, the speech component may be present in certain speech intervals only, which speech intervals are interrupted by non-speech intervals (in which the ambient sound does not include the speech component). Moreover, the speech component may be defined to be attributed to a dominant voice (i.e. , a voice having a certain loudness and / or originating from a close vicinity to the hearing instrument) or to a particular voice (i.e. the voice of one particular individual speaker). In these cases, the noise component may contain speech of one or more other speakers, e.g. babble of many distant voices in a party environment. Finally, the ambient sound may consist of the speech component and the noise component (i.e. only include one speech component and one noise component), or include at least one further component, a music component and / or another speech component attributed to another voice.

[0024] The method also comprises a signal processing step in which the first input audio signal is modified to produce a processed audio signal. The entire signal processing step may be performed by the signal processor of the at least one hearing instrument. Alternatively, only part of the signal processing step is performed by the signal processor whereas another part of the signal processing step is performed externally, e.g., by a functional part of the hearing app. In binaural versions of the hearing system, the signals processors of said two hearing instruments may perform the signal processing step together in a work-sharing manner.

[0025] Finally, in an outputting step, the processed signal is output as perceivable sound to a user of the at least one hearing instrument, by an output transducer of the at least one hearing instrument.

[0026] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 6

[0027] The signal processing step further includes a speech enhancing step (SE step). In said SE step, a speech enhancement model (SEM) is applied to the first input audio signal to produce an enhanced audio signal in which the noise component of the ambient sound is suppressed while the speech component is retained. Preferably, the SEM includes a (deep) neural network, in particular a convolutional neural network. Also preferably, the SEM is implemented in the at least one hearing instrument, in particular as a (hardware and / or software) part of the signal processor. Optionally, in case of binaural versions of the hearing system, SEMs may be implemented in both hearing instruments of the system. Alternatively, within the scope of the invention, only one of the two hearing instruments of the binaural hearing system may be equipped with the SEM. In a less preferred alternative, however still within the scope of the invention, the SEM may be implemented externally to the at least on hearing instrument, e.g. as a part of a hearing app of the system.

[0028] The SEM may be applied directly to the first input audio signal produced by the at least one first input audio transducer. Alternatively, the SEM may be applied indirectly to the first input audio signal; i.e. it may be applied to an audio signal derived from the first input audio signal by pre-processing.

[0029] In embodiments of the invention, the captured sound signal may be fed to the SEM as a (raw or pre-processed) time domain signal. In an alternative embodiment, features such as a time-frequency representation may be extracted from the captured sound signal which features are then fed as input signals to the SEM.

[0030] Prior to or during normal operation of the hearing system, the SEM is trained for suppressing the noise component using a machine learning algorithm. Preferably, the training is performed offline, i.e. not on the hearing system itself. In other words, preferably, the SEM is provided to the hearing system in a pre-trained state.

[0031] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 7

[0032] According to the invention, an embedding signal that includes information characterizing the noise component is supplied to the SEM to configure the latter for suppressing the noise component.

[0033] Using such embedding signals (also referred to as “embeddings” or “fingerprints”) is known per se, e.g., from reference [4],

[0034] In general, “embeddings” are auxiliary inputs to a model which provide additional information supporting the model in performing its task. In the present case, the embedding is a “noise embedding” or “noise fingerprint” defining the noise the SEM is supposed to remove from its input signal. So far, noise embeddings have been developed for applications such as remote voice transmission (e.g. telephone, video conference or radio applications) or audio recording in which the embeddings can benefit of the high numerical power of stationary computer systems. Embeddings have not yet been used for SE in hearings systems. As they normally add further complexity to already complex SE solutions, known SEMs using embeddings are not suited for implementation in hearing systems for their large numerical requirements. The present inventors nevertheless suggest implementation of a SEM using a noise embedding in a hearing system as they found out that using the noise embedding allows for reducing the complexity of the SEM, as compared to a SEM not using embeddings, without having to accept a poorer SE performance. In other words, it has been shown by the present inventors that the added complexity induced by providing embeddings can be compensated by reducing complexity of the SEM, without losing SE performance, thus allowing use of noise embeddings in a hearing system. However, vis-a-vis a conventional SEM not using embeddings, the solution according to the invention has the advantage of having improved adaptability to diverse acoustic environments.

[0035] Within the scope of the invention, the embedding signal may be provided in different ways and by using different hardware platforms.

[0036] In a first preferred embodiment of the invention, in an embedding selecting step, the embedding signal to be provided to the SEM is selected from a plurality of pre-

[0037] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 8 generated embedding signals. The selection is made based on acoustic scene detection, i.e. a sound classifier is used to analyze the noise component in the ambient sound and to classify the noise component into one of several predetermined noise classes, based on this analysis. The selection of the embedding signal to be provided to the SEM is made based on the classification. In particular, one pregenerated embedding signal is stored for each predefined noise class, and the pre-generated embedding signal attributed to the sound class into which the noise component is classified, is selected and used for configuring the SEM. The sound classifier may be applied to the first input audio signal, i.e. the ambient sound captured by the at least one input transducer of the at least one hearing instrument. Alternatively, the sound classifier may be applied to a second input audio signal in which the ambient sound is captured by a second input transducer which is not part of the at least one hearing instrument. For instance, the hearing app may use a microphone of a smartphone on which it is installed, for capturing the ambient sound of the user’s environment and, thus, producing the second input audio signal. This embodiment has the advantage of being most easy to implement in the hearing system and requiring little numerical resources during operation of the hearing system. Within the scope of the invention, the pre-generated embeddings may be stored in the at least one hearing instrument or externally to the hearing instrument, e.g. in the system’s hearing app or in a data cloud. Preferably, the embedding selecting step is repeated whenever the classification of the acoustic scene changes, i.e. whenever the sound classifier changes the noise class into which the noise component of the ambient sound is classified.

[0038] In a second preferred embodiment of the invention, the embedding signal is generated “online” (i.e. during operation of the hearing instrument), in an embedding generating step. Herein, an embedding generation model (EGM) is applied to the first input audio signal or to a second input audio signal (as mentioned above). The EGM analyses the noise component of the ambient sound to generate the embedding signal. Preferably, similar to the SEM, the EGM includes a (deep) neural network, in particular a convolutional neural network. Within the scope of the invention, the EGM may be implemented in the at least one hearing instrument, in particular as a (hardware and / or software) part of the signal processor. However, in a

[0039] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 9 preferred implementation of the method, the EGM is implemented as part of the hearing app and is, thus, configured to run on an external computer, in particular the user’s smartphone.

[0040] Hence, in a variant of said second preferred embodiment of the invention, the embedding signal may be generated on the at least one hearing instrument itself. An advantage of this embodiment is that it is suited for stand-alone operation of a single hearing instrument. In other words, it does not require a connection of the hearing instrument to another device or structure.

[0041] On the other hand, implementations generating the embedding externally to the at least one hearing instrument have the advantage of not being bound to the limited numerical capacities and battery power of a hearing instrument. Instead, the larger numerical capacity and battery power of an external device, e.g. the user’s smartphone, can be used for generating the embedding signal.

[0042] Preferably, in further variants of said first and second preferred embodiment of the invention, the embedding signal is only selected or generated in non-speech intervals when the ambient sound does not contain the speech component. To this end, in a voice activity detection step, a voice activity detector (VAD) is applied to the first input audio signal or to a second input audio signal (as described above). The VAD analyzes the ambient sound to detect said non-speech intervals. Based on the output of the VAD, the embedding selecting step or the embedding generating step are only performed during detected non-speech intervals. In other words, the embedding selecting step or the embedding generating step is not performed during speech intervals (i.e. when and as long as the VAD detects the speech component in the ambient sound). To conclude, the EGM is run only when the background noise to be suppressed can be captured directly from the environment.

[0043] Since, normally, the noise background does not change rapidly, the EGM may be run intermittently with large interruptions (pauses) as compared to the sampling rates required for normal hearing aid signal processing. Thus, in a preferred

[0044] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 10 embodiment of the method, the SEM updates the enhanced audio signal with a first sampling rate (i.e. , more precisely, the SEM produces the enhanced audio signal as a temporal sequence of values wherein subsequent values the enhanced audio signal are produced with the first sampling rate), and EGM updates the embedding signal with a second sampling rate (i.e., more precisely, the EGM produces the embedding signal as a temporal sequence of values wherein subsequent values the embedding signal are produced with the second sampling rate), wherein the first sampling rate exceeds the second sampling rate (embedding generation rate) by a factor of at least 102, preferably 103, in particular 104. For instance, the EGM is run once every second or longer interval, whereas the SEM produces output values at the sampling rate of the hearing instrument, e.g., at 48 kHz.

[0045] Preferably, the SEM comprises an encoder part that reduces the first input audio signal or the audio signal derived therefrom to a first intermediate signal of lower temporal and / or spectral resolution (as compared to the input signal to the SEM) and at least one decoder part that expands the first intermediate signal to a second intermediate signal of higher temporal and / or spectral resolution (as compared to the first intermediate signal). In a preferred implementation of the method, the first input audio signal or audio signal derived therefrom is reduced to the first intermediate signal by the encoder part. The first intermediate signal is then expanded to the second intermediate signal by the at least one decoder part.

[0046] Thus, preferably, the SEM is formed with a so-called “bottleneck” represented by the low-dimensional first intermediate signal. The bottleneck structure, together with training of the encoder part and the at least one decoder part, effectively forces the SEM to discard unwanted noise components of the first input signal. Speech enhancement models of this kind are known per se, for instance from references [3] and [5], the content of which is incorporated herein by reference. Preferably, a SEM of this kind, only supplemented by the embedding, is used for the method and system according to the present invention.

[0047] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 11

[0048] It an advantageous embodiment of the method, in the speech enhancing step thereof, the embedding signal is applied in the “bottleneck” of the SEM, i.e. to the first intermediate signal, in particular by adding (in particular by subtracting, which is considered a special form of adding) the embedding signal to the first intermediate signal. Correspondingly, in this embodiment, the dimensionality of the embedding signal is adapted to the dimensionality of the first intermediate signal. The dimensionality of the embedding signal is, thus, lower than that of the input signal to the SEM. For instance, the input signal to the SEM could be a 48 kHz speech signal of several seconds length while the embedding is a vector with a fixed length of, e.g., 128, which is constant for the entire embedding signal (but may, of course, be different for different embedding signals). In an alternative implementation, the first intermediate signal and the embedding signal are combined with a multi-head attention block (e.g. as described in [6]). Preferably, the multi-head attention block has four attention heads, where the first intermediate signal is used as query and the embedding signal is used as both key and value for the attention mechanism. The output of the multi-head attention block is then added to the first intermediate signal, and the result is used as input to the at least one decoder.

[0049] It an advantageous embodiment of the method and system according to the invention, the EGM and the encoder part of the SEM are implemented as neural networks of a same type differing only in the number of neurons and / or layers.

[0050] Herein, different from conventional realizations SEMS using embeddings, the EGM is formed with a higher complexity as compared to the encoder part of the SEM in that the EGM has more neurons and / or more layers than the encoder part of the SEM.

[0051] Generally, the hearing system according to the second aspect of the information is configured to automatically perform the method according to the first aspect of the invention. To this end, the hearing system further includes a SEM (as described above) and means for supplying the embedding signal to the SEM configuring the latter for suppressing the noise component in the ambient sound.

[0052] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 12

[0053] As the hearing system according the second aspect of the invention is configured to automatically perform the method according to the first aspect of the invention, each embodiment or variation of said method corresponds to an embodiment or variation of the hearing system. Hence, any disclosure related to the method also applies, mutatis mutandis, to the hearing system, and vice-versa.

[0054] Preferably, the signal processor of the at least one hearing instrument is designed as a digital electronic device. It may be a single unit or consist of a plurality of subprocessors. In accordance with the invention, the signal processor or at least one of said sub-processors may be a programmable device (e.g., a micro-controller).

[0055] In this case, the functionality mentioned above or part of said functionality is implemented as software (in particular firmware) which is stored in executable form in a storage of the signal processor. Also, the signal processor or at least one of said sub-processors may be a non-programmable device (e.g., an ASIC). In this case, the functionality mentioned above or part of said functionality is implemented as hardware circuitry.

[0056] In particular, both the SEM and the means for supplying the embedding signal may be realized by software or hardware (i.e. , non-programmable electronic circuitry such as an ASIC) or a combination of both. Within the scope of the invention, the SEM and the means for supplying the embedding signal may be implemented in the at least one hearing instrument, externally to the at least one hearing instrument or distributed to the at least one hearing instrument and an external part of the hearing system (such as a hearing app).

[0057] As can be seen from the above, the main effect of the invention is saving numerical power and energy consumption for operating the SEM in a hearing system by supplying a noise embedding to the SEM. The noise embedding supports the operation of the SEM and, thus, allows to reduce the complexity of the SEM without losing SE performance which, in turn, allows for reduced numerical power and energy consumption for operating the SEM. In particular, in embodiments in which the embedding signal is generated during operation of the hearing system, numerical power and energy consumption for operating the SEM may be saved by

[0058] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 13 designing the EGM with a greater complexity as compared to the encoder part of the SEM.

[0059] Numerous features of the various embodiments of the invention described above further support this main effect by allowing to supply the embedding signal in a way that does not, or at least not heavily, utilize the numerical capacity and the energy supply of the at least hearing instrument. This effect is achieved,

[0060] • by selecting the embedding signal from pre-generated embedding signals (instead of generating the embedding signal in the hearing instrument, during operation of the latter),

[0061] • by generating the embedding signal externally to the hearing instrument, i.e. , in an external part of the hearing system such as by the hearing app (instead of generating the embedding signal in the hearing instrument, during operation of the latter),

[0062] • by generating the embedding signal with an embedding generation rate that is much lower than the sampling rate of the SEM, and / or

[0063] • by generating the embedding signal in detected speech intervals only.

[0064] Unless specified otherwise, the above teachings can be applied alone or in combination.

[0065] Subsequently, exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings in which

[0066] Fig. 1 shows a schematic representation of a first exemplary embodiment of a hearing system comprising a hearing instrument to be worn at the ear of a user and a hearing app which is installed on smartphone of the user, wherein a speech enhancement model (SEM) and a means for supplying an embedding signal to the SEM

[0067] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 14 are implemented in the hearing instrument, as a part of a signal processor thereof;

[0068] Fig. 2 shows a schematic representation (similar to that of Fig. 1 ) of a second exemplary embodiment of the hearing system wherein the SEM is implemented in the hearing instrument, wherein the means for supplying the embedding signal to the SEM is implemented as a functional part of the hearing app; and

[0069] Fig. 3 and 4 show block diagrams of different embodiments of the functional structure of the hearing system of Fig. 1 or 2.

[0070] In the figures, like reference numerals always indicate like parts, structures and elements unless otherwise indicated.

[0071] Fig. 1 shows a hearing system 2 comprising a hearing instrument 4 that is configured to be worn in or at one of the ears of a user 5 (Fig. 3,4). Preferably, the hearing instrument 4 is a hearing aid, i.e. a hearing instrument configured to support the hearing of a hearing-impaired user. As shown in Fig. 1 , by way of example, the hearing instrument 4 may be designed as a Behind-The-Ear (BTE) hearing instrument. However, in further embodiments of invention, the hearing instrument 4 may be designed in any of the further configurations described above or known in the art; e.g. as a RIC or ITE device. Optionally, the system 2 comprises a second hearing instrument (not shown) to be worn in or at the other ear of the user to provide binaural support to the user 5.

[0072] The hearing instrument 4 comprises, inside a housing 6, two microphones 8 as (first) input transducers and a receiver 10 as an output transducer. The hearing instrument 4 further comprises a battery 12 and a signal processor 14. Preferably, the signal processor 14 comprises both a programmable sub-unit (such as a microprocessor) and a non-programmable sub-unit (such as an ASIC).

[0073] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 15

[0074] The signal processor 14 is powered by the battery 12, i.e. , the battery 12 provides an electric supply voltage U to the signal processor 14.

[0075] During normal operation of the hearing instrument 4, the microphones 8 capture an airborne ambient sound from an environment of the hearing instrument 4. The microphones 8 convert the airborne sound into a (first) input audio signal I (also referred to as the “captured sound signal”), i.e., an electric signal containing information on the captured sound. The first input audio signal I is fed to the signal processor 14. The signal processor 14 processes the first input audio signal I, e.g., to provide a directed sound information (beamforming), to perform noise reduction and dynamic compression, and to individually amplify different spectral portions of the first input audio signal I based on audiogram data of the user 5 to compensate for the user-specific hearing loss. The signal processor 14 emits a processed audio signal 0, i.e., an electric signal containing information on the processed sound to the receiver 10. The receiver 10 converts the processed audio signal 0 into processed airborne sound that is emitted into the ear canal of the user 5, via a sound channel 16 connecting the receiver 10 to a tip 18 of the housing 6 and a flexible sound tube (not shown) connecting the tip 18 to an earpiece inserted in the ear canal of the user 5.

[0076] Further to the at least one hearing instrument 4, the hearing system 2 comprises a software application (subsequently denoted “hearing app” 20), that is installed on a mobile phone 22 of the user 5. Herein, the mobile phone 22 is not a part of the hearing system 2. Instead, it is only used by the hearing system 2 as an external resource providing computing power, data storage (memory) and communication services.

[0077] The hearing instrument 4 and the hearing app 20 exchange data via a wireless link 24, e.g., based on the Bluetooth standard. To this end, the hearing app 20 accesses a wireless transceiver (not shown) of the mobile phone 22, in particular a Bluetooth transceiver, to send data to the hearing instrument 4 and to receive data from the hearing instrument 4. The hearing app 20 also includes functions to remote control, configure and update the hearing instrument 4.

[0078] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 16

[0079] The hearing system 2 also includes am speech enhancement model (SEM 26) and a means 28 (subsequently described in further detail) for supplying an embedding signal E to the SEM 26. In the example of Fig. 1 , both the SEM 26 and the means 28 are realized as software components being installed in executable form in the signal processor 14 of the hearing instrument 4.

[0080] Fig. 2 shows an alternative arrangement of the hearing system 2 which, unless indicated otherwise, corresponds to the embodiment shown in Fig. 1. However, in the embodiment of Fig. 2, different from the embodiment of Fig. 1 , the means 28 are implemented as a part of the hearing app 20. In this case, the means 28 provide the embedding signal E to the SEM 26 via the wireless link 24.

[0081] Fig. 3 show a schematically simplified sketch of the functional structure of an embodiment of hearing system 2. As shown in the figure, the microphones 8 of the hearing instrument 4 capture an ambient sound from an environment of the user 5 which ambient sound, at least in certain temporal speech intervals, contains a speech component S and a noise component N. The noise component N may be traffic noise (as schematically shown in Fig. 3) or any other sound that is not interrelated with a dominant voice to which the user 5 is assumed to listen. In particular, the noise component N may contain babble of multiple distant voices.

[0082] In the course of the signal processing performed by the hearing instrument 4, an analysis filter bank 30 or a similar frequency analysis unit is applied to the first input audio signal I of the microphones 8 of the hearing instrument 4, either directly or via one or more pre-processing unit(s) 32 that, e.g., apply amplification and / or beamforming to the first input audio signal I. In its raw or pre-processed form, the first input audio signal I is a variable signal depending on time t: I = l(t).

[0083] The analysis filter bank 30 converts the (raw or pre-processed) first input audio signal I into a digital time-frequency representation X. The time-frequency representation X of the first input audio signal I can be regarded as a discrete mathematical function depending on a discrete time variable (referred to as time frames

[0084] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 17 k) and a discrete frequency variable (referred to a frequency bands f). In other words, any value of the time-frequency representation X is attributed to a specific time frame k and a specific frequency band f: X -> X(k,f).

[0085] The SEM 26 is implemented as a deep neural network. Unless indicated otherwise, in the preferred implementation, the structure of the SEM 26 corresponds to that of the speech enhancement algorithm DEEPFILTERNET 2 disclosed in [5], the content of which is incorporated herein by reference. In other words, DEEPFILTERNET 2, is used as the basic structure of the SEM 26. Thus, as suggested in [5], the time-frequency representation X is fed in parallel to analyzers 34 and 36 of the SEM 26. Herein, the analyzer 34 extracts a signal Xerb containing ERB (equivalent rectangular bandwidth) features from the time-frequency representation X. In other words, it maps the time-frequency representation X to the ERB frequency scale: Xerb = Xerb(k, ferb), wherein ferb denotes an ERB frequency band and, thus, recovers the speech envelope. The analyzer 36 extracts a signal Xdf containing complex features from the time-frequency representation X. The signal Xdf is a function of time frame k and a discrete frequency variable (subsequently referred to a “frequency band”) fdf: Xdf = Xdf(k, fdf). The signal Xdf represents the periodicity of the speech component S. In the preferred implementation, the ERB frequency bands ferb and the frequency bands fdf have different values as compared to the original frequency bands f of the filter bank 30. Preferably, the ERB frequency bands ferb and the frequency bands fdf arise from different combinations (i.e. sums or averages) or selections of the original frequency bands f. In particular, fdf only comprises a lower range of the frequency bands f (without summation over the latter), e.g. the frequency bands f up to a frequency of 5 kHz.

[0086] The two signals Xerb and Xdf are used an input signal to the “heart” of the SEM 26 which consists of three deep convolutional networks, i.e. an encoder part (encoder 38) and two decoder parts (decoders 40 and 42).

[0087] The encoder 38 reduces its input signal, i.e. the signals Xerb and Xdf, to a first intermediate signal Z of lower dimensionality as compared to the combination (i.e. concatenation) of the two signals Xerb and Xdf. More specifically, the first intermediate

[0088] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 18 signal Z has a lower frequency resolution as compared to the combination of the two signals Xerb and Xdf.

[0089] According to reference [5], the decoders 40 and 42 are provided to expand the first intermediate signal Z to a second intermediate signal of higher dimensionality as compared to the first intermediate signal Z. Herein, the second intermediate signal is structured into two parts, i.e. a frequency-dependent real-value gain Gerb output by the decoder 40 and having the dimensionality of the signal Xerb (Xerb = Xerb(k, ferb)) and deep filtering coefficients Cnof order n output by the decoder 42 (where, preferably, n is chosen between 1 and 7, e.g., set to 5). Thus, the first intermediate signal Z forms a bottleneck of low dimensionality between the encoder 38 and the decoders 40 and 42.

[0090] As described in reference [5], the real-value gain Gerb and the time-frequency representation X of the first input audio signal I are fed to a gain applicator 44 of the SEM 26, which applies the real-value gain Gerb (or, more precisely, its interpolation to the frequency scale of the original frequency bands f) to the time-frequency representation X to receive a short-time spectrum YG (YG = Yc(k,f)).

[0091] The short-time spectrum YG and the deep filtering coefficients Cnare fed to a deep filter 46 of the SEM 26. In the preferred implementation, the deep filter 46 corresponds to a filter disclosed in reference [7] which is incorporated herein by reference. It applies the deep filtering coefficients Cnto the short-time spectrum YG to produce an enhanced audio signal Y (Y = Y(k,f)) which is the output signal of the SEM 26 and in which the noise component N of the ambient sound is suppressed while the speech component S is retained.

[0092] The enhanced audio signal Y is fed to a synthesis filter bank 48, either directly or via one or more post-processing unit 50 that, e.g., apply frequency dependent amplification, dynamic compression, etc., to the enhanced audio signal Y. The synthesis filter bank 48 converts the (raw or post-processed) enhanced audio signal Y into the time dependent processed signal 0 (0 = O(t)) which is then output by the receiver 10.

[0093] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 19

[0094] The decoder 40 outputs a new real-value gain Gerb for every frame k. Similarly, the decoder 42 outputs new coefficients Cnfor every frame k. However, the output coefficients Cnare used for an n-tap filter. This means, n frames of the spectrum YG are multiplied with the filter coefficients Cn. These values are summed up (practically implemented using matrix multiplication). Thus, even though n frames of the spectrum YG are needed as an input of each time step, only a single frame of the enhanced audio signal Y is produced in every time step. So, the decoder 42 is designed to work on a multi-frame scale; i.e. it predicts coefficients Cnof an n-tap filter which is applied to the spectrum YG to produce a new frame of the enhanced audio signal Y.

[0095] Differing from the DEEPFILTERNET 2 disclosed in [5], the SEM 26 uses the embedding signal E as an additional input, wherein the embedding E configures the SEM 26 for suppressing the noise component N. In the preferred implementation, the embedding signal E has the same dimensionality as the first intermediate signal Z and is applied to the latter, by an adder 52. More precisely, the adder 52 subtracts the embedding signal E from the first intermediate signal Z. Thus, different from the DEEPFILTERNET 2, the first intermediate signal Z is fed to the decoders 40 and 42 indirectly via the adder 52 that modifies the first intermediate signal Z by measure of the embedding signal E.

[0096] In the embodiment of Fig. 3, the means 28 for supplying the embedding signal E is implemented as an embedding generation model (EGM 54). The EGM 54 has a same structure as a first part of the SEM 26 including the analyzers 34, 36 and the encoder 38; i.e. the EGM 54 includes a first analyzer 56 having a same structure and function as the analyzer 34, a second analyzer 58 having a same structure and function as the analyzer 36, and a convolutional neural network 60 that resembles the encoder 38. In fact, the neural network 60 differs from the encoder 38 only in its complexity. It has a higher complexity as compared to the encoder part 38, i.e. a larger number of neurons and / or a larger number of layers. In a preferred implementation, the neural network 60 of the EGM 54 consists of four convolutional layers with 128 channels followed by a GRU (Gated Recurrent Unit) layer

[0097] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 20 with a hidden size of 64 or larger, whereas the encoder 38 uses four convolutional layers with 16 channels and a GRU with hidden size of 16 and is thus much smaller than the neural network 60.

[0098] In the embodiment of Fig. 3, the EGM 54 receives its input signal from a microphone 62 that also captures the ambient sound in the environment of the user 5 which ambient sound includes the speech component S and the noise component N. The microphone 62 produces a second input audio signal I’ that is fed to an analysis filter bank 64 that corresponds to the analysis filter bank 30. The analysis filter bank 64 converts the time-dependent second input audio signal I’ (I’ = l’(t)) into a time-frequency representation X’ (X’ = X’(k,f)), similar to the time-frequency representation X). Said time-frequency representation X’ is fed to the analyzers 56 and 58. The analyzer 56 extracts from the time-frequency representation X’ a signal X’erb (X’erb = X’erb(k,ferb)), similar to the signal Xerb, that contains ERB (equivalent rectangular bandwidth) features. The analyzer 58 extracts a signal X’df (X’df = X’df(k,fdf)), similar to the signal Xdf, that contains complex features of the time-frequency representation X’. The neural network 60 reduces its input signal, i.e. the signals X’erb and X’df, to the low-dimensional embedding signal E.

[0099] The neural networks of the SEM 26 and the EGM 54 (i.e. encoder 38, the decoders 40, 42 and the neural network 60) are trained together to effectively suppress the noise component N in the first input audio signal I and, thus, produce the enhanced audio signal Y. The training process is performed preceding to the normal operation of the hearing system 2, during development of the latter. Thus, the hearing system 2 is manufactured with the SEM 26 and the EGM 54 being provided in a pre-trained state. Preferably, the SEM 26 and the EGM 54 are trained using gradient descent or a variant such as stochastic gradient descent or Adam. Herein, the EGM 54 is specifically trained to analyze the noise component N in the ambient sound.

[0100] In order to support analysis of the noise component N by the EGM 54, preferably, the EGM 54 is operated in non-speech intervals only, i.e. during time intervals in which the ambient sound does not contain the speech component S. To this end,

[0101] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 21 the hearing system 2 comprises a voice activity detector (VAD 66) that analyzes the first input audio signal I to detect presence (and absence) of the speech component S. The VAD 66 controls the EGM 54, e.g. via a control signal C, such that the EGM 54 is activated when the VAD 66 detects absence of speech, and deactivated, when the VAD 66 detects presence of speech.

[0102] In a variation of the embodiment of Fig. 3, the hearing system 2 may be implemented such that the VAD 66 receives its input from the microphone 62. In this case, preferably, the time-frequency representation X’ produced by the analysis filter bank 64 is fed to VAD 66.

[0103] Optionally, as a further measure to reduce the numerical requirements and energy consumption of the hearing system 2, the EGM 54 is operated intermittently such that it produces the embedding signal E with an (embedding generation) rate that is much lower than the sampling rate with which the SEM 26 updates the enhanced audio signal Y. For example, the EGM 54 is run once every second, whereas the SEM 26 produces the enhanced audio signal Y at the sampling rate of the hearing instrument 4, e.g., at 48 kHz.

[0104] The embodiment of the hearing system 2 according to Fig. 3 may be implemented in the arrangement of Fig. 1 or in the arrangement of Fig. 2. In the latter case, preferably, the EGM 54 uses the microphone of the mobile phone 22 as the microphone 62, and the analysis filter bank 64 is implemented as a part of the hearing app 20. Preferably, the VAD 66 is implemented as a part of the hearing app 20 (especially in embodiments in which it receives its input from the microphone of the mobile phone 22). In another embodiment, the VAD 66 is implemented as a part of the hearing instrument 4.

[0105] In another variation of the embodiment of Fig. 3, the hearing system 2 may be implemented such that the EGM 54 receives its input from the microphones 8 of the hearing instrument 4 (instead of from the microphone 62). In this case, preferably, the time-frequency representation X produced by the analysis filter bank 30 is fed

[0106] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 22 to the EGM 54. In this variation, the microphone 62 and the analysis filter bank 64 are not necessary and, preferably, not provided to the hearing system 2.

[0107] In the preferred implementation of the hearing system 2, all parts of the SEM 26 and the EGM 54 as well as the filterbanks 30, 48 and 64, the pre-processing unit 32, the post-processing unit 50 and the VAD 66 are implemented as software modules. However, as an alternative, any of these units may be realized as hardware circuitry.

[0108] Fig. 4 shows a further embodiment of the hearing system 2 which, unless indicated otherwise, corresponds to the embodiment shown in Fig. 3. However, in the embodiment of Fig. 4, the means 28 for supplying the embedding signal E include a sound classifier 70 and a selection unit 72 which replace the EGM 54 of Fig. 3. The sound classifier 70 indirectly analyzes the second input audio signal I’ produced by the microphone 62 to identify the current noise component N of the ambient sound with one of several predetermined noise classes K. In fact, it receives the time-frequency representation X’ as an input.

[0109] The sound classifier 70 delivers an indication of the sound class K that is found to correspond to the current noise component N to the selection unit 72. The selection unit 72 accesses a data storage 74 in which a plurality of different pre-generated embedding signals E are stored, wherein said plurality of pre-generated embedding signals E include one embedding signal E for each predetermined sound class K. The selection unit 72 selects the pre-generated embedding signal E from the data storage 74 that corresponds to the sound class K delivered from the sound classifier 70 and transfers the selected embedding signal E to the SEM 26, for configuration of the latter. The selection unit 72 repeats this embedding selecting step and, thus, updates the embedding signal E fed to the SEM 26, whenever the sound classifier 70 changes the noise class K into which the noise component N is classified.

[0110] The embodiment of the hearing system 2 according to Fig. 4 may be implemented in the arrangement of Fig. 1 or in the arrangement of Fig. 2. In the latter case,

[0111] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 23 preferably, the microphone of the mobile phone 22 is used as the microphone 62, and the analysis filter bank 64 is implemented as a part of the hearing app 20.

[0112] In variation of the embodiment of Fig. 4, the hearing system 2 may be implemented such that the sound classifier 70 receives its input from the microphones 8 of the hearing instrument 4 (instead of from the microphone 62). In this case, preferably, the time-frequency representation X produced by the analysis filter bank 30 is fed to the sound classifier 70. In this variation, the microphone 62 and the analysis filter bank 64 are not necessary and, preferably, not provided to the hearing system 2.

[0113] In the preferred implementation of the hearing system 2, the sound classifier 70 and a selection unit 72 are implemented as software modules. However, as an alternative, any of these units may be realized as hardware circuitry. The data storage 74 may be part of the hearing instrument 4. As an alternative, the hearing app 20 may use a part of the data storage of the mobile phone 22 as the data storage 74.

[0114] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the invention as shown in the specific examples without departing from the spirit and scope of the invention as broadly described in the claims. The present examples are, therefore, to be considered in all aspects as illustrative and not restrictive.

[0115] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 24

[0116] REFERENCES

[0117] [1] Y. Ephraim, D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,” IEEE Trans, on acoustics, speech, and signal processing, vol. 32, no. 6, pp. 1109-1121 , 1984;

[0118] DOI: 10.1109 / TASSP.1984.1164453

[0119] [2] P. Wang and D. Wang, “Enhanced spectral features for distortion-independent acoustic modeling,” in INTERSPEECH, 2019, pp. 476-480;

[0120] DOI: 10.21437 / lnterspeech.2O19-1493

[0121] [3] H. Schroeter, et al., “DEEPFILTERNET: A low complexity speech enhancement framework for full-band audio based on deep filtering”, in ICASSP 2022-2022 IEEE Int. Conf, on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 7407-7411 ; DOI:

[0122] 10.1109 / ICASSP43922.2022.9747055

[0123] [4] S. Liu, et al., “N-HANS: A neural network-based toolkit for in-the-wild audio enhancement”, Multimedia Tools and Appl. , vol. 80, no. 18, pp. 28365-28389, 2021 ; DOI: 10.1007 / s11042-021 -11080-y

[0124] [5] H. Schroeter, et al., “DEEPFILTERNET2: Towards real-time speech enhancement on embedded devices for full-band audio”, 2022 Int. Workshop on Acoustic Signal Enhancement (IWAENC), 2022, pp. 1 -5; DOI:

[0125] 10.48550 / arXiv.2205.05474

[0126] [6] A. Vaswani, et al., ..Attention Is All You Need", 31st Conf, on Neural Information Proc. Systems (NIPS 2017), 2017; DOI: 10.48550 / arXiv.1706.03762

[0127] [7] H. Schroter, et al., “CLCNet: Deep learning-based noise reduction for hearing aids using complex linear coding”, in IEEE Int. Conf, on Acoustics, Speech and Signal Processing (ICASSP), 2020; DOI: 10.48550 / arXiv.2001 .10218

[0128] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 25

[0129] LIST OF REFERENCE NUMERALS

[0130] 2 hearing system

[0131] 4 hearing instrument

[0132] 5 user

[0133] 6 housing

[0134] 8 microphone

[0135] 10 receiver

[0136] 12 battery

[0137] 14 signal processor

[0138] 16 sound channel

[0139] 18 tip

[0140] 20 hearing app

[0141] 22 mobile phone

[0142] 24 wireless link

[0143] 26 SEM (speech enhancement model)

[0144] 28 means (for supplying an embedding signal to the SEM)

[0145] 30 (analysis) filter bank

[0146] 32 pre-processing unit

[0147] 34 analyzer

[0148] 36 analyzer

[0149] 38 encoder

[0150] 40 decoder

[0151] 42 decoder

[0152] 44 gain applicator

[0153] 46 deep filter

[0154] 48 (synthesis) filter bank

[0155] 50 post-processing unit

[0156] 52 adder

[0157] 54 EGM (embedding generation model)

[0158] 56 analyzer

[0159] 58 analyzer

[0160] 60 neural network

[0161] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024 FDST Patentanwalte, Nurnberg Seite 26

[0162] 62 microphone

[0163] 64 (analysis) filter bank

[0164] 66 VAD (voice activity detector)

[0165] 70 sound classifier

[0166] 72 selection unit

[0167] 74 data storage

[0168] C control signal

[0169] C’ control signal

[0170] Cndeep filter coefficient

[0171] E embedding signal

[0172] Gerb real-value gain

[0173] I (first) input audio signal

[0174] I’ (second) input audio signal

[0175] K noise class

[0176] N noise component

[0177] 0 processed audio signal

[0178] U supply voltage

[0179] S speech component

[0180] X time-frequency representation

[0181] X’ time-frequency representation

[0182] Xerb signal

[0183] Xdf signal

[0184] X’ erb signal

[0185] X’df signal

[0186] Y enhanced audio signal

[0187] YG short-time spectrum

[0188] Z (first) intermediate signal

[0189] (\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024

Claims

FDST Patentanwalte, Nurnberg Seite 27CLAIMS1 . Method for operating a hearing system (2) including at least one hearing instrument (4), the method comprising:- a sound capturing step in which ambient sound from an environment of the hearing instrument (4) is captured by at least one first input transducer (8) of the hearing instrument (4) to produce a first input audio signal (I), wherein the ambient sound includes a speech component (S) and a noise component (N);- a signal processing step in which the first input audio signal (I) is modified, at least in part by a signal processor (14) of the hearing instrument (4), to produce a processed audio signal (0); and- an outputting step in which, by an output transducer (10) of the hearing instrument (4), the processed signal (0) is output as perceivable sound to a user of the hearing instrument (4); wherein the signal processing step further includes a speech enhancing step in which- a speech enhancement model (26) is applied to the first input audio signal (I), or an audio signal (X) derived therefrom to produce an enhanced audio signal (Y) in which the noise component (N) of the ambient sound is suppressed while the speech component is (S) retained; and- an embedding signal (E) that includes information characterizing the noise component is supplied to the speech enhancement model (26) to configure the speech enhancement model (26) for suppressing the noise component (N).

2. Method according to claim 1 , wherein the speech enhancement model (26) includes a neural network (38, 40, 42), in particular a convolutional neural network.

3. Method according to claim 1 or 2, wherein the speech enhancement model (26) is implemented in the hearing instrument (4).(\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024FDST Patentanwalte, Nurnberg Seite 284. Method according to any one of claims 1 to 3, further comprising an embedding selecting step in which- a sound classifier (70) is applied to the first input audio signal (I) or to a second input audio signal (I’) in which the ambient sound is captured by a second input transducer (62) of an external device (22), wherein the sound classifier (70) analyzes said noise component (N) to classify the noise component (N) into one of several predetermined noise classes (K), and- one of a plurality of pre-generated and stored embedding signals (E) is selected for being supplied to the speech enhancement model (26) based on the classification.

5. Method according to any one of claims 1 to 3, further comprising an embedding generating step in which an embedding generation model (54) is applied to the first input audio signal (I) or to a second input audio signal (I’) in which the ambient sound is captured by a second input transducer (62) of an external device (22), wherein the embedding generation model (54) analyses the noise component (N) of the ambient sound to generate the embedding signal (E) during operation of the hearing instrument (4).

6. Method according to claim 4 or 5, further comprising a voice activity detection step in which- a voice activity detector (66) is applied to said first input audio signal (I) or to said second input audio signal (I’), wherein the voice activity detector (66) analyzes the ambient sound to detect non-speech intervals in which the speech component (S) is not contained in the ambient sound, and- the embedding selecting step or the embedding generating step is only performed during detected non-speech intervals.

7. Method according to claim 5 or 6, wherein the embedding generation model (54) includes a neural network (60), in particular a convolutional neural network.(\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024FDST Patentanwalte, Nurnberg Seite 298. Method according to any one of claims 5 to 7, wherein the embedding generation model (54) is implemented as part of a software application (20) installed on an external computer, in particular a smartphone (22).

9. Method according to any one of claims 5 to 8,- wherein the speech enhancement model (26) updates the enhanced audio signal (Y) with a first sampling rate;- wherein the embedding generation model (54) updates the embedding signal (E) with a second sampling rate; and- wherein the first sampling rate exceeds the second sampling rate by a factor of at least 102, preferably 103, in particular 104.

10. Method according to any one of claims 2 to 9, wherein, in the speech enhancing step,- by an encoder part (38) of the speech enhancement model (26), the first input audio signal (I) or audio signal (X) derived therefrom is reduced to a first intermediate signal (Z) of lower temporal and / or spectral resolution and,- by at least one decoder part (40, 42) of the speech enhancement model (26), the first intermediate signal (Z) is expanded to a second intermediate signal (Gerb, Cn) of higher temporal and / or spectral resolution.11 . Method according to claim 10, wherein, in the speech enhancing step, the embedding signal (E) is applied to the first intermediate signal (Z), in particular by adding the embedding signal (E) to the first intermediate signal (Z).

12. Method according to any one of claims 5 or 11 , wherein the embedding generation model (54) and the encoder part (38) of the speech enhancement model (26) are implemented as neural networks of a same type differing only in the number of neurons and / or layers, wherein(\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024FDST Patentanwalte, Nurnberg Seite 30 the embedding generation model (54) has more neurons and / or more layers than the encoder part (38) of the speech enhancement model (26).

13. Hearing system (2) including at least one hearing instrument (4), the at least one hearing instrument (4) comprising:- at least one first input transducer (8) configured to capture an ambient sound from an environment of the hearing instrument (4) to produce a first input audio signal (I), wherein the ambient sound includes a speech component (S) and a noise component (N);- a signal processor (14) configured to modify the first input audio signal (I) to produce a processed audio signal (0); and- an output transducer (10) configured to output the processed audio signal (0) to a user of the hearing instrument (4); wherein the hearing system (2) further includes- a speech enhancement model (26) configured to produce an enhanced audio signal (Y) from the first input audio signal (I), or an audio signal (X) derived therefrom, by suppressing the noise component (N) of the ambient sound while retaining the speech component (S); and- means (28) for supplying an embedding signal (E) that includes information characterizing the noise component (N) to the speech enhancement model (26) to configure the speech enhancement model (26) for suppressing the noise component (N).

14. Hearing system (2) according to claim 13, wherein the speech enhancement model (26) includes a neural network (38, 40,42), in particular a convolutional neural network.

15. Hearing system (2) according to claim 13 or 14, wherein the speech enhancement model (26) is implemented in the hearing instrument (4), in particular as a part of the signal processor (14).

16. Hearing system (2) according to any one of claims 13 to 15, wherein the means (28) for supplying the embedding signal (E) include(\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024FDST Patentanwalte, Nurnberg Seite 31- a sound classifier (70) to which the first input audio signal (I) or to a second input audio signal (I’) are fed, wherein the second input audio signal (I’) contains the ambient sound captured by a second input transducer (62) of an external device (22), wherein the sound classifier (70) is configured to analyze the noise component (N) in the ambient sound to classify the noise component (N) into one of several predetermined noise classes (K); and- a selection unit (72) configured to select one of a plurality of predetermined embedding signals (E) for being supplied to the speech enhancement model (26), based on the classification.

17. Hearing system (2) according to any one of claims 13 to 15, wherein the means (28) for supplying the embedding signal (E) include an embedding generation model (54) to which the first input audio signal (I) or to a second input audio signal (I’) are fed, wherein the second input audio signal (I’) contains the ambient sound captured by a second input transducer (62) of an external device (22), wherein the embedding generation model (54) is configured to analyze the noise component (N) of the ambient sound to generate the embedding signal (E) during operation of the hearing instrument (4).

18. Hearing system (2) according to claim 16 or 17, wherein the means (28) for supplying the embedding signal (E) include a voice activity detector (66) to which said first input audio signal (I) or said second input audio signal (I’) are fed, wherein the voice activity detector (66) is configured to analyze the ambient sound to detect non-speech intervals in which the speech component (S) is not contained in the ambient sound, wherein the hearing system (2) is configured to activate- the sound classifier (70) and the selection unit (72) or- the embedding generation model (54) only during detected non-speech intervals.

19. Hearing system (2) according to claim 17 or 18,(\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024FDST Patentanwalte, Nurnberg Seite 32 wherein the embedding generation model (54) includes a neural network (60), in particular a convolutional neural network.

20. Hearing system (2) according to any one of claims 17 to 19, wherein the embedding generation model (54) is implemented as part of a software application (20) to be installed on an external computer, in particular a smartphone (22).21 . Hearing system (2) according to any one of claims 17 to 20,- wherein the speech enhancement model (26) is configured to update the enhanced audio signal (26) with a first sampling rate;- wherein the embedding generation model (54) is configured to update the embedding signal (E) with a second sampling rate; and- wherein the first sampling rate exceeds the second sampling rate by a factor of at least 102, preferably 103, in particular 104.

22. Hearing system (2) according to any one of claims 14 to 21 , wherein the speech enhancement model (26) includes- an encoder part (38) configured to reduce the first input audio signal (I), or audio signal (X) derived therefrom to a first intermediate signal (Z) of lower temporal and / or spectral resolution and- at least one decoder part (40,42) configured to expand the first intermediate signal (Z) to a second intermediate signal (Gerb, Cn) of higher temporal and / or spectral resolution.

23. Hearing system (2) according to claim 22, wherein the speech enhancement model (26) is implemented such that the embedding signal (E) is applied to the first intermediate signal (Z), in particular by adding the embedding signal (E) to the first intermediate signal (Z).

24. Hearing system (2) according to any one of claims 17 or 23, wherein the embedding generation model (54) and the encoder part (38) of the speech enhancement model (26) are implemented as neural networks of(\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024FDST Patentanwalte, Nurnberg Seite 33 a same type differing only in the number of neurons and / or layers, wherein the embedding generation model (54) has more neurons and / or more layers than the encoder part (38) of the speech enhancement model (26).(\\fs2012\gsi-software\winpat5\document\amt\3980667 docx) letzte Speicherung: 15 Oktober 2024

Citation Information

Patent Citations

  • Method for customizing audio signal processing of a hearing device and hearing device

    EP4345656A1