Method of operating an audio device system and an audio device system
The use of a time-domain FIR filter with a learned encoder and masker network in audio devices addresses latency issues in neural network processing, enabling real-time speech enhancement and noise suppression with adjustable latency and audiological control, improving user experience in complex acoustic scenarios.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- WS AUDIOLOGY AS
- Filing Date
- 2025-10-14
- Publication Date
- 2026-04-23
AI Technical Summary
Existing audio device systems, particularly hearing aids, face challenges in effectively suppressing noise and providing real-time speech enhancement due to algorithmic delays caused by end-to-end neural network processing, limiting their ability to handle complex hearing situations like the 'cocktail party' scenario.
Implementing a time-domain convolution using a Finite Impulse Response (FIR) filter controlled by a learned encoder and masker network, allowing for adjustable latency and audiological control, while decoupling latency from the encoder-decoder structure, and incorporating post-processing to minimize artifacts.
Enables low-latency, real-time speech enhancement with improved noise suppression and sound source separation, allowing for dynamic adjustments based on user preferences and environmental conditions, enhancing user experience in challenging acoustic environments.
Smart Images

Figure EP2025079553_23042026_PF_FP_ABST
Abstract
Description
[0001] METHOD OF OPERATING AN AUDIO DEVICE SYSTEM AND AN AUDIO
[0002] DEVICE SYSTEM
[0003] The present invention relates to a method of operating an audio device system. The present invention also relates to an audio device system adapted to carry out said method.
[0004] BACKGROUND OF THE INVENTION
[0005] An audio device system may comprise one or two audio devices. In this application, an audio device should be understood as a small, battery-powered, microelectronic device designed to be worn in or at an ear of a user. The audio device generally comprises an energy source such as a battery or a fuel cell, at least one microphone, a microelectronic circuit comprising a digital signal processor, and an acoustic output transducer. The audio device is enclosed in a casing suitable for fitting in or at (such as behind) a human ear.
[0006] If the audio device is capable of amplifying an ambient sound signal and hereby alleviate to some extent a hearing deficit of the user but without e.g. being able to take into account a specific hearing loss of the user, then such an audio device may be denoted e.g. a personal sound amplification product. If on the other hand the audio device is specifically adapted to alleviate or rather compensate a user’s hearing loss then the audio device should be denoted a hearing aid.
[0007] However, audio device systems generally and (at least) both the above mentioned types may additionally benefit from various speech enhancement features such as speaker separation as will be explained in the following.
[0008] According to variations the mechanical design of an audio device may resemble those of hearing aids and as such traditional hearing aid terminology may be used to describe various mechanical implementations of audio devices that are not hearing aids. As the name suggests, Behind-The-Ear (BTE) hearing aids are worn behind the ear. To be more precise, an electronics unit comprising a housing containing the major electronics parts thereof is worn behind the ear. An earpiece for emitting sound to the hearing aid user is worn in the ear, e.g. in the concha or the ear canal. In a traditional BTE hearing aid, a sound tube is used to convey sound from the output transducer, which in hearing aid terminology is normally referred to as the receiver, located in the housing of the electronics unit and to the ear canal. In more recent types of hearing aids, a conducting member comprising electrical conductors conveys an electric signal from the housing and to a receiver placed in the earpiece in the ear. Such hearing aids are commonly referred to as Receiver-In-The-Ear (RITE) hearing aids. In a specific type of RITE hearing aids the receiver is placed inside the ear canal. This category is sometimes referred to as Receiver-In-Canal (RIC) hearing aids. In-The-Ear (ITE) hearing aids are designed for arrangement in the ear, normally in the funnel-shaped outer part of the ear canal. In a specific type of ITE hearing aids the hearing aid is placed substantially inside the ear canal. This category is sometimes referred to as Completely-In-Canal (CIC) hearing aids or Invisible-In-Canal (IIC). This type of hearing aid requires an especially compact design in order to allow it to be arranged in the ear canal, while accommodating the components necessary for operation of the hearing aid.
[0009] Generally, a hearing aid system according to the invention is understood as meaning any device which provides an output signal that can be perceived as an acoustic signal by a user or contributes to providing such an output signal, and which has means which are customized to compensate for an individual hearing loss of the user or contribute to compensating for the hearing loss of the user.
[0010] Within the present context an audio device system may comprise a single audio device (a so called monaural audio device system) or comprise two audio devices, one for each ear of the user (a so called binaural audio device system). Furthermore, the audio device system may comprise at least one additional device (which in the following may also be denoted an external device despite that it is part of the audio device system), such as a smart phone or some other computing device having software applications adapted to interact with other devices of the audio device system. However, the audio device system may also include a remote microphone system (which generally can also be considered a computing device) comprising additional microphones and / or may even include a remote server providing abundant processing resources and generally these additional devices will also include link means adapted to operationally connect to the various other devices of the audio device system. Despite the advantages that contemporary audio device - and especially hearing aid - systems provide, some users may still experience hearing situations that are difficult. A critical element when seeking to alleviate such difficulties is the audio device systems ability to suppress noise.
[0011] Speech enhancement has been a significant interest to the signal processing community already from the 1960's and earlier, see e.g. Y. Ephraim and D. Malah: “Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator,” IEEE Trans. Acoustics, Speech and Signal Processing, vol. ASSP-32, no. 6, pp. 1109-1121, December 1984.
[0012] The basic problem concerns the estimation of a clean underlying speech signal from one or more noisy / degraded signals. A more general problem is that of speaker separation, in which one attempts to separate a mixture of (potentially noisy) speech signals into separate clean speech signals.
[0013] Hearing impairment limits the ability to cope with background noise. Therefore, a noise reduction (speech enhancement) system is usually seen as an integral part of a hearing aid. When asked about desired improvements of their hearing aids, users often request more effective systems for attenuating (or otherwise managing) background noise.
[0014] Traditionally, speech enhancement systems have been based on manually devised statistical models of speech and noise, which in turn lead to signal processing based methods of optimally removing noise from noisy speech. These approaches are limited by the ability to accurately describing speech and noise by statistical models, while keeping these models simple enough to allow closed form derivations of the relevant estimators.
[0015] One particularly difficult hearing situation is the so called cocktail party situation where multiple speakers are present at the same time and typically positioned close together.
[0016] It has therefore been suggested to provide separation of speakers in order to suppress undesired speakers. Traditionally this has been provided using e.g. various beamforming techniques and more recently speaker separation (which in the following may also be denoted source separation or sound source separation) has been demonstrated based on specifically trained neural networks.
[0017] Recently, machine learning (artificial intelligence) has revolutionized speech enhancement (along with many other scientific disciplines). Speech enhancement is essentially a simple regression problem, where a clean signal is estimated from a noisy counterpart. Plentiful training data can easily be obtained by mixing clean speech recordings with separate noise recordings (which in the following may also be called synthesizing speech in noise signals). Proposed methods vary primarily according to their model architecture (how the neural network is configured), the data used for training, and the objective (loss) optimized during training.
[0018] A variety of cost functions may be used to train a neural network for speech enhancement in the context of the present invention. These cost functions are suitable for optimizing the neural network to minimize the difference between the output of the audio device system and a target signal representing only the desired sound source(s). Examples of such cost functions include:
[0019] - Mean Squared Error (MSE), which is defined as:
[0020] MSE = (1 / T) * S (s(t) - s(t))2where, s(t) is the target clean speech signal, s(t) is the output signal from the system and T is the number of time samples.
[0021] This cost function directly measures the squared error between the desired clean speech and the system output, and is commonly used in regression-based training of neural networks.
[0022] - another example of a cost function is the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR), The SI-SDR loss is defined as:
[0023] SI-SDR = 10 * logw ( | |as| |2 / 1 |as - s||2) where: s is the target signal, s is the estimated signal and a is a scaling factor to align the energy of the signals.
[0024] This loss function is robust to gain variations and aligns well with perceptual quality, making it suitable for training neural networks in audio enhancement tasks. One particularly successful neural-network-based method is the TASNet, see e.g.: Y. Luo and N. Mesgarani, "TaSNet: Time-Domain Audio Separation Network for Real- Time, Single-Channel Speech Separation," 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, 2018, pp. 696-700. Herein it is suggested to model the signal in the time-domain using an encoder-decoder framework and perform the source separation on nonnegative encoder outputs, whereby the separation problem is reduced to estimation of source masks on encoder outputs that are then synthesized by the decoder.
[0025] More specifically, it is suggested to use a low-complexity encoder and decoder module to generate a down sampled multidimensional representation of an audio signal that lends itself more effective to speech enhancement. This principle is similar to how a filterbank is applied in conventional hearing aid processing. However, the encoder and decoder parameters are learnt during the training procedure of the system, such as to generate a maximally efficient representation. Furthermore, the encoder and decoder can involve more steps than a conventional filterbank, e.g. as illustrated in above mentioned TASNet paper.
[0026] Reference is now made to Fig. 1 for illustrating the basic building blocks of a TASNet neural network in the context of an audio device system 100. It is attempted to follow the notation of the paper: Yi Luo and Nima Mesgarani, “Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,” IEEE / ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 27, no. 8, pp. 1256-1266, 2019.
[0027] The TASNet neural network considers a single channel input, x(t) (which in the following may also be denoted an audio input signal or an input signal), that is provided by the acoustical electrical input transducer 101, for t = 0,. . ,,T, which is split into overlapping chunks Xk E RlxLfor k = 1, . . . , T . These frames are given by Xk = [x(AAT), (kM + 1), . . . , (kM + L— 1 )] for down sampling factor Aland frame length L. We drop the index k for notational convenience.
[0028] The encoder 102 transforms each frame according to w = H(xU) where U E RLxNis a linear transformation and H(-) is an optional non-linear function. The decoder, similarly, uses a linear transformation, to map the encoded representation, w, back to the time domain x = wV, with V E RNxL. The resulting segments are overlap-added to form the final output waveform.
[0029] Following the encoder 102, a masker-network 103 is used to compute a mask, d =W), which is multiplied, using a multiplier 104, onto the encoded signal representation (i.e. the output from the encoder 102), w = d ■ w. This serves the purpose of attenuating unwanted parts of the signal while maintaining desired ones. Next the output from the multiplier 104 is provided to the decoder 105 and finally to the electrical-acoustical output transducer 106, in order to provide a sound signal with the desired parts of the signal from the input transducer being maintained while the unwanted parts are being attenuated.
[0030] See e.g. Cem Subakan and Mirco Ravanelli and Samuele Cornell and Frederic Lepoutre and Francois Grondin, "Resource-Efficient Separation Transformer" 2022 or any of the two above mentioned TASNet references for additional examples of how such a masker network can be structured.
[0031] The system is trained end-to-end, meaning that the encoder, masker, and decoder are trained together.
[0032] This could, for example, be done by minimizing a mean-squared-error loss function. Let the input be a sum of speech and noise, x = s n We can then train the system to minimize the distance between the output and the clean speech in isolation min\ \s s\ |2, where s is the output of the TASNet.
[0033] When applying a TASNet in a real-time system, such as a hearing aid, it results in a delay. Firstly, it produces a purely algorithmic delay (i.e., unrelated to processing speed). This happens because the system must collect the entirety of a frame before processing it. E.g., when processing the frame Xk = [x(kM), x(kM + 1), , , , x(kM + L - 7)], we must wait for sample kM + L 1 to arrive. Only when this has occurred can we start processing the entire frame and begin presenting the user with output samples from the processed frame, starting with sample kM, thus at a delay of L samples. On top of this, we have to add the delay from the computational time involved in processing the frame (i.e., from the encoder, masker, and decoder blocks). Consequently, there is a need for an audio device system (such as a hearing aid system) capable of both providing neural network based speech enhancement and a low signal processing delay.
[0034] SUMMARY OF THE INVENTION
[0035] The invention is set out in the appended set of claims.
[0036] BRIEF DESCRIPTION OF THE DRAWINGS
[0037] By way of example, there is shown and described a preferred embodiment of this invention. As will be realized, the invention is capable of other embodiments, and its several details are capable of modification in various, obvious aspects all without departing from the invention. Accordingly, the drawings and descriptions will be regarded as illustrative in nature and not as restrictive. In the drawings:
[0038] Fig. 1 illustrates highly schematically an audio device system according to an embodiment of the prior art;
[0039] Fig. 2 illustrates highly schematically an audio device system; according to an embodiment of the present invention;
[0040] Fig. 3 illustrates highly schematically a hearing aid system; according to an embodiment of the present invention; and
[0041] Fig. 4 illustrates highly schematically a hearing aid system; according to a more advanced embodiment of the present invention.
[0042] DETAILED DESCRIPTION
[0043] In the present context the term “audio signal” will generally be construed to mean an electrical (analog or digital) signal representing a sound. A beamformed signal (either monaural or binaural) is one example of such an electrical signal representing a sound. Another example is an electrical signal wirelessly streamed to the audio device system. However, the audio signal may also be internally generated by the audio device system.
[0044] More specifically the term “audio input signal” will generally be construed to mean an electrical signal representing a sound from the sound environment, but the term “audio input signal” may also be construed to mean an electrical signal representing a beamformed audio signal. In the following a beamformed audio signal may also be denoted a signal derived from the sound environment.
[0045] Similarly, the term “audio output signal” will generally be construed to mean an electrical signal representing a sound to be output by an electrical-acoustical output transducer of an audio device of an audio device system.
[0046] Furthermore, in the present context the terms “sound source signal” and “separated sound source signal” may be used interchangeably, since both terms are used to describe signals that primarily represent a single sound source - typically in the form of a human speaker.
[0047] In a similar manner a sound source signal may or may not be specifically denoted a “latent space sound source signal” when considered in an embodiment comprising an encoder-decoder neural network.
[0048] While the term “sound source” does not represent the same as a “sound source signal” then the terms can sometimes be considered interchangeable, e.g. with respect to selecting a specific (e.g. a first) sound source signal, because in case a first sound source signal is selected then this necessarily implies that the corresponding sound source can likewise be considered selected.
[0049] More specifically this means that e.g. in the context of training a neural network it is considered implicit that e.g. when an input signal (or a target signal) is described as comprising e.g at least one desired sound source and / or at least one noise source then this means that the input signal comprises a signal from at least one of these sources.
[0050] Finally the terms source and sound source may also be considered interchangeable. In the present context the term signal processing is to be understood as any type of hearing aid system related signal processing that includes at least: noise reduction, speech enhancement and hearing compensation.
[0051] In the following the terms “masker” and “masker network” may be used interchangeably.
[0052] 1. Fig. 2 embodiment: Audio device system
[0053] Reference is now made to Fig. 2, which highly schematically illustrates an audio device system 200 according to an embodiment of the invention. The audio device system 200 is adapted to provide speech enhancement similar to the TASNet audio device system of Fig. 1, but also distinguishes the Fig. 1 system, e.g. in that the time-domain multiplication of the audio device system 100 of Fig. 1 is replaced by a time-domain convolution in the audio device system 200 of Fig. 2.
[0054] Thus, the audio device system 200 enhances the input signal that is provided by the acoustical electrical input transducer 101 by applying a time-varying filter 204, such as a Finite Impulse Response (FIR) filter, directly in the time domain. The FIR-filter is controlled by a masker 202 that operates in the domain of a learned encoder 201 (which in the following may simply be denoted encoder). However, it is noted that the encoder 201 is only optional since it is not a strictly necessary part of the invention, but the encoder 201 is advantageous because it allows the masker 202 to operate at a lower sample rate. The encoder 201 can be the same as in a TASNet (as explained above) and generate a representation from frames of the input signal, w = H(xU). Subsequently, the masker 202 outputs a partly or fully specified FIR-filter, z = g(w), with z E RlxQ. This filter specification does not need to have the same dimensionality as either the encoder output (R1XN) or the final FIR-filer (RlxD). The specification could be (but is not limited to):
[0055] • A full FIR filter specified as an impulse response.
[0056] • An amplitude response that can be converted to a minimum-phases filter.
[0057] • A complex frequency response.
[0058] • A post-processing block 203, v = h(z), with z E RlxD, handles any relevant postprocessing to the masker output. Thus comprising at least one of: o Converting the masker output to a valid FIR filter. This could include the addition of relevant signal processing domain knowledge, leading to a FIR-filter with favorable properties (such as a minimum-phase filter). o Smoothing across time to avoid audible artifacts, which represents one way to modify the audio device system 200 after training of the neural networks, in order to introduce audiological / psychoacoustic domain knowledge. Thus more specifically the post processing block is configured to apply temporal smoothing to the FIR filter coefficients generated by the neural network. This smoothing reduces abrupt changes in the filter response over time, thereby minimizing audible artifacts and improving the perceptual continuity of the output signal. The smoothing may be implemented using moving average filters, exponential decay functions, or other suitable techniques known in the art. o Introducing adjustable handles. By doing so, it becomes possible to finetune the sound of the system 200 without having to re-train the neural network(s) entirely from scratch at every design iteration. More specifically these adjustable handles or parameter adjustment interface can enable parameters as gain, bandwidth, and attenuation thresholds to be tuned in real-time or during calibration. This enables audiologists or system designers to fine-tune the system response based on user preferences or environmental conditions, while preserving the integrity of the trained neural network.
[0059] We expect the masker network 202 to produce a time-varying filter representation at the same rate as the encoder output. The postprocessing block 203 is responsible for up-sampling this representation to the full sampling rate. In effect, the post-processing block will have to output M v-vectors for each input.
[0060] Thus it is noted that the post-processing block 203, while being advantageous may still be omitted e.g. for reasons of saving processing power or simply as a result of the masker output being of so high quality that post-processing is not required. In fact this is true for all embodiments herein that comprises the post processing block (203, 303 and 403).
[0061] Anyway, the resulting FIR filter (204) is convolved onto the input signal provided by the acoustical electrical input transducer 101, leading to an enhanced signal, s(t) = x(t) * v(t). Similar to the conventional TASNet, this system is trained end-to-end, i.e., the encoder 201 and masker 202 are trained together. The postprocessing-block 203 and the time-domain convolution (provided by the FIR filter 204 ) must also be included in the training stage, but are not trained as such, since they contain no trainable parameters.
[0062] It is a specific advantage of the present invention that the time-domain convolution (as provided by the digital filter 204, e.g. in the form of a FIR filter) is the only operation applied on the path from input to output. This means that the algorithmic latency can be varied from 0 to D-l samples, purely by the specification of the FIR filter, v(t). Notably, this makes it possible for the masker-network (or the post-processing system) to control the latency of the system (as opposed to the fixed latency of a TASNet). The latency of the overall system is entirely decoupled from the specification of the encoder, which determines the latency of a TASNet.
[0063] Thus overall the proposed audio device system, according to Fig. 2, comes with a number of advantages, compared to the TASNet audio device system according to Fig. 1 :
[0064] One advantage is greater control of latency. Latency can be varied freely, without connection to the latency incurred form the use of an encoder / decoder- structure. The latency may even be time-varying, so that a lower latency may automatically occur in situations that require less processing of the input signal.
[0065] Another advantage is the ability to have at least some audiological control over the output of the masker network. The output of the masker (e.g. 103 in Fig. 1) in a conventional TASNet is represented in the unknown learned domain of the encoder (102 in Fig. 1). This makes it virtually impossible to post-process the masker output by using domain knowledge from audiology, psychoacoustics, and signal processing.
[0066] However, in the proposed method according to the embodiment of Fig. 2 the output of the masker network (202) is an interpretable representation of a FIR filter that can be post-processed before being applied. This allows for using domain knowledge to limit artifacts, control the activity level of the system, to do real-world tweaking of the system by adjusting parameters on the fly, and potentially many other benefits.
[0067] According to one specifically advantageous aspect the post processing block 203 is adapted to ensure that the digital filter 204 is a minimum phase FIR filter. Obviously, the post processing block (303, 403) according to the other embodiments can be adapted in a similar manner to ensure that the digital filters (304 and 404-a and 404-b) are also minimum phase filters.
[0068] 2. Fig. 3 / Hearing Aid System
[0069] Reference is now made to Fig. 3, which highly schematically illustrates a hearing aid system 300 according to another embodiment of the invention.
[0070] This embodiment comprises two input transducers 101-a and 101-b that each provides an input signal to both a spatial processing block 306 and an encoder 301. It is noted that the various embodiments of the present invention can be implemented both with a single input transducer and with two (or more) input transducers, e.g. by including input transducers from at least one external device, such as a smart phone or a contralateral audio device or a contra-lateral hearing aid as part of a binaural audio device system or binaural hearing aid system as will be clear for the skilled person. Thus, it will likewise be clear for the skilled person how to adapt the spatial processing block dependent on the number and type of input transducers.
[0071] Additionally, this embodiment comprises a hearing loss compensation block 305 that is adapted to compensate or at least alleviate the hearing loss of a specific user’s hearing loss, basically by ensuring that relevant sounds are amplified above the specific user's hearing loss. Obviously, the skilled person will also be able to figure out how to implement that, since standard hearing aid technology can be applied with very few - if any - adaptations required.
[0072] Thus, the hearing aid system 300 is similar to the audio device system 200 with respect to the main idea being to use the (digital) FIR filter 304 to apply a frequency dependent gain to provide e.g. sound source separation, wherein the frequency dependent gain is determined by a neural network comprising an encoder 301, a masker 302 and optionally (but preferably) a post processing block 303. Generally, this can be implemented using various different types of neural networks or more general machine learning components, e.g. (but not limited to) fully connected layers, convolutional layers, recurrent neural network layers such as GRU or LSTM or attention layers. Method of training the Hearing Aid system of Fig. 3
[0073] According to the embodiment of Fig. 3, the encoder 301, masker 302 and post processing block 303 are initially trained alone (i.e. without e.g. the hearing loss compensation block 305). It is noted that in the following the encoder 301, masker 302 and post processing block 303 may be denoted the enhancement system.
[0074] According to the embodiment of Fig. 3, the training is carried out using e.g. the well known technique based on (in a first step) synthesizing a plurality of training audio signals, wherein each of said plurality of synthesized training audio signals comprises at least one of: at least one desirable sound source, and at least one type of noise (i.e. noise sources) such as motor noise, machine noise, street noise and speakers that the hearing aid system user has no intention of listening to, and wherein said desirable sound sources and the noise sources are located at different positions around the user. In a subsequent step said enhancement system (i.e. the encoder 301, the masker 302 and the post processing block 303) is trained (using said plurality of synthesized training audio signals) to control the FIR filter to provide an output signal representing at least one desired sound source when the input signal comprises such at least one desired sound source.
[0075] Beamforming
[0076] Thus according to the embodiment of Fig. 3 the training comprises providing as input to the encoder 301 both the microphone signals from the two input transducers 101-a and 101-b and the output from the spatial processing block 306. In an alternative embodiment only the output from the spatial processing block 306 is provided as input to the encoder.
[0077] Thus, it is noted that according to the embodiment of Fig. 3 the enhancement system is provided with information about the input to (and output from) the spatial processing block 306 but has no impact on it).
[0078] In a further embodiment, the spatial processing block is configured to perform adaptive beamforming based on environmental cues or user intent. For example, the beamforming may be dynamically adjusted to focus on a speaker located in a specific direction or to suppress dominant noise sources identified in the acoustic scene. Additionally, the spatial processing block may provide metadata or directional cues to the neural network, enabling context-aware enhancement. This metadata may include estimated source locations, signal-to-noise ratios, or confidence scores, which can be used by the encoder or masker to refine the filtering strategy.
[0079] The integration of spatial processing with neural network-based enhancement allows the system to leverage both physical microphone array geometry and learned signal representations, resulting in improved separation of overlapping sound sources and enhanced speech intelligibility in complex environments.
[0080] According to any of the disclosed embodiments, at least one desired sound source can be selected using a variety of different methods. One method is based on determining the at least one speaker participating in a conversation that the audio or hearing aid system user is paying attention to (as is further described in e.g. the patent application WO-A1-2023085577). Another method is based on selecting the at least one desired sound based on speaker recognition or based on a desired position (e.g. represented by a specific direction relative to the user of the hearing.
[0081] Thus according to the embodiment of Fig. 3, the hearing loss compensation block 305 is not involved in the training process, instead it is just added subsequently, i.e. down stream of the digital filter. In a similar manner the remaining hearing aid blocks such as the feedback cancellation system are just added subsequently. Thus after the training has been carried out, the trained encoder 301, masker 302 and post processing block 303 are simply integrated in the full hearing aid system 300 (wherein most hearing aid blocks such as e.g. the feedback cancellation system are omitted in Fig. 3 (and Fig. 4) for reasons of clarity).
[0082] 3. Fig. 4 Hearing Aid System with Neural Network controlled Beam Forming
[0083] Reference is now made to Fig. 4, which highly schematically illustrates a hearing aid system 400 according to an embodiment of the invention.
[0084] This embodiment comprises two input transducers 101-a and 101-b that each provides an input signal to respective digital filters 404-a and 404-b and to an encoder 401.
[0085] Thus this embodiment distinguishes the embodiment according to Fig. 3, in that the enhancement system of Fig. 4 (in the form of the encoder 401, masker 402 and post processing block 403) is trained to control two digital filters (401 -a and 401-b) that each receive an input signal from a respective one of the two input transducers 101-a and 101-b and wherein the output signals from the two digital filters (401 -a and 401-b) are subsequently combined in the summation unit 407. Hereby the beamforming properties of the Fig. 4 embodiment are controlled by the enhancement system, that consequently can be trained to further improve the sound source separation in at least some acoustical scenarios.
[0086] Thus according to the embodiment illustrated in Fig. 4, the audio device system comprises two input transducers, each connected to a respective digital filter. These digital filters are controlled by the neural network, which is trained to adapt their filter coefficients based on the spatial characteristics of the input signals. The outputs of the two digital filters are combined to form a beamformed signal.
[0087] This configuration allows the neural network to perform spatial filtering directly, enabling dynamic beamforming that enhances a desired sound source while suppressing interfering signals. The neural network may be trained end-to-end to optimize both the individual filter responses and the combined output for improved speech enhancement in complex acoustic environments. According to a variation of the embodiments according to respectively Fig. 2, Fig. 3 and Fig. 4, the encoder (201, 301 and 401) and masker (201, 302 and 402) are replaced by a Deep Neural Network (DNN).
[0088] Thus the inventors have found that the use of an encoder - masker structure is not required in order to obtain a high quality sound source separation.
Claims
CLAIMS1. A method of operating an audio device to provide speech enhancement comprising the steps of:- providing a stream of input samples from an acoustical-electrical input transducer or derived from at least one acoustical-electrical input transducer;- providing said stream of samples to a main signal path comprising a first digital filter wherefrom a stream of filtered output samples is provided to an output transducer and converted into sound; and- providing said stream of input samples to an analysis branch comprising a neural network trained to provide speech enhancement of said input samples provided to the main signal path by controlling said first digital filter, wherein said neural network has been trained using a plurality of synthesized training signals, each comprising at least one desired sound source and at least one noise source, and corresponding target signals representing only the at least one desired sound source, and wherein the training comprises minimizing a cost function representing the difference between the output signal resulting from filtering the input signal by said first digital filter and the corresponding target signal.
2. The method according to claim 1, wherein said desired sound source is silence if the input signal only comprises at least one noise source.
3. The method according to any of the preceding claims, wherein said neural network comprises an encoder neural network and a masker neural network.
4. The method according to any of the preceding claims, said analysis branch comprises a post processing block configured to ensure that said first digital filter is adapted to operate as a minimum phase or low latency digital filter and wherein said post-processing block in non-trainable and based on input from the neural network.
5. The method according to claim 4, wherein the post processing block comprises a smoothing function adapted to reduce temporal artifacts by smoothing the FIR filter coefficients across time.
6. The method according to claim 4 or 5, wherein the post processing block further comprises a parameter adjustment interface allowing modification of filter behaviour without retraining the neural network.
7. The method according to any of the preceding claims, wherein said speech enhancement comprises sound source separation.
8. The method according to any of the preceding claims, wherein the spatial processing block is configured to perform adaptive beamforming adapted to enhance signals originating from a desired direction relative to the user based on at least one of environmental conditions or user intent.
9. The method according to claim 8, wherein the spatial processing block is further configured to provide metadata comprising directional cues or source location estimates to the neural network for context-aware enhancement.
10. The method according to any of the preceding claims, wherein the audio device comprises two input transducers, and the neural network is trained to control a respective digital filter for each input transducer, separate from said first digital filter, and wherein the outputs of said respective digital filters are combined to provide a beamformed signal.
11. The method according to claim 10, wherein the neural network is trained to adapt the filter coefficients of each of said digital filters to enhance a desired sound source based on spatial separation of input signals.
12. The method according to claim 10 or 11, wherein the neural network is trained to jointly optimize the filter coefficients of each of said digital filters to enhance a desired sound source based on spatial separation of input signals.
13. The method according to any of the preceding claims, wherein at least some of said input and target signals have been provided by synthesizing a mix of signals comprising at least one sound source and / or at least one noise source.
14. An audio device system comprising:- at least one acoustical-electrical input transducer configured to provide a stream of input samples;- a main signal path comprising a first digital filter configured to receive said18 input samples and provide a stream of filtered output samples to an output transducer;- an analysis branch comprising a neural network trained to provide speech enhancement by controlling said first digital filter, via filter coefficient adaptation, based on said input samples; wherein the neural network of said analysis branch has been trained using synthesized training signals comprising at least one desired sound source and / or at least one noise source, and corresponding target signals representing only the desired sound source.
15. The audio device system according to claim 14, further comprising a post processing block configured to convert the output of the neural network into at least one valid minimum-phase or low-latency FIR filter.
16. The audio device system according to claim 14-15, wherein the input samples are derived from a spatial processing block configured to perform beamforming based on signals from at least two input transducers.
17. The audio device system according to any of the claims 14-16 wherein the neural network is trained to control a respective digital filter for each of at least two input transducers, separate from said first digital filter, and wherein the outputs of said respective digital filters are combined to provide a beamformed signal.
18. The audio device system according to any of claims 14-17, wherein said audio device system is a hearing aid system.
19. The audio device system according to any of claims 14-18, wherein the neural network has been trained by minimizing a cost function representing the difference between the output signal and a target signal representing only at least one desired sound source.
Citation Information
Patent Citations
System and method for assisting selective hearing
US12069470B2
Estimating an optimized mask for processing acquired sound data
US20240212701A1
Low-latency noise suppression
US20240331716A1
Audio processing based on target signal-to-noise ratio
WO2024205944A1