Training method of algorithm for extracting at least one required component

By using the encoder training algorithm, the masked intra-domain data elements and the encoder parameters are optimized, the inaccurate problem of target extraction of artificial neural networks in speech enhancement and noise reduction is solved, and the extraction accuracy is improved and suitable for small devices.

CN120048273APending Publication Date: 2025-05-27OTICON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411717622.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, when using artificial neural networks for speech enhancement and noise reduction, there is a mismatch between the training data and the actual use environment, resulting in inaccurate target extraction.

Method used

By using the encoder training algorithm, masked in-domain data elements are obtained, encoder parameters are determined to optimize prediction of noisy components, ensuring that the data elements used by the algorithm during runtime match the training data.

Benefits of technology

Improves the accuracy of target extraction and maintains the reasonable size of the algorithm in devices such as hearing aids, suitable for limited available memory and processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048273A_ABST
    Figure CN120048273A_ABST
Patent Text Reader

Abstract

A method of training an algorithm for extracting at least one desired component, the algorithm comprising an encoder and a first decoder, each of the encoder and the first decoder comprising at least one parameter, where the method comprises training the encoder and the first decoder, comprising: obtaining at least a portion of masked intra-domain data elements, wherein at least a portion of the masked intra-domain data elements comprise noise components; values of at least one parameter of the encoder and at least one parameter of the first decoder are determined using the at least partially masked intra-domain data elements to optimize prediction of noisy components in the at least one masked portion of the at least partially masked intra-domain data elements by the first algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for training an algorithm, such as an artificial neural network, for extracting at least one desired component, usually a desired component of a sound signal.

[0002] The present invention relates to the field of speech enhancement and noise reduction in, for example, hearing aids, earphones, headsets, hands-free telephone systems, mobile phones, teleconferencing systems, classroom amplification systems, etc. In particular, the present invention relates to speech enhancement and noise reduction using artificial neural networks. Background Art

[0003] The sound signal heard by a person usually includes a noisy component including a combination of a target component and a noise component. The target component is a part of the sound signal generated by a desired sound source and is, for example, a speech signal, and the noise component is a part of the sound signal generated by at least one noise source or the reflection of the sound signal generated by the desired sound source on many surfaces.

[0004] Since the intelligibility of the target component is poor in a noisy environment, especially for hearing aid users, extracting the target component through a hearing aid is very important to help its users understand speech in a noisy environment.

[0005] Speech enhancement and noise suppression in hearing aids can be improved using neural networks that are trained to suppress the noise component of the sound signal and preserve the target component of the sound signal (eg, speech).

[0006] It is known to train a neural network using methods based on supervised learning. Such supervised training requires a clean target signal paired with a noisy signal to represent the same acoustic situation (the noisy signal thus includes a clean target signal and a noise signal).

[0007] However, it is usually not feasible to record a clean target signal and a noisy signal simultaneously. Instead, a pair of simulated (ie artificially generated) sound signals is used.

[0008] The sound signal is generated by adding a clean speech signal to a noise signal to generate a mixed signal. The clean signal and the noise signal are both taken from a publicly available database, for example. However, neither the clean signal nor the mixed signal has the specific characteristics of the sound signal picked up by the microphone in the hearing aid (the sound signal is therefore called an out-of-domain data element).

[0009] In other words, when the neural network is deployed in a hearing aid, there is a mismatch (or difference) between the sound signal used for training and the sound signal used by the trained neural network.

[0010] The mismatch may be caused by different noise types, or different acoustic properties, such as different coloration of the sound due to microphone type, placement type, distance from the target talker, or reverberation patterns.

[0011] This results in inaccurate object extraction by the trained neural network. Therefore, a solution to improve the accuracy of object extraction is needed. Summary of the invention

[0012] Training method of the first algorithm

[0013] One aspect of the present invention is to provide a training method for an algorithm (hereinafter referred to as "first algorithm") for extracting at least one desired component (hereinafter referred to as "desired component") of a sound signal. The first algorithm includes an encoder, and the encoder includes at least one parameter.

[0014] The method of the present invention includes training an encoder, which includes obtaining at least a portion of masked in-domain data elements, wherein at least a portion of the masked in-domain data elements include noise components.

[0015] The training of the encoder further includes determining a value of at least one parameter of the encoder using at least a portion of the masked in-domain data elements.

[0016] The value of at least one parameter of the encoder is determined to optimize the prediction of the first algorithm for the noisy component in at least a masked portion of the at least partially masked in-domain data elements.

[0017] In-domain data elements are data elements based on sound signals obtained in the same or similar sound environment / acoustic setting as the sound signals used by the trained first algorithm, for example having the same or similar noise type and / or the same or similar acoustic properties, for example the sound has the same or similar coloration due to the same or similar microphone / input transducer type, microphone placement type, head and torso characteristics of the individual wearing the microphone (e.g., a hearing aid user), distance from the target talker, or reverberation pattern.

[0018] Thus, the first algorithm is trained based on data elements that match the data elements used during runtime after its training (the data elements used are, for example, sound signals picked up by a microphone in a hearing aid). As a result, the accuracy of target extraction is improved.

[0019] Furthermore, the improvement in accuracy is achieved while keeping the algorithm at a reasonable size, which is advantageous when the trained algorithm is stored in a small battery-powered device with limited available memory and processing power, such as a hearing aid.

[0020] Furthermore, masking of data elements enables using in-domain data elements comprising noisy components for training, i.e. determining parameters / weights of the encoder (and decoder) of the first algorithm, which is easy to obtain since this does not mean obtaining noisy signals simultaneously with potentially clean target signals.

[0021] In the present invention, the first algorithm comprises arithmetic and / or algebraic steps for extracting the target component (in this case, the desired component) from the noisy component of the sound signal when training the first algorithm.

[0022] The first algorithm generally includes a machine learning model such as an artificial neural network. The artificial neural network may include a multi-layer perceptron and / or a feedforward neural network and / or a convolutional neural network and / or a recurrent neural network.

[0023] The sound signal (and the in-domain data elements used in the training method) includes a noise component, which includes a combination of a target component and a noise component. The target component represents the portion of the sound signal generated by at least one desired sound source. The noise component represents the portion of the sound signal generated by at least one noise source or an unwanted signal portion such as a high-order reflection (late reverberation) originating from the desired sound source.

[0024] Examples of sound signals generated by at least one desired sound source include parts of speech signals, mixtures of speech signals, music, alarms, notification sounds, voice keywords, and tones generated by the desired sound source. The desired sound source may be a person or an output transducer.

[0025] Examples of sound signals generated by at least one noise source include portions of speech signals, mixtures of speech signals, music, alarms, tones, sound artifacts such as microphone noise and quantization noise, room reverberation, sound reflections, howling sounds due to feedback, transient sounds, wind noise, noise caused by microphone processing, echoes, and ambient sounds such as ambient noise generated by at least one noise source.

[0026] The sound signal is usually picked up using the hearing aid's input transducer, which is configured to measure the acoustic energy of the sound and convert it into an electrical signal. The electrical signal can also be generated by an auxiliary device and received by the hearing aid.

[0027] The sound signal can be represented in different time-related domains. For example, the sound signal can be represented in the time domain. For example, the sound signal can be represented in the frequency domain. For example, the sound signal can be represented in the time-frequency domain.

[0028] In an embodiment, the sound signal includes a noisy component including a desired component and a noise component, and extracting the desired component of the sound signal includes:

[0029] - amplifying the desired component while the power of the noise component remains the same or is reduced compared to the power of the noise component before extraction;

[0030] - amplifying the desired component and the noise component, wherein the amplification of the desired component is greater than the amplification of the noise component. This may be the case, for example, where the audibility of the noise component is still important;

[0031] - attenuating the noise component while the power of the desired component remains the same or increases compared to the power of the desired component before extraction;

[0032] - attenuating a noise component and a desired component, wherein the attenuation of the noise component is greater than the attenuation of the desired component;

[0033] - Increase the signal-to-noise ratio (SNR), which is defined as the ratio between the power of the desired component and the power of the noise component.

[0034] In an example, extracting the desired component includes spatial filtering (beamforming) of the noisy signal, ie, a directional microphone having a larger gain in the direction of the target sound source than at least one noise source.

[0035] The first algorithm comprises an encoder, wherein the encoder comprises at least one parameter, typically a plurality of parameters. The encoder also comprises a procedure for arithmetic and / or algebraic steps to compress the input to the encoder.

[0036] The input to the encoder is referred to as an input data element, which may be, for example, at least a portion of a masked in-field data element. The input data element may be pre-processed before being provided to the encoder, for example, by filtering, beamforming, and / or at least one mathematical transformation (e.g., absolute value, exponentiation, logarithm, scaling, addition / subtraction, modulus, trigonometric function).

[0037] In an example, the input data element includes a plurality of values. For example, each value may represent an amplitude or magnitude or phase of a portion of a sound signal. For example, each value may represent a discrete Fourier transform coefficient of a portion of a sound signal. The values ​​of the input data elements may be stacked into one or more input vectors. In one or more examples, the dimension of each input vector is equal to the number of values ​​of the input data elements.

[0038] The encoder generates an encoded data element. In a typical example, the encoded data element includes the same or fewer data bits than the input data element.

[0039] In an example, the encoded data element provided by the encoder includes a plurality of values. The values ​​of the encoded data element may be stacked into an encoded vector. The dimension of the encoded vector is equal to the number of the values ​​of the encoded data element.

[0040] In the example, the output of the encoder is a code vector, and the length (i.e., the dimension) of the code vector is equal to or shorter than the dimension of the input vector. For example, the dimension of the input vector is M in , the dimension of the encoding vector is M out , so that M out ≤M in .

[0041] In an example, the encoder is also a feature extractor. For example, the encoder can extract relevant features from its input and discard irrelevant features that are not helpful in extracting the desired components. For example, the relevant features can be characteristics of speech (e.g., speech harmonics, fundamental frequency, spectral shape) and / or characteristics of noise (e.g., harmonics, spectral shape, modulation).

[0042] In an example, the first algorithm includes a machine learning model such as an artificial neural network, wherein the encoder is a first part of the machine learning model.

[0043] In another example, the first algorithm includes multiple machine learning models, where the encoder is one of the machine learning models.

[0044] In this specification, a "data element" may be at least one time period of a sound signal including at least one component (e.g., a noisy component, a target component, a desired component, a noise component, a mixed component). In another example, a data element may include two time periods, wherein the first time period is a portion of the sound signal including the desired component, and the second time period is a portion of the sound signal including the mixed component.

[0045] For example, when the data element includes an amplitude representation of a time-frequency segment of a sound signal, the data element can be represented by a real-valued number. In another example, when the data element includes a time-frequency domain representation of a time-frequency segment of a sound signal, the data element can be represented by a complex-valued number.

[0046] As noted above, in-domain data elements are data elements based on sound signals obtained in the same or similar sound environment as the sound signals used by the trained first algorithm, e.g., having the same or similar noise type and / or the same or similar acoustic properties, such as the sound having the same or similar coloration due to the same or similar microphone type, microphone placement type (including the head and torso acoustics of the individual wearing the microphone, such as a hearing aid user), distance from a target talker, or reverberation pattern.

[0047] The sound environment may depend on the microphone, in particular, may depend on the acoustics of the microphone's location (eg, in a room or outdoors), the placement of the microphone on people and / or objects, the distance between the microphone and a target source, and / or the microphone type.

[0048] For example, the sound signal on which the in-domain data element is based may be collected by an input transducer of a hearing aid, which is configured to measure sound signals having certain acoustic properties, such as acoustic properties caused by the placement of the hearing aid on the user, the user's head and torso acoustics, the position of the input transducer on the hearing aid, the frequency response of the input transducer, and / or other factors that may affect the response of the sound signal picked up by the input transducer.

[0049] Thus, if the sound signal used by the trained first algorithm and the sound signal used to generate the in-domain data element were collected by a similar input transducer of a similar hearing aid with similar acoustic properties, then the data element may be considered an in-domain data element.

[0050] In the example, the in-domain data element is a portion of a sound signal including a noise component.

[0051] In an example, a comprehensive database including an in-domain database provides in-domain data elements for training an encoder of the first algorithm.

[0052] In this specification, a masked portion of at least one data element in a field is a portion of the data element in the field that is removed, hidden, or replaced with at least one predetermined value (preferably a value of zero).

[0053] In an example, a partially masked in-field data element is an in-field data element that includes at least one masked portion (eg, at least one portion is replaced with a zero value).

[0054] In the context of a training algorithm, the term "training" means a training procedure in which the values ​​of parameters of a first algorithm are adjusted, for example, to optimize (e.g., minimize) a loss (i.e., value / objective) function. The training procedure may include a first loss function. The first loss function is used to mathematically measure the performance of the first algorithm in solving the objective.

[0055] Training of the encoder comprises adjusting the value of at least one parameter (eg all parameters) of the encoder to optimize the prediction of the noisy component in at least a masked portion of at least a portion of the masked in-domain data elements, ie to optimize the prediction of what the noisy component in that portion was before masking.

[0056] In an example, the first loss function measures the difference between at least one predicted noisy component in at least one masked portion and a corresponding noisy component of at least one in-domain data element.

[0057] In an example, a plurality of in-domain data elements are used in a training method of a first algorithm. The in-domain data elements are typically acquired by an audio system including at least one input transducer that picks up at least a portion of a sound signal. At least a portion of the sound signal may be stored locally on the audio system and / or may be uploaded and stored on a server (e.g., in a comprehensive database) along with audio system information (e.g., location of the audio system, current time of the audio system, user information, device configuration). The server may generate the in-domain data elements from the stored sound signals. The audio system and the server may be configured to exchange data wirelessly or using a wired connection and / or through an auxiliary device, such as a smart phone, a computer, an electronic storage device serving as an intermediate communication interface.

[0058] The audio system typically comprises at least one hearing aid, thereby enabling obtaining in-domain data elements representing real usage conditions of the first algorithm, i.e. having characteristics representing, for example, user head filtering of the sound and / or processing of the sound by an input transducer of the hearing aid. The audio system may also be a hearing aid system, a headphone system, a loudspeaker system, a microphone system, or a video conferencing system.

[0059] In an example, training of the encoder further includes obtaining at least a portion of masked out-of-domain data elements including a mixture component including a target component and a noise component.

[0060] Training of the encoder further includes using at least a portion of the masked out-of-domain data elements to determine a value of at least one parameter of the encoder to further optimize the first algorithm's prediction of a target component in at least a masked portion of at least a portion of the masked out-of-domain data elements.

[0061] In an example, a plurality of masked out-of-domain data elements are used to train an encoder, and values ​​of a plurality of parameters are determined during the training of the encoder.

[0062] Out-of-domain data elements are data elements based on sound signals obtained in a sound environment that may be substantially different from the sound signals used by the trained first algorithm, such as having different noise types and / or different acoustic properties, such as the sound having different coloration due to different microphone types, microphone placement types, distances from the target talker, or reverberation patterns.

[0063] For example, the sound signal on which the out-of-domain data element is based may be collected by an input transducer of a smart phone having certain acoustic properties, such as the position of the input transducer on the smart phone, the frequency response of the input transducer, and other factors that may affect the response of the sound signal picked up by the input transducer.

[0064] In the example, the mixed component is simulated / generated by combining / mixing the noise component and the target component.

[0065] At least a portion of the masked out-of-range data elements are out-of-range data elements having a noisy component including at least a masked portion.

[0066] In an example, the out-of-domain data element is obtained from at least one public database including a plurality of sound signals.

[0067] In the example, the out-of-domain data element is based on at least two sound signals, and the at least two sound signals are obtained simultaneously. Specifically, the two sound signals can be obtained through two input transducers, the first input transducer is close to at least one target sound source, and the second input transducer is located at a place farther away from the target sound source, such as close to at least one noise sound source. The sound signal picked up by the first input transducer can be used to provide the target component, and the sound signal picked up by the second input transducer can be used to provide the noise component of the out-of-domain data element.

[0068] The sound signal obtained at the same time can be stored locally on the audio system including the first input transducer and the second input transducer, the audio system being, for example, a speaker system, a microphone system or a video conferencing system, and / or can be uploaded and stored on the server together with the audio system information. The audio system and the server can be configured to exchange data wirelessly or using a wired connection and / or through an auxiliary device, the auxiliary device being, for example, a smart phone, a computer, an electronic storage device serving as an intermediate communication interface.

[0069] In an example, at least a portion of the masked out-of-domain data elements is a partially masked spectrogram. In an example, at least a portion of the masked in-domain data elements is a partially masked spectrogram.

[0070] In this specification, the term "spectrogram" refers to a frequency representation of a sound signal over time. A spectrogram is an example of a sound signal represented in the time-frequency domain. In an example, the spectrogram includes the amplitude of the time-frequency domain representation of the sound signal. In other examples, the spectrogram can be obtained using a short-time Fourier transform (STFT) or other filter bank implementation. The time-frequency decomposition can be a linear decomposition, or the frequency decomposition can be a non-linear decomposition such as a logarithmic. In an example, the spectrogram includes the compressed amplitude (e.g., logarithmic value) of the time-frequency domain representation of the sound signal. In an example, the spectrogram is a complex-valued time-frequency representation of the sound signal.

[0071] A partially masked spectrum refers to a spectrum whose at least one portion or position is removed, or hidden, or replaced with at least one predetermined value (preferably zero). The portion of the spectrum can be a patch, such as a portion that can be represented by at least one single number (real or complex) value.

[0072] In an example, at least a portion of the masked out-of-domain data elements can be obtained by dividing a first out-of-domain data element including a mixed component including a target component and a noise component into a plurality of first patches and masking a predetermined percentage of the plurality of first patches to obtain at least a masked portion of at least a portion of the masked out-of-domain data elements.

[0073] In an example, at least a portion of the masked domain data elements can be obtained by dividing the domain data elements including at least a noisy component into a plurality of second patches and masking a predetermined percentage of the plurality of second patches to obtain at least a masked portion of the at least a portion of the masked domain data elements.

[0074] In an example, the predetermined percentage is a value between 5% and 50%. More specifically, the predetermined percentage may be a value between 20% and 40%. In an example, the predetermined percentage is 40%. In an example, the decision of which patches will be masked is randomized. The predetermined percentage may depend on the resolution across time and frequency. In an example, the minimum area of ​​a patch is given by a certain period of time multiplied by a certain range of frequencies (e.g., a frequency range corresponding to a 1 / 3 octave band).

[0075] In an example, the masked in-domain or out-of-domain patches are uniformly distributed across time and frequency of the spectrogram. In an example, the masked in-domain or out-of-domain patches are unevenly distributed across time and frequency of the spectrogram. In an example, the size of the masked in-domain or out-of-domain patches is a function of their center frequency.

[0076] In an example, the first algorithm includes a first decoder and a second decoder. The second decoder may also be trained using at least a portion of the masked out-of-domain data elements during training of the encoder, the encoder and the second decoder being trained to optimize the first algorithm's prediction of a target component in at least a masked portion of the at least a portion of the masked out-of-domain data elements.

[0077] The first decoder may also be trained using at least a portion of the masked in-domain data elements during training of the encoder, the encoder and first decoder being trained to optimize the first algorithm's prediction of noisy components in at least a masked portion of the at least a portion of the masked in-domain data elements.

[0078] The first decoder may also be trained using at least a portion of the masked out-of-domain data elements during training of the encoder, the encoder and first decoder being trained to optimize the first algorithm's prediction of a target component in at least a masked portion of the at least a portion of the masked out-of-domain data elements.

[0079] Each of the first decoder and the second decoder comprises a procedure of arithmetic and / or algebraic steps to decompress the input of the decoder, the decompressed input being called the decoded signal.

[0080] In an example, the decoded signal of the first decoder and / or the second decoder is a sound signal. In an example, the decoded signal of the second decoder is a predicted target component of at least one masked portion of at least one out-of-domain data element. In an example, the decoded signal of the first decoder is a predicted noisy component of at least a portion of the masked in-domain data element.

[0081] The input of the first decoder and / or the second decoder comprises coded data elements or processed coded data elements, wherein the processing comprises filtering, beamforming, and / or at least one mathematical transformation (e.g., absolute value, exponentiation, logarithm, scaling, addition / subtraction, modulus, trigonometric function).

[0082] In the example where the coded data element is a coded vector, each coded vector element in the coded vector represents the value of the coded data element. In this example, the output of a first decoder receiving the coded vector is a first decoded vector, and the output of a second decoder receiving the coded vector is a second decoded vector.

[0083] In an example, each of the first and second decoded vectors has more vector elements than the encoded vector, preferably as many elements as the input vector to the encoder.

[0084] In an example, the decoded signal of the first decoder and / or the second decoder is at least one target activity determination indicating the presence of the target component. For example, the target activity determination may include a value between "0" and "1", where the value "1" indicates that the target component is present in the input of the decoder, and the value "0" indicates that the target component is not present. The values ​​between "0" and "1" (excluding "0" and "1") may be interpreted as the probability of the presence of the target component. For example, a value of 0.6 indicates that there is a 60% chance that the target component is present in the input of the decoder.

[0085] In an example, the decoded signal is at least one noise-only determination indicating that only noise is present (i.e., no target component is present). For example, the noise-only determination may include a value between "0" and "1", where a value of "1" indicates that only noise is present at the input of the decoder, and a value of "0" indicates that only noise is not present (i.e., the target component is present). Values ​​between "0" and "1" (excluding "0" and "1") may be interpreted as probabilities that only noise is present. For example, a value of 0.6 indicates that there is a 60% chance that only noise is present at the input of the decoder.

[0086] The decoded signal typically includes more data bits than the input to the decoder or as many data bits as the input to the encoder.

[0087] In an example, during training of the encoder, each of the first and second decoders includes at least one artificial neural network.

[0088] In an example, during training of the encoder, the first algorithm comprises an artificial neural network having two parts, wherein the encoder is a first part of the artificial neural network and the first or second decoder is a second part of the artificial neural network.

[0089] In another example, during training the encoder, the first algorithm includes an artificial neural network having three parts, wherein the encoder is a first part of the artificial neural network, the first decoder is a second part of the artificial neural network, and the second decoder is a third part of the artificial neural network.

[0090] In another example, during training of the encoder, the first algorithm includes two artificial neural networks, wherein the encoder is a first artificial neural network and the first or second decoder is a second artificial neural network.

[0091] In another example, during training of the encoder, the first algorithm includes three artificial neural networks, wherein the encoder is a first artificial neural network, the first decoder is a second artificial neural network, and the second decoder is a third artificial neural network.

[0092] In an example, training the first algorithm includes discarding the first decoder and / or the second decoder after training the encoder.

[0093] In the example, when determining the value of at least one parameter of an encoder using at least a portion of a masked out-of-domain data element, the value of the at least one parameter is determined by minimizing a first loss function of the difference between a measured predicted target component and a corresponding target component of at least a portion of the masked out-of-domain data element (i.e., the target component before masking in at least a masked portion of the masked out-of-domain data element).

[0094] In the example, a first sub-portion of the first algorithm comprising an encoder and a first decoder aims at predicting a target component from at least a masked portion of at least a portion of masked out-of-domain data elements.

[0095] In another example, a second sub-portion of the first algorithm including an encoder and a second decoder aims to predict the noisy component from at least a masked portion of at least a portion of the masked in-domain data elements.

[0096] In the example, the second sub-part of the first algorithm comprising the encoder and the second decoder aims at predicting the mixing component from at least a masked portion of at least a portion of the masked out-of-domain data elements.

[0097] In the example, when determining the value of at least one parameter of an encoder using at least a portion of masked in-domain data elements, the value of the at least one parameter is determined by minimizing a second loss function that determines the difference between a predicted noisy component and a corresponding noisy component of at least a portion of the masked in-domain data elements (i.e., the noisy component before masking in at least a masked portion of the masked in-domain data elements).

[0098] In an example, the first loss function and / or the second loss function determines a difference. Determining the difference includes determining a magnitude loss by providing a logarithm of a sum of squared differences between an absolute value of a predicted target component and a corresponding target component of at least a portion of the masked out-of-domain data elements (for the first loss function) or between a predicted noisy component and a corresponding noisy component of at least a portion of the masked in-domain data elements (for the second loss function).

[0099] In this example, determining the amplitude loss includes calculating X n,f The first quantity (e.g., the corresponding noisy component or the corresponding target component) is denoted by The logarithm of the sum of the squared differences between the second quantity (e.g., the predicted noisy component or the predicted target component) and the predicted target component. The magnitude loss is given by:

[0100]

[0101] In the example, the variable n represents the time index and f represents the frequency index.

[0102] In an example, determining the amplitude loss includes calculating the logarithm of the sum of the exponential differences between the absolute values ​​of the first quantity and the second quantity, see the following formula:

[0103]

[0104] Here, β is an exponent which may be selected, for example, as β=1, β=2, β=1 / 2, and so on.

[0105] In the example, determining the difference also includes determining the phase loss by providing the logarithm of the sum of weighted squares (or weighted exponent β) of the difference between a normalized predicted target component and a normalized corresponding target component of at least a portion of the masked out-of-domain data elements (for a first loss function) or between a normalized predicted noisy component and a normalized corresponding noisy component of at least a portion of the masked in-domain data elements (for a second loss function).

[0106] In an example, determining the difference further comprises determining the magnitude-phase loss by providing a weighted sum between the magnitude loss and the phase loss, wherein the weighting comprises at least one weighting factor.

[0107] In an example, determining the phase loss includes calculating the logarithm of the sum of weighted squared differences between the absolute values ​​of the corresponding target component and the extracted / predicted target component (for the first loss function), weighted by the absolute value of the corresponding target component. The phase loss is given by:

[0108]

[0109] In the example, determining the magnitude-phase loss includes calculating the magnitude loss and the phase loss and linearly combining them using a weight factor λ. The magnitude-phase loss is given by:

[0110]

[0111] In other examples, measuring the difference may be performed by calculating a signal-to-noise ratio (SNR), a scale-invariant signal-to-noise ratio (SI-SNR), a cross-entropy between at least two quantities, and / or a relative entropy.

[0112] In an example, training the first algorithm further includes adding a third decoder to the first algorithm, the third decoder including at least one parameter, and training the third decoder using the trained encoder.

[0113] In an example, the third decoder is the second decoder and the second decoder is not discarded. In an example, at least one parameter of the third decoder is initialized by copying a parameter from the second decoder to the third decoder before discarding the second decoder.

[0114] Training the third decoder includes obtaining a second out-of-domain data element including a mixture of at least one target component and at least one noise component.

[0115] Training the third decoder includes determining a value of at least one parameter of the third decoder using at least one second out-of-domain data element. The value of at least one parameter of the third decoder is determined to optimize the first algorithm's prediction of the target component of the second out-of-domain data element.

[0116] The third decoder may have the same structure as the first decoder and / or the second decoder. In addition, the input to the third decoder may be a coded data element (or a processed version of the coded data element) provided by a trained encoder, having the same structure as the coded data element input to the first decoder and / or the second decoder. The output of the third decoder may be a decoded signal having the same structure as the decoded signal output by the first decoder and / or the second decoder.

[0117] In an example, when determining at least one parameter of a third decoder using at least one second out-of-domain data element, the at least one parameter is determined by minimizing a third loss function measuring a difference between a predicted target component and a corresponding target component of the at least one second out-of-domain data element.

[0118] In an example, the trained first algorithm includes an artificial neural network having two parts, wherein the trained encoder is the first part of the artificial neural network and the trained third decoder is the second part of the artificial neural network.

[0119] In another example, the trained first algorithm includes two artificial neural networks, wherein the trained encoder is a first artificial neural network and the trained third decoder is a second artificial neural network.

[0120] In the example, the first algorithm including the encoder and the third decoder aims at extracting the target component of the second out-of-domain data element.

[0121] In an example, the training method of the first algorithm is executed by a processor of at least one hearing aid and / or at least one computer.

[0122] Training method of the second algorithm

[0123] A second aspect of the present invention provides a training method for a second algorithm for extracting at least one target component of a sound signal.

[0124] The training method of the second algorithm includes extracting at least one target component from at least one second domain data element including the target component using the trained first algorithm.

[0125] The training method of the second algorithm also includes using at least one data element in the second domain and at least one target component extracted from the at least one data element in the second domain by the first algorithm to determine the value of at least one parameter of the second algorithm so as to optimize the prediction of at least one target component extracted from the at least one data element in the second domain.

[0126] In an example, the second algorithm includes at least one artificial neural network. In an example, the first algorithm and the second algorithm each include at least one artificial neural network.

[0127] In an example, the first algorithm is used to train the second algorithm by extracting a target component from at least one data element in the second domain. The extracted target component can then be used as a training target for training the second algorithm. In an example, the parameters of the first algorithm remain fixed and the second algorithm is trained during deployment.

[0128] In an embodiment, the second algorithm comprising an artificial neural network aims to predict a target component of a data element in the second domain.

[0129] The fourth loss function is used to mathematically measure the performance of the second algorithm in solving the objective. In an example, a scale-invariant signal-to-noise ratio (SI-SNR) is used as the fourth loss function for training the second algorithm.

[0130] In the context of training the first algorithm and / or the second algorithm, the term "training" includes a training procedure including an optimization procedure. For example, the optimization procedure includes a closed-form solution in which the partial derivatives of the first, second, third, and / or fourth loss functions with respect to at least one parameter of the first algorithm and / or the second algorithm are set to zero and solved for the at least one parameter.

[0131] In the example, the optimization procedure includes an iterative optimization algorithm. The iterative optimization algorithm includes the following steps:

[0132] - a first step comprises calculating the partial derivative of the first, second, third and / or fourth loss function with respect to at least one parameter of the first algorithm and / or the second algorithm; and

[0133] - The second step is to adjust or update at least one parameter. For example, the iterative optimization algorithm may include one of the following optimizers: stochastic gradient descent, Adam optimizer, AdaGrad optimizer, RMSprop optimizer.

[0134] In another example, the optimization procedure includes a combinatorial optimization algorithm such as a genetic algorithm, one advantage of which is that no partial derivatives of parameters of the first algorithm and / or the second algorithm need to be calculated.

[0135] In an example, training the first algorithm and / or the second algorithm includes using a comprehensive database of data elements in a plurality of domains. The comprehensive database can be split into at least one of the following data sets: a training set, a validation set, and a test set.

[0136] The training may be used to train parameters of the first algorithm and / or the second algorithm so that the first algorithm and / or the second algorithm have improved performance in solving their objectives. The validation set may be used to fine tune parameters of the first algorithm and / or the second algorithm and to evaluate the performance of the first algorithm and / or the second algorithm. The test set may be used to evaluate the performance of the first algorithm and / or the second algorithm for unseen in-domain data elements, i.e., in-domain data elements that were not used during training and validation.

[0137] In the example, the training set is divided into multiple batches, i.e., multiple subsets, so that the partial derivative calculation at each update step is based on one of the multiple batches.

[0138] In an example, training the first algorithm and / or the second algorithm includes at least one round, wherein one round refers to a complete traversal of the entire training set. During one round, the first algorithm and / or the second algorithm has seen and processed multiple (preferably all) domain data elements in the training set.

[0139] In an example, the batches are randomly shuffled for each round including the first round, so that the calculation of partial derivatives and subsequent updating of parameters of the first algorithm and / or the second algorithm do not follow a specific order of the batches.

[0140] In an example, the training set is used to train the first algorithm and / or the second algorithm, and the verification set is used to measure the performance of the first algorithm and / or the second algorithm during training. In an example, a mutual verification training scheme is used to train the algorithms, wherein both the training set and the verification set are used to adjust the parameters of the first algorithm and / or the second algorithm.

[0141] A third aspect of the present invention is to provide a method for extracting at least one desired component of a sound signal using a first algorithm, wherein the first algorithm is trained according to the training method of the first algorithm previously described, or to provide a method for extracting at least one sound target component of a sound signal using a second algorithm, wherein the second algorithm is trained according to the training method of the second algorithm previously described.

[0142] In an example, the first algorithm and / or the second algorithm is used to perform speech enhancement in a hearing aid by extracting a target component of a sound signal.

[0143] Hearing aid comprising the first algorithm or the second algorithm

[0144] A fourth aspect of the present invention is to provide a hearing aid comprising at least one input transducer, the at least one input transducer being configured to receive at least one first sound signal comprising a desired component and / or a noise component from an acoustic environment. The input transducer provides at least one electrical signal representing the at least one first sound signal.

[0145] The hearing aid comprises a first algorithm trained according to the training method of the first algorithm described above, or a second algorithm trained according to the training method of the second algorithm described above. The first algorithm or the second algorithm is configured to extract at least one desired component of at least one first sound signal.

[0146] The hearing aid includes at least one output transducer configured to output a second sound signal based on at least one extracted target component.

[0147] Hearing aids

[0148] An aspect of the invention is the use of the trained first and / or second algorithm in a hearing aid.

[0149] The hearing aid may be adapted to provide frequency dependent gain and / or level dependent compression and / or frequency transposition of one or more frequency ranges to one or more other frequency ranges (with or without frequency compression) to compensate for hearing impairment of the user. The hearing aid may include a signal processor for enhancing an input signal and providing a processed output signal. The signal processor may use the trained first or second algorithm to, for example, perform speech enhancement.

[0150] The hearing aid may comprise an output unit for providing a stimulus perceived by the user as an acoustic signal based on the processed electrical signal. The output unit may comprise a vibrator of a bone conduction hearing aid. The output unit may comprise an output transducer. The output transducer may comprise a receiver (speaker) for providing the stimulus as an acoustic signal to the user (e.g. in an acoustic (air conduction based) hearing aid). The output transducer may comprise a vibrator for providing the stimulus as a mechanical vibration of the skull to the user (e.g. in a bone attached or bone anchored hearing aid). The output unit may (in addition or as an alternative) comprise a (e.g. wireless) transmitter for transmitting the sound picked up by the hearing aid (e.g. via a network, e.g. in telephone operating mode, or in a headset configuration) to another device, such as a remote communication partner.

[0151] The hearing aid may comprise an input unit for providing an electrical input signal representing a sound. The input unit may comprise an input transducer such as a microphone for converting the input sound into the electrical input signal. The input unit may comprise a wireless receiver for receiving a wireless signal comprising or representing a sound and providing the electrical input signal representing the sound.

[0152] The wireless receiver and / or transmitter may be configured to receive and / or transmit electromagnetic signals in the radio frequency range (3kHz to 300GHz), for example. The wireless receiver and / or transmitter may be configured to receive and / or transmit electromagnetic signals in the optical frequency range (e.g., infrared light 300GHz to 430THz or visible light such as 430THz to 770THz), for example.

[0153] A hearing aid may include a directional microphone system adapted to spatially filter sounds from the environment so as to enhance a target sound source among a plurality of sound sources in the local environment of a user wearing the hearing aid. The directional system may be adapted to detect (e.g., adaptively detect) from which direction a particular portion of the microphone signal originates. This may be implemented in a variety of different ways, such as described in the prior art. In hearing aids, microphone array beamformers are typically used to spatially attenuate background noise sources. The beamformer may include a linear constrained minimum variance (LCMV) beamformer. Many beamformer variants can be found in the literature. Minimum variance distortionless response (MVDR) beamformers are widely used in microphone array signal processing. Ideally, the MVDR beamformer keeps the signal from the target direction (also called the line of sight) unchanged, while maximally attenuating sound signals from other directions. The generalized sidelobe canceler (GSC) structure is an equivalent representation of the MVDR beamformer, which provides computational and digital representation advantages over direct implementation in its original form.

[0154] Most sound signal sources (except the user's own voice) are relatively small compared to the size of the hearing aid, such as the distance d between the two microphones of a directional system. micLocated away from the user. The typical microphone distance in a hearing aid is on the order of 10 mm. The minimum distance to the user's sound source of interest (e.g., sound from the user's mouth or sound from an audio transmission device) is 0.1 m (>10 d mic ) level. For such a minimum distance, the hearing aid (microphone) will be in the acoustic near field of the sound source and the level differences of the sound signals incident on the respective microphones may be significant. The typical distance of the communication partner is greater than 1m (>100d mic ). The hearing aid (microphone) will be in the acoustic far field of the sound source, and the level difference of the sound signal incident on the corresponding microphone is not obvious. The arrival time difference of the sound incident in the direction of the microphone axis (for example, in front of or behind a normal hearing aid) is ΔT = d mic / v sound =0.01 / 343[s]=29μs, where v sound It is the speed of sound in air at 20°C (343 m / s).

[0155] The hearing aid may comprise an antenna and a transceiver circuit which enables a wireless link to be established to an entertainment device (e.g. a television), a communication device (e.g. a telephone), a wireless microphone, a separate (external) processing device, or another hearing aid, etc. The hearing aid may thus be configured to wirelessly receive a direct electrical input signal from another device. Similarly, the hearing aid may be configured to wirelessly transmit a direct electrical output signal to another device. The direct electrical input or output signal may represent or comprise an audio signal and / or a control signal and / or an information signal.

[0156] In general, the wireless link established by the antenna and the transceiver circuit of the hearing aid may be of any type. The wireless link may be a link based on near field communication, such as an inductive link based on inductive coupling between antenna coils of a transmitter part and a receiver part. The wireless link may be based on far field electromagnetic radiation. Preferably, the frequency used to establish the communication link between the hearing aid and the other device is lower than 70 GHz, such as in the range from 50 MHz to 70 GHz, such as higher than 300 MHz, such as in the ISM range higher than 300 MHz, such as in the 900 MHz range or in the 2.4 GHz range or in the 5.8 GHz range or in the 60 GHz range (ISM = Industrial, Scientific and Medical, such standardized ranges are defined, for example, by the International Telecommunication Union ITU). The wireless link may be based on standardized or dedicated technologies. The wireless link may be based on Bluetooth technology (such as low power Bluetooth technology, such as LE Audio) or ultra-wideband (UWB) technology.

[0157] The hearing aid may consist of or may form part of a portable (i.e. configured to be wearable) device, e.g. a device comprising a local energy source such as a battery, e.g. a rechargeable battery. The hearing aid may for example be a low weight, easily wearable device, e.g. having a total weight of less than 100 g, e.g. less than 20 g, e.g. less than 5 g.

[0158] The hearing aid may include a "forward" (or "signal") path between the input and output of the hearing aid for processing audio signals. A signal processor may be located in the forward path. The signal processor may be adapted to provide a frequency-dependent gain according to the specific needs of the user (e.g. hearing loss). The hearing aid may include an "analysis" path having functional parts for analyzing signals and / or controlling processing of the forward path. Part or all of the signal processing in the analysis path and / or the forward path may be performed in the frequency domain, in which case the hearing aid includes appropriate analysis and synthesis filter banks. Part or all of the signal processing in the analysis path and / or the forward path may be performed in the time domain.

[0159] The analog electrical signal representing the acoustic signal can be converted into a digital audio signal in an analog-to-digital (AD) conversion process, where the analog signal is sampled at a predetermined sampling frequency or sampling rate f. s Sampling, f s For example, in the range from 8 kHz to 48 kHz (adapted to the specific needs of the application) at discrete time points t n (or n) provides digital samples x n (or x[n]), each audio sample is passed through a predetermined N b The bit represents the sound signal at t n The value of N b For example, in the range from 1 to 48 bits, such as 24 bits. Each audio sample thus uses N b bit quantization (resulting in 2 Nb different possible values). A digital sample x has a 1 / f s The time length, such as 50μs, for f s = 20kHz. Multiple audio samples can be arranged in time frames. A time frame can include 64 or 128 audio data samples. Other frame lengths can be used depending on the actual application.

[0160] The hearing aid may include an analog-to-digital (AD) converter to digitize an analog input (e.g., from an input transducer such as a microphone) at a predetermined sampling rate, such as 20 kHz. The hearing aid may include a digital-to-analog (DA) converter to convert the digital signal into an analog output signal, such as for presentation to a user via an output transducer.

[0161] The hearing aid, such as the input unit and / or the antenna and transceiver circuitry, may comprise a transform unit for transforming a time domain signal into a signal in a transform domain (e.g. frequency domain or Laplace domain, Z transform, wavelet transform, etc.). The transform unit may consist of or include a time-frequency (TF) transform unit for providing a time-frequency representation of the input signal. The time-frequency representation may comprise an array or mapping of corresponding complex or real values ​​of the signal concerned in a specific time and frequency range. The TF transform unit may comprise a filter bank for filtering the (time-varying) input signal and providing a plurality of (time-varying) output signals, each output signal comprising a distinct frequency range of the input signal. The TF transform unit may comprise a Fourier transform unit (e.g. a discrete Fourier transform (DFT) algorithm, a short-time Fourier transform (STFT) algorithm, or a similar algorithm) for transforming the time-varying input signal into a (time-varying) signal in the (time-)frequency domain. The minimum frequency f is taken into account by the hearing aid. min To the maximum frequency f max The frequency range of may include a portion of the typical human hearing range from 20 Hz to 20 kHz, for example a portion of the range from 20 Hz to 12 kHz. Typically, the sampling rate f s Greater than or equal to the maximum frequency f max twice, that is, f s ≥2f max The signals of the forward path and / or analysis path of the hearing aid may be split into NI frequency bands (e.g. of uniform width), wherein NI is, for example, greater than 5, such as greater than 10, such as greater than 50, such as greater than 100, such as greater than 500, at least parts of which are processed separately. The hearing aid may be adapted to process the signals of the forward and / or analysis path in NP different frequency channels (NP≤NI). The frequency channels may be of uniform or non-uniform width (e.g. the width increases with frequency), overlapping or non-overlapping.

[0162] The hearing aid may be configured to operate in different modes, such as a normal mode and one or more specific modes, which may be selected by a user or may be automatically selected, for example. The operating mode may be optimized for a specific acoustic situation or environment, such as a communication mode, e.g., a telephone mode. The operating mode may include a low power mode, in which the functionality of the hearing aid is reduced (e.g., to save energy), such as disabling wireless communication and / or disabling specific features of the hearing aid.

[0163] The hearing aid may include a plurality of detectors configured to provide status signals related to the current network environment of the hearing aid (e.g., the current acoustic environment), and / or to the current state of a user wearing the hearing aid, and / or to the current state or operating mode of the hearing aid. Alternatively or additionally, one or more of the detectors may form part of an external device that communicates with the hearing aid (e.g., wirelessly). The external device may, for example, include another hearing aid, a remote control, an audio transmission device, a phone (e.g., a smartphone), an external sensor, etc.

[0164] One or more of the plurality of detectors may operate on a full-band signal (time domain). One or more of the plurality of detectors may operate on a band-split signal ((time-)frequency domain), for example in a limited number of frequency bands.

[0165] The plurality of detectors may include a level detector for estimating the current level of the signal of the forward path. The detector may be configured to determine whether the current level of the signal of the forward path is above or below a given (L-)threshold. The level detector acts on the full-band signal (time domain). The level detector acts on the band-split signal ((time-)frequency domain).

[0166] The hearing aid may comprise a voice activity detector (VAD) for estimating whether (or with what probability) an input signal (at a given point in time) comprises a voice signal. In the present specification, a voice signal may be taken to mean a signal comprising speech from a human being. It may also include other forms of vocalizations (such as singing) produced by a human speech system. The voice activity detector unit may be adapted to classify the user's current acoustic environment as a "voice" or a "no-voice" environment. This has the following advantage: time periods comprising an electric microphone signal of human vocalizations (such as speech) in the user's environment may be identified and thus separated from time periods comprising only (or mainly) other sound sources (such as artificially produced noise). The voice activity detector may be adapted to detect the user's own voice as "voice" as well. As an alternative, the voice activity detector may be adapted to exclude the user's own voice from the detection of "voice".

[0167] The hearing aid may include a target activity detector for determining whether (or with what probability) an input signal (at a given point in time) includes a target signal. In this specification, the target signal may include a speech signal from a human, an alarm, music, a notification sound, and a ringtone.

[0168] The hearing aid may include a self-voice detector for estimating whether (or with what probability) a particular input sound (e.g. voice, such as speech) originates from the voice of a user of the hearing system. The microphone system of the hearing aid may be adapted to be able to distinguish the user's own voice from the voice of another person and possibly from unvoiced sounds.

[0169] The plurality of detectors may include a motion detector, such as an accelerometer. The motion detector may be configured to detect movement of the user's facial muscles and / or bones, such as due to speech or chewing (eg, jaw movement), and provide a detector signal indicative of the movement.

[0170] The hearing aid may comprise a classification unit configured to classify the current situation based on an input signal from (at least part of) the detector and possibly other inputs. In this specification, a "current situation" may be defined by one or more of the following:

[0171] a) the physical environment (e.g. including the current electromagnetic environment, such as the presence of electromagnetic signals (including audio and / or control signals) intended or not intended to be received by the hearing aid, or other properties of the current environment other than acoustics);

[0172] b) Current acoustic conditions (input level, feedback, etc.);

[0173] c) the user’s current mode or state (motion, temperature, cognitive load, etc.);

[0174] d) The current mode or status of the hearing aid and / or another device communicating with the hearing aid (selected program, time elapsed since last user interaction, etc.).

[0175] The classification unit may be based on or may comprise a neural network, such as a recurrent neural network, such as a trained neural network.

[0176] Hearing aids may include acoustic (and / or mechanical) feedback control (e.g., suppression) or an echo cancellation system. Adaptive feedback cancellation has the ability to track changes in the feedback path over time. It is usually based on a linear time-invariant filter to estimate the feedback path, but the filter weights are updated over time. The filter updates can be calculated using a stochastic gradient algorithm, including some form of the least mean square (LMS) or normalized LMS (NLMS) algorithm. They all have the property of minimizing the difference signal in terms of mean squares, with NLMS additionally normalizing the filter updates to the square of the Euclidean norm of a reference signal.

[0177] The hearing aid may also include other appropriate functions for the application in question, such as compression, noise reduction, etc.

[0178] A hearing aid may comprise a hearing instrument, such as a hearing instrument adapted to be located at the ear of a user or to be located fully or partially in the ear canal, such as an earphone, a headset, an ear protection device or a combination thereof. A hearing system may comprise a loudspeaker amplifier (comprising a plurality of input transducers (such as a microphone array) and a plurality of output transducers such as one or more loudspeakers, and one or more audio (and possibly video) transmitters, such as for use in an audio conferencing situation), such as comprising a beamformer filter unit, such as to provide a plurality of beamforming capabilities.

[0179] application

[0180] In one aspect, there is provided an application of a hearing aid as described above, in detail in the "Detailed Description" section and in the claims. The application may be provided in a system including one or more hearing aids (such as hearing instruments), headphones, headsets, active ear protection systems, etc., such as a hands-free telephone system, a teleconferencing system (such as including a speaker amplifier), a broadcasting system, a karaoke system, a classroom amplification system, etc.

[0181] When appropriately replaced by corresponding processes, part or all of the structural features of the apparatus described above, described in detail in the "Detailed Description" or defined in the claims may be combined with the implementation of the method, and vice versa. The implementation of the method has the same advantages as the corresponding apparatus.

[0182] Computer readable medium or data carrier

[0183] The present invention further provides a tangible computer-readable medium (data carrier) storing a computer program including program code (instructions), which, when executed on a data processing system (computer), enables the data processing system to execute (implement) at least part (such as most or all) of the steps of the method described above, described in detail in the "Specific Implementation Method" and defined in the claims.

[0184] As an example but not limitation, the aforementioned tangible computer readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage device, or any other medium that can be used to execute or store the desired program code in the form of instructions or data structures and can be accessed by a computer. As used herein, disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and blue-ray disks, wherein these disks usually reproduce data magnetically, while these disks can reproduce data optically with lasers. Other storage media include storage in DNA (e.g., in synthetic DNA chains). Combinations of the above disks should also be included in the scope of computer readable media. In addition to being stored on tangible media, computer programs can also be transmitted via transmission media such as wired or wireless links or networks such as the Internet and loaded into a data processing system to run at a location different from the tangible media.

[0185] Computer Programs

[0186] In addition, the present application provides a computer program (product) comprising instructions, which, when executed by a computer, causes the computer to execute (the steps of) the method described above, described in detail in the “Detailed Description” and defined in the claims.

[0187] Data processing system

[0188] On the one hand, the present invention further provides a data processing system comprising a processor and a program code, wherein the program code enables the processor to perform at least part (such as most or all) of the steps of the method described above, described in detail in the "Specific Implementation Method" and defined in the claims.

[0189] Hearing System

[0190] In another aspect, there is provided a hearing aid comprising the above description, the detailed description in the "Detailed Description of the Invention" and the claims, and a hearing system comprising an auxiliary device.

[0191] The hearing system may be adapted to establish a communication link between the hearing aid and the auxiliary device so that information (eg control and status signals, possibly audio signals) can be exchanged or forwarded from one device to the other.

[0192] The auxiliary device may include or may consist of a remote control, a smart phone, or other portable or wearable electronic device such as a smart watch or the like.

[0193] The auxiliary device may consist of or comprise a remote control for controlling the functions and operation of the hearing aid. The functions of the remote control are implemented in a smartphone, which may run an APP enabling control of the functions of the audio processing device via the smartphone (the hearing aid comprises a suitable wireless interface to the smartphone, e.g. based on Bluetooth or some other standardized or proprietary solution).

[0194] The auxiliary device may be constituted by or include an audio gateway device, which is suitable for receiving multiple audio signals (e.g. from an entertainment device such as a TV or music player, from a telephone device such as a mobile phone, or from a computer such as a PC, a wireless microphone, etc.) and is suitable for selecting and / or combining appropriate signals (or signal combinations) from the received audio signals for transmission to the hearing aid.

[0195] The auxiliary device may consist of or may comprise a further hearing aid.The hearing system may comprise two hearing aids adapted to implement a binaural hearing system, eg a binaural hearing aid system.

[0196] APP

[0197] On the other hand, the present invention also provides a non-transient application called APP. The APP includes executable instructions configured to run on an auxiliary device to implement a user interface for the hearing aid or hearing system described above, described in detail in the "Detailed Description of the Invention" and defined in the claims. The APP can be configured to run on a mobile phone such as a smart phone or another portable device that enables communication with the hearing aid or hearing system.

[0198] definition

[0199] In this specification, a hearing aid, such as a hearing instrument, refers to a device suitable for improving, enhancing and / or protecting the hearing ability of a user by receiving an acoustic signal from the user's environment, generating a corresponding audio signal, possibly modifying the audio signal, and providing the possibly modified audio signal as an audible signal to at least one ear of the user. The audible signal may be provided, for example, in the form of an acoustic signal radiated into the user's outer ear and / or an acoustic signal transmitted to the user's inner ear as mechanical vibrations through the bone structure of the user's head and / or through parts of the middle ear.

[0200] The hearing aid may be configured to be worn in any known manner, such as as a unit worn behind the ear (with a tube directing the radiated acoustic signal into the ear canal or with an output transducer such as a loudspeaker arranged close to or in the ear canal), as a unit arranged wholly or partly in the auricle and / or the ear canal, as a unit connected to a fixed structure implanted in the skull such as a vibrator, etc. The hearing aid may comprise a single unit or several units communicating with each other (e.g. acoustically, electrically or optically). The loudspeaker may be arranged in a housing together with the other components of the hearing aid, or it may itself be an external unit (possibly in combination with a flexible guiding element such as a dome-shaped element).

[0201] The hearing aid may be adapted to the needs of a specific user, such as hearing loss. The configurable signal processing circuit of the hearing aid may be adapted to apply frequency- and level-dependent compression amplification of the input signal. The customized frequency- and level-dependent gain (amplification or compression) may be determined during the fitting process by the fitting system based on the user's hearing data, such as an audiogram, using the basic principles of fitting (e.g. adaptation to speech). The frequency- and level-dependent gain may, for example, be embodied in a processing parameter, uploaded to the hearing aid, for example, via an interface to a programming device (fitting system), and used by a processing algorithm executed by the configurable signal processing circuit of the hearing aid.

[0202] "Hearing system" refers to a system including one or two hearing aids. "Binaural hearing system" refers to a system including two hearing aids and adapted to provide audible signals to the two ears of a user in a coordinated manner. A hearing system or binaural hearing system may also include one or more "auxiliary devices" that communicate with the hearing aids and affect and / or benefit from the functions of the hearing aids. The aforementioned auxiliary devices may include at least one of the following: a remote controller, a remote microphone, an audio gateway device, an entertainment device such as a music player, a wireless communication device such as a mobile phone (e.g., a smart phone) or a tablet computer or another device, for example, including a graphical interface. Hearing aids, hearing systems or binaural hearing systems may be used, for example, to compensate for the loss of hearing ability of hearing-impaired persons, enhance or protect the hearing ability of normal hearing persons, and / or transmit electronic audio signals to people. Hearing aids or hearing systems may, for example, form part of or interact with a broadcasting system, an active ear protection system, a hands-free telephone system, a car audio system, an entertainment (e.g., television, music playback or karaoke) system, a teleconferencing system, a classroom amplification system, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0203] Various aspects of the present invention will be best understood from the detailed description below in conjunction with the accompanying drawings. For clarity, the drawings are schematic and simplified, and only the details necessary for understanding the present invention are given, while other details are omitted. Throughout the specification, the same reference numerals are used for the same or corresponding parts. The various features of each aspect may be combined with any or all features of the other aspects. These and other aspects, features and / or technical effects will be apparent from and illustrated in conjunction with the following figures, in which:

[0204] Figure 1 is a schematic representation of a hearing aid configured to use a first or second algorithm trained according to an example of the present invention;

[0205] Figure 2A is a schematic representation of units and data for training an encoder of a first algorithm according to an example of the present invention;

[0206] Figure 2B A block diagram of a masking unit for masking in-domain or out-of-domain data elements according to an example of the present invention is shown;

[0207] Figure 2C A flowchart showing steps for training an encoder of a first algorithm using in-domain data elements according to an example of the present invention;

[0208] Figure 3A A schematic representation of units and data of an encoder and a first decoder for training a first algorithm using in-domain or out-of-domain data elements according to an example of the present invention;

[0209] Figure 3Bis a schematic representation of units and data of an encoder and a second decoder for training a first algorithm using out-of-domain data elements according to an example of the present invention;

[0210] Figure 3C is a schematic representation of units and data of a third decoder for training a first algorithm using out-of-domain data elements according to an example of the present invention;

[0211] Figure 4 A schematic representation of units and data for training a first algorithm using a database of in-domain data elements and a database of out-of-domain data elements according to an example of the present invention;

[0212] Figure 5 A flowchart showing steps for training a first algorithm using in-domain and out-of-domain data elements according to an example of the present invention;

[0213] Figure 6 is a schematic representation of units and data for training a second algorithm using a trained first algorithm and in-domain data elements according to an example of the present invention;

[0214] Figure 7 A flow chart showing steps for training a second algorithm using a trained first algorithm and in-domain data elements according to an example of the present invention is shown.

[0215] By the detailed description given below, the further scope of application of the present invention will be apparent. However, it should be understood that while the detailed description and specific examples show the preferred embodiments of the present invention, they are only provided for illustrative purposes. For those skilled in the art, based on the following detailed description, other embodiments of the present invention will be apparent. DETAILED DESCRIPTION

[0216] The detailed description proposed below in conjunction with the accompanying drawings serves as a description of a variety of different configurations. The detailed description includes specific details for providing a thorough understanding of a number of different concepts. However, it is apparent to those skilled in the art that these concepts can be implemented without these specific details. Several aspects of the apparatus and method are described by a number of different blocks, functional units, modules, components, circuits, steps, processes, algorithms, etc. (collectively referred to as "elements"). Depending on the specific application, design limitations or other reasons, these elements can be implemented using electronic hardware, computer programs or any combination thereof.

[0217] The electronic hardware may include microelectromechanical systems (MEMS), (e.g., application specific) integrated circuits, microprocessors, microcontrollers, digital signal processors (DSPs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), gating logic, discrete hardware circuits, printed circuit boards (PCBs) (e.g., flexible PCBs), and other suitable hardware configured to perform a number of different functions described in this specification, such as sensors for sensing and / or recording physical properties of the environment, device, user, etc. Computer programs shall be broadly construed as instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, programs, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0218] Hearing aid configured to use a first or second algorithm

[0219] Figure 1 An example of a hearing aid HA configured to use a first algorithm ALG1 or a second algorithm ALG2 trained according to the present invention is shown. In this example, the hearing aid HA is a BTE (i.e. "behind the ear") type hearing aid. However, other types of hearing aids with similar functional elements, such as RITE (i.e. "receiver in the ear") type hearing aids, may also be used, and the corresponding description is similar. The hearing aid HA according to the present invention comprises a BTE part having a loudspeaker and an ITE (i.e. "in the ear") part connected by an acoustic propagation element. The BTE part is suitable for being located at or behind the ear of a user, and the ITE part is suitable for being located in or at the ear canal of a user.

[0220] The BTE part and the ITE part are connected (eg electrically connected) via a connection element IC including an acoustic propagation channel, such as a hollow tube.

[0221] The BTE section includes first and second input transducers such as microphones (M 1 and M 2 ), which is used to pick up sound from the environment of the user wearing the hearing aid (see sound field S IN ). Each input transducer generates an electrical signal representative of the sound from the environment picked up at the location of each respective input transducer.

[0222] The BTE part also includes a signal processor DSP configured to receive the input converter M 1 ,M 2An electrical signal is received. The signal processor DSP is adapted to provide a frequency dependent gain and / or a level dependent compression and / or a frequency transposition of one or more frequency ranges to one or more other frequency ranges (with or without frequency compression) in order to, for example, compensate for a hearing impairment of a user. The signal processor DSP further comprises a trained first algorithm ALG1 or a trained second algorithm ALG2 configured to extract a target component.

[0223] The hearing aid HA (here the BTE part) may also comprise two (eg individually selectable) wireless receivers WLR 1 ,WLR 2 , for providing a corresponding directly received auxiliary audio input and / or control or information signal. The wireless receiver may be configured to receive signals from another hearing device (e.g. a binaural hearing system) or from any other communication device such as a telephone, e.g. a smartphone, or from a wireless microphone or a T-coil. The wireless receiver may be able to receive (and possibly also transmit) audio and / or control or information signals. The wireless receiver may be based on Bluetooth or similar technology, or may be based on near field communication (e.g. inductive coupling).

[0224] The BTE part comprises a substrate SUB on which a plurality of electronic components such as a memory MEM and a signal processor DSP are mounted. The signal processor DSP forms part of an integrated circuit, such as a (mainly) digital integrated circuit. The memory MEM is configured to store at least a time segment of an electrical signal received from an input transducer, the stored time segment being referred to as an intra-domain data element.

[0225] The BTE portion includes an output transducer SP, which provides an enhanced output signal as a stimulus that can be perceived as sound by the user based on the enhanced (e.g., amplified, frequency-shaped, extracted target component) audio signal from the signal processor DSP or a signal derived therefrom. Figure 1 In the example of FIG. 1 , the output transducer is in the form of a loudspeaker (receiver) SP for converting an electrical signal into an acoustic signal. The loudspeaker SP is configured to play sound to the connected element. The loudspeaker is connected to the BTE via internal wiring in the BTE part (see, for example, the schematic diagram of the BTE part as W). x The hearing device is connected to the relevant electronic circuits of the hearing device, for example to a signal processor DSP.

[0226] Alternatively or additionally, the enhanced audio signal from the signal processor DSP may be further processed and / or passed to another device depending on the specific application.

[0227] The BTE part may also comprise a battery BAT, such as a rechargeable battery, for powering the electronic components of the BTE part and possibly the ITE part, if any.

[0228] The ITE part comprises an ear mold and is used to enable a considerable sound pressure level to be transmitted to the eardrum of a user (e.g. a user with severe to profound hearing loss). In addition, the ITE part comprises a through opening so that the sound can be transmitted to the eardrum of the user via the connecting element (see sound field S OUT ).

[0229] Method for training the first algorithm

[0230] Figure 2A An exemplary block diagram of units and data for training an encoder ENC of a first algorithm ALG1 using an in-domain data element ID is shown. Training the encoder ENC of the first algorithm ALG1 comprises determining the parameters of the encoder

[0231] The first algorithm ALG1 receives at least a portion of masked intra-domain data elements PM-ID. The at least a portion of masked intra-domain data elements PM-ID is provided by a masking unit MU. The masking unit MU is configured to receive at least one intra-domain data element ID including a noise component and thereafter use a masking procedure, for example in combination with Figure 2B The described masking procedure masks a data element ID within at least one field.

[0232] The encoder ENC receives at least a portion of the masked intra-domain data element PM-ID and generates at least one encoded data element. The at least one encoded data element is received by a pre-trained decoder configured to generate at least one prediction PRED-NSY of at least a masked portion of at least a portion of the masked intra-domain data element PM-ID.

[0233] Figure 2B An exemplary block diagram of a masking unit MU for masking at least one in-domain data element ID or at least one out-of-domain data element OOD is shown. The masking unit MU is configured to receive at least one in-domain data element ID or at least one out-of-domain data element OOD, wherein the at least one in-domain data element ID includes a noise component, and the at least one out-of-domain data element OOD includes a mixed component including a target component and a noise component. In this example, the at least one in-domain data element ID or the at least one out-of-domain data element OOD includes at least one spectrogram.

[0234] The masking unit is configured to use a masking procedure, wherein the masking procedure includes randomly selecting a plurality of patches of a plurality of different sizes of the data element ID within at least one field, and then replacing the value within each patch with at least one predetermined value, such as zero.

[0235] If provided with at least one in-domain data element ID, the masking unit MU generates at least a portion of masked in-domain data elements PM-ID; or, if provided with at least one out-of-domain data element OOD, the masking unit MU generates at least a portion of masked out-of-domain data elements PM-OOD.

[0236] The partially masked intra-domain data element includes at least one masked portion, and typically includes a plurality of masked portions, in which noise components are masked.

[0237] Similarly, a partially masked out-of-domain data element includes at least one masked portion, and typically a plurality of masked portions, wherein a mixed component is masked.

[0238] Figure 2C An exemplary flow chart of steps for training an encoder ENC of a first algorithm ALG1 using at least one intra-domain data element ID is shown. The training starts with a start step 200, in which the parameters of the encoder Initial values ​​for the parameters of the encoder may be selected by sampling values ​​from a probability distribution, setting these values ​​to predetermined values ​​such as zero, or may be copied from a pre-trained encoder such as transfer learning.

[0239] In step 210, the first algorithm ALG1 receives at least a portion of the masked intra-domain data elements PM-ID. At least a portion of the masked intra-domain data elements PM-ID are received by the encoder in the first algorithm ALG1. At least a portion of the masked intra-domain data elements PM-ID are propagated through the first algorithm ALG1, which includes propagating through the encoder ENC. The first algorithm generates at least one prediction PRED-NSY of the masked part of the at least a portion of the masked intra-domain data elements PM-ID, i.e., a prediction of the noisy element part masked by the masked part. The parameters of the encoder are determined by updating these parameters in step 220 using an iterative optimizer such as an Adam optimizer using at least one prediction PRED-NSY of the masked part and at least one intra-domain data element ID. To calculate the gradient or partial derivative of each parameter, a loss function such as an amplitude-phase loss function is used. The amplitude-phase loss function is given by the following formula:

[0240]

[0241] in

[0242]

[0243] Among them, Y ID Indicates at least one data element ID in the domain. represents at least one predicted intra-domain data element PRED-NSY. In this embodiment, n represents a time index and f represents a frequency window index of at least one intra-domain data element including the spectrogram. The amplitude-phase loss function is a linear combination of the amplitude loss function and the phase loss function. The weight factor λ is used to control the significance of the phase. The training of the first algorithm ends at the end step 230 after the predetermined stop condition is met.

[0244] In the following Figure 3A , Figure 3B , Figure 3C In the description of the illustrated embodiment, for the sake of clarity, we use the following notation. The out-of-domain data element OOD includes a target component denoted by S, a noise component denoted by V, and a OOD The out-of-domain data element OOD is characterized by the fact that all three elements can be observed in isolation. The mixed component can be synthetically simulated or generated as a linear combination of the target component and the noise component, for example:

[0245] Y OOD =S+V

[0246] The target component and / or the noise component may be preprocessed before the linear combination, such as filtering, amplification / attenuation, frequency shaping, and beamforming. Figure 3A , Figure 3B , Figure 3C In an exemplary embodiment of the present invention, the target component, the noise component and the mixed component are all spectrograms in the time-frequency domain having a time domain.

[0247] The domain database of multiple domain data element IDs is defined as a set of domain data element IDs, denoted as Here, express The noisy component of the data element in the nth domain, n = 1, 2, ..., N ID . N ID Indicates the number of data element IDs in the domain database.

[0248] The external database of multiple external data elements is defined as a set / collection of external data elements, denoted as Here, express The mixed component of the nth out-of-domain data element, S (n) express The corresponding target component of the nth out-of-domain data element, n = 1, 2, ..., N OOD . N OOD Indicates the number of foreign data elements in the foreign database. In addition, for clear distinction, only foreign data elements are mixed. yes A subset of Defined as containing only The elements of the mixture, excluding S (n) ,b=1,2,…,N OOD .

[0249] Training the first algorithm ALG1 includes two training phases. The first training phase combines Figure 3A and Figure 3B Describe, the second training stage combines Figure 3C Give a description.

[0250] Figure 3A An exemplary block diagram of units and data of an encoder ENC and a first decoder DEC1 for training a first algorithm ALG1 using at least one in-domain data element ID or at least one out-of-domain data element OOD is shown. The masking unit MU receives at least one in-domain data element ID or at least one out-of-domain data element OOD and generates at least a portion of masked in-domain data elements PM-ID or at least a portion of masked out-of-domain data elements PM-OOD accordingly. The masking unit MU generates a binary mask (or a soft mask, for example, where the value of the mask is between "0" and "1", including "0" and "1"), denoted as M, and applies it to at least one in-domain data element ID or at least one out-of-domain data element OOD to generate at least a portion of masked in-domain data elements PM-ID or at least a portion of masked out-of-domain data elements PM-OOD accordingly. For at least one out-of-domain data element OOD, this is given by the following formula:

[0251]

[0252] in, PM-OOD refers to at least a portion of the masked out-of-domain data elements, ⊙ is the element-by-element multiplication between the two matrices (i.e. and M are represented by matrices of equal size). For at least one data element in the field, this is given by:

[0253]

[0254] in, Refers to a data element PM-ID within a field that is at least partially masked.

[0255] exist Figure 3A In the example of , the first algorithm comprises an encoder ENC, a first decoder DEC1 and a second decoder DEC2.

[0256] The parameter set of the encoder is denoted as The parameter set of the first decoder is denoted as The parameter set of the second decoder is denoted as The encoder transformation is recorded as The transformation of the first decoder is denoted as The transformation of the second decoder is recorded as The first algorithm ALG1 comprises two function combinations, wherein the first function combination comprises an encoder ENC and a first decoder DEC1. The first function combination generates at least one predicted mixed component PRED-NSY from at least a portion of masked out-of-domain data elements PM-OOD or at least a portion of masked in-domain data elements PM-ID. For partially masked out-of-domain data elements PM-OOD, it is given by the following formula:

[0257]

[0258] in, is the mixed component PRED-NSY predicted from the partially masked out-of-domain data element PM-OOD by the first functional combination of the first algorithm ALG1, In addition, the first function combination generates at least one predicted noisy component PRED-NSY from at least a portion of the masked in-domain data elements PM-ID, that is,

[0259]

[0260] in, is the noisy component PRED-NSY predicted from the partially masked intra-domain data element PM-ID by the first functional combination of the first algorithm ALG1.

[0261] exist Figure 3A In FIG. 1 , it is shown that the parameters of the encoder ENC and the first decoder DEC1 can be determined using partially masked in-domain data elements PM-ID and partially masked out-of-domain data elements PM-OOD. For training the encoder and the first decoder using at least a partially masked in-domain data elements PM-ID, the amplitude-phase loss function is used to measure the difference between at least one predicted noisy component PRED-NSY and the noisy component of at least one in-domain data element ID. This is given by:

[0262]

[0263] Similarly, for training the encoder ENC and the first decoder DEC1 using at least partially masked out-of-domain data elements PM-OOD, the amplitude-phase loss function is used to measure the difference between at least one predicted mixing component PRED-NSY and the mixing component of at least one out-of-domain data element OOD.

[0264]

[0265] The full noisy amplitude-phase loss function can be described by the following equation:

[0266]

[0267] Here, ∪ refers to the union of two sets, so A database of data elements in multiple domains Combined with multiple databases containing only data elements outside the noise domain. Fully noisy magnitude-phase loss function It can also be called the Masked Noise Spectrogram Prediction (MSP) loss.

[0268] exist Figure 3B In FIG. 1 , a second functional combination including an encoder ENC and a second decoder DEC2 is shown for generating at least one predicted target component PRED-TRG from at least a portion of the masked out-of-domain data elements PM-OOD, namely

[0269]

[0270] in, is the target component PRED-TRG predicted from the partially masked out-of-domain data elements PM-OOD by the second functional combination of the first algorithm ALG1.

[0271] exist Figure 3B In FIG. 1 , it is shown that the parameters of the encoder ENC and the second decoder DEC2 can be determined using at least a portion of masked out-of-domain data elements PM-OOD. For training the encoder ENC and the second decoder DEC2 using at least a portion of masked out-of-domain data elements PM-OOD, the amplitude-phase loss function is used to measure the difference between at least one predicted target component PRED-TRG and the target component of at least one out-of-domain data element OOD. This is given by:

[0272]

[0273] The full target amplitude-phase loss function can be described by the following formula:

[0274]

[0275] The parameters of the encoder ENC, the first decoder DEC1 and the second decoder DEC2 are updated jointly for multiple in-domain data elements ID and out-of-domain data elements OOD. Therefore, the full amplitude-phase loss function can be described by the following formula:

[0276]

[0277] The first stage of training the encoder ENC, the first decoder DEC1 and the second decoder DEC2 thus solves the following optimization problem overall:

[0278]

[0279] in, refer to The independent variables of the minimizer function (i.e. the parameters of the encoder ENC, the first decoder DEC1 and the second decoder DEC2), Refers to Θ (pre) is a combined parameter set of the encoder ENC, the first decoder DEC1 and the second decoder DEC2. To solve the optimization problem, an iterative optimization algorithm is used. When the parameters of the encoder ENC, the first decoder DEC1 and the second decoder DEC2 have been determined using the iterative optimization algorithm, the first stage of training the first algorithm ALG1 ends.

[0280] exist Figure 3C , an exemplary block diagram of a second training phase of the first algorithm ALG1 is shown. The second training phase comprises training the first algorithm ALG1 using at least one out-of-domain data element OOD. In the second training phase, the first and second decoders DEC1, DEC2 from the first training phase are discarded. The first algorithm ALG1 comprises a trained encoder ENC, wherein the parameters of the trained encoder ENC are obtained from the first phase. The first algorithm ALG1 further comprises a third decoder DEC3. The encoder ENC and the third decoder DEC3 are configured to extract a target component from a mixed component of at least a portion of the masked out-of-domain data element PM-OOD.

[0281] The parameters of the third decoder DEC3 are determined using another loss function. The loss function includes the scale-invariant signal-to-noise ratio SI-SNR. The loss function for the second stage is given by:

[0282]

[0283] The SI-SNR between the extracted target component and the corresponding target component of at least a portion of the masked out-of-domain data element PM-OOD is given by:

[0284]

[0285] in,

[0286]

[0287] After training the first algorithm, the parameters of the encoder ENC and the third decoder DEC3 may then be fixed.The first algorithm ALG1 may then be used to extract target components from a sound signal picked up, for example, by an input transducer of a hearing aid.

[0288] Figure 4An exemplary overview of data and units for training a first algorithm using a database ID-DB of in-domain data elements and a database OOD-DB of out-of-domain data elements is shown.

[0289] The training of the first algorithm ALG1 includes a first training phase ST1 and a second training phase ST2 , wherein the first training phase ST1 is performed before the second training phase ST2 .

[0290] The intra-domain database ID-DB including a plurality of intra-domain data element IDs is used to train the first algorithm. Specifically, the intra-domain database ID-DB provides intra-domain data element IDs used to train the first algorithm in the first training phase.

[0291] Furthermore, an out-of-domain database OOD-DB including a plurality of out-of-domain data elements OOD is used for training the first algorithm ALG1.

[0292] In the first training stage ST1, the pre-trained model PTM of the first algorithm ALG1 uses the masked noise spectrogram to predict the MSP loss function LOSS1 by Figure 3A and Figure 3B The pre-trained model PTM includes an encoder ENC. After the first training stage ST1, the trained encoder ENC from the pre-trained model PTM is used in the second training stage ST2. The parameters of the trained encoder ENC are fixed in the second training stage ST2. In the second training stage ST2, the fine-tuned model FTM of the first algorithm is obtained using a speech enhancement loss function SE, such as an SI-SNR loss function LOSS2.

[0293] Figure 5 Shown is the Figure 3A , Figure 3B and Figure 3C An exemplary flow chart of the steps of training a first algorithm ALG1 using at least one in-domain data element ID and at least one out-of-domain data element OOD. The training starts at a start step 500, where the parameters of the encoder ENC, the first decoder DEC1, the second decoder DEC2 and the third decoder DEC3 are initialized. The initial values ​​of the parameters of the encoder ENC, the first decoder DEC1, the second decoder DEC2 and the third decoder DEC3 can be selected by sampling values ​​from a probability distribution, setting these values ​​to predetermined values, such as zero, or can be copied from a pre-trained encoder, such as transfer learning.

[0294] like Figure 5 As shown in FIG, training the encoder ENC and the decoders DEC1, DEC2, DEC3 can be divided into a first training phase 560 and a second training phase 565. In the first training phase 560, the encoder ENC, the first decoder DEC1 and the second decoder DEC2 are trained.

[0295] At step 501, at least one in-domain data element ID and / or at least one out-domain data element OOD is received from an in-domain database or an out-domain database. The at least one in-domain data element ID and / or at least one out-domain data element OOD is then masked by a masking unit in step 505 to provide at least a portion of masked in-domain data element PM-ID and / or at least a portion of masked out-domain data element PM-OOD.

[0296] In step 510, a conditional statement is used to determine whether at least one provided data element for the first algorithm ALG1 contains a partially masked intra-domain data element PM-ID. If "yes", i.e. the provided data element contains a partially masked intra-domain data element PM-ID, then in step 515, at least one parameter of the encoder ENC and the first decoder DEC1 is updated using at least a portion of the masked intra-domain data element (and the extra-domain data element, if provided).

[0297] If at least one of the provided data elements does not include any partially masked intra-domain data element PM-ID, the training proceeds to step 520, where a conditional statement determines whether the first decoder DEC1 should be updated. This decision is predetermined. If the conditional statement returns "yes", then at step 515, at least one parameter of the encoder ENC and the first decoder DEC1 is updated using at least one out-of-domain data element OOD.

[0298] If "no", at least one parameter of the encoder ENC and the second decoder DEC2 is updated using the out-of-domain data element OOD in step 525. After updating the parameters in step 515 or 525, in step 530, a conditional statement is used to decide whether the first training stage 560 is ended. This decision can be made based on a predetermined stop condition.

[0299] In the second training phase 565, the parameters of the encoder ENC are kept fixed, the first and second decoders DEC1, DEC2 are discarded, and the third decoder DEC3 is trained. In the first step 540 of the second training phase, at least one out-of-domain data element is received from the out-of-domain database. In step 545, at least one parameter of the third decoder DEC3 is updated using the at least one out-of-domain data element OOD.

[0300] In step 550, a conditional statement is used to determine whether the training of the second training stage 565 is finished. If the training of the second training stage is not finished, the process returns to step 540. If the training of the second training stage 565 is finished, in step 555, the parameters of the first algorithm ALG1 are determined and fixed.

[0301] Training method of the second algorithm

[0302] Figure 6 An exemplary block diagram of units and data for training a second algorithm ALG2 using a trained first algorithm ALG1 and at least one intra-domain data element ID is shown. The first algorithm ALG1 comprises a trained encoder ENC and a trained third decoder DEC3.

[0303] The first algorithm ALG1 receives at least one intra-domain data element ID to generate a predicted target component PRED-TRG. The second algorithm ALG2 receives the target component PRED-TRG predicted by the first algorithm ALG1 and at least one intra-domain data element ID.

[0304] The second algorithm ALG2 comprises a second artificial neural network configured to extract a target component PRED-TRG2 given at least one intra-domain data element ID. The second algorithm ALG2 uses the target component PRED-TRG predicted by the first algorithm ALG1 and at least one intra-domain data element ID to train the second algorithm ALG2. The second algorithm ALG2 is trained using a scale-invariant signal-to-noise ratio SI-SNR loss function.

[0305] Figure 7 An exemplary flow chart of steps for training a second algorithm ALG2 using a trained first algorithm ALG1 and at least one intra-domain data element ID is shown. The training of the second algorithm ALG2 begins at a start step 700, where parameters of the second algorithm ALG2 are initialized. The initial values ​​of the parameters of the second algorithm ALG2 may be selected by sampling values ​​from a probability distribution, setting these values ​​to predetermined values, such as zero, or may be copied from a pre-trained encoder, such as transfer learning.

[0306] At least one intra-domain data element ID is received from the intra-domain database at step 710. The trained first algorithm ALG1 receives the at least one intra-domain data element ID and provides the extracted target component at step 720. At step 730, the second algorithm ALG2 receives the at least one intra-domain data element ID and the extracted / predicted target component PRED-TRG from the trained first algorithm ALG1 and uses them to update at least one parameter of the second algorithm ALG2.

[0307] In step 740, a predetermined conditional statement is used to determine whether the training of the second algorithm ALG2 is finished. If "no", step 710 is repeated by receiving at least one new intra-domain data element ID. If "yes", the training of the second algorithm ALG2 is finished and ends in step 750, fixing and saving the parameters of the second algorithm ALG2.

[0308] The structural features of the apparatus described above, described in detail in the “Detailed Description of the Invention” and defined in the claims may be combined with the steps of the method of the present invention when appropriately replaced by corresponding processes.

[0309] Unless expressly stated, the singular forms "one", "the" used herein include the plural form (i.e., have the meaning of "at least one"). It should be further understood that the terms "having", "including" and / or "comprising" used in the specification indicate the presence of the described features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or combinations thereof. It should be understood that, unless expressly stated, when an element is referred to as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate intervening element. As used herein, the term "and / or" includes any and all combinations of one or more listed related items. Unless expressly stated, the steps of any method disclosed herein do not have to be performed in the exact order disclosed.

[0310] It should be appreciated that reference to "an embodiment" or "embodiment" or "aspect" or features that "may" include in this specification means that the specific features, structures or characteristics described in conjunction with the embodiment are included in at least one embodiment of the present invention. In addition, the specific features, structures or characteristics may be appropriately combined in one or more embodiments of the present invention. The foregoing description is provided to enable those skilled in the art to implement the various aspects described herein. Various modifications will be apparent to those skilled in the art.

[0311] The claims are not limited to the various aspects shown herein, but rather have the full scope consistent with the claim language wherein, unless expressly stated otherwise, elements referred to in the singular do not mean "one and only one" but rather "one or more." Unless expressly stated otherwise, the term "some" means one or more.

Claims

1. A method for training a first algorithm for extracting at least one desired component of a sound signal, the first algorithm comprising an encoder and a first decoder, each of the encoder and the first decoder comprising at least one parameter, wherein the method comprises training the encoder and the first decoder, comprising: Obtaining at least a portion of the masked in-domain data elements, wherein at least a portion of the masked in-domain data elements include noise components; The at least a portion of the masked intra-domain data elements are used to determine values ​​of at least a parameter of the encoder and at least a parameter of the first decoder to optimize the first algorithm's prediction of noisy components in at least a masked portion of the at least a portion of the masked intra-domain data elements.

2. The method according to claim 1, wherein: Training the encoder and the first decoder further comprises: obtaining at least a portion of the masked out-of-domain data elements, which includes a mixed component including a target component and a noise component; The at least a portion of the masked out-of-range data elements are used to determine values ​​of at least a parameter of the encoder and at least a parameter of the first decoder to optimize the first algorithm's prediction of a target component in at least a masked portion of the at least a portion of the masked out-of-range data elements.

3. The method according to claim 2, wherein: At least a portion of the masked out-of-domain data elements are partially masked spectrograms, and / or at least a portion of the masked in-domain data elements are partially masked spectrograms.

4. The method according to claim 2 or 3, wherein: Obtaining at least a portion of the masked out-of-domain data elements includes: - dividing the first out-of-domain data element including the mixed component including the target component and the noise component into a plurality of first patches; and - masking a predetermined percentage of the plurality of first patches to obtain at least a masked portion of at least a portion of masked out-of-domain data elements; and / or The data elements in the domain that are at least partially masked include: - dividing the first in-domain data elements including at least one noisy component into a plurality of second patches; and - masking a predetermined percentage of the plurality of second patches to obtain at least a masked portion of the data elements within at least a portion of the masked field.

5. The method according to any one of claims 2 to 4, wherein: During training of the encoder, the first algorithm includes a first decoder and a second decoder, wherein, - the second decoder is also trained using at least a portion of the masked out-of-domain data elements during training of the encoder, the encoder and the second decoder being trained to optimize the first algorithm's prediction of the target component in at least a masked portion of the at least a portion of the masked out-of-domain data elements; and -The first decoder is also trained using at least a portion of the masked in-domain data elements during training of the encoder, the encoder and the first decoder being trained to optimize the first algorithm's prediction of noisy components in at least a masked portion of at least a portion of the masked in-domain data elements and / or of target components in at least a masked portion of at least a portion of the masked out-of-domain data elements.

6. The method according to claim 5, wherein: The first decoder and / or the second decoder are discarded after training the encoder.

7. The method according to any one of claims 2 to 6, wherein: When determining the value of at least one parameter of an encoder using at least a portion of masked out-of-domain data elements, the value of the at least one parameter is determined by minimizing a first loss function that measures the difference between a predicted target component and a corresponding target component of at least a portion of the masked out-of-domain data elements.

8. The method according to any one of claims 1 to 7, wherein: When determining a value of at least one parameter of an encoder using at least a portion of the masked in-domain data elements, the value of the at least one parameter is determined by minimizing a second loss function that determines a difference between a predicted noisy component and a corresponding noisy component of the at least a portion of the masked in-domain data elements.

9. The method according to claim 7 or 8, wherein: Determining the difference includes determining the amplitude loss by providing the logarithm of the sum of squared differences between the absolute values ​​of the predicted target component and the corresponding target component of at least a portion of the masked out-of-domain data elements or between the predicted noisy component and the corresponding noisy component of at least a portion of the masked in-domain data elements.

10. The method according to claim 9, wherein: Determining differences also includes: - determining the phase loss by providing the logarithm of the sum of weighted squared differences between the normalized predicted target component and the normalized corresponding target component of at least a portion of the masked out-of-domain data elements or between the normalized predicted noisy component and the normalized corresponding noisy component of at least a portion of the masked in-domain data elements; - determining the magnitude-phase loss by providing a weighted sum between the magnitude loss and the phase loss, wherein the weighting comprises at least one weighting factor.

11. The method according to any preceding claim, further comprising: adding a third decoder to the first algorithm, the third decoder comprising at least one parameter; obtaining a second out-of-domain data element including a mixture of at least one target component and at least one noise component; and Training a third decoder using the trained encoder includes using the second out-of-domain data element to determine a value of at least one parameter of the third decoder to optimize the first algorithm's prediction of a target component of the second out-of-domain data element.

12. A method according to any preceding claim, wherein: The method is performed by a processor of at least one hearing aid and / or at least one computer.

13. A method for training a second algorithm for extracting at least one desired component of a sound signal, comprising: Extracting at least one target component from at least one second domain data element including the target component using a first algorithm trained according to any one of the methods of claims 1 to 12; The value of at least one parameter of the second algorithm is determined using the at least one second domain data element and the at least one target component extracted from the at least one second domain data element by the first algorithm to optimize the prediction of the at least one target component extracted from the at least one second domain data element.

14. A method for extracting at least one desired component of a sound signal using a first algorithm, the first algorithm being trained according to any one of the methods of claims 1-12, or a method for extracting at least one desired component of a sound signal using a second algorithm, the second algorithm being trained according to the method of claim 13.

15. A hearing aid, comprising: at least one input transducer configured to receive at least one first sound signal including a desired component and / or a noise component from an acoustic environment and to provide at least one electrical signal representative of the at least one first sound signal; A first algorithm trained according to any one of claims 1 to 12 or a second algorithm trained according to the method of claim 13, wherein the first algorithm or the second algorithm is configured to extract at least one desired component of at least one first sound signal; At least one output transducer is configured to output a second sound signal based on the at least one extracted desired component.