Hearing aid with feedback control based on machine learning
By training a machine learning model for the feedback control system, predicting intermediate signals and performing post-processing, the artifact problem in existing acoustic feedback cancellation systems is solved, achieving faster convergence speed and lower steady-state error, and improving signal quality and intelligibility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2026-03-10
AI Technical Summary
Existing machine learning-based acoustic feedback cancellation systems may introduce artifacts when predicting signals without feedback, and they consume a lot of computational resources, making it difficult to achieve a balance between convergence speed and steady-state performance without affecting sound quality.
By training a machine learning model for the feedback control system, intermediate signals are predicted and post-processed to generate feedback-free signals. A layer structure consisting of convolutional layers, a first fully connected layer, and long short-term memory layers is used to estimate and post-process the feedback path transfer function, thereby reducing the impact of artifacts.
Without compromising sound quality, it achieves faster convergence speed and lower steady-state error, improving signal quality and intelligibility while reducing computational load.
Smart Images

Figure CN121645112A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of hearing aids, and more particularly to hearing aids with machine learning (ML) feedback control functionality. This application also relates to a method for training a machine learning model used in the feedback control system of a hearing aid, and related hearing aids. Background Technology
[0002] Acoustic feedback cancellation is an important and challenging task in audio processing systems, aiming to mitigate the impact of feedback loops on system stability and sound quality. Acoustic feedback occurs when a sound signal emitted by a speaker is picked up by a microphone, amplified, and then played back through the same speaker, forming a continuous loop (e.g., a feedback loop). This can lead to undesirable effects such as echo, ringing, reverberation, and howling.
[0003] Current technologies rely on adaptive filtering algorithms, but these algorithms face challenges in balancing fast convergence with low steady-state error. Specifically, the performance of these techniques (such as variable step size algorithms) has reached a bottleneck: in real-world environments, they are unable to accurately estimate acoustic feedback or respond quickly to changes in the acoustic feedback path (e.g., in hearing aid (HA) scenarios, a response within hundreds of milliseconds is required to maintain stability).
[0004] Machine learning techniques (such as deep neural networks (DNNs)) have been integrated as optimal step-size predictors (e.g., estimators) into filter-based acoustic feedback cancellation (AFC) and acoustic echo cancellation (AEC) systems. Specifically, existing machine learning-based AFC systems of this type, if configured to directly predict signals without feedback (e.g., signals with feedback correction), may introduce artifacts into the signals without feedback if the prediction results are not optimal. Summary of the Invention
[0005] When machine learning (ML)-based acoustic feedback cancellation (AFC) systems are configured to directly predict the signal of interest (e.g., rather than predicting intermediate signals), sound quality and speech intelligibility may be compromised. Furthermore, training such existing machine learning-based AFC systems can consume significant computational resources.
[0006] There is an urgent need for a machine learning-based feedback control system (e.g., a feedback control system that includes a trained machine learning model) that can address the shortcomings of existing technologies and achieve an ideal balance between convergence speed and steady-state performance without affecting sound quality.
[0007] Embodiments of the present invention provide a method for training a machine learning model for a feedback control system. In other words, training this machine learning model enables the machine learning-based feedback control system to achieve an ideal balance between convergence speed and steady-state performance without affecting sound quality. Specifically, the machine learning model is trained to predict intermediate signals, and the signal of interest is determined based on these intermediate signals. For example, when the prediction is not optimal, the intermediate signal can be post-processed to provide a more accurate signal of interest. In other words, unlike traditional adaptive filtering techniques and existing machine learning-based AFC techniques, embodiments of the present invention enable post-processing of intermediate signals (e.g., the output of the machine learning model) to generate an improved version without feedback.
[0008] Training methods
[0009] This invention provides a method, executed by an electronic device, for training a machine learning (ML) model used in a hearing aid feedback control system. The feedback control system includes a machine learning model.
[0010] This method involves performing multiple training iterations.
[0011] Each training iteration in multiple training iterations includes acquiring training data.
[0012] The training data includes the training input signal and the processed signal. The training input signal includes external input signal components and feedback input signal components. The external input signal components represent the sound from a known simulated acoustic environment of the auto-audio system. The feedback input signal components represent the acoustic and / or mechanical feedback from the auto-audio system's feedback path.
[0013] The post-training signal represents one or more processing algorithms applied to the training feedback correction input signal. The training feedback correction input signal represents a feedback-corrected version of the training input signal. For example, the training feedback correction input signal indicates an acoustic and / or mechanical feedback-corrected version of the training input signal. In other words, the post-training signal can be a processed version of the training feedback correction input signal. In one or more exemplary methods, the training data for each training iteration (e.g., acquisition) in multiple training iterations can be understood as a training sequence.
[0014] Each training iteration in a series of training iterations includes acquiring target data, which contains the training feedback path transfer function. The training feedback path transfer function represents the impulse response of the hearing aid feedback path (FBP). The target data can be considered as reference data.
[0015] Each training iteration in multiple training iterations includes determining an estimate of the training feedback path transfer function based on the training data.
[0016] Each training iteration in multiple training iterations includes updating the machine learning model based on the target data (e.g., the training feedback path transfer function) and the estimated value of the training feedback path transfer function.
[0017] The machine learning model contains a convolutional layer, a first fully connected (FC) layer, and a long short-term memory (LSTM) layer in the following order.
[0018] In one or more exemplary methods, the method is performed by an electronic device (e.g., a computer). For example, the method may be performed by an external device (e.g., a device external to a hearing aid). For example, the training method may be a computer-implemented method. For example, the method includes performing multiple training iterations (e.g., multiple rounds) by an electronic device (e.g., a computer).
[0019] In one or more exemplary methods, the training method can be performed during an offline training phase. In other words, a known simulated acoustic environment can be understood as an environment that models (e.g., simulates) the actual use environment of a hearing aid. For example, such a known simulated acoustic environment can be generated through computer simulation. For example, the offline training phase can be understood as a process of representatively modeling a real-world scenario (e.g., a complex acoustic scenario) in a computer simulation.
[0020] In one or more exemplary methods, the method may be performed using training data and target data from multiple known simulated acoustic environments. For example, the training data and target data may be provided via computer simulation of a hearing aid in one or more acoustic environments used to model (e.g., simulate) a real-world scene, the hearing aid including a known feedback system (e.g., an ideal feedback system). In other words, the method may generate training data and target data via computer simulation (e.g., by an electronic device). For example, the step of acquiring training data and target data includes generating training data and target data via computer simulation. For example, the method may include retrieving training data and target data from the memory of an electronic device (e.g., a computer). Optionally, the method may also include retrieving training data and target data from the memory of the hearing aid.
[0021] For example, training and target data can be generated using a known feedback control system (e.g., an ideal feedback control system), for example in static feedback scenarios and / or dynamic feedback scenarios (e.g., scenarios where the dynamic feedback path changes). For example, training and target data can be obtained through computer simulation of a hearing aid in a known simulated acoustic environment, the hearing aid containing a known (e.g., ideal) feedback control system.
[0022] Training and target data can represent the characteristics of a known feedback control system, such as a feedback control system capable of immediate response to changes in the feedback path, without the convergence time required by an adaptive filter. For example, a known feedback control system can refer to a feedback control system capable of immediate and accurate response to feedback changes because the training feedback path transfer function (e.g., the acoustic feedback of a source auto-audiophone feedback path) is known. Alternatively, a known feedback control system can also be understood as a feedback control system with a known feedback path transfer function, such as a measured feedback path transfer function (e.g., measured in a real acoustic environment) or a simulated (e.g., synthetically generated) feedback path transfer function. In other words, a machine learning model can be trained using both measured and synthetically generated feedback path transfer functions. Such synthetic feedback path transfer functions can be generated (e.g., through computer simulation) to match real acoustic environments.
[0023] In one or more exemplary methods, training a machine learning model using training data and target data obtained from the hearing aid in a known simulated environment, based on a hearing aid containing a known feedback control system (e.g., an ideal feedback control system), avoids the machine learning model learning from poorly performing feedback control systems (e.g., current-level feedback control systems, which suffer from defects such as slow convergence and incorrect response to changes in the feedback path). In other words, an ideal feedback control system avoids many defects of current-level feedback control systems (e.g., slow convergence, incorrect response).
[0024] Embodiments of this application can provide machine learning models trained on known (e.g., ideal, perfect) feedback control systems, without requiring decorrelation processing. When generating training and target data, the aim is to minimize artifacts in the output signal (e.g., the signal after training processing) when the feedback path changes abruptly. The training and target data can be understood as data used to train machine learning models (e.g., training the feedback control system of a hearing aid).
[0025] In one or more exemplary methods, training and target data can be generated via computer simulation to reflect the characteristics of a known (e.g., ideal) feedback control system. For example, data for training machine learning models can be generated using a known feedback control system in both static and dynamic feedback scenarios (e.g., scenarios with changing dynamic feedback paths).
[0026] In one or more exemplary methods, the external input signal component of the training input signal may include one or more of the following signals: white noise, speech signal, and music signal. In other words, the training input signal may include any one or more of the following mixed signals: white noise, speech signal, and music signal. For example, the external input signal may include data from multiple sound sources. The multiple sound sources may include one or more of the following: noise, speech, music, and sounds recorded from everyday life scenarios as input sound from the hearing aid. The external input signal component may be a portion of the training input signal that is independent of feedback. In other words, the external input signal component (e.g., denoted as x(n)) may be the desired signal that the hearing aid signal processing unit needs to process.
[0027] In one or more exemplary methods, the training data may include multiple training input signals and a trained processed signal. For example, each training input signal includes an external input signal component and a feedback input signal component. The feedback input signal component represents the acoustic feedback of the source hearing aid feedback path, wherein the hearing aid includes multiple feedback paths. For example, the target data may include multiple feedback path transfer functions, each feedback path transfer function representing the impulse response of a corresponding feedback path among the multiple feedback paths.
[0028] In one or more exemplary methods, the method (e.g., each training iteration in multiple training iterations) includes retrieving training data and target data from the memory of a hearing aid (e.g., in a known simulated environment) or from the memory of an electronic device (such as a computer). The electronic device may be an electronic device that performs the training method. The electronic device may also be another electronic device different from the electronic device performing the training method.
[0029] In one or more exemplary methods, the applied processing algorithm (e.g., technique) may include one or more of the following: noise reduction algorithms (e.g., related to beamforming and / or post-filtering), compression algorithms (e.g., related to providing gain that varies with frequency and level), transform domain algorithms (e.g., frequency domain transform algorithms), spatial sound processing algorithms, and any other applicable processing algorithms. For example, a compression algorithm may be considered as a processing algorithm for compensating for a user's hearing impairment. For example, a transform domain algorithm may be considered as an algorithm for supporting processing in the transform domain (e.g., multiple frequency bands).
[0030] Optionally, the signal after training processing can represent a frequency-dependent and / or level-dependent gain function applied to the training feedback correction input signal. In other words, the frequency-dependent and / or level-dependent gain function can represent a function that models one or more processing algorithms. The frequency-dependent and / or level-dependent gain function can be understood as a time-varying and / or frequency-dependent function. The frequency-dependent and / or level-dependent gain function can be understood as a forward path gain function.
[0031] In one or more exemplary methods, the feedback path of an acoustic and / or mechanical feedback source auto-auditory is defined from the output unit (e.g., the output transducer) to the input unit (e.g., the input transducer). For example, the feedback path transfer function (e.g., the training feedback path transfer function) can be denoted as h(n) = [h1(n), h2(n), ..., h L (n)] T , where L represents the length of the feedback path impulse response, and n represents (e.g., discrete) time index.
[0032] In one or more exemplary methods, the processed signal (e.g., the training processed signal, denoted as u(n)) can be considered as an output signal that will be converted into an acoustic signal (e.g., sound) by a speaker (e.g., contained in a hearing aid output unit). The training processed signal may be an amplified and / or processed version of an external input signal component.
[0033] In one or more exemplary methods, the feedback input signal component (e.g., denoted as v(n)) can be viewed as a filtered version of the trained signal. For example, the trained signal can be filtered using a training feedback path transfer function to obtain the aforementioned filtered version.
[0034] In one or more exemplary methods, the training input signal (e.g., denoted as y(n)) can be considered as a signal picked up by an input transducer (e.g., a microphone) included in the hearing aid input unit. In one or more exemplary methods, the training input signal is disrupted by acoustic and / or mechanical feedback. The training input signal can be denoted as y(n) = v(n) + x(n).
[0035] In one or more exemplary methods, a normal mode (e.g., an inference phase) can be executed after a training mode (e.g., a training phase). In other words, after the training phase, multiple weights of the machine learning model can be fixed. In normal mode, the machine learning model is trained (e.g., multiple weights remain fixed). However, in training mode, multiple weights of the machine learning model can be updated based on training data and target data. In other words, the machine learning model can be trained in training mode by updating the machine learning model (e.g., updating multiple weights). For example, the training method can be executed in a training mode (e.g., for a hearing aid).
[0036] In one or more exemplary methods, the machine learning model is configured to receive a training input signal and a trained processed signal as input. Optionally, the machine learning model may be configured to receive an input signal derived from (e.g., based on) the training input signal and the trained processed signal as input. In one or more exemplary methods, the machine learning model is configured to output an estimate of the training feedback path transfer function.
[0037] In one or more exemplary methods, a machine learning model may include an input layer, multiple hidden layers, and an output layer. For example, the input layer may contain convolutional layers. For example, the multiple hidden layers may contain a first fully connected (FC) layer and a long short-term memory (LSTM) layer. For example, the output layer may contain pooling layers.
[0038] In one or more exemplary methods, a fully connected layer may include a feedforward (FF) layer. For example, a machine learning model may include a deep neural network (DNN).
[0039] Embodiments of the present invention can provide hearing aids with improved signal quality and signal intelligibility because the machine learning model used in the hearing aid feedback control system is configured to determine (e.g., provide) an estimate of the training feedback path transfer function (e.g., the feedback path impulse response), thereby allowing post-processing of this estimate. Post-processing of this estimate helps eliminate artifacts caused by suboptimal predictions from the machine learning model, thus improving signal quality and signal intelligibility. In other words, the training method provided by embodiments of the present invention trains the machine learning model to estimate the feedback impulse response, rather than directly estimating the feedback correction input signal (e.g., providing an alternative to conventional adaptive filtering techniques and existing machine learning techniques). The estimate of the training feedback path transfer function may contain defects caused by the aforementioned suboptimal predictions; these defects can be removed from the training feedback path transfer function during post-processing, or retained if the defects are negligible (e.g., have a minor impact).
[0040] Embodiments of the present invention provide a machine learning-based feedback control method (e.g., a training method) that operates frame-by-frame to directly determine estimates of the feedback path impulse response. In other words, embodiments of the present invention require the machine learning model to include a convolutional layer, a first fully connected (FC) layer, and a long short-term memory (LSTM) layer in the following order. This layer structure (and corresponding signal processing) achieves faster convergence and lower steady-state error, thereby improving the balance between convergence speed and steady-state error. In other words, embodiments of the present invention provide a faster-converging and more stable feedback cancellation (e.g., control) system.
[0041] Furthermore, the estimates of training feedback path transfer functions can be of an acceptable size, thereby reducing the computational cost of machine learning models (e.g., during the training and / or inference phases). For example, a set of feedback path transfer functions (e.g., impulse responses) is largely determined by the positions of the input transducers (e.g., microphones) and the output transducers (e.g., speakers), without requiring a large representation space (e.g., a wide variety of signals), and therefore can be more easily estimated (e.g., represented) even when using size-constrained machine learning models (e.g., deep neural networks (DNNs)).
[0042] The method used to train the machine learning model for the hearing aid feedback control system can be called the deep feedback compensation method (e.g., reference). Figure 9 This can be represented as DFC, DFC(M), or DFC(S). A feedback control system that includes a trained machine learning model of the hearing aid (e.g., a machine learning-based feedback control system) can be called a deep feedback compensation (DFC) system.
[0043] In one or more exemplary methods, the machine learning model may further include one or more of the following: a second fully connected layer and a third fully connected layer. In one or more exemplary methods, each of the first, second, and third fully connected layers includes a fully connected class technique representing an activation function. In other words, at least one feedforward layer includes an activation function. In one or more exemplary methods, the activation function may introduce nonlinear characteristics into the machine learning model (e.g., a machine learning model layer containing the activation function). In one or more exemplary methods, at least one feedforward layer may be configured to receive input data from the previous layer of the machine learning model. For example, the activation function may be one or more of the following: a rectified linear unit (RELU) function, a hyperbolic tangent (tanh) function, and a SoftMax function. For example, the layer containing the activation function is configured to pass the output based on the above processing to the next layer of the machine learning model.
[0044] In one or more exemplary methods, the activation function includes one or more of the following: the sigmoid function, the hyperbolic tangent (tanh) function, the corrected linear unit (ReLU) function, the leaky corrected linear unit (leaky ReLU) function, the swish function, and the Gaussian error linear unit (GELU) function. For example, the sigmoid function can be expressed as f(x) = σ(x) = (1 + e^(-1 / 2)) / 2. -x ) -1 Here, x represents the input data. The sigmoid function maps input data to the range of 0 to 1. For example, the hyperbolic tangent function can be expressed as f(x) = tanh(x) = (e^(x-1) / (x-1)). x -e -x ) / (e x +e -xThe hyperbolic tangent function maps input data to the range of -1 to 1. For example, the ReLU function can be expressed as f(x) = max(0,x), where x represents the input data. When the input data is negative, the ReLU function converts it to 0 (e.g., if x ≤ 0, then f(x) = 0); when the input data is positive, the ReLU function directly outputs the input data (e.g., if x > 0, then f(x) = x). For example, the leaky ReLU function can be expressed as f(x) = max(αx,x), where x represents the input data and α represents a (small) positive constant; when the input data is positive, the leaky ReLU function directly outputs the input data (e.g., if x > 0, then f(x) = x). When the input data is negative, the leaky ReLU function converts the input data into a linear form (e.g., if x ≤ 0, then f(x) = αx). For example, the swish function can be represented as xσ(x), where x represents the input data (e.g., the input data of the layer containing the activation function), and σ(x) represents the sigmoid function. The swish function can be viewed as a sigmoid-based function. For example, the GELU function can be represented as xΦ(x), where x represents the input data, and Φ(x) represents the standard Gaussian cumulative distribution function.
[0045] For example, activation functions can also include arctangent functions (e.g., f(x) = atan(x)), which map the input data to the range of -π / 2 to π / 2. Activation functions can also include softplus functions (e.g., denoted as f(x) = ln(1 + e^(-π / 2)) x For example, activation functions may include one or more of the following: sigmoid-based functions, ReLU-based functions, exponential linear unit (ELU)-based functions, square root linear unit (SRLU)-based functions, SoftMax functions, and any other applicable activation functions.
[0046] In one or more exemplary methods, each layer of the machine learning model may include an activation layer. In one or more exemplary methods, a recurrent neural network may include one or more of the following: Long Short-Term Memory (LSTM) layers and Gated Recurrent Unit (GRU) layers. In one or more exemplary methods, a convolutional neural network (CNN) may include convolutional layers.
[0047] For example, multiple hidden layers may include one or more of the following: a first fully connected layer, a long short-term memory (LSTM) layer, and a second and a third fully connected layer (e.g., at least one). For example, the output layer may include one or more of the following: a second fully connected layer, a third fully connected layer, and a pooling layer.
[0048] In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying a preprocessing technique to the training input signal to determine a first preprocessed signal. In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying the preprocessing technique to a processed training signal to determine a second preprocessed signal. In one or more exemplary methods, determining the first preprocessed signal (e.g., applying the preprocessing technique to the training input signal) includes applying a Fourier transform-based technique to the training input signal to determine a frequency domain training input signal. In one or more exemplary methods, determining the second preprocessed signal (e.g., applying the preprocessing technique to a processed training signal) includes applying a Fourier transform-based technique to the processed training signal to determine a frequency domain processed training signal.
[0049] For example, Fourier transform-based techniques include one or more of the following: Discrete Fourier Transform (DFT) and Short-Time Fourier Transform (STFT). For example, applying Fourier transform-based techniques to a training input signal and a processed training signal includes providing time-frequency representations of the training input signal and the processed training signal. For example, the training input signal and the processed training signal are time-domain signals. For example, both the frequency-domain training input signal and the frequency-domain processed training signal are frequency-domain signals. A frequency-domain signal may contain multiple frequency components (e.g., frequency bands), and each frequency component may contain multiple time components (e.g., at least one time component).
[0050] In one or more exemplary methods, the training input signal and the processed training signal can be represented as y(n) and u(n), respectively, where n represents a time index (e.g., a discrete-time index). In one or more exemplary methods, each of the frequency-domain training input signal and the processed frequency-domain training signal comprises a set of frames and a set of frequency bins (e.g., frequency units). For example, the frequency-domain training input signal and the processed frequency-domain training signal can be represented as Y(m,k) and U(m,k), respectively, where m = 1, 2, ..., M, and k = 1, 2, ..., K. For example, M represents the set of frames (e.g., the number of frames); K represents the set of frequency bins (e.g., the number of frequency bins, i.e., frequency units). For example, each of the frequency-domain training input signal and the processed frequency-domain training signal comprises multiple time-frequency bins (e.g., time-frequency units).
[0051] In one or more exemplary methods, determining a first preprocessed signal includes determining a normalized version of the frequency-domain training input signal. In other words, determining the first preprocessed signal includes performing a normalization operation on the frequency-domain training input signal. In one or more exemplary methods, determining a second preprocessed signal includes determining a normalized version of the frequency-domain training processed signal. In other words, determining the second preprocessed signal includes performing a normalization operation on the frequency-domain training processed signal.
[0052] In one or more exemplary methods, each of the frequency-domain training input signal and the frequency-domain training processed signal may be normalized based on the energy of the frequency-domain training processed signal. In other words, determining the respective normalized versions of the frequency-domain training input signal and the frequency-domain training processed signal may include determining the energy of the frequency-domain training processed signal (e.g., the energy associated with the frequency-domain training processed signal, i.e., the energy of the training processed signal in the frequency domain).
[0053] For example, the normalized version of the frequency domain training input signal can be expressed as: Where m = 1, 2, ..., M, k = 1, 2, ..., K; the normalized version of this frequency domain training input signal can be obtained through... This is calculated. For example, the normalized version of the signal after frequency domain training processing can be expressed as... Where m = 1, 2, ..., M, k = 1, 2, ..., K; the normalized version of the signal after frequency domain training processing can be obtained through... Calculated. For example, the energy of the signal after frequency domain training (e.g., the energy associated with the signal after frequency domain training) can be obtained through... Determine (e.g., calculate).
[0054] In one or more exemplary methods, the normalized version of the frequency domain training input signal includes a first principal component and a first component. In one or more exemplary methods, the normalized version of the frequency domain training processed signal includes a second principal component and a second component. In other words, when applying preprocessing techniques to the frequency domain training input signal and the frequency domain training processed signal respectively, it may include performing a decomposition operation on each of the frequency domain training input signal and the frequency domain training processed signal. For example, performing a decomposition operation on the frequency domain training input signal means decomposing it into a first principal component (e.g., a first amplitude or a first real value) and a first component (e.g., a first phase or a first imaginary value). For example, the first principal component and the first component may be represented as follows: and For example, performing a decomposition operation on a signal after frequency domain training processing decomposes it into a second principal component (e.g., a second amplitude or a second real value) and a second component (e.g., a second phase or a second imaginary value). For instance, the second principal component and the second component can be represented as follows: and
[0055] For example, the first principal component can be the first logarithmic magnitude (e.g. For example, the first logarithmic magnitude can be obtained through... Calculated. For example, the second principal component can be the second logarithmic magnitude (e.g. For example, the second logarithmic magnitude can be obtained through... The calculation yielded the result.
[0056] For example, the first component can be the first phase (e.g.) For example, the first phase can be achieved through... Calculated. For example, the second component can be the second phase (e.g. For example, the second phase can be achieved through... The calculation yielded the result.
[0057] For example, the first principal component is a matrix of size (e.g., dimension) M×K, such as matrix It can contain multiple elements There are a total of MK amplitude elements. For example, the first component is a matrix of size (e.g., dimension) M×K, such as... matrix It can contain multiple elements There are a total of MK phase elements.
[0058] For example, the second principal component is a matrix of size (e.g., dimension) M×K, such as matrix Contains multiple elements There are a total of MK amplitude elements. For example, the second component is a matrix of size (e.g., dimension) M×K, such as... matrix Contains multiple elements There are a total of MK phase elements.
[0059] In one or more exemplary methods, the first preprocessed signal can be viewed as a matrix containing a first principal component and a first component (e.g., a matrix). (The size is M×2K). In one or more exemplary methods, the second preprocessed signal can be viewed as a matrix containing a second principal component and a second component (e.g., matrix). (The size is M×2K). In one or more exemplary methods, the first preprocessed signal and the second preprocessed signal have the same size.
[0060] In one or more exemplary methods, each of the first preprocessed signal and the second preprocessed signal can be regarded as a feature signal (e.g., a feature matrix).
[0061] For example, preprocessing techniques include one or more of the following: Fourier transform-based techniques, normalization operations, and decomposition operations. Preferably, the preprocessing techniques simultaneously include Fourier transform-based techniques, normalization operations, and decomposition operations.
[0062] For example, machine learning models are configured to process a limited range of data by applying preprocessing techniques to both the training input signal and the post-training processed signal. Applying preprocessing techniques to both the training input signal and the post-training processed signal can be viewed as a data representation transformation. This data representation transformation is particularly important, for example, when the machine learning model contains layers with (e.g., non-linear) activation functions.
[0063] In one or more exemplary methods, the convolutional layer includes convolutional techniques. In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying convolutional techniques to a first preprocessed signal and a second preprocessed signal to determine a first machine learning processed signal. A convolutional layer can be considered a single layer. A convolutional layer may contain an activation function. A convolutional layer may not contain an activation function.
[0064] In one or more exemplary methods, a convolutional layer includes a filter (e.g., a kernel) of a given size. For example, the filter includes a set of weights. In one or more exemplary methods, applying convolutional techniques to a first preprocessed signal and a second preprocessed signal includes applying (e.g., employing) a filter of a given size to the first and second preprocessed signals in both a time dimension (e.g., a set of frames) and a frequency dimension (e.g., a set of frequency bins).
[0065] For example, the number of channels in a filter (e.g., a kernel) can be the same as the number of channels in the input data (e.g., two channels, corresponding to the first preprocessed signal and the second preprocessed signal, respectively). In other words, the size of the filter can be 2×I×J, where I represents the number of rows in the filter (e.g., the number of frames) and J represents the number of columns in the filter (e.g., the number of frequency bins). A filter of size 2×I×J can be considered as two sub-filters (i.e., the first sub-filter and the second sub-filter), each of size I×J, and convolved with the first preprocessed signal and the second preprocessed signal, respectively. The size of each of the two sub-filters is smaller than the size of the corresponding preprocessed signal.
[0066] For example, each of the first and second sub-filters is a two-dimensional (2D) filter (e.g., a two-dimensional array of weights). In one or more exemplary methods, applying convolution-like techniques to the first and second preprocessed signals includes applying convolution operations based on two-dimensional (2D) filters to each of the first and second preprocessed signals.
[0067] For example, applying convolution-like techniques to the first and second preprocessed signals includes: performing a dot product (e.g., element-wise multiplication) on an array (e.g., a set of elements, i.e., a slice of the first sub-filter size) in the first preprocessed signal with the first sub-filter. Alternatively, applying convolution-like techniques to the first and second preprocessed signals includes: applying the first sub-filter (e.g., acting on multiple slices of the first sub-filter size of the first preprocessed signal) across the entire size range of the first preprocessed signal to obtain a first filtered signal.
[0068] For example, applying convolution-like techniques to the first and second preprocessed signals includes: performing a dot product (e.g., element-wise multiplication) on an array (e.g., a set of elements, i.e., a slice of the second sub-filter size) in the second preprocessed signal that has the same size as the second sub-filter. Alternatively, applying convolution-like techniques to the first and second preprocessed signals includes: applying the second sub-filter (e.g., acting on multiple slices of the second sub-filter size of the second preprocessed signal) across the entire size range of the second preprocessed signal to obtain a second filtered signal.
[0069] For example, convolutional techniques can be applied to the first and second preprocessed signals, including summing the first and second filtered signals (e.g., summing the outputs of the two channels of a filter with a size of 2×I×J).
[0070] In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes determining the size of a first machine learning processed signal to be the same as the size of a first preprocessed signal and a second preprocessed signal. For example, determining the first machine learning processed signal includes zero-padding the first and second preprocessed signals respectively in a time dimension (e.g., a set of frames) and a frequency dimension (e.g., a set of frequency bins). For example, zero-padding the first and second preprocessed signals respectively ensures that the size of the first machine learning processed signal is the same as the size of both the first and second preprocessed signals.
[0071] Embodiments of the present invention can implement causal systems (which are crucial, for example, in hearing aid applications) by ensuring that the context information of each frame is derived only from previous frames through appropriate padding in the time dimension (e.g., padding with zeros at the beginning of the first and second preprocessed signals). For example, the “context” of a frame refers to the information contained in that frame (e.g., the content of the frame).
[0072] In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes concatenating a first machine learning processing signal, a first preprocessed signal, and a second preprocessed signal along the frequency dimension to determine a second machine learning processing signal. For example, determining the estimate of the training feedback path transfer function includes determining the second machine learning processing signal based on the first machine learning processing signal, the first preprocessed signal, and the second preprocessed signal. For example, the second machine learning processing signal can be considered to include the first machine learning processing signal (e.g., denoted as C) and the first preprocessed signal (e.g., denoted as Y). pre ) and the second preprocessed signal (e.g., denoted as U) pre A matrix. For example, the second machine learning processing signal can be represented as a matrix [CU]. pre Y pre ]∈R m×6K or [CY] pre U pre ]∈R M×6K The size is M×6K. For example, the first machine learning processing signal is concatenated with the first preprocessed signal and the second preprocessed signal in the frequency dimension (e.g., a set of frequency bins). For example, the first machine learning processing signal (e.g., C), the first preprocessed signal (e.g., Y) pre ) and the second preprocessed signal (e.g., U pre The dimensions of both are M×2K. For example, the second machine learning processing signal can be represented by a matrix (e.g., C) representing the first machine learning processing signal and a matrix (e.g., Y) representing the first preprocessed signal. pre ) and a matrix representing the second preprocessed signal (e.g., U pre This is obtained by horizontally splicing the pieces together.
[0073] In one or more exemplary methods, the first fully connected layer includes a first fully connected class technique. In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying the first fully connected class technique to a second machine learning processing signal to determine a third machine learning processing signal. In one or more exemplary methods, the first fully connected layer includes an activation function. The activation function of the first fully connected layer may be a ReLU-based activation function, such as a leaky ReLU function (or any other activation function mentioned above in this invention). The first fully connected layer may be a feedforward (FF) layer. In one or more exemplary methods, determining the third machine learning processing signal includes applying the aforementioned activation function to the second machine learning processing signal.
[0074] Optionally, determining an estimate of the training feedback path transfer function includes applying a first fully connected layer technique to the first machine learning processed signal, the first preprocessed signal, and the second preprocessed signal to determine the third machine learning processed signal. For example, the first fully connected layer is configured to receive the first machine learning processed signal, the first preprocessed signal, and the second preprocessed signal as input. In other words, the first fully connected layer may be configured to receive a second machine learning processed signal (e.g., a signal containing the first machine learning processed signal, the first preprocessed signal, and the second preprocessed signal) as input.
[0075] For example, the first machine learning processing signal is a poor estimate of the training feedback path transfer function. For example, when the first fully connected layer receives a second machine learning processing signal (e.g., containing the first machine learning processing signal, a first preprocessed signal, and a second preprocessed signal) as input, it can provide a better estimate of the training feedback path transfer function (compared to the first machine learning processing signal).
[0076] In one or more exemplary methods, the Long Short-Term Memory (LSTM) layer incorporates LSTM-like techniques. In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying LSTM-like techniques to a third machine learning processing signal to determine a fourth machine learning processing signal. In one or more exemplary methods, the fourth machine learning processing signal is the estimate of the training feedback path transfer function.
[0077] An LSTM layer can be considered as a single layer. An LSTM layer may or may not contain an activation function.
[0078] In one or more exemplary methods, a Long Short-Term Memory (LSTM)-like technique is applied to the third machine learning processing signal, including determining multiple dependencies between frames. For example, the LSTM layer is configured to utilize (e.g., manage) a set of dependencies (e.g., long-term dependencies) between frames in the third machine learning processing signal. For example, the size of the fourth machine learning processing signal is smaller than the size of the second machine learning processing signal. The fourth machine learning processing signal can be used as the output of a machine learning model.
[0079] In one or more exemplary methods, the LSTM layer can be replaced with a gated recurrent unit (GRU) layer. For example, in the layer structure described above, an LSTM layer may be a better choice than a GRU layer.
[0080] In one or more exemplary methods, the second fully connected layer includes a second fully connected class technique. In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying the second fully connected class technique to a fourth machine learning processing signal to determine a fifth machine learning processing signal. In one or more exemplary methods, the fifth machine learning processing signal is the estimate of the training feedback path transfer function.
[0081] In one or more exemplary methods, the second fully connected layer includes an activation function. The activation function of the second fully connected layer may be a ReLU-based activation function, such as a leaky ReLU function (or any other activation function mentioned above in this invention). In one or more exemplary methods, determining the fifth machine learning processing signal includes applying the aforementioned activation function to the fourth machine learning processing signal. The fifth machine learning processing signal may be the output of a machine learning model. For example, the fifth machine learning processing signal has a higher accuracy in estimating the training feedback path transfer function than the fourth machine learning processing signal.
[0082] By employing a layer structure (and corresponding signal processing) in the order of convolutional layer, first fully connected layer, LSTM layer and second fully connected layer, faster convergence speed and lower steady-state error can be achieved, thereby improving the balance between convergence speed and steady-state error.
[0083] In one or more exemplary methods, the third fully connected layer includes a third fully connected class technique. In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying the third fully connected class technique to a fifth machine learning processing signal to determine a sixth machine learning processing signal. In one or more exemplary methods, the sixth machine learning processing signal is the estimate of the training feedback path transfer function.
[0084] In one or more exemplary methods, the third fully connected layer includes an activation function. The activation function of the third fully connected layer may be a tangent-based activation function, such as the hyperbolic tangent function (or any other activation function mentioned above in this invention). In one or more exemplary methods, determining the sixth machine learning processing signal includes applying the aforementioned activation function to the fifth machine learning processing signal. The sixth machine learning processing signal may be the output of a machine learning model.
[0085] In one or more exemplary methods, the sixth machine learning processing signal contains an appropriate estimate of the training feedback path transfer function for each frame. In one or more exemplary methods, the sixth machine learning processing signal may be represented as a matrix h. ′ m ∈R M×P, where P represents the number of coefficients (e.g., taps). For example, the sixth machine learning processing signal has a higher accuracy in estimating the training feedback path transfer function than the fifth machine learning processing signal.
[0086] Employing a layer structure (and corresponding signal processing) in the order of convolutional layers, a first fully connected layer, an LSTM layer, a second fully connected layer, and a third fully connected layer can achieve faster convergence and lower steady-state error, thereby improving the balance between convergence speed and steady-state error. For example, the sixth machine learning processing signal is of moderate size, allowing for the use of scale-constrained machine learning models (e.g., low-complexity machine learning models). For instance, such machine learning models can provide appropriate (e.g., accurate) estimates of the training feedback path transfer function while maintaining computational efficiency.
[0087] In one or more exemplary methods, the machine learning model further includes a pooling layer. In one or more exemplary methods, the pooling layer includes pooling-like techniques. In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes applying pooling-like techniques to a sixth machine learning processing signal to determine a seventh machine learning processing signal. In one or more exemplary methods, the seventh machine learning processing signal is the estimate of the training feedback path transfer function.
[0088] In one or more exemplary methods, the first preprocessing signal, the second preprocessing signal, the first machine learning processing signal, the second machine learning processing signal, the third machine learning processing signal, the fourth machine learning processing signal, the fifth machine learning processing signal, and the sixth machine learning processing signal are all associated with the same plurality of frames (e.g., time frames, such as M frames).
[0089] In one or more exemplary methods, the pooling layer is an average pooling layer. For example, applying pooling techniques to the sixth machine learning processing signal includes performing a moving average operation over multiple frames (e.g., past frames and the current frame, N frames in total). Pooling techniques may include moving average operations. For example, the pooling layer is used to smooth the sixth machine learning processing signal (i.e., the estimate of the training feedback path transfer function). Optionally, the pooling layer may also be a max pooling layer. The seventh machine learning processing signal may be the output of a machine learning model.
[0090] In one or more exemplary methods, the seventh machine learning processing signal can be represented as a matrix. The size is (M-(N-1))×P. For example, the seventh machine learning processing signal can be determined as:
[0091]
[0092] The parameter N can be significant in controlling convergence speed and estimation accuracy (e.g., the estimation accuracy of the training feedback path transfer function). In other words, the value of parameter N is chosen to enable improvements in both convergence speed (e.g., the speed of convergence to the target data) and estimation accuracy.
[0093] The seventh machine learning signal may have a higher accuracy in estimating the transfer function of the training feedback path than the sixth machine learning signal. The seventh machine learning signal may have an accuracy comparable to the sixth machine learning signal.
[0094] By employing a layer structure (and corresponding signal processing) in the order of convolutional layer, first fully connected layer, long short-term memory (LSTM) layer, second fully connected layer, third fully connected layer and pooling layer, faster convergence speed and lower steady-state error can be achieved, thereby improving the balance between convergence speed and steady-state error.
[0095] In one or more exemplary methods, determining an estimate of the training feedback path transfer function includes providing a first preprocessed signal and a second preprocessed signal as input to a machine learning model. For example, determining an estimate of the training feedback path transfer function includes providing a training input signal and a training-processed signal as input to the machine learning model. For example, the machine learning model is configured to output an estimate of the training feedback path transfer function.
[0096] In one or more exemplary methods, updating the machine learning model includes determining a training error signal based on an estimate of the training feedback path transfer function and the training feedback path transfer function. In one or more exemplary methods, determining the training error signal includes determining a loss function (e.g., a cost function) based on the estimate of the training feedback path transfer function and the training feedback path transfer function. For example, the loss function may quantify the difference between an estimate of the training feedback path transfer function (e.g., output by the machine learning model) and the training feedback path transfer function. The training feedback path transfer function (e.g., contained in the target data) may be considered as a reference (e.g., target) feedback path transfer function. In one or more exemplary methods, the training error signal may represent the training loss associated with the machine learning model. For example, minimizing this training loss (e.g., reducing the training error signal) may indicate that the estimate of the training feedback path transfer function is reasonable (e.g., meets requirements, is sufficiently accurate). In one or more exemplary methods, the loss function may be one or more of the following: mean squared error (MSE), mean absolute error (MAE), binary cross-entropy (BCE) loss function, normalized Euclidean systematic distance (NESD), and any other applicable loss function. For example, an arbitrary distance metric between the estimated value of the training feedback path transfer function and the training feedback path transfer function can be used as a loss function.
[0097] In one or more exemplary methods, each training iteration in multiple training iterations may include: determining an estimate of the feedback input signal component based on an estimate of the training feedback path transfer function. In one or more examples, each training iteration in multiple training iterations may include: determining a training feedback correction input signal (e.g., by subtracting the estimate of the feedback input signal component from the training input signal) based on the training input signal and the estimated feedback input signal component. In one or more exemplary methods, each training iteration in multiple training iterations may include: determining a target feedback correction input signal based on the training input signal and the training feedback input signal component. For example, since the training feedback path transfer function is known, the training feedback input signal component can also be determined. For example, the training feedback input signal component can be determined based on the training feedback path transfer function. For example, updating the machine learning model may include: determining a training error signal based on the training feedback correction input signal and the target feedback correction input signal. For example, a loss function (e.g., a cost function) can be determined based on the training feedback correction input signal and the target feedback correction input signal.
[0098] In one or more exemplary methods, for each frame, the training loss associated with the NESD loss function can be determined as follows:
[0099]
[0100] in, h represents the estimated value of the training feedback path transfer function (e.g., the fourth, fifth, sixth, or seventh machine learning processing signal). m This represents the training feedback path transfer function (e.g., contained in the target data). In other words, determining the training error signal (e.g., training loss) may include calculating the normalized Euclidean systematic distance (NESD) between the training feedback path transfer function and an estimate of the training feedback path transfer function.
[0101] In one or more exemplary methods, the training loss (e.g., total training loss) associated with the NESD loss function can be determined as:
[0102]
[0103] The total training loss is calculated for each frame (e.g., excluding the first N-1 frames) in the entire sequence. m The training loss is obtained by averaging a series of training losses (e.g., each training loss is associated with a frame).
[0104] In one or more exemplary methods, updating a machine learning model includes updating multiple weights of the machine learning model based on a training error signal and employing a learning rule. For example, multiple weights of the machine learning model may be updated (e.g., adjusted) when the training loss is minimized. For example, multiple weights of the machine learning model may not be updated (e.g., not adjusted) when the training loss is not minimized. The updated (e.g., adjusted) weights may be stored in memory associated with the machine learning model.
[0105] In one or more exemplary methods, updating multiple weights of a machine learning model using a learning rule includes: adjusting the multiple weights when it is determined that the training loss is less than or equal to a training loss threshold. In one or more exemplary methods, not updating multiple weights of a machine learning model using a learning rule includes: not adjusting the multiple weights when it is determined that the training loss is greater than a training loss threshold.
[0106] For example, the method includes performing a pre-training process and a fine-tuning process. The training phase may include a pre-training process and a fine-tuning process. For example, performing the pre-training process includes executing the training method (e.g., the steps described above) during the first set of training iterations in a series of training iterations. For example, the machine learning model is a pre-trained machine learning model.
[0107] For example, in each training iteration of the first set of training iterations, target data associated with the pre-training process is acquired. For example, the target data associated with the pre-training process (e.g., target data used to pre-train a machine learning model) may contain a feedback path transfer function determined through computer simulation (e.g., a synthetically generated feedback path transfer function). In other words, the training feedback path transfer function may be a synthetically generated feedback path transfer function. For example, the synthetically generated feedback path transfer function can be considered as a training sequence.
[0108] For example, in each training iteration of the first set of training iterations, training data associated with the pre-training process is acquired. For example, the training data associated with the pre-training process (e.g., training data used to pre-train a machine learning model) may include training input signals (e.g., including external input signal components and feedback input signal components) and a post-training processed signal. For example, the external input signal component of the training input signal associated with the pre-training process represents sound in an acoustic environment simulated by computer simulation. For example, the feedback input signal component of the training input signal associated with the pre-training process represents synthetically generated acoustic and / or mechanical feedback (e.g., generated by computer simulation). In other words, the generation of this feedback input signal component may be based on a synthetically generated feedback path transfer function. The training data may be determined based on the synthetically generated feedback path transfer function (e.g., impulse response). For example, each training iteration of the first set of training iterations may include generating target data containing a synthetic feedback path transfer function through computer simulation. For example, each training iteration of the first set of training iterations may include generating training data based on a synthetically generated feedback path transfer function. For example, the training input signal associated with the pre-training process and the post-training processed signal associated with the pre-training process may be considered as a training sequence.
[0109] For example, performing the fine-tuning process includes executing the training method (e.g., the steps described above) during the second set of training iterations in a series of training iterations. For example, the machine learning model is a fine-tuned machine learning model and is ready to be deployed in a hearing aid for the inference phase.
[0110] For example, in each training iteration of the second set of training iterations, target data associated with the fine-tuning process is acquired. For example, the target data associated with the fine-tuning process (e.g., target data used to fine-tune a machine learning model) may include a measured feedback path transfer function, such as one relevant to a real-world environment (e.g., the feedback path transfer function measured when or after the hearing aid is used in a real-world user scenario). In other words, the training feedback path transfer function can be a measured feedback path transfer function. For example, the measured feedback path transfer function can be considered as a training sequence.
[0111] For example, in each training iteration of the second set of training iterations, training data associated with the fine-tuning process is acquired. For example, the training data associated with the fine-tuning process (e.g., training data used to fine-tune a machine learning model) may include training input signals (e.g., including external input signal components and feedback input signal components) and the trained signal. For example, the external input signal component of the training input signal associated with the fine-tuning process represents sound in an acoustic environment simulated by computer simulation, or sound in the acoustic environment (e.g., a real environment) in which the user uses the hearing aid. For example, the feedback input signal component of the training input signal associated with the fine-tuning process represents measured acoustic and / or mechanical feedback (e.g., from a real-world scenario). In other words, the generation of the feedback input signal component may be based on a measured feedback path transfer function. The training data may be determined based on a measured feedback path transfer function (e.g., impulse response). For example, each iteration of the second set of training iterations may include generating target data containing a measured feedback path transfer function through computer simulation. For example, each iteration of the second set of training iterations may include generating training data based on a measured feedback path transfer function. For example, the training input signal associated with the fine-tuning process and the training post-processing signal associated with the fine-tuning process can be regarded as a training sequence.
[0112] For example, the method includes training a machine learning model using measured feedback path transfer functions (e.g., impulse responses of each feedback path associated with a real acoustic environment) and synthesized feedback path transfer functions (e.g., impulse responses of each feedback path associated with a simulated acoustic environment). Optionally, the method may also train the machine learning model using only measured feedback path transfer functions or only synthesized feedback path transfer functions. The method may include training the machine learning model using training data and target data generated through computer simulation. The method may include training the machine learning model using training data and target data measured from a real acoustic environment. The method may include training the machine learning model using mixed data of training data and target data measured from a real acoustic environment and training data and target data generated from a simulated acoustic environment.
[0113] In one or more exemplary methods, the method may include performing a verification process. The training phase may include a pre-training process, a fine-tuning process, and a verification process. For example, performing the verification process includes executing the training method (e.g., the steps described above) during a third set of training iterations in a plurality of training iterations. For example, the machine learning model is a fine-tuned machine learning model.
[0114] For example, in each training iteration of the third set of training iterations, target data associated with the validation process is acquired. This target data associated with the validation process can be called validation target data. For example, validation target data (e.g., used to validate the fine-tuned machine learning model) can contain either a synthetically generated feedback path transfer function or a measured feedback path transfer function. For example, a synthetically generated feedback path transfer function can be considered a validation sequence. For example, a measured feedback path transfer function can be considered a validation sequence.
[0115] For example, in each training iteration of the third set of training iterations, training data associated with the validation process is acquired. This training data associated with the validation process can be referred to as validation data. For example, validation data (e.g., used to validate a machine learning model) can include training input signals and post-training signals. The training input signals associated with the validation process can be referred to as validation input signals, and the two terms are used interchangeably. The post-training signals associated with the validation process can be referred to as validation post-processing signals, and the two terms are used interchangeably.
[0116] For example, the verification input signal includes an external input signal component and a feedback input signal component. The external input signal component associated with the verification process can be called the verification external input signal component, and the two terms are used interchangeably. The feedback input signal component associated with the verification process can be called the verification feedback input signal component, and the two terms are used interchangeably.
[0117] For example, the external input signal component for verification represents sound in an acoustic environment simulated by a computer, or sound in the acoustic environment (e.g., a real-world environment) in which the user is using the hearing aid. For example, the feedback input signal component for verification represents measured acoustic and / or mechanical feedback (e.g., from a real-world scenario) or synthetically generated acoustic and / or mechanical feedback. For example, in each training iteration of the third set of training iterations, verification data can be generated based on a measured feedback path transfer function or a synthetically generated feedback path transfer function. For example, both the verification input signal and the processed verification signal can be considered as a verification sequence.
[0118] In one or more exemplary methods, the method may include generating a training dataset prior to performing multiple training iterations. For example, the training dataset may contain data from multiple computer simulations (e.g., data from multiple known simulation environments, such as data from multiple ideal feedback control systems) and data measured from a real environment. In other words, the training dataset may contain multiple training input signals and multiple trained processed signals. The training data to be used in each training iteration can be obtained from the generated training dataset. For example, the training dataset may contain a set of training sequences and a set of validation sequences.
[0119] In one or more exemplary methods, the method may include: generating a target dataset prior to performing multiple training iterations. For example, the target dataset may contain data from multiple computer simulations (e.g., data from multiple known simulation environments, such as data from multiple ideal feedback control systems). In other words, the target dataset may contain multiple training feedback path transfer functions (e.g., multiple experimentally tested training feedback path transfer functions and / or multiple synthetically generated training feedback path transfer functions). The target data to be used in each training iteration can be obtained from this generated target dataset. For example, the target dataset may contain a set of pre-trained sequences and a set of validation sequences.
[0120] hearing aids
[0121] A hearing aid includes an input unit, a signal processing unit, and an output unit.
[0122] The input unit is configured to provide an electrical input signal representing the sound in the environment in which the hearing aid user is located. The electrical input signal includes an external input signal component and a feedback input signal component. The external input signal component represents the sound in the environment in which the hearing aid is located. The feedback input signal component represents acoustic and / or mechanical feedback originating from the feedback path from the output unit of the hearing aid to the input unit of the hearing aid.
[0123] The signal processing unit is configured to provide a processed signal by applying one or more processing algorithms to the feedback correction input signal. The feedback correction input signal is a feedback-corrected version of the electrical input signal.
[0124] The output unit is configured to output an audible signal to the hearing aid user based on the processed signal.
[0125] The hearing aid (e.g., also) includes a feedback control system comprising a trained machine learning (ML) model. The feedback control system is configured to determine an estimate of a feedback input signal component based on an estimate of a feedback path transfer function. The feedback path transfer function represents the impulse response of the feedback path. The trained machine learning model is configured to provide an estimate of the feedback path transfer function based on the electrical input signal and the processed signal. The machine learning model is trained according to the methods disclosed herein. The feedback control system is further configured to determine a feedback correction input signal based on the electrical input signal and the estimated values of the feedback input signal components.
[0126] Therefore, an improved hearing aid can be provided.
[0127] Embodiments of the present invention can provide hearing aids with improved signal quality and signal intelligibility because the trained machine learning model is configured to determine (e.g., provide) an estimate of the training feedback path transfer function (e.g., the impulse response of the feedback path), thereby allowing post-processing of the estimate. Post-processing of the estimate helps eliminate artifacts caused by suboptimal predictions from the machine learning model, thus improving signal quality and signal intelligibility. In other words, the method provided by embodiments of the present invention trains a machine learning model to estimate the feedback impulse response, rather than directly estimating the feedback correction input signal (e.g., providing an alternative to conventional adaptive filtering techniques and existing machine learning techniques). The estimate of the training feedback path transfer function may contain defects caused by the aforementioned suboptimal predictions, which can be removed from the training feedback path transfer function during post-processing, or retained if these defects are negligible (e.g., have a minor impact).
[0128] In one or more example hearing aids, the hearing aid is configured to provide frequency-varying gain and / or level-varying compression and / or frequency shifting (with or without frequency compression) from one or more frequency ranges to one or more other frequency ranges to compensate for the user's hearing loss. For example, the aforementioned frequency-varying gain, and / or level-varying compression, and / or frequency shifting functions can be implemented through one or more processing algorithms, and / or gain functions that vary with frequency and / or level. For example, the signal processing unit is configured to enhance the feedback correction input signal and provide the processed signal.
[0129] In one or more example hearing aids, an output unit is configured to provide (e.g., generate) stimuli perceived by the user as acoustic signals (e.g., sounds) based on a processed signal. The output unit may include a vibrator of a bone conduction hearing aid. The output unit may include an output transducer. The output transducer may include a receiver (e.g., a speaker) configured to provide the stimulus as an acoustic signal to the user (e.g., in an acoustic (air conduction-based) hearing aid). The output transducer may include a vibrator for providing the stimulus as mechanical vibrations of the skull to the user (e.g., in a bone-attached or bone-anchored hearing aid). The output unit may (additionally or alternatively) include a transmitter (e.g., wireless) for transmitting sound picked up by the hearing aid (e.g., via a network, such as in telephone operation mode) to another device, such as a remote communication partner.
[0130] In one or more example hearing aids, the input unit may include an input converter (e.g., a microphone) configured to convert input sound (e.g., sound in the hearing aid user's environment) into an electrical input signal. The input unit may include a wireless receiver configured to receive a wireless signal representing sound in the hearing aid user's environment and provide an electrical input signal representing said sound.
[0131] In one or more example hearing aids, the wireless receiver and / or transmitter may be configured to receive and / or transmit electromagnetic signals in a radio frequency range (e.g., 3 kHz to 300 GHz). In one or more example hearing aids, the wireless receiver and / or transmitter may be configured to receive and / or transmit electromagnetic signals in an optical frequency range (e.g., infrared light 300 GHz to 430 THz or visible light such as 430 THz to 770 THz).
[0132] In one or more example hearing aids, the hearing aid may include a directional microphone system adapted to spatially filter sound from the environment to enhance a target sound source among multiple sound sources in the local environment of the hearing aid wearer. The directional system may be adapted to detect (e.g., adaptive detection) the direction from which a specific portion of the microphone signal originates. This can be achieved, for example, in a variety of different ways described in the prior art. In hearing aids, microphone array beamformers are commonly used to spatially attenuate background noise sources. Beamformers may include linearly constrained minimum variance (LCMV) beamformers. Many beamformer variations can be found in the literature. Minimum variance distortionless response (MVDR) beamformers are widely used in microphone array signal processing. Ideally, an MVDR beamformer keeps the signal from the target direction (also known as the line of sight) unchanged while attenuating sound signals from other directions to the greatest extent possible. A generalized sidelobe canceller (GSC) structure is an equivalent representation of an MVDR beamformer, offering computational and digital representation advantages over a direct implementation of the original form.
[0133] Most sound sources (except for the user's own voice) are relatively small compared to the size of a hearing aid, such as the distance d between the two microphones in a directional system. mic Located away from the user. Typical microphone distance in hearing aids is in the 10mm range. The minimum distance for sound sources of interest to the user (e.g., sound from the user's mouth or sound from the audio transmission device) is 0.1m (>10mm). mic At this minimum distance, the hearing aid (microphone) will be in the acoustic near field of the sound source, and the level difference of the sound signal incident on the corresponding microphone may be significant. Typical distances for communication partners are greater than 1m (>100d). mic The hearing aid (microphone) will be located in the acoustic far field of the sound source, and the level difference of the sound signal incident on the corresponding microphone is not significant. The arrival time difference of the sound incident along the microphone axis (e.g., in front of or behind a normal hearing aid) is ΔT = d. mic / v sound =0.01 / 343[s] = 29μs, where v sound The speed of sound in air at 20°C is 343 m / s.
[0134] Hearing aids may include antenna and transceiver circuitry that enables the establishment of wireless links to entertainment devices (e.g., televisions), communication devices (e.g., telephones), wireless microphones, separate (external) processing devices, or other hearing aids. The hearing aid can thus be configured to wirelessly receive direct electrical input signals from another device. Similarly, the hearing aid can be configured to wirelessly transmit direct electrical output signals to another device. The direct electrical input or output signals may represent or include audio signals and / or control signals and / or information signals.
[0135] Generally, the wireless link established by the antenna and transceiver circuitry of a hearing aid can be of any type. The wireless link can be a near-field communication-based link, such as an inductive link based on inductive coupling between the antenna coils of the transmitter and receiver sections. The wireless link can also be based on far-field electromagnetic radiation. Preferably, the frequency used to establish the communication link between the hearing aid and another device is below 70 GHz, for example, in the range from 50 MHz to 70 GHz, or above 300 MHz, for example, in the ISM range above 300 MHz, or in the 900 MHz range, or in the 2.4 GHz range, or in the 5.8 GHz range, or in the 60 GHz range (ISM = Industrial, Scientific and Medical, such standardized ranges are defined, for example, by the International Telecommunication Union ITU). The wireless link can be based on standardized or proprietary technologies. The wireless link can be based on Bluetooth technology (e.g., Bluetooth Low Energy technology, such as LE Audio) or Ultra Wideband (UWB) technology.
[0136] Hearing aids may be constituted by or be part of a portable (i.e., configured to be wearable) device, such as a device that includes a local power source, such as a battery, for example a rechargeable battery.
[0137] Hearing aids can be, for example, low-weight, easy-to-wear devices, with a total weight of less than 100g, less than 20g, or less than 5g.
[0138] A hearing aid may include a "forward" (or "signal") path between the input and output of the hearing aid for processing audio signals. A signal processing unit may be located in this forward path. The signal processing unit may be configured to provide frequency-varying gain according to the user's specific needs (e.g., hearing loss). The hearing aid may include an "analysis" path having functionalities for analyzing signals and / or controlling the processing of the forward path. Some or all of the signal processing in the analysis path and / or the forward path may be performed in the frequency domain, in which case the hearing aid includes appropriate analysis and synthesis filter banks. Some or all of the signal processing in the analysis path and / or the forward path may be performed in the time domain.
[0139] Analog electrical signals representing sound signals can be converted into digital audio signals during analog-to-digital (AD) conversion, where the analog signal is sampled at a predetermined sampling frequency or sampling rate f.s Perform sampling, f s For example, in the range from 8kHz to 48kHz (to suit specific application needs) at discrete time points t n (or n) provides digital samples x n (or x[n]), each audio sample passes through a predetermined N b Bit represents the acoustic signal at t n The value of N at time b For example, in the range of 1 to 48 bits, such as 24 bits. Each audio sample therefore uses N. b Bit quantization (resulting in 2^n voltammetry of audio samples) Nb (Number of different possible values). The numerical sample x has 1 / f s The duration of the time, such as 50 μs, for f s =20kHz. Multiple audio samples can be arranged in time frames. A time frame can include 64 or 128 audio data samples. Other frame lengths can be used depending on the application.
[0140] Hearing aids may include analog-to-digital (AD) converters to digitize analog inputs (e.g., from an input converter such as a microphone) at a predetermined sampling rate such as 20 kHz. Hearing aids may also include digital-to-analog (DA) converters to convert digital signals into analog output signals, for example, for presentation to the user via an output converter.
[0141] Hearing aids, such as input units and / or antenna and transceiver circuitry, may include transformation units for converting time-domain signals into signals in a transform domain (e.g., frequency domain or Laplace domain, Z-transform, wavelet transform, etc.). The transformation unit may constitute or include a time-frequency (TF) conversion unit for providing a time-frequency representation of the input signal. The time-frequency representation may include an array or mapping of corresponding complex or real values of the signal involved over a specific time and frequency range. The TF conversion unit may include a filter bank for filtering the (time-varying) input signal and providing multiple (time-varying) output signals, each output signal comprising a distinctly different frequency range of the input signal. The TF conversion unit may include a Fourier transform unit (e.g., a Discrete Fourier Transform (DFT) algorithm, a Short-Time Fourier Transform (STFT) algorithm, or a similar algorithm) for converting the time-varying input signal into a (time-)frequency signal. The hearing aid considers a frequency range from the minimum frequency f. min up to the maximum frequency f max The frequency range can include a portion of the typical human hearing range from 20Hz to 20kHz, such as a portion of the range from 20Hz to 12kHz. Typically, the sampling rate f... s Greater than or equal to the maximum frequency f max twice, that is, f s ≥2f maxThe signals from the forward and / or analytical pathways of the hearing aid can be divided into NI (e.g., uniformly wide) frequency bands, where NI is, for example, greater than 5, greater than 10, greater than 50, greater than 100, or greater than 500, and at least some of these bands are processed individually. The hearing aid may be adapted to process the signals from the forward and / or analytical pathways (NP≤NI) on NP different channels. The channels may have consistent or inconsistent widths (e.g., width increases with frequency), and may overlap or not overlap.
[0142] In one or more example hearing aids, the hearing aid includes an analysis filter bank (e.g., at least one analysis filter bank) configured to provide an electrical input signal (e.g., at least one electrical input signal) in a time-frequency domain representation. Signal processing along the forward path from (e.g., at least one) input transducer to the output transducer can be performed in the time-frequency domain (m,k), where m represents a time (frame) index and k represents a frequency bin (e.g., index). The analysis filter bank may incorporate Fourier transform-based techniques, such as the Short-Time Fourier Transform (STFT) algorithm.
[0143] Hearing aids can be configured to operate in different modes, such as a normal mode and one or more specific modes, which may be user-selectable or automatically selected. Operating modes can be optimized for specific acoustic conditions or environments, such as communication modes like telephone modes. Operating modes may include low-power modes, where the hearing aid's functionality is reduced (e.g., for energy saving), such as disabling wireless communication and / or disabling specific features of the hearing aid.
[0144] Hearing aids may include multiple detectors configured to provide status signals relating to the hearing aid's current network environment (such as the current acoustic environment), and / or the current state of the user wearing the hearing aid, and / or the current state or operating mode of the hearing aid. Alternatively or additionally, one or more detectors may form part of an external device that communicates with the hearing aid (e.g., wirelessly). External devices may include, for example, another hearing aid, a remote control, an audio transmission device, a telephone (e.g., a smartphone), external sensors, etc.
[0145] One or more of a plurality of detectors can operate on a full-band signal (time domain). One or more of a plurality of detectors can operate on a band-split signal ((time-)frequency domain), for example, in a finite number of frequency bands.
[0146] Multiple detectors may include level detectors for estimating the current level of the signal in the forward path. Detectors can be configured to determine whether the current level of the signal in the forward path is above or below a given (L-) threshold. Level detectors operate on full-band signals (time domain). Level detectors operate on band-split signals ((time-)frequency domain).
[0147] Hearing aids may include a voice activity detector (VAD) for estimating whether (or with what probability) an input signal (at a given point in time) includes a voice signal. In this specification, a voice signal may be understood to include speech signals from humans. It may also include other forms of vocalization produced by the human speech system (such as singing). The voice activity detector unit may be adapted to classify the user's current acoustic environment as a "voice" or "no-voice" environment. This has the advantage that time periods including electrophonic signals of human vocalizations (such as speech) in the user's environment can be identified and thus separated from time periods that include only (or primarily) other sound sources (such as artificially generated noise). The voice activity detector may be adapted to also detect the user's own voice as "voice." Alternatively, the voice activity detector may be adapted to exclude the user's own voice from the detection of "voice."
[0148] Hearing aids may include a self-voice detector for estimating whether (or with what probability) a particular input sound (such as speech) originates from the user's voice. The microphone system of the hearing aid may be adapted to distinguish the user's own voice from another person's voice and possibly from non-voice sounds.
[0149] Multiple detectors may include motion detectors such as accelerometers. Motion detectors may be configured to detect movements of a user’s facial muscles and / or bones, such as those caused by speech or chewing (e.g., jaw movements), and provide detector signals indicating those movements.
[0150] The hearing aid may include a classification unit configured to classify the current situation based on input signals from (at least partially) a detector and possibly other inputs. In this specification, "current situation" may be defined by one or more of the following:
[0151] a) Physical environment (including the current electromagnetic environment, such as the presence of electromagnetic signals (including audio and / or control signals) that are planned or unplanned to be received by the hearing aid, or other properties of the current environment that are different from acoustics);
[0152] b) Current acoustic conditions (input level, feedback, etc.);
[0153] c) The user's current mode or state (movement, temperature, cognitive load, etc.);
[0154] d) The current mode or state of the hearing aid and / or another device communicating with the hearing aid (selected program, time elapsed since the last user interaction, etc.).
[0155] The classification unit may be based on or may include neural networks, such as recurrent neural networks, or trained neural networks.
[0156] A feedback control system can be viewed as acoustic (and / or mechanical) feedback control (such as suppression) or echo cancellation system. Acoustic (and / or mechanical) feedback originating from the auto-auditor feedback path can be considered as feedback sound (e.g., represented by feedback input signal components) generated by the output transducer (e.g., a loudspeaker) and leaked through the feedback path to the input transducer (e.g., a microphone).
[0157] Hearing aids may also include other suitable functions for the applications involved, such as compression and noise reduction.
[0158] Hearing aids may include hearing instruments, such as hearing instruments adapted to be located in the user's ear or wholly or partially in the ear canal.
[0159] use
[0160] On the one hand, the uses of the hearing aids described above, in detail in the Detailed Description section, are also provided. Use in systems comprising one or more hearing aids (e.g., hearing instruments) is also possible.
[0161] Computer-readable media or data carrier
[0162] The present invention further provides a tangible computer-readable medium (data carrier) storing a computer program including program code (instructions), which, when the computer program is run on a data processing system (computer), causes the data processing system to perform (implement) at least some (such as most or all) of the steps of the methods described above, in detail in the "Detailed Description" and as defined in the claims.
[0163] By way of example, but not limitation, the aforementioned tangible computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to execute or store required program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks include compact discs (CDs), laser discs, optical discs, digital multipurpose discs (DVDs), floppy disks, and Blu-ray discs, wherein these disks typically magnetically copy data while simultaneously being optically copied using lasers. Other storage media include those stored in DNA (e.g., in synthetic DNA strands). Combinations of the aforementioned disks should also be included within the scope of computer-readable media. In addition to being stored on tangible media, computer programs may also be transmitted via transmission media such as wired or wireless links or networks such as the Internet and loaded into data processing systems to run at locations other than tangible media.
[0164] Computer program
[0165] In addition, this application provides a computer program (product) including instructions that, when run by a computer, cause the computer to perform the steps of the methods (methods) described above, in detail in the "Detailed Description" section, and as defined in the claims.
[0166] Data processing system
[0167] In one aspect, the present invention further provides a data processing system, including a processor and program code, the program code causing the processor to perform at least some (such as most or all) of the steps of the methods described above, in detail in the "Detailed Description" section, and as defined in the claims.
[0168] Hearing system
[0169] On the other hand, hearing aids and hearing systems including assistive devices are provided, including those described above, described in detail in the "Detailed Description" section, and defined in the claims.
[0170] Hearing systems can be adapted to establish a communication link between hearing aids and assistive devices so that information (such as control and status signals, possibly audio signals) can be exchanged or forwarded from one device to another.
[0171] The auxiliary device may include a remote control, a smartphone, or other portable or wearable electronic device such as a smartwatch or may be composed of such devices.
[0172] The assistive device may consist of or include a remote control for controlling the functions and operation of the hearing aid. The remote control functionality is implemented in a smartphone, which may run an app that enables control of the audio processing device via the smartphone (the hearing aid includes a suitable wireless interface to the smartphone, such as Bluetooth or some other standardized or proprietary solution).
[0173] The assistive device may be constituted by or include an audio gateway device, which is adapted to receive multiple audio signals (e.g., from entertainment devices such as TVs or music players, from telephone devices such as mobile phones, or from computers such as PCs, wireless microphones, etc.) and to select and / or combine appropriate signals (or combinations of signals) from the received audio signals to transmit to the hearing aid.
[0174] The assistive device may be composed of or may include another hearing aid. The hearing system may include two hearing aids adapted to implement a binaural hearing system, such as a binaural hearing aid system.
[0175] APP
[0176] On the other hand, the present invention also provides a non-transitory application called an APP. The APP includes executable instructions configured to run on an assistive device to implement a user interface for the hearing aid or hearing system described above, in detail in the "Detailed Description," and as defined in the claims. The APP can be configured to run on a mobile phone, such as a smartphone, or another portable device enabled to communicate with said hearing device or hearing system.
[0177] definition
[0178] In this specification, a hearing aid, such as a hearing instrument, refers to a device suitable for improving, enhancing, and / or protecting a user's hearing ability, which achieves this by receiving sound signals from the user's environment, generating corresponding audio signals, possibly modifying the audio signals, and providing the possibly modified audio signals as audible signals to at least one ear of the user. The audible signals may be provided, for example, as sound signals radiated into the user's outer ear, and / or as sound signals transmitted as mechanical vibrations through the bone structures of the user's head and / or through portions of the middle ear to the user's inner ear.
[0179] Hearing aids can be configured to be worn in any known manner, such as as a unit worn behind the ear (having a tube that directs radiated sound signals into the ear canal or having an output transducer, such as a speaker, arranged close to or within the ear canal), as a unit wholly or partially arranged in the auricle and / or ear canal, or as a unit connected to a fixed structure implanted in the skull, such as a vibrator. Hearing aids may include a single unit or several units that communicate with each other (e.g., acoustically, electrically, or optically). The speaker may be housed within the housing along with other components of the hearing aid, or it may be an external unit (possibly combined with a flexible guiding element such as a dome-shaped element).
[0180] Hearing aids can be adapted to the specific needs of users, such as those with hearing loss. The configurable signal processing circuitry of a hearing aid can be adapted to apply frequency- and level-variable compression and amplification of the input signal. Customized frequency- and level-variable gain (amplification or compression) can be determined during the fitting process by the fitting system based on the user's hearing data, such as an audiogram, using basic fitting principles (e.g., speech adaptation). This frequency- and level-variable gain can be reflected, for example, in processing parameters, uploaded to the hearing aid via an interface to a programming device (fitting system), and used by a processing algorithm executed by the hearing aid's configurable signal processing circuitry.
[0181] A “hearing system” refers to a system that includes one or two hearing aids. A “binaural hearing system” refers to a system that includes two hearing aids and is adapted to work together to provide audible signals to both of a user’s ears. A hearing system or a binaural hearing system may also include one or more “assistive devices” that communicate with the hearing aids and influence and / or benefit from the functionality of the hearing aids. The aforementioned assistive devices may include at least one of the following: a remote control, a remote microphone, an audio gateway device, an entertainment device such as a music player, a wireless communication device such as a mobile phone (e.g., a smartphone), or a tablet computer, or another device, such as one that includes a graphical interface. Hearing aids, hearing systems, or binaural hearing systems may be used, for example, to compensate for hearing loss in persons with hearing impairments, enhance or protect the hearing ability of persons with normal hearing, and / or transmit electronic audio signals to persons. Hearing aids or hearing systems may, for example, be part of or interact with broadcasting systems, active ear protection systems, hands-free telephone systems, car audio systems, entertainment (e.g., television, music playback, or karaoke) systems, teleconferencing systems, classroom amplification systems, etc. Attached Figure Description
[0182] Various aspects of the invention will be best understood from the following detailed description taken in conjunction with the accompanying drawings. For clarity, these drawings are schematic and simplified, showing only the details necessary for understanding the invention while omitting other details. Throughout the specification, the same reference numerals are used for the same or corresponding parts. Features of each aspect may be combined with any or all features of other aspects. These and other aspects, features, and / or technical effects will be apparent from and illustrated in the following figures, wherein:
[0183] Figure 1 An exemplary hearing aid comprising a feedback cancellation system employing an adaptive filter is schematically shown;
[0184] Figure 2 An exemplary hearing aid according to the present invention is illustrated schematically;
[0185] Figure 3 A flowchart is shown of an exemplary method, performed by a hearing aid, for determining an estimate of the feedback path transfer function according to the present invention;
[0186] Figures 4A-4B An exemplary training structure for a machine learning model according to the present invention is illustrated schematically;
[0187] Figure 5 A flowchart is shown of an exemplary method for training a machine learning (ML) model for a hearing aid feedback control system according to the present invention;
[0188] Figure 6 A block diagram of an exemplary electronic device according to the present invention is shown;
[0189] Figure 7 An exemplary impulse response of the hearing aid feedback path according to the present invention is shown;
[0190] Figure 8 A table showing exemplary configurations of the layers of a machine learning model according to the present invention is provided;
[0191] Figure 9 A graph illustrating an exemplary training loss according to the present invention is shown;
[0192] Figure 10 An exemplary table of mean and standard deviation is shown for a machine learning-based feedback control system comprising a trained machine learning model according to the present invention.
[0193] The further applicability of the invention will become apparent from the detailed description given below. However, it should be understood that while the detailed description and specific examples illustrate preferred embodiments of the invention, they are given for illustrative purposes only. Other embodiments of the invention will become apparent to those skilled in the art based on the following detailed description. Detailed Implementation
[0194] The detailed description below, taken in conjunction with the accompanying drawings, serves as a description of various different configurations. This detailed description includes specific details to provide a thorough understanding of several different concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. Several aspects of the apparatus and method are described by various different blocks, functional units, modules, elements, circuits, steps, processes, algorithms, etc. (collectively, “elements”). Depending on the specific application, design constraints, or other reasons, these elements may be implemented using electronic hardware, computer programs, or any combination thereof.
[0195] Electronic hardware may include microelectromechanical systems (MEMS), (e.g., application-specific integrated circuits), microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), gating logic, discrete hardware circuits, printed circuit boards (PCBs) (e.g., flexible PCBs), and other suitable hardware configured to perform the various functions described in this specification, such as sensors for sensing and / or recording the physical properties of the environment, devices, users, etc. Computer programs should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, programs, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description languages, or other names.
[0196] Figure 1An exemplary hearing aid 300 according to the present invention is illustrated schematically. The hearing aid 300 includes a feedback control system 308 (e.g., a feedback cancellation system). Figure 1 A feedback cancellation system 308 employing an adaptive filter is shown.
[0197] The hearing aid 300 includes a forward path and a feedback path 302. For example, the feedback path 302 may contain an impulse response represented by a feedback path transfer function 302B (e.g., denoted as h(n) = [h1(n), h2(n), ..., h...). L (n)] T Where L represents the length of the feedback path impulse response, and n represents (e.g., discrete) time index. For example, the feedback path transfer function 302B can be regarded as a time-varying transfer function of the feedback path.
[0198] The forward path includes an input unit (e.g., an input transducer 304, such as a microphone) configured to provide (e.g., pick up) an electrical input signal 304A representing the sound in the environment where the hearing aid 300 user is located. In other words, the microphone (e.g., the input transducer) can be configured to acquire (e.g., pick up) sound from the environment where the hearing aid 300 is located and provide an electrical input signal 304A representing that sound.
[0199] The electrical input signal 304A (e.g., y(n) = x(n) + v(n)) includes an external input signal component 301 (e.g., x(n)) and a feedback input signal component 302A (e.g., v(n), where n represents a time index). The external input signal component 301 represents the sound in the environment in which the hearing aid 300 is located. The feedback input signal component 302A represents acoustic and / or mechanical feedback originating from a feedback path 302, which is a path from the output unit (e.g., output transducer 310) of the hearing aid 300 to the input unit (e.g., input transducer 304).
[0200] The forward path includes a signal processing unit 306 configured to generate a processed signal 306A by applying one or more processing algorithms 306 to the feedback correction input signal 305A. Optionally, the signal processing unit 306 may be configured to generate the processed signal 306A by applying a gain function (e.g., g(n), where n represents the time index) that varies with frequency and / or level to the feedback correction input signal 305A (e.g., denoted as e(n), where n represents the time index). For example, the gain function that varies with frequency and / or level can be considered as the time-varying transfer function of the forward path. The gain function that varies with frequency and / or level can represent the impulse response of the forward path of the hearing aid 300.
[0201] In one or more example hearing aids, both the gain function and the feedback path transfer function 302B, which vary with frequency and / or level, are vectors. Each of the gain function and the feedback path transfer function 302B may contain multiple elements (e.g., L elements), each element representing the response to external changes over time (e.g., the response to an impulse). For example, each of the multiple elements can be considered as a sampled value of the impulse response at time index n.
[0202] For example, in dynamic feedback scenarios (e.g., when the feedback path changes dynamically), the impulse responses of the gain function and the feedback path transfer function 302B, which vary with frequency and / or level, may change over time (e.g., making the impulse responses dependent on the time index n). Conversely, in static feedback scenarios, the impulse responses of the gain function and the feedback path transfer function 302B, which vary with frequency and / or level, may remain constant, meaning that neither of their impulse responses depends on the time index n.
[0203] The forward path includes an output unit configured to output an audible signal 310A to the hearing aid user based on the processed signal 306A. The processed signal 306A can be converted into an acoustic signal by an output transducer 310 (e.g., a loudspeaker). In other words, the output transducer 310 can be configured to provide a stimulus signal that the user can perceive as sound based on the processed signal 306A.
[0204] In one or more example hearing aids, the feedback control system 308 may be implemented as an adaptive filter. For example, the adaptive filter is configured to model a real and practically unknown feedback path (e.g., feedback path 302 represented by feedback path transfer function 302B). The feedback control system 308 is configured to determine an estimate 308BA of the feedback input signal component 302A based on an estimate 308AA of the feedback path transfer function 302B. For example, the estimate 308BA of the feedback input signal component 302A may represent (e.g., include) the estimate 308AA of the feedback path transfer function 302B. For example, the feedback control system 308 is configured to estimate the impulse response of the feedback path 302. The estimate 308AA of the feedback path transfer function 302B may be represented as... Where L represents the length of the feedback path impulse response, and n represents (e.g., discrete) time index.
[0205] For example, the feedback control system 308 includes an adaptive algorithm 308A (e.g., a least mean square (LMS) estimation algorithm) and a variable filter 308B (e.g., a time-varying filter). For example, the adaptive algorithm 308A is configured to determine an estimate 308AA of the feedback path transfer function 302B based on the feedback correction input signal 305A and the processed signal 306A. For example, the adaptive algorithm 308A is configured to provide the estimate 308AA of the feedback path transfer function 302B to the variable filter 308B. For example, the adaptive algorithm 308A is configured to use multiple filter coefficients (e.g., ... The estimated value 308AA of the feedback path transfer function 302B is provided in the form of L filter coefficients. For example, the adaptive algorithm 308A is configured to repeatedly determine the estimated value 308AA of the feedback path transfer function 302B (e.g., multiple filter coefficients), for example, during multiple feedback changes that may occur during the user's use of the hearing aid 300.
[0206] For example, the variable filter 308B is configured to provide an estimate 308BA of the feedback input signal component 302A based on the estimate 308AA of the feedback path transfer function 302B and the processed signal 306A. In other words, the variable filter 308B can be configured to determine the estimate 308BA of the feedback input signal component 302A by filtering the processed signal 306A using multiple filter coefficients.
[0207] In one or more example hearing aids, a feedback control system 308 is configured to determine a feedback correction input signal 305A based on an estimate 308BA of an electrical input signal 304A and a feedback input signal component 302A. In other words, the feedback control system may include a combining unit 305 (e.g., a summing unit) configured to combine the electrical input signal 304A and the estimate 308BA of the feedback input signal component 302A. For example, the combining unit 305 is configured to determine the feedback correction input signal 305A (e.g., by subtracting the estimate 308BA of the feedback input signal component 302A from the electrical input signal 304A.) The feedback correction input signal 305A (e.g., e(n)) can be regarded as an estimate of the external input signal component 301 (e.g., x(n)).
[0208] For example, the feedback control system 308 is implemented as an adaptive filter, such as using conventional adaptive filtering techniques. For instance, the adaptive algorithm 308A can use a step size parameter to determine (e.g., continuously update) an estimate 308BB of the feedback path transfer function 302B. Choosing this step size parameter can be a challenging task (e.g., requiring in-depth and comprehensive analysis) because it controls the convergence (e.g., convergence speed), robustness, and steady-state performance (e.g., estimation error in steady state) of the adaptive algorithm 308A.
[0209] In other words, while traditional adaptive filtering techniques can determine an estimate of the feedback transfer function, they may not be able to accurately model the acoustic and / or mechanical feedback paths, nor can they respond as quickly as required by real-world environments (e.g., actual scenarios) to changes in the acoustic feedback path (e.g., response times need to be less than a few hundred milliseconds to maintain hearing aid stability). The limitation of traditional adaptive filtering techniques lies in the difficulty of ensuring an ideal balance between convergence speed and steady-state performance, especially in real-world environments.
[0210] Embodiments of the present invention provide a machine learning-based feedback control system for predicting (e.g., estimating) feedback impulse responses, thus offering an alternative to conventional adaptive filtering techniques. Compared to conventional adaptive filtering techniques, the machine learning-based feedback control system offers significant advantages: it achieves faster convergence and lower steady-state error, thereby improving the balance between convergence speed and steady-state error. In other words, embodiments of the present invention can estimate the feedback path transfer function more quickly and robustly (e.g., more accurately), enabling reliable application of this machine learning-based feedback control system in real-world environments. For example, the machine learning-based feedback control system provided by the present invention does not require step-size control to adjust convergence, robustness, and steady-state error, as is done in conventional adaptive filtering techniques.
[0211] Figure 2 An exemplary hearing aid 400 according to the present invention is illustrated schematically. The hearing aid 400 includes a feedback control system 408 (e.g., a feedback cancellation system). Figure 2 A feedback cancellation system 408 employing a trained machine learning model 408A is illustrated. The feedback cancellation system 408 includes the trained machine learning model 408A. For example, the feedback control system 408 is a machine learning-based feedback control system. For example, the trained machine learning model can replace... Figure 1 The adaptive algorithm 308A in the text.
[0212] The hearing aid 400 includes a forward path and a feedback path 402. For example, the feedback path 402 may contain an impulse response represented by a feedback path transfer function 402B (e.g., denoted as h(n) = [h1(n), h2(n), ..., h...). L (n)] T Where L represents the length of the feedback path impulse response, and n represents (e.g., discrete) time index. For example, the feedback path transfer function 402B can be regarded as a time-varying transfer function of the feedback path.
[0213] The forward path includes an input unit (e.g., an input transducer 404, such as a microphone) configured to provide (e.g., pick up) an electrical input signal 404A representing the sound in the environment where the hearing aid 400 user is located. In other words, the microphone (e.g., the input transducer) may be configured to acquire (e.g., pick up) sound from the environment where the hearing aid 400 is located and provide an electrical input signal 404A representing that sound.
[0214] The electrical input signal 404A (e.g., y(n) = x(n) + v(n)) includes an external input signal component 401 (e.g., x(n)) and a feedback input signal component 402A (e.g., v(n), where n represents a time index). The external input signal component 401 represents the sound in the environment in which the hearing aid 400 is located. The feedback input signal component 402A represents acoustic and / or mechanical feedback originating from the feedback path 402, which is the path from the output unit (e.g., output transducer 410) of the hearing aid 400 to the input unit (e.g., input transducer 404).
[0215] The forward path includes a signal processing unit 406 configured to generate a processed signal 406A by applying one or more processing algorithms 406 to the feedback correction input signal 405A. Optionally, the signal processing unit 406 may be configured to generate the processed signal 406A by applying a gain function (e.g., g(n), where n represents the time index) that varies with frequency and / or level to the feedback correction input signal 405A (e.g., denoted as e(n), where n represents the time index). For example, the gain function that varies with frequency and / or level can be considered as a time-varying transfer function of the forward path. The gain function that varies with frequency and / or level can represent the impulse response of the forward path of the hearing aid 400.
[0216] In one or more example hearing aids, both the gain function and the feedback path transfer function 402B, which vary with frequency and / or level, are vectors. Each of the gain function and the feedback path transfer function 402B may contain multiple elements (e.g., L elements), each element representing the response to external changes over time (e.g., the response to an impulse). For example, each of the multiple elements can be considered as a sampled value of the impulse response at time index n.
[0217] For example, in dynamic feedback scenarios (e.g., where the feedback path changes dynamically), the impulse responses of the gain function and the feedback path transfer function 402B, which vary with frequency and / or level, may change over time (e.g., making the impulse responses dependent on the time index n). Conversely, in static feedback scenarios, the impulse responses of the gain function and the feedback path transfer function 402B, which vary with frequency and / or level, may remain constant, meaning that neither of their impulse responses depends on the time index n.
[0218] The forward path includes an output unit configured to output an audible signal 410A to the hearing aid user based on the processed signal 406A. The processed signal 406A can be converted into an acoustic signal by an output transducer 410 (e.g., a loudspeaker). In other words, the output transducer 410 can be configured to provide a stimulus signal that the user can perceive as sound based on the processed signal 406A.
[0219] The feedback control system 408 includes a trained machine learning model 408A. For example, the feedback control system 408 includes a trained machine learning model 408A and a machine learning filter 408B (e.g., a time-varying filter). The trained machine learning model 408A is configured to provide an estimate 408BB of the feedback path transfer function 402B based on the electrical input signal 404A and the processed signal 406A. In other words, the feedback control system 408 can be configured to estimate the impulse response of the feedback path 402.
[0220] The estimated value 408AA of the feedback path transfer function 402B can be expressed as: Where L represents the length of the feedback path impulse response, and n represents (e.g., discrete) time index.
[0221] The feedback control system 408 is configured to determine an estimate 408BA of the feedback input signal component 402A based on an estimate 408AA of the feedback path transfer function 402B. For example, the estimate 408BA of the feedback input signal component 402A may represent (e.g., include) the estimate 408AA of the feedback path transfer function 402B. The estimate 408AA of the feedback path transfer function 402B may be expressed as... Where L represents the length of the feedback path impulse response, and n represents (e.g., discrete) time index.
[0222] For example, a trained machine learning model 408A is configured to provide an estimate 408AA of the feedback path transfer function 402B to a variable filter 408B. For example, the trained machine learning model 408A is configured to use multiple filter coefficients (e.g., The estimated value 408AA of the feedback path transfer function 402B is provided in the form of L filter coefficients. For example, the trained machine learning model 408A is configured to repeatedly determine (e.g., inference) the estimated value 408AA of the feedback path transfer function 402B (e.g., multiple filter coefficients), for example, during multiple feedback changes that may occur during the user's use of the hearing aid 400.
[0223] For example, the variable filter 408B is configured to provide an estimate 408BA of the feedback input signal component 402A based on the estimate 408AA of the feedback path transfer function 402B and the processed signal 406A. In other words, the variable filter 408B can be configured to determine the estimate 408BA of the feedback input signal component 402A by filtering the processed signal 406A using multiple filter coefficients.
[0224] In one or more example hearing aids, a feedback control system 408 is configured to determine a feedback correction input signal 405A based on an estimate 408BA of an electrical input signal 404A and a feedback input signal component 402A. In other words, the feedback control system may include a combining unit 405 (e.g., a summing unit) configured to combine the electrical input signal 404A and the estimate 408BA of the feedback input signal component 402A. For example, the combining unit 405 is configured to determine the feedback correction input signal 405A (e.g., by subtracting the estimate 408BA of the feedback input signal component 402A from the electrical input signal 404A.) The feedback correction input signal 405A (e.g., e(n)) can be regarded as an estimate of the external input signal component 401 (e.g., x(n)).
[0225] For example, Figure 1 The feedback control system 308 can be considered as an adaptive filter that includes an adaptive algorithm and a variable filter, while Figure 2 The feedback control system 408 can be considered as a machine learning-based feedback control system comprising a trained machine learning model (e.g., trained machine learning model 408A) and a variable machine learning filter (e.g., variable machine learning filter 408B). For example, the present invention aims to replace adaptive algorithms (e.g., [missing information]) with a trained machine learning model (e.g., trained machine learning model 408A). Figure 1 An adaptive filter (308A) is used to improve the balance between convergence speed and steady-state error.
[0226] Embodiments of the present invention provide a hearing aid (e.g., hearing aid 400) that includes a machine learning-based feedback control system (e.g., including a trained machine learning model, such as trained machine learning model 408A). The machine learning model is based on... Figure 5Method 100 is used for training. The machine learning model (e.g., trained machine learning model 408A) comprises, in the following order: convolutional layers (e.g., ... Figure 4B The convolutional layer 820), the first fully connected (FC) layer (e.g. Figure 4B The first fully connected layer (824) and the Long Short-Term Memory (LSTM) layer (e.g.) Figure 4B The LSTM layer 826). For example, each layer of the machine learning model (e.g., the trained machine learning model 408A) may contain an activation function. Optionally, the machine learning model (e.g., the trained machine learning model 408A) may also contain pooling layers. For example, the machine learning model (e.g., the trained machine learning model 408A) includes, in the following order: convolutional layers (e.g., ... Figure 4B The convolutional layer 820), the first fully connected (FC) layer (e.g. Figure 4B The first fully connected layer 824), Long Short-Term Memory (LSTM) layer (e.g. Figure 4B LSTM layer 826), second fully connected layer (e.g. Figure 4B The second fully connected layer 828), the third fully connected layer (e.g. Figure 4B The third fully connected layer (830) and pooling layers (e.g.) Figure 4B Pooling layer 832).
[0227] This invention requires the machine learning model to be trained using the layers (e.g., and corresponding activation functions) in the above order, which is different from the adaptive algorithm in the adaptive filter (e.g. Figure 1 The implementation of the adaptive algorithm 308A differs significantly. For example, the implementation of the adaptive algorithm (e.g., Figure 1 The adaptive algorithm 308A requires selecting a step size parameter to determine the estimated value of the feedback transfer function, and in real-world environments, selecting this step size parameter is challenging.
[0228] Figure 3 A flowchart is shown of an exemplary method 600, performed by a hearing aid according to the present invention, for determining an estimate of a feedback path transfer function. The hearing aid is one disclosed herein, such as... Figure 2 The hearing aid 400.
[0229] Method 600 includes acquiring (S602) an electrical input signal representing the ambient sound of the hearing aid. The hearing aid may acquire the electrical input signal from its input unit.
[0230] The electrical input signal includes an external input signal component and a feedback input signal component. The external input signal component represents the sound in the environment in which the hearing aid is located. The feedback input signal component represents the acoustic and / or mechanical feedback originating from the feedback path, which is the feedback path from the hearing aid's output unit (e.g., the output transducer, such as...). Figure 2The output transducer 410) is connected to the input unit (e.g., the input transducer, such as... Figure 2 The path of the input transducer 404.
[0231] For example, the electrical input signal representing the ambient sound of the hearing aid can be considered as a signal from the actual acoustic environment, such as the acoustic environment in which the user of the hearing aid is located (e.g., the environment in which the hearing aid is being used). In other words, the hearing aid can operate in a normal operating mode. The hearing aid can perform method 600 when it is in normal operating mode.
[0232] Method 600 includes determining (S604) an estimate of the feedback path transfer function based on the electrical input signal and a first processed signal. For example, the first processed signal can be considered as the processed signal determined in the (r-1)th iteration of multiple iterations R. The first processed signal can be considered as a previously determined processed signal. The feedback path transfer function represents the impulse response of the feedback path. In other words, method 600 may include determining the impulse response of the feedback path (e.g., the feedback path at the r-th iteration).
[0233] Method 600 includes determining (S606) the estimated value of the feedback input signal component based on the estimated value of the feedback path transfer function.
[0234] A hearing aid's feedback control system can be configured to provide estimates of the feedback input signal components. This feedback control system includes a trained machine learning model. The estimate of the feedback path transfer function can be determined (e.g., provided) by the trained machine learning model. This machine learning model is based on... Figure 5 The method was obtained by training 100.
[0235] Method 600 includes determining (S608) a feedback correction input signal (e.g., the feedback correction input signal at the r-th iteration) based on estimates of the electrical input signal and feedback input signal components. The feedback correction input signal is a feedback-corrected version of the electrical input signal (e.g., a version of the electrical input signal corrected by acoustic and / or mechanical feedback). The feedback control system may be configured to determine (e.g., provide) estimates of the feedback input signal components.
[0236] Method 600 includes generating (S610) a second processed signal by applying one or more processing algorithms (e.g., a gain function varying with frequency and / or level) to the feedback correction input signal. For example, the second processed signal can be considered as a processed signal determined in the r-th iteration of multiple iterations R. The second processed signal can be considered as the current processed signal. Optionally, method 600 includes applying a gain function varying with frequency and / or level to the feedback correction input signal. The signal processing unit of the hearing aid can be configured to generate the second processed signal. The signal processing unit of the hearing aid (e.g., Figure 2The signal processing unit 406 can be configured to apply one or more processing algorithms (e.g., a gain function that varies with frequency and / or level) to the feedback correction input signal.
[0237] For example, Method 600 is a machine learning (ML) inference method. In other words, the estimate of the feedback path transfer function can be the output of the inference (e.g., derivation) of the machine learning model. For example, determining the estimate of the feedback path transfer function involves feeding an electrical input signal and a processed signal into a trained machine learning model.
[0238] For example, during the training phase, machine learning models may introduce artifacts that could be perceived as excessively strong distortion by hearing aid users.
[0239] Method 600 may include controlling the estimate of the feedback path transfer function (e.g., multiple coefficients of the estimate) to avoid excessive artifact removal problems associated with the machine learning model. For example, method 600 includes controlling the estimate by applying a smoothing technique (e.g., causal exponential smoothing) to the estimate of the feedback path transfer function. This smoothing technique can be considered a post-processing technique. For example, applying a smoothing technique to the estimate of the feedback path transfer function can mitigate distortion in the estimate (e.g., distortion caused by suboptimal predictions from the trained machine learning model).
[0240] The final estimate of the feedback path transfer function at time index n can be expressed as:
[0241]
[0242] in, This represents the estimated value of the feedback path transfer function at time index n (e.g., the output of a trained machine learning model). This represents the final estimate of the feedback path transfer function at time index n-1, where α∈[0,1] represents the weighting parameters. For example, This represents the final estimated value at time index n. For example, It can be viewed as a smoothed version of the feedback path transfer function estimate. The weighting parameters can be user-adjustable.
[0243] For example, the estimated value of the feedback input signal component can be determined based on the final estimated value of the feedback path transfer function.
[0244] For example, similar to parameter N in equation (1), the weighting parameter α can be used to further control the balance between convergence speed and steady-state error. For example, unlike parameter N, the weighting parameter α can be adjusted during the inference phase without retraining the machine learning model.
[0245] For example, the parameter N (e.g., the average pooling parameter) in formula (1) can be set to 50 frames. For example, the weighting parameter α can be (e.g., approximately) equal to 0.5.
[0246] Method 600 includes outputting (S612) an audible signal to a hearing aid user based on a second processed signal. The output unit of the hearing aid can be configured to output the audible signal. For example, method 600 includes generating a stimulus signal that the hearing aid user can perceive as sound based on the processed signal.
[0247] For example, the size of the electrical input signal may differ from the size of the training input signal (e.g., the training input signal used in the training phase). Similarly, the size of the processed signal may differ from the size of the training processed signal (e.g., the training processed signal used in the training phase). For example, the size of the estimate of the feedback path transfer function (e.g., provided by the machine learning model in the inference phase) may be the same as the size of the estimate of the training feedback path transfer function (e.g., provided by the machine learning model in the training phase). In other words, the size of the feedback path transfer function may be the same as the size of the training feedback path transfer function.
[0248] Each of the multiple test input signals may contain M = 997 frames. Each of the multiple test processed signals may contain M = 997 frames. Each of the multiple test feedback path transfer functions may contain 64 coefficients (e.g., taps).
[0249] Figures 4A-4B An exemplary training structure 800 for a machine learning (ML) model 814 according to the present invention is illustrated schematically.
[0250] exist Figure 4A In one embodiment, the training structure 800 includes a preprocessing unit 801, a machine learning model 814, a loss function unit 816, and a weight determination unit 818. The machine learning model 814 can be considered as a machine learning model unit (e.g., its output is determined by the machine learning model 814).
[0251] In one or more examples, the preprocessing unit 801 is configured to determine a first preprocessed signal 812AA by applying preprocessing techniques to the training input signal 802. The training input signal 802 includes an external input signal component and a feedback input signal component. The external input signal component represents the sound of a known simulated acoustic environment of the source auto-audio system. The feedback input signal component represents the acoustic and / or mechanical feedback of the feedback path of the source auto-audio system.
[0252] In one or more examples, the preprocessing unit 801 is configured to determine a second preprocessed signal 812BA by applying preprocessing techniques to the training-processed signal 804. The training-processed signal 804 represents the result of applying one or more processing algorithms to the training feedback correction input signal. The training feedback correction input signal is a feedback-corrected version of the training input signal. For example, the training feedback correction input signal represents a version of an electrical input signal corrected by acoustic and / or mechanical feedback.
[0253] In one or more examples, the preprocessing unit 801 includes at least one of the following: a first Fourier transform unit 808A, a second Fourier transform unit 808B, a first normalization unit 810A, a second normalization unit 810B, a first decomposition unit 812A, and a second decomposition unit 812B.
[0254] In one or more examples, a first Fourier transform unit 808A is configured to determine a frequency domain training input signal 808AA by applying a first Fourier transform-based technique to the training input signal 802. In one or more examples, a second Fourier transform unit 808B is configured to determine a frequency domain training processed signal 808BA by applying a second Fourier transform-based technique to the training processed signal 804. The first Fourier transform-based technique may be the same as the second Fourier transform-based technique.
[0255] In one or more examples, the first normalization unit 810A is configured to determine a normalized version 810AA of the frequency domain training input signal 808AA. In one or more examples, the second normalization unit 810B is configured to determine a normalized version 810BA of the frequency domain training processed signal 808BA.
[0256] In one or more examples, the first decomposition unit 812A is configured to decompose the normalized version 810AA of the frequency domain training input signal 808AA into a first principal component and a second primary component. In other words, the normalized version 810AA of the frequency domain training input signal 808AA may contain a first principal component and a second primary component. In one or more examples, the second decomposition unit 812B is configured to decompose the normalized version 810BA of the frequency domain training processed signal 808BA into a second principal component and a second primary component. In other words, the normalized version 810BA of the frequency domain training processed signal 808BA may contain a second principal component and a second primary component.
[0257] For example, the first preprocessed signal 812AA can be the frequency domain training input signal 808AA. For example, the first preprocessed signal 812AA can be a normalized version 810AA of the frequency domain training input signal 808AA. For example, the first preprocessed signal 812AA can include a first principal component and a first primary component. For example, the second preprocessed signal 812BA can be the frequency domain training processed signal 808BA. For example, the second preprocessed signal 812BA can be a normalized version 810BA of the frequency domain training processed signal 808BA. For example, the second preprocessed signal 812BA can include a second principal component and a second primary component.
[0258] In one or more examples, the machine learning model 814 is configured to acquire (e.g., receive) a first preprocessed signal 812AA and a second preprocessed signal 812BA from the preprocessing unit 801. Optionally, the machine learning model 814 may be configured to acquire a training input signal 802 and a training-processed signal 804 (e.g., training data) from the preprocessing unit 801. In other words, the machine learning model 814 may be configured to acquire signals derived (e.g., determined, calculated) from the training data from the preprocessing unit 801.
[0259] In one or more examples, the machine learning model 814 is configured to determine an estimate 814A of the training feedback path transfer function 806 based on a first preprocessed signal 812AA and a second preprocessed signal 812BA. Optionally, the machine learning model 814 is configured to determine an estimate 814A of the training feedback path transfer function 806 based on a training input signal 802 and a training-processed signal 804 (e.g., training data).
[0260] In one or more examples, the loss function unit 816 is configured to acquire a training feedback path transfer function 806 (e.g., target data). The training feedback path transfer function 806 represents the impulse response of the hearing aid feedback path. For example, the training feedback path transfer function 806, the training input signal 802, and the training processed signal 804 may be acquired from the memory of an electronic device (e.g., a computer) or the memory of the hearing aid.
[0261] In one or more examples, the loss function unit 816 is configured to determine the training error signal 816A based on the estimate 814A of the training feedback path transfer function 806 and the training feedback path transfer function 806.
[0262] In one or more examples, the weight determination unit 818 is configured to update (e.g., determine) multiple weights 818A of the machine learning model 814 based on the training error signal 816A (using a learning rule). The updated weights may be stored in the machine learning model module 814A (e.g., memory). The machine learning model 814 is updated based on the training feedback path transfer function 806 (e.g., target data) and the estimated value 814A of the training feedback path transfer function 806.
[0263] Figure 4B An exemplary structure of machine learning model 814 is shown. Machine learning model 814 includes a convolutional layer 820, a first fully connected (FC) layer 824, and a long short-term memory (LSTM) layer 826 in the following order.
[0264] In one or more examples, the machine learning model further includes at least one of the following layers: a second fully connected layer 828 and a third fully connected layer 830. In one or more examples, each of the first fully connected layer 826, the second fully connected layer 828, and the third fully connected layer 830 contains a fully connected class technique specifying the activation function. In one or more examples, the activation function includes at least one of the following: the Sigmoid function, the hyperbolic tangent (tanh) function, the rectified linear unit (ReLU) function, the leaky ReLU function, the Swish function, and the Gaussian error linear unit (GELU) function. In one or more examples, the activation function may also include at least one of the following: a Sigmoid-based function, a ReLU-based function, an exponential linear unit (ELU)-based function, a square root linear unit (SRLU)-based function, a SoftMax function, and any other applicable activation function.
[0265] In one or more examples, the convolutional layer 820 incorporates convolutional techniques. In one or more examples, the convolutional layer 820 is configured to determine a first machine learning processing signal 820A by applying convolutional techniques to a first preprocessed signal 812AA and a second preprocessed signal 812BA.
[0266] In one or more examples, the machine learning model 814 includes a splicing unit 822 configured to determine a second machine learning processing signal 822A by splicing a first machine learning processing signal 820A, a first preprocessing signal 812AA, and a second preprocessing signal 812BA in the frequency dimension.
[0267] In one or more examples, the first fully connected layer 824 includes a first fully connected class technique. In one or more examples, the first fully connected layer 824 is configured to determine a third machine learning processing signal 824A by applying the first fully connected class technique to a second machine learning processing signal 822A.
[0268] In one or more examples, LSTM layer 826 incorporates LSTM-like techniques. In one or more examples, LSTM layer 826 is configured to determine a fourth machine learning processing signal 826A by applying LSTM-like techniques to a third machine learning processing signal 824A. In one or more examples, the fourth machine learning processing signal 826A may serve as an estimate 814A of the training feedback path transfer function 806.
[0269] In one or more examples, the second fully connected layer 828 incorporates a second fully connected class technique. In one or more examples, the second fully connected layer 828 is configured to determine a fifth machine learning processing signal 828A by applying the second fully connected class technique to a fourth machine learning processing signal 826A. In one or more examples, the fifth machine learning processing signal 828A serves as an estimate 814A of the training feedback path transfer function 806.
[0270] In one or more example methods, the third fully connected layer 830 includes a third fully connected class technique. In one or more examples, the third fully connected layer 830 is configured to determine a sixth machine learning processing signal 830A by applying the third fully connected class technique to a fifth machine learning processing signal 828A. In one or more examples, the sixth machine learning processing signal 830A may serve as an estimate 814A of the training feedback path transfer function 806.
[0271] In one or more examples, the machine learning model further includes a pooling layer 832. In one or more examples, the pooling layer 832 incorporates pooling-class techniques. In one or more examples, the pooling layer 832 is configured to determine a seventh machine learning processing signal 832A by applying pooling-class techniques to a sixth machine learning processing signal 830A. In one or more examples, the seventh machine learning processing signal 832A may serve as an estimate 814A of the training feedback path transfer function 806.
[0272] Figure 5 A flowchart of an exemplary method 100 for training a machine learning (ML) model in a hearing aid feedback control system according to the present invention is shown.
[0273] In one or more example methods, method 100 is provided by an electronic device (e.g., Figure 6The electronic device 200 performs the operation. For example, the electronic device disclosed herein can be considered as a computer. For example, the electronic device disclosed herein can be considered as any device that includes a processing unit (e.g., a central processing unit and / or a local processing unit). The electronic device may include at least one of the following: hearing aids, smartphones (e.g., mobile phones), portable electronic devices, wearable electronic devices (e.g., smartwatches), hands-free phones, tablets, computers, and any other suitable electronic devices.
[0274] In one or more example methods, method 100 is a computer-implemented method for training a hearing aid (e.g., Figure 2 The machine learning model used in the feedback control system of the hearing aid 400.
[0275] The inference phase can follow the training phase. In other words, the inference method can be executed after method 100 (e.g., the training method for a machine learning model), such as... Figure 3 Method 600. During the training phase, multiple weights associated with the machine learning model can be (continuously) updated. During the inference phase, the machine learning model has been trained (e.g., multiple weights are fixed) and is ready for deployment.
[0276] Method 100 includes performing (S102) multiple training iterations.
[0277] Each training iteration in multiple training iterations includes acquiring (S104) training data.
[0278] The training data includes the training input signal and the processed signal. The training input signal includes external input signal components and feedback input signal components. The external input signal components represent the sound of the source auto-audio system in a known simulated acoustic environment. The feedback input signal components represent the acoustic and / or mechanical feedback of the source auto-audio system's feedback path.
[0279] The post-training signal represents the result of applying one or more processing algorithms to the training feedback correction input signal. The training feedback correction input signal is a feedback-corrected version of the training input signal. For example, this training feedback correction input signal represents a version of an electrical input signal corrected by acoustic and / or mechanical feedback.
[0280] Each training iteration in the multiple training iterations includes acquiring (S106) target data, which contains a training feedback path transfer function representing the impulse response of the hearing aid feedback path.
[0281] Each training iteration in multiple training iterations includes determining the estimated value of the training feedback path transfer function based on the training data (S108).
[0282] Each training iteration in multiple training iterations includes updating the machine learning model (S110) based on the estimated value of the target data and the training feedback path transfer function.
[0283] The machine learning model comprises, in the following order: a convolutional layer, a first fully connected (FC) layer, and a long short-term memory (LSTM) layer. In one or more example methods, the machine learning model may further include at least one second fully connected layer and a third fully connected layer. In one or more example methods, each of the first, second, and third fully connected layers contains a fully connected class technique specifying an activation function. In one or more example methods, the activation function includes at least one of the following: a sigmoid function, a hyperbolic tangent (tanh) function, a modified linear unit (ReLU) function, a leaky ReLU function, a Swish function, and a Gaussian error linear unit (GELU) function. In one or more example methods, the activation function may further include at least one of the following: a sigmoid-based function, a ReLU-based function, an exponential linear unit (ELU)-based function, a square root linear unit (SRLU)-based function, a SoftMax function, and any other suitable activation function.
[0284] In one or more example methods, determining the estimate of the training feedback path transfer function (S108) includes: determining a first preprocessed signal (S108A) by applying a preprocessing technique to the training input signal. In one or more example methods, determining the estimate of the training feedback path transfer function (S108) includes: determining a second preprocessed signal (S108B) by applying a preprocessing technique to the trained signal.
[0285] In one or more example methods, determining (S108A) the first preprocessed signal includes: determining (S108AA) the frequency domain training input signal by applying a Fourier transform-based technique to the training input signal. In one or more example methods, determining (S108B) the second preprocessed signal includes: determining (S108BA) the frequency domain training processed signal by applying a Fourier transform-based technique to the training processed signal.
[0286] In one or more example methods, determining (S108A) the first preprocessed signal includes: determining (S108AB) a normalized version of the frequency domain training input signal. In one or more example methods, the normalized version of the frequency domain training input signal includes a first principal component and a second component. In one or more example methods, determining (S108B) the second preprocessed signal includes: determining (S108BB) a normalized version of the frequency domain training processed signal. In one or more example methods, the normalized version of the frequency domain training processed signal includes a second principal component and a second component.
[0287] In one or more example methods, a first preprocessed signal and a second preprocessed signal are provided as input to a machine learning model. In other words, determining an estimate of the training feedback path transfer function may include providing the first and second preprocessed signals as input to the machine learning model.
[0288] In one or more example methods, the convolutional layer incorporates convolutional techniques. In one or more example methods, determining (S108) the estimate of the training feedback path transfer function includes: determining (S108C) the first machine learning processing signal by applying convolutional techniques to the first preprocessed signal and the second preprocessed signal.
[0289] In one or more example methods, determining the estimate of the training feedback path transfer function (S108) includes determining the second machine learning processing signal (S108D) by concatenating the first machine learning processing signal, the first preprocessed signal, and the second preprocessed signal in the frequency dimension.
[0290] In one or more example methods, the first fully connected layer includes a first fully connected class technique. In one or more example methods, determining (S108) the estimate of the training feedback path transfer function includes: determining (S108E) the third machine learning processing signal by applying the first fully connected class technique to the second machine learning processing signal.
[0291] In one or more example methods, the LSTM layer incorporates LSTM-like techniques. In one or more example methods, determining the estimate of the training feedback path transfer function (S108) involves determining the fourth machine learning processing signal (S108F) by applying LSTM-like techniques to the third machine learning processing signal. In one or more example methods, the fourth machine learning processing signal is the estimate of the training feedback path transfer function.
[0292] In one or more example methods, the second fully connected layer incorporates a second fully connected class technique. In one or more example methods, determining the estimate of the training feedback path transfer function (S108) includes: determining the fifth machine learning processing signal (S108G) by applying the second fully connected class technique to the fourth machine learning processing signal. In one or more example methods, the fifth machine learning processing signal is the estimate of the training feedback path transfer function.
[0293] In one or more example methods, the third fully connected layer incorporates a third fully connected class technique. In one or more example methods, determining the estimate of the training feedback path transfer function (S108) includes: determining the sixth machine learning processing signal (S108H) by applying the third fully connected class technique to the fifth machine learning processing signal. In one or more example methods, the sixth machine learning processing signal is the estimate of the training feedback path transfer function.
[0294] In one or more example methods, the machine learning model further includes a pooling layer. In one or more example methods, the pooling layer incorporates pooling-like techniques. In one or more example methods, determining the estimate of the training feedback path transfer function (S108) includes determining the seventh machine learning processing signal (S108I) by applying pooling-like techniques to the sixth machine learning processing signal. In one or more example methods, the seventh machine learning processing signal is the estimate of the training feedback path transfer function.
[0295] In one or more example methods, updating the (S110) machine learning model includes: determining (S110A) a training error signal based on an estimate of the training feedback path transfer function and the training feedback path transfer function. In one or more example methods, updating the (S110) machine learning model includes: updating multiple weights of the (S110B) machine learning model using a learning rule based on the training error signal.
[0296] Figure 6 A block diagram of an exemplary electronic device 200 according to the present invention is shown. The electronic device 200 includes a memory 201, a processor 202, and an interface 203. The electronic device 200 can be configured to perform... Figure 5 The method disclosed in the document.
[0297] For example, electronic device 200 can be considered as any device that includes a processing unit (e.g., a central processing unit and / or a local processing unit). Electronic devices may include at least one of the following: hearing aids, smartphones (e.g., mobile phones), portable electronic devices, wearable electronic devices (e.g., smartwatches), hands-free phones, tablets, computers, and any other suitable electronic devices.
[0298] Electronic device 200 is configured (e.g., via processor 202) to perform multiple training iterations.
[0299] Electronic device 200 is configured to acquire training data for each training iteration in a plurality of training iterations (e.g., via interface 203 and / or memory 201).
[0300] The training data includes the training input signal and the processed training signal. The training input signal includes external input signal components and feedback input signal components. The external input signal components represent the sound of a known simulated acoustic environment for the source auto-audio system. The feedback input signal components represent the acoustic and / or mechanical feedback of the source auto-audio system's feedback path. The processed training signal represents the result of applying one or more processing algorithms to the training feedback-corrected input signal. The training feedback-corrected input signal is a feedback-corrected version of the training input signal (e.g., a version of the training input signal corrected by acoustic and / or mechanical feedback).
[0301] Electronic device 200 is configured to acquire target data (e.g., via interface 203 and / or memory 201) for each training iteration in a plurality of training iterations, the target data containing a training feedback path transfer function representing the impulse response of the hearing aid feedback path.
[0302] For example, electronic device 200 is configured to retrieve training data and target data from memory 201 (e.g., the electronic device's own memory) and / or the memory of an external device (e.g., a device outside electronic device 200). The external device may be a hearing aid and / or a server device.
[0303] Electronic device 200 is configured to determine an estimate of the training feedback path transfer function based on training data for each training iteration in a plurality of training iterations (e.g., via processor 202).
[0304] Electronic device 200 is configured to update the machine learning model (e.g., via processor 202) for each training iteration in multiple training iterations based on the target data and an estimate of the training feedback path transfer function.
[0305] The machine learning model comprises, in the following order: a convolutional layer, a first fully connected (FC) layer, and a long short-term memory (LSTM) layer. Optionally, the machine learning model may also include at least one of the following: a second fully connected layer, a third fully connected layer, and a pooling layer.
[0306] For example, electronic device 200 includes Figures 4A-4B The training structure 800. For example, processor 202 contains... Figures 4A-4B The training structure is 800.
[0307] Interface 203 can be configured for wired or wireless communication. If wireless communication is used, interface 203 can be implemented using a wireless communication system, such as a short-range wireless communication system like Wi-Fi, Bluetooth, Zigbee, IEEE 802.11, IEEE 802.15, or infrared communication. Optionally, interface 203 may include a connector for wired communication, such as a cable, to establish a wired connection. This connector can connect electronic device 200 to auxiliary equipment to establish a wired connection.
[0308] Optionally, processor 202 can be configured to execute Figure 5 Any operation disclosed in the [system name] (e.g., any one or more of the following: S102, S104, S106, S108, S108A, S108AA, S108AB, S108B, S108BA, S108BB, S108C, S108D, S108E, S108F, S108G, S108H, S108I, S110, S110A, S110B). The operation of electronic device 200 may be embodied as an executable logic program (e.g., lines of code, software program, etc.), which is stored in a non-transitory computer-readable medium (e.g., memory 201) and executed by processor 202.
[0309] Furthermore, the operation of electronic device 200 can be considered as a method configured to be executed by electronic device 200. While the aforementioned functions and operations can be implemented via software, they can also be implemented via dedicated hardware or firmware, or a combination of hardware, firmware, and software. Memory 201 can take at least one of the following forms: buffer, flash memory, hard disk, removable media, volatile memory, non-volatile memory, random access memory (RAM), and any other suitable device. In a typical configuration, memory 201 may include non-volatile memory for long-term data storage and volatile memory used as system memory for processor 202. Memory 201 exchanges data with processor 202 via a data bus. Control lines and an address bus may also exist between memory 201 and processor 202. Figure 6 (Not shown in the image). Memory 201 is a non-transitory computer-readable medium.
[0310] The memory 201 can be configured to store training data (e.g., training input signals, training processed signals), target data (e.g., training feedback path transfer function), an estimate of the training feedback path transfer function, training error signals, and multiple weights in a portion of its storage area.
[0311] Figure 7 Exemplary impulse responses 902, 904 of the hearing aid feedback path according to the present invention are schematically illustrated.
[0312] For example, the target data (e.g., training feedback transfer function) acquired (e.g., generated) for each training iteration in multiple training iterations may include either a synthetically generated feedback path transfer function or a measured feedback path transfer function. For example, a machine learning model to be used in a hearing aid feedback control system may be trained using a synthetically generated feedback path transfer function (e.g., for multiple training iterations). For example, a machine learning model to be used in a hearing aid feedback control system may be trained using a measured feedback path transfer function (e.g., for multiple training iterations). For example, a machine learning model to be used in a hearing aid feedback control system may be trained by combining a synthetically generated feedback path transfer function (e.g., during pre-training) with a measured feedback path transfer function (e.g., during fine-tuning) (e.g., for multiple training iterations).
[0313] The feedback path transfer function represents the impulse response of the feedback path.
[0314] For example, impulse response 902 is a measured impulse response, i.e., an impulse response represented by a measured feedback path transfer function (such as an impulse response derived from a real acoustic environment). For example, impulse response 904 is a synthesized impulse response, i.e., an impulse response represented by a synthesized feedback path transfer function (such as an impulse response derived from a simulated acoustic environment). Figure 7 The differences between impulse response 902 (e.g., measured impulse response) and impulse response 904 (e.g., synthesized impulse response) can be highlighted.
[0315] For example, the machine learning model to be used in a hearing aid feedback control system can be trained (e.g., for multiple training iterations) using a training dataset (e.g., containing multiple training input signals and multiple trained processed signals) and a target dataset (e.g., containing multiple training feedback transfer functions).
[0316] Each of the multiple training input signals can contain M = 997 frames. Each of the multiple training processed signals can contain M = 997 frames. Each of the multiple training feedback path transfer functions can contain 64 coefficients (e.g., taps).
[0317] For example, for a particular training iteration in a series of training iterations, the training process of the machine learning model to be used in the hearing aid feedback control system may employ the following: one of multiple training feedback transfer functions, one corresponding training input signal from multiple training input signals (e.g., the training signal whose feedback input signal component is determined based on the training feedback transfer function), and one corresponding training processed signal from multiple training processed signals (e.g., the processed signal obtained by applying one or more processing algorithms to the signal determined based on the corresponding training input signal and the training feedback transfer function, such as the training feedback correction input signal).
[0318] Figure 8 Table 500 illustrates exemplary configurations of the layers of a machine learning model according to the present invention. For example, Table 500 lists the input size, output size, and number of trainable parameters (e.g., denoted as "#parameters" in Table 1) for each layer of the machine learning model. For example, the machine learning model can also be trained with other configurations (e.g., other input sizes, output sizes, and number of parameters).
[0319] The number of trainable parameters for a given layer refers to the number of parameters (e.g., variable parameters) of one or more filters in that layer. For example, the number of trainable parameters can be viewed as the number of weights (e.g., including bias terms) learned during training. In other words, it can be viewed as the number of weights adjusted by the machine learning model during the training phase (e.g., training operating mode). For example, each of several weights updated based on the training error signal can be considered a trainable parameter of the machine learning model.
[0320] The machine learning model comprises, in the following order: a convolutional layer, a first fully connected (FC) layer, and a long short-term memory (LSTM) layer. Optionally, the machine learning model may also include at least one of the following: a second fully connected layer, a third fully connected layer, and a pooling layer.
[0321] For example, the convolutional layer is configured to determine a first machine learning processing signal by applying convolutional techniques to a first preprocessed signal and a second preprocessed signal. For example, the convolutional layer is configured to output the first machine learning processing signal as an M×130 matrix (where M represents the number of frames). For example, the convolutional layer is configured to receive the first and second preprocessed signals as input. For example, the convolutional layer is configured to receive the first and second preprocessed signals as input as a matrix of size 2×M×130. In other words, each of the first and second preprocessed signals is M×130 in size (e.g., existing as a matrix of that size). For example, the convolutional layer contains filters (e.g., convolutional kernels) of size 4×5, corresponding to 4 frames (e.g., the temporal dimension) and 5 frequency bins (e.g., the frequency dimension).
[0322] In one or more examples, a convolutional layer may contain 41 trainable parameters, such as 41 weights (including bias terms) that need to be determined and updated when training a machine learning model.
[0323] For example, the first fully connected layer is configured to determine the third machine learning processing signal by applying a first fully connected class technique to the second machine learning processing signal. The second machine learning processing signal can be obtained by concatenating the first machine learning processing signal, the first preprocessed signal, and the second preprocessed signal in the frequency dimension (e.g., across k = 130 frequency bins and / or frequency indices). For example, the first fully connected layer is configured to receive the second machine learning processing signal in the form of an M×390 matrix as input and output the third machine learning processing signal in the form of an M×390 matrix.
[0324] In one or more examples, the first fully connected layer may contain 152,000 trainable parameters (denoted as 152k), such as the 152,000 weights (including bias terms) that need to be determined and updated when training a machine learning model. For example, "152k" can be understood as 152,000 trainable parameters.
[0325] For example, the LSTM layer is configured to determine the fourth machine learning processing signal by applying LSTM-like techniques to the third machine learning processing signal. For example, the LSTM layer is configured to receive the third machine learning processing signal as an M×390 matrix as input. For example, the LSTM layer is configured to output the fourth machine learning processing signal as an M×256 matrix.
[0326] In one or more examples, an LSTM layer may contain 663,000 trainable parameters (denoted as 663k), such as the 663,000 weights (including bias terms) that need to be determined and updated when training a machine learning model. For example, "663k" can be understood as 663,000 trainable parameters.
[0327] For example, the second fully connected layer is configured to determine the fifth machine learning processing signal by applying a second fully connected class technique to the fourth machine learning processing signal. For example, the second fully connected layer is configured to receive the fourth machine learning processing signal as an input in the form of an M×256 matrix. For example, the second fully connected layer is configured to output the fifth machine learning processing signal in the form of an M×128 matrix.
[0328] In one or more examples, the second fully connected layer may contain 32,900 trainable parameters (denoted as 32.9k), such as the 32,900 weights (including bias terms) that need to be determined and updated when training a machine learning model. For example, "32.9k" can be understood as 32,900 trainable parameters.
[0329] For example, the third fully connected layer is configured to determine the sixth machine learning processing signal by applying a third fully connected class technique to the fifth machine learning processing signal. For example, the third fully connected layer is configured to receive the fifth machine learning processing signal as an input in the form of an M×128 matrix. For example, the third fully connected layer is configured to output the sixth machine learning processing signal as an M×64 matrix.
[0330] In one or more examples, the third fully connected layer may contain 8300 trainable parameters (denoted as 8.3k), such as the 8300 weights (including bias terms) that need to be determined and updated when training a machine learning model. For example, "8.3k" can be understood as 8300 trainable parameters.
[0331] The machine learning model may also include pooling layers. For example, a pooling layer is configured to determine a seventh machine learning processing signal by applying pooling-like techniques to a sixth machine learning processing signal. For example, the pooling layer is configured to receive the sixth machine learning processing signal as an M×64 matrix as input. For example, the pooling layer is configured to output the seventh machine learning processing signal as an (M-(N-1))×64 matrix, where N represents a parameter value used to control the convergence speed and the accuracy of estimations (e.g., estimations of the training feedback path transfer function). For example, the pooling layer may not contain trainable parameters.
[0332] For example, the machine learning model contains a total of 856,000 trainable parameters (denoted as 856k). For example, the 856,000 weights (including bias terms) that need to be determined and updated during the training phase can be understood as 856,000 trainable parameters.
[0333] For example, the machine learning model can be trained using multiple training input signals (e.g., using 5375 voices for training and 1344 voices for validation, each voice lasting 10 seconds).
[0334] For example, training speech can be considered as an external input signal component used to generate both the training input signal and the post-training signal, both of which are related to the pre-training and / or fine-tuning processes. For example, the training input signal related to the pre-training process can be determined based on a synthetic feedback transfer function. For example, the training input signal related to the fine-tuning process can be determined based on a measured feedback transfer function.
[0335] For example, the machine learning model can be trained (e.g., pre-trained, validated) using multiple synthetically generated feedback transfer functions (e.g., impulse responses). For example, generating multiple synthetic feedback transfer functions can simulate real-world scenarios (e.g., scaling processes ensure that the maximum amplitude response is random and uniformly distributed across the entire frequency range within the [-20, -10] dB range). For example, 10,000 feedback transfer functions have been generated (e.g., through computer simulation) and randomly combined with 1,344 speech samples to generate 100,000 training sequences and 20,000 validation sequences. The validation sequences can be used to validate the fine-tuned machine learning model. The training sequences can be viewed as training data (e.g., training input signals and / or trained processed signals) related to the pre-training and / or fine-tuning processes, and target data (e.g., synthetically generated feedback path transfer functions and / or measured path transfer functions) related to the pre-training and / or fine-tuning processes.
[0336] For example, the speech used for verification can be considered as an external input signal component used to generate both the training input signal and the processed training signal, both of which are relevant to the verification process. For instance, the training input signal relevant to the verification process can be determined based on a measured feedback transfer function or a synthesized feedback path transfer function.
[0337] For example, multiple measured feedback transfer functions (e.g., impulse responses) can be used to train a machine learning model (e.g., pre-training, fine-tuning, validation). For example, the multiple measured feedback transfer functions could include 1010 measured feedback transfer functions (each associated with a specific combination of hearing aids and earpieces). For example, the 1010 measured feedback transfer functions could be divided into three groups: 753 for training the machine learning model (e.g., pre-training and fine-tuning), 107 for validating the machine learning model, and 200 for testing the machine learning model (e.g., for testing the validated final machine learning model). For example, generating multiple measured feedback transfer functions can avoid large variations within the data (e.g., each measured feedback transfer function includes an amplitude response that is randomly scaled within the range of [-20, -10] dB). For example, 753 training sequences, 107 validation sequences, and measured and / or synthesized feedback transfer functions can be randomly combined to generate 10,000 training sequences and 3,000 validation sequences.
[0338] For example, the generation of training and validation sequences can be achieved through closed-loop hearing aid simulation. For instance, a signal processing unit (e.g., a signal processing unit associated with the hearing aid simulation) can be configured to determine the post-training signal (e.g., the post-training signal associated with the pre-training, fine-tuning, and validation processes) by applying a gain function that varies with frequency and / or level to the training feedback correction input signal (e.g., the training feedback correction input signal associated with the pre-training, fine-tuning, and validation processes). For example, the generation of the gain function that varies with frequency and / or level can simulate a real-world environment (e.g., a complex environment), such as ensuring maximum loop gain (e.g., max). ω (|H(ω,n)·G|, where H(ω,n) represents the frequency response of the feedback transfer function) is uniformly distributed in the range of [-6,0] dB.
[0339] For example, the training and validation sequences can be split into non-overlapping sequences of 2 seconds each, ultimately generating 52,648 training sequences and 6,547 validation sequences. Performing such a split can reduce the number of frames (e.g., M) across which the training loss is averaged during each gradient update (e.g., training iteration) during the training phase. Each of the multiple training input signals (e.g., training input signals associated with the pre-training, fine-tuning, and validation processes) can contain M = 997 frames. Each of the multiple post-training signals (e.g., post-training signals associated with the pre-training and / or fine-tuning processes) can contain M = 997 frames. Each of the multiple training feedback path transfer functions (e.g., measured feedback transfer functions and / or synthesized feedback transfer functions) can contain 64 coefficients (e.g., taps). The (final) validation sequence may differ from the (final) training sequence.
[0340] Figure 9 A graph 700 illustrating an exemplary training loss according to the present invention is shown. The horizontal axis 708A (e.g., the X-axis) represents time (in seconds, ranging from 0 to 15 seconds). The vertical axis 708B (e.g., the Y-axis) represents the training loss (in decibels (dB)). Figure 9 The training loss associated with the NESD loss function is illustrated, for example, by presenting the NESD curve after averaging across the test dataset. In other words, the average NESD over time reflects the accuracy of the impulse response estimation.
[0341] Curve 702 illustrates the NESD training loss associated with time-domain broadband adaptive feedback control techniques (e.g., denoted as "TD-AFC," a traditional adaptive filtering technique). In other words, curve 702 reflects the average performance of a feedback control system incorporating an adaptive filter based on TD-AFC (e.g., an adaptive algorithm employing TD-AFC).
[0342] Curve 704 illustrates the NESD training loss associated with frequency-domain adaptive feedback control techniques (e.g., denoted as "FD-AFC," a traditional adaptive filtering technique). In other words, curve 704 reflects the average performance of a feedback control system incorporating an adaptive filter based on FD-AFC (e.g., an adaptive algorithm employing FD-AFC).
[0343] Both FD-AFC and TD-AFC employ the Normalized Least Mean Square (NLMS) algorithm. For example, FD-AFC only performs adaptive adjustments in the frequency range above 1000 Hz to avoid low-frequency artifacts.
[0344] Curves 706, 708, and 710 illustrate the NESD training loss associated with the DFC technique. In other words, these curves reflect the average performance of a feedback control system that includes a machine learning model trained using the DFC technique (e.g., a machine learning model trained using the training method provided in this invention).
[0345] Curve 706 illustrates that the feedback control system incorporates a machine learning model, which is trained using a synthetically generated feedback transfer function (e.g., as target data) during the training phase and tested using a measured feedback transfer function during the testing phase (similar to the inference phase). This method is called the "DFC(S) method".
[0346] Curve 708 illustrates that the feedback control system includes a machine learning model, which uses a measured feedback transfer function in both the training and testing phases (similar to the inference phase). This method is called the "DFC(M) method".
[0347] Curve 710 illustrates that a feedback control system incorporates a machine learning model. This model is pre-trained using a synthetically generated feedback transfer function (e.g., as target data) during the training phase, fine-tuned using a measured feedback transfer function (also as target data), and tested using the measured feedback transfer function during the testing phase (similar to the inference phase). This method is called the "DFC method." In other words, curve 710 reflects the average performance of a feedback control system that incorporates a machine learning model trained with both synthetic and measured feedback transfer functions, and tested using measured feedback transfer functions.
[0348] For example, a measured feedback transfer function can represent a real-world measured feedback transfer function (e.g., the impulse response of a real-world measured feedback path).
[0349] For example, the test dataset contains multiple test input signals, multiple processed test signals, and multiple test feedback path transfer functions. The test dataset can be generated through computer simulation, such as simulating complex acoustic scenarios (e.g., near-unstable scenarios). The multiple test feedback path transfer functions may include synthetically generated feedback path transfer functions. For example, the data contained in the test dataset may differ from the training and target data used in the training phase. Each of the multiple test input signals may contain M = 997 frames. Each of the multiple processed test signals may contain M = 997 frames. Each of the multiple test feedback path transfer functions may contain 64 coefficients (e.g., taps).
[0350] For example, the test dataset contains 100 speech samples (e.g., input signals) from 6 speakers (e.g., users), each 15 seconds long. For example, the feedback path changes at 7.5 seconds in these 100 speech samples. For example, the test dataset contains real-world data, i.e., data obtained from real-world environments. For example, the test speech (e.g., from the test dataset) can be considered as an external input signal component used to generate both the test input signal and the test processed signal, both of which are related to the testing process. For example, the test input signal can be determined based on a measured feedback transfer function. For example, the test dataset can contain 200 test sequences (e.g., for testing a validated machine learning model).
[0351] For example, a validated machine learning model can be tested using a test dataset. Testing the validated machine learning model may include executing the training method provided by this invention during the fourth training iteration in a series of training iterations. For example, the test dataset may contain training data (e.g., test sequences) and target data (e.g., test sequences) related to the testing process. For example, the training data related to the testing process may contain test input signals and test processed signals. For example, the target data related to the testing process may contain measured feedback path transfer functions (e.g., test sequences). A measured feedback path transfer function corresponding to a specific speech in 100 speech samples may represent the feedback path change that occurs at 7.5 seconds in that specific speech. The test sequence may differ from the validation sequence and training sequence. For example, the machine learning model of this invention may be trained using training sequences, validation sequences, and test sequences.
[0352] Figure 9 This indicates that, compared to traditional adaptive filtering techniques, feedback control systems incorporating machine learning models trained using the DFC method (e.g., trained with synthetic and experimental feedback transfer functions and tested with experimental feedback transfer functions) exhibit not only smaller steady-state errors but also significantly faster convergence speeds after abrupt changes in the feedback path. Specifically, in Figure 9 In the embodiments (e.g.) Figure 9 In a specific simulated scenario, the FD-AFC method takes an average of about 3 seconds to converge, while the DFC method converges in less than 0.1 seconds, making the latter about 30 times faster. Furthermore, the steady-state error of the DFC method is reduced by approximately 2 dB compared to the steady-state error of the FD-AFC method.
[0353] For example, the fast reconvergence characteristic of the DFC method (which has been verified) makes changes in the feedback path difficult for users to perceive in most cases; while for the FD-AFC method (especially the TD-AFC method), changes in the feedback path can introduce severe sound artifacts. Furthermore, Figure 9 This also demonstrates the superiority of the training method provided by this invention. In other words, training a machine learning model using a synthesized feedback path transfer function (e.g., during pre-training) and a measured feedback path transfer function (e.g., during fine-tuning) can achieve faster convergence speed and lower steady-state error, thereby improving the balance between convergence speed and steady-state error (e.g., compared to traditional adaptive filtering techniques and / or existing machine learning-based feedback control methods).
[0354] Figure 9 This indicates that when training machine learning models using only synthetically generated feedback path transfer functions (e.g., the DFC(S) method, see curve 706), model performance degrades significantly. This significant performance degradation may stem from a mismatch between the training data, the target data, and the test dataset.
[0355] also, Figure 9 It also shows that when training machine learning models using only the experimental feedback path transfer function (e.g., the DFC(M) method, see curve 708), the model performance is not optimal.
[0356] Embodiments of the present invention provide methods for training machine learning models using synthetically generated feedback path transfer functions and experimentally measured feedback path transfer functions. Embodiments of the present invention also provide methods for training machine learning models using only synthetically generated feedback path transfer functions. Finally, embodiments of the present invention provide methods for training machine learning models using only experimentally measured feedback path transfer functions.
[0357] Figure 9 The results show that the machine learning model trained using the synthetically generated feedback path transfer function and the measured feedback path transfer function can improve the convergence speed by 30 times and reduce the steady-state error by 2 dB under normal working mode (e.g., inference stage) and when the feedback path changes rapidly.
[0358] from Figure 9It can be observed that the model trained using both synthetic and experimental feedback path transfer functions performs better in feedback path transfer function estimation compared to training the machine learning model using only synthetic or experimental feedback path transfer functions. Nevertheless, the approach of training the machine learning model using only synthetic or experimental feedback path transfer functions can still serve as an alternative to traditional adaptive filtering techniques and / or existing machine learning techniques (such as TD-AFC and / or FD-AFC).
[0359] Figure 10 Table 730 illustrates an exemplary mean and standard deviation (STD) of a machine learning-based feedback control system comprising a trained machine learning model according to the present invention. The machine learning model is based on... Figure 5 The method was obtained by training 100.
[0360] For example, Table 730 lists the mean and standard deviation of the following systems: TD-AFC systems (e.g., using TD-AFC techniques), FD-AFC systems (e.g., using FD-AFC techniques), and machine learning-based feedback control systems that include trained machine learning models (e.g., using deep feedback compensation (DFC) techniques). A machine learning-based feedback control system that includes trained machine learning models can be called a DFC system. Figure 5 Method 100 can be called DFC technology.
[0361] For example, "full sequence" can be understood as the entire test dataset with feedback path changes. For example, "no path change" can be understood as the portion of the test dataset that does not contain feedback path changes. For example, "with path change" can be understood as the portion of the test dataset that contains feedback path changes.
[0362] TD-AFC and FD-AFC technologies are benchmark (e.g., current level of development) technologies used for comparison with the DFC technology provided in this invention. TD-AFC and FD-AFC technologies can be considered as traditional adaptive filtering technologies.
[0363] Table 730 provides the Perceptual Speech Quality Assessment (PESQ) scores for the feedback transfer function estimates of TD-AFC, FD-AFC, and DFC technologies.
[0364] For example, the estimated feedback transfer functions of TD-AFC, FD-AFC, and DFC technologies can be compared with the reference feedback transfer function (e.g., the training feedback transfer function) of an ideal feedbackless control system—in which the processed signal is generated from the external input signal components and is unaffected by the feedback input signal components and / or their estimated values. In other words, the estimated feedback transfer functions of TD-AFC, FD-AFC, and DFC technologies can also be compared with the reference feedback transfer function (e.g., the training feedback transfer function) of a known feedback system (e.g., a system with a known feedback transfer function, such as a system containing a reference feedback transfer function).
[0365] Figure 10 This indicates that TD-AFC technology scored the lowest. This low score was caused by the degradation of sound quality in the low-frequency range.
[0366] Although the DFC system is configured to operate across the entire frequency range (which can lead to significant degradation in the quality of the input signal at low frequencies), this system still mitigates this problem. While FD-AFC technology was observed to also mitigate the aforementioned problem by not implementing feedback adaptation (e.g., feedback control) at low frequencies, the PESQ score of the DFC system is still superior to that of the FD-AFC system employing FD-AFC technology (e.g., a higher score).
[0367] For example, when TD-AFC, FD-AFC, and DFC systems are required to respond to changes in the feedback path (e.g., in static and / or dynamic feedback scenarios), significant differences in their PESQ scores can be observed. Specifically, in scenarios involving changes in the feedback path (e.g.) Figure 9 Within the time interval of 7.4 seconds to 9.4 seconds, the PESQ score of DFC technology differs from that of the FD-AFC system by 0.96; in the same scenario ( Figure 9 During the time interval of 7.4 seconds to 9.4 seconds, the PESQ score of DFC technology differed from that of TD-AFC system by 2.68.
[0368] Furthermore, Table 730 shows that the standard deviation (e.g., and / or variance) of the PESQ score of the TD-AFC system and the FD-AFC system is higher than that of the DFC system. This result demonstrates the superiority of the DFC system provided by the present invention over the TD-AFC system and the FD-AFC system.
[0369] The term "processed version" may, for example, cover features extracted from the original audio signal. The term "processed version" may also, for example, cover the original audio signal that has been subjected to a processing algorithm that applies gain or attenuation and / or delay to the original audio signal, resulting in a modified audio signal (preferably enhanced in some way, such as noise reduction relative to the target signal, or simply delayed).
[0370] When appropriately replaced by a corresponding process, the structural features of the apparatus described above, in detail in the "Detailed Description" section, and as defined in the claims can be combined with the steps of the method of the present invention.
[0371] Unless explicitly stated otherwise, the singular forms “a” and “the” used herein include the plural forms (i.e., meaning “at least one”). It should be further understood that the terms “having,” “comprising,” and / or “including” as used in the specification indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that, unless explicitly stated otherwise, when an element is referred to as “connected” or “coupled” to another element, it may be a direct connection or coupling to the other element, or there may be intermediate inserting elements. The term “and / or” as used herein includes any and all combinations of one or more of the listed related items. Unless explicitly stated otherwise, the steps of any method disclosed herein do not necessarily have to be performed in the exact order disclosed.
[0372] It should be understood that references to "an embodiment," "an embodiment," "an aspect," or "may" in this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Furthermore, particular features, structures, or characteristics may be suitably combined in one or more embodiments of the invention. The foregoing description is provided to enable those skilled in the art to implement the various aspects described herein. Various modifications will be apparent to those skilled in the art.
[0373] The claims are not limited to the aspects shown herein, but encompass the full scope consistent with the language of the claims, wherein, unless expressly stated, an element referred to in the singular does not mean "one and only one," but rather "one or more." Unless expressly stated, the term "some" means one or more.
Claims
1. A method performed by an electronic device for training a machine learning model for use in a hearing aid feedback control system, the method comprising performing a plurality of training iterations, each training iteration of the plurality of training iterations comprising: obtaining training data, the training data comprising: a training input signal, the training input signal comprising an external input signal component and a feedback input signal component; wherein the external input signal component represents sound originating from a known simulated acoustic environment of a hearing aid, and the feedback input signal component represents acoustic and / or mechanical feedback originating from a feedback path of the hearing aid; a training processed signal, the training processed signal representing a result of applying one or more processing algorithms to a training feedback corrected input signal, the training feedback corrected input signal being a version of the training input signal that has been feedback corrected; obtaining target data, the target data comprising a training feedback path transfer function representing an impulse response of the feedback path of the hearing aid; determining an estimate of the training feedback path transfer function based on the training data; updating a machine learning model based on the target data and the estimate of the training feedback path transfer function; wherein the machine learning model comprises, in the following order, a convolutional layer, a first fully connected layer, and a long short-term memory layer.
2. The method of claim 1, wherein, the machine learning model further comprises at least one of a second fully connected layer and a third fully connected layer; each of the first, second, and third fully connected layers comprises a fully connected class of techniques that specifies an activation function; the activation function comprises at least one of a sigmoid function, a hyperbolic tangent function, a rectified linear unit (ReLU) function, a leaky ReLU function, a Swish function, and a Gaussian error linear unit function.
3. The method of claim 1, wherein, determining the estimate of the training feedback path transfer function comprises: determining a first pre-processed signal by applying a pre-processing technique to the training input signal; determining a second pre-processed signal by applying the pre-processing technique to the training processed signal.
4. The method of claim 3, wherein: determining the first pre-processed signal comprises determining a frequency domain training input signal by applying a Fourier transform based technique to the training input signal; determining the second pre-processed signal comprises determining a frequency domain training processed signal by applying the Fourier transform based technique to the training processed signal.
5. The method of claim 4, wherein: determining the first pre-processed signal comprises determining a normalized version of the frequency domain training input signal, wherein the normalized version of the frequency domain training input signal comprises a first principal component and a first secondary component; determining the second pre-processed signal comprises determining a normalized version of the frequency domain training processed signal, wherein the normalized version of the frequency domain training processed signal comprises a second principal component and a second secondary component.
6. The method of claim 3, wherein, the convolutional layer comprises a convolutional class of techniques; wherein determining the estimate of the training feedback path transfer function comprises: determining a first machine learning processed signal by applying the convolutional class of techniques to the first pre-processed signal and the second pre-processed signal.
7. The method of claim 6, wherein, The first fully connected layer comprises a first fully connected class technique; wherein determining the estimated value of the training feedback path transfer function comprises: determining a third machine learning processed signal by applying the first fully connected class technique to the first machine learning processed signal, the first pre-processed signal, and the second pre-processed signal.
8. The method of claim 7, wherein, The long short-term memory layer comprises a long short-term memory class technique; wherein determining the estimated value of the training feedback path transfer function comprises: determining a fourth machine learning processed signal by applying the long short-term memory class technique to the third machine learning processed signal; wherein the fourth machine learning processed signal is the estimated value of the training feedback path transfer function.
9. The method of claim 8, wherein, The second fully connected layer comprises a second fully connected class technique; wherein determining the estimated value of the training feedback path transfer function comprises: determining a fifth machine learning processed signal by applying the second fully connected class technique to the fourth machine learning processed signal; wherein the fifth machine learning processed signal is the estimated value of the training feedback path transfer function.
10. The method of claim 9, wherein, The third fully connected layer comprises a third fully connected class technique; wherein determining the estimated value of the training feedback path transfer function comprises: determining a sixth machine learning processed signal by applying the third fully connected class technique to the fifth machine learning processed signal; wherein the sixth machine learning processed signal is the estimated value of the training feedback path transfer function.
11. The method of claim 10, wherein, The machine learning model further comprises a pooling layer, the pooling layer comprising a pooling class technique; wherein determining the estimated value of the training feedback path transfer function comprises: determining a seventh machine learning processed signal by applying the pooling class technique to the sixth machine learning processed signal; wherein the seventh machine learning processed signal is the estimated value of the training feedback path transfer function.
12. The method of claim 3, wherein, Determining the estimated value of the training feedback path transfer function comprises: providing the first pre-processed signal and the second pre-processed signal as inputs to the machine learning model.
13. The method of claim 1, wherein, Updating the machine learning model comprises: determining a training error signal based on the estimated value of the training feedback path transfer function and the training feedback path transfer function; updating a plurality of weights of the machine learning model using a learning rule based on the training error signal.
14. A hearing aid, comprising: an input unit configured to provide an electrical input signal representing sound in an environment in which a hearing aid user is located; wherein the electrical input signal comprises an external input signal component representing sound in an environment in which the hearing aid is located and a feedback input signal component representing acoustic and / or mechanical feedback originating from a feedback path from an output unit of the hearing aid to an input unit of the hearing aid; a signal processing unit configured to provide a processed signal by applying one or more processing algorithms to a feedback-corrected input signal; wherein the feedback-corrected input signal is a feedback-corrected version of the electrical input signal; an output unit configured to output an audible signal to the hearing aid user based on the processed signal; wherein the hearing aid comprises a feedback control system comprising a trained machine learning model; the feedback control system is configured to: determine an estimate of a feedback input signal component based on an estimate of a feedback path transfer function; wherein the feedback path transfer function represents an impulse response of the feedback path; the trained machine learning model is configured to provide the estimate of the feedback path transfer function based on the electrical input signal and the processed signal; the machine learning model is a machine learning model trained according to the method of claim 1; determine the feedback correction input signal based on the electrical input signal and the estimate of the feedback input signal component.