Hearing device comprising a recurrent neural network and method for processing an audio signal
Patent Information
- Application Number
- CN202210067599.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-20
- Filing Date
- 2022-01-20
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-01-20
AI Technical Summary
低功率要求限制了例如在助听器中可行的神经网络的类型和大小,因而限制了适合由神经网络实施的功能任务
Smart Images

Figure CN114827859B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing such as audio and / or image processing, and more particularly to the use of neural networks in audio processing, especially to algorithms for implementing learning algorithms or machine learning techniques in (e.g., portable) audio processing devices. Background Technology
[0002] So-called learning algorithms, such as neural networks, have found increasing applications in all classes of signal processing tasks, including (natural) speech processing in hearing aids or headphones. In hearing aids, noise reduction is a key feature for addressing the problem of providing or increasing a user's (e.g., hearing-impaired users') perception of speech and other sounds in the environment as close to acceptable as possible.
[0003] In hearing aids, headphones, or similar portable audio processing devices, a) small size, b) low latency, and c) low power consumption are important design parameters. Therefore, the "power efficiency" of the processing algorithm is paramount. Low power requirements limit, for example, the type and size of neural networks feasible in hearing aids, thus limiting the functional tasks suitable for implementation by neural networks. In practice, the number of network parameters (at current processing power in hearing aids) is limited to the 10,000 level. This number is expected to increase over time (e.g., due to advancements in the field of integrated circuits). Summary of the Invention
[0004] This application relates to improvements in audio or video processing that utilize neural networks, particularly recurrent neural networks, such as delta (Δ) recurrent neural networks (delta RNNs) like long short-term memory (LSTM) type RNNs, or gated recurrent units (GRUs), or modifications thereof. The improvements are exemplified in the field of audio processing, particularly in hearing devices such as hearing aids or headphones.
[0005] The above improvements can be applied to any problem related to data processing, such as representing time-varying data, like audio or video processing. They may be particularly advantageous in applications with limited processing power, and for example, in applications where the (significant) change in data from one time period to the next is limited, for example, less than 20% of the data. Such examples appear in video images moving from one frame to the next, or in audio data, for example, by moving from one time frame to the next (or within a time frame, from one channel to an adjacent channel, e.g., see...). Figure 7 The spectrum is represented by .
[0006] Exemplary hearing devices, such as hearing aids or headphones, may include an input unit for providing at least one electrical input signal in time-frequency representation k,t, where k and t are the frequency index and time index, respectively, and k represents a sub-band channel, k = 1, ..., K, K > 1. The at least one electrical input signal represents sound, which may include a target signal component and a noise component. The input unit may, for example, include an input converter such as a microphone. The hearing device input unit may include an analog-to-digital converter and / or a filter bank. The hearing device input unit may include one or more beamformers for providing a spatially filtered signal based on at least one electrical input signal (e.g., including at least two electrical input signals) or a signal derived therefrom. The hearing device may also include a signal processor connected to the input unit and configured to receive at least one electrical input signal or a signal derived therefrom (e.g., a spatially filtered signal and / or a feedback-corrected signal). The signal processor may include a noise reduction system configured to reduce the noise component in the at least one electrical input signal or a signal derived therefrom. The signal processor can be configured to determine a corresponding gain value G(k,t) in the time-frequency representation, which reduces noise components relative to the target signal components when applied to at least one electrical input signal or a signal derived therefrom.
[0007] Signal processor and corresponding method configured to execute an improved gated loop unit ("improved GRU").
[0008] The signal processor may include a neural network comprising at least one layer, defined as a delta recurrent neural network (delta RNN) including memory. A delta RNN may, for example, include a long short-term memory (LSTM) type RNN or a gated recurrent unit (GRU) or a modification thereof. At least one layer may be implemented as a GRU (or a modified GRU) including memory, such as in the form of hidden states (see, for example, see...). Figure 1A The signal processor is configured to execute a recurrent neural network comprising at least one layer, implemented as a GRU (e.g., a modified GRU according to the invention). The GRU (or the modified GRU according to the invention) provides an output vector o(k,t) based on the input vector x(k,t) and the hidden state vector h(k,t-1), wherein the output o(k,t) at a given time step t is stored as the hidden state h(k,t) and used to compute the output vector o(k,t+1) at the next time step t+1. The noise reduction system may include a neural network (e.g., implemented via a neural network). The parameters of the neural network may have been trained with multiple training signals. The signal processor may be configured to compute, at a given time t, the changes in the input vector x(k,t) and the hidden state vector h(k,t-1) from one time step t-1 to the next time step t, respectively. and in and These are the estimated values for x and h, respectively. and It can be equal to (at least) the values of x and h from an earlier time step. (Estimated value) and It can be stored in memory and used during the calculation of time step t. Estimated value and This can be equal to the last values of x and h, respectively, leading to a threshold change (from one time step to the next). The signal processor can also be configured such that the number of channels used for updating the input vector x(k,t) and the hidden state vector h(k,t-1) at the given time t is limited to the peak (i.e., maximum) value N. p (or N) p,x N p,oh (See below) the number of, where N p Less than (or equal to) N ch (or N) ch,x N ch,oh ), where N ch (or N) p,x N p,oh The total number of processing channels is N. This modified GRU is referred to below as the Peak GRU (or Peak GRU RNN) (see further below). The Peak GRU can be seen as a modified version of the Delta GRU described in [Neil et al.; 2018]. The number of channels actually processed can be limited to less than N through additional energy-saving measures. p A channel, for example, discarding N at a given time step. p The processing / updating of values among the maximum values that are less than or equal to the threshold. Estimated value. and These can be equal to the final values of x and h, respectively, which result in N at the time t' involved. p The change in value among the largest changes.
[0009] For the input vector (x)(N) ch,x ) and output (hidden state) vector (o, h, where h(t) = o(t))(N ch,oh The number of channels (nodes) N ch (N ch,x N ch,oh ) can be equal to (N) ch =N ch,x =N ch,oh (or different.) Similarly, for peak GRU, the number of peaks N p (N p,x N p,oh For the input vector (x)(N) p,x ) and output (hidden state) vector (o,h)(N p,oh ) can be equal to (N) p =Np,x =N p,oh (or different.) The number K of sub-band channels provided by the input unit can be equal to (K = N) ch Or the number of processing channels N that differs from the delta GRU / peak GRU layer. ch (N ch,x N ch,oh (This also applies to the baseline GRU). Therefore, although the channel index of a delta GRU / peak GRU layer is generally denoted as a single (k, equal to the sub-band channel index), in practice it can differ (and the input vector (x) and output vector (o, h) of the delta GRU / peak GRU layer can differ). The (maximum) range of the sub-band channel index k is, for example, 1 ≤ k ≤ K. If the number of input and output nodes of the delta GRU / peak GRU layer is the same, the (maximum) range of the delta GRU / peak GRU channel index k is, for example, 1 ≤ k ≤ N. ch If the number of input and output nodes of the Delta GRU / Peak GRU layer is different, then 1 ≤ i ≤ N ch,x and 1≤j≤N ch,oh Instead of using k as the general “channel index”, the indices i and j can be used for the input vector and the output (and hidden state) vector, respectively, as shown in equations (5)', (6)', (7)' and (8)' below for peak GRU.
[0010] The number of peaks N in peak GRU p (N p,x N p,oh The number of peaks N can be fixed. p (N p,x N p,oh It can be dynamically determined, for example, based on at least one electrical input signal (e.g., based on its characteristics such as modulation, level, estimated signal-to-noise ratio, etc.).
[0011] The number of input nodes can be equal to, for example, the number of channels in the input signal (an input vector can represent, for example, one frame of the input signal) or a multiple of the number of channels. The number of input nodes can range from, for example, 16 to 500. The number of output nodes can be equal to, for example, the number of channels in the input signal or a fraction of that number. The number of output nodes can range from, for example, 1 to 500, such as 1 to 150. The number of input nodes and output nodes can be the same.
[0012] The number of layers in a neural network can be greater than 2. The number of "hidden" layers can be greater than or equal to 2, for example, greater than or equal to 3, and in the range of 2 to 10.
[0013] The number of layers in the modified gated recurrent unit implemented according to the present invention can be greater than or equal to 1, for example, greater than or equal to 2, for example, in the range of 1 to 3, for example, in the range of 1 to 10. All layers of the neural network can be modified GRU layers.
[0014] The improved gated loop unit according to the invention can be used, for example, in audio or video processing applications, such as in applications where (low) power consumption is an important parameter, such as in wearable electronic devices like hearing aids or headphones or handheld video processing devices.
[0015] In this invention, the use of the modified GRU is illustrated in conjunction with noise reduction (e.g., SNR-gain conversion) in audio processing devices such as hearing aids. However, the modified GRU can be used in other applications, such as self-voice detection, wake-word detection, keyword detection, and voice activity detection. Furthermore, the modified GRU can be used in video processing, for example, processing video images frame by frame. The similarity between audio-images (spectral maps) and other (ordinary, video) images is obvious. However, audio-images (spectral maps) differ from ordinary (video) images in that they possess a temporal dimension. Given that the input from an audio processing application to a neural network might be a time frame comprising the "spectrum" of the audio signal at a given moment, and that audio-images (spectral maps) are constructed through cascaded sequential time frames, the input from a video processing application to a neural network might be a sequence of images, where each image represents a specific moment, and where the sequence of images provides the temporal dimension.
[0016] Hearing devices including improved gating loop units
[0017] On one hand, a hearing device, such as a hearing aid or headphones, is provided. The hearing device can be configured to be worn by a user in or in the ear, or fully or partially implanted in the head of the user's ear. It includes an input unit for providing at least one electrical input signal in a time-frequency representation and a signal processor including a neural network configured to provide a corresponding gain value G(k,t) in the time-frequency representation to reduce noise components in the at least one electrical input signal. The neural network includes at least one layer defined as a modified Mendelta-controlled recurrent unit (referred to as a peak GRU), which includes a memory of the form of a hidden state vector h, wherein the output vector o is provided by the peak GRU according to the input vector x and the hidden state vector h, wherein the output o(j,t) of the peak GRU at a given time step t is stored as a hidden state vector h(j,t) and used to compute the output o(j,t+1) of the next time step t+1. The signal processor is configured such that for the input vector x(t) and the hidden state vector h(t-1) at a given time t, the N of the peak GRU... ch (N ch,x N ch,oh The number of update channels in a processing channel is limited to the peak value N.p (N p,x N p,oh ), where N p (N p,x N p,oh (less than N) ch (N ch,x N ch,oh The signal processor is configured to execute a neural network including a modified gated recurrent unit according to the invention. A method for operating the hearing device is also disclosed.
[0018] First hearing device
[0019] In one aspect of this application, a hearing device is provided. The hearing device may include:
[0020] - An input unit for providing at least one electrical input signal in a time-frequency representation of k,t, where k and t are the frequency exponent and time exponent, respectively, and k represents a sub-band signal, k = 1, ..., K, and the at least one electrical input signal represents sound and includes a target signal component and a noise component; and
[0021] -Signal processor, including
[0022] --SNR estimator for providing a target signal-to-noise ratio (SNR) estimate SNR(k,t) of the at least one electrical input signal or a signal derived therefrom in the time-frequency representation;
[0023] --SNR- gain converter for converting the target signal-to-noise ratio estimate SNR(k,t) into the corresponding gain value G(k,t) in the time-frequency representation;
[0024] The signal processor includes a neural network comprising at least one layer defined as a gated recurrent unit (GRU), the GRU including a memory in the form of a hidden state vector h, wherein the output vector o(t) is provided by the GRU based on the input vector x(t) and the hidden state vector h(t-1), wherein the output o(t) at a given time step t is stored as the hidden state h(t) and used to compute the output vector o(t+1) at the next time step t+1. The hearing device can be configured such that at least the SNR-gain converter is implemented via the neural network, and at least one layer defined as a GRU is implemented as a modified GRU, wherein the signal processor is configured to compute, at a given time t, the changes in the input vector x(t) and the hidden state vector h(t-1) from one time t-1 to the next time t. and in and Let x(i,t-1) and h(j,t-2) be the estimated values, respectively, where i and j refer to the i-th and j-th input neurons in the hidden state, respectively, and 1≤i≤N. ch,x and 1≤j≤N ch,oh , where N ch,x and N ch,oh The number of processing channels for the input vector x and the hidden state vector h are respectively, and the signal processor is further configured such that for the input vector x(t) and the hidden state vector h(t-1) at a given time t, the N of the improved gated loop unit is... ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p.x Less than N ch,x and N p,oh Less than N ch,oh .
[0025] The input vector of the neural network may be based on or include the target signal-to-noise ratio (SNR) estimate SNR(k,t) of the at least one electrical input signal or a signal derived therefrom. The output vector of the neural network includes the gain (G(k,t)) of the denoising algorithm.
[0026] This can provide improved hearing devices.
[0027] The method of operating the first hearing device is also disclosed (by converting the structural features of the first hearing device into process features).
[0028] Second hearing device
[0029] In one aspect of this application, a hearing device is provided. The hearing device includes:
[0030] - An input unit for providing at least one electrical input signal in a time-frequency representation of k,t, where k and t are the frequency index and time index, respectively, and k represents the channel, k = 1, ..., K, K > 1; the at least one electrical input signal represents sound and includes a target signal component and a noise component; and
[0031] - A signal processor connected to the input unit and configured to receive at least one electrical input signal or a signal derived therefrom, the signal processor being configured to determine a corresponding gain value G(k,t) in the time-frequency representation, reducing the noise component relative to the target signal component when the gain value is applied to the at least one electrical input signal or one or more signals derived therefrom, wherein the signal processor includes a neural network comprising at least one layer defined as a gated recurrent unit, the gated recurrent unit being in the form of a modified gated recurrent unit;
[0032] The signal processor is configured to calculate, at a given time t, the changes in the input vector x(t) and the hidden state vector h(t-1) from one time t-1 to the next time t. and in and Let x(i,t-1) and h(j,t-2) be the estimated values, respectively, where i and j refer to the i-th and j-th input neurons in the hidden state, respectively, and 1≤i≤N. ch,x and 1≤j≤N ch,oh , where N ch,x and N ch,oh These represent the number of processing channels for the input vector x and the hidden state vector h, respectively; and
[0033] -The signal processor is further configured such that, for the input vector x(t) and the hidden state vector h(t-1) at a given time t, the N of the improved gated loop unit is... ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p.x Less than N ch,x and N p,oh Less than N ch,oh .
[0034] The input vector of the neural network may be based on or include the at least one electrical input signal or a signal derived therefrom. The output vector of the neural network includes the gain (G(k,t)) of the noise reduction algorithm.
[0035] The method of operating the second hearing device is also disclosed (by converting the structural features of the second hearing device into process features).
[0036] Third hearing device
[0037] In one aspect of this application, a hearing device is provided. The hearing device includes:
[0038] - An input unit for providing at least one electrical input signal in a time-frequency representation of k,t, where k and t are the frequency index and time index, respectively, and k represents the channel, k = 1, ..., K, K > 1; the at least one electrical input signal represents sound and includes a target signal component and a noise component; and
[0039] - A signal processor connected to the input unit and configured to receive at least one electrical input signal or a signal derived therefrom, the signal processor including
[0040] --Target signal estimator, used to provide an estimate of the target signal;
[0041] --Noise estimator, used to provide an estimate of the noise level;
[0042] --A gain estimator for providing a corresponding gain value based on a target signal estimate and a noise estimate, wherein the gain estimator includes a neural network, wherein the weights of the neural network have been trained with multiple training signals, and wherein the output of the neural network includes a real-valued or complex-valued gain or separate real-valued gain and real-valued phase;
[0043] The signal processor is configured to calculate, at a given time t, the changes in the input vector x(t) and the hidden state vector h(t-1) from one time t-1 to the next time t. and in and Let x(i,t-1) and h(j,t-2) be the estimated values, respectively, where i and j refer to the i-th and j-th input neurons in the hidden state, respectively, and 1≤i≤N. ch,x and 1≤j≤N ch,oh , where N ch,x and N ch,oh These represent the number of processing channels for the input vector x and the hidden state vector h, respectively; and
[0044] -The signal processor is further configured such that, for the input vector x(t) and the hidden state vector h(t-1) at a given time t, the N of the improved gated loop unit is... ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p.x Less than N ch,x and N p,oh Less than N ch,oh .
[0045] The input vector may be based on or include estimates of the target signal and estimates of the noise, or signals derived from them. The output vector of the neural network includes the gain (G(k,t)) of the denoising algorithm.
[0046] The method of operating the third hearing device is also disclosed (by converting the structural features of the third hearing device into process features).
[0047] Characteristics of hearing devices (and corresponding methods)
[0048] The following feature plan can be combined with the first, second and third hearing devices mentioned above.
[0049] The hearing device can be adapted to further discard N at a given time step. p,x and Np,oh The processing of values among the maximum values that are less than or equal to the threshold limits the number of channels processed to less than N. p,x and N p,oh One channel.
[0050] Hearing devices can be adapted to make the estimated value and These can be equal to the final values of x and h, respectively, which result in N values at the time t' involved. p,x and N p,oh The change in value among the largest changes.
[0051] The improved gated recurrent unit according to the present invention is called a "peak GRU" or "peak GRU RNN". The term "number of peaks N" refers to the number of peaks at a given time t. p,x and N p,oh In this specification, N means the parameters related to the input vector and hidden state at time t (or for hidden state h, t-1). p,x and N p,oh The maximum value. Number of peaks N p,x and N p,oh Can be equal (N) p,x =N p,oh The channel index is denoted as i when combined with the input vector x, and as j when combined with the hidden state vector h (to indicate the number of input nodes N). ch,x The number of output (or hidden state) nodes N is different from that of a delta GRU RNN or peak GRU RNN layer. ch,oh At that time, the processing channel index ranged from 1 to N. ch,x and from 1 to N ch,oh (Varies independently). However, for simplicity, this is generally not applied to all expressions in this invention, where k can be used as a common "channel index". Estimates of the input vector x(i,t) and hidden state vector h(j,t-1). Memorizing from one time step to the next (e.g., making...) and At time step t, Δx(i,t) and Δh(j,t-1) can be determined, for example, see [link to relevant documentation]. Figure 1C Estimated value and It can be equal to the values of x and h at least one earlier time step. (Estimated value) and It can be stored in memory and used in the calculation of time step t. Estimated value and These can be equal to x and h, respectively, the final values that cause the "threshold" change (from one time step to the next), i.e., not (necessarily) the difference since the last time step. Estimated value and Can be equal to x and h respectively, resulting in N at the time t' involved. p (N p,x N p,oh The final value of the change among the maximum changes. The term "N" refers to the change in the value of the modified gated recurrent unit (i.e., peak GRU) for a given input vector x(t) and hidden state vector h(t-1) at time t. ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p.x Less than N ch,x and N p,oh Less than N ch,oh "means (at least) different from the N" p The maximum value of Δx(i,t) (N) ch,x -N p,x ) values and Δh(j,t-1) (N) ch,oh -N p,oh The value is set to 0 in the calculation of a given time step (t) of the peak GRU unit. In other words, the maximum number of processing channels at a given time step is N. p The term "number of update channels" in this specification refers to those channels through which the output value o(k,t) and thus the hidden state value h(k,t) can change (compared to the corresponding value at a previous time step) at a given time step. Each neuron (channel) transmits its value only when the absolute values of the changes (Δx(k,t) and Δh(k,t-1)) "exceed the threshold," for example, exceeding the threshold and / or in N. p (N p,x N p,oh When among the ) maximum values. In other words, the estimated value and It can be equal to the value at the last change that satisfies the peak (and / or threshold) criterion (and can only be stored in that case).
[0052] estimated value and These can be viewed as states. These states are stored respectively in the input of the i-th neuron and the hidden state of the j-th neuron at the last change. The current input x(i,t) and state h(j,t) are compared with these values to determine Δx and Δh respectively. Afterwards, and The value will only be updated when it crosses a threshold (see [Neil et al.; 2018]), or, in the case of this invention, in and When the value is in the peak and / or above the threshold.
[0053] The signal processor can be configured to determine the estimates of the input vector and the hidden state vector as follows:
[0054]
[0055]
[0056] Where i and j refer to the i-th and j-th input neurons in the hidden state, respectively, and 1 ≤ i ≤ N. ch,x and 1≤j≤N ch,oh As mentioned, the number of peak values can be equal, N p,oh =N p,x =N p The number of channels can be equal, N ch,x =N ch,oh =N ch .
[0057] The signal processor can be configured to determine the changes in the values of the input vector and the hidden state vector as:
[0058]
[0059]
[0060] If the number of input nodes differs from the number of output or hidden state nodes in the peak GRU RNN layer, then in the expression above, i and j satisfy 1 ≤ i ≤ N. ch,x and 1≤j≤N ch,oh If the number of peaks differs for the input vector, output vector, and therefore the hidden state vector, the delta value above should be correspondingly related to N in the expression above. p,x and N p,oh related.
[0061] The input unit may include multiple input converters and a beamformer filter, wherein the beamformer filter is configured to provide at least one electrical input signal based on signals from the multiple input converters. The beamformer filter may provide at least one electrical input signal as a spatially filtered signal based on signals from the multiple input converters and pre-determined or adaptively determined beamformer weights.
[0062] The hearing device may include a speech activity detector configured to estimate whether or with what probability an input signal at a given time point includes a speech signal and provide a speech activity control signal indicating the result. An estimate of the current noise level may be provided, for example, when the speech activity control signal indicates a time when speech is absent. The output of the speech activity detector may be used by a beamforming filter. The absolute or relative acoustic transfer function of each input transducer from the sound source to the hearing device (when worn by a user) may be estimated when the speech activity control signal indicates, for example, speech different from that from the user.
[0063] The hearing device may include an output unit configured to provide output stimulation to a user based on at least one electrical input signal.
[0064] The input unit may include at least one input converter, such as a microphone. The input unit may include at least one analog-to-digital converter for providing at least one electrical input signal as a digitized signal. The input unit may include at least one analysis filter bank configured to provide at least one electrical input signal (in time-frequency representation (k,t) or (k,l) or (k,m), where t, l, and m are time exponents) as a sub-band signal. The output unit may include a synthesis filter bank configured to convert the processed time-frequency represented signal (sub-band signal) into a time-domain signal. The output unit may include a digital-to-analog converter for converting the digitized signal (including digital samples) into an analog electrical signal. The output unit may include an output converter configured to provide output stimulation to a user. The output converter may include a loudspeaker or a vibrator. The output converter may include an implanted portion, such as a multi-electrode array configured to electrically stimulate the cochlear nerve at the user's ear.
[0065] The signal processor can be configured to apply a gain value G(k,t) provided by an SNR-gain converter to at least one electrical input signal or a signal derived therefrom. The signal processor may include a combination unit (such as a multiplier) to apply the gain value G(k,t) to at least one electrical input signal IN(k,t) (or a signal derived therefrom) and provide a noise-reduced signal. The signal processor can be configured to process at least one electrical input signal (or a signal derived therefrom) or a noise-reduced signal and provide a processed signal. The signal processor can be configured to apply a compression algorithm configured to compensate for a user's hearing impairment.
[0066] Other configurations and input features are also possible. Peak GRU can also be used for direction of arrival estimation, feedback path estimation, (self)voice activity detection, or other scenario classification.
[0067] The signal processor can be configured to discard N at a given time t. p The absolute value of each channel or Less than threshold Θ pThe processing of the channels provides a combination of Delta GRURNN and Peak GRURNN algorithms. This has the advantage of not processing small values when the selected peaks (numerical values) are small. This maintains an upper limit on processing power while allowing for lower processing power when signal variations are small. The number of peaks can differ for the input vector and the hidden state vector. Similarly, the threshold Θ... p The input vector and hidden state can be different (e.g., Θ and Θ respectively). p,x and Θ p,oh ). In changes (e.g.) In the asymmetric case, two different thresholds (Θ) can be applied. p,x+ ,Θ p,x- ), for example, depending on Is it greater than or less than 0? Threshold (Θ) p or Θ p,x+ ,Θ p,x- For example, it may depend on the given detected sound scene. The aforementioned appropriate threshold can be estimated during training. The feature of discarding channels whose values are less than the threshold can also be combined with the statistical RNN (StatsRNN) described below.
[0068] Peak quantity N p,x and N p,oh It can be adaptively determined based on at least one electrical input signal. The at least one electrical input signal can, for example, be evaluated over a time period. Number of peak values N p (or N respectively) p,x and N p,oh The number of peaks may remain constant for a given acoustic environment (e.g., for a given program on a hearing aid). Hearing aids or headphones may include an acoustic environment classifier. The number of peaks at a given time may depend on a control signal from the acoustic environment classifier that indicates the current acoustic environment. The number of peaks N p (or N respectively) p,x and N p,oh The changes in the input (Δx) and the changes in the hidden state (Δh) can be different.
[0069] The parameters of a neural network can be trained using multiple training signals. A neural network including the peak GRU RNN algorithm according to the present invention can be trained using the baseline GRU RNN algorithm, thereby providing the optimal weights (e.g., weight matrix W) of the trained network. xr W hr W xc W hc W xu W hu See Figure 1A , 1CThe optimized weights can be stored in the hearing device, and peak GRU constraints can be applied to the trained network. The neural network can be trained based on the estimated signal-to-noise ratio as an example of the input obtained from a mixture of noisy inputs and its corresponding (known) output as a vector across the frequencies of the noise-reduced input signal that mainly contains the desired signal. The neural network can also be trained based on an example of the (digitized) electrical input signal IN(k,t), for example directly from the analytical filter bank (see, for example). Figure 3A , 3B In the FB-A network, the corresponding SNR and appropriate gain are known (or, if the SNR estimator is part of a neural network, its appropriate gain is known). Figure 3B ).
[0070] Hearing devices may consist of or include air-conduction hearing aids, bone-conduction hearing aids, cochlear implant hearing aids, or combinations thereof. Hearing aids may be configured to be worn by a user in or in the ear, or may be wholly or partially implanted in the head within the user's ear. Hearing devices may consist of or include headphones.
[0071] Hearing devices may include hardware modules particularly suited for processing elements of gated loop units in a vectorized manner. These hardware modules may, for example, be part of an integrated circuit, such as a digital signal processor (DSP) for the hearing device. The hardware modules may be configured to perform operations on a set of values (a vector) at a time, for example, N. pro Multiple groups, such as N pro = Four elements. The decision to process or discard all elements of this group can be based on logical operations, such as logical operations involving the values of the individual elements of the group (vector). The decision to process or discard may, for example, include whether at least one (e.g., most) of these values is above a threshold (e.g., in the top N). p (Among the values). Different peak counts N p (N p,x N p,oh This can be used for both the input vector and the hidden state vector. If so, the entire vector (all N) pro For example, four values are all processed. If not, all N... pro The values are discarded. In the audio processing example of this invention, this group N... pro Each value can represent a subset of the elements in a vector representing the spectrum (each element of the vector represents the value of a frequency band of the spectrum at a given time point). It is important to consider how the frequency bands are divided into (sub)vectors. Following the original order (frequency bands 1, 2, 3, 4, etc., N) chBand grouping (e.g., 64, 128, or 512 bands) can lead to the loss of information, such as at lower frequencies, if most values within a vector are too small. Therefore, regrouping can be performed (e.g., based on experimentation).
[0072] Hearing devices may be adapted to provide frequency-varying gain and / or level-varying compression and / or frequency shifting (with or without frequency compression) from one or more frequency ranges to one or more other frequency ranges to compensate for a user's hearing loss. Hearing devices may include a signal processor for amplifying the input signal and providing a processed output signal.
[0073] Hearing devices may include an output unit for providing stimulation, perceived as an acoustic signal by a user, based on processed electrical signals. The output unit may include multiple electrodes of a cochlear implant (for CI-type hearing aids) or a vibrator of a bone conduction hearing aid. The output unit may include an output transducer. The output transducer may include a receiver (speaker) for providing the stimulation as an acoustic signal to the user (e.g., in acoustic (air conduction-based) hearing aids or headphones). The output transducer may also include a vibrator for providing the stimulation as mechanical vibrations of the skull to the user (e.g., in bone-attached or bone-anchored hearing aids).
[0074] The hearing device includes an input unit for providing at least one electrical input signal representing sound. The input unit may include an input transducer, such as a microphone, for converting the input sound into an electrical input signal. The input unit may include a wireless receiver for receiving wireless signals that include or represent sound and providing an electrical input signal representing said sound. The wireless receiver may, for example, be configured to receive electromagnetic signals in the radio frequency range (3 kHz to 300 GHz). The wireless receiver may, for example, be configured to receive electromagnetic signals in the optical frequency range (e.g., infrared light 300 GHz to 430 THz or visible light such as 430 THz to 770 THz).
[0075] Hearing devices may include directional microphone systems adapted to spatially filter sound from the environment, thereby enhancing a target sound source among multiple sound sources in the local environment of the user wearing the hearing device. The directional system may be adapted to detect (e.g., adaptive detection) the direction from which a specific portion of the microphone signal originates. This can be achieved, for example, in a variety of different ways described in the prior art. In hearing devices, microphone array beamformers are commonly used to spatially attenuate background noise sources. Many beamformer variations can be found in the literature. Minimum variance distortion-free response (MVDR) beamformers are widely used in microphone array signal processing. Ideally, an MVDR beamformer keeps the signal from the target direction (also known as the line of sight) unchanged while attenuating sound signals from other directions to the greatest extent possible. A generalized sidelobe canceller (GSC) structure is an equivalent representation of an MVDR beamformer, offering computational and digital representation advantages over a direct implementation of the original form. In a binaural configuration, the directional signal may be based on the microphones from the hearing instrument. The estimated SNR may depend on the binaural signals.
[0076] A hearing device may include an antenna and transceiver circuitry (such as a wireless receiver) for wirelessly receiving direct electrical input signals (such as audio signals) from another device, such as an entertainment device (e.g., a television), a communication device (e.g., a telephone), a wireless microphone, or another hearing device. Generally, the wireless link established by the antenna and transceiver circuitry of the hearing device can be of any type. The wireless link can be established between two devices, such as between a communication device and a hearing device, or between two hearing devices, such as via a third intermediary device (e.g., a processing device, such as a remote control, smartphone, etc.). Preferably, the frequency used to establish the communication link between the hearing device and the other device is below 70 GHz, for example, in the range from 50 MHz to 70 GHz, or above 300 MHz, for example, in the ISM range above 300 MHz, or in the 900 MHz range, or in the 2.4 GHz range, or in the 5.8 GHz range, or in the 60 GHz range (ISM = Industrial, Scientific and Medical, such standardized ranges are defined, for example, by the International Telecommunication Union ITU). The wireless link may be based on standardized or proprietary technologies. Wireless links can be based on Bluetooth technology (such as Bluetooth Low Energy).
[0077] Hearing devices can be portable (i.e., configured to be wearable) devices or integral to them, such as devices that include a local power source, such as a battery, for example a rechargeable battery. Hearing devices can be, for example, lightweight, easy-to-wear devices, for example having a total weight of less than 100g, for example less than 20g.
[0078] A hearing aid may include a forward or signal path between an input unit (such as an input converter, for example a microphone or microphone system and / or a direct electrical input (such as a wireless receiver)) and an output unit such as an output converter. A signal processor may be located in this forward path. The signal processor may be adapted to provide frequency-varying gain according to the specific needs of the user. The hearing aid may include an analysis path having functionalities for analyzing the input signal (such as determining level, modulation, signal type, acoustic feedback estimate, etc.). Some or all of the signal processing of the analysis path and / or signal path may be performed in the frequency domain. Some or all of the signal processing of the analysis path and / or signal path may be performed in the time domain.
[0079] Analog electrical signals representing sound signals can be converted into digital audio signals during analog-to-digital (AD) conversion, where the analog signal is sampled at a predetermined sampling frequency or sampling rate f. s Perform sampling, f s For example, in the range from 8kHz to 48kHz (to suit specific application needs) at discrete time points t n (or n) provides digital samples x n (or x[n]), each audio sample passes through a predetermined N b Bit represents the acoustic signal at t n The value of N at time b For example, in a range from 1 to 48 bits, such as 24 bits. Each audio sample therefore uses N. b Bit quantization (resulting in 2^n voltammetry of audio samples) Nb (Number of different possible values). The numerical sample x has 1 / f s The duration of time, such as 50 μs, for f s =20kHz. Multiple audio samples can be arranged in time frames. A time frame can include 64 or 128 audio data samples. Other frame lengths can be used depending on the application.
[0080] Hearing devices may include analog-to-digital (AD) converters to digitize analog inputs (e.g., from an input converter such as a microphone) at a predetermined sampling rate, such as 20 kHz. Hearing devices may also include digital-to-analog (DA) converters to convert digital signals into analog output signals, for example, for presentation to a user via an output converter.
[0081] Hearing devices, such as input units and / or antenna and transceiver circuitry, include time-frequency (TF) conversion units for providing a time-frequency representation of the input signal. The time-frequency representation may include an array or mapping of corresponding complex or real values of the signal in question over a specific time and frequency range. The TF conversion unit may include a filter bank for filtering the (time-varying) input signal and providing multiple (time-varying) output signals, each output signal encompassing a distinctly different frequency range of the input signal. The TF conversion unit may include a Fourier transform unit for converting the time-varying input signal into a (time-)frequency signal. The hearing device considers a frequency range starting from the minimum frequency f. min up to the maximum frequency f max The frequency range can include a portion of the typical human hearing range from 20Hz to 20kHz, such as a portion of the range from 20Hz to 12kHz. Typically, the sampling rate f... s Greater than or equal to the maximum frequency f max twice that, i.e., f s ≥2f max The signals from the forward and / or analytical pathways of the hearing aid can be divided into NI (e.g., uniformly wide) frequency bands, where NI is, for example, greater than 5, greater than 10, greater than 50, greater than 100, or greater than 500, and at least some of these bands are processed individually. The hearing aid can be adapted to process the signals from the forward and / or analytical pathways (NP≤NI) on NP different channels. The channels can have consistent or inconsistent widths (e.g., width increases with frequency), and can overlap or not overlap.
[0082] The hearing device can be configured to operate in different modes, such as a normal mode and one or more specific modes, which may be user-selectable or automatically selected. Operating modes can be optimized for specific acoustic conditions or environments. Operating modes may include low-power modes, where the functionality of the hearing device is reduced (e.g., for energy saving), such as disabling wireless communication and / or disabling specific features of the hearing device. Operating modes may include communication modes such as telephone mode.
[0083] The hearing device may include multiple detectors configured to provide status signals relating to the hearing device's current network environment (such as the current acoustic environment), and / or the current state of the user wearing the hearing device, and / or the current state or operating mode of the hearing device. Alternatively or additionally, one or more detectors may form part of an external device that communicates with the hearing device (e.g., wirelessly). The external device may include, for example, another hearing device, a remote control, an audio transmission device, a telephone (e.g., a smartphone), external sensors, etc.
[0084] One or more of a plurality of detectors can operate on a full-band signal (time domain). One or more of a plurality of detectors can operate on a band-split signal ((time-)frequency domain), for example, in a finite number of frequency bands.
[0085] Multiple detectors may include level detectors for estimating the current level of the signal in the forward path. Detectors may be configured to determine whether the current level of the signal in the forward path is above or below a given (L-) threshold. Level detectors operate on full-band signals (time domain). Level detectors operate on band-split signals ((time-)frequency domain).
[0086] The hearing device may include a voice activity detector (VAD) for estimating whether (or with what probability) the input signal (at a specific point in time) includes a voice signal. In this specification, the voice signal may include speech signals from humans. It may also include other forms of vocalization produced by the human speech system (such as singing). The voice activity detector unit may be adapted to classify the user's current acoustic environment as a "voice" or "no-voice" environment. This has the advantage that time periods including electrophonic signals of human vocalizations (such as speech) in the user's environment can be identified and thus separated from time periods that include only (or primarily) other sound sources (such as artificially generated noise). The voice activity detector may be adapted to also detect the user's own voice as "voice." Alternatively, the voice activity detector may be adapted to exclude the user's own voice from the detection of "voice."
[0087] Hearing devices may include a self-voice detector for estimating whether (or with what probability) a particular input sound (such as speech) originates from the user of the hearing device system. The microphone system of the hearing device may be adapted to distinguish the user's own voice from another person's voice and possibly from non-voice sounds.
[0088] Multiple detectors may include motion detectors, such as accelerometers. Motion detectors may be configured to detect movements of the user's facial muscles and / or bones, such as those caused by speech or chewing (e.g., jaw movements), and provide detector signals indicating these movements. Sensor signals (or signals derived from sensor signals) may also be used as input features for the peak GRU.
[0089] The hearing device may include a classification unit configured to classify the current situation based on input signals from (at least partially) a detector and possibly other inputs. In this specification, "current situation" may be defined by one or more of the following:
[0090] a) Physical environment (including the current electromagnetic environment, such as the presence of electromagnetic signals (including audio and / or control signals) that are planned or unplanned to be received by the hearing device, or other properties of the current environment that are different from acoustics);
[0091] b) Current acoustic conditions (input level, feedback, etc.); and
[0092] c) The user's current mode or state (movement, temperature, cognitive load, etc.);
[0093] d) The current mode or state of the hearing device and / or another device communicating with the hearing device (selected program, time elapsed since the last user interaction, etc.).
[0094] The classification unit may be based on or include neural networks (e.g., recurrent neural networks), such as trained neural networks.
[0095] Hearing devices may include acoustic (and / or mechanical) feedback control (e.g., suppression) or echo cancellation systems.
[0096] Hearing devices may also include other suitable functions for the applications involved, such as compression, noise reduction, voice activity detection (e.g., self-voice detection and / or estimation), keyword detection, etc.
[0097] Hearing devices may include hearing aids, such as hearing instruments, such as hearing instruments adapted to be located at the user's ear or wholly or partially in the ear canal, such as headphones, headsets, ear protection devices or combinations thereof.
[0098] application
[0099] On the one hand, applications of the hearing device described in detail in the "Detailed Description" section and defined in the claims are provided. Applications can be provided in devices or systems including one or more hearing aids (such as hearing instruments), headphones, headsets, active ear protection systems, etc., such as hands-free telephone systems, teleconferencing systems (e.g., including loudspeaker amplifiers), broadcasting systems, karaoke systems, classroom amplification systems, etc.
[0100] Operating methods of audio or video processing devices
[0101] On the one hand, a method of operating an audio or video processing device, such as a hearing device like a hearing aid or headphones, is provided. The audio or video processing device includes at least an input unit and a signal processor for processing the output of the input unit and providing a processed output. The signal processor includes a neural network comprising at least one layer implemented as a modified gated recurrent unit (modified GRU), the modified GRU including a memory in the form of a hidden state vector (h(t-1)). The method includes:
[0102] - Provide at least one electrical input signal in time-frequency representation k,t through the input unit, where k and t are the frequency index and time index respectively, and k represents the channel, k = 1, ..., K, and the at least one electrical input signal represents sound or image data;
[0103] - Provide an input vector x(t) to at least one layer implemented as a gated recurrent unit (GRU) based on the at least one electrical input signal or a signal derived therefrom;
[0104] - The signal processor calculates the changes in the input vector x(t) and the hidden state vector h(t-1) from time t-1 to the next time t at a given time t. and in and Let x(i,t-1) and h(j,t-2) be the estimated values, respectively, where i and j refer to the i-th and j-th input neurons in the hidden state, respectively, and 1≤i≤N. ch,x and 1≤j≤N ch,oh , where N ch,x and N ch,oh These represent the number of processing channels for the input vector x and the hidden state vector h, respectively; and
[0105] - By using a signal processor, the N of the gated recurrent unit can be improved for a given time t, given the input vector x(t) and the hidden state vector h(t-1). ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p.x Less than N ch,x and N p,oh Less than N ch,oh ;
[0106] - The signal processor calculates the output vector o(t) of at least one layer implemented as a gated recurrent unit (GRU) at a given time t based on the input vector x(t) and the hidden state vector h(t-1), wherein the output o(t) at a given time step t is stored as the hidden state h(t) and used to calculate the output vector o(t+1) at the next time step t+1.
[0107] - wherein the processed output is determined by a signal processor based on the output vector or a signal derived therefrom; and
[0108] -The processed output is used to control processing in the audio or video processing device and / or transmitted by the audio or video processing device to another device.
[0109] On the other hand, a method of operating a video processing apparatus, such as a portable video processing apparatus, is provided. The video processing apparatus includes at least an input unit and a signal processor for processing the output of the input unit and providing a processed output. The signal processor includes a neural network comprising at least one layer implemented as a modified gated recurrent unit (modified GRU), the modified GRU including a memory in the form of a hidden state vector (h(t-1)). The method includes:
[0110] - A series of digital images representing a video sequence are provided through the input unit, each image being associated with a specific time (t), and subsequent images representing the image at a subsequent time (t+1). Each image includes a plurality of pixels, which together constitute the image, wherein the image change from one time to the next time is represented by the change of one or more of the plurality of pixels.
[0111] - Provide an input vector x(t) to at least one layer implemented as a gated recurrent unit (GRU) based on the image or signal derived therefrom associated with the specific time (t);
[0112] - The signal processor calculates the changes in the input vector x(t) and the hidden state vector h(t-1) from time t-1 to the next time t at a given time t. and in and Let x(i,t-1) and h(j,t-2) be the estimated values, respectively, where i and j refer to the i-th and j-th input neurons in the hidden state, respectively, and 1≤i≤N. ch,x and 1≤j≤N ch,oh , where N ch,x and N ch,oh These represent the number of processing channels for the input vector x and the hidden state vector h, respectively; and
[0113] - By using a signal processor, the N of the gated recurrent unit can be improved for a given time t, given the input vector x(t) and the hidden state vector h(t-1). ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p.x Less than N ch,x and N p,oh Less than N ch,oh ;
[0114] - The signal processor calculates the output vector o(t) of at least one layer implemented as a gated recurrent unit (GRU) at a given time t based on the input vector x(t) and the hidden state vector h(t-1), wherein the output o(t) at a given time step t is stored as the hidden state h(t) and used to calculate the output vector o(t+1) at the next time step t+1.
[0115] - wherein the processed output is determined by a signal processor based on the output vector or a signal derived therefrom; and
[0116] -The processed output is a revised version of the continuous digital images representing the video sequence, or is used to control processing in the video processing device and / or transmitted by the video processing device to another device.
[0117] The aforementioned methods can be used for data processing, such as data representing time-varying processes, such as audio or video processing, or video sequences represented by multiple subsequent images (or frames). The aforementioned methods may be particularly advantageous in applications with limited processing power, and for example, in applications where the (significant) change in data from one time step to the next is limited, for example, less than 20% of the data. Such examples appear in video images from one frame to the next or in audio data, for example, represented by a spectrogram from one time frame to the next. In audio processing, the methods of the present invention can be used for tasks such as noise reduction (as illustrated in this application), voice activity detection, self-voice detection, self-voice estimation, wake word detection, keyword detection, etc.
[0118] Operating methods of the first and second hearing devices
[0119] On the one hand, a method for operating a hearing device, such as a hearing aid or headphones, is provided. The method includes:
[0120] - Provide at least one electrical input signal in time-frequency representation k,t, where k and t are the frequency exponent and time exponent, respectively, and k represents the channel, k = 1, ..., K. The at least one electrical input signal represents sound and includes the target signal component and noise component; and
[0121] or
[0122] --Provide a target signal-to-noise ratio (SNR) estimate SNR(k,t) for the at least one electrical input signal or a signal derived therefrom in the time-frequency representation; and
[0123] --Convert the target signal-to-noise ratio estimate SNR(k,t) into the corresponding gain value G(k,t) in the time-frequency representation; or
[0124] --Convert the at least one electrical input signal into the corresponding gain value G(k,t) in the time-frequency representation;
[0125] The method further includes:
[0126] - Provide a neural network comprising a layer containing at least one gated recurrent unit, the gated recurrent unit including a memory in the form of a hidden state vector h, wherein the output vector o(t) is provided by the gated recurrent unit based on the input vector x(t) and the hidden state vector h(t-1), wherein the output o(t) at a given time step t is stored as the hidden state h(t) and used to compute the output vector o(t+1) at the next time step t+1.
[0127] The method may include converting a target signal-to-noise ratio estimate SNR(k,t) into a corresponding gain value G(k,t) in the time-frequency representation or converting the at least one electrical input signal into a corresponding gain value G(k,t) in the time-frequency representation via the neural network, wherein at least one layer defined as a gated recurrent unit is implemented as a modified gated recurrent unit, and wherein the method further includes:
[0128] - At a given time t, determine the changes of the input vector x(t) and the hidden state vector h(t-1) from time t-1 to the next time t. and in and Let x(i,t-1) and h(j,t-2) be the estimated values, respectively, where i and j refer to the i-th and j-th input neurons in the hidden state, respectively, and 1≤i≤N. ch,x and 1≤j≤N ch,oh , where N ch,x and N ch,oh These represent the number of processing channels for the input vector x and the hidden state vector h, respectively.
[0129] -The signal processor is further configured such that, for the input vector x(t) and the hidden state vector h(t-1) at a given time t, the N of the improved gated loop unit is... ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p.x Less than N ch,x and N p,oh Less than N ch,oh .
[0130] When appropriately replaced by a corresponding process, some or all of the structural features of the apparatus described above, in detail in the "Detailed Description," or as defined in the claims can be combined with the implementation of the method of the present invention, and vice versa. The implementation of the method has the same advantages as the corresponding apparatus.
[0131] The input vector (x(t)) may be based on or include the target signal-to-noise ratio (SNR) estimate SNR(k,t) of the at least one electrical input signal or a signal derived therefrom.
[0132] This can provide improved hearing devices.
[0133] The method of operating the first hearing device is also disclosed (by converting the structural features of the first hearing device into process features).
[0134] The method of the present invention may include, for example, training the parameters of a neural network with multiple training signals before the hearing device is operated.
[0135] Computer-readable media or data carrier
[0136] The present invention further provides a tangible computer-readable medium (data carrier) storing a computer program including program code (instructions), which, when the computer program is run on a data processing system (computer), causes the data processing system to perform (implement) at least some (such as most or all) of the steps of the methods described above, in detail in the "Detailed Description" and as defined in the claims.
[0137] By way of example, but not limitation, the aforementioned tangible computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to execute or store required program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks include compact discs (CDs), laser discs, optical discs, digital multipurpose discs (DVDs), floppy disks, and Blu-ray discs, wherein these disks typically magnetically copy data while simultaneously being optically copied using lasers. Other storage media include those stored in DNA (e.g., in synthetic DNA strands). Combinations of the aforementioned disks should also be included within the scope of computer-readable media. In addition to being stored on tangible media, computer programs may also be transmitted via transmission media such as wired or wireless links or networks such as the Internet and loaded into data processing systems to run at locations other than tangible media.
[0138] Computer program
[0139] In addition, this application provides a computer program (product) including instructions that, when run by a computer, cause the computer to perform the steps of the methods (methods) described above, in detail in the "Detailed Description" section, and as defined in the claims.
[0140] Data processing system
[0141] In one aspect, the present invention further provides a data processing system, including a processor and program code, the program code causing the processor to perform at least some (such as most or all) of the steps of the methods described above, in detail in the "Detailed Description" section, and as defined in the claims.
[0142] Hearing system
[0143] On the other hand, a hearing device and a hearing system including auxiliary devices are provided, including those described above, described in detail in the "Detailed Description" section, and defined in the claims.
[0144] Hearing systems can be adapted to establish a communication link between hearing devices and assistive devices so that information (such as control and status signals, possibly audio signals) can be exchanged or forwarded from one device to another.
[0145] Auxiliary devices may include remote controls, smartphones, or other portable or wearable electronic devices such as smartwatches.
[0146] The auxiliary device may consist of or include a remote control for controlling the functions and operation of the hearing device. The remote control functions are implemented in a smartphone, which may run an app that enables control of the audio processing device via the smartphone (the hearing device includes a suitable wireless interface to the smartphone, such as Bluetooth or some other standardized or proprietary solution).
[0147] The auxiliary device may be constituted by or include an audio gateway device, which is adapted to receive multiple audio signals (e.g., from an entertainment device such as a TV or music player, from a telephone device such as a mobile phone, or from a computer such as a PC) and to select and / or combine appropriate signals (or combinations of signals) from the received audio signals for transmission to the hearing device.
[0148] The assistive device may consist of another hearing aid or include another hearing device. The hearing system may include two hearing devices suitable for implementing a binaural hearing system, such as a binaural hearing aid system.
[0149] APP
[0150] On the other hand, the present invention also provides a non-transitory application called an APP. An APP includes executable instructions configured to run on an assistive device to implement a user interface for the hearing device or hearing system described above, in detail in the "Detailed Description," and as defined in the claims. The APP can be configured to run on a mobile phone, such as a smartphone, or another portable device enabled to communicate with said hearing device or hearing system. Attached Figure Description
[0151] Various aspects of the invention will be best understood from the following detailed description taken in conjunction with the accompanying drawings. For clarity, these drawings are schematic and simplified, showing only the details necessary for understanding the invention while omitting other details. Throughout the specification, the same reference numerals are used for the same or corresponding parts. Features of each aspect may be combined with any or all features of other aspects. These and other aspects, features, and / or technical effects will be apparent from and illustrated in the following figures, wherein:
[0152] Figure 1A A first graphical illustration shows the computations required to implement a basic gated cyclic unit (GRU);
[0153] Figure 1B A second graphical illustration shows the computations required to implement the basic GRU;
[0154] Figure 1C A graphical illustration shows the computations required to implement the Delta GRU layer;
[0155] Figure 2A This schematically illustrates the method for converting signal-to-noise ratio to gain (see...). Figure 3A An exemplary neural network that attenuates noise (using the SNR2G module in the network) includes a gated recurrent unit (GRU) layer.
[0156] Figure 2B This illustrates the conversion of an input signal representing sound into gain (see [link]). Figure 3B An exemplary neural network that attenuates noise (using the IN2G module in the network) includes a gated recurrent unit (GRU) layer.
[0157] Figure 3A The diagram schematically illustrates the input section of a hearing aid or headphones, which includes a noise-canceling system and comprises an SNR estimator and an SNR-gain module, the latter implemented via a neural network (e.g., ...). Figure 2A (as shown in the image);
[0158] Figure 3B The diagram schematically illustrates the input section of a hearing aid or headphones, which includes a noise-canceling system and comprises an input signal-gain module, the latter implemented via a neural network (e.g., ...). Figure 2B (as shown in the image);
[0159] Figure 4A A first embodiment of a hearing aid or earphone according to the present invention is illustrated schematically, comprising, as shown in the diagram... Figure 3A The input section shown;
[0160] Figure 4B A second embodiment of a hearing aid or earphone according to the present invention is illustrated schematically, comprising, as shown in the figure below. Figure 3B The input section shown;
[0161] Figure 5A A third embodiment of a hearing aid or earphone according to the present invention is illustrated schematically;
[0162] Figure 5B A fourth embodiment of a hearing aid or earphone according to the present invention is illustrated schematically;
[0163] Figure 6 Another embodiment of the hearing aid or headphones according to the present invention is shown;
[0164] Figure 7 The spectrum of the speech signal is shown;
[0165] Figure 8 The training setup of the neural network for the SNR-gain estimator according to the present invention is illustrated schematically.
[0166] Figure 9 The hardware module for parallel processing of vectorized data is illustrated schematically.
[0167] Figure 10 The image shows the training dataset used for a neural network including StatsRNN (StatsGRU layers). A magnified view of a portion of the logarithmic histogram of the data;
[0168] Figure 11A An embodiment of a hearing device according to the present invention is shown, wherein the input to the neural network includes separate target and noise estimates or corresponding magnitude responses of target and noise estimates, or at least noise estimates, or a mixture of noise estimates and noisy input;
[0169] Figure 11B An embodiment of the hearing device according to the invention is shown, wherein the input of the neural network includes the output of a target-preserving beamformer (representing a target estimate) and the output of a target-cancelling beamformer (representing a noise estimate).
[0170] The further applicability of the invention will become apparent from the detailed description given below. However, it should be understood that while the detailed description and specific examples illustrate preferred embodiments of the invention, they are given for illustrative purposes only. Other embodiments of the invention will become apparent to those skilled in the art based on the following detailed description. Detailed Implementation
[0171] The detailed description below, taken in conjunction with the accompanying drawings, serves as a description of various different configurations. This detailed description includes specific details to provide a thorough understanding of several different concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. Several aspects of the apparatus and method are described by various different blocks, functional units, modules, elements, circuits, steps, processes, algorithms, etc. (collectively, “elements”). Depending on the specific application, design constraints, or other reasons, these elements may be implemented using electronic hardware, computer programs, or any combination thereof.
[0172] Electronic hardware may include microelectromechanical systems (MEMS), (e.g., application-specific integrated circuits), microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), gating logic, discrete hardware circuits, printed circuit boards (PCBs) (e.g., flexible PCBs), and other suitable hardware configured to perform the various functions described in this specification, such as sensors for sensing and / or recording the physical properties of the environment, devices, users, etc. Computer programs should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, programs, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description languages, or other names.
[0173] In this invention, a (e.g., deep) neural network is used to determine the gain of a post-filter in a single-channel noise reduction system for a hearing aid or headphones. The post-filter is configured to reduce residual noise in a spatially filtered signal. The spatially filtered signal is, for example, determined to be a linear combination of multiple electrical input signals, such as those from the microphone of a hearing aid or headphones. Post-filters with this purpose have been implemented in various ways, for example as Zener filters (or modifications thereof). The use of neural networks to implement post-filters has been proposed in the prior art, see, for example, EP3694229A1.
[0174] This invention relates to recurrent neural network (RNN) architectures (single-layer or multi-layer) that include memory nodes and so-called gated recurrent units (GRUs). While functionally attractive, such networks can be computationally demanding and are generally impractical for low-power applications such as portable electronic devices, for example, audio processing devices like hearing aids or headphones. In an attempt to limit the processing complexity of such neural networks, recurrent neural network (RNN) architectures called delta networks have been proposed (see, for example, [Neil et al.; 2018]). In a delta recurrent neural network (delta RNN), each neuron transmits its value only if the change at the time of its activation exceeds a threshold. This is particularly efficient when the signal to be processed is relatively stable over time (changes relatively slowly). This is, for example, the case with some audio and video signals.
[0175] Even delta RNNs may not be suitable for use in low-power (portable) devices such as hearing aids because the computational load performed by the network could be considerable. However, the primary concern is that the computational load is unknown, as it will depend on the network's input signal and threshold parameters.
[0176] This invention proposes an improvement to Delta RNN.
[0177] Below, we describe the modified so-called peak GRU RNN. But first, we outline the key elements of the baseline GRU and Delta RNN.
[0178] Baseline GRU, Delta GRU, or Peak GRU RNNs can be used as single-layer (recurrent) neural networks.
[0179] A baseline deep neural network (DNN) consists of several layers. A deep neural network includes at least one (hidden) layer (hereinafter referred to as a DNN) between the input layer and the output layer. At least one of these layers includes a gated recurrent unit (GRU), such as the peak GRU according to the present invention.
[0180] The baseline GRU performs computations on all inputs (see, for example, see below). Figure 1A The peak GRU scheme according to the present invention can skip some input calculations. Since the present invention focuses on low power, it seeks a method to reduce computational load and thus save energy while still maintaining sufficient performance (audio quality).
[0181] Sparsity (such as zeros) in network activations or input data (i.e., sparsity in network parameters (weights, biases, or nonlinear functions), or sparsity in the network's input data) is a property that can be used to achieve high power efficiency. By ignoring computations involving zeros, computational complexity (and thus computational power) can be reduced.
[0182] The following is a basic section of the gated recurrent unit (GRU) (adapted from section 2.1 of [Neil et al.; 2018]):
[0183] Gated Cyclic Unit (GRU)
[0184] The GRU neuron model has two gates (reset gate r and update gate u) and a candidate hidden state c. The reset gate r determines the amount of information from the previous hidden state that will be added to the candidate hidden state c. The update gate u determines the extent to which the activation vector h should be updated by the candidate hidden state c to enable long-term memory. The GRU formula is expressed as follows:
[0185] r(t)=σ[W xr x(t)+W hr h(t-1)+b r (1)
[0186] u(t)=σ[W xu x(t)+W hu h(t-1)+b u (2)
[0187] c(t) = tanh[W xc x(t)+r(t)⊙(W hch(t-1))+b c (3)
[0188] h(t)=(1-u(t))⊙h(t-1)+u(t)⊙c(t) (4)
[0189] Where x is the (external) input vector, h is the activation vector, W is the weight matrix, b is the bias, σ represents the logistic sigmoid function, and ⊙ indicates element-wise multiplication. The data flow of the GRU node model is... Figure 1A , 1B The illustration is shown in the figure (adapted from [Kostadinov; 2017]).
[0190] Figure 1A This represents all the computations required for a GRU node, as written in equations (1)-(4).
[0191] The parameter x(t) is the (external) input, and the parameter h(t) is the layer that will be used as a possible next layer (e.g., a fully connected layer, see [link]). Figure 2A , 2B The output of the GRU (in) Figure 1A , 1B Also denoted as o(t) in 2, see [reference]. Figure 1A The dashed curve arrow denoted as h(t) represents the path from the output o(t) to the internal input h(t-1). In other words, the parameter h reflects the network's memory of past output values. In the figure, the input and output parameters (external (x,o) and internal (h)) are shown in bold to indicate "eigenvectors". In the illustrated example of this invention, the "features" of the "eigenvectors" are embodied in the values of the parameters involved at different frequencies k, k = 1, ..., K (where K may be equal to or different from the number of processing channels N of the GRU). ch For example, see Figure 2A , 2B .exist Figure 2A In the example, the input vector of the neural network, which includes (peak) GRU units as intermediate layers, is a (Kx1) input vector SNR(t) at time t, which represents the spectrum of the signal-to-noise ratio (SNR) estimate at time t, where the vector SNR(t) T Each element of the GRU (SNR(1,t),…,SNR(K,t)) represents the SNR estimate at time t and frequency k, k=1,…,K. The input vector x(t) of the GRU can have the same (Kx 1) or different (N) elements. ch,x The dimension of x 1) depends on the number of output nodes in the preceding layers of the neural network (e.g., x 1). Figure 2A , 2BThe fully connected layer (NN layer) in the GRU can either be a single layer or have the dimension of the input data (if the GRU layer is the first layer and there are no other layers preceding it). Meanwhile, while the GRU processing is not yet complete, the output o(t) = h(t) is used as the internal input h(t-1) for the next time step of the GRU (see, for example, [link to relevant documentation]). Figure 1B The computation time series t-1, t, t+1 of the GRU unit are folded. Any type of neural network (feedforward, recursive, long short-term memory, gated recurrent units, etc.) can provide input and / or output layers to the GRU (including peak GRU). Figure 2A , 2B In this, the fully connected feedforward neural network is labeled as the input and output layers (NN layers), with GRU RNN (e.g., the peak GRU RNN according to the present invention) as the intermediate layer.
[0192] Figure 1A This illustrates how the GRU computes the output (using equations (1)-(4) above). The main components are the two gates (reset gate r and update gate u) and the candidate state c, which are described in more detail below.
[0193] Update Gate u
[0194] In the update gate u, the input x(t) is coupled to its corresponding matrix W. xu (W xu x(t) represents the weight multiplication. In the example of this invention, the weight matrix W xu W hu The matrix is a KxK matrix, where K is the number of frequency bands of the input signal of the GRU (e.g., K = N). ch However, in general, the dimension of the weight matrix adapts to the specific situation based on the dimensions of the GRU's input and output vectors. For example, if there are 64 inputs and 256 outputs (the dimension of the hidden states), then X(W) = ... xu Each element in the kernel (matrix) will have a dimension of 64x256, and the H matrix (W... hu The dimension of each of the matrices (e.g., W) will be 256x256. The same rule applies to the matrix W. hu (W hu The product of h(t-1) and h(t-1) is h(t-1), where h(t-1) contains information from the previous time step. The two results are added together (W). xu x(t)+W hu h(t-1))(and possibly additional deviation parameter b) u Apply the sigmoid activation function (σ) to restrict the result to between 0 and 1 (u(t) = σ[W)). xu x(t)+W hu h(t-1)+b u(See equation (2)). Other activation functions other than sigmoid can also be applied.
[0195] Reset gate r
[0196] The reset gate r is used to determine how much previous information should be forgotten. Its formula is the same as that of the update gate, except for the weight matrix (W). xr W hr ) and possible deviation parameters (b) r Besides the difference (r(t)=σ[W) xr x(t)+W hr h(t-1)+b r See equation (1).
[0197] Candidate state c
[0198] Similar to the previous one, execution and input (W) xc x(t)) and W hc Multiply by h(t-1) (see equation (3)). Then, the element-wise product with reset gate (r(t)⊙(W)) is multiplied. hc h(t-1) determines what information will be removed from previous time steps. The closer r is to 0, the more information will be forgotten. Finally (adding a bias b is optional). c Then, the application restricts the result to the tanh activation function (c(t) = tanh[W)) in the range of -1 to 1. xc x(t)+r(t)⊙(W hc h(t-1))+b c ]) Activation functions other than tanh can also be applied.
[0199] Hidden state h
[0200] The final step is to compute the hidden state h(t), which retains information from the current cell. To obtain h(t), an update gate (u) is needed. The update gate determines what to retain from the current memory content (c) and the previous memory content h(t-1). Therefore, we need to perform element-wise multiplication again, (1-u(t))⊙h(t-1) and u(t)⊙c(t), this time using the update gate u(t). Finally, we sum the results to obtain h(t) (h(t) = (1-u(t))⊙h(t-1) + u(t)⊙c(t)), which will be used as input to the next layer and the next time step (see, for example, [link to previous steps]). Figure 1B ).
[0201] The baseline idea for Gated Recurrent Units (GRUs) is common to both Delta GRU RNN and Peak GRU RNN methods. Peak GRU RNNs are derived from Delta GRU RNN methods, and they share many computational similarities.
[0202] Figure 1B A second graphical illustration shows the computations required to implement the basic GRU, where, as shown... Figure 1A The GRU units shown and described above are "folded out" to indicate the time series (t-1, t, t+1) and illustrate the presence of memory in the GRU (as opposed to feedforward neural networks). Three consecutive GRUs ("units") representing times t-1, t, and t+1 are linked together to illustrate how the output value (o = h) is used as input to the next unit (the next time step). The input (feature) vector x can be a physical entity vector (e.g., a time frame of an audio or video signal), or it can be used to receive output vectors from (possibly) preceding neural network layers (which can be GRU layers or any other type of neural network layer). Similarly, the output (feature) vector o can be a result vector used in audio or video applications (e.g., a frequency-varying gain vector applied to a (noisy) input signal), or it can be used to provide input to another neural network layer, which can be a GRU layer or any other type of neural network layer.
[0203] The following is the basic part of Delta GRU RNN (adapted from part 2.2 of [Neil et al.; 2018]).
[0204] Delta GRU RNN algorithm
[0205] The Delta GRU RNN algorithm reduces memory access and arithmetic operations by leveraging the temporal stability of RNN inputs, states, and outputs. Computations associated with neuron activations that have changed only slightly from their previous time steps can be skipped. Skipping individual neurons saves all related weight matrices (e.g., see the GRU weight matrix described above). xr W hr W xc W hc W xu W hu The multiplication of entire columns and the reading of corresponding weight elements. In matrix-vector multiplication (MxV) between the neuron activation vector (e.g., x or h) and the weight matrix (W), zero elements in the vector result in a zero portion that does not contribute to the final result.
[0206] The maximum element is determined and placed in the delta vector (Δx, Δh). The minimum element is skipped in the hardware (set to 0 in the equation). These delta vectors are then used for multiplication with the noise reduction. Thus, instead of the original (e.g., for external input x(t)) W of the reset state r... xr x(t), we will get W xr Δx(t), where Δx(t) is sparse. The same applies to the activation vector h.
[0207] The key feature of Delta GRU RNNs is that they update the neuron's output only when the neuron's activation vector changes more than the delta threshold Θ. To skip computations related to any small Δx(t), a delta threshold Θ is introduced to determine when delta vector elements can be ignored. Changes in the neuron's activation vector are only remembered when they are greater than Θ. Furthermore, to prevent error from accumulating over time, only the last activation value whose change is greater than the delta threshold is remembered. This is defined by the following set of equations (node exponents omitted):
[0208]
[0209]
[0210]
[0211]
[0212] The final change Remembered and used for the next time period (t+1, see Figure 1C The internal input of ) and the delta vectors Δx(t) and Δh(t-1) are estimated using previous values. and The calculation is performed at each time step t. The Θ thresholds in equations (5) and (6) or (7) and (8) can be the same, or the same across frequencies. However, the Θ values can be different, for example, for x(Δx) (equations (5), (7)) and h(Δh) (equations (6), (8)). The current input x i,t (x(t)) (where i is the i-th element (node) of the input vector x) and state h j,t (h(t)) (where j is the j-th element (node) of the hidden state vector h) will be compared (e.g., subtracted) with these values to determine the corresponding Δ value. Afterwards, and The value will only be updated when it crosses the threshold. The exponents i and j have been omitted in the equations (5)-(8) above.
[0213] In the modified Delta GRU, asymmetric thresholds can be applied. In another modified Delta GRU, different thresholds can be applied across different neurons in a layer.
[0214] Next, using equations (5), (6), (7) and (8), the traditional GRU equation set (above (1)-(4)) can be transformed into its delta network version:
[0215] M r (t)=Wxr Δx(t)+W hr Δh(t-1)+M r (t-1) (9)
[0216] M u (t)=W xu Δx(t)+W hu Δh(t-1)+M u (t-1) (10)
[0217] M xc (t)=W xc Δx(t)+M xc (t-1) (11)
[0218] M hc (t)=W hc Δh(t-1)+M hc (t-1)) (12)
[0219] r(t)=σ[M r (t)] (13)
[0220] u(t)=σ[M u (t)] (14)
[0221] c(t) = tanh[M xc (t)+r(t)⊙M hc (t)] (15)
[0222] h(t)=(1-u(t))⊙h(t-1)+u(t)⊙c(t) (16)
[0223] Among them, M r M u M xc and M hc These refer to the stored (memory) values of the reset gate r, update gate u, candidate state c, and hidden state h, respectively, where M r (0)=b r M u (0)=b u M xc (0)=b c and M hc (0) = 0. The above treatment of the time exponent t (denoted as equations (9) to (18)) is in Figure 1C The graph is shown in the middle. The value M for the previous time point (t-1) is shown on the left side of the graph. r (t-1),M u (t-1),M xc (t-1),M hc(t-1) is stored in the preprocessing cycle (usually overwritten by the new value of the parameter used in the subsequent processing cycle).
[0224] In a Delta GRU, there might be, for example, six different weight matrices (other representations are also possible). These weight matrices could be W... xr W hr W xu W hu W xx and W hc (see Figure 1A x and h indicate whether the matrix is related to the input x or h, and u, r, and c represent update, reset, and candidate, respectively. The subscripts are used to distinguish what kind of input-matrix pair the output comes from. If we write W... xr We know that this matrix will be used in the computation with input vector x and is related to the reset gate (see, for example, M obtained in equation (9)). r (t) is used in the calculation of the reset gate (r) in equation (13).
[0225] In summary, the Delta GRU RNN algorithm sets a specific threshold Θ, such as 0.1 (see Θ in equations (5)-(8) above), which will be used for comparison with the element value. If the value of the currently processed element (the absolute value of the subtraction) is lower than this threshold, the element will not be used for further computation at that time step, but will be skipped (set to 0). Therefore, the number of elements to be further processed can vary from one time step to another depending on the element value.
[0226] Peak GRU RNN
[0227] The following describes a peak GRU RNN according to the present invention, which is an improved version of GRU and Delta GRU RNN.
[0228] The advantage of the Peak GRU RNN algorithm lies in its deterministic reduction of computation (i.e., it always knows in advance how many operations will be performed by the algorithm given its configuration). This is achieved by setting a hard limit, which determines how many values (peaks) in each input feature vector of the Peak GRU RNN (e.g., in each time frame of an audio or video signal) will be processed. Instead of setting a threshold for the rate of change of the elements of the input vector (as in Delta GRU), a predetermined number (N) of elements are selected to be processed in each iteration. p Therefore, only the elements of the matrix involved (such as W) are needed. xr W hr W xu W hu W xc W hc N pThe column. In each iteration, the N with the largest rate of change (delta) of the input vector. p Each element is selected for further processing. This allows for the flexibility of not having to handle the multiplication of the entire vector matrix, and is deterministic (total N). p (dot product). Peak GRURNN (and Delta GRURNN, e.g., if there are no zeros in the (delta) input vector) can perform full vector matrix multiplication (e.g., N). p =N ch ) or its subsets, down to zero multiplication (e.g., N) p =0). Threshold N p The number of the largest absolute values / quantities that will (or can) be processed for the (delta) input vector (Δx(t), Δh(t-1)) corresponding to the peak GRU RNN layer. As noted above, the number of peaks N p (N p,x N p,oh For the input vector (x)(N) p,x ) and output (and hidden state) vectors (o,h)(N p,oh ) can be equal (N) p =N p,x =N p,oh (or different)
[0229] Update and (e.g., N) p The neuron corresponding to the largest Δ(Δx,Δh) is assumed to be of equal importance. However, Δ can be weighted by an "importance" factor, which relates to the change in the value function obtained during training, see the "Training" section below.
[0230] The weight matrix (e.g., W) of the peak GRU RNN (and similarly, the delta GRU RNN and the baseline GRU RNN) xr W hr W xu W hu W xc W hc This is usually fixed to a value optimized during the training process.
[0231] The lower the peak count, the less computation and memory access is required, and therefore the lower the power consumption.
[0232] The improvement stems from equations (5) to (8) of the Delta RNN algorithm. Instead of the requirements of the Delta RNN algorithm, it satisfies... and The criterion for values greater than the threshold Θ is determined by the peak GRU RNN algorithm according to the present invention at a given time point t. and N pOnly the maximum values are processed. At different time points…,t-1,t,t+1…, and N p The maximum value can be associated with elements of different vectors / positions (indices) (e.g., different frequencies in the example of this invention) in the input vector and the hidden state vector.
[0233] Using absolute values |·| allows us to compare quantities based on their magnitudes, and the delta vectors are then assigned the actual results of the subtraction (corresponding to their magnitudes above a threshold). This also works for delta GRUs (but the number of values processed can be different).
[0234] Instead of equations (5) and (6) for delta GRU RNNs, the following expression can be used for peak GRU RNNs:
[0235]
[0236]
[0237] Where i and j refer to the i-th input and the hidden state of the j-th neuron (where 1≤i≤K and 1≤j≤K, and K is the number of frequency bands in the exemplary (input and output feature vectors) application scenario (audio processing) of this invention).
[0238] Similarly, equations (7) and (8) for delta GRU RNNs can be used for peak GRU RNNs, including the predetermined number of elements (N) to be processed in each iteration. p Related improvements:
[0239]
[0240]
[0241] Elements with different indices in a vector can be selected at different time steps. For example, at time step (t-1), we can have the following exemplary vector, where the boldbody values are selected (based on the two maximum values of their magnitudes, N). p =2,N ch =5): [1,0,-9,8,-4], and at time step t, we can have the vector [2,-5,0,1,-1], where the two other exponents in the vector are chosen because they (numerically) represent the maximum value.
[0242] For x and h, we choose the maximum value separately, which means we will have N p Each Δx peak and N p N' (distinct values) p The peak value of Δh.
[0243] In summary, peak GRU RNNs use a different type of limit operation (however, the subtraction in equations (5)-(8) above is retained). It uses a hard limit equal to the number of elements that will always (or at most) be processed at each time step. The elements to be processed are selected based on their absolute values. For example, if we set the peak number to 48 in a 64-input-node system, 64-48 = 16 of the smallest absolute values will always be skipped at each time step.
[0244] Explain the two methods (or (referred to as the "subtracted element") and (or Examples of values for (referred to as ABS (subtracted) elements) are shown below (by simplifying by using only 10 elements, N) ch =10):
[0245] Subtracted elements: [0, -0.6, 0.09, 0.8, 0.1, 0, 0, -1.0, 0.05, 0.22]
[0246] Elements after ABS (subtraction): [0, 0.6, 0.09, 0.8, 0.1, 0, 0, 1.0, 0.05, 0.22]
[0247] Delta GRU RNN
[0248] If, for example, the threshold Θ = 0.06, elements with values less than or equal to 0.06 will be set to 0 (e.g., zeros with an underscore). 0 As shown in the figure: [ 0 -0.6, 0.09, 0.8, 0.1, 0 , 0 -1.0 0 ,0.22].
[0249] Peak GRU RNN
[0250] If, for example, we want to process N peak values p =6, then we will skip 10–6 = 4 smallest absolute values (set to 6). 0 ): [ 0 -0.6, -0.09, 0.8, 0.1, 0 , 0 -1.0 0 ,0.22].
[0251] In this example (at this moment), the Delta GRU method and the Peak GRU method are computationally equally efficient. This illustrates that comparing the difficulty of the two methods is based solely on the number of data values removed by the respective methods (at a given moment). This depends on Θ and N, respectively. p The values and the data processed by the corresponding algorithms are also considered. Therefore, without knowing the dataset or how it affects the final result, it's impossible to say which method is better simply by looking at how many values are removed. If we set the number of peaks to, for example, 3 in this example, the peak GRU RNN method would appear to be more efficient.
[0252] Potential advantages of peak GRU
[0253] The Delta GRU RNN algorithm operates by thresholding the actual values that each element is compared to. Therefore, if all values are above the threshold (the threshold is not set high enough), all elements will be processed, and all the additional computations compared to the baseline GRU will contribute to higher power consumption rather than lower power consumption. The Delta GRU RNN is expected to work well in environments where changes between time steps are not so abrupt (e.g., quieter environments).
[0254] Generally, Delta GRU sets a threshold for the rate of change of the elements in the input vector, while Peak GRU selects a predetermined number of elements from the input vector, where the selected elements have the largest rate of change.
[0255] If the number of skipped peaks is not high enough, the peak GRU RNN algorithm may also contribute to increased power consumption. However, in general, the peak GRU RNN algorithm:
[0256] - A specific number of elements are always removed, regardless of whether they are above or below a certain numerical threshold. Therefore, compared to Delta GRU RNN, Peak GRU RNN is a deterministic method, meaning we can always predefine how much computation will be performed. This is a crucial aspect for low-power devices such as hearing instruments;
[0257] Because the Peak GRU RNN algorithm doesn't use a static threshold like the Delta GRU RNN, but instead works only with the number of the first few elements (i.e., the "numerical threshold" is dynamic and adjusted at each time step), it is a robust method for preprocessing and can process datasets without existing analysis of the dataset values. If we need to apply preprocessing to the data, causing the data to change from one representation to another (e.g., quantization, or some simple filtering, normalization, etc.), the Delta GRU threshold will no longer work, and it will have to be remapped. However, for Peak GRU, no additional adjustment is needed because the order of the data elements will remain the same after preprocessing.
[0258] Similarly, like the Delta GRU RNN algorithm, the Peak GRU RNN algorithm saves computation compared to the Baseline GRU RNN algorithm.
[0259] Example: SNR - Gain Estimator
[0260] Figure 2A This schematically illustrates, for example, the conversion of signal-to-noise ratio to gain in a hearing aid (see...). Figure 3A An exemplary neural network (using the SNR2G module in the example). This neural network is a deep neural network that includes gated recurrent unit (GRU) layers as hidden layers. Figure 2A The deep neural network comprises corresponding fully connected feedforward neural network layers (NN layers) as input and output layers, with GRU RNNs (GRU, such as the peak GRU RNN according to the invention) as intermediate layers. In the example shown, the "features" of the "input feature vector" of the input layer (NN layer) are represented by the signal-to-noise ratio (SNR)(k,t) at different frequencies k, k = 1, ..., K. The input vector is written as (Kx1) input vector SNR(t). T =(SNR(1,t),…,SNR(K,t)), superscript T The term refers to its transpose and can be viewed as the spectrum representing an estimate of the signal-to-noise ratio (SNR) at time t. Figure 2A In the example, the output of the first input layer (NN layer) is denoted as x(t) and is the input vector of the GRU unit (GRU). The GRU input vector x(t) can be a (Kx1) vector. The GRU layer provides an output vector o(t), which can also be a (Kx1) vector. (The last sentence appears to be incomplete and requires further context.) Figure 1A , 1B As described in 1C, the GRU includes memory of past outputs from the GRU, because the output o(t) = h(t) is used as the internal input h(t-1) for the next time step (see 1C). Figure 2A The dashed arrow from o(t) to h(t-1) (see, for example, the dashed arrow from o(t) to h(t-1)). Figure 1B (where the time series t-1, t, t+1 of the GRU unit computation are not folded). The output vector o(t) of the GRU layer is fed to the output layer (NN layer), which provides an output vector G(t) representing the "gain spectrum" corresponding to the input SNR vector SNR(t). T =(G(1,t),…,G(K,t)). In Figure 4A In an exemplary hearing aid, a gain G(t) = G(k,t), k = 1, ..., K, is applied to the signal in the forward (audio) path (thus providing, for example, an enhanced (noise-reduced) signal).
[0261] The multiple dashed arrows in the (fully connected, feedforward) input and output layers (NN layers) of the SNR2G estimator indicate that the gain in the k-th channel (at a given time point t) depends on at least one estimated SNR value, such as some or all of the K channels, i.e., for example...
[0262] G(k,t)=f(SNR(1,t),…,SNR(k,t),…,SNR(K,t))
[0263] This characteristic may also be inherent in GRU layers. In other words, the (deep) neural network according to the invention can be optimized to find the best mapping from a set of SNR estimates (SNR(k,t)) across frequencies to a set of frequency-varying gain values (G(k,t)).
[0264] Figure 2B This illustrates the conversion of an input signal (directly) representing sound into gain (see...). Figure 3B An exemplary neural network (using the IN2G module in the example) to attenuate noise includes a gated recurrent unit (GRU) layer. Instead of using an SNR estimate as input, Figure 2B The neural network directly takes the signal from the filter bank (or its processed version). Figure 2B Therefore, it can be considered to represent the entire noise reduction system (NRS).
[0265] exist Figure 2A , 2B In this design, fully connected feedforward neural network layers (NN layers) serve as input and output layers “around” the GRU RNN (GRU). Any type of neural network (feedforward, recursive, long short-term memory, gated recurrent unit, etc.) can provide input and / or output layers to the GRU (e.g., the peak GRU according to the invention). Furthermore, the GRU does not need to be a hidden layer. Additionally, three or more layers can be used to implement an SNR2G estimator (or other functional units of an audio or video processing device).
[0266] GRU RNN (GRU) can be a peak GRU RNN according to the present invention. In this specification, the number of peaks in the peak GRU RNN algorithm corresponds to the total number of processing channels (N). ch The number of channels (N) processed by the RNN (at a given time point) p ), N p ≤N ch The number of (input or output) channels (e.g., N) ch (or N) ch,x N ch,oh The number N can be any number (greater than 2), for example, between 2 and 1024, between 4 and 512, or between 10 and 500. This means we can determine the value of N based on N.p Value processing N ch N ch -1,N ch -2,… down to 0 channels. Figure 2A , 2B In the context of the peak GRU layer, the number of channels N ch K can be equal to the number of bandwidths in the filter bank FB-A, or it can be greater than or less than K depending on the characteristics of the input and output layers (NN layers) of the neural network. In the exemplary context of the present invention, the neural network is preferably configured to receive an input vector and provide an output vector, both of which have K dimensions (as the time frame of the input signal IN(k,t) (or SNR(k,t)) and the corresponding (estimated) noise reduction gain G(k,t)). K is the number of bandwidths in the processing path of the hearing device (e.g., 24 or 64, etc.), in which the neural network is implemented (see, for example, Figures FIG. 4A, 4B, 5A, 5B). As mentioned above, intermediate layers (e.g., peak GRU layers) can of course have other numbers of nodes (less than or greater than the number of input and output nodes).
[0267] Number of channels N to be processed ch Reduce to N p ≤N ch (e.g. N) p <N ch For example, considering (updating) the number of units (k,t) at any given point in time will lead to energy savings, but naturally, it also affects the resulting audio quality. Therefore, it is necessary to find a reasonable number of peaks, for example, based on the current acoustic environment. One option is to choose a fixed number N, for example, based on simulations of different acoustic environments. p (For example, using speech intelligibility (SI) as a metric to make the number of peaks N) p (The given choice meets the conditions). Therefore, the optimal number of peaks N for each environment can be found. p (Each peak count is applied, for example, in different hearing aid programs specifically designed for a given sound environment or hearing situation, such as speech in noise, speech in quiet, at a party, with music, in a car, on an airplane, in an open office, in an auditorium, in a church, etc.) Alternatively, a general peak count N can be found. p The peak number that provides the maximum average SNR in all environments.
[0268] Peak quantity N p The number of peaks N can be constant for a given application or adaptively determined based on the input signal (e.g., evaluated over a time period). pFor a given acoustic environment (e.g., for a given program of a hearing aid), it can remain constant. A hearing aid or headset may include an acoustic environment classifier. The number of peaks at a given time may depend on a control signal from the acoustic environment classifier that indicates the current acoustic environment.
[0269] To optimize the number of peaks for a given application, simulations can be performed using an appropriate dataset, for example, by increasing the number of peaks we want to skip and observing how it affects SNR, estimated gain, speech intelligibility metrics (and / or other metrics). For both peak GRU RNN and delta GRU RNN, peak GRU RNN (and delta GRU RNN) perform exactly the same as baseline GRU when we process all peaks and when the threshold is 0, respectively. As we start dropping some values, at some point these metrics will gradually begin to deteriorate until they reach a point where metrics such as SNR become too low and unacceptable.
[0270] The optimal number of layers in a neural network depends on the dataset, the complexity of the problem we are trying to solve, the number of neurons per layer, and the type of signal we feed into the network. Three layers may be sufficient, or four or five layers may be necessary. In the example of this invention, the number of neurons per layer is kept at the number of channels K. However, this is not mandatory.
[0271] Figure 3A The diagram schematically illustrates the input section of a hearing aid or headphones, which includes a noise-canceling system and comprises an SNR estimator and an SNR-gain module, the latter implemented via a neural network (e.g., ...). Figure 2A (As shown in the diagram). The input section includes an input unit comprising at least one input converter (here, a microphone M) for providing at least one electrical input signal IN(t) and at least one analysis filter bank FB-A for providing at least one electrical input signal IN(t) in time-frequency representation of IN(k,t), where k and t are the frequency and time exponents, respectively. The frequency exponent k represents the channel, k = 1,…,K (e.g.,…). Figure 2A The number of channels can vary in different parts of the device, for example, more in the forward audio path than in the analysis path (or vice versa), see, for example, see Figure 4A,4B(exponent k,k'). At least one electrical input signal (IN(t),IN(k,t)) represents sound and may include a target signal component and a noise component. The target signal component is the signal component originating from a sound source that the user (currently) of the hearing aid or headphones may be interested in (e.g., speech from people around the user). The input section also includes a noise reduction system NRS (typically part of the analysis path of the hearing aid or headphones) that aims to reduce the noise component in at least one electrical input signal (or the signal derived therefrom, such as a spatially filtered (beamforming) signal). The noise reduction system is configured to provide a gain G(k,t) (k = 1,…,K) for application to the electrical input signal (IN(k,t) or the signal derived therefrom), see, for example, see Figure 4A The noise reduction gain G(k,t) (k=1,…,K) is suitable for reducing (attenuating) the noise component while preserving the target signal component unchanged (or attenuating the target signal component less). The noise reduction system NRS includes a signal-to-noise ratio estimator SNR-EST and a signal-to-noise ratio to gain converter SNR2G. The signal-to-noise ratio estimator SNR-EST receives at least one electrical input signal (IN(k,t)) and provides an estimate of the signal-to-noise ratio of that electrical input signal (SNR(k,t), k=1,…,K). The SNR estimate (SNR(k,t)) can be based on any existing technical method, such as determining it as the observed (available) noisy electrical input signal IN(k,t) (including a mixture of the target signal S and noise N, IN(k,t)=S(k,t)+N(k,t)), for example, picked up by one or more microphones, see, for example, see Figure 3A The SNR estimate is the ratio of the power of the noisy signal IN(k,t) at a given time point (e.g., at a given time frame) to the power estimate of the noise signal. In other words, or t0 refers to the time before the "current" time t, where the noise estimate (update) can be performed, for example, when the noise level is estimated to be absent in the electrical input signal IN(k,t). Preferably, t0 can be the last time index before the current time index, where the noise level has already been estimated.
[0272] The signal-to-noise ratio estimator (SNR-EST) can be implemented as a neural network. The SNR-EST can be included in the signal-to-noise ratio to gain converter (SNR2G) and implemented via a recurrent neural network according to the invention. Figure 3A In an exemplary embodiment, this would mean that the noise reduction system module NRS will be implemented as a recurrent neural network according to the invention, for example, receiving its input vector directly from the analysis filter bank FB-A as IN(k,t), for example, frame by frame (where t represents the time frame index).
[0273] The signal-to-noise ratio estimator (SNR-EST) provides an SNR estimate (SNR(k,t)) to the signal-to-noise ratio to gain converter (SNR2G, RNN). For the corresponding channels (k = 1, ..., K), consecutive time frames of the signal-to-noise ratio (SNR(k,t)) are used as input vectors for the SNR to gain converter (SNR2G), which is implemented as a deep neural network, particularly a peak GRU recurrent neural network according to the present invention. The neural network (RNN) of the SNR to gain converter (SNR2G, RNN) includes an input layer, multiple hidden layers, and an output layer (see, for example, [link to relevant documentation]). Figure 2A The output vector from the output layer includes a frequency-varying gain G(k,t), which is configured to be applied to (e.g., digitized) electrical input signals (or signals derived therefrom) to provide a noise-reduced signal (see, for example, [link to relevant documentation]). Figure 4A The signal OUT(k',t) in the vector G(k,t) can be written as (vector) G(t) = G(k,t), k = 1, ..., K.
[0274] Figure 3B The diagram schematically illustrates the input section of a hearing aid including a noise reduction system (NRS), which includes an input signal-gain module (IN2G) implemented via a neural network (e.g., as shown in the diagram). Figure 2B (as shown in the image).
[0275] Figure 4A A first embodiment of a hearing aid (or, if the speaker is replaced by a transmitter, the microphone path of an earphone) according to the present invention is shown, comprising as follows: Figure 3A The input section is shown. The SNR-to-gain converter SNR2G can be pressed as shown. Figure 2A The implementation shown is as follows.
[0276] In this invention, a hearing aid or headphone including a noise reduction system is described, the noise reduction system including an SNR-to-gain conversion module (SNR2G) implemented as a recurrent (possibly deep) neural network (RNN) according to the invention. The hearing aid includes a forward (audio) signal path, which includes at least one input converter (e.g., a microphone M) providing at least one electrical input signal IN(t) representing sound in the hearing aid or headphone environment. The forward path also includes an analysis filter bank FB-A for converting the (time-domain) electrical input signal IN(t) to K channels IN(k,t), where k = 1, ..., K are frequency exponents and t is a time (frame) exponent. The hearing aid also includes an analysis path, which includes a noise reduction system NRS (see dashed box) configured to reduce noise in the (noisy) electrical input signal (IN(k,t), or a signal derived from it) to provide the user with a target signal (e.g., speech from a communication partner) presumably present in the noisy electrical input signal, thereby providing better quality audio to the user. The noise reduction system includes an SNR estimator, SNR-EST, configured to estimate the signal-to-noise ratio (SNR(k,t)) of the corresponding channel (k) of the (frequency domain) electrical input signal IN(k,t), see, for example, [link to relevant documentation]. Figure 3A The analysis pathway can have its own analysis filter bank (FBA), such as... Figure 4A As shown in the diagram. This is suitable, for example, when the number of channels in the analysis path (e.g., K) is different from (greater than or less than) the number of channels in the forward (audio) path (e.g., K'). If the number of channels is the same as in the forward path (K = K'), the analysis path can use the same analysis filter bank (FB-A) as the forward path. If the number of channels in the analysis path is less than the number of channels in the forward path, a "bandwidth" unit can be introduced between the analysis filter bank (FB-A) of the forward path and the input of the noise reduction system NRS. The bandwidth summation unit can be adapted to combine multiple channels of the forward path into a single channel of the analysis path, such that the resulting number of channels K in the analysis path is less than the number of channels K' in the forward path.
[0277] An example of a three-layer recurrent neural network for implementing an SNR-to-gain converter (SNR2G, RNN) is included in Figure 2A As shown in the diagram, the three layers are a) a fully connected input layer, b) a hidden GRU layer (e.g., a peak GRU RNN layer), and c) a fully connected output layer. The SNR-to-gain module utilizes information across different channels to improve the noise reduction system by setting the gain estimate of the k-th channel to depend not only on the SNR in the k-th channel but also on the SNR estimates of multiple adjacent channels, such as all channels.
[0278] Figure 4B A second embodiment of the hearing aid according to the present invention is illustrated schematically, comprising, as shown in the figure... Figure 3BThe input section shown in the figure includes a direct transformation (typically attenuation) of the subband signal (IN(k,t)) from the analysis filter bank FB-A to the noise reduction gain G(k,t).
[0279] Figure 4A and 4B Each of the hearing device embodiments shown includes a forward path comprising multiple combining units (here, multiplication units 'x') for applying a frequency-varying gain G(k,t) to the input signal IN(k',t) of the forward (audio) path. If the number of channels K of the analysis path providing the frequency-varying gain G(k,t) differs from the number of channels K' of the forward path providing the input signal IN(k',t), an implicit "bandwidth distribution" (or "bandwidth summation") unit is included between the output (G(k,t)) of the noise reduction system NRS and the combining unit ('x') to adapt the gain signal G(k,t) to the number of channels of the input signal IN(k',t). The output of the combining unit ('x') is the resulting noise-reduced signal OUT(k',t). The forward path also includes a synthesis filter bank (FB-S) for converting the sub-band signal OUT(k',t) into a time-domain signal OUT(t). The forward pathway also includes an output transducer (here, a loudspeaker (SPK)) for converting the output signal OUT(t) into a stimulus that can be perceived by the user as sound (here, an acoustic signal including vibrations in the air). The output transducer may include a vibrator providing bone conduction stimulation or a multi-electrode array of a cochlear implant providing electrical stimulation of the cochlear nerve.
[0280] Hearing aids / earphones may include additional circuitry to implement other functions of the hearing aids / earphones, such as an audio processor for applying a gain that varies with frequency and level to the signal in the forward (audio) path, thereby, for example, compensating for the user's hearing loss. The audio processor may be located, for example, in the forward path, between the combining unit ('x') and the synthesized filter bank (FB-S). Hearing aids may also include analog-to-digital converters and digital-to-analog converters, provided they are appropriate for the application in question. Hearing aids may also include antenna and transceiver circuitry to enable the hearing aid or earphone to communicate with other devices (e.g., contralateral devices, such as the contralateral hearing aid or earphone portion), for example, establishing a link to a distant communication partner, such as via a mobile phone.
[0281] Figure 5A A second embodiment of the hearing aid according to the present invention is shown. Figure 5A Implementation examples and Figure 4AThe implementation is similar, but includes additional input converters (M1, M2) and beamformer BF (see dashed box). The hearing aid HD includes two input converters (microphones (M1, M2)) that provide corresponding (time-domain) electrical input signals (IN1, IN2). Each electrical input signal undergoes analog-to-digital conversion before being presented to the analysis filter bank (FB-A), which provides corresponding (time-varying) sub-band signals IN1(k) and IN2(k) representing the sound in the user's environment, as indicated by the thick arrow denoted as K. Figure 5A The hearing aid includes a beamformer BF adapted to provide a spatially filtered (beamforming) signal YBF(k) based on input signals IN1, IN2 (or signals derived therefrom) and (fixed and / or adaptively updated) beamformer weights. The beamformer includes a fixed beamformer module Fx-BF that provides multiple fixed beamformers (here, two, C1, C2, e.g., a target-preserving beamformer and a target-cancelling beamformer) based on the input signals IN1, IN2. The beamformer also includes an adaptive beamformer ABF and a voice activity detector VAD. The voice activity detector is configured to provide a control signal VA indicating whether (or with what probability) the input signal (here, a signal from fixed beamformer C1, e.g., a target-preserving beamformer) includes a voice signal at a given time point. The adaptive beamformer ABF is configured to provide a spatially filtered signal YBF(k) based on signals from the fixed beamformers C1, C2. An adaptive beamformer (ABF) is adapted, for example, to update its filter weights based on a control signal (VA) from a voice activity detector. An estimate of the noise field around the user can be determined, for example, when no voice is present. The algorithm for adaptively updating the filter weights of the adaptive beamformer (ABF) can be, for example, a minimum variance distortion-free response (MVDR) algorithm or a similar algorithm, such as one based on statistical methods and one or more constraints.
[0282] The hearing aid also includes the noise reduction system NRS according to the invention, for example, combined with Figure 3A The noise reduction system described above. The noise reduction system NRS provides the post-filter gain G(k) based on the output of the fixed beamformers (C1, C2) (in Figure 3A and 4A Let G(k,t) be the denoting factor. Signals from fixed beamformers (e.g., target-preserving beamformers (C1) and target-cancelling beamformers (C2)) can form the basis for estimating the signal-to-noise ratio (SNR) on a time-frequency (k,t) basis. For example, combining... Figure 2A and 3AThe SNR estimate can be provided as input to an SNR-gain estimator (SNR2G). Alternatively, the signal from a fixed beamformer (e.g., a target-preserving beamformer (C1) and a target-cancelling beamformer (C2)) can be directly fed into a neural network for estimating the appropriate gain (G(k)), such as in combination with... Figure 11B As described. Figure 5A As indicated by the dashed arrow, the beamforming signal YBF(k) can also be used as input to the noise reduction system (e.g., to improve the SNR estimate). The post-filter gain G(k) is applied in the combining unit ('X') to K sub-bands of the spatially filtered signal YBF(k), thereby providing a noise-reduced signal YNR(k). The noise reduction system NRS and the combining unit ('X') provide a (single-channel) post-filter (PF, see below). Figure 5A The function of the dashed box in the middle.
[0283] Figure 5B A fourth embodiment of a hearing aid according to the present invention is schematically illustrated, wherein the noise reduction system includes a combined beamformer and a noise reduction system (post-filter). In this embodiment, the noise reduction system may be implemented as a recurrent neural network (RNN), such as... Figure 2B , 3B As shown, one of the microphone signals (here, M1, for example, the front microphone of the BTE section of a hearing aid) is taken as the input to the neural network. The resulting gain G(k) is applied to the two sub-band signals IN1(k) and IN2(k). Thus, the noise-reduced electrical input signals are summed in the combining unit ('+') to form the beamforming signal. The beamformer-noise reduction unit (BF-NR) may include a fixed or adaptive beamformer to transform the signal Y... NR (k) provides a spatially filtered (and noise-reduced) signal.
[0284] like Figure 5A and 5B As shown in the embodiments, the hearing aid HD may also include an audio processor PRO for applying additional processing algorithms to the signal in the forward (audio) path (e.g., YNR(k)), such as a compression amplification algorithm to compensate for the user's hearing loss. Similarly, algorithms for processing feedback control, voice interface, etc., may be applied in the audio processor PRO. The processor PRO provides the resulting output signal OUT(k), which is fed to a synthesis filter bank FB-S. The synthesis filter bank FB-S converts the sub-band output signal OUT(k) into a time-domain signal OUT, which is fed to the output converter (here, a speaker) of the hearing aid. In other embodiments, other output converters may be suitable, such as the vibrator of a bone conduction hearing aid or the electrode array of a cochlear implant hearing aid, or the wireless transmitter of an earphone. Figure 5A , 5BIn the embodiment representing the microphone path of the headphones, the beamformer is configured to pick up the user's voice based on the microphone signal and fixed and / or adaptively updated beamformer weights. In this case, an additional speaker path may be provided for representing the sound from a distant communication partner. This is in... Figure 6 As shown in the image.
[0285] Figure 6 An embodiment of a hearing aid or headphone (HD) according to the present invention is shown. Figure 6 An embodiment of a headset or hearing aid is shown, which includes self-voice estimation and the option to transmit the self-voice estimation to another device, as well as receiving sound from the other device for presentation to the user via a speaker, for example, mixed with sound from the user's environment. Figure 6 An embodiment of a hearing device (HD), such as a hearing aid or headphones, is shown, comprising two microphones (M1, M2) configured to provide an electrical input signal (IN1, IN2) representing sound in the environment of a user wearing the hearing device. The hearing device also includes a spatial filter DIR and a self-voice DIR, each spatial filter providing spatially filtered signals (ENV and OV, respectively) based on the electrical input signal. The spatial filter DIR may, for example, implement a target-preserving, noise-cancelling beamformer for a target signal in the acoustic far field relative to the user. The spatial filter self-voice DIR implements a self-voice beamformer directed towards the user's mouth (its activation is controlled, for example, by a self-voice presence control signal and / or a telephone mode control signal). In the telephone operation mode of the hearing aid (or the normal operation mode of the headset), the user's self-voice is picked up by microphones M1 and M2 and spatially filtered by the self-voice beamformer of the spatial filter "Self-Voice DIR," thus providing the signal OV. Optionally, this signal is fed to the transmitter Tx via the self-voice processor OVP for transmission (through a cable or wireless link to another device or system, such as a telephone; see the dashed arrow labeled "To Telephone" and the telephone symbol). In the telephone operation mode of the hearing aid (or the normal operation mode of the headset), the signal PHIN can be received from another device or system (such as a telephone, as shown by the telephone symbol and the dashed arrow labeled "From Telephone") via the (wired or wireless) receiver Rx. When the distant speaker is active, the signal PHIN contains the speech from the distant speaker, for example, transmitted over a telephone line (e.g., entirely or partially wireless, but typically at least partially cable-based). The "remote" telephone signal PHIN can be selected in the combination unit (here, the selector / mixer SEL-MIX) or mixed with the ambient signal ENV from the spatial filter DIR. The selected or mixed signal PHENV is fed to the output converter SPK (such as a speaker or the vibrator of a bone conduction hearing device) to be presented to the user as sound. (Optionally, such as...) Figure 6As shown, the selected or mixed signal PHENV can be fed to the processor PRO, thereby applying one or more processing algorithms to the selected or mixed signal PHENV to provide the processed signal OUT, which is fed to the output converter SPK. Figure 6 An example of this could be an earphone in which the received signal PHIN can be selected to be presented to the user without being mixed with ambient signals. Figure 6 An example of this can represent a hearing aid in which the received signal PHIN can be mixed with an ambient signal before being presented to the user (so that the user can retain a sense of their surroundings; of course, this is also suitable for headphone applications, depending on the usage). Furthermore, in the hearing aid, the processor PRO can be configured to compensate for hearing loss in the user of the hearing device (hearing aid).
[0286] The noise reduction system according to the present invention (e.g.) Figure 3A , 3B The noise reduction system (NRS) can be included in the "self-voice path" and / or the "speaker path". In the self-voice path, the noise reduction system can be implemented in the "self-voice DIR" module or the self-voice processor OVP. In the speaker path, the noise reduction system can be implemented in the DIR module.
[0287] Figure 7 A spectrum is shown, illustrating how the signal varies with time at different frequencies and how the signal has been attenuated by the neural network according to the invention. More specifically, Figure 7 The spectrum of the speech signal is shown, along with the post-filter gain applied to it (brighter light, less attenuation (speech); darker light, more attenuation (noise)). The gray scale corresponds to attenuation values in decibels from 0 dB (bright) to 12 dB (dark). The spectrum shows the relationship between the magnitudes of 512 channels representing the frequency range of the speech signal from 0 to 10 kHz and the time of 1200 time frames of the signal. (The last sentence appears to be incomplete and possibly refers to a different topic.) Figure 7 It is evident that the speech is "sparse" in the spectrogram representation, for example, <20% of the TF window contains the target speech. Therefore, this is beneficial to the current scheme, where unchanging TF windows can be ignored when updating the neural network. The same applies to video images.
[0288] train
[0289] Figure 8 The training setup of the neural network for the SNR-gain estimator according to the present invention is illustrated schematically. The figure and part of the following description are taken from EP3694229A1.
[0290] Generally, a neural network including the peak GRU RNN algorithm according to the present invention can be trained with a baseline GRU RNN to provide optimal weights (e.g., weight matrix W) for the trained network. xr W hr W xc W hc W xu W hu See Figure 1A , 1C Peak GRU constraints can be applied to trained networks.
[0291] Neural networks can be trained based on the estimated signal-to-noise ratio as an example of an input obtained from a mixture of noisy inputs and its corresponding output as an example of a vector across the frequencies of a noise-reduced input signal that primarily contains the desired signal. Neural networks can also be trained based on an example of a (digitized) electrical input signal IN(k,t), for example directly from an analytical filter bank (see, for example). Figure 3A , 3B In the FB-A), the corresponding SNR and / or appropriate gain are known (or, if the SNR estimator is part of a neural network, its appropriate gain is known).
[0292] A neural network (RNN) can be trained based on an estimated signal-to-noise ratio as an example of the input obtained from mixing noisy inputs and its corresponding output as a cross-frequency vector of noise reduction gain (attenuation). When applied to an electrical input signal (a digital representation), it provides a signal that primarily contains the desired (target) signal. A peak GRU RNN can be trained using a regular GRU, i.e., the trained weights and biases are passed from the GRU to the peak GRU.
[0293] As described in EP3694229A1, an SNR-to-gain converter (e.g., see SNR2G in this figure) may include a neural network, wherein the weights of the neural network have been trained using multiple training signals (e.g., see...). Figure 8 The SNR estimator SNR-EST, which provides input to the SNR-to-gain converter SNR2G, can be implemented using conventional methods, such as without artificial neural networks or other algorithms based on supervised or unsupervised learning. It can also be implemented using neural networks.
[0294] Figure 8 The training setup of a neural network for the SNR-gain estimator (SNR2G) according to the present invention is illustrated schematically. Figure 8The database DB-SN is shown, comprising appropriate examples (exponents q, q = 1, ..., Q) of time intervals for clean speech S, each interval being, for example, longer than 1 second, e.g., in the range of 1 second to 20 seconds. The database may include each time interval of S(k,t) represented by time-frequency, where k is the frequency exponent and t is the time exponent. The database may include examples of corresponding noise N (e.g., different types of noise and / or different noise amounts (levels) for the p-th speech segment), e.g., represented by time-frequency, N(k,t). A given mixture of clean speech S and noise N is associated with a known SNR and a correspondingly known optimal gain G-OPT. This data can be taken as “ground truth” data for training algorithms and providing optimized weights for neural networks. Clean speech S q (k,t) and noise N q Different corresponding time intervals of (k,t) can be presented separately (in parallel) to a given combination S for speech and noise. q (k,t),N q (k,t) provides the optimal gain G-OPT q The (k,t) module OPTG. Similarly, the clean speech S q (k,t) and noise N q Different time intervals corresponding to (k,t) can be mixed, and the mixed signal IN q (k,t) can be presented to a given combination S for speech and noise. q (k,t),N q (k,t) provides a noisy (mixed) input signal IN. q The estimated SNR (SNR-EST) for (k,t) q The SNR estimator SNR-EST for (k,t) is given. The estimated SNR (SNR-EST) is... q (k,t) is fed to an SNR-gain estimator SNR2G implemented as a neural network, such as a recurrent neural network according to the present invention, which provides a corresponding estimated gain G-EST. p (k,t). The corresponding optimal and estimated gains (G-OPT) q (k,t),G-EST q (k,t) is fed into the value function module LOSS, which provides a measure of the current "value" ("error estimate"). This "value" or "error estimate" is iteratively fed back to the neural network module SNR2G to modify the neural network parameters until an acceptable error estimate is achieved. This replaces reliance on separate SNR estimators (e.g., ...). Figure 3A As shown in Figures 4A and 5A, the neural network can be configured to provide a noise reduction gain G(k,t) directly from the noisy input signal IN(k,t), for example, see [reference 4A, 5A]. Figure 3B 4B, 5B. In this case, Figure 8 The adaptive training procedure should be included as a neural network.
[0295] Training data can be fed to the neural network frame by frame (SNR(k,t) followed by SNR(k,t+1), or IN(k,t) followed by IN(k,t+1)), where a time step represents the frame length (e.g., N). s (e.g. N) s =64) samples / frame divided by the sampling frequency f s (e.g. f) s =20kHz), thus providing an exemplary frame length of 3.2ms) or a portion thereof (in the case of overlapping time frames). Similarly, the output G(k,t) of the neural network can be transmitted frame by frame (G(k,t) is followed by G(k,t+1)).
[0296] The neural network can be initialized randomly and then iteratively updated. The optimized network parameters (e.g., weights and biases) for each node can be determined based on the neural network output G-EST. p (k,t) and the optimal gain G-OPT q (k,t) is found using standard, iterative stochastic gradients, such as steep descent and steep ascent methods, or by using a backpropagation implementation that minimizes the value function, such as mean squared error (see signal ΔG). q (k,t)). The value function (such as mean square error) is calculated across many training pairs of the input signal (q = 1, ..., Q, where Q can be ≥10, such as ≥50, such as ≥100 or greater).
[0297] Optimized neural network parameters can be stored in the SNR-gain estimator SNR2G implemented in the hearing device and used to obtain frequency-varying input SNR values, such as from the "posterior SNR" (simple SNR, e.g., (S+N) / <n>) or from "prior SNR" (improved SNR, e.g. <s> / <n>Alternatively, the gain that varies with frequency can be determined from either of these (where <●> refers to the estimate).
[0298] It may be advantageous to use different thresholds across neurons (within a given layer) for selecting which neurons will be processed. The threshold for a given neuron may be adjusted during training.
[0299] If there is more than one RNN layer, the threshold can be different between the layers.
[0300] Other training methods can be used, for example, see [Sun et al.; 2017]. See also the "Statistical RNN (StatsRNN)" section below.
[0301] The neural network preferably has a K-dimensional input vector and output vector (the same as the time frame of the input signal IN(k,t) (or SNR(k,t)) and the corresponding (estimated) noise reduction gain G(k,t)), where K is the number of frequency bands in the analysis path of the hearing device in which the neural network is implemented (see, for example, [link to relevant documentation]). Figure 4A (4B, 5A, 5B). K may be equal to or different from the number of frequency bands in the forward (audio) path of the hearing device.
[0302] Figure 9 The diagram schematically illustrates a hardware module configured to act on a set of values (vectors) at a time, exemplified here as multiple sets of four elements. In hardware, computation is typically performed group by group (vector), and in the case of peak GRU RNNs and delta GRU RNNs, either no processing is performed or the entire vector is processed. Therefore, a hardware module adapted to the peak GRU RNN algorithm could, for example, be adapted to compute the average of each group and select the largest vector group based on such average.
[0303] Furthermore, elements can be reorganized so that adjacent frequency bands are expanded across different groups (to prevent loss of basic information from lower frequency bands when the entire vector group is discarded). Memory ( Figure 9 The four vertical rectangles in the diagram are used to illustrate the entire vector / rectangle (e.g., columns of the weight matrix) corresponding to the Δx / Δh vector. Figure 9 The input is denoted as Δx / Δh. Therefore, if four elements in either the Δx or Δh vector are skipped (they are 0), then the four weight columns will also be skipped. That is, the entire weight column is skipped, not just the weights themselves. Figure 9 A single element within the horizontal rectangle of the memory.
[0304] When processing elements in hardware, several methods are possible. We can process the elements one by one sequentially, or we can, for example, perform vectorized operations in groups of four (operating on a group of values (vectors) at a time, providing speedup). In this case, we will need to decide whether we want to process or discard all four elements, as they will be received and output as vectors. This can be based on determining whether most / part of these values are above a threshold / in the top N of the vector. p Of the values, if yes, the entire vector will be processed. If no, all values will be discarded. It is important to consider how the frequency bands are grouped into vectors. Grouping the frequency bands in their original order (band 1, 2, 3, 4, etc.) may result in the loss of information, such as lower frequencies, if most values within the vector are too small. Therefore, regrouping may be considered (e.g., based on experimentation).
[0305] Combination of Peak GRU and Delta GRU
[0306] Furthermore, the peak GRU RNN and delta GRU RNN methods can be combined; that is, we can first determine whether the corresponding value is higher than a certain delta threshold, and then apply another filter based on peak GRU RNN (and vice versa). This combines the advantages of both methods, namely, updating only when the threshold is exceeded, but not updating more than N. p One neuron.
[0307] The combination of peak GRU RNN and delta GRU RNN can include the following selection: if (based on the delta GRU threshold) N is not reached in a layer p The "saved" computational resources can be moved to it, which has exceeded N. p Another layer (or time step) (assuming the neural network has more than one peak GRU layer).
[0308] Statistical RNN (StatsRNN) (peak and / or threshold N) p (Training and setup of Θ)
[0309] Another approach to obtaining / approximating the first few elements of a peak-weighted GRU RNN is to look at the data statistically to investigate whether an element should be set to 0. and The calculation determines whether there are usable statistical properties. We can separately calculate x and h (the absolute values of the difference) based on the training dataset (across all acoustic environments or individually for each environment) to create histograms and explore whether the data exhibits random behavior (this is used...). Figure 10 The example illustration shows data from a training dataset used for a neural network that includes StatsRNN (StatsGRU layers). (A magnified view of a portion of the logarithmic histogram of the data). In this example, the very left thin black vertical line corresponds to the first window containing ~31% zeros. The X (horizontal) and y (vertical) axes correspond to the threshold (histogram bar boundaries) and percentage, respectively, i.e., how many elements in the entire training dataset are within each bar. If the data is not random, such as... Figure 10 As in the previous example, the threshold can be statistically determined separately for x and h (using a histogram window), corresponding to the percentage of the first n elements that should be processed. This approach retains the idea of peak GRU RNNs and provides valuable insights into a given dataset. StatsRNN defines the initial thresholds analytically for x and h individually. This is a crucial consideration because, as shown in [Neil et al.; 2018], vectors have different sparsities, and their individual processing can contribute to further improvements. Therefore, leveraging prior knowledge about the data will lead to better and more deterministic algorithms with consistent performance, as well as a better determination of the data word length necessary for hardware execution.
[0310] Furthermore, we can apply adaptive threshold settings, meaning the thresholds for x and h are not static values, but can be adjusted, for example, from one time step to another (or across all environments or individually for each environment). If too many elements need to be processed at the current time step t (too many elements in the vector are above the threshold, e.g., 80 out of 100 instead of 50) using the current threshold, the threshold can be increased at the next time step. The new threshold can again be determined based on the histogram, i.e., we can take the boundary of the adjacent histogram bar to the right of the current boundary (a larger value) or any other boundary of the histogram greater than the current boundary. Regarding the elements themselves, we can continue further calculations with all elements or perform additional filtering to obtain (or approximate) the required number of elements (50 out of 100, decreasing by 30 elements from the initial 80). This filtering can be based on classification and selection of the maximum value, selection of the first / last n elements above the threshold, selection of random elements above the threshold, etc. As an alternative, elements above the threshold but not selected can be prioritized in the next time step.
[0311] Similarly, if too few elements need to be processed at the current time step t (e.g., 20 out of 100 instead of 50), the threshold can be lowered in the next time step and can be determined based on the histogram. We can take the boundary of the adjacent histogram bar to the left of the current boundary (smaller value) or any other boundary of the histogram smaller than the current boundary. Regarding the elements themselves, we can continue further calculations with several of the obtained elements or perform additional element selection to obtain (or close to) the required number of elements (in this example, an additional 30 elements, thus obtaining 50 out of 100). This selection can be based on classification and selecting the maximum value, setting a threshold that is already high at the current time step, and selecting, for example, the first / last / random n elements above the threshold from the initial set of elements discarded using the initial threshold.
[0312] The initial threshold itself can (ideally) be based on a histogram (or can be randomized), which will provide the best starting point and a more predictable way of applying the touch threshold. The same x and h thresholds can be used for all acoustic environments, or different x and h thresholds can be applied individually to each environment. This approach can be implemented in both software and hardware primarily for StatsRNN, but can also be used for DeltaRNN (and its improved versions).
[0313] Figure 11A and 11B Different exemplary embodiments of hearing devices such as hearing aids are shown, wherein the gain estimator (TE-NE2Gain) is implemented by a neural network NN, such as a recurrent or convolutional neural network, such as a deep neural network, preferably including a modified gated recurrent unit according to the invention.
[0314] Figure 11A An embodiment of the hearing device according to the invention is shown, wherein the input of the neural network NN implementing the noise reduction gain estimator (TE-NE2Gain) is not the SNR (e.g., as shown in the figure). Figure 2A (as in 3A, 4A) or electrical input signals not from the input converter (such as...) Figure 2B (as in 3B, 4B), including separate target estimators (TE) and noise estimators (NE), or corresponding magnitude responses of the target and noise estimators, or at least the noise estimator, or a mixture of the noise estimator and noisy input. Figure 11A In this embodiment, the target estimate (TE) and noise estimate (NE) are estimated based on a signal (IN(t)) from a single input microphone (M). The microphone (M) provides a time-domain electrical input signal (IN(t), where t represents time) (e.g., digitized via an analog-to-digital converter, where appropriate). The hearing device includes a filter bank (including an analysis filter bank (FB-A) and a synthesis filter bank (FB-S)) that enables processing in the hearing device to be performed in the time-frequency domain (k,m), where k and m are the frequency exponent and the time exponent, respectively. The analysis filter bank (FB-A) is connected to the microphone (M) and provides the electrical input signal IN(t) in time-frequency representation IN(k,m). The hearing device includes an output converter, in this case a loudspeaker (SPK), for converting the time-domain output signal (OUT(t)) into a stimulus that can be perceived as sound by the user. The synthesis filter bank (FB-S) converts the processed time-frequency signal (OUT(k,m)) into a time-domain output signal (OUT(t)). The hearing device includes a target and noise estimator (TE-NE) for providing estimates of the target component (TE(k,m)) and the noise component (NE(k,m)) of the electrical input signal (IN(k,m)). The target estimate (TE(k,m)) and the noise estimate (NE(k,m)) are input to a gain estimator (TE-NE2Gain), which is implemented by a neural network (NN), such as a recurrent or convolutional neural network like a deep neural network, preferably including a modified gated recurrent unit according to the invention. The output of the neural network (NN) is a gain value (G(k,m)) representing the attenuation used for noise reduction when the signal applied to the positive path is, in this case, the electrical input signal (IN(k,m)). The gain (G(k,m)) is applied to the electrical input signal (IN(k,m)) in a combining unit (in this case, a multiplication unit (X)). The output of the combining unit (X) is the processed output signal (OUT(k,m)), which is fed to the synthesis filter bank (FB-S) to be converted to the time domain and presented to the user (and / or transmitted to another device or system, for example, for further processing).
[0315] The estimated gain can be applied to the target estimate (TE(k,m)) instead of the output of the analysis filter bank (the electrical input signal IN(k,m)). This is particularly suitable when the target estimate is a beamforming signal. Instead of applying the noise reduction gain (G(k,m)) from the neural network (NN) to the electrical input signal (IN(k,m)), it can be applied to a further processed version of the electrical input signal, such as as described below (see...). Figure 11B ).
[0316] In the case of multiple microphones (see, for example) Figure 11B The target estimate (TE) and noise estimate (NE) can be obtained from beamformer filters (BFa) that provide a linear combination of microphone signals (IN1, IN2), such as target enhancement beamformers and target cancellation beamformers (the latter having a zero orientation approximately pointing towards the target). Figure 11B An embodiment of the hearing device according to the invention is shown, wherein the inputs of the neural network (NN) are the output of the target-preserving beamformer (representing the target estimate, TE(k,m)) and the output of the target-cancelling beamformer (representing the noise estimate, NE(k,m)).
[0317] Figure 11B The hearing device includes two microphones (M1, M2), each providing a time-domain electrical input signal (IN1(t), IN2(t)). The hearing device includes two analysis filter banks (FB-A) connected to the respective microphones for providing electrical input signals (IN1(t), IN2(t)) in time-frequency representation (IN1(k,m), IN2(k,m)). The electrical input signals (IN1(k,m), IN2(k,m)) are fed to a beamformer filter (BFa) comprising a target-preserving beamformer and a target-cancelling beamformer. The target-preserving beamformer and the target-cancelling beamformer provide the target estimate (TE) and the noise estimate (NE), respectively. Figure 11A In this embodiment, s, the target estimate (TE), and the noise estimate (NE) are used as inputs to a neural network (NN) that implements the noise reduction gain estimator (TE-NE2Gain). The noise reduction gain estimator (TE-NE2Gain) provides a representation of the signal applied to the forward path, which is here the beamforming signal (Y). BF The gain value (G(k,m)) attenuation for noise reduction when (k,m) is applied. Beamforming signal (Y) BF (k,m) is provided by the beamforming filter (BFb), for example as TE-βNE, where β is an adaptively determined parameter (determined in module BFb), see, for example, US2017347206A1. The beamforming signal is determined based on the target-preserving beamformer and the target-cancelling beamformer, as well as the current electrical input signals (IN1(k,m), IN2(k,m)) (possibly and for use with a voice activity detector to distinguish between speech and noise).
[0318] Beamformers can be fixed or adaptive. Multiple target cancellation beamformers can be used simultaneously as input features for a neural network (NN). For example, two target cancellation beamformers, each with a single zero direction but multiple zeros pointing to different possible targets.
[0319] The magnitudes, squares, or logarithms of the target and noise estimates can be used as inputs to a neural network (NN). The output of a neural network (NN) can include real-valued or complex-valued gain, or separate real-valued gain and real-valued phase.
[0320] The maximum noise reduction provided by a neural network can be controlled by the level of the neural network's input, or modulation (such as SNR), or sparsity. Sparsity can be represented, for example, by the degree of temporal and / or frequency overlap between the background noise and the (target) speech.
[0321] exist Figure 11B In the embodiments described, a noise estimate based on a single target-cancelled beamformer is shown. However, several noise estimates can be provided as input features to a neural network. Different noise estimates can be composed of or provided by different target-cancelled beamformers, each with zeros pointing in a specific direction. But the noise estimates (used as input to the neural network) can also be based on other characteristics besides spatial characteristics, such as a noise floor estimator, for example, based on the modulation of the input signal.
[0322] When appropriately replaced by a corresponding process, the structural features of the apparatus described above, in detail in the "Detailed Description" section, and as defined in the claims can be combined with the steps of the method of the present invention.
[0323] Unless explicitly stated otherwise, the singular forms "a" and "the" as used herein include the plural forms (i.e., meaning "at least one"). It should be further understood that the terms "having," "comprising," and / or "including" as used in the specification indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that, unless explicitly stated otherwise, when an element is referred to as "connected" or "coupled" to another element, it can be a direct connection or coupling to the other element, or there may be intermediate inserting elements. The term "and / or" as used herein includes any and all combinations of one or more of the listed related items. Unless explicitly stated otherwise, the steps of any method disclosed herein do not necessarily need to be performed in the exact order disclosed.
[0324] It should be understood that references to "an embodiment," "an embodiment," "an aspect," or "may" in this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Furthermore, particular features, structures, or characteristics may be suitably combined in one or more embodiments of the invention. The foregoing description is provided to enable those skilled in the art to implement the various aspects described herein. Various modifications will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects.
[0325] The claims are not limited to the aspects shown herein, but encompass the full scope consistent with the language of the claims, wherein, unless expressly stated, an element referred to in the singular does not mean "one and only one," but rather "one or more." Unless expressly stated, the term "some" means one or more.
[0326] The peak GRU RNN algorithm is exemplified in the field of audio processing, particularly within the framework of noise reduction systems for hearing aids or headphones. However, this algorithm can be applied to other tasks in hearing aids or headphones or other audio processing devices, such as direction of arrival estimation, feedback path estimation, (self)voice activity detection, keyword detection, or other (e.g., acoustic) scene classification. Furthermore, this algorithm can also be applied to other domains besides audio processing, such as images or data containing a certain amount of redundancy (e.g., changes relatively slowly over time, such as relative to the sampling time (t)). s =1 / f s , where f s This refers to the processing of other data, such as financial data and climate data, for the sampling frequency of this data.
[0327] References
[0328] ·[Neil et al.; 2018] Daniel Neil, Jun Haeng Lee, Tobi Delbruck, Shih-ChiiLiu, DeltaRNN: A Power-efficient Recurrent Neural Network Accelerator, published in FPGA'18: Proceedings of the 2018ACM / SIGDA International Symposium on Field-Programmable Gate Arrays, February 2018, Pages 21–30, https: / / doi.org / 10.1145 / 3174243.3174261;
[0329] ·EP3694229A1(Oticon)12.08.2020;
[0330] ·[Kostadinov; 2018];
[0331] ·https: / / towardsdatascience.com / understanding-gru-networks-2ef37df6c9be.< / n> < / s> < / n>
Claims
1. A hearing device configured for wear by a user in or in the ear, or wholly or partially implanted in the head at the user's ear, said hearing device comprising: - An input unit for providing at least one electrical input signal in time-frequency representation of k, t, where k and t are the sub-band channel index and time index, respectively, k=1, …, K, where K is the number of sub-band channels, and the at least one electrical input signal represents sound and includes the target signal component and noise component; and - Signal processor, including -- SNR estimator for providing a target signal-to-noise ratio (SNR) estimate SNR(k,t) of the at least one electrical input signal or a signal derived therefrom in the time-frequency representation; -- SNR-gain converter for converting the target signal-to-noise ratio estimate SNR(k,t) into the corresponding gain value G(k,t) in the time-frequency representation; The signal processor includes a neural network comprising at least one layer defined as a gated recurrent unit, the gated recurrent unit comprising a memory in the form of a hidden state vector h, wherein the output vector o(t) is provided by the gated recurrent unit based on the input vector x(t) and the hidden state vector h(t-1), wherein the output o(t) at a given time step t is stored as the hidden state h(t) and used to compute the output vector o(t+1) at the next time step t+1; At least one of the SNR-gain converters is implemented via the neural network, and at least one layer defined as a gated recurrent unit is implemented as a modified gated recurrent unit, wherein the signal processor is configured to calculate, at a given time t, the changes Δx(i,t) in the input vector x(t) and the hidden state vector h(t-1) from one time t-1 to the next time t. and Δh(j,t-1) = ,in 𝑡−1 and 𝑗, x(i,t-1) and The estimated value of (j,t-2), where i and j refer to the i-th input neuron and the j-th hidden neuron, respectively, and x(i,t-1) is the input vector of the i-th input neuron at time t-1. (j,t-2) is the hidden state vector of the j-th neuron at time t-2, where 1 ≤ i ≤ N. ch,x and 1 ≤ j ≤ N ch,oh , where N ch,x and N ch,oh The number of processing channels for the input vector x and the hidden state vector h are respectively, and the signal processor is further configured such that for the input vector x(t) and the hidden state vector h(t-1) at a given time t, the N of the improved gated loop unit is... ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p,x Less than N ch,x and N p,oh Less than N ch,oh .
2. The hearing device according to claim 1, wherein, The signal processor is configured to determine the estimates of the input vector and the hidden state vector as follows: in To improve the peak count of the gated recurrent unit, i and j refer to the i-th input neuron and the j-th hidden neuron, respectively.
3. The hearing device according to claim 2, wherein, The signal processor is configured to determine the changes in the values of the input vector and the hidden state vector as: 。 4. The hearing device according to claim 1, wherein, The input unit includes multiple input converters and a beamformer filter, wherein the beamformer filter is configured to provide at least one electrical input signal based on signals from the multiple input converters.
5. The hearing device of claim 1, comprising a voice activity detector configured to estimate whether or with what probability an input signal at a given time point includes a voice signal and to provide a voice activity control signal indicating a specified result.
6. The hearing device of claim 1, further comprising an output unit configured to provide output stimulation to a user based on at least one electrical input signal.
7. The hearing device according to claim 1, wherein, The signal processor is configured to apply the gain value G(k,t) provided by the SNR-gain converter to at least one electrical input signal or a signal derived therefrom.
8. The hearing device according to claim 1, wherein, The signal processor is configured to discard N at a given time t. p The absolute value of each channel or Less than threshold Θ p The processing of the channels, among which To improve the peak count of the gated loop unit.
9. The hearing device according to claim 1, wherein, Peak quantity N p,x and N p,oh It is adaptively determined based on at least one electrical input signal.
10. The hearing device according to claim 1, wherein, The parameters of the neural network have been trained using multiple training signals.
11. The hearing device according to claim 1, comprising an air conduction hearing aid, a bone conduction hearing aid, a cochlear implant hearing aid, or a combination thereof.
12. The hearing device according to claim 1, comprising or including headphones.
13. The hearing device according to claim 1, comprising a hardware module adapted to process the elements of the gated loop unit in a vectorized manner.
14. The hearing device according to claim 1, comprising an air conduction hearing aid, a bone conduction hearing aid, a cochlear implant hearing aid, or a combination thereof.
15. A method for operating a hearing device, the method comprising: - Provide at least one electrical input signal in terms of time and frequency k, t, where k and t are the sub-band channel index and time index, respectively, k=1, …, K, where K is the number of sub-band channels, and the at least one electrical input signal represents sound and includes the target signal component and noise component; and - Provide a target signal-to-noise ratio (SNR) estimate SNR(k,t) for the at least one electrical input signal or a signal derived therefrom in the time-frequency representation; - Convert the target signal-to-noise ratio estimate SNR(k,t) into the corresponding gain value G(k,t) in the time-frequency representation; - Provides a neural network comprising a layer containing at least one gated recurrent unit, the gated recurrent unit including a memory in the form of a hidden state vector h, wherein the output vector o(t) is provided by the gated recurrent unit based on the input vector x(t) and the hidden state vector h(t-1), wherein the output o(t) at a given time step t is stored as the hidden state h(t) and used to compute the output vector o(t+1) at the next time step t+1; The conversion of the target signal-to-noise ratio estimate SNR(k,t) into the corresponding gain value G(k,t) in the time-frequency representation is implemented through the neural network, wherein at least one layer defined as a gated recurrent unit is implemented as a modified gated recurrent unit, and wherein the method further includes: - At a given time t, determine the changes Δx(i,t) in the input vector x(t) and the hidden state vector h(t-1) from time t-1 to the next time t. and Δh(j,t-1) = ,in 𝑡−1 and 𝑗, x(i,t-1) and The estimated value of (j,t-2), where i and j refer to the i-th input neuron and the j-th hidden neuron, respectively, and x(i,t-1) is the input vector of the i-th input neuron at time t-1. (j,t-2) is the hidden state vector of the j-th neuron at time t-2, where 1 ≤ i ≤ N ch,x and 1 ≤ j ≤ N ch,oh , where N ch,x and N ch,oh The number of processing channels for the input vector x and the hidden state vector h are respectively, and the signal processor is further configured such that for the input vector x(t) and the hidden state vector h(t-1) at a given time t, the N of the improved gated loop unit is... ch,x and N ch,oh The number of update channels in each processing channel is limited to the peak number N. p,x and N p,oh , where N p,x Less than N ch,x and N p,oh Less than N ch,oh .
16. The method of claim 15, further comprising training the parameters of a neural network with a plurality of training signals.
Citation Information
Patent Citations
A hearing device comprising a noise reduction system
EP3694229A1
Hearing aid comprising a beam former filtering unit comprising a smoothing unit
US20170347206A1
Audio processing with neural networks
CN109074820A
Hearing device comprising a noise reduction system
US20200260198A1