Stable hybrid auditory filterbanks
The hybrid auditory filterbank addresses the interpretability and adaptability challenges by combining fixed and trainable filters, enhancing stability and adaptability, resulting in improved speech enhancement performance.
Patent Information
- Application Number
- EP2024196325
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-24
- Publication Date
- 2026-02-25
AI Technical Summary
Existing audio processing techniques face challenges in achieving a balance between interpretability and adaptability, particularly in encoder-mask-decoder settings, with fixed transforms being inflexible and data-driven methods being unstable and hard to interpret.
A hybrid auditory filterbank is introduced, combining fixed and trainable filters to enhance stability and adaptability, using an auditory filterbank structure with trainable convolutional layers and a penalty term to maintain tightness, ensuring numerical stability and efficient reconstruction.
The hybrid auditory filterbank achieves improved performance in speech enhancement tasks by maintaining stability and adaptability, outperforming traditional methods in perceptual quality and computational efficiency.
Smart Images

Figure IMGAF001_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention generally relates to the field of audio processing, and more specifically to audio-related deep learning tasks such as speech processing and speech enhancement.BACKGROUND
[0002] Time-frequency transforms, such as the short-time Fourier transform (STFT) and the constant-Q transform (CQT) have been the standard choice for extracting features from audio signals in machine-learning applications for decades (see Reference [1]; references are listed at the end of the description). Meanwhile, with the availability of large data sets and increasing computing power, data-driven feature extraction has shown to be capable of keeping up with classical methods, even outperforming them in many tasks [2, 3]. While both approaches are being used in applications, they come with different pros and cons. On the one hand, fixed transforms are typically under control and interpretable but may not uniformly suit any task at hand. On the other hand, data-driven features are typically flexible and adaptive but hard to train, barely interpretable, and potentially unstable.
[0003] The disagreement on the superiority of either of the two approaches is particularly noticeable in an encoder-mask-decoder setting. Using neural networks, this classical signal processing technique has been modernized in various forms in the last few years [4, 5, 6]. In time-frequency domain approaches, the encoder is typically a fixed time-frequency transform (e.g., STFT [5], mel-frequency spectrogram, Gammatone filters [7]), and the mask is estimated by a neural network that is trained on the coefficients of the transformed input signals [5, 8, 7]. In time domain approaches, the encoder is typically a convolutional layer with 1-D filters that is optimized together with the other parameters of the model and the waveform of the signals is used directly [3, 4]. These two paradigms have often been in competition with each other [6] and the dichotomy in this discussion indicates that the optimal design of the encoder remains an unsolved problem [9].
[0004] While many works refer to a distinction between end-to-end approaches - learning from time domain data directly - and predetermined coefficient domain - e.g. time-frequency - methods, some approaches aim to bridge this gap, and fuse feature engineering with feature learning. In these hybrid constructions, encoders are optimized in a more structured manner using domain knowledge, e.g., optimizing only certain parameters of prototype filter functions, such as center frequencies, bandwidth, and amplitude [10, 11, 12].
[0005] One approach to this was recently introduced in the context of fitting given, fixed filterbanks with multiresolution neural networks
[13] . There, learning filter coefficients is done on different levels of resolution of the input signal using a wavelet decomposition. This writes as a structured pair-wise composition of wavelet filters with trainable filters by convolution.
[0006] It is an objective of the present invention to provide techniques for improved trainable filterbanks, in particular in an encoder-decoder setting for speech-related and / or other audio-related deep learning tasks, thereby overcoming the above-mentioned disadvantages of the prior art at least in part.SUMMARY OF THE INVENTION
[0007] The above-mentioned objective is solved by the subject-matter defined in the independent claims. Advantageous modifications of embodiments of the present invention are defined in the dependent claims as well as in the description and the drawings.
[0008] One aspect of the present invention relates to a filterbank, particularly an auditory filterbank such as a learned auditory filterbank, even more particularly a hybrid auditory filterbank. It may be implemented on a data processing apparatus. A filterbank may generally be understood as a collection of bandpass filters that decompose an input signal into multiple frequency bands. Accordingly, the filterbank corresponding to certain aspects of the invention may comprise a plurality of filters configured to decompose an input audio signal into a plurality of sub-bands.
[0009] The hybrid auditory filterbank may be configured for audio processing. In other words, the hybrid auditory filterbank may be configured to perform one or more audio processing tasks, including by way of example but without limitation: Speech enhancement and / or noise suppression: Speech enhancement serves to improve the intelligibility and quality of speech signals in noisy environments. This may involve isolating and / or enhancing the desired speech signal, i.e., speech data present in the input audio data, while suppressing noise data or otherwise unwanted data present in the input audio data. In terms of noise suppression, the filterbank may analyze the sub-bands to identify and reduce noise components that are not part of the speech signal. This may be achieved through techniques like spectral subtraction or masking, where the noise profile is estimated and subtracted from the speech signal in the frequency domain. In general, in any type of speech processing task according to aspects of the invention disclosed herein, the input audio signal may comprise speech data and noise data, and the noise data may comprise, e.g., everyday noises. Acoustic echo cancellation: This serves to eliminate echo and / or reverberations from the input audio signal, particularly in telecommunication systems like speakerphones and video conferencing. The primary goal of acoustic echo cancellation is to improve the clarity of the speech signal by removing the delayed version of the speaker's voice that is picked up by the microphone after reflecting off surfaces in the environment. Speech recognition: This may involve converting raw audio into a format that can be processed effectively in an automated way by recognition systems, such as to be analyzed by machine-learning algorithms or classifiers. To this end, the filterbank may extract frequency-domain features like Mel-frequency cepstral coefficients (MFCCs) from speech signals. Audio classification: The filterbank may be used as front-end feature extractor for tasks like audio scene classification. Audio source separation: The filterbank may enable separating a mixed audio signal into its individual sources. Audio coding: The filterbank may be used in sub-band coding to efficiently compress audio signals.
[0010] It may be provided that the filters, or at least selected ones of the filters, of the hybrid auditory filterbank are each based on a fixed filter and a trainable filter (hence the hybrid characteristic of the filterbank). This way, the filters of the hybrid auditory filterbank can advantageously combine certain characteristics of fixed filters and trainable filters, thereby conceptually fusing fixed filterbank design with filterbank learning.
[0011] It may be provided that the fixed filters each comprise a filter of an auditory filterbank. An auditory filterbank may be implemented as an audlet. An audlet may be understood as an analysis-synthesis system for audio applications, which is typically based on an oversampled filterbank with filters distributed on auditory frequency scales, following auditory filter bandwidths. Technical background information about audlets can be found in
[20] . Accordingly, the hybrid auditory filterbank according to this aspect differs from the construction in
[13] in that it uses an auditory filterbank or audlet instead of a wavelet. A wavelet is not an auditory frequency scale. More precisely, it is known that the auditory frequency perception only follows a logarithmic progression (like a wavelet) in the high frequency domain. Using an auditory filterbank or audlet instead of a wavelet is particularly beneficial in speech processing tasks because it helps to focus on the perceptually important parts of speech, e.g., higher resolution in lower frequencies, which can be fine-tuned via the hybrid construction.
[0012] It may be provided that the trainable filters have been trained to perform an audio processing task and / or to improve the stability, preferably the numerical stability, even more preferably the tightness, of the hybrid auditory filterbank. Tightness may be understood as the numerical stability of the transformation in the sense of energy norm preservation. Accordingly, this aspect helps to stabilize the hybrid auditory filterbank throughout the training, in particular by enforcing or at least improving tightness. A crucial aspect of an encoder-mask-decoder architecture is that encoder and decoder have to be coordinated well with one another. In the absence of a mask, the encoder-decoder pair should ideally yield the identity. The decoder is called a dual for the encoder. If the encoder is its own dual, it is said to be tight. Tightness may be understood equivalent to norm preservation, which provides a strong notion of stability for the encoder since small changes in the input are always under control. Moreover, tightness can be understood as a property in which the encoder is its own inverse, i.e., applying the transposed filterbank as a decoder yields the identity. This is a particularly useful property for encoder-mask-decoder architectures since no inverse has to be computed or learned. In audio signal processing, the assumption of tightness has a long history [14, 15, 16] and is slowly making its way into the deep learning regime
[17] , providing stable signal representations that are robust against noise and / or adversarial examples.
[0013] It may be provided that the fixed filters impose an inductive bias on the hybrid auditory filterbank in terms of at least one audio processing characteristic. The at least one audio processing characteristic may comprise: center frequencies, bandwidths, filter shapes, impulse response supports, or any combination thereof. It may be provided that the at least one audio processing characteristic is preserved during a training of the trainable filters.
[0014] It may be provided that the hybrid auditory filterbank is configured for speech processing. The audio processing task may be a speech processing task, in particular a speech enhancement task. The input audio signal may comprise speech data and, optionally, noise data, as already described above.
[0015] It may be provided that the trainable filters are trainable in terms of convolution. It may be provided that the trainable filters each comprise a one-dimensional kernel of a convolutional layer, in particular a conv1d layer. It may be provided that the trainable filters are randomly initialized.
[0016] It may be provided that the filters are formed by a composition, preferably a convolution, of the fixed filters and the trainable filters. It may be provided that the convolution is a channel-wise or filter-wise convolution of the fixed filters and the trainable filters.
[0017] It may be provided that the trainable filters have been trained such that the audio processing task is improved and the stability of the hybrid auditory filterbank is improved by optimizing the condition number of the hybrid auditory filterbank.
[0018] Another aspect of the invention concerns an encoder for audio processing, preferably only for audio processing, comprising a hybrid auditory filterbank in accordance with any of the aspects disclosed herein.
[0019] Another aspect of the invention concerns a decoder for audio processing, preferably only for audio processing, comprising a transposed version of a hybrid auditory filterbank in accordance with any of the aspects disclosed herein. This way, the decoder can achieve (almost) perfect reconstruction in a computationally very efficient manner. In particular, no inverse decoder has to be computed, which saves processing resources and therefore decreases the computational load.
[0020] Another aspect of the invention concerns a method of audio processing. The method may comprise receiving an input audio signal, and decomposing the input audio signal into a plurality of sub-bands using a hybrid auditory filterbank in accordance with any of the aspects disclosed herein.
[0021] Another aspect of the invention concerns a method of training a hybrid auditory filterbank in accordance with any of the aspects disclosed herein. The method may comprise training the filters of the hybrid auditory filterbank. The training may be such that an audio processing task can be performed and / or the stability of the hybrid auditory filterbank is improved, preferably by optimizing the condition number of the hybrid auditory filterbank.
[0022] Another aspect of the invention concerns a data processing apparatus. The data processing apparatus may comprise means for implementing a hybrid auditory filterbank in accordance with any of the aspects disclosed herein and / or an encoder in accordance with any of the aspects disclosed herein and / or a decoder in accordance with any of the aspects disclosed herein and / or for carrying out a method in accordance with any of the aspects disclosed herein.
[0023] Another aspect of the invention concerns a computer program. Another aspect of the invention concerns a computer-readable medium having stored thereon a computer program.
[0024] The computer program may comprise instructions which, when the program is executed by a computer, cause the computer to implement a hybrid auditory filterbank in accordance with any of the aspects disclosed herein and / or an encoder in accordance with any of the aspects disclosed herein and / or a decoder in accordance with any of the aspects disclosed herein and / or for carrying out a method in accordance with any of the aspects disclosed herein.
[0025] The methods may be computer-implemented methods. The methods may be implemented on a data processing apparatus. The data processing apparatus may comprise a memory. The memory may be stored on a storage medium. The methods may be implemented by instructions stored on the storage medium of the data processing apparatus that, when executed, implement the method(s). Such a data processing apparatus may comprise one or more computers.
[0026] The terms used herein should generally be construed as understood by the average person skilled in the art, unless explicitly indicated otherwise. The following explanations may guide the understanding: The term "artificial Intelligence" (AI) should be understood as referring to a branch of computer science that aims to develop machines or software capable of intelligent behavior, typically with the goal to mirror or surpass human intelligence in specific tasks. AI systems are designed to perform complex tasks such as reasoning, learning, perception, problem-solving, and understanding natural language. These systems can typically adapt to new situations and improve their performance over time. The goal of AI is to create systems that can function autonomously and interact with their environment in a human-like manner.
[0027] The term "machine learning" (ML) should be understood as a subset of artificial intelligence that focuses on the development of algorithms and statistical models that enable computers to perform specific tasks without using explicit instructions. Instead, machine-learning systems learn and make predictions or decisions based on data. Machine-learning algorithms build a mathematical model based on sample data, known as training data, to make predictions or decisions without being explicitly programmed to perform the task. Machine learning can be employed in a variety of applications, including image and speech recognition, medical diagnosis, predictive analytics, and many more, where it enables systems to learn from and adapt to new data independently.
[0028] The term "machine-learning algorithm" should be understood as a computational procedure that is designed to analyze data, learn from it, and identify patterns or make decisions based on the input data without being explicitly programmed for the task. Machine-learning algorithms leverage statistical techniques to enable systems to improve their performance on a specific task with more data over time. Machine-learning algorithms are the foundation upon which machine-learning models are built, providing the methods or processes through which data is transformed into actionable insight. Examples of machine-learning algorithms include linear regression, decision trees, support vector machines, and neural networks, among others.
[0029] The term "machine-learning model" should be understood as referring to the output generated when a machine-learning algorithm is trained on a dataset. It represents the knowledge or understanding gained by the algorithm from the data, encapsulating the learned patterns or predictions. Essentially, a machine-learning model is what enables predictions or decisions based on new, unseen data, based on the learning it has derived from the training process. The machine-learning model is typically defined by its parameters, which may be adjusted during the training phase to minimize the difference between the predicted outcome and the actual outcome. Although, strictly speaking, "machine-learning algorithm" and "machine-learning model" have distinct definitions, it is not uncommon for these terms to be used interchangeably in casual discourse. This usage stems from the close relationship between algorithms and models in the workflow of machine-learning projects, where the algorithm is the means of creating the model. Therefore, these terms may be used synonymously herein unless the distinction is decisive.
[0030] The term "artificial neural network" (ANN), or "neural network" (NN) in short, should be understood as a machine-learning or deep-learning model or algorithm. Neural networks are generally inspired by the human brain and typically comprise interconnected nodes or neurons organized into layers. Neural networks can be used to process data and learn from examples, enabling them to perform tasks such as image recognition, natural language processing, and more. A neural network typically comprises an input layer, one or more hidden layers, and an output layer. Through a process called training, neural networks can learn to perform specific tasks by adjusting their internal parameters, or "weights", based on labeled or unlabeled data.
[0031] The term "training" should be understood as referring to the process of teaching and / or optimizing a machine-learning model to make predictions or decisions, by exposing it to data for which the outcomes are known. The training process typically involves feeding a training dataset into a machine-learning algorithm, which then uses statistical analysis to learn the patterns or relationships within the data. During training, the algorithm iteratively adjusts the parameters of the model to minimize the difference between the predicted outcomes and the actual outcomes in the training data. This adjustment process is typically guided by a loss function, which measures the accuracy of the model's predictions. The goal of training is to produce a model that accurately represents the underlying structure of the data, enabling it to make reliable predictions about new, unseen data. Supervised learning involves training a model on a labeled dataset, where each example in the training data is paired with the correct output. The model learns to predict the output from the input data. Unsupervised learning involves training a model on data without labeled responses. The model tries to find patterns and relationships in the data on its own. Semi-supervised learning combines both labeled and unlabeled data during the training process, which can be beneficial when acquiring a fully labeled dataset is costly or impractical.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The invention may be better understood by reference to the following drawings: Fig. 1:A schematic block diagram of a hybrid auditory filterbank according to embodiments of the invention. Fig. 2:An example of selections of real and imaginary parts of filters (top) and their frequency responses (bottom) from three different filterbanks. From left to right: An auditory filterbank with center frequencies on the mel scale, a random filterbank with σ 2< = (TJ) -1< , and a hybrid auditory filterbank according to embodiments of the invention as the channel-wise composition of the previous two. Fig. 3:An encoder-mask-decoder architecture according to embodiments of the invention. Fig. 4:An example of the log magnitude responses of three encoders for a speech signal. From left to right: An auditory filterbank, a randomly initialized filterbank, and a hybrid filterbank according to embodiments of the invention as channel-wise composition of the previous two. While the random responses are hard to interpret, the hybrid responses are comparable to the fixed ones with the possibility to be fine-tuned in a data-driven manner. DETAILED DESCRIPTION
[0033] In the following, representative embodiments illustrated in the accompanying drawings will be explained. It should be understood that the illustrated embodiments and the following descriptions refer to examples which are not intended to limit the embodiments to one preferred embodiment.
[0034] Fig. 1 illustrates a schematic block diagram of a hybrid auditory filterbank 100 according to an embodiment of the invention. As can be seen the hybrid auditory filterbank 100 takes an input audio signal 104 as input and decomposes it into a plurality of sub-bands 106 as output. The hybrid auditory filterbank 100 comprises a plurality of filters 102. Each filter 102 is based on a fixed filter 108 and a trainable filter 110.
[0035] In the following, various characteristics and features of the hybrid auditory filterbank 100 will be described:LEARNING TIGHT HYBRID FILTERBANKS WITH INDUCTIVE AUDITORY BIAS
[0036] Let x ∈ ℝ N be an audio signal, such as the input audio signal 104. A convolutional layer Φ with 1D kernels w j ∈ ℝ T , T ≤ N decomposes x into J > 1 sub-bands via convolution. The output is represented as the array Φ x n j = x ∗ w j n = ∑ k = 0 T − 1 w j k x n − k mod N , also referred to as responses of Φ for x. In the context of classical signal processing this corresponds to an oversampled finite impulse response (FIR) filterbank
[18] . Besides all common linear time-frequency transforms, such as the STFT and the CQT, also adaptive or adapted auditory-related time-frequency representations [19, 20] can be envisioned and implemented in this way. One obstacle to a successful and functional implementation of such customized filterbanks, however, is stability.Tight Filterbank Frames
[0037] A filterbank Φ forms a frame for ℝ N if there are positive constants A ≤ B such that A ⋅ ‖ x ‖ 2 ≤ ‖ Φ x ‖ 2 ≤ B ⋅ ‖ x ‖ 2 holds for any x ∈ ℝ N
[21] . The numbers A, B are called the frame bounds. This Lipschitz-type inequality guarantees that the filterbank decomposition is invertible and well-conditioned, i.e., stable. The optimal bounds in (1) are given by the smallest and largest eigenvalue of the associated frame operator Φ T< Φ, where Φ T< denotes the transposed filterbank of Φ. These values determine the numerical stability of Φ via the condition number κ = B / A
[22] . Hence, a filterbank with A = B has optimal stability properties and is called tight. For a tight filterbank Φ, the following are equivalent
[22] . ‖ Φ x ‖ 2 = A ⋅ ‖ x ‖ 2 for all x ∈ ℝ N Φ T Φ = A ⋅ I N κ = B / A = 1
[0038] Property (i) says that the filterbank is norm preserving. This is convenient as the energy level of the encoder responses is always under control. In particular, Φ is most robust to small perturbations, which has been shown to be a crucial property in the context of adversarial examples
[17] .
[0039] Property (ii) is especially interesting in an encoder-decoder regime: If the encoder filterbank Φ is tight, then the transposed filterbank Φ T< as a decoder yields perfect reconstruction
[23] . Hence, no inverse decoder has to be computed, which may decrease the computational complexity significantly.
[0040] Property (iii) coincides with the classical definition of optimal stability of the linear operator associated with Φ. To make encoder filterbanks with trainable weights benefit from (i) and (ii), we propose to minimize κ during training in parallel with the objective function, as will be explained in more detail further below.Encoder Design: Hybrid Auditory Filterbanks
[0041] Following the idea of multiresolution neural networks
[13] , in one embodiment of the hybrid auditory filterbank 100, the filters 102 are composed from a fixed filterbank 108 and trainable filters 110 in terms of convolution. Note that for the sake of simplicity, the reference numeral 108 may refer to an individual filter of the fixed filterbank or to the fixed filterbank as a whole, and similarly the reference numeral 110 may refer to an individual filter of the trainable filterbank or to the trainable filterbank as a whole. Letting Ψ denote the fixed filterbank 108 with filters ψ j and Φ the trainable filterbank 110 with filters w j with trainable weights, then we define the hybrid filterbank Φ Ψ as the filterbank 100 with filters (w j * ψ j ) for every j. Hence, any signal x is decomposed as Φ ψ x n j = x ∗ w j ∗ ψ j n .
[0042] When initializing the filter entries of Φ at random, e.g., w j ~ (0, σ 2< I T ), the hybrid encoder can be interpreted as a random filterbank with an inductive bias. This bias is inherited from the characteristics of Ψ, and may embrace band limitation, a structured scale of center frequencies, and the like. These characteristics are preserved during the optimization of Φ. If Ψ is an auditory filterbank, we call Φ Ψ a hybrid auditory filterbank. Fig. 4 illustrates filters and their frequency responses of a hybrid auditory filterbank 100 in accordance with an embodiment of the invention.
[0043] While using an auditory encoder filterbank alone is already beneficial for audio processing tasks such as speech enhancement, the hybrid construction further allows data-driven fine-tuning by the trainable weights, and to coordinate with the audio processing task at hand, e.g. a masking model for a source separation or denoising task.Stability and κ-penalization
[0044] A random filterbank with i.i.d. (independent and identically distributed) standard normal weights forms a so-called random tight frame
[24] , i.e., it is tight in expectation
[25] , E ‖ Φ x ‖ 2 = JTσ 2 ‖ x ‖ 2 .
[0045] For a random hybrid filterbank we have the following theorem: Let Ψ be a tight filterbank with frame bound A Ψ and Φ a random filterbank with length-T filters. The associated hybrid filterbank Ω= Φ Ψ is a random tight frame with E ‖ Ω x ‖ 2 = A ψ Tσ 2 ‖ x ‖ 2 .
[0046] Hence, Φ Ψ naturally inherits the stability properties of Ψ and Φ. However, in any setting where the encoder filterbank is trainable, it is not guaranteed that it remains stable during training. To overcome this issue, for a given loss function (x, x̃) we propose to penalize the condition number κ = B / A of Φ by minimizing L β x x ˜ = L x x ˜ + β ⋅ κ , with a scaling factor β > 0. The gradient of κ is taken with respect to the filter entries of the trainable encoder filterbank. This is similar to the approach in
[17] , but less restrictive.
[0047] The computation of κ can be done efficiently if no striding is included. Denoting by ŵ j the discrete Fourier transforms (DFT) of the filters w j , which have been zero-padded to have length N, then Φ T< Φ is diagonalized as Φ T< Φ = U * ΣU, where Σ = diag ∑ k = 1 J w ^ k 2 and U is the unitary DFT matrix. Hence, the frame bounds coincide with the smallest and largest eigenvalue of Σ, given by A = min 0 ≤ k ≤ N − 1 ∑ j = 1 J w ^ j k 2 , B = max 0 ≤ k ≤ N − 1 ∑ j = 1 J w ^ j k 2 .
[0048] From this representation and a basic calculus argument, we can deduce that the gradient of κ is well-defined as long as the filterbank forms a frame. Hence, using Fast Fourier Transform (FFT) methods makes the computation of κ and its gradient fast enough to include it in the iterative procedure of gradient descent without significant loss of speed.
[0049] In applications, convolution is usually performed with a stride to reduce redundancy, i.e., a hop-size in the sliding filter. If Φ is tight with the stride, the theory works as explained above.MODEL IMPLEMENTATION FOR SPEECH ENHANCEMENT AND TRAINING
[0050] In the following, the concepts described above, including the hybrid auditory filterbank and the κ-penalization method, are applied to a speech enhancement task, i.e., given a noisy signal x noisy = x + n we aim to suppress the noise signal n via an encoder-mask-decoder model.Encoder / Decoder Design
[0051] We compare four different encoder configurations. Each one comes with 256 channels. 1. STFT (baseline): A STFT with Hann window of length 512 and a hop-size of 256. The associated filterbank has a condition number κ = 2. 2. Audlet: A tight auditory filterbank Ψ computed with the routine audfilters
[20] from the LTFAT toolbox
[26] (see Fig. 2 on the left). The filters are smoothed and cut to a length of 512 samples and a hop-size of 128 is used. This filterbank is comparable to a CQT transform with frequency-adaptive bandwidths and the center frequencies following the mel scale. 3. Conv1d: A randomly initialized trainable filterbank Φ with filters of length 32 and a hop-size of 8. This setting is reminiscent of the encoder used in Conv-TasNet [4]. 4. Hybrid audlet: A randomly initialized hybrid auditory filterbank Φ Ψ as one exemplary embodiment of the invention, which is composed of Ψ and a trainable filterbank Φ with filters of length 11 and a hop-size of 1.
[0052] For all cases, the decoder is the transposed filterbank of the encoder and is not learned, i.e., the weights are shared between encoder and decoder. Using κ-penalization (see Equation (5) above), this will always be very close to a dual for the encoder. To benefit from fast convolution on graphics processing units (GPU), certain embodiments implement some or all encoder decompositions via Pytorch's conv1d
[27] . In certain embodiments, complex convolution is implemented separately on real and imaginary part.Mask Model Architecture
[0053] Based on the log magnitude responses of the encoder, the central part of the model of certain embodiments computes a mask that is applied to the responses before being decoded. Following the simple and effective architecture proposed in [5] the mask consists of a feedforward layer, two gated recurrent unit (GRU) layers, and another three feedforward layers, as illustrated in Fig. 3. In the illustrated embodiment, the last layer uses a sigmoid activation function, all others are activated by a rectified linear unit (ReLU). In total, the mask model has 2.78m trainable parameters.
[0054] In the case of the hybrid filterbank, note that the trainable filters are preferably applied before the log magnitude is taken, which is to be distinguished from an architecture, where a convid layer is applied after the log magnitude for a fixed filterbank is taken.Training
[0055] For any given encoder filterbank Φ we are performing empirical risk minimization of a target signal x and the enhanced noisy signal x̃ with respect to a mixed compressed spectral loss on the encoder responses (φ being the phase of x in the representation Φx): MCS x x ˜ = γ ⋅ ‖ Φ x c e jφ − Φ x ˜ c e j φ ˜ ‖ 2 + 1 − γ ⋅ ‖ Φ x c − Φ x ˜ c ‖ 2 .
[0056] This loss is a generalization of the one introduced in [5], where it is computed with STFT coefficients. For fixed encoders, we found that it is crucial to design the objective function based on the representation that is also used to estimate the mask. For trainable encoders, this representation changes with training. Following [5] we choose compression and weighting terms as c = γ = 0.3 which has been found to perform best for the proposed mask model in terms of the highest PESO score
[28] . When using κ-penalization described in Equation (5), we minimize MCS β x x ˜ = MCS x x ˜ + β ⋅ κ .
[0057] By experimental exploration for our setting, we identified β = 10 -5< as a good value that is sufficiently small to not interfere with the minimization of the objective and sufficiently large to produce tightness consistently.
[0058] As optimizer, we use AdamW
[29] with a learning rate of 10 -4< and validate every 10 epochs. The batch size is 32. The model with the highest PESQ score on the validation set is chosen for being evaluated on an unseen test set, where we report the performance in terms of PESO and SI-SDR
[30] . This covers a perceptive, as well as a physical evaluation metric of the enhanced signals.Dataset
[0059] We use the CHiME-2 WSJ-0 dataset
[31] which consists of 7138 (train), 2418 (dev), and 1998 (test) speech utterances in English, from which we take 5 s excerpts, respectively. The sampling rate is 16 kHz. Every sample consists of a reverberate speech signal and a noise signal, added with a signal-to-noise ratio (SNR) ranging from -6 up to 9 dB in steps of 3 dB. The target signal is the corresponding reverberate speech signal. Using this relatively small dataset with short signal lengths allows to compare different models in short training times.TEST RESULTS AND DISCUSSIONGeneral
[0060] One benefit of certain embodiments disclosed herein lies in the enhanced usability of trainable filterbanks. The fixed filterbank can be flexibly chosen to fit the problem at hand. Construction, implementation, and training of the proposed hybrid filterbank is straightforward. Then, by enforcing tightness via κ-penalization, one or more of the following benefits arise: the encoder output level is under control and easily adjustable the output level is under control the decoder does not have to be computed stability: small perturbations have small effects, therefore not being susceptible to adversarial attacks.
[0061] In all our experiments, κ-penalization did not negatively influence the optimization of the main objective function, and we did not observe a noticeable loss in computational time.Speech Enhancement
[0062] The outcome of the speech enhancement task aligns very well with our intuition, as demonstrated in the following table: Fixed filterbankTrainable filterbankObjectivePESOSI-SDRκParametersSTF (baseline)-MCS3.199.8520audlet-MCS3.239.5810-conv1dMCS2.6611.693.28.1k-conv1dMCS β 2.7711.9918.1kaudletconv1dMCS3.388.861.22.8kaudletconv1dMCS β 3.398.6812.8k
[0063] The above results of a speech enhancement benchmark on the CHiME-2 WSJ-0 dataset compare the MCS (mixed compressed spectral loss; see Equation (7) above), PESQ (perceptual evaluation of speech quality), SI-SDR (scale-invariant signal-to-distortion ratio), and κ (condition number of the encoder; lower is better). The following conclusions can be drawn: As expected, the audlet encoder outperforms the STFT in terms of PESO. The hybrid audlet filterbank yields the highest PESO overall. Conv1d with random initialization yields the best SI-SDR.
[0064] As can be seen, the hybrid auditory filterbank according to certain embodiments not only outperforms all other models in some aspects, it can also arrive with an optimal condition number at the end of training due to κ-penalization. We note that κ-penalization has only little effect on the relatively short trainable filters here. For conv1d, the effect is larger (x = 3.2). Although not significantly, κ-penalization yields better scores in all the cases.
[0065] Conv1d takes a long time to train due to the small stride values. Furthermore, it is very sensitive to hyperparameters such as filter length, stride, learning rate, and the hyperparameter β. On the contrary, the hybrid filterbanks according to certain embodiments learn fast and work in many hyperparameter configurations.
[0066] We conjecture that the high SI-SDR score by conv1d comes from the fact that gradient descent treats every filter equally since there are no restrictions. Hence, the contributions of the filters average out over the different bands. As a result, the MCS resembles an energy measure. This is visible in the center plot of Fig. 4.CONCLUSION
[0067] Certain embodiments disclosed herein introduce a paradigm for designing stable trainable hybrid encoders with desirable properties for audio feature extraction, such as band limitation, and fixed center frequencies of the filters. The properties can be set in advance and persist throughout training. Moreover, a frame theoretic perspective provides the theoretical backbone for defining a simple and effective stabilization mechanism that keeps any trainable filterbank very close to tight throughout training. The implications are that the filterbank is norm preserving and can be inverted by its transpose. In a speech enhancement task, the hybrid filterbank according to certain embodiments manages to outperform the performance of the STFT and a randomly initialized conv1d layer in terms of the PESO score. In the same task, we successfully showed the applicability of the proposed gradient-based tightening procedure. While the present disclosure focuses on demonstrating the methods and concepts in a specific application, it shall be noted that they may well be applied to other tasks and domains, outlining the universality of the disclosed techniques.REFERENCES
[0068] [1] M. J. Palakal and M. J. Zoran, "Feature extraction from speech spectrograms using multilayered network models," in IEEE International Workshop on Tools for Artificial Intelligence: Architectures, Languages and Algorithms. IEEE Computer Society, 1989, pp. 224-230. [2] J. Lee, J. Park, K. L. Kim, and J. Nam, "SampleCNN: End-to-end deep convolutional neural networks using very small filters for music classification," Applied Sciences, vol. 8, no. 1,2018. [3] T. Sainath, R. J. Weiss, K. Wilson, A. W. Senior, and O. Vinyals, "Learning the speech front-end with raw waveform CLDNNs," in Proc. Interspeech, 2015. [4] Y Luo and N. Mesgarani, "Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation," IEEE / ACM Trans. Audio, Speech and Lang. Proc., vol. 27, no. 8, p. 1256-1266, 2019. [5] S. Braun and I. Tashev, "Data augmentation and loss normalization for deep noise suppression," in International Conference on Speech and Computer. Springer, 2020, pp. 79-86. [6] J. Heitkaemper, D. Jakobeit, C. Boeddecker, L. Drude, and R. Haeb-Umbach, "Demystifying TasNet: A dissecting approach," in Proc. ICASSP, 2020. [7] D. Ditter and T. Gerkmann, "A multi-phase gammatone filterbank for speech separation via Tasnet," in Proc. ICASSP, 2020. [8] Q. Li, F. Gao, H. Guan, and K. Ma, "Real-time monaural speech enhancement with short-time discrete cosine transform," arXiv, vol. abs / 2102.04629, 2021. [9] M. Dörfler, T. Grill, R. Bammer, and A. Flexer, "Basic filters for convolutional neural networks applied to music: Training or design?" Neural Computing and Applications, vol. 32, pp. 941-954, 2020.
[10] M. Ravanelli and Y. Bengio, "Speaker recognition from raw waveformwith SincNet," in IEEE Spoken Language Technology Workshop (SLT), 2018, pp. 1021-1028.
[11] N. Zeghidour, O. Teboul, F. de Chaumont Quitry, and M. Tagliasacchi,"LEAF: A learnable frontend for audio classification," in Proc. ICML, 2021.
[12] H. Seki, K. Yamamoto, and S. Nakagawa, "A deep neural network integrated with filterbank learning for speech recognition," in Proc. ICASSP, 2017.
[13] V. Lostanlen, D. Haider, H. Han, M. Lagrange, P. Balazs, and M. Ehler, "Fitting auditory filterbanks with multiresolution neural networks," in Proc. WASPAA, 2023.
[14] H. Bölcskei and F. Hlawatsch, "Noise reduction in oversampled filter banks using predictive quantization," IEEE Transactions on Information Theory, vol. 47, pp. 155-172, 2001.
[15] G. Yu, S. Mallat, and E. Bacry, "Audio denoising by time-frequencyblock thresholding," IEEE Transactions on Signal Processing, vol. 56, no. 5, pp. 1830-1839, 2008.
[16] P. Balazs, M. Dörfler, M. Kowalski, and B. Torrésani, "Adaptedand adaptive linear time-frequency representations: a synthesis point of view," IEEE Signal Processing Magazine (special issue: Time-Frequency Analysis and Applications), vol. 30, no. 6, pp. 20-31, 2013.
[17] M. Cisse, P. Bojanowski, E. Grave, Y. Dauphin, and N. Usunier, "Parseval networks: Improving robustness to adversarial examples," Proc. ICML, 2017.
[18] H. Bölcskei, F. Hlawatsch, and H. G. Feichtinger, "Frame-theoretic analysis of oversampled filter banks," IEEE Trans. Signal Processing, vol. 46, no. 12, pp. 3256-3268, 1998.
[19] T. Necciari, P. Balazs, N. Holighaus, and P. Sondergaard, "The ERBlet transform: An auditory-based time-frequency representation with perfect reconstruction," in Proc. ICASSP, 2013.
[20] T. Necciari, N. Holighaus, P. Balazs, Z. , P. Majdak, and O. Derrien, "Audlet filter banks: A versatile analysis / synthesis framework using auditory frequency scales," Applied Sciences, vol. 8, no. 1, 2018.
[21] O. Christensen, An Introduction to Frames and Riesz Bases, ser. Applied and Numerical Harmonic Analysis. Birkhäuser Boston, 2002.
[22] P. G. Casazza and G. Kutyniok, Finite frames: Theory and applications. Springer, 2012.
[23] P. Balazs, N. Holighaus, T. Necciari, and D. Stoeva, Frame Theory for Signal Processing in Psychoacoustics. Springer International Publishing, 2017, pp. 225-268.
[24] M. Ehler, "Preconditioning Filter Bank Decomposition Using Structured Normalized Tight Frames," Journal of Applied Mathematics, vol. 2015, pp. 1 - 12, 2015.
[25] D. Haider, V. Lostanlen, P. Balazs, and M. Ehler, "Instabilities in convnets for raw audio," arXiv:2309.05855, 2023.
[26] P. L. Sondergaard, B. Torrésani, and P. Balazs, "The Linear Time Frequency Analysis Toolbox," International Journal of Wavelets, Multiresolution Analysis and Information Processing, vol. 10, no. 4, 2012.
[27] K. W. Cheuk, H. Anderson, K. Agres, and D. Herremans, "nnAudio: An on-the-fly GPU audio to spectrogram conversion toolbox using 1D convolutional neural networks," IEEE Access, vol. 8, pp. 161 981-162 003, 2020.
[28] ITU-T, "Perceptual evaluation of speech quality (pesq), an objective method for end-to-end speech quality assessment of narrowband telephone networks and speech codecs," Feb. 2001.
[29] I. Loshchilov and F. Hutter, "Decoupled weight decay regularization," in Proc. ICLR, 2019.
[30] J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, "SDR - half-baked or well done?" arXiv:1811.02508, 2018.
[31] E. Vincent, J. Barker, S. Watanabe, J. Le Roux, F. Nesta, and M. Matassoni, "The second 'CHiME' Speech Separation and Recognition Challenge: Datasets, tasks and baselines," in Proc. ICASSP, 2013.
[0069] Although specific exemplary embodiments of the invention have been described, the person skilled in the art will readily understand that alternative embodiments may comprise only individual aspects, components, building blocks, or subsets thereof, which may provide their individual benefits as disclosed herein.
[0070] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
[0071] Embodiments of the invention may be implemented on a computer system. The computer system may be a local computer device (e.g. personal computer, laptop, tablet computer or mobile phone) with one or more processors and one or more storage devices or may be a distributed computer system (e.g. a cloud computing system with one or more processors and one or more storage devices distributed at various locations, for example, at a local client and / or one or more remote server farms and / or data centers). The computer system may comprise any circuit or combination of circuits. In one embodiment, the computer system may include one or more processors which can be of any type. As used herein, processor may mean any type of computational circuit, such as but not limited to a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a graphics processor, a digital signal processor (DSP), multiple core processor, a field programmable gate array (FPGA), or any other type of processor or processing circuit. Other types of circuits that may be included in the computer system may be a custom circuit, an application-specific integrated circuit (ASIC), or the like, such as, for example, one or more circuits (such as a communication circuit) for use in wireless devices like mobile telephones, tablet computers, laptop computers, two-way radios, and similar electronic systems. The computer system may include one or more storage devices, which may include one or more memory elements suitable to the particular application, such as a main memory in the form of random-access memory (RAM), one or more hard drives, and / or one or more drives that handle removable media such as compact disks (CD), flash memory cards, digital video disk (DVD), and the like. The computer system may also include a display device, one or more speakers, and a keyboard and / or controller, which can include a mouse, trackball, touch screen, voice-recognition device, or any other device that permits a system user to input information into and receive information from the computer system.
[0072] Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a processor, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some one or more of the most important method steps may be executed by such an apparatus.
[0073] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a non-transitory storage medium such as a digital storage medium, for example a floppy disc, a DVD, a Blu-Ray, a CD, a ROM, a PROM, and EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0074] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0075] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine-readable carrier. Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier. In other words, an embodiment of the present invention is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0076] A further embodiment of the present invention is, therefore, a storage medium (or a data carrier, or a computer-readable medium) comprising, stored thereon, the computer program for performing one of the methods described herein when it is performed by a processor. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory. A further embodiment of the present invention is an apparatus as described herein comprising a processor and the storage medium.
[0077] A further embodiment of the invention is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may, for example, be configured to be transferred via a data communication connection, for example, via the internet.
[0078] A further embodiment comprises a processing means, for example, a computer or a programmable logic device, configured to, or adapted to, perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0079] A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0080] In some embodiments, a programmable logic device (for example, a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
Claims
1. A hybrid auditory filterbank (100) configured for audio processing implemented on a data processing apparatus, wherein the hybrid auditory filterbank (100) comprises: a plurality of filters (102) configured to decompose an input audio signal (104) into a plurality of sub-bands (106); wherein the filters (102) are each based on a fixed filter (108) and a trainable filter (110); wherein the fixed filters (108) each comprise a filter of an auditory filterbank; and wherein the trainable filters (110) have been trained to perform an audio processing task and to improve the stability of the hybrid auditory filterbank (100).
2. The hybrid auditory filterbank (100) of claim 1, wherein the fixed filters (108) impose an inductive bias on the hybrid auditory filterbank (100) in terms of at least one audio processing characteristic; wherein the at least one audio processing characteristic comprises: center frequencies, bandwidths, filter shapes and / or impulse response supports.
3. The hybrid auditory filterbank (100) of claim 2, wherein the at least one audio processing characteristic is preserved during a training of the trainable filters (110).
4. The hybrid auditory filterbank (100) of any one of the preceding claims, wherein the hybrid auditory filterbank (100) is configured for speech processing, wherein the audio processing task is a speech processing task, in particular a speech enhancement task, and wherein the input audio signal (104) comprises speech data and, optionally, noise data.
5. The hybrid auditory filterbank (100) of any one of the preceding claims, wherein the trainable filters (110) are trainable in terms of convolution.
6. The hybrid auditory filterbank (100) of any one of the preceding claims, wherein the trainable filters (110) each comprise a one-dimensional kernel of a convolutional layer.
7. The hybrid auditory filterbank (100) of any one of the preceding claims, wherein the trainable filters (110) are randomly initialized.
8. The hybrid auditory filterbank (100) of any one of the preceding claims, wherein the filters (102) are formed by a composition, preferably a convolution, of the fixed filters (108) and the trainable filters (110).
9. The hybrid auditory filterbank (100) of any one of the preceding claims, wherein the trainable filters (110) have been trained such that the audio processing task is improved and the stability of the hybrid auditory filterbank (100) is improved by optimizing the condition number of the hybrid auditory filterbank (100).
10. An encoder for audio processing comprising the hybrid auditory filterbank (100) of any one of claims 1-9.
11. A decoder for audio processing comprising a transposed version of the hybrid auditory filterbank (100) of any one of claims 1-9.
12. A method of audio processing, comprising: receiving an input audio signal (104); and decomposing the input audio signal (104) into a plurality of sub-bands (106) using the hybrid auditory filterbank (100) of any one of claims 1-9.
13. A method of training the hybrid auditory filterbank (100) of any one of claims 1-9, comprising: training the filters (102) of the hybrid auditory filterbank (100) such that an audio processing task can be performed and the stability of the hybrid auditory filterbank (100) is improved, preferably by optimizing the condition number of the hybrid auditory filterbank (100).
14. A data processing apparatus comprising means for implementing the hybrid auditory filterbank (100) of any one of claims 1-9 and / or the encoder of claim 10 and / or the decoder of claim 11 and / or for carrying out the method of claim 12 and / or 13.
15. A computer program or a computer-readable medium having stored thereon a computer program, the computer program comprising instructions which, when the program is executed by a computer, cause the computer to implement the hybrid auditory filterbank (100) of any one of claims 1-9 and / or the encoder of claim 10 and / or the decoder of claim 11 and / or for carrying out the method of claim 12 and / or 13.