Hearing aid system
Patent Information
- Application Number
- EP2024709042
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-01
- Filing Date
- 2024-03-01
- Publication Date
- 2026-01-07
AI Technical Summary
Audio processing networks with encoder-masker-decoder architectures face limitations in generalization abilities, failing to effectively separate speakers in real-world scenarios due to over-fitting and poor performance on new datasets.
A variational inference approach is introduced to learn stochastic encodings and masks, using a likelihood with scale-invariance properties, and incorporating a probabilistic framework for improved generalization, rate-distortion analysis, and uncertainty quantification through multitasking and adaptive priors.
The proposed solution enhances speaker separation performance by improving generalization to new datasets, reducing distortion, and providing robustness and uncertainty quantification, leading to better real-world applicability of hearing aid systems.
Smart Images

Figure EP2024055472_06092024_PF_FP
Abstract
Description
[0001] HEARING AID SYSTEM Technical Field The present inventive concept relates to an audio processing network with an encoder-masker-decoder audio processing block. In particular, the present inventive concept relates to an audio processing network with an encoder-masker-decoder audio processing block adapted to provide source separation after decoding. Background Audio processing networks with encoder-masker-decoder architectures are effective speaker separation systems, but their generalization abilities have deficiencies that limit their real-world viability. Summary It is thereby an object of the present inventive concept to improve the speaker separation. Building on existing deterministic models, we propose a variational inference approach to learn stochastic encodings as well as stochastic masks that can be applied to the encodings for separation. To facilitate this, we introduce a likelihood with scale-invariance properties similar to a commonly used separation objective. We show that the approach improves generalization to new datasets while also improving overall test performance. The probabilistic framework further enables a wide range of modeling possibilities; we consider three aspects in particular: rate-distortion analysis on speaker separation generalization, the influence of prior specification on performance, and the relation between quantified uncertainty and performance. This and other objects are achieved by the features set out in the appended independent claims, with embodiments set out in the dependent claims. Accordingly, a first aspect of the inventive concept is provided by a hearing aid system. The hearing aid system comprises an encoder-masker-decoder audio processing block. The encoder-masker-decoder audio processing block is adapted to provide source separation, wherein said audio processing block comprises at least one of a stochastic encoder and a stochastic mask that has been learnt / taught / trained using variational inference, wherein said mask is applied to the encoder output in order to provide the source separation after decoding. The audio processing block is a software structure implemented e.g. using circuits in a hearing aid system. The audio processing block may e.g. be implemented in a smart phone or a separate audio processing unit, the smart phone or separate audio processing unit thereby forming part of the hearing aid system. Alternatively, the audio processing block may e.g. be implemented in (dedicated) processing circuits of the hearing aid system. The audio processing block comprises three parts. Audio is received and input into the first part, the encoder. This encoder may be implemented in any number of ways known to a person skilled in the art. The encoder may be a neural network trained on a dataset using a suitable distribution prior. The end result of the encoder may be encoded and / or compressed audio in a specific encoded space. The received audio may be a mono-channel audio signal and the goal of the audio processing block may be to separate different speakers in this audio signal into separate audio streams. Additionally or alternatively, the received audio may comprise noise and the goal of the audio processing block may be to identify and suppress the noise. Further goals and implementations of the audio processing block are defined by the attached claims. After the encoder encodes the received audio into encoded audio, the second part is the masker. The masker receives the encoded audio and changes it (using different mathematical operators and transformations) so to that source separation is achieved after decoding. The actual mask used by the masker is based on a stochastic estimation that is reiterated and refined using variational inference. The mask used may depend on certain factors, such as the type of input or the usage scenario. This may involve training a neural network such as a recurrent neural network to estimate the mask. The third part is the decoder, which receives the masked audio and decodes it into source separated audio. The source separation may be speaker separation or noise suppression. These applications are especially relevant for hearing aid users and are well- suited for the proposed processing block. The system may comprise a variational evidence lower bound. The encoder may be a variational auto-encoder (VAE) configured to optimize an evidence lower bound, ELBO, being a standard VAE ELBO with an added Kullback-Leibler divergence term. By using such an architecture, the encoder may be trained in a manner that avoids over-fitting to a specific training set, resulting in a more generalizeable / robust system. The variational evidence lower bound ^^^^^^^^may be given as: ^^^^^^^^= − ^^^^^^^^− ^^^^^^^^− ^^^^^^^^, wherein ^^^^^^^^is a distortion of single sources, ^^^^^^^^is a divergence of the mixing encodings from their prior, and ^^^^^^^^is a divergence of masks from their prior. This ELBO has shown to be especially beneficial. Priors used by the encoder may depend on the sources to be separated. This would allow for a more specialised and accurate encoder. The audio processing block may be adapted to perform multitasking by combining single source autoencoding and source mixture autoencoding. This enables the model to better adapt to changes in the system, such as using different hardware, such as microphones, than the model was trained using. The multitasking may provide uncertainty quantification by using input mixture density modeling to estimate performance without knowledge of targets. This enables a more robust system. The encoder-masker-decoder audio processing block may be a sepformer or be based on LSTM blocks. Such an architecture may be beneficial in specific embodiments, for example sepformers are especially good at speech separation. The encoder-masker-decoder audio processing block may be trained based on a probabilistic version of the scale invariant signal distortion. Such training may improve the source separation. The probabilistic version of the scale invariant signal distortion may be the scale invariant Bayesian linear regression observation model. This model has shown to be especially beneficial. Description of drawings The invention will be described in further detail with reference to preferred aspects and the accompanying drawing, in which: Figs.1a-d show schematic views of EMD models according to four embodiments, such as TasNets depicted in (a), share an encoder-decoder structure with AEs in (b) and their variational inference (VI) counterparts VAEs in (d). VI- EMDs, like VI-TasNets shown in (c), are robust, variational extensions. VI-EMDs learn distributions—over encodings z of audio and multiplicative masks m—by reconstructing single sources s from the input mixture x. Deterministic parametrizing networks left, mid, right: encoder (φ), masker (ψ), decoder (θ) respectively. Variables: explanatory (input), stochastic (round), deterministic (diamond). ⊙: Hadamard product. Fig.2 shows a multitasking VI-EMD showing the autoencoding tasks according to an embodiment. Single source autoencoding tasks are added with arrows going to and from za, zband the corresponding parts of the encoder / decoder blocks, and a mixture autoencoding task is added with arrows along the bottom. The use of + operator in mixing saand sbis a simplification to the Mix(◦)-mixing process. Fig.3 shows RD curves for speaker separation on synthetic (top) and real data (bottom) according to an embodiment. Stars denote optimum distortion within a dataset. LM: LibriMix, V: VCTK, L / M / H: low / mid / high noise setting. Fig.4 shows separation performance versus mixture AE for LibriMix (top) and VCTK (bottom) according to an embodiment. Fig.5 shows example outputs of a VI-TasNet according to an embodiment. From left to right: input mixture, mixture encoding sample, masks samples, estimates, and ground truth single sources. Based on an audio mixture input, the inference network provides an encoding of the mixture. These encodings are here visualized as the log-value of a sample from a gamma approximate posterior. The latent dimensions have been sorted using an agglomerative clustering over the full sentence input (for visualization purposes solely). The encodings are processed by the masker, and the masker provides distributions for a fixed set of speakers, here two. The top and bottom masks shown are samples from beta mask approximate posteriors for two different sources. The generative network sees the multiplication of the encodings and the masks to parametrize distributions of the estimated separated signals, and we visualize a sample from this distribution alongside the known ground truth. Fig.6 shows a comparison of encodings and spectrogram representation according to an embodiment. Like Fig.5, the latents were sorted with a clustering. From top to bottom: ground truth spectrograms, estimated signal spectrograms, latents, estimated time-series, ground truth time-series. From left to right: input mixture, source A, and source B. The latents shown for the mixture are the output of the encoder network, while the latents shown for the single sources are the masked mixture encodings. These results are not for a multitasking VI-TasNet, so no visualizations are present for the mixture in estimated spectrograms or time-series. Fig.7 shows rates versus distortion according to an embodiment. Rate and distortions values are normalized by T, and the rates are normalized additionally by K. Fig.8 shows rate versus negative SI-SDRi according to an embodiment. Left: normalized encoding rate. Center: normalized masker rate. Right: normalized total rate. Fig.9 shows an example of the synthetic Gaussian pulses dataset according to an embodiment. Top row: spectrogram of single sources targets in isolation. Bottom row: input mixture to the model. Fig.10 shows RD curves for various test conditions performance of models with varying target rates trained in a particular version of the synthetic Gaussian pulses separation test according to an embodiment. Note that the x-axis and y-axis are shared across all plots. Details in the text. NA: noise amplitude, OS: (number of) overlapping sines, OT: (number of) over-tones, AMN: amplitude-modulated noise, BPFN: band-pass filtered noise, GS: Gauss-pulse scale, FR: frequency range. Description of embodiments In the following, terms such as a / an / the and comprising are intended to be interpreted as non-limiting. Any processor or computing unit may be implemented as an electronic circuit and electronic connections may be wired or wireless unless otherwise explicitly stated. Unless explicitly specified, a wireless connection may be implemented in any number of standards known to a person skilled in the art, such as Wi-Fi, Bluetooth®, Zigbee, 4G / LTE, 5G and so on. Fig.1 shows an embodiment of a hearing aid system according to the invention compared to the prior art. Fig.1a shows a prior art TasNet, which has shown to be effective at source separation within a training set, but has proven to be difficult to make generalizeable / robust on real-world data. Fig.1b shows a prior art auto-encoder (AE) and Fig.1d shows a variational auto-encoder (VAE). By changing the deterministic operations in the auto-encoder of Fig.1b to stochastic operations in Fig.1d, the VAE has been shown to be more robust. Inspired by this, the hearing aid system according to the invention uses the encoder-masker-decoder audio processing block of Fig.1c, which has replaced the deterministic operations of Fig. 1a with stochastic operations, leading to improved robustness. In other embodiments of the invention, the encoder-masker-decoder audio processing block may be used in other networks / architectures than the VI-TasNet as shown in the specific embodiment of Fig.1c. The embodiment of Fig.1c only shows a separation of two sources, however the principle may be generalized to any number of sources. Speaker separation, and its application towards solving the cocktail party problem, has been studied for many decades, especially from the perspective of classical (digital) signal processing. Re-framing the problem as a deep learning (DL) supervised learning problem has enabled many advances in how well speaker separation can be done. A series of models following an encoder-masker-decoder (EMD) architecture performs particularly well. These models operate on a raw waveform representation, building upon time-domain audio separation networks (TasNets) (see Fig.1a). These models separate speech in a learned encoding space. Separation is done by estimating masks that are applied to the mixture encoding such that decoding the resulting masked encodings provides estimates of the component sources in isolation. The original TasNet relied on recurrent neural network (RNN) architectures to estimate masks. In certain scenarios, a convolutional variant was shown to outperform an ideal ratio mask baseline (which has access to ground truth signals in producing the masks) for both objective distortion measures and perceived, subjective audio quality. Later work has improved performance by, for example, introducing improvements to the masking network such as using dual- path recurrent neural networks or using attention mechanisms. TasNets learn to separate speech in a manner that does not generalize well beyond the training dataset and task, limiting their utility in a real-world setting; their inter-dataset generalization is poorer relative to the intra-dataset generalization, i.e., poorer on test data from an unseen dataset as opposed to test data from the same dataset that the training data came from. Similarly, presenting mixing procedures different from the training task reduces generalization. It has been shown that training on a dataset, LibriMix, shows higher performance on the WHAM! test set, which is a noisy extension of WSJ0-2mix, than on the LibriMix test set. However, performance reductions of about 1-2 dB are found when testing on, e.g., the larger, more diverse VCTK. In standard supervised learning, a mapping is learned from a high- dimensional input to a low-dimensional supervision label. DL models with sufficient capacity will generally be able to learn near-perfect mappings for simple problems on the training data, but if these models are not adequately constrained, the implicit representation of the input will not generalize. This over-fitting is caused by a model learning to map uninformative characteristics of training data points to their labels. Models like TasNets can, for the same reasons, fail to generalize if not properly regularized, even though the supervisory signal is more high-dimensional. Generative models, on the other hand, aim to learn the distribution of data. By tasking the model with generating data, representations can be learned without explicit guidance from a label or supervisory signal. The generative task provides a learning signal to infer patterns that are robust (generalize well) such that the learned representations can be usefully applied to another task of interest. Methods for deep generative modeling include, among many others, directly modeling the likelihood of the data diffusion models, or using variational inference (VI) with variational auto- encoders (VAEs). A VAE is a latent variable model that aims to learn a representation of high-level features of data in a scalable manner through amortized approximate inference of variational distributions. Given data, x, a VAE optimizes an evidence lower bound (ELBO), where θ, φ are the encoder and decoder parameters, respectively, qφ(z|x) is a variational approximation to a true, but intractable, posterior pθ(z|x), Eqφ(z|x) denotes an expectation w.r.t. the variational distribution, and DKL is the Kullback- Leibler (KL) divergence. The distributions pθ(x|z) and qφ(z|x) are parameterized by the decoder / generative network and by the encoder / inference network, respectively. The encoder maps from a given data point, x, to a latent representation, z, and the decoder learns to map from the latents back to the data. Often the parameterized distributions are Gaussians, and an isotropic Gaussian is used for the prior, p(z). By optimizing LVAE, we ensure that the model learns—in some sense—a well-behaved, compressed representation of x (ensured by the “KL divergence term”) while still being able to reconstruct the data (ensured by pθ(x|z), the “reconstruction term”). VAEs have been successfully applied in modeling data in a wide variety of domains, such as computer vision, chemistry, natural language, and astronomy. VAEs have also been applied for time-domain audio modeling in general; a high-capacity encoder-decoder can, for example, learn representations that enable voice conversion and has learned representations that correlate with high-level representations such as phonetic content. Balancing the terms in the ELBO can provide different behaviors of the representation. These terms are the negative log-likelihood of the data, or distortion, D, and the deviations from the prior, i.e., the KL between the approximate posterior and the prior, or rate, R. Different trade-offs can be achieved, for example, with a re- weighting as in β-VAEs. Targeting lower rates under isotropic Gaussian priors can provide better disentanglement by some measures and, notably, affects generalization. We can visualize these trade-offs as rate-distortion (RD) curves, and the RD-trade-off has been discussed, feasible and realizable models, as well as the relation to the entropy of the data and the capacity of the model. Notably, the information bottleneck principle shows how there is an optimum rate when seeking to improve generalization. In the following, we explore whether a generative approach can characterize and improve the generalization of EMDs, such as TasNets, leveraging that these networks have an encoder-decoder structure, similar to auto-encoders (AEs) (see Fig.1). Specifically, our contributions over the prior art are: i) a variational inference encoder-masker-decoder (VIEMD) framework, alongside which we introduce (a) a separation ELBO, LS, and (b) a scale- invariant Bayesian linear regression (BLR) observation model, ii) introducing rate-distortion analysis as a method of understanding generalization in speaker separation, and a comparison between the introduced stochastic VIEMDs and their deterministic counterparts, and iii) introducing a multitasking version of the VI-EMDs and an approach to do uncertainty quantification through input mixture density modeling to estimate performance without knowledge of targets. Speaker separation is the task of recovering a set of clean, single sources, S = {s0, ... , sN}, from a noisy, mixture of the components, x, under some mixture generating function, Mix. A simple expression with a set of speakers and a single noise source of this could be x = Mix(s1, ... , sN, n) = n +∑ ^{^^^^^^^=1}^^^^^^^^, where x, si, n are time-series (for example, x = [x0, ... , xt, ... , xT]) of length T, and n is some interference / noise. In this work, for Mix, we consider a simple additive mixture process for a mono-channel audio signal. In an EMD (see Figure 1a), an encoder (fφ(x) = z) provides an encoding (z) for an input mixture (x). The encoding is fed to a masker (hψ(z) = M) which provides a set of masks (M = {m0, ... ,mi}) that will be applied to the encoding. Applying a given mask (mi) to the encoding provides a masked encoding ( ^^̂^^i= mi⊙ z). Finally, the decoder (¨si= gθ(ˆzi)) outputs estimates of the single sources (¨si) based on the masked encodings. We propose a probabilistic modeling variant of EMDs. We introduce variational inference TasNets (VI-TasNets) that closely resemble a well-studied (deterministic, Conv-)TasNet. Note that we refer to the Conv-TasNet as the deterministic TasNet, or just TasNet, for brevity. We stress that the approach can readily be used for any EMD architecture, notably also, e.g., improved successors to TasNets, which we demonstrate by applying the VI-EMD framework to SuDoRMRF. Like TasNets, VI-TasNets use simple, singlelayer, convolutional encoders and decoders. For TasNet, the outputs of the encoder directly provide encodings, and in VI-TasNet, the same outputs instead parameterize encoding distributions. Similarly, while the TasNet masker and decoder directly output masks and the estimated sources, the VI-TasNet parameterizes distributions over the masks and the estimated sources. An overview of the model is shown on Fig.1c. Example model outputs are given as visualizations later in the description. Encoder and masker distributions: For the encoder distribution, we consider a large, over-complete K-dimensional latent space (K = 512), matching the TasNets. For an input time-series x of length T to the encoder, we obtain a distribution of a latent time-series of length T′. The sampling frequency of the two is related through 5 the strides of the encoder-decoder structure; we choose a lower latent space tick, such that T / T ′ = 8. Note that the input time-series is one-dimensional, x ∈ ^^^^1× ^^^^, while the latents have K dimensions per latent time step. Commonly, an isotropic Gaussian prior imposing less co-varying, smaller magnitude latents is used. This prior can be prohibitively restrictive, and using more expressive priors can improve 10 learning. Other characteristics can be imposed by using non-Gaussian priors, such as a directional or nonnegative distribution. We opt for factorized Gaussian approximate posterior for the encodings parametrized by fφ; similarly, we opt for log- normal masks parametrized by hψ (mimicking non-negative properties of a ReLU activated TasNet mask): 15 where, for example, ^^^^^^^^^^^^,., time-series of length T′ with distribution parameters (location and scale) for the k’th latent dimension output. That is, we explore a model going beyond the standard isotropic Gaussian prior using log-normal mask priors. For further specifics and prior specifications, we refer to later in the description. Also 20 later, we discuss other choices of approximating families, priors, and their performance. Decoder distribution: TasNets are usually trained towards maximizing the scale-invariant signal distortion ratio (SISDR), which for a known source s and estimated source ˆs is defined as: SI-SDR(s, ˆs) = 10 log10 (||αs||2 / ||αs − ˆs||2), where 25 α = designates the 2-norm. The α coefficient rescales the target such that the error in the denominator between the estimated source and target source is measured in a way that is invariant to the overall scale (power) of the time- series. While the scale-invariance is a central part of the formulation of SI-SDR, it notably measures the error in a logarithmically scaled manner. Often an SI-SDR 30 improvement (SI-SDRi) is reported, giving the increase in SI-SDR by using the processing over using the input mixture as the estimate of the single source. A similar objective function to SI-SDR that does not re-scale the target but retains the logarithmic scaling of the errors is the logarithm of the mean-squared error (MSE). A log-MSE has been shown to enable training of TasNets in a manner similar to the SI- SDR, while a standard MSE does not achieve comparable results. That is, training a TasNet with an MSE loss (as opposed to using a SI-SDR or log-MSE) does not provide performant speaker separation in TasNets. In standard VAEs, a typical choice would be to parameterize a per time-step Gaussian for the generative network: where μ., σ2. Are time-series output of the decoder, gθ. Given a fixed unit variance, this would correspond to an MSE loss, which, as discussed, does not train performant TasNets. In an alternative embodiment discussed later, we discuss a multivariate Cauchy objective (MVC) since the log-likelihood conceptually enables us to minimize a log- error similar to the log-MSE. Since the SI-SDR is not a likelihood, we are interested in developing a likelihood that is similarly invariant to a rescaling of the time-series. Where the SI-SDR uses α, we take the view that a similar factor γ is a regression coefficient. Using Bayesian linear regression, we wish to model both γ and a noise- scale parameter, σ2. For the target time-series s with steps st, we consider a BLR model in which the approximation ˜s takes the role of predictor: Conjugate priors for σ2and γ take the form p(γ, σ2)= p(σ2)p(γ| σ2), where p(σ2) is an inverse-gamma distribution with parameters a0 and b0, Inv-Gamma(a0, b0), and the conditional prior distribution for the regression coefficient is a normal distribution with a mean μ0, and variance σ2λ0−1, where λ is a scalar prior precision for the regression coefficient. Update rules for these parameters enable us to determine a likelihood for given time-series given the input ˜s when integrating over γ and σ2, This provides an analytical expression for the likelihood which we use in the optimization of VI-EMDs as the distortion measure. For further details, see later in the description where we show a comparison to the SI-SDR and how they optimize related quantities. Separation evidence lower bound: Fig.1 shows a two-speaker scenario, matching the data used for the various experiments (i.e., in the figure we have S = {sa, sb}). We assume a joint distribution over single sources (S), their mixture (x), corresponding masks (M= {m0, ... ,mN}), and the mixture encoding (z) like follows: with VI- EMDs takes the form (derivation later in the description): where DS is the distortion of the single sources (how well they are reconstructed), Rz is the divergence of the encodings from their prior (the encoding rate), and Rm is the divergence of the masks from their prior (the mask rate). Different from direct auto- encoding as in VAE, a VI-EMD reconstructs single source components instead of the original input to the encoder. To explore the RD trade-off, we use various modified losses based on the separation ELBO during training by means of free bits, adaptive re-weighing of the rate, or a β-coefficient. We discuss these aspects in greater detail later. In particular, we also discuss the representation learning aspects of RD curves, how these trade-offs relate to compression, and how these trade-offs are also related to further quantities regarding matching the overall data distribution (similar to a standard generative adversarial learning objective), and how, e.g., over- completeness and sparsity fits into VAE-based representation learning. Flexible priors: We can consider priors that are more flexible than the standard, static ones. Later, we discuss a learnt auto-regressive flow prior. While such a prior constitute a powerful, domain-agnostic approach to a more flexible prior, we also introduce a domain-inspired adaptive prior which promotes that mixture encodings resemble an addition of single source encodings. Such a prior is in part motivated by the separation task, seeing as the separation-by-masking is assuming a similar underlying generation mechanism. In itself, the use of a masker in the encoded space with element-wise masks on the mixture encoding can be thought of as an inductive bias, or architectural prior, under which the model is learning to separate. The addition of random variables corresponds to a convolution of their distributions. For the encoding distributions, we use we have closed-form convolutions (to be described more later). Considering encodings of single sources for the known true single sources, si, in Fig.1, and using a notational shorthand for the encodings of the single sources, we define an adaptive prior for the mixture encoding as: where ⊛ denotes a convolution operation, which amounts to enforcing that the mixture encoding is a sum of the single source encodings z = za+ zb. We add corresponding rate terms to reflect the single source encoding divergence from a prior. This is needed to use the adaptive prior to ensure that we actively promote that qφ,aand qφ,bresemble the distributions that we convolve in making the adaptive prior. Fig.2 shows a multitasking VI-EMD as another alternative embodiment of the encoder-masker-decoder audio processing block of the invention. Compared to the embodiment of Fig.1c, Fig.2 has added single source autoencoding tasks and a mixture autoencoding task. Together, these can be used to improve the source separation of the system when deployed to real-world scenarios. Multitasking VI-TasNet with autoencoding objectives: VAE-based semi- supervised learning enables learning from large amounts of unlabeled data, such that the representation can be used to efficiently learn how to solve a specific task using fewer, but labelled, examples. A model that utilizes the described adaptive prior already obtains encodings of the single sources. We can utilize the single sources as a target signal, too, combining the VI-TasNet separation task with a single source AE task, sharing the encoder and decoder parameters. This amounts to adding a distortion term stemming from decoding the single source encodings instead of masked mixture encodings. We can further augment such a multitasking model with a mixture AE task, especially relevant since we do not have the true single sources in a real-world scenario. We reconstruct the original mixture from the obtained mixture encodings and add a mixture reconstruction distortion term. Adding mixture AE also provides a quantification of how well the model fits the mixture, and we investigate whether the mixture AE performance is indicative of separation performance, which would provide a principled uncertainty quantification for the separation system. Fig.2 shows the multitasking VI-TasNet. The orange arrows enable the auto-encoding task on the single sources, whereas the purple arrows enable the mixture autoencoding task, which can be done without any knowledge of the single sources. When using the adaptive prior, that the visualized za and zb are added (their distributions convolved) in making the adaptive prior for the mixture encoding, z. Comparison with later TasNet variant: Finally, we can investigate how the variational inference procedure affects later developments of EMD models. We choose to investigate the “successive downsampling and resampling of multiresolution features” (SuDoRMRF) model, as this model—like the original (Conv)TasNets—seeks to have a minimal computational and memory footprint. The SuDoRMRF model uses a more efficient, convolutional separation module relying on an architecture reminiscent of U-net. We isolate the effect of the VI framework and enable a comparison to the (VI-)TasNet results by only replacing the masking network, and otherwise keeping everything the same. For further details on this, we refer to later in the description. Later EMDs similar to TasNets remain competitive models for speaker separation. Improvements include improved dilated temporal convolution blocks and multi-scale modelling, specialized RNN architectures, and utilizing transformers / attention mechanisms in the masker. Other separation networks rely on, e.g., explicitly modeling speakers and perform separation using a learned speaker stack, or—in line with others—use specialized RNNs. Early work separated using a clustering approach, and attractor-based systems have been proposed. Beyond the currently considered anechoic problems, recent works have more explicitly considered reverberant environments and spatialized problems, for instance by incorporating neural beamformers. While we use VI to improve generalization, pre-training tasks have been considered; some show improvements with speech enhancement pre-training, and a range of self-supervised learning approaches. Other improvements to the training procedures include MixIT and ReMixIT. Spectral speaker separation VAEs: Speaker separation in the frequency domain using VAEs has seen many studies in recent years including also multi- channel models, and models integrating visual information. More general considerations on modelling audio spectrograms with VAEs are known. While these methods employ VAEs towards speaker separation, they operate in a spectral representation (often using the short-term Fourier transform), where the presently considered work is concerned with directly modeling the time-domain audio waveform. The difference in TasNets modeling spectral or temporal representations is substantial. It can be beneficial to use spectral losses to optimize AEs even if operating in the time domain. Hybrid TasNets are proposed that use both spectral and temporal representations. Such hybrids provide an avenue for leveraging the extensive study of spectrogram VAEs with the present work. Time-domain modelling: Deep learning speech separation in the time-domain has been done using WaveNets, e.g. in a generative modeling framework, or in discriminative, non-autoregressive variants. As audio modeling with, e.g., generative adversarial networks or flows improves, similar concepts are applied in speech separation and enhancement including both adversarial or flowbased methods. In all cases, the models produce speech separation with high-quality outputs but they rely on a high model complexity (compared to TasNets) to achieve these results. Structured State Space sequence (S4) models have improved audio modeling, but have yet to be applied to speaker separation. VAE applied to time-domain audio notably include, e.g., VAEs with WaveNet decoders, and more recently NaturalSpeech, used for text to speech. Time-domain VAEs are less studied for speech separation (speaker separation and / or speech enhancement). One example is the variance constrained (VC) AE for speech enhancement. The VC AE focuses on a different task and does not optimize a VAE objective, and the VC AE model differs from the present work in that the VI-TasNet relies on the same architecture as the well-studied EMDs / TasNets (using masking of the encodings for separation). Fig.3 shows RD curves for speaker separation on synthetic (top) and real (bottom) data. Stars denote optimum distortion within a dataset. LM: LibriMix, V: VCTK, L / M / H: low / mid / high noise setting. The results show different trade-offs for rate, distortion and generalization that match the information bottleneck perspective and further informs how the present inventive concept improves generalisation by optimizing rate and distortion in a principled manner. Synthetic data: We introduce a simple synthetic dataset for source separation; the dataset constitutes a controlled, simplified speaker separation problem, further details later in the description. The datasets consist of “2-speaker” mixtures, where the sources are generated as a number of overlapping randomly-generated sinusoidal Gauss pulses in a speaker-specific frequency range with overtones. On this dataset, we fit a series of VI-TasNets using an adaptive re-weighting of the total rate (encoding rate and the masking rate) towards a desired target rate. For the synthetic data, we reduce the capacity of both the encoder-decoder and the masker to fit the complexity of the problem by reducing the number of filters / channels, but otherwise we use the same architecture overall as are used in later experiments. LibriMix / VCTK data: We consider N = 2 talker mixtures with a sampling frequency of 8 kHz from the LibriMix dataset. We use the same processing and splits and consider mixtures of length matching the shortest single source (the “min” mode). We train models either on the clean or the noisy 100 hour variant (Libri2Mix train-100). Single replicates are reported for these models considered, and randomness in the initialization is not characterized; further specifics on training / compute-requirements and model / results limitations will be discussed later. We evaluate the performance of these models in their ability to generalize to both a “familiar” Libri2Mix test set and to the “unfamiliar” VCTK-2mix test. RD and generalization In Fig.3 (top) we show test-set RD curves for models trained on a given level of Gaussian additive noise (dashed, black, medium noise) on the synthetic data (lines are running means of distortion as a function of sorted rates). The curves show the expected trade-off of poor distortion at low rates. Note that rates are normalized by the number of latent dimensions. The models with the highest rates (above 1.25 nats / dim) display poorer separation performance than models with slightly lower rates, indicating that high-rate models learn sub-optimal representations. We evaluate the same models in lower (full, bottom) and higher noise settings (dotted, top), and we see that an optimal rate lower than the highest rates exists in all conditions. Furthermore, the distortion difference between optimum and high-rate models grows in a low noise setting compared to the training domain. This is evidence of how expressive models that optimize single source distortion over rate produce poorer generalization. In particular, these results show that the variational framework for separation displays trade-offs for rate, distortion and generalization that match the information bottleneck perspective. This, in turn, provides an avenue for improving the generalization of EMD models, since we can learn models that jointly optimize rate and distortion in a principled manner. Further details and conditions are discussed later. We investigate the same concept in the real speaker separation data. We evaluate test-RD curves for models trained on clean LibriMix. These models were trained using various levels of free-bits, details later. Models with very low rates and higher distortions were trained but omitted in the visualization for clarity. In Fig.3 (bottom), we see that models with rates that are too low perform poorly, and we see indications that the model with the lowest distortion in LibriMix does not achieve the best distortion in VCTK, where instead a lower-rate model performs the best. The deterministic models will solely optimize a distortion and do not quantify the informational rate, and so the results for generalization (from a LibriMix-trained model to VCTK) are consistent with the need for considering models that can quantify and optimize the distortion and rate jointly to achieve improved generalization. Table 1: SI-SDRi (and performance drop from LibriMix to VCTK) for models trained on the 100 hr noisy 2-speaker LibriMix data. Models are not trained on the clean condition, it is an unseen test condition, similar to how VCTK is an unseen dataset / domain.†: using only the SuDoRMRF separation module. N: noisy, C: clean. Improved generalization to unseen data and conditions: We train a VI-TasNet with Gaussian encodings and lognormal masks using the BLR likelihood optimizing an adaptively reweighted ELBO towards a target total rate. As the deterministic baseline, we train a TasNet using the standard SI-SDR objective. Having trained the models on the noisy LibriMix, we contrast performance seen / unseen conditions (noisy, clean) and seen / unseen datasets (LibriMix / VCTK). The performance of these models is shown in Table 1. We see that the VI-TasNet is a strict improvement over the deterministic TasNet both in seen and unseen datasets and conditions. Specifically, noisy LibriMix / VCTK test-performance is improved by 0.37 and 0.55 dB respectively, with a 2 %-points lower relative drop. Similar generalization improvements to the unseen clean condition are observed with the VI-TasNet. While this improvement comes at negligible increases in parameter count and inference time computations, it does increase the training time. Later we show results for TasNets trained with both SI-SDR and BLR on the clean condition, which we compare to VI-TasNet with different priors. Similarly, when we evaluate the framework on a different masking network, the SuDoRMRF, we see consistent, but small improvements in generalization. Fig.4 shows separation performance versus mixture AE for LibriMix (top) and VCTK (bottom) using the multitasking VI-TasNet of Fig.2. Better performance corresponds to the bottom left of the plots. The plots and related text to follow show that the multitasking VI-TasNet of Fig.2 may be self-correcting when deployed on a hearing aid system. Uncertainty quantification: We train a multitasking VITasNet on the clean condition. Importantly, besides separation, this model also learns to do mixture AE. We leverage the mixture AE task to quantify how well the shared encoder-decoder structure models the input mixture (input density). We investigate whether the mixture AE ELBO is informative of the separation performance, as measured by the single source distortions in the separation task. Fig.4 shows the separation performance measured as BLR estimated single source distortion as a function of the negative mixture ELBO for each mixture in the LibriMix and VCTK test sets— each dot is a mixture, and contours are from a kernel density estimator used solely for visualization. Lower values for both values indicate improved performance (either better separation for the distortion or higher evidence for the mixture AE). Mixture AE is informative of the separation task performance, and the lowest distortions also model the mixture the best—conversely, when the mixture is poorly modelled, the separation distortion increases. We note that there is a slightly increased number of mixtures (dots) at lower mixture AE and higher distortion for VCTK, indicating mixtures for which the input density is a poorer predictor of separation performance. The positive correlation effectively provides a model that can quantify its uncertainty in performing the separation task. The mixture is available in a real scenario, and so if the AE task is performing poorly, the results suggest that this would be indicative of poor separation performance, too. While the two quantities are not normally distributed or linearly related, we can quantify the correlation as a Pearson correlation coefficient, or we can quantify it with a non-parametric Spearman rank correlation coefficient, instead. Here, we report both and see a weak to moderate correlation that decreases slightly in the new domain (Pearson correlation rp, Spearman, rs, and p <1e−50, n = 3000 per data set and condition): rp= .37, rs= .45 (LibriMix) and rp= .28, rs= .35 (VCTK). Limitations: A limitation of these models is their requirement of separation into a fixed number of speakers; this is a central problem addressed in other work. We can consider speaker separation in quiet, in simplistic noise (additive Gaussian noise), and more realistic noise (using realistic recordings of background noise of speech). Realistic data and problems are key to developing models viable for actual use. Standard benchmarks often investigated, such as WSJ0-2, often lack diversity in the speakers and recording conditions, have unrealistic mixing process (e.g., with too high overlaps), have no consideration of reverberation, use a fixed number speakers, or are based on non-ecological speech material (i.e., speech recorded, for example, while reading written material aloud as opposed to during a natural conversation). Later datasets have since addressed some of these issues, for example by extending WSJ with realistic, reverberant noises in realistic conditions, integrating multiple corpora, a more diverse set of speakers, sparse speaker overlaps, and varying number of speakers and sound types. In this work, we relied on LibriMix, since it both includes realistic noises (from WHAM!) and a large, diverse set of speakers based on LibriSpeech. While we focus on mono-channel separation, multi-channel processing that can utilize spatial information will be more relevant in reverberant environments. Broader impact: Improved speaker separation systems will facilitate the improvement of hearing aids and thus improve the treatment of hearing loss. DL systems in hearing aids, however, will likely impose increased demands on, e.g., internet connectivity or hardware capabilities, meaning these improvements will likely reach listeners with access to more resources first. Improved speaker separation improves automatic speech recognition systems and, for example, transcription automation. Probabilistic models that more explicitly impose priors enable interpretability, while also potentially improving learning in the face of more scarce data (e.g., smaller non-English corpora). Generative modelling enables the production of deep fakes and systems that can, for example, mimic a given speaker’s voice can be used to conduct fraud and produce fake media. While many of the examples shown use a VI-TasNet, the principles of the inventive encoder-masker-decoder audio processing block concept are applicable to different models. For example, by using attention blocks and transformer layers in the masker, a sepformer model is created with the same improved robustness and speaker separation results. This is done by changing the masker from using convolution blocks based on variational formulation to variational formulations of attention blocks and transformer layers. Attention blocks uses info from the signal itself to find what to attend to in each kernel. It is noted that the main motivator for using a VI-TasNet for the plots presented in the Figs. is limited time and resources to test other embodiments, it is not to say that it is better or worse than any other embodiment of the inventive encoder-masker-decoder audio processing block concept. Non-limiting embodiments presented herein: We have presented variational inference encoder-masker-decoder models, and particular instantiations in the variational inference time-domain audio separation network, VI-TasNet, and VI- SuDoRMRF. The VI-EMDs effectively learns to separate audio while learning distributions of latent encodings and latent masks in a manner that improves tests performances on seen conditions and test data from seen datasets, but also improves generalization to unseen conditions and new datasets. The VI-TasNet uses a Bayesian linear regression likelihood, which enables likelihood-based training with scale-invariance similar to scale-invariance signal-distortion-ratio. The probabilistic formulation of the model provides the means of imposing priors on the learning, we discuss an adaptive prior and provide results in the supplementary on how various priors impact the separation performance. We show how the generalization of VI- TasNet can be characterized using rate distortion tradeoffs; we show indications that, while trading off increased rates sometimes improves performance within the condition and within the same dataset, this does not necessarily generalize to a new dataset. Lastly, we show that a multitasking VI-TasNet performing mixture autoencoding can quantify its uncertainty in a manner informative of the separation performance without requiring access to the single sources. A derivation of VI-TasNet evidence lower bound, building on standard result of the ELBO for standard VAEs: We start by considering the standard VAE evidence lower bound (ELBO), and use this as a starting point for showing the slightly more notationally involved bound for the VI-TasNet. Conceptually, the extra latents simply add an extra KL divergence term in the VI-TasNet formulation. We will write the Kullback-Leibler divergence between distributions a(x) and b(x) as (note that we are using lower x here to denote a stochastic variable): VAE evidence lower bound: In practice, we consider data from a particular dataset, D, and consider an empirical data distribution pD(x) = (where ||D|| denotes the cardinality of the dataset), which we hope reflects some true data distribution. The derivations below follow for a single sample from pD(x), and we derive a lower bound on the evidence conditioned on that sample, L(θ, φ; x), dependent on the generative and inference network parameters. We optimize the bound for such samples, but this can be also be extended to considering an expectation over the dataset and in a batch setting, too, such that we can extend this to consider a total bound over the dataset, L(θ, φ) = Ex∼pD(x)[L(θ, φ; x)]. We consider random variables x and z (adopting lower case notation for this). Note that pθ(x, z) = pθ(z|x)pθ(x) = pθ(x|z)pθ(z), and e.g. pθ(x) = pθ(x,z) / pθ(z|x) (assuming here that the denominator is non-zero everywhere), where the subscript θ denotes the dependency on generative network parameters. An expression for the ELBO can be arrived at by introducing an expectation over the variational distribution qφ(z|x) (dependent on the inference network parameters φ) and re-arranging:
[0002] Where we used that the KL divergence between the approximate posterior over the latents z and the true (but intractable) posterior is a non-negative quantity, such that the model evidence is lower bounded by the expression LVAE. This expression can be rewritten as: where the first term corresponds to the negative distortion (the distortion is the negative log-likelihood), and the second term corresponds to a rate (the KL- divergence between approximate posterior and prior). VI-TasNet evidence lower bound: We denote a set of single sources of time- series S = {s0, ... , sN}. Fig.1 shows a two-speaker scenario, matching the data used for the various experiments (i.e., on the figure we have S = {sa, sb}). We assume a joint distribution over single sources (S), their mixture (x), corresponding masks (M= {m0, ... ,mN}), and the mixture encoding (z) like follows: where we assume that We also assume that these distributions factorize over the temporal dimension (latent or original data space), as e.g. shown for the mask prior later. With repeated application of the product rule, we have that: The assumption for the set of masks might be worth exploring, seeing as e.g. self-consistency in the deterministic TasNet under the sometimes used softmax / sum- to-one-like constraints enforce a dependency between masks. The formulation with pθ(x|S) could accommodate some stochastic mixing process, which is not the case for the data and model we are considering (For the LibriMix data considered, the dataset is static, and while the mixing process randomly samples a SNR in a particular range during the creation of the dataset, this mixing SNR does not change after the dataset is made), and instead x = Mix(S) is a deterministic mapping from sources to a mixture, or pθ(x|S) = δ(x − Mix (S)). Realistic mixture generation functions are not noise-free or additive, but should e.g. take into account a reverberant environment with different spatial locations of speakers and multiple noise sources. In this setting, multi-channel recordings are of value in enabling resolving different spatial locations, and deep learning systems in general can benefit from utilizing systems that have traditional been used to improve performance. Following a similar derivation of the ELBO for the VAE, we start by introducing an expectation over the variational approximation to the latent encodings arising from the inference network with parameters φ, and an expectation over the variational approximation to the latent masks arising from the masker network with parameters ψ. We also, for the purpose of illustration, include an expectation over the mixing process, initially: Dropping the mixing process expectation and using x instead of Mix(S) to highlight it as the input to the inference network, the expression can be written as: Splitting the factors within the logarithm out as addends, and noting that the expression with encodings does not depend on the mask latent, we get: Since the divergences are non-negative quantities, measuring how close the variational approximation for the latent encodings and latent masks are to the true posteriors, the last term is a lower bound on the evidence over the set of speakers, L(θ, φ, ψ; S). We can write this as: We can also split the bound in three contributing terms: Here, DS is the distortion of the single sources (how well they are reconstructed), Rz is the divergence of the mixture encoding from their prior (the encoding rate), and RM is the divergence of the masks from their prior (the mask rate). Note that since we assume the single sources conditioned on their mask and the mixture encoding factorize, the distortion is the sum of the log-likelihoods for each individual source. As is common when training VAEs, we estimate the expectation with single samples from the approximate posteriors for z and M, and have defined priors for these latents, which enable us to evaluate the objective. For the models we consider, we are also using a prior on the masks that is independent of z, both for the log-normal masker (used previously) and the beta masker discussed here: where k indicates a particular latent dimension, and t′ indicates the latent time-step, and i is speaker index. We hypothesize that e.g. a dependency on the strength of the prior mask based on the “energy” in the encodings (the distance from zero in the Gaussian case, or the concentration parameter in the gamma case) might improve learning. In the case of the multi-tasking VI-TasNet, we hypothesize that using a mask prior that explicitly depend on the relative energies in single source encodings at particular times and latent dimensions could be useful, too.
[0003] Figs.5 and 6 show different visualizations of model outputs. Fig.5 uses a mask in two parts in accordance with the inventive concept to separate two sources from an input mixture. This is compared to the true sources used to create the input mixture. Fig.6 uses a similar encoder-masker-decoder audio processing block to separate two sources (middle and right columns) from an input mixture (left column). Visualizations of model outputs: We visualize the outputs the model with a gamma approximate posterior in Table 3 (introduced later in the description). Fig.5 shows the example output of a network for approximately 50 ms of input audio. Similarly, Fig.6 shows a comparison over approximately four seconds of input mixture of the learnt representation (encodings) compared to a spectrogram.
[0004] Priors, rate-distortion and modified VI-TasNet ELBO: Priors in VAEs: A central part of the VI is the prior distributions used for the latent variables. The most common prior for a simple VAE latent encoding is an isotropic Gaussian distribution (i.e., a multivariate normal distribution with identity matrix covariance). Conceptually, it can be argued that this prior forces the VAE to learn latent dimensions with activations that generally tend to zero and that do not co-vary. The disentanglement is obtained by implicitly penalizing off-diagonal elements in the approximate posterior covariance. This prior is, in some scenarios, prohibitively restrictive. The benefits of more flexible priors are underlined by the improvements seen e.g. using a VampPrior or by learning the prior distributions using normalizing flows. Other choices than the isotropic Gaussian prior and Gaussian approximate posterior provide tools for enforcing other characteristics on the learned encoding. For instance, by using von Mises-Fisher distributions, a learned representations reside on a unit hyper-sphere, forcing latents describing directions without considerations of magnitudes. Similarly, a non-negative encoding (parts based, akin to non-negative matrix factorization) can be achieved using log- normal, gamma or Weibull distributions. Rate-distortion analysis: Optimizing a modified loss different from the ELBO facilitates adjustment of the trade-off between accurate generation (i.e., reconstruction) and deviation from the prior (i.e., more tightly constrained). For instance, for an isotropic Gaussian prior, a lower rate is related to more disentangled latents). Specifically, we can adjust the prioritization of the KL-term with a coefficient β (β-VAEs). For β < 1, the model is less restricted by the prior, freeing the model to produce better reconstructions, and for β > 1 the models are forced to learn representations more aligned with the prior. This trade-off can be thought of as a trade-off between a distortion D, the decoded negative log-likelihood, and a rate R, the KL divergence between encoding approximate posterior and prior, since −L = D + R. While higher capacity models can generally achieve better model evidences (i.e., higher ELBOs), it is only up to a limit of the complexity of the data. An unconstrained (β ≪ 1) model with sufficient capacity could reconstruct the inputs perfectly (D = 0) to the limit of the entropy of data by having a high R (i.e., the “auto- decoding limit”). Similarly, a tightly constrained model with sufficient capacity for encoding can map to something with R = 0 (i.e., the auto-encoding limit) but high D. There is a gap between models that are feasible (i.e., within auto-encoding and - decoding limits) and models that are realizable. For realizable models, there exists an optimal trade-off (in terms of lowest ELBO) between rate and distortion, but the relative capacity of the encoder and decoder alter the optimal trade-off. The trade- offs can be visualized using RD curves (i.e., phase diagrams in the RD-plane) by optimizing different trade-offs (e.g., using β). Work in exploring rate-regularization and the role of the prior (e.g., concerning generalization) shows how an isotropic Gaussian prior might not be the best inductive bias in general. The information bottleneck principle provides a framework for understanding representation learning and generalization. In particular—under a finite data sample—an optimal rate exists to minimize a generalization gap. The information bottleneck principle applies to deep learning-based variational inference, and an optimal rate exists to improve model test performance (generalization). Representation learning and compression: The mutual information between a learned latent representation and the observed data is lower bounded by the difference between the entropy of the data and distortion and upper bounded by the rate. The RD trade-offs made are comparable to trade-offs in data compression. From the perspective of lossy compression, the RD trade-off can be thought of as reducing the complexity of the latents (i.e., more compression) at the expense of poorer reconstructions (i.e., increased distortion)—or the other way around. Note that, for VAEs, this analogy is less directly related to dimensionality of the latents, and more so a matter of prior divergence. In fact, latent variable models can be turned into (lossless) compression models, and for these models, the rate term is indeed related to the achievable compression rate. Lastly, recent studies of RD analysis have shown how an inherent trade-off exists between not just rate and distortion, but also a quantity measuring divergence between the encoder-decoder induced distribution over the data and the true data distribution. This can be linked to e.g. the generative adversarial network objectives and naturalness—or potentially overall audio quality in the speaker separation setting. Mostly, VAEs learn representations that are a compression in terms of e.g. dimensionality of the learned representation. However, in some settings learning over-complete representations (i.e., higher dimensionality of representation than input signal), under constraints or regularized, can be a sensible approach to representation learning, and there exist general approaches to learning representations through over-complete AEs in relation to robustness of the representations, such as sparse, contractive, or denoising AEs. Free bits and reweighting: Attempting to optimize the ELBO without any modifications often yields models effectively stuck in a local minimum of low rate, without a driving force of distortion sufficient to overcome the loss incurred of moving away from the prior. This is often true for simple, standard VAEs, but is, in particular, the case with high-dimensional over-complete representations like the ones for the VI-TasNet. One solution to this is to e.g. introduce the rate term gradually, using KL annealing. Instead of annealing, we can opt for a“free information” approach to enable learning. When we measure the rate in bits / shannon, introducing free bits means that the model is always penalized some value of bits, providing a set “budget” (we measure the information in nats, here). This causes the model to effectively have the freedom to operate without being penalized within this budget. This should allow the model to trade off distortion and rate early on in training—and even go beyond the free budget. Since we have rate terms both for the encodings and the mask, we can investigate the effect of free bits, λ, for both individually, denoted by a subscript, and we will denote by rz,k,t′the rate contribution from encoding dimension k at latent time step t′, and rmi,k,t′the rate contribution from the k’th dimension in the mask for speaker i at latent time step t′. We also make use of the re-weighting introduced with β-VAEs, and use the same subscript notation to denote encoding or mask specific re-weightings. We can write the modified ELBO as, using both β (a multiplicative factor on the rate term) and λ (free bits), with slight notation abuse (letting the output of maximum(◦, ◦) be the largest of the arguments): Dynamic adaption towards target: We have also used a form of ELBO modification which uses adaption of the βz and βm terms based on two fixed, target rates for the sum over all K and T′ rates of encodings and sum over all N, K and T′ rates of the masks. We can combine the two, as well, and consider a total rate (sum of mask and encoding rates, with one, shared adapted β value). This is inspired by a target rate and automatic penalty weighting. It has been shown how an objective which directly optimizes towards a target rate for a VAE learns a better model in a synthetic experiment with a known ground truth generative process. We investigated using a known approach which can actively promote increased rates with the gradients (not just penalize), but for reasons of stability, we opted for using a version that incorporated an adaption. We note that, with ways of ensuring stability in losses that promote increased rates, we might see better performances. Dieleman et al. (2021) showed that adaptively increasing or decreasing a re- weighting of a regularizing term enables optimizing towards a desired target value for a quantity of interest in an auto-encoder setting (in their work this is not variational auto-encoders, and not a rate / KL-term). When using adaptive re-weighting, we can use the same formulation as in their Sec.3.1.3, Eq.7 (incorporated herein), for both adapting the βz and βm, towards target rates for Rz and RM, or adapting a target total rate (the sum of the two rates) and sharing the adapted β weight. Compared to the free-bits and the fixed β approach, this approach has the benefit that we can control the rates towards a very specific point on the RD-curve. Starting with a low initial value for the adapted weight also provides something similar to annealing in the beginning of training. We note that this dynamic must be adjusted to the training; if too quick or too slow adaption happens, the learning can be needlessly slowed down. If e.g. early stopping or learning rate annealing is used, a poorly specified adaption rate can cause the model to prematurely lower the learning rate or stop too early—even if the monitored objective is not the re-weighted ELBO but e.g. SI-SDR. Numerical experimental results for RD-curve in clean condition: We now provide further details on the RD-curve analysis in the main paper for the real (i.e. no synthetic) results on the clean condition of LibriMix and VCTK. Previously in Table 1, we opted to primarily discuss a model with Gaussian encodings and log-normal masks to align the most with how a “standard” TasNet operates. In the results ^^^^ presented here, and on Fig.3 (bottom), we used a with beta-distribution maskers, also shown in Table 3, with the BLR objective. This was chosen because this formulation of the model, initially, more readily put information in the encodings (over a Gaussian encoding formulation) when using the free-bits and fixed β setup which we used for these earlier experiments; that is, with fewer optimization steps, e^^^^^ th^^^ ^^^^^^^^^^^^-model saw higher rates than the Gaussians. The VI-TasNet in Table 1 was ultimately trained using the adaptive re-weighting scheme (instead of a fixed), which lessened this advantage of the gamma distribution formulation over the Gaussian, causing us to opt for a formulation more closely resembling the linear encoder outputs from a deterministic TasNet in the later experiments. Fig.7 shows rates versus distortion for different models, as well as how the models generalize from the LibriMix test set to the VCTK dataset. The results show a trade-off between generalization, rate and distortion. Parameters for the different models are given in Table 2 below. In this section, we present results where we modified the VI-TasNet ELBO with varying levels of free bits and with varying (but static for a given model) weights on both the masker rate and encoding rate. Fig.7 shows a visualization where the different models fall in the RD-plane, and how they generalize from the LibriMix test set to the VCTK dataset. The RD-plane plots show how the VI-TasNet display an expected trade-off between rate and distortion; generally, by increasing the rate, the distortion is reduced. The generalization from LibriMix to VCTK push the RD-curve trade-off up (poorer rate) and to the right (poorer reconstruction), for a generally overall poorer performance. Increases in rate alone is primarily driven by the encoding rates (the masking rates are largely unchanged between test sets). The increased rates (contributing to increased ELBOs) indicate that the differences in the datasets mostly affect the encoder and decoder, and less so the masker. The model with the lowest distortion on the LibriMix also has a higher rate, but this trade-off did not result in an improved distortion on VCTK, indicating a relation between the overall rate and the dataset generalization gap. Table 2: Rate distortions for VI-TasNets with gamma encoding and BLR- likelihood for various levels of free-bits, λ, and β-values (as introduced in the most recent equation, but since we only use β < 1, we give the reciprocal in the table). We show numerical results in Table 2 that are the basis for the visualizations in Fig.7 (which in turn is the full version of the bottom figure in Fig.3 previously discussed) and Fig.8. We note that the un-modified ELBO results in a model stuck at a low masking and encoding rate, with poor distortion / SI-SDRi. This model corresponds to a point in the far right-hand side and bottom of the RD plane, and for visualization purposes, it was left out on the RD curves. Note that all values are normalized by T in the plots and that the rates are normalized by the number of latent dimensions (K), too. The table gives the average mask rate (over N = 2 speakers as in all problems considered here), and total rate is Rz+ RM= Rz+ N · Rm. The last three rows in Table 2 provide a comparison between a model that has overall low encoding and masking β-values with one that penalizes each more (≈ 0.1 versus ≈ 0.01). Penalizing both relatively little (with βm = βz = 1 / 128) resulted in the highest encoding and mask rates shown in the table, whereas increasing the penalty on either lowered the rates for both. We show one run for each model configuration, as also discussed later, and so to further resolve the uncertainty in these estimates multiple runs would be needed. The improved performance (in terms of SI-SDR) from lower β-values is presumably largely attributable to the learning of an over-complete representation, but might additionally partially be due to a need to compensate for the differences in overall scale in the distortions and rates considered. The distortion for a very poor model is at approximately −2 nats, a poor model at approximately −2.8, and the best models at about −3.1, whereas the same models e.g. have encoding rates ranging from about 4 to 200 nats. We use a continuous output distribution; there is a difference in differential entropy / continuous entropy and the (discrete) entropy, and we note that e.g. a discrete output distribution would enable the use of theoretical results concerning the relationships between the entropy of the modelled data to the rate and distortion. We hypothesize that a discrete output distribution, e.g. a discretized (mixture of) logistic(s), might provide a suitable alternative to counter this rate-distortion scale difference, but the standard variants lack the scale-invariance and logarithmically scaled error measurement; for this, the Cauchy distribution (discussed later, too) provides an alternative solution worth considering, although it does not have scale-invariance. Fig.8 shows the same models as Fig.7, but instead plots rates versus negative SI-SDR. We provide a view of the RD curves where the negative SI-SDR replaces the BLR distortion on Fig.8. Later, we discuss how the BLR objective, while similar to the SI-SDR, reweights terms of the objective depending on the length of the signals considered. This, in part, causes the shift to the right on the distortion axis on Fig.7 in going from LibriMix to VCTK, since the overall average length of sentences in the two datasets is different. The overall conclusions regarding the shape and trade-offs are, however, still valid, e.g. supported by the Fig.8 with SI-SDR, which does not have this T-dependency.
[0005] Distributions: With N(x; μ, σ2) we denote a (univariate) Gaussian with scalar mean μ and scalar variance σ2. The gamma distribution density function we write as Gamma(x; α, β) = (βα / Γ(α)) xα−1exp (−βx), where Γ denotes the gamma function. Note that the α and β here has no connection to the β-VAE, nor the scaling in the SI-SDR. We refer to the α for the gamma distribution as the concentration, and β as the gamma rate. The encoder distributions are specified as: are time-series of length T′ with distribution parameters for the k’th latent dimension output. Similarly, is a concentration parameter time-series output. We also define a Gaussian prior, (and the equivalent log-normal version), and a gamma prior Masker distribution: The output of the masker parameterizes stochastic masks for all N speakers, where we opted for a variational approximation using either a beta distribution or a log-normal distribution: where κψ,0n,k,., κψ,1n,k,.are time-series of parameters output from the masking network, hψ, that correspond to the k’th latent dimension for the n’th source. The κψ,0models increasing the likelihood of the mask being closer to 0, the same but for a value of 1. The log-normal parameters are transformed versions of corresponding Gaussian parameters. We opt for a flat, uniform prior for the Beta-distributed masks, such that but note that the inductive bias of an e.g. Jeffrey’s prior towards either exclusion or inclusion of encoding elements is worthwhile investigating. For the log-normal masks, we use a standard lognormal distribution: LN(0, 1) prior. Convolutions of standard distributions: For Gaussians, we have that: for gamma distributions with one fixed
[0006] Scale-invariance: We consider time series of length T, s ∈ R1×T, and we approximate true single source speakers, s, with the approximation ˜s. In training TasNets, a standard approach is to optimize the SI-SDR (or minimize the negative SI-SDR). First, we provide a view of this as related to projection of estimates onto the true sources. Following this, we introduce a likelihood invariant to a scaling. Negative scale-invariant signal-distortion-ratio: The negative SI-SDR (NSISDR, for convenience) is (Here we let ∥◦∥ be the 2-norm and ^◦, ◦^ is the inner product operator, such that where aiis the i’th element of the column-vector a and Σiimplies a sum over all indices.): We can rewrite the NSISDR as: Furthermore, we can consider a version where we rescale s and ˜s to be unit vectors: and its estimate, the NSISDR is related to the squared projection of ˆe˜sonto ˆes. Bayesian linear regression likelihood: Alternatively, we can consider a loss using a Bayesian linear regression likelihood. We are interested in a likelihood which is invariant to a re-scaling of the whole time-series. The SI-SDR handles this using the α, and in the following we take the view that a similar factor γ is a regression coefficient, which we will marginalize out. In classic Bayesian linear regression, we are interested in modelling an unknown regression coefficient, γ, and an unknown noise-scale parameter, σ2. For the target time-series s with steps st, we consider a linear regression model, where the approximation ˜s takes the role of predictor (to avoid confusion with β-VAEs and the SI-SDR α we will denote the regression coefficient γ. The design matrix we consider (analogous to X ∈ Rn×k) is ˜s with k = 1 predictor variables. Note that for the scalars considered, we remove various transpositions, and e.g. replace det (Λ0) with simply the scalar value λ0.): st = ˜stγ + ^t, where ^t ∼ N(0, σ2). Here ˜st, β ∈ Rk×1=1×1are a scalar predictor and a regression coefficient, not vectors. The entire predictor time-series ˜s of length T corresponds to a design matrix of size T ×1. The corresponding likelihood is proportional to: . The least-squares estimate of the γ coefficient is given by: . Conjugate priors for σ2and γ take the form p(γ, σ2)= p(σ2)p(γ|σ2), where p( σ2) is an inverse-gamma distribution with parameters a0and b0, Inv-Gamma(a0, b0), and the conditional prior distribution for the regression coefficient is a normal distribution with a mean μ0, and variance σ2λ0−1, where λ is a scalar prior precision for the regression coefficient. Updates rules based on this take the form: . For this, the model, m, evidence is given by: and the log-likelihood, log p (s|m), becomes: Using this log-likelihood as the objective for the generative network (the decoder), we will refer to as using the Bayesian linear regression (BLR) likelihood. Up to a constant (depending on prior parameters and for a fixed length T), the negative of the above is equal to: .
[0007] Relationship between the BLR objective and SI-SDR: If we make some simplifications by assuming e.g. a weak prior, we can relate the BLR to SI-SDR. For λ0≪ ˜s⊤˜s andλ0≪ s⊤s, we have that the previous equation is approximately equal to (note that the λ0 plays much the same role as a constant added to ensure numerical stability of the log operation): Under the same assumption on λ0, we have that μT and λT are: And so, inserting this μ2TλTinto the expression for the loss, we have: . Up to a constant arising from the 1 / 2 factor within the log in the second term, this is equal to: . When T is large (e.g. as during training T = 3 s · 8 kHz = 24000), the second term dominates, which, in isolation, looks like: . W.r.t. the model parameters the first term is constant, ignoring this we have: and we see that the objectives are based on measuring the square of the inner product between the directions of the target and the estimate, albeit in slightly different manners. This difference stems partly from the difference in the view of re-scaling the target to fit the estimate, or the other way around (compare α with ˆγ). We note, especially, that the BLR objective varies with T in how the power of the estimated signal is taken into account. This dependency on T stems from the update rule for the atfor the inverse gamma distribution, and thus, in part, from an i.i.d. assumption on ^t, and it is worthwhile considering a model that addresses this differently.
[0008] Multivariate Cauchy objective: Besides using the BLR objective, we also consider a model which can model a per time-step scale (similar to the per time-step Gaussian with a modelled scale / variance), but which measures the error in a log- manner. We consider a Cauchy distribution, since the log-likelihood conceptually enables us to minimize a log-error similar to the log-MSE. A T dimensional Student-t with one degree of freedom (ν = 1) corresponds to a multivariate (T-dimensional) Cauchy (MVC) distribution, with a density function: where we model ˜s using a mean time series, μ, and scale matrix, Σ. The MVC model we consider will only parametrize a diagonal scale matrix. For an identity scale matrix, the log-likelihood of the MVC is proportional to the log of one plus the squared difference between ˜s and μ (i.e. similar to a log-MSE objective). We present and discuss some early results for using this likelihood later in the description.
[0009] Data, model and training specifics: General specifics on data foundation We use the open-source LibriMix, which combines speech from LibriSpeech (CC BY 4.0) and noise samples from WHAM! (CC BY-NC 4.0). A dataset with VCTK speech (CC BY 4.0) mixed with WHAM! Noise has been previously introduced, which we also use. Impacts of improved generative models: With increasingly powerful generative models on audio, as previously discussed, the problems of misuse of models, and misuse of available data to clone a voice without consent should be taken into consideration. While the presented VI-TasNet does not enable e.g. voice conversion, it is a generative model. The use of the VI-TasNet is solely focused on recreating, as closely as possible, the original speech of the single sources in the mixture. A key aspect is the level of temporal abstraction on which the generative model works; higher-capacity models operate on a considerably higher temporal abstraction in the latent variables than a VI-TasNet. Simplistically, a VI-TasNet learns something akin to an efficient version of the average speech spectrum with the addition of some knowledge of phase, whereas models with higher temporal abstraction will learn more high-level components of speech, like phonemes, words, sentences and speaker identity, allowing them to also reproduce or coherently alter such aspects of audio. We further refer to Dieleman et al. (2021, Sec.6.1, incorporated herein) for a discussion on such considerations concerning e.g. imitating voice identities in the datasets. Consent and identity: It should be noted that only the VCTK dataset was explicitly made with an aim related to voice synthesis. The LibriSpeech data is a curation of speech from the LibriVox project, which collects free public domain audiobooks. Volunteer audiobook recorders are instructed that audio enters the public domain, and examples are currently provided to inform the volunteers of various potential consequences. The VCTK speakers are given an anonymous numeric ID, which is available alongside their age, gender, accents, and region of England that they came from. The LibriVox speakers are identified by the name under which the reader is registered in LibriVox alongside the sex specified. The LibriSpeech dataset and the VCTK dataset, however, both contain audio recordings of speech, which can be used to identify a person. For the WHAM! noise dataset, it stated that the noise datasets “have been processed to remove any segments containing intelligible speech”. Written material foundation: The content for the VCTK dataset was curated from relatively recent newspaper articles (“3000 articles of the Scottish Herald newspapers”, and additionally “The Rainbow passage”, and “Accent elicitation passage from the Speech Accent Archive”); while the creation was based on coverage optimisation, we expect a very limited amount of explicitly offensive material, although the study does not mention filtering on a such a parameter. LibriMix, which uses LibriSpeech, is built using LibriVox. For LibriVox, the written content is old books in the public domain. The content of the LibriVox books, being older books, contain sentiments prevalent at the time of writing of the original works; this is an especially important consideration if learning e.g. a generative language model based on the data, but the VI-TasNet under consideration does not enable that higher level of representation learning of underlying language. It is important to note the limitations of training on a dataset of predominantly older material mainly from (non-conversational) English book reciting. The representation learnt of such audio does not reflect well the diversity of spoken English, and it is very unlikely that a model trained on English performs well on different languages altogether, and poorer representation fit for “non-standard” English and non-English causes reduced separation performance. Testing on VCTK enables testing of the algorithm on a dataset explicitly attempting to include diversity in British dialects, and further efforts in this direction could include evaluating on larger datasets with e.g. more nationalities (such as the VoxCeleb). Realistic mixtures: Speech separation algorithms like TasNet perform poorly on more sparsely overlapping data. The mixtures in LibriMix are both densely overlapping and are unrelated sentences from various audiobooks. Towards a better understanding of the speech separation performance, evaluation on more realistic, conversation-like dataset would be valuable; the sparsely overlapping version of LibriMix, SparseLibriMix test set, represents one such dataset. This present study focus on simple mixtures (“anechoically mixed”). Future investigations should address how e.g. reverberant environments affect the learning (using e.g. WHAMR! as considered for LibriMix in, using simulated room impulse responses, or even actual, reverberant / realistic mixtures). Various training specifics: We aligned with the Asteroid training recipes for the ConvTasNet. We used a smaller batch size of 4 (limited by the memory of available hardware) for the LibriMix experiments in the clean condition (that is, for results in Table 2 and Table 3). Another deviation from the Asteroid recipes is a reduced learning rate to 3 · 10−4. Training the variational network parameters with higher learning rates for that particular batch size resulted in training instabilities. For the results on the noisy condition (in Table 1), we used four GPUs in parallel with the same batch size of 4 with an effective batch size of 16. We could alleviate this instability for higher learning rates by increasing the batch size, but this, however, reduces the experimental throughput significantly. For the noisy condition VI-TasNet model, the total rate per time step (i.e. (Rz + RM) / T ) was adaptively optimized towards a value of 256 nats. This is a heuristically chosen hyper-parameter, and it is worth tuning and exploring, e.g. by resolving the RD-curves for the problem. Here, it was chosen through initial exploratory runs in trying to balance on the one hand being too restrictive while on the other hand ensuring that the value is imposing the needed regularization. Too low of a target rate would result in never reaching a performance comparable to the TasNet. Similarly, too high target rates would result in distributions collapsing onto deterministic distribution, using the freedom to remove all variance. For the SuDoRMRF results, we use only the separation modules of the SuDoRMRF implementation in Asteroid (or more specifically, their SuDORMRFImproved). We compare a deterministic version and a variational inference version (SuDoRMRF and VI-SuDoRMRF). The only changes for these models (compared to the TasNet and VI-TasNet models) were the networks called in the masker module (the separation module) that parametrizes the distribution over the masks (or directly outputs the masks in the deterministic setting). The SuDoRMRF was started with the package’s standard parameters matching the improved version configuration, and otherwise, the same setup was used as the VI- TasNet. To not include too many new factors, we did not, for instance, use the sum- to-one masking activation used in. We also did a similar experiment with the Sepformer using the SpeechBrain implementation of the Sepformer for the separation models. With the size of the Sepformer, we needed to reduce the batch size. We used the reduced learning rate and magnitude of gradient clipping reported in the original paper and used the same configuration as in the original paper as available on the SpeechBrain repository for the masking network, and otherwise the same parameters for all other VI-EMD- related parameters (i.e., the same as the TasNet / VI-TasNet configurations). The VI- Sepformer, notably, used the same target total rate of 256 nats, which caused the VI-Sepformer to more quickly be strictly more regularized than the VI-versions of TasNet or SuDoRMRF. Seeing the Sepformer is a considerably more expressive model (on parameter count it is more than 5–10x larger than the TasNet and SuDoRMRF) and operating on a different mechanism (convolutional versus attention), it is unsurprising that the model has significantly different rate-distortion trade-offs than the TasNet and SuDoRMRF. The VI-Sepformer with the same target total rate was heavily over-regularized (i.e., closer to the first rows in Table 2). We concluded that it would be beyond the scope of this paper to also contrast these additional trade-offs, especially considering the increasing training times of the larger models, but we consider it a promising future investigation to illuminate how generalization and RD-trade-offs are related to model architectures. Learning rate annealing: We trained with increased “patience” in the learning rate scheduler (50 epochs / passes over the full training dataset before reducing the learning rate at validation SI-SDR plateaus). This is, we believe, the cause of improved results on the deterministic TasNet reported compared to the Asteroid repository reported results (13.0 dB SI-SDRi versus the performance shown in Table 3 of 14.4 dB using Libri2Mix train-100 on the clean separation task). Similarly, Asteroid reports an SI-SDR improvement of 10.8 dB on the same noisy LibriMix (2 speakers, 8 Khz, min-mode, 100 hr), whereas the increased patience improved the deterministic baseline to a performance of 11.6 dB) in the noisy condition as reported in the main paper Table 1 (and the corresponding expanded table in Table 4). While we otherwise use the same architecture as provided in Asteroid (time-dilated convolutions architecture) masker, we saw some stability improvements by normalizing the summed skip connections by the depth of the masking network. Computational resources: A single experiment with the VI-TasNet doing 200 epochs over the 100 hour Libri2Mix training set with a batch size of 4 on a single GPU took approximately 5-6 days on an NVIDIA GeForce GTX 1080 Ti (or approximately 120-144 GPU hours). The adaptive model and flow models took slightly longer with extra encoding / decoding of single sources and mixture for the adaptive model and more model parameters in the learnt flow. The seven models in Table 3 alongside the eight models in Fig.7 (numerical results in Table 2) thus required approximately 2000 GPU hours total, disregarding an equivalent, at least, amount of GPU hours in developing the model before running the experiments. As previously mentioned, the reported results are all performances of one run of the model (one random seed). Accordingly, we have no basis for comparing the model performances rigorously, e.g. evaluating whether the flow prior is statistically performing better than the other variants of the VI-TasNet. While such a characterization is valuable information, we balanced the available time and compute against the value of resolving various priors. For the noisy condition results (in Table 1), the TasNet and VI-TasNet were trained on four GPUs for 1000 epochs (280 and 300 hours, respectively), or about 1.2k GPU hr per model. The TasNet, even with the increased patience, converged faster and could likely be stopped as early as 75–100 hours, whereas the VI-TasNet could likely have been stopped at about 200 hours. The best validations loss checkpoints (used in testing) were in both models, however, from the last 50 epochs of the 1000 epochs. We stress that the results are, under the compute available, single repetitions of a model training, and the results are limited in not addressing the uncertainty in final model performance given the stochastic initialization and optimization. Given the significant compute involved in training a single of these models, we chose to focus on the more challenging noisy condition for larger models (presented in Table 1), as we expected differences between deterministic and VI-based models to be bigger in this condition. Some experiments were done in the clean condition (presented in e.g. Table 2, Table 3, Fig.3, and Fig.4), and we chose to retain these findings and present them as initially done instead of repeating the experiments in the noisy condition. Parameter counts and model size: The VI-TasNet has a parameter count (and prediction time complexity) similar to the base TasNet counterpart we consider which has 5.1 million parameters. When using the gamma encoding, no extra parameters are added to the encoder. In the results for Table 3, using a Gaussian encoding or the MVC objective adds an extra dimension to the encoder or decoder output, respectively, which adds 8192 parameters to model the variances (i.e. a very negligible about 1–2‰ relative increase in parameters). The extra parameters in the masking network to enable two parameters per time and latent dimension increases the total model parameter count to 5.2 million trainable parameters, for the Gaussian encoding. The presented flow prior model, due to the large size of the latent space and relatively large size of chosen MADE configurations, has a (potentially needlessly large) total of 17.8 million trainable parameters. We chose to focus on a model closely aligned with a (deterministic, Conv-)TasNet to more readily compare to a well-studied model. Notably, the standard TasNet has a single-layer encoder and decoder—and can benefit from deeper structures. We saw in initial exploratory investigations that a VI-TasNet will similarly benefit from deeper encoders and decoders, possibly to a greater extent than a TasNet depending on how restrictive the utilized prior is. For a Gaussian prior and posterior, we are essentially forcing a linear mapping from windows of 16 samples of raw audio to closely resemble a Gaussian. Even slightly deeper, non-linear mappings might be beneficial. Approximate KL-divergences: In evaluating the rate terms, we have analytical expressions for the KL-divergence between two Gaussians, two betas or two gammas, but this is not the case for e.g. the flow prior. In training the models, we initially used an approximation of the KL-divergence, as is common practice in training VAEs, and this was used in the clean condition results (Table 2 and Table 3). With the KL-divergence as defined, a single sample Monto Carlo estimate corresponds to estimating the divergence using a single sample drawn from the approximate posteriors of the latent variables to determine the expectation of the log density ratios. While these approximations are generally close to the analytical expressions, we saw that the model learned better using the approximation than (when available) when the analytical expression. Investigating this, we saw that in some cases the estimated KL saturates when the parameters of the variational distribution are much lower than the prior value (e.g. trying to have very little activation in a gamma latent dimension), whereas the analytical expression does not. While an unintended consequence of using the estimated KL-divergence, this behaviour enabled the model to perform better and could indicate that a prior more flexible in allowing to “turn off” latents is useful for the VI-TasNet specification considered. In the synthetic and noisy condition experiments, we used the analytical expressions, facilitated by the use of using adaptive re-weighting, instead of the free bits and static re-weighting used in the clean condition results. While the noisy condition results use an adaptive re-weighting, the training also included a (potentially inconsequential) free encoding nat (1.0 nat) across all encoding dimensions as well as one free mask nat across all mask dimensions and speaker masks. Numerical stability: dequantization, initialization, clamping, and margin loss: Since the audio signal is a 16-bit audio discrete-time signal, we used a (potentially inconsequential) uniform dequantization. Traditionally, float representations of audio are scaled to be in the [−1, 1] range. Using a standard PyTorch initialization for the encoder and decoder resulted in initial estimates that were much outside this range, so we used a uniform initialization on [−10−2; 10−2] for the encoder and decoder weights. While values outside the [−1, 1] range is not a problem for the scale- invariance objective (it simply re-scales the signals), it is inconvenient to have the model operate in this range needlessly. We employ a margin loss entirely similar to the one used in Dieleman et al. (2021, incorporated herein). We saw little to no effect from dequantization, and its use was dropped for the noisy condition and synthetic experiments (these experiments are later than the clean condition results). The synthetic experiments also do not use a margin loss. The parameterizations of the various distributions rely on transforming the direct output of the networks in some manner; for the gamma concentration parameter and Gaussian scale, we e.g. pass the network output through a softplus function, and clamp it to a minimum value of 10−6to have the required strictly positive parameter. The adaptive KL re-weighting was in the noisy condition started at a factor of 10−9and adapted by a factor of 10−4at each step (otherwise using the same formulation and e.g. minimum threshold for change as Dieleman et al. (2021, incorporated herein)). For the synthetic experiments, a higher initial value and adaption rate of 10−6and 10−2, respectively, were used, to accommodate the faster training of simpler models. The data was standardized based on the mean and variance of the waveforms across time and across all mixtures in the dataset, and the same standardization was used in testing on both LibriMix and VCTK (i.e. LibriMix train set values of mean and variance were used for standardization). Autoregressive flow prior: The learnable autoregressive flow (AF) prior introduced with the variational lossy autoencoder (VLAE) is equivalent to the inverse- autoregressive flow (IAF). We investigate how this type of more flexible AF prior can be used in the VI-TasNet, by learning a mapping, or flow, from a base distribution (“noise source”), u(^), to the latent encodings, pξ(z). We use a Gaussian base distribution for ^, and learn a series of flows. These flows together make up a mapping z = ωξ(^), and they use a series of invertible mappings parametrized by autoregressive networks with parameters ξ. While one such possible mapping is a series of affine transformations, where a scaling and a translation are learnt, we adopt the approach from the VLAE to use a mean-only flow. Similarly, we also make use of a series of masked autoencoder density estimations (MADE) networks as the autoregressive networks parameterizing the flows. Sampling frequency: We chose to work with the 8 kHz version of the datasets to reduce the computational requirements. Investigations on whether the findings presented hold for increased sampling frequencies would be important for real-world use since many use cases require a higher audio quality than achievable with 8 kHz. Dataset size: Similar to the choice of the 8 kHz variant of the data, we worked with the 100 hr (smaller) version of LibriMix to reduce computational requirements. Using the larger versions of the datasets, or more datasets, tends to increase performance; for instance, Asteroid reports 10.8 dB SI-SDRi in the noisy, 2-speaker LibriMix condition when trained on 100 hrs, but a performance of 12.0 dB SI-SDRi when trained on the larger 360 hrs variant of LibriMix. Within the dataset (i.e. on LibriMix), this means that the TasNet trained on 360 hrs of data matches the performance of a VI-TasNet trained on only 100 hrs of data (this, of course, does not say anything about the generalization to new domains or conditions of VI-TasNet versus TasNets on larger datasets). For scenarios where examples are scarce, we would propose the investigation of VI-TasNets—and especially the multitasking version for learning from more abundant audio without available single sources. Characterizing the performance as a function of dataset sizes / number of examples with single sources (exploring learning curves) was not the focus of the present study; in such scenarios, we stress that the baseline would not solely be a deterministic TasNet, but rather a model trained with methods such as MixIt and methods with similar aims.
[0010] Synthetic experiment: Fig.9 shows an example of a synthetic Gaussian pulses dataset. Top row: spectrogram of single sources targets in isolation. Bottom row: input mixture to the model. These are used to test the encoder-masker-decoder audio processing block. Gaussian pulse mixtures: We construct a simple source separation problem which makes mixtures of single sources that themselves are overlapping sinusoidal Gaussian pulses from a specific frequency region with overtones. We specify a frequency range for the “fundamental frequency of given speaker”. In the experiments shown, this was set to 300-400 Hz for “Speaker / Target 0” and 100-200 Hz for “Speaker / Target 1”. For an example set of targets and input, see Fig.9. To create one of the synthetic single sources, we create three sinusoidal pulses with a Gaussian envelope and add them to create one single source in the mixture. For each pulse, we sample a fundamental frequency in the given “speaker’s” range. We also sample an overall amplitude in a given range (this range is shared between targets), a phase, and pulse delay within a specified time axis of 3 seconds so at least half of the pulse is within the segment. Additionally, we add four overtones to the sampled fundamental sine wave and add Gaussian noise. We create distinct examples / datasets by keying the randomness / sampling procedures to integer base seed / ranges; for the replicates in the synthetic experiments, the first replicate had indices ranging from [100000, 102048[ for 2048 training examples, the next 128 indices were validation examples, and after that came 1024 indices for each test data configuration (e.g. different noise levels). Similarly, the next replicate had training examples from [200000, 202048[, and so on. While the frequency region of the fundamental frequencies is not overlapping, the inclusion of the overtones results in a problem where regions of the spectrogram will share energy between the two targets. Model and optimization: We reduce the number of encoding dimensions with respect to the models we consider on LibriMix from 512 to 32. As available in the Asteroid framework, we similarly reduce the masker bottleneck channels down from 128 to 16, the skip channels from 128 to 8, the hidden channels from 512 to 16. We retain the same number of blocks per repeat but reduce the number of repeats / cycles to 1. We parametrize Gaussian encodings and lognormal masks and we use standard priors for both (mean / location and standard deviation / scale of 0.0 and 1.0, respectively). We use a BLR likelihood, and we re-weigh the rate terms to match varying levels of rates to resolve the RD-curve using the adaptive re-weighting previously described. We use an initial value for the adaptive factor for the total rate (sum of both rate and encodings) of 10−6and adapt with 1 % at each step if needed (δ = 0.01). We use a learning rate of 3 · 10−4(with the same higher patience scheduling as described for larger models) and optimize the model using a permutation invariance evidence lower bound loss (i.e. while using PIT, we minimize the negative ELBO). The models are trained for a maximum of 800 epochs (parses over the 2048 data examples), with potential early stopping if no improvement in validation SI-SDR is seen for 60 consecutive epochs. We do not use the model with the highest validation SI-SDR for the evaluation but instead use the last model to make the RD curves, since the adaptation can produce early models that had high rates with better performance than later, more tightly regularized versions. Monitoring the actual loss (modified ELBO) instead of the SI-SDR is non-trivial, because the adaptive reweighing continuously changes the values, e.g. potentially increasing the loss in periods where a rate is being re-weighted towards lower rate values without it necessarily indicating a plateau and a needed “early stop”. We expected to find the models could over-fit, which would be evident as a (significant, especially for higher rates) gap between training and validation / test performances. However, with the specifications detailed above, the models did not display significant over-fitting to the training data, even for the highest rates. This, we hypothesize, is a consequence of simultaneously (i) having reduced the complexity of the model (fewer latents, fewer filters, etc.) and (ii) having employed regularizing elements (such as early stopping and learning rate annealing). Provided that e.g. a more over-complete model was trained without learning rate annealing, we hypothesize that the generalization behaviour and rate-generalization trade-offs would be even more pronounced. Different priors on clean LibriMix: In this section, we present earlier results for TasNets and VI-TasNets trained on the clean version of LibriMix (min-mode, 2 speakers, 8 kHz, clean, 100 hrs). As the deterministic baseline, we trained a TasNet using both the standard SI-SDR objective and the BLR. We compare these TasNets to four variations of the VI-TasNet to investigate the effect of encoding posterior and priors, all using the BLR objective, and all using a beta masker. Firstly, we train a VI- TasNet with a Gaussian encoder distribution and prior, and a similar gamma version. In addition to these, we train a model that uses a Gaussian approximate posterior with a flow prior. We train a VI-TasNet that uses a gamma posterior in conjunction with adaptive prior, which also uses the multitasking objective. Finally, we train a VI- TasNet with a more flexible decoder distribution, the MVC, using a gamma encoder and prior. The performance of these models is shown in Table 3. Training a deterministic TasNet with the BLR objective (unsurprisingly) reduced the SI-SDR performance compared to directly optimizing SI-SDR, but the BLR does produce reasonable SI-SDRi scores while additionally providing actual probabilities / normalized densities. The VI-TasNets learn to perform the separation task well, albeit not—in this earlier version, without e.g. adaptive re-weighting—to the same performance as the deterministic counterparts. These models are different from the models outperforming the TasNets in Table 1 in the amount of (target total) rate they achieve and the masker distribution used. Here, the best performing VI- TasNet uses the flow-based prior, but both the simple Gauss and gamma models attain performances near the 13 dB SI-SDRi mark. The multitasking, adaptive model and the MVC model perform the poorest of the VI-TasNet. These models were not converged within a 200 epoch limit (about 150 GPU hours of training), and they would likely see improved performances with longer training. While the TasNets here display the highest difference between the LibriMix and VCTK test sets (generalization gap), this result does not support this being attributed to differences in variational versus deterministic models, seeing as the TasNets also display higher overall SI-SDRi. For the models in Table 3, the differences in model performance are smaller in the VCTK-2mix test, where e.g. the difference between the BLR TasNet and the flow-based VI-TasNet is 0.38 dB, as opposed to their difference of 0.73 on LibriMix. We note, in contrast to the findings in this section, that other configurations of the VI-TasNet models (such as the ones presented in the main paper in Table 1) produce both better overall performance of the VI-TasNet and notably also better generalization to VCTK. Table 3: Model SI-SDR improvements in dB for LibriMix and VCTK test sets in noise-free / clean condition. qNφ / qΓφ: Gaussian / gamma approximate posterior, respectively; similar notation for Gaussian and gamma prior; pξ: flow prior; pφ,S: adaptive prior. SI-SDR is parenthesized to denote it as an objective rather than a likelihood. In the main paper, we report results for the model with Gaussian encodings and log-normal masks (to align with the most standard TasNet formulation), but we here show the viability of considering other distributions. With these results, we show that the VI-TasNets support the incorporation of different types of structure in the latent encodings and different observation models (decoder distributions, likelihoods) and that the choice of these affects performance. The gamma formulation can produce a non-negative encoding, which could potentially draw strengths from a parts-based representation, and similarly, we show that the BLR and MVC are viable avenues of exploration for imposing certain characteristics (like scale-invariance) in a manner compatible with variational inference. The flow model is expensive during training, but we do not need to evaluate the prior during a call to the model if we are using it to do separation after training is completed. In this case, training time complexity might potentially be traded off for increased performance We would, however, need it if we wanted to do uncertainty estimation. The variational inference formulation enables future work to incorporate various other well-known approaches from probabilistic modelling. Some examples are: we could address problems with an unknown number of speakers using a stick- breaking / Dirichlet process for the masks distribution; we could address the permutation problem by modelling target and estimated speakers with a mixture; or, we could provide a stronger learning signal to the masks by incorporating knowledge of the single source encodings as masking targets through an adaptive mask prior. The adaptive model in Table 3 (second to last row) is, importantly, an example of a functioning multitasking model. We showed how such a model, and its input density estimates, can be useful from an uncertainty quantification perspective with Fig.4. It is worthwhile stressing, however, that such a multitasking model can also learn directly from mixtures without reference single target sources in isolation. The results shown here do not investigate the effect of the masker distribution choice, but we note that these results consider a beta masker (more similar to a sigmoidally gated TasNet mask), whereas e.g. Table 1 consider log-normal masks (more similar to a ReLU gated / rectified TasNet mask), showing how both are viable possibilities even if they have very different ways of masking.
[0011] Further synthetic rate-distortion results: In this section, we provide further results on the RD-curve analysis of the synthetic problem and models considered in Fig.3 and Fig.9. We train VI-TasNets with the adaptive rate-regularization towards a range of rates on the 2-speaker synthetic problem. The target rates log-spaced from 10−3to 103. Fig.10 shows RD curves for various test conditions performance of encoder- masker-decoder audio processing block models with varying target rates trained in a particular version of the synthetic Gaussian pulses separation test as discussed in relation to Fig.9. In Fig.10, we show how a model trained on a domain characterized by a particular level of noise amplitude, “NA” of 0.5—this corresponds to a standard deviation of Gaussian noise added on top of the single sources after additive mixing. The test performance on this seen / familiar-condition data is shared across all plots in black. A line is drawn as the running average distortion as a function of (sorted) rates, and a dot is drawn for each model. Alongside the black line that is the same in each plot, we show a corresponding line which shows the performance of the model evaluated on different test conditions. Firstly, in going from top to bottom and from left to right, we start with the training domain and then ranging from models with no noise to increasing amounts of noise, all the way to a noise amplitude of 1.0 (a factor 2 above the training domain of 0.5). We see that models with too high rates generalize less well to low noise settings, but the same is not immediately the case for higher noise settings. Following this, we see how changing the number of over-lapping sines (“OS”)—by either adding or removing one from the training domain amount of 3 sines— produces either slightly better separation or slightly poorer separation, but no clear differences for high rates versus other models in their generalization abilities. A similar conclusion holds for the number of overtones (“OT”); removing them altogether makes it easier, whereas adding 1 (from five to six) or doubling (five to ten) makes the problem harder. We see that a slow (2 Hz) amplitude modulation of the noise signal (“AMN”) or band-pass filtering of the noise (“BPFN”, to have the noise only be in the region of the speaker Gaussian pulse frequencies) both produce slightly easier problems. We can control the width of the Gaussian pulses with a scale (“GS”). Evaluating with pulses that can be slightly longer and shorter ([0.01, 0.5] versus [0.05, 0.3] in the training domain), does not significantly change the performance in expectation. Testing on shorter pulses ([0.01, 0.05]), however, is a harder task, and only longer ([0.3, 0.5]) is an easier one. Lastly, we can change the frequency ranges (“FR”) that define the speakers, to be either slightly expanded (from 100–200 Hz and 300–400 Hz to 75–225 Hz / 275–425 Hz), nearly overlapping, or actually overlapping. In each case, this produces models with poorer distortion, but no clear difference in optimal versus higher rate models in the generalization abilities. While the target total rates in some cases were as high as 1000, no model achieve rates over 150 nats. We hypothesize that the limited capacity coupled with e.g. learning rate annealing, early stopping and stochastic optimization might produce models that do not over-fit to the same extent and thus do not produce very high rates, even if the model has the freedom to do it. Generally, we have found that, when the models can achieve the target rate they do so with a higher consistency across replicates, whereas the models that cannot achieve the set rate tend to display a larger variance in the final expected rate over the test set and a similarly large variance in achieved distortion (some might do well, others do very poorly).
[0012] All metrics LibriMix / VCTK evaluation: In Table 4 we provide extra numerical results for the Table 1. The full table shows (in addition to the already reported SI- SDRi), the BLR objectives and the SI-SDR. The table also highlights what is a familiar (“intra-dataset” or “intra-condition”) evaluation versus an unfamiliar one (“inter-”). These extra metrics show how the VI-TasNet and VI-SuDoRMRF is an improvement both in SI-SDR, SI-SDRi and BLR over the deterministic counterpart on both familiar and unfamiliar datasets and conditions. The only exception is that the BLR better for the deterministic model in the LibriMix conditions. Table 4: TasNet and VI-TasNet on noisy separation task. Full version of Table 1 with SI-SDR and BLR metrics, and indication of whether the performance is an inter- and intra- dataset or condition evaluation.
[0013] Enumerated Example Embodiments EEE1. A hearing aid system comprising an encoder-masker-decoder audio processing block, adapted to provide source separation, wherein said audio processing block has been learnt using variational inference wherein said mask is applied to the encoder output in order to provide the source separation after decoding. EEE2. The hearing aid system according to EEE1, wherein said source separation is speaker separation or noise suppression. EEE3. The hearing aid system according to EEE1 or EEE2 further comprising a variational evidence lower bound. EEE4. The hearing aid system according to EEE3 wherein the variational evidence lower bound ^^^^^^^^is given as: ^^^^^^^^= − ^^^^^^^^− ^^^^^^^^− ^^^^^^^^, wherein ^^^^^^^^is a distortion of single sources, ^^^^^^^^is a divergence of the mixing encodings from their prior, and ^^^^^^^^is a divergence of masks from their prior. EEE5. The hearing aid system according to any preceding EEE, wherein priors used by the encoder depend on the sources to be separated. EEE6. The hearing aid system according to any preceding EEE, wherein said audio processing block is adapted to perform multitasking by combining single source autoencoding and source mixture autoencoding. EEE7. The hearing aid system according to EEE6, wherein the multitasking provides uncertainty quantification by using input mixture density modeling to estimate performance without knowledge of targets. EEE8. The hearing aid system according to any preceding EEE, wherein the encoder-masker-decoder audio processing block is a sepformer or is based on LSTM blocks. EEE9. The hearing aid system according to any preceding EEE, wherein the encoder-masker-decoder audio processing block is trained based on a probabilistic version of the scale invariant signal distortion. EEE10. The hearing aid system according to EEE9, wherein the probabilistic version of the scale invariant signal distortion is the scale invariant Bayesian linear regression observation model.
Claims
CLAIMS 1. A hearing aid system comprising an encoder-masker-decoder audio processing block, adapted to provide source separation, wherein said audio processing block comprises at least one of a stochastic encoder and a stochastic mask that has been taught using variational inference, wherein said mask is applied to the encoder output in order to provide the source separation after decoding.
2. The hearing aid system according to claim 1, wherein said source separation is speaker separation or noise suppression.
3. The hearing aid system according to claim 1 or 2, wherein the encoder is a variational auto-encoder, VAE, configured to optimize an evidence lower bound, ELBO, being a standard VAE ELBO with an added Kullback-Leibler divergence term.
4. The hearing aid system according to claim 3, wherein the ELBO used is: − ^^^^^^^^− ^^^^^^^^− ^^^^^^^^, where ^^^^^^^^is a distortion of single sources, ^^^^^^^^is a divergence of the mixture encoding from their prior, and ^^^^^^^^is a divergence of masks from their prior.
5. The hearing aid system according to any one of the preceding claims, wherein priors used by the encoder depends on the source to be separated.
6. The hearing aid system according to any one of the preceding claims, wherein the encoder-masker-decoder audio processing block is adapted to perform multitasking by combining single source autoencoding and source mixture autoencoding.
7. The hearing aid system according to claim 6, wherein the multitasking provides uncertainty quantification by using input mixture density modelling to estimate performance without knowledge of targets.
8. The hearing aid system according to any one of the preceding claims, wherein the encoder-masker-decoder audio processing block is a sepformer or is based on LSTM blocks.
9. The hearing aid system according to any one of the preceding claims, wherein the encoder-masker-decoder audio processing block is trained based on a probabilistic version of the scale invariant signal distortion.
10. The hearing aid system according to claim 9, wherein the probabilistic version of the scale invariant signal distortion is the scale invariant Bayesian linear regression observation model.