Frame-level permutation-invariant training for source separation
By combining frame-level permutation-invariant training (tPIT) with cluster-level training, the permutation ambiguity problem in deep learning models is solved, achieving high-quality sound source separation and speaker coherence, and improving the accuracy and efficiency of speech source separation.
Patent Information
- Application Number
- CN202180070431.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-13
- Filing Date
- 2021-10-13
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2041-10-13
AI Technical Summary
Existing deep learning models suffer from permutation ambiguity in speech source separation, which makes it impossible to clearly attribute extracted speech frames to specific speakers. Furthermore, while frame-level permutation invariant training (tPIT) improves separation quality, it fails to maintain speaker coherence.
By employing frame-level permutation-invariant training (tPIT) combined with clustering, the accuracy of sound source separation in each frame is ensured by minimizing the difference function and maximizing the separation criterion. Furthermore, permutation ambiguity is resolved at the clustering level, achieving end-to-end sound source separation.
It achieves higher quality sound source separation, ensures speaker coherence, and improves the accuracy and efficiency of the separation model.
Smart Images

Figure CN116348953B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to the following priority applications: European Patent Application No. 21151297.5, filed January 13, 2021; U.S. Provisional Application No. 63 / 126,085, filed December 16, 2020; and Spanish Patent Application No. P202031039, filed October 15, 2020, each of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to the field of audio processing. Specifically, this disclosure relates to techniques for source separation (e.g., speaker separation) using deep learning models or systems, and frameworks for training deep learning models or systems for source separation. Background Technology
[0004] In the following text, the terms speaker, speaker, or speech source separation will be used as examples of sound source separation. It should be understood that this disclosure should not be construed as limited to speaker, speaker, or speech source separation, but rather generally relates to any kind of sound source separation.
[0005] Speech source separation can be performed within a deep learning (DL) framework. One of the main challenges of this framework is the permutation ambiguity problem, which can prevent the explicit attribution of extracted speech frames to a single speaker. This problem can be addressed through utterance-level permutation-invariant training (uPIT) of a deep learning-based system performing speech source separation. However, uPIT is inferior to frame-level permutation-invariant training (tPIT) in terms of the quality of speech source separation. Therefore, there is often a trade-off between the quality of source separation and speaker coherence along the frame sequence.
[0006] Therefore, there is a need for a method to train a deep learning-based system for sound source separation (e.g., speech source separation) that achieves improved sound source separation quality while still allowing the extracted sound source signals to be explicitly attributed to the actual sound source. Summary of the Invention
[0007] In view of this, the present disclosure provides a method for training a deep learning-based system for sound source separation, a method for performing sound source separation using a deep learning-based system, and corresponding apparatus, computer program, and computer-readable storage medium.
[0008] According to one aspect of this disclosure, a method is provided for training a deep learning-based system for sound source separation. Training may refer to determining the parameters of a deep learning model (e.g., a neural network) used to implement the system. Furthermore, training may refer to iterative training. The system may include a separation level for extracting sound source representations frame-by-frame from a representation of an audio signal. The audio signal may be a mixed signal comprising multiple sound sources (e.g., a speaker). The system may also include a clustering level for generating a vector for each frame indicating an assignment permutation (or permutation assignment) to the corresponding sound source (e.g., a label or other identifier of the sound source). The assignment permutation may indicate a one-to-one assignment of the extracted frame to a sound source (label). The vector generated by the clustering level may correspond to an embedding vector. Thus, the clustering level can map frames of the mixed audio signal to a lower-dimensional space, the so-called embedding space, depending on the assignment permutation chosen for each frame. The clustering at the clustering level can be used to resolve ambiguities in the sound source permutation. The representation of the audio signal and the representation of the sound sources may be waveform-based representations. The method may include obtaining a representation of a mixed audio signal and representations of at least two reference audio signals as input. The mixed audio signal and the reference audio signals may be waveform audio signals. The reference audio signals may be or relate to ground truth signals. The representation may be a waveform-based representation. The mixed audio signal may include at least two sound sources. The reference audio signals may correspond to individual sound sources included in the mixed audio signal. The method may further include inputting the representation of the mixed audio signal and the representations of the at least two reference audio signals into a separation stage, and training the separation stage to extract representations of sound sources from the representation of the mixed audio signal in such a way that, for each frame, a difference function is minimized. The difference function may be based on the difference between frames of the extracted sound source representations and frames of the reference audio signal representations. For each frame, the assignment permutations of the extracted sound source representations and the reference audio signal representations are selected to minimize the minimum difference function. Therefore, the training of the separation stage may be frame-level permutation-invariant training (tPIT). The method may further include inputting a representation of the mixed audio signal, and for each frame of the mixed audio signal representation, inputting the frame of the extracted source representation along with an indication of the assigned permutations selected for the corresponding frame of the mixed audio signal representation to a clustering level, and training the clustering level to generate vectors that indicate how the frames of the extracted source representations are assigned permutations to the corresponding sources in such a way that the separation between groups of vectors of the mixed audio signal frames is maximized. The frame vectors can be grouped according to the corresponding assigned permutations indicated by these vectors. The difference function and separation criterion may be implementations of loss functions used in the training of the separation level and clustering level, respectively.
[0009] With the configuration described above, the proposed method can apply tPIT to waveform-based audio signals. Using tPIT instead of uPIT allows for higher quality source separation. Furthermore, in addition to the separation level, source coherence across the frame sequence can be ensured by providing a clustering level, which is typically a problem with tPIT frames. Specifically, reordering the extracted frames according to the embedding vectors generated at separation time by the clustering level produces output streams, each containing only a single source, with no source exchange between output streams.
[0010] In some embodiments, the difference function may indicate a combined difference (e.g., a combination of differences) between frames of the extracted source representation and frames of the reference audio signal representation. For each extracted source representation, the combined difference may include the difference between frames of the extracted source representation and the corresponding frames of the reference audio signal representation.
[0011] In some embodiments, the cluster level can be trained such that the separation criterion is optimized for each frame of the representation of the mixed audio signal. The separation criterion can be based on the Euclidean distance between vectors and / or groups of vectors. Specifically, for a given frame of the representation of the mixed audio signal, the separation criterion can be based on the Euclidean distance between the vector indicating the assignment permutation of the frame and the vector groups of other frames in the representation of the mixed audio signal. The distance between a vector and a group of vectors (clustering) (e.g., Euclidean distance) can be defined as the distance between the vector and the centroid (mean) of that group.
[0012] The separation criterion derived using Euclidean distance allows for cluster-level training to be performed in a memory-efficient manner, and further enables training when using waveform-based representations of mixed audio signals with high temporal resolution.
[0013] In some embodiments, the optimization separation criterion may correspond to maximizing the following expression for a given frame of the representation of the mixed audio signal:
[0014]
[0015] Here e i It is a vector of a given frame, c k Let be the centroid of the vector group formed by the k-th distributional permutation, P be the total number of distributional permutations, and d(-,-) be the squared Euclidean distance. The centroid of the vector group can be the mean of the group. The squared Euclidean distance can be calculated from... Given, where w is an optional scaling parameter and b is an optional offset. For example, w and b can be learnable parameters.
[0016] In some embodiments, the system may further include a transformation stage for transforming the mixed audio signals into a representation of the mixed audio signals. The mixed audio signals may be transformed to a signal space, which is a waveform-based signal space. In other words, the signal space may not be a frequency domain signal space. Specifically, the signal space may not be based on a spectrogram. The signal space may be referred to as a generalized signal space. If the system includes a transformation stage, it may also include an inverse transformation stage that transforms the extracted source representation back to the waveform domain, for example, into a waveform audio signal of the extracted audio source.
[0017] Therefore, the current embodiments of this disclosure allow for deep learning-based sound source separation in an end-to-end manner (i.e., waveform input / waveform output). Advantageously, waveform-based representations can include feature space representations tailored for general sound source separation applications, particularly for speaker separation. Furthermore, the encoder and / or decoder (i.e., transform stages and / or inverse transform stages) can be deep learning-based and can be trained to meet specific requirements. This is typically not the case for spectrogram-based representations.
[0018] In some embodiments, the system may further include a transform stage for transforming the mixed audio signal into a representation of the mixed audio signal. Specifically, the transform may involve one of the following: segmenting the mixed audio signal into multiple frames (e.g., short frames) in the time domain; projecting the mixed audio signal onto a deep learning-based encoding in a latent feature space optimized for source separation; Mel space encoding; and deep learning-based problem-agnostic speech encoding. If applicable, the transform stage (and / or inverse transform stage) may be jointly trained with the separation stage and / or clustering stage. As noted, the system may also include an inverse transform stage that transforms the extracted source representation back into the waveform domain. In the case of Mel encoding, the inverse transform stage may include or correspond to a MelGAN decoder.
[0019] In some embodiments, representing the mixed audio signal may involve segmenting the mixed audio signal into waveform frames (e.g., short waveform frames). A separation level can then be trained to determine, for each frame of the mixed audio signal, the frame from which the extracted sound source is located from the frames of the mixed audio signal in a manner that minimizes the following loss function:
[0020]
[0021] Where t indicates the frame, l indicates the number of samples within the frame, L is the total number of samples within the frame, n is the label of the extracted sound source, N indicates the total number of extracted sound sources, est represents the frame of the extracted sound source, and ref represents the frame of the reference audio signal. It is a permutation map for labels n=1, ..., N, indicating the label of the reference audio signal. For each frame, the permutation map... The option chosen to produce the minimum loss function is the one from which the loss function can be minimized. For each possible permutation of labels n = 1, ..., N, there may be a corresponding permutation map. Since for N labels, the number of possible permutations is N!, there can be N! permutation maps, meaning that index k can run from 1 to N!.
[0022] For example, the total number of samples within frame L is in the range of 2-16 samples. At a typical sampling rate of 8000 Hz, a frame with 2-16 samples corresponds to a duration of 0.25-2 milliseconds. However, this disclosure is not limited to a sampling rate of 8000 Hz, and other sampling rates, such as 44.1 kHz or 48 kHz, can be used. For such sampling rates, L can also be chosen to correspond to 2-16 samples. Alternatively, L can be chosen so that the frame corresponds to a duration of 0.25-2 milliseconds, for example, 11-88 samples per frame for a sampling rate of 44.1 kHz; and 12-96 samples per frame for a sampling rate of 48 kHz.
[0023] Segmenting into (short) time frames corresponds to a particularly simple waveform-based representation. Although this representation may result in relatively high temporal resolution, the proposed method enables the training of clustering levels in a memory-efficient manner for deep learning-based systems.
[0024] In some embodiments, the representation of the mixed audio signal may involve a latent feature space representation of the mixed audio signal that can be generated by a pre-trained deep learning-based encoder. Then, a separation level can be trained to determine the extracted sound source frame from the frames of the mixed audio signal for each frame in a manner that minimizes the following loss function:
[0025] .
[0026] Where t indicates the frame, f indicates the features in the latent feature space, n indicates the labels of the extracted sound sources and the reference audio signal, and N is the total number of extracted sound sources. V indicates the frame representing the extracted sound source, and V indicates the frame representing the reference audio signal. It is a permutation map for labels n=1, ..., N, indicating the label of the reference audio signal. The representation of the reference audio signal can be determined based on the representation of the mixed audio signal and a set of masks, where the set of masks has been determined based on the reference audio during the pre-training of a deep learning-based encoder. For each frame, the permutation map... It is selected to produce the minimum loss function, and the resulting loss function can be minimized.
[0027] Waveform-based latent feature space representations can be trained specifically for sound source separation (e.g., speaker separation), thus enabling accurate and efficient sound source separation. The proposed method allows for the training of deep learning-based systems in such frameworks in a memory-efficient manner.
[0028] In some embodiments, the method may further include a pre-trained deep learning-based encoder. The deep learning-based encoder can be adapted to generate a latent feature space representation of the input audio signal. Pre-training the deep learning-based encoder may involve inputting a mixed audio signal and a reference audio signal into the deep learning-based encoder, which generates latent feature space representations of the mixed audio signal and the reference audio signal. For each reference audio signal, a mask of the reference audio signal is generated based on the latent feature space representation of the reference audio signal using a softmax function. A representation of the extracted audio signal is determined by applying the generated mask to the latent feature space representation of the mixed audio signal. A deep learning-based decoder is applied to the representation of the extracted audio signal to obtain an estimate of the extracted audio signal, wherein the deep learning-based decoder is adapted to perform the inverse operation of the deep learning-based encoder, and the deep learning-based encoder is trained such that the signal-to-distortion ratio of the decoded estimate of the extracted audio signal to the corresponding reference audio signal is minimized.
[0029] In some embodiments, the representation of the mixed audio signal may involve a Mel feature space representation that can be generated from the mixed audio signal by a Mel encoder. The Mel feature space representation may be related to the Mel frequency cepstral representation. The corresponding Mel decoder and separation stage may be jointly trained to minimize a joint loss function that includes contributions from the Mel transform and contributions related to the difference function. The Mel decoder may include or correspond to a MelGAN.
[0030] In some embodiments, the representation of the mixed audio signal may involve a problem-agnostic speech encoder (PASE), and the feature space representation may be generated from the mixed audio signal by a deep learning-based PASE encoder. The PASE decoder and separation level corresponding to the PASE encoder can be jointly trained to minimize a joint loss function that includes contributions from the PASE transform and contributions associated with the difference function.
[0031] Both the Mel feature space and the PASE feature space are specifically tailored for speaker separation, thus enabling accurate and efficient separation. The proposed method provides an efficient framework for implementing these feature spaces.
[0032] In some embodiments, the sound source can be associated with a speech source. In other words, the sound source can be associated with the speaker, such as in a teleconference.
[0033] Maintaining speaker continuity is particularly important in audio applications involving multiple speakers, such as teleconferences or video conferences. The proposed method reliably ensures speaker continuity.
[0034] In some embodiments, the system's separation level can be based on one of the Conv-TasNet, DPRNN, WaveNet, and Demucs architectures.
[0035] According to another aspect, a method for sound source separation using a deep learning-based system is provided. The system includes a separation stage for extracting sound source representations frame-by-frame from a representation of an audio signal. The system may also include a clustering stage for generating, for each frame, a vector indicating the allocation permutation of the extracted sound source representations to the corresponding sound sources. The audio signal representation may be waveform-based. The method may include generating a representation of a mixed audio signal by applying a waveform-based transform to the mixed audio signal. This representation is waveform-based. The mixed audio signal may include at least two sound sources. The method may also include inputting the representation of the mixed audio signal into the separation stage to generate frames of the extracted sound source representations. The method may further include inputting the frames of the extracted sound source representations into the clustering stage to generate, for each frame of the mixed audio signal, a vector indicating the allocation permutation of the extracted sound source representations to the corresponding sound sources selected for the corresponding frames of the mixed audio signal. The method may further include clustering the vectors generated at the cluster level according to a clustering algorithm, and determining the corresponding assignment permutation (its estimate) selected for frames of the mixed audio signal based on the vector's affiliation to different clusters, each cluster representing a specific assignment permutation. The method may further include assigning the extracted representation frames to a set of audio streams according to the determined assignment permutations. The method may further include applying an inverse waveform-based transform to the set of audio streams to generate a waveform audio stream of the extracted sound sources.
[0036] According to another aspect, a computer program is provided. This computer program may include instructions that, when executed by a processor, cause the processor to perform all steps of the methods described herein.
[0037] According to another aspect, a computer-readable storage medium is provided. This computer-readable storage medium can store the aforementioned computer program.
[0038] According to another aspect, an apparatus is provided, including a processor and a memory coupled to the processor. The processor is adapted to perform all steps of the methods described throughout this disclosure.
[0039] It should be understood that the apparatus features and method steps can be interchanged in various ways. In particular, the details of the disclosed method can be implemented by the corresponding apparatus and vice versa, as will be understood by those skilled in the art. Furthermore, any statements made above regarding the method should be understood to apply equally to the corresponding apparatus and vice versa. Attached Figure Description
[0040] The following explanation of exemplary embodiments of this disclosure is based on the accompanying drawings, wherein...
[0041] Figure 1 is a schematic diagram of a standard deep learning-based framework for speaker source separation.
[0042] Figure 2 This is a schematic diagram of discourse-level permutation-invariant training for a deep learning-based framework used for speaker source separation.
[0043] Figure 3 This is a schematic diagram of frame-level permutation-invariant training for a deep learning-based framework used for speaker source separation.
[0044] Figure 4 This is a schematic high-level illustration of an example of a method for training a deep learning-based system for sound source separation according to embodiments of the present disclosure.
[0045] Figure 5 This is a flowchart illustrating an example of a method for training a deep learning-based system for sound source separation according to embodiments of the present disclosure.
[0046] Figure 6 This is a schematic diagram illustrating an example of training a separation level of a deep learning-based system based on a first example of waveform-based representation of an audio signal according to embodiments of the present disclosure.
[0047] Figure 7 This is a schematic diagram illustrating an example of training a separation level of a deep learning-based system based on a second example of waveform-based representation of an audio signal according to embodiments of the present disclosure.
[0048] Figure 8 This is a schematic diagram illustrating a third example of training a deep learning-based system for a separation level based on a waveform-based representation of an audio signal according to embodiments of the present disclosure.
[0049] Figure 9 This is a schematic diagram illustrating another example of training a separation level of a deep learning-based system based on a third example of waveform-based representation of an audio signal according to embodiments of the present disclosure.
[0050] Figure 10 This is a schematic diagram illustrating a fourth example of training a deep learning-based system for a separation level based on waveform-based representation of an audio signal according to embodiments of the present disclosure.
[0051] Figure 11 This is a schematic diagram illustrating another example of training a separation level of a deep learning-based system for a fourth example of waveform-based representation of an audio signal according to embodiments of the present disclosure.
[0052] Figure 12 This is a schematic diagram illustrating an example of training a deep learning-based system at the cluster level according to embodiments of the present disclosure.
[0053] Figure 13 This is a flowchart illustrating an example of the proposed method for sound source separation using a deep learning-based system according to embodiments of the present disclosure.
[0054] Figure 14 This is a schematic diagram of an example of an apparatus for performing a method according to an embodiment of the present disclosure. Detailed Implementation
[0055] The accompanying drawings and the following description refer to preferred embodiments only by way of illustration. It should be noted that, from the discussion below, alternative embodiments of the structures and methods disclosed herein will readily be considered as feasible alternatives that can be employed without departing from the claimed principles.
[0056] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. It should be noted that, wherever feasible, the same or similar reference numerals may be used in the drawings and may indicate the same or similar functions. The drawings depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein.
[0057] Speaker source separation (e.g., deep learning-based multi-speaker source separation) aims to separate each speaker's voice from their mixed signal (mixed audio signal), which is modeled as the sum of the individual speaker signals. Deep learning-based sound source separation (e.g., multi-speaker source separation) aims to use deep neural networks to separate each sound source (e.g., each speaker's voice) from their mixed signal.
[0058] Some methods rely on processing spectrograms or waveforms, typically based on the structure schematically illustrated in Figure 1. Within this standard deep learning-based framework, spectrograms can be processed (using fixed transforms, such as STFTs) or waveform-based signals (e.g., using learnable transforms). Typically, spectrogram-based models use signal processing-driven transforms, while waveform-based models can jointly optimize learnable transforms with the rest of the network. It is worth noting that Figure 1 only shows the case of two speakers, but the standard deep learning-based framework can often be easily extended to more than two speakers or sound sources.
[0059] In the spectrogram-based approach, the input mixed signal (mixed audio signal) 10 is first transformed into a time-frequency representation 22 using an established signal processing transform 20, such as a Short-Time Fourier Transform (STFT). Next, a separator 40 implemented using a deep neural network is employed to learn a set of masks 23-1, 23-2, one for each speaker. The mask value for each time-frequency segment represents the degree of presence of each speaker in that segment. These estimated masks 23-1, 23-2 are then multiplied by the mixed representation 22 to obtain the separated speaker spectrograms 42, 44. Finally, an inverse transform 25, such as an inverse STFT (ISTFT), is used to obtain the separated waveforms 46, 48. In the spectrogram-based model, only the separator 40 is learnable (trainable), while the signal transforms 20, 25 are, for example, STFT / ISTFT.
[0060] In the waveform-based approach, both the signal transform 20 and its inverse transform 25 are learnable, meaning they can be jointly optimized with the separator 40 during training. The forward transform 20, typically referred to as the encoder, usually uses a one-dimensional convolutional layer to transform the mixed signal 10 to a high-dimensional latent space (e.g., a 512-dimensional latent space), while the decoder (the learnable inverse transform) 25 transforms the latent space back into the waveform domain.
[0061] Compared to spectrogram-based methods, waveform-based models operate at a higher temporal resolution (i.e., spectrum frames compared to waveform samples), which makes waveform-based models computationally demanding.
[0062] The most useful speaker separation models are trained to be speaker-independent. Therefore, the identity (e.g., label) of the separated speaker is unknown prior to the training. Training such speaker-independent separation models encounters the permutation ambiguity problem, which will be described below.
[0063] More specifically, in order to train the speaker separation model (e.g., the model in Figure 1), the model's parameters are iteratively updated by minimizing a loss function defined as the difference between the estimated signal and the reference (true) signal.
[0064] Taking a two-speaker example, for a mixed signal x, there are two reference speech signals (ref1, ref2) available for calculating the loss function, and the model separates two speech estimates (est1, est2). However, since the trained model is assumed to be speaker-independent (rather than designed for two specific speakers), the speaker's identity is unknown. Therefore, the corresponding reference identity is unknown when calculating the loss function. Consequently, there are two efficient methods to assign the estimates to the references (defined as permutations or assigned permutations) when calculating the loss function, i.e. as well as The loss associated with each permutation can be (illustratively) defined as:
[0065] (1)
[0066] (2)
[0067] Specifically, loss can be defined as follows:
[0068] (1a)
[0069] (2a)
[0070] It should be understood that equations (1) and (2) describe general principles for handling different permutation assignments when calculating the loss, and these equations should not be construed as limiting this disclosure. For example, the square of the difference term |•|, or its logarithm, or the logarithm of the squared difference term, etc., can be used instead. It should also be understood that equations (1), (1a), (2), and (2a) are schematic equations that omit any indices indicating frames, samples, etc., and a specific implementation of the loss will include such indices and an appropriate sum of these indices.
[0071] To further illustrate this problem, it might be useful to consider a speaker-related source separation formula without permutation ambiguity. In such an example, all mixed signals always contain the same two speakers. Therefore, during training, est1 can always be assigned to ref1 (speaker 1) and est2 to ref2 (speaker 2), corresponding to equation (1) above. As a result, the network will remember the difference between the two speakers, and during separation, the model can always output est1 as speaker 1 and est2 as speaker 2.
[0072] However, for speaker-independent tasks aimed at separating mixtures of any two speakers, the training set must contain a variety of speakers, including, for example, different / same genders, ages, pitches, loudnesses, etc. Therefore, for each specific mixture signal, there is ambiguity between Equations (1) and (2) as a loss when choosing. Furthermore, if arbitrary assignment permutations are used, model training often fails to converge due to incorrect or inappropriate objectives.
[0073] As a major consequence of permutation ambiguity during training, the resulting speaker separation model may suffer from speaker identity swapping in certain parts of the separated signal.
[0074] A widely used method for training deep neural networks for speech source separation is based on selecting a permutation in equation (1) or (2) that provides a smaller loss, and then minimizing that loss to update the network. This method is called Permutation Invariant Training (PIT).
[0075] There are two basic types of PIT: utterance-level PIT (uPIT) (see, for example, Morten Kolbaek et al., Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks, https: / / arxiv.org / pdf / 1703.06284.pdf) and frame-level PIT (tPIT) (see, for example, Dong Yu et al., Permutation Invariant training of deep models for speaker-independent multi-talker speech separation, https: / / arxiv.org / pdf / 1607.00325.pdf). These methods compute permutation errors at different time scales. In uPIT, the error associated with each permutation in Equations (1) and (2) is computed over the entire signal (i.e., utterance) length. Thus, the optimal frame-level permutation is trained and strengthened by uPIT, just like the optimal utterance-level permutation. However, this goal is not always achieved through the training process. Figure 2 illustrates the training progress of uPIT. In the early stages, speaker identities (represented by different patterns) may frequently switch between frames in the separated signals 210 and 220. uPIT attempts to enhance global speaker coherence, where frames est1 210 and 230 should ideally all be associated with one speaker (e.g., all with vertical solid line patterns), and frames est2 220 and 240 should ideally all be associated with another speaker (e.g., all with vertical dashed line patterns). However, in the example shown, speaker switching still exists in local frames (e.g., in the first three frames of est1 230 and est2 240) after uPIT training. That is, due to the suboptimal nature of uPIT in each local frame, it leads to speaker switching in certain parts of the separated signals est1 and est2. Nevertheless, because uPIT operates at the utterance level, it remains a very popular training loss for both waveform-based and spectrogram-based models.
[0076] Figure 3 illustrates the training progress of tPIT. Unlike uPIT, tPIT's training process independently finds the optimal permutation for each frame of the mixed signal by calculating the local error in each frame. Therefore, the optimal permutation at each frame is independent of other frames. This local optimum criterion ensures high-quality source separation at each frame. Samples (or spectral frequency components) within each frame of an output accurately belong to the same speaker. However, tPIT does not enforce the estimation of global speaker coherence for est1310 and est2320, and is therefore not a good method for cross-frame speaker tracking (i.e., speaker identity, speaker coherence). To ensure cross-frame speaker coherence, tPIT requires ground truth signals (reference audio signals) 330 and 340 during both training and testing. During training, smaller frame-level permutation errors are minimized, resulting in the aforementioned high local speaker separation accuracy. During inference (i.e., at separation or test time), tPIT can use truth signals 330, 340 to reassemble frames with the same permutation identity into utterances to achieve reordered estimates est1 350 and est2 360. However, in practice, truth signals are not available during testing.
[0077] Given the above, there is a trade-off between uPIT and tPIT—uPIT is designed to maintain speaker coherence along the estimate (while compromising the quality of separation), while tPIT is designed to partially separate the speaker (while failing to maintain speaker coherence along the estimate).
[0078] Typically, speech source separation is preferably performed in an end-to-end manner (i.e., waveform input, waveform output). As mentioned above, deep learning methods for speech source separation suffer from either suboptimal separation quality or the so-called source substitution ambiguity problem (e.g., speaker substitution ambiguity). The source substitution ambiguity problem remains unresolved, especially for training waveform-based separation models that are independent of the source (e.g., independent of the speaker).
[0079] This disclosure resolves the permutation ambiguity problem, particularly for waveform-based separation models. It is applicable to the training of any waveform-based model where permutation ambiguity may occur.
[0080] In a general sense, this disclosure uses a tPIT separation level, followed by a second level implementing a speaker tracking model. This second level can be a clustering level as described below. The methods and systems of this disclosure operate in the "generalized signal space" of the waveform model, independent of the spectrogram. Since waveform-based models are computationally demanding, this disclosure proposes a memory-efficient loss function to achieve linear memory complexity O(T) instead of quadratic complexity O(T). 2This disclosure relies on a general three-level framework to efficiently train speaker tracking models (e.g., cluster-level). Therefore, this disclosure depends on a general three-level framework to reduce or completely avoid permutation errors in waveform-based separation models.
[0081] • The first stage sets up a pair of encoders and decoders, performing signal transformations into a “generalized signal space” that is very flexible in its setup, as well as inverse signal transformations from the “generalized signal space” that is very flexible in its setup.
[0082] • The second stage uses frame-level permutation-invariant training (tPIT) to perform high-quality speaker separation at each time step in the transformed domain.
[0083] • The third level tracks (permutations) identities over time and addresses speaker permutation ambiguity through clustering models.
[0084] Therefore, a portion of this disclosure can be summarized as relating to tPIT, followed by clustering in the generalized signal space for waveform models, and specifically to the efficient training of the tPIT separation level, followed by the speaker tracking level for waveform-based models. Combining waveform-based models in the “generalized signal space” for tPIT with speaker tracking allows input data to be projected (or mapped) into a wide range of (pre-trained or fixed) latent spaces or other feature spaces, thus providing additional flexibility and adaptability for source separation. It should be noted that the use of tPIT and clustering for waveform-based models has not been applied in conventional methods.
[0085] Figure 4 is a schematic high-level illustration of an example of the proposed method for training a deep learning-based system for sound source separation (e.g., speech source separation). Training the system can be equivalent to (iteratively) determining (or tuning) the parameters of the deep learning model (e.g., a deep neural network) used to implement the system.
[0086] The system comprises three levels: Generalized Signal Transform (top diagram), tPIT Separation (middle diagram), and Clustering (bottom diagram). In some implementations, the Generalized Signal Transform stage can be configured such that the system operates directly on waveform audio signals (e.g., segmented waveform signals). In some such cases, it can be said that this transform stage is not present in the system.
[0087] More generally, the system includes a transform stage 420 (optional in some implementations), a separation stage 440, and a clustering stage 450. If the system includes a transform stage 420, it also includes an inverse transform stage 425. The system is understood to be suitable for performing source separation on a mixed audio signal 410 of the input system. Here, the mixed audio signal 410 is a waveform audio signal. It can include multiple different sound sources (e.g., a speaker). For example, it can be obtained as a mixture of signals from different sound sources.
[0088] In a general sense, transform stage 420 is typically adapted to transform input (waveform) audio signal 405 into representation 422 of that (waveform) audio signal. Separation stage 440 is adapted to extract representations of sound sources frame by frame from the representation of the input audio signal (e.g., mixed audio signal 410). Clustering stage 450 is adapted to generate a vector (e.g., embedding vector) for each frame, which indicates the permutation of the extracted sound source representation frames to the corresponding sound sources (or sound source labels).
[0089] The following will refer to Figure 5 The flowchart describes an example of the method 500 for training the above-described system for sound source separation. It includes at least a separation level 440 and a clustering level 450.
[0090] Step S510 of method 500 is to obtain a representation of the mixed audio signal and representations of at least two reference audio signals as input. As described above, the representation is a waveform-based representation. The mixed audio signal includes at least two sound sources, and the reference audio signals correspond to the respective sound sources included in the mixed audio signal. Therefore, the reference audio signals are used as the ground truth signals for training the separation stage.
[0091] For example, the above representation can be obtained by applying a transform level to the mixed audio signal and the reference audio signal.
[0092] Step S520 involves inputting a representation of the mixed audio signal and representations of at least two reference audio signals into a separation stage, and training the separation stage to extract representations of sound sources from the representation of the mixed audio signal. This is accomplished as follows: for each frame, a difference function is minimized. Here, the difference function represents the loss function in the context of training the separation stage. Furthermore, the difference function is based on the difference between the frame of the extracted sound source representation and the frame of the reference audio signal representation. In some implementations, the difference function indicates a combined difference (e.g., a combination of differences) between the frame of the extracted sound source representation and the frame of the reference audio signal representation. For each extracted sound source representation (i.e., for each estimated sound source), the combined difference includes the difference between the frame of the extracted sound source representation and the corresponding frame of the reference audio signal representation. Thus, for the example case of two sound sources, there will be two difference terms in the combined difference. Examples of the difference function (loss function) in the two-sound-source case are given in equations (1) and (2) above. In this context, for each frame, this allocation permutation of the extracted source representation and the reference audio signal representation is chosen to minimize the difference function, resulting in a smaller (or typically, minimal) difference function. Therefore, the separation level is trained within the tPIT framework.
[0093] Step S530 is as follows: it takes a representation of the mixed audio signal as input, and for each frame of the mixed audio signal representation, it inputs the extracted source representation frame and the indication of the assigned permutations already selected for the corresponding frame of the mixed audio signal representation to a clustering level, and trains the clustering level to generate vectors indicating the permutation assignments from the extracted source representation frames to the corresponding source. This is achieved in such a way that the separation between the vector groups of the frames of the mixed audio signal is maximized. Specifically, the frame vectors are grouped according to the corresponding assigned permutations indicated by these vectors. In other words, the frame vectors are grouped according to the known permutation labels already used for the frames.
[0094] return Figure 4 The details of transform stage 420 will now be described. Typically, transform stage 420 is adapted to transform a mixed (waveform) audio signal 410 (or any input waveform audio signal 405) into a representation 422 of the mixed audio signal. This representation is waveform-based. Therefore, the mixed audio signal is transformed to a signal space that is a waveform-based signal space. For example, the signal space may not be a frequency domain signal space, nor may it be based on a spectrogram (i.e., the transformation may not involve determining a spectrogram). In other words, the transformation may not involve time-frequency transformation. In particular, the transformation may be independent of STFT.
[0095] Figure 4 illustrates the transformation of the input waveform audio signal 405 to a set of generalized features 422. These features may, for example, relate to (short) frames of the input waveform audio signal or to features in a latent feature space. For ease of reference, the signal space may be referred to as the generalized signal space in the remainder of this disclosure. The inverse transform stage 425 is adapted to transform the representation of the audio signal 424 back to the waveform domain, i.e., to the waveform audio signal 430. Specifically, the inverse transform stage 425 can be used to transform the extracted sound source representation back to the waveform domain.
[0096] By considering the transformation stages as described above, the system operates on waveform-based representations of audio signals or directly on waveform audio signals. Therefore, this disclosure applies a combined tPIT and clustering strategy to waveform-based models. The stages referred to as tPIT stages in conventional models, which do not operate in the STFT domain, are now divided into two stages: a signal transformation stage and a tPIT separation stage. This partitioning or refactoring enables more general signal transformation applicable to waveform-based models.
[0097] The following describes an example of a generalized transformation performed by a transformation stage.
[0098] • (i) Short waveform frames (time domain). In this case, the signal transform simply divides the waveform into short frames, and the inverse transform reconstructs the waveform by overlapping and adding the frames. The true value for the tPIT loss is defined directly on the waveform samples in the time domain.
[0099] • (ii) Pre-trained encoder-decoder optimized for speaker separation. In this case, the encoder / decoder is first pre-trained using an ideal mask (i.e., no deep learning-based separator) to project the waveform into a space optimized for speaker separation. Next, the encoder / decoder is fixed, and a separator model is trained in this learning space using tPIT.
[0100] • (iii) Mel spatial coding and MelGAN decoder. The Mel domain originates from human perception and is therefore a natural choice for speaker source separation. tPIT targets the Mel features of the ground truth waveform. MelGAN (see, e.g., Kundan Kumar et al., MelGAN: generative adversarial networks for conditional waveform synthesis, https: / / arxiv.org / pdf / 1910.06711.pdf) is a DL-based generative model trained to accurately reconstruct the signal phase from the separated Mel features. The advantage of separation in the Mel domain is that it decouples the separation task from phase modeling, thus allowing for the use of low temporal resolution.
[0101] • (iv) Problem-agnostic speech encoder and decoder based on deep learning. PASE (see, e.g., Santiago Pascual et al., Learning problem-agnostic speech representations from multiple self-supervised tasks, https: / / arxiv.org / pdf / 1904.03416.pdf) is a deep learning-based speech feature extractor. The PASE encoder is pre-trained to learn general deep features that encode different levels of signal extraction. The PASE decoder is trained to perform inverse transform. PASE has the same low temporal resolution advantage as Mel features, but may be more expressive than Mel.
[0102] Consistent with the above, the transform operations applied to transform level 420 may involve one of the following: segmenting the mixed audio signal into multiple frames in the time domain (see example scheme (i)), projecting the mixed audio signal onto a deep learning-based encoding in a latent feature space optimized for source separation (see example scheme (ii)), Mel space encoding (see example scheme (iii)), and deep learning-based problem-agnostic speech encoding (see example scheme (iv)). Where applicable, the transform level (and / or inverse transform level) may be jointly trained with the separation level and / or clustering level.
[0103] As described above, the system may also include an inverse transform stage 425, which transforms the extracted sound source representation back into the waveform domain. In the case of Mel encoding, the inverse transform stage may include or correspond to a MelGAN decoder.
[0104] Although implementation options for the generalized transform are listed in the example schemes (i) through (iv) above, this disclosure should not be construed as being limited to these implementation examples. Any kind of (waveform-based) transform can be employed within the generalized framework for speech source separation of this disclosure. For this reason, the proposed source separation framework is referred to herein as operating on a "generalized signal space" basis.
[0105] For the first two generalized transformations above (example schemes (i) and (ii)), a high (sample-based) temporal resolution is required. This results in a large number of frames for clustering. The pairwise similarity loss described in Equation (10) below would become too memory-intensive to train the clustering model at such a high temporal resolution. This disclosure proposes a memory-efficient similarity loss that makes training the clustering model feasible.
[0106] Next, the separation stage 440 is described in more detail. Typically, the separation stage 440 is adapted to extract (waveform-based) source representations frame-by-frame from the (waveform-based) representation of the input audio signal 410. As mentioned above, the input audio signal can be a mixed audio signal including multiple sound sources. The separation stage 440 outputs frames representing the extracted audio sources 442, 444. During training, the representations of the reference audio signals 416-1, 416-2 (i.e., the representations of the truth signals) are used to iteratively train the separation stage 440 to minimize the loss function (tPIT loss function). The representations of the reference audio signals 416-1, 416-2 can be obtained by applying the transform stage 420 to the reference audio signals (truth signals) 415-1, 415-2. Depending on the applicable assignment permutation, the loss function for a given frame of the mixed audio signal can be determined by comparing the frames representing the extracted sound sources 442, 444 from the given frame of the mixed audio signal representation with the corresponding representations of the ground truth signals 416-1, 416-2. There is such a loss function for each possible assignment permutation, and the assignment permutation that produces the minimum loss function is finally selected. The separation level 440 is then trained to minimize this loss function.
[0107] The source representations extracted from the output of separation stage 440 are frame streams of 442 and 444 as estimates est1 and est2. Due to the source ambiguity trained in the tPIT framework, each stream may contain frames associated with different sources.
[0108] The following describes an example of training the separation stage 440 in the tPIT framework for different transformations performed by the transformation stage 425.
[0109] • (i) Short waveform frames (time domain). Figure 6 is a schematic high-level illustration of an example of a method for training a deep learning-based system for sound source separation, in the case where the representation of a mixed audio signal involves segmenting the mixed audio signal into (short) waveform frames. In this case, transform stage 620 involves segmenting the input audio signal 605 into (short) waveform frames 622 (by... Figure 6 The segmentation level (shown by the sparse vertical dashed line pattern in the image) is used. Each frame contains samples of short waveform segments. The mixed waveform 610 and the true signals 615-1, 615-2 (ref1, ref2) are segmented into very short frames. No signal transformation to another space is required. The mixed frames are processed by the speaker separation model (separation level 640), and the output frames 642, 644 (est1, est2) are also compared with the true signal frames 616-1, 616-2 in the time domain using tPIT. For frame t in the case of two speakers, the tPIT loss for each permutation is calculated as follows:
[0110] (3)
[0111] (4)
[0112] Where L indicates the frame length in the sample, the separation model is selected and the smaller loss is minimized. For example, the total number of samples in frame L is in the range of 2-16 samples. Alternatively or additionally, L can be selected according to the sampling rate, such that the frame corresponds to a duration of 0.25-2 ms.
[0113] In equations (3) and (4), the square of the difference term |•|, or its logarithm, or the logarithm of the squared difference term, can be used instead. During separation, the reordered frames (in the order predicted by the clustering model, not shown here) are overlapped and added in inverse transform stage 625 to reconstruct each separated discourse. In this sense, inverse transform stage 625 can be said to relate to overlapped and additive stages.
[0114] In other words, tPIT operates directly on short waveform frames of the mixed audio signal 610 in the time domain. Reference audio signals (truth signals) 615-1 and 615-2 are also segmented by transform stage 620 into (short) waveform frame sequences 616-1 and 616-2. Separation stage 640 generates frames of the extracted audio sources 642 and 644 (in... Figure 6 (Represented by vertical dashed and solid lines). In the tPIT framework, the extracted frames of audio sources 642 and 644 are compared with frames 616-1 and 616-2 of the reference audio signal, and the loss function (difference function) is minimized.
[0115] Therefore, the separation level 640 can be trained to determine the extracted frames of sound sources 642 and 644 from the frames of the mixed audio signal 610 for each frame of the mixed audio signal 610 in a manner that minimizes the following loss function:
[0116] (5)
[0117] Where t indicates the frame, l indicates the number of samples within the frame, L is the total number of samples within the frame, n is the label of the extracted sound source, N is the total number of extracted sound sources, est represents the frame of the extracted sound source, and ref represents the frame of the reference audio signal. It is a permutation map for labels n=1, ..., N, indicating the label of the reference audio signal. Alternatively, the square of the difference term |•|, or its logarithm, or the logarithm of the squared difference term, etc., can be used. As described above, for each frame, the permutation map... It is chosen to produce the minimum loss function. Due to its construction method, the loss function can be called the difference function.
[0118] It should be noted that the loss function of equation (5) is applicable to the general case where N sound sources are included in the mixed audio signal 610.
[0119] Here, and in the remainder of this disclosure, the reordering (or reorganization) of frames is understood in the sense of reassigning frames to different sound sources. For example, at a given time, frames est1 and est2 can be swapped separately. However, reordering does not mean reordering along the timeline, i.e., swapping frames at different times.
[0120] • (ii) Pre-trained encoder-decoder optimized for speaker separation. Instead of directly computing tPIT in the time domain, tPIT can be performed in a pre-trained latent space optimized for speaker separation. In this case, the above representation of the mixed audio signal can be said to be related to the latent feature space representation generated by the (pre-trained) deep learning-based encoder. A high-level representation of this case is schematically shown in Figure 7. Encoder 720 / decoder 725 is pre-trained using an ideal mask to learn an optimized latent space for speaker separation. The separator 740 is then trained in this space using tPIT, where encoder 720 and decoder 725 are frozen / fixed (in...). Figure 7 The fixed encoder 20 is indicated by the dashed box.
[0121] The upper part of Figure 7 illustrates the scheme for the encoder / decoder components used in pre-training the source separation network (source separation system). During the pre-training of the encoder / decoder, an ideal mask is generated from the ground-value encoder features (i.e., representations of the generated reference audio signals 715-1, 715-2) via a softmax function 721. The decoder 725 transforms the separated latent features V1(t,f), V2(t,f) of each speaker back into the waveform domain to estimate est1, 726-1 and est2, 726-2. The loss (loss function) used to learn the encoder / decoder transform is based on optimizing (e.g., minimizing) the signal-to-distortion ratio of the estimated speech sources est1 and est2 relative to the ground-value speech signals ref1, ref2. For more details on how to pre-train the encoder / decoder, see Efthymios Tzinis et al., Two-step sound source separation: training on learned latent targets, https: / / arxiv.org / pdf / 1910.09804.pdf.
[0122] Therefore, the pre-trained deep learning-based encoder 720 involves inputting the mixed audio signal 710 and reference audio signals 715-1, 715-2 into the deep learning-based encoder 720, and generating latent feature space representations of the mixed audio signal and the reference audio signal through the deep learning-based encoder 720. Furthermore, for each reference audio signal 715-1, 715-2, a mask for the reference audio signal 715-1, 715-2 is generated based on the latent feature space representation of the reference audio signal using a softmax function 721. The generated mask is applied to the latent feature space representation of the mixed audio signal to determine the representation of the extracted audio signal. The deep learning-based decoder 725 is applied to the representation of the extracted audio signal to obtain estimates of the extracted audio signals 726-1, 726-2, wherein the deep learning-based decoder 725 is adapted to perform the inverse operation of the deep learning-based encoder 720. In this framework, the deep learning-based encoder 720 (and correspondingly, the deep learning-based decoder 725) is trained (iteratively) in such a way that the signal distortion ratio of the decoded estimates of the extracted audio signals 726-1, 726-2 and the corresponding reference audio signals 715-1, 715-2 is minimized.
[0123] Figure 7The lower figure depicts the scheme of the separator 740 used for the actual training of the source separation model (source separation system). It should be noted that this training is performed within the tPIT framework, i.e., the training is tPIT. Specifically, after pre-training the encoder / decoder, their parameters are set to freeze (indicated by dashed box 720), and the separator component (separation stage 740) acquires the mixed latent space representation X(t,f) (representation of the mixed audio signal 710) provided by the encoder 720. The separation stage 740 learns the features of the estimated signals 742, 744 in the latent space. and The goal is to make and Features of true signals 716-1 and 716-2 in the latent space and Match as closely as possible. Here, and This is the pre-trained latent target derived from the ideal mask. For this example, the tPIT permutation error at each time step t can be calculated as:
[0124] (6)
[0125] (7)
[0126] Alternatively, the square of the difference |•|, or its logarithm, or the logarithm of the squared difference can be used. Again, according to the tPIT paradigm, a smaller loss is minimized at each time step. Minimizing the tPIT loss ensures that all frequency components in the separated frames always belong to the same speaker. After the reconstructed frames (described below), the pre-trained decoder 725 converts them back to the waveform domain (i.e., waveform or waveform audio signal).
[0127] Therefore, the separation level 740 is trained to determine the extracted sound source 742, 744 (representations) frames from the frames of the mixed audio signal (representation) for each frame of the mixed audio signal 710 in a manner that minimizes the following loss function:
[0128] (8)
[0129] Where t indicates the frame, f indicates the features in the latent feature space, n indicates the labels of the extracted sound sources and the reference audio signal, and N is the total number of extracted sound sources. V indicates the frame representing the extracted sound source, and V indicates the frame representing the reference audio signal. It is a permutation map for labels n=1, ..., N, indicating the label of the reference audio signal. Alternatively, the square of the difference term |•|, or its logarithm, or the logarithm of the squared difference term, etc., can be used. As described above, for each frame, the permutation map... It is chosen to produce the minimum loss function. Due to its construction method, the loss function can be called the difference function.
[0130] As described above, the representation of the reference audio signal can be determined based on the representation of the mixed audio signal and a set of masks. This set of masks can have already been determined based on the reference audio signal during the pre-training of a deep learning-based encoder.
[0131] It should be noted that the loss function of equation (8) is applicable to the general case of N sound sources included in the mixed audio signal 710.
[0132] • (iii) Mel spatial encoding and MelGAN decoder. In this example scheme, the representation of the mixed audio signal involves a Mel feature space representation (e.g., a Mel frequency cepstral representation) that can be generated from the mixed audio signal by a Mel encoder. Mel filter banks are derived from human perception and are therefore a natural choice for speaker source separation. Since the Mel spectrogram itself does not model the phase, MelGAN is a recently proposed generative model (see, for example, Kundan Kumar et al., MelGAN: generative adversarial networks for conditional waveform synthesis, https: / / arxiv.org / pdf / 1910.06711.pdf) and is trained to accurately reconstruct the signal phase from Mel features. Decoupling amplitude and phase modeling for speaker separation has significant advantages. The separation model only needs to focus on the separation task in the spectral amplitude domain (leaving phase modeling to MelGAN), whereas in the previous two example schemes (i) and (ii), the separator must perform both separation and phase modeling simultaneously. Therefore, by utilizing Mel features, speaker separation can work at a significantly reduced temporal resolution while still accurately reconstructing the phase. Therefore, operating in the Mel feature space is much cheaper computationally than operating directly in the waveform domain.
[0133] A high-level representation of the Mel / MelGAN case is schematically shown in Figure 8. The upper figure relates to the training of MelGAN 825, and the lower figure relates to the training of the separation model (separation level) 840. MelGAN 825 is first trained using Mel filter bank features 822 generated by the Mel encoder 820 from a clean training set (input waveform audio signal 805) (frames in Figure 8 with sparse vertical dashed line patterns). Then, the separation model 840 is trained in the Mel domain using tPIT. Figure 8 The frames with vertical solid lines and vertical dashed lines correspond to the separate frames 842 and 844.
[0134] One way to train MelGAN 825 is to use clean, distortion-free data, as shown in the example in Figure 8. MelGAN training is completely independent of speaker separation training. After separation using tPIT, the Mel features of each speaker are converted back into waveforms by the generator part of MelGAN 825.
[0135] However, due to the distortion contained in the separated speech, a mismatch can occur between the training data of MelGAN and the output data of the separation model if trained alone. Therefore, as an alternative, MelGAN 925 and the tPIT model (i.e., the separation level 940) can be trained jointly. Figure 9 schematically illustrates this joint training scheme of tPIT model 940 and MelGAN 925. In this case, before sending the frames of the separated Mel features 942, 944 to MelGAN 925, these frames must be reordered 945 into corresponding output utterances 946, 948 based on the optimal tPIT permutation, so that the MelGAN input to MelGAN925 will have the same speaker order (speaker assignment) as the ground truth waveforms 915-1, 915-2. Therefore, the GAN loss can be calculated correctly. The ground truth waveforms 915-1, 915-2 are also input to the MelGAN decoder 925 for training. The total loss of the joint training is given by the following formula:
[0136] (9)
[0137] The tPIT loss is determined based on the frames of the separated Mel features 942, 944 and the frames of Mel features 916-1, 916-2 of the truth signals 915-1, 915-2 generated by the Mel encoder 920 from the truth signals 915-1, 915-2. For example, the tPIT loss can be determined similarly to equations (6) and (7) or equation (8).
[0138] At the start of joint training, the outputs 942 and 944 of the separate model 940 will exhibit many permutation errors, which may cause MelGAN 925 to have difficulty converging. To mitigate this issue, the tPIT model 940 can be warmed up until it generates reasonably correct outputs before starting joint training with MelGAN 925.
[0139] Therefore, the Mel decoder 925 and the separation stage 940, corresponding to the Mel encoder 920 used to generate the representation of the mixed audio signal, can be jointly trained to minimize a joint loss function that includes contributions from the Mel transform (e.g., contribution GAN_loss in Equation (9)) and contributions related to the difference function (e.g., contribution tPIT in Equation (9)). The Mel decoder 925 includes or corresponds to MelGAN.
[0140] Generally, it can be said that part of this disclosure relates to combining tPIT with generative adversarial networks (GANs) for speaker source separation. It should also be noted that while the Mel / MelGAN framework is one of the preferred implementations of this disclosure, any vocoder (deep learning-based or fixed) can be employed in the context of this disclosure. Therefore, example case (iii) should not be construed as being limited to Mel / MelGAN frames.
[0141] • (iv) Deep Learning-Based Problem-Agnostic Speech Encoder (PASE) and Decoder. In this example scheme, the representation of the mixed audio signal is related to a PASE feature space representation that can be generated from the mixed audio signal by a deep learning-based PASE encoder. The Mel feature space is a handcrafted speech feature space based on filter banks. By using a deep learning-based speech encoder, more powerful features can be utilized. PASE (see, for example, Santiago Pascual et al., Learning problem-agnostic speech representation from multiple self-supervised tasks, https: / / arxiv.org / pdf / 1904.03416.pdf) is a data-driven, deep learning-based feature extractor. PASE features share the same advantage as Mel features, namely low temporal resolution. Unlike Mel features, PASE provides a deep, hierarchical abstraction of the underlying signal processing characteristics that are important for speaker separation, such as speaker ID, pitch, phonemes, etc. Therefore, computing PASE features may be more expressive than computing Mel features.
[0142] A high-level representation of the PASE encoder / decoder configuration is schematically shown in Figure 10. The upper part of the figure shows the PASE encoder 1020 and the PASE decoder 1025. The PASE encoder 1020 transforms the input audio signal 1005 into a PASE feature space representation 1022 of the input audio signal 1005. After its processing, the PASE feature space representation 1024 may be transformed back into the output waveform audio signal 1030 by the PASE decoder 1025.
[0143] The PASE encoder 1020 and PASE decoder 1025 can be pre-trained first. Figure 10 The frames with sparse vertical dashed line patterns represent the PASE features of the data used to train the PASE encoder 1020 and PASE decoder 1025, namely the PASE feature space representation 1022 of the input audio signal 1005. In the tPIT step shown in the lower figure, i.e., during the training of the separation stage 1040, the PASE encoder 1020 is frozen (indicated by the dashed block) as a feature extractor. The separation stage 1040 is trained to minimize the loss function (tPIT loss) based on the frames of the separated PASE features 1042, 1044, and the frames of the PASE features 1016-1, 1016-2 of the truth signals 1015-1, 1015-2 generated by the (fixed) PASE encoder 1020 from the truth signals 1015-1, 1015-2. For example, the tPIT loss can be determined similarly to the tPIT loss in the Mel / MelGAN case described above. That is, for example, it can be determined similarly to equations (6) and (7) or (8). The frames with vertical solid and vertical dashed line patterns in Figure 10 correspond to the separated frames, i.e., the frames representing the extracted sound sources 1042 and 1044.
[0144] The combined PASE and tPIT architecture of this example scheme is similar to its Mel counterpart. Figure 10 In this model, the PASE encoder 1020 and PASE decoder 1025 are pre-trained and frozen in the subsequent tPIT stage. For training details, see Santiago Pascual et al. For the PASE decoder 1025, a similar architecture and training method can be used as in the Mel / MelGAN case. Therefore, PASE features can be inverted into the waveform domain via a vocoder (e.g., in a similar way to how MelGAN does for Mel spectrograms).
[0145] Alternatively, the PASE encoder can be pre-trained in the first step. Subsequently, the source separation model (separation level 1040) and the PASE decoder 1025 can be jointly trained in a later step (i.e., the tPIT step). Figure 11 The scheme for this joint training of the tPIT separator 1140 and the PASE decoder 1125 is illustrated schematically.
[0146] This time, only the PASE encoder 1120 was pre-trained, as shown in the upper part of Figure 11. Figure 11 The frames with sparse vertical dashed line patterns represent the PASE features of the data used to train the PASE encoder 1120, i.e., the PASE feature space representation 1122 of the input audio signal 1105. In the tPIT step shown in the lower figure, i.e., during the training of the separation stage 1140, the PASE encoder 1120 is frozen (represented by the dashed box) as a feature extractor. The separation stage 1140 and the PASE decoder 1125 are jointly trained to minimize the loss function of the joint training. Before sending the frames of the separated PASE features 1142, 114 to the PASE decoder 1125, these frames must be reordered 1145 into corresponding output utterances 1146, 1148 based on the optimal tPIT permutation, so that the PASE decoder input of the PASE decoder 1125 will have the same speaker order (speaker assignment) as the true waveforms 1115-1, 1115-2. The loss function of the tPIT step can be derived similarly to Equation (9). Its tPIT contribution can be determined as described above based on frames of separated PASE features 1142, 1144, and frames of PASE features 1116-1, 1116-2 of truth signals 1115-1, 1115-2 generated by the (fixed) PASE encoder 1120 from truth signals 1015-1, 1015-2. For example, the tPIT loss can again be determined similarly to the tPIT loss in the Mel / MelGAN case described above. Figure 11 The frames with vertical solid lines and vertical dashed lines correspond to the separated frames, that is, the frames represented by the extracted sound sources 1142 and 1144.
[0147] Therefore, the PASE decoder 1125 and the separation stage 1140, corresponding to the PASE encoder 1120 used to generate the representation of the mixed audio signal, are jointly trained to minimize the following joint loss function, which includes contributions from the PASE transform (e.g., contributions similar to the contribution GAN_loss in Equation (9)) and contributions related to the difference function (e.g., contributions similar to the contribution tPIT in Equation (9)).
[0148] The separation models in all the example schemes (i) to (iv) above can use a wide range of existing network architectures, such as the Conv-TasNet architecture (e.g., see Yi Luo et al., Conv-TasNet: Surpassing IdealTime-Frequency Magnitude Masking for Speech Separation, https: / / arxiv.org / pdf / 1809.07454.pdf), the DPRNN architecture (e.g., see Yi Luo et al., Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation, https: / / arxiv.org / pdf / 1910.06379.pdf), the WaveNet architecture (e.g., see Aaron van denOord et al., WaveNet: a generative model for raw audio, https: / / arxiv.org / pdf / 1609.03499.pdf), or the Demucus architecture (e.g., see Alexandre Defossez et al., Music sources separation in the waveform domain, (https: / / hal.archives-ouvertes.fr / hal-02379796 / document).
[0149] Next, we will refer to Figure 4 and Figure 12 The clustering level 450 of the deep learning-based system for speech separation is described. It should be noted that Figure 12 is essentially the same as the lower part of Figure 4.
[0150] Since tPIT aims to locally separate sound sources or speakers, it cannot preserve source coherence along the estimation. To preserve source coherence along the estimation, a clustering model (implemented by clustering levels 450 and 1250) is trained to predict the optimal permutation over time. In general, clustering levels 450 and 1250 generate a vector (embedding vector) for each frame of the mixed audio signal, indicating the permutation from the extracted source representation frame to the corresponding source. The permutation can indicate a one-to-one assignment of the extracted frame to the source (label). Thus, depending on the permutation chosen for each frame, clustering levels 450 and 1250 map the frames of the mixed audio signal to a low-dimensional space, the so-called embedding space. The clustering results of clustering levels 450 and 1250 can be used to resolve ambiguities in source permutations.
[0151] Clustering levels 450 and 1250 are arranged to receive input from separation levels 440 and 1240. They take as input frames the representations (est1, est2) of the extracted audio sources 442, 444, 1242, 1244 from separation levels 440 and 1240. During training, clustering levels 450 and 1250 further take as input frames the representations of the mixed audio signals 421 and 1221 and the indications of the permutations (permutation labels) already used for each frame in the tPIT step. Training clustering levels 450 and 1250 involves training a clustering model that generates embedding vectors 453, 455, 1253, 1255 representing the permutation identifier for each frame of the mixed input audio signal. Ideally, frames with the same permutation should be close together in the embedding space, while frames with different permutations should be far apart. In other words, embedding vectors indicating the same permutation should be grouped into corresponding groups or clusters 452, 454, 1252, 1254. To achieve this, pairwise similarity loss can be used to train clustering models 450, 1250, for example...
[0152] (10)
[0153] In Equation (10), each row of W stores the embedding vector of the frame, and each row of V is a one-hot vector of the permutation label of the frame (e.g., [0, 1] or [1, 0] in the case of two speakers), which is estimated by tPIT after training the separation level. The term calculates the cosine distance between the embedding vector of each frame and the embedding vectors of all frames. Typically, pairwise similarity loss can be based on the (cosine) distance between the embedding vector of each frame and the embedding vectors of all frames. Therefore, the dimensions of W and V depend on the temporal resolution of the chosen mixed audio signal representation. If a sentence has T frames, the memory complexity of this loss is O(n log n). During testing, permutation labels are predicted by clustering the embedding vectors, and frames are reordered to form coherent speech by the speaker.
[0154] Therefore, during training, the representation of the mixed audio signal, and for each frame of the mixed audio signal representation, the extracted source representation frame along with an indication of the assigned permutation selected for the corresponding frame of the mixed audio signal representation, are input to clustering levels 450 and 1250. Using this input, clustering levels 450 and 1250 are trained to generate the aforementioned embedding vectors 453, 455, 1253, and 1255 indicating the assigned permutations of the extracted source representation frames to the corresponding source sources, in such a way that the separation between groups 452, 454, 1252, and 1254 of the embedding vectors 453, 455, 1253, and 1255 of the mixed audio signal frames is maximized. It should be understood that the frame embedding vectors 453, 455, 1253, and 1255 are grouped according to the corresponding assigned permutations indicated by these embedding vectors (i.e., the known assigned permutations used by tPIT). The separation (separation criterion) can be an implementation of the loss function used to train clustering level 1250. It can also be called similarity loss.
[0155] The following describes more details of training clustering levels 450 and 1250. As mentioned above, clustering levels 450 and 1250 (or the clustering model) take the output frames from the tPIT level (separation level) and the mixed frames as input. The clustering level further takes known assignment permutations from the tPIT level as input. For example, depending on which domain the tPIT level operates in, these frames can be waveform samples, optimized latent spatial features, Mel features, or PASE features. Since the training of the separation level is tPIT, the separated frames are usually unordered, which means that the speaker identity (represented by the different patterns in frames 1242 and 1244 in Figure 12) may vary throughout the input stream.
[0156] However, the optimal permutation labels estimated by the tPIT-level (separation) model during training are known. These labels can now be used to train a clustering model to output the embedding vector for each frame. The training objective is designed such that the embeddings (embedding vectors) of frames with the same permutation label in the utterance should be close to each other, while the embeddings (embedding vectors) of frames with different permutations should be far apart.
[0157] If the separation model (separation level) operates in the temporal domain or the latent space of optimization (example schemes (i) and (ii) above), a high temporal resolution is required to perform separation and phase modeling simultaneously at the tPIT level, which results in a large number of frames. In this case, the pairwise similarity loss of Equation (10) can become very (memory) expensive when training the clustering model. This problem can be addressed by training the clustering model using a reformulated version of the generalized end-to-end (GE2E) loss (see, for example, Li Wan et al., Generalized end-to-end loss for speaker verification, https: / / arxiv.org / pdf / 1710.10467.pdf). To compute the reformulated GE2E loss, for each frame of the mixed audio signal (representation), the squared Euclidean distance between the embedding vector of that frame and the average embedding (centroid) of its own cluster, and the distance vector between the embedding vector of that frame and the average embedding of other clusters are computed. In this sense, training the clustering level to minimize the reformulated GE2E loss corresponds to optimizing the separation criterion for each frame of the mixed audio signal (representation). Typically, the separation criterion can be based on the (squared) Euclidean distance between vectors (embedded vectors) and / or groups of vectors (embedded vectors).
[0158] Each cluster consists of embedding vectors with the same permutation label (assignment permutation). As mentioned above, during training, the permutation labels estimated by the tPIT level (separation level) are known, and these known permutation labels can be used to define which frames belong to each cluster. In other words, the known permutation labels can be used to group the embedding vectors of frames from the mixed audio signal into different clusters, one cluster per permutation label. In other words, different clusters of embedding vectors are defined based on the permutation labels. The clustering level is then trained so that the defined clusters are sufficiently separated from each other.
[0159] Consistent with the above, training is performed by maximizing the probability that an embedding vector belongs to its own cluster (i.e., by optimizing the separation criterion described above). The probability is defined considering the fundamental distance relative to the centroid of the cluster. For example, for the embedding vector e of the j-th frame in the i-th cluster (i.e., the cluster with permutation label i), ij Training clusters may involve maximizing the following:
[0160] (11)
[0161] Where P is the total number of permutations, and c k is the mean (centroid) of the k-th cluster, and d is the squared Euclidean distance. For example, the squared Euclidean distance can be defined as:
[0162] (12)
[0163] Where w and b are optional parameters: scaling and offset. For example, parameters w and b can be learnable parameters.
[0164] Therefore, typically, for a given frame of a mixed audio signal (representation), the separation criterion can be based on the (squared) Euclidean distance between the vector (embedded vector) indicating the assignment permutation of that frame and the vectors (embedded vectors) of other frames in the mixed audio signal representation. Here, the distance between a vector and a vector group (cluster) can be defined as the distance between the vector and the centroid (mean) of that group. The centroid of a vector group can be the mean of that group. For example, the separation distance can be based on the Euclidean distance or squared Euclidean distance between the frame's vector and (previously determined) vector groups, where each group corresponds to one of the possible assignment permutations and (only) includes all those vectors indicating each assignment permutation. Optimizing the separation criterion might correspond to maximizing the probability of equation (11).
[0165] It should be noted that this disclosure proposes to adjust or reformulate the GE2E loss for speech source separation. Contrary to common practice, this disclosure uses squared Euclidean distance instead of cosine distance to define the loss. Euclidean distance significantly improves performance and robustness. Furthermore, contrary to the conventional use of GE2E loss, in equation (11)... This indicates "permutation embedding" rather than "speaker embedding".
[0166] Using the separation criterion defined above can significantly reduce computational complexity. Specifically, when using the separation criterion described above (instead of the similarity loss of Equation (10)), it can be seen that only the distance P needs to be calculated for each frame, instead of the distance T, where P is the total number of permutations (e.g., N! for N sources in a mixed audio signal) and T is the total temporal resolution of the signal. The value of P is typically relatively small (e.g., 2 in the case of two speakers). Therefore, the separation criterion (e.g., the reformulated GE2E loss) provides linear memory complexity O(T), which is crucial for making training feasible on a GPU.
[0167] On the other hand, if the input frames of the clustering model are Mel or PASE features (or other vocoder feature sets), these features already have relatively low temporal resolution. Therefore, in these cases, pairwise losses (e.g., similarity loss as defined in Equation (10)) or proposed separation criteria (e.g., modified GE2E loss as defined in Equation (11)) can be used to train the clustering model.
[0168] Therefore, by using the proposed separation criteria (e.g., modified GE2E loss) or employing low temporal resolution features, this disclosure helps reduce the computational complexity of waveform-based source separation models trained using frame-level permutation-invariant training. This can be seen as an important step towards developing speech source separation models faster and easier.
[0169] During separation, when attempting to isolate the speaker (or, in general, the sound source) from a given mixture (mixed audio signal), this mixed waveform is first passed through a selected signal transform and then through a separation model to obtain separated but unordered frames. A clustering model then generates a sequence of permutation embeddings (e.g., a sequence of embedding vectors), one permutation embedding per frame. Next, an unsupervised clustering method (e.g., K-means, Gaussian Mixture Model (GMM), Hidden Markov Model (HMM), etc.) is applied to the permutation embedding sequence to separate it into clusters (or groups), one cluster per permutation. Through this clustering operation, the permutations used for each frame can be determined (e.g., predicted or estimated) (assigned permutations). The extracted frames are then reorganized according to the determined permutation labels so that each output stream will contain only one speaker. Finally, an inverse transform converts each feature sequence back into a waveform.
[0170] The following will refer to Figure 13 The flowchart describes a corresponding example of method 1300 for sound source separation (e.g., speaker separation) using the aforementioned deep learning-based system. As described above, the system includes a separation level and a clustering level. The separation level extracts the representation of the sound sources frame by frame from the representation of the audio signal. The clustering level generates a vector (embedding vector) for each frame, which indicates the permutation of the extracted sound source representation from the frame to the corresponding sound source. It should be noted that the representation of the audio signal is based on a waveform representation.
[0171] Step S1310 of method 1300 is a step of generating a representation of the mixed audio signal by applying a waveform-based transformation operation to the mixed audio signal. The mixed audio signal includes at least two sound sources.
[0172] Step S1320 is the step of inputting a representation of the mixed audio signal into the separation stage to generate a frame of the extracted sound source representation.
[0173] Step S1330 is to input the extracted sound source representation frames into the clustering level to generate a vector for each frame of the mixed audio signal, the vector indicating the allocation permutation of the extracted sound source representation frames to the corresponding sound sources selected for the corresponding frames of the mixed audio signal.
[0174] Step S1340 involves clustering the vectors generated at the cluster level according to a clustering algorithm, and determining the corresponding assignment permutation (estimated or predicted) selected for the frames of the mixed audio signal based on the vectors' affiliation to different clusters. Each cluster represents a specific assignment permutation.
[0175] Step S1350 is the step of assigning the extracted representation frames to a set of audio streams according to the determined assignment permutation.
[0176] Finally, step S1360 is the step of applying the inverse of the waveform-based transformation to the set of audio streams to generate a waveform audio stream of the extracted sound source.
[0177] Example computing device
[0178] Methods for training deep learning-based systems for sound source separation and methods for performing sound source separation using such systems have been described. Furthermore, this disclosure also relates to apparatus for performing these methods. An example of such apparatus 1400 is described in... Figure 14 The diagram is schematically shown. Apparatus 1400 may include a processor 1410 (e.g., a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), one or more application-specific integrated circuits (ASICs), one or more radio frequency integrated circuits (RFICs), or any combination thereof), and a memory 1420 coupled to the processor 1410. The processor may be adapted to perform some or all of the steps of the methods described throughout this disclosure. To perform a method for training a deep learning-based system, apparatus 1400 may, for example, receive a mixed audio signal 10 and a reference audio signal 15 as input. In this case, apparatus 1400 may output parameters 1430 (e.g., parameters of a deep neural network) for setting separation, clustering, and / or (inverse)transformation stages. To implement a method for performing actual source separation, apparatus 1400 may, for example, receive a mixed audio signal 10 as input. In this case, apparatus 1400 may output audio signals of the separated sound sources (e.g., speakers).
[0179] Device 1400 may be a server computer, client computer, personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, smartphone, network device, network router, switch, or bridge, or any machine capable of executing instructions (sequence or other) specifying the actions to be taken by the device. Furthermore, although Figure 14 Only a single device 1400 is illustrated herein, but this disclosure should relate to any collection of devices, individually or in combination, that execute instructions to perform any one or more of the methods discussed herein.
[0180] This disclosure further relates to a program (e.g., a computer program) that includes instructions which, when executed by a processor, cause the processor to perform some or all of the steps of the methods described herein.
[0181] Furthermore, this disclosure relates to a computer-readable (or machine-readable) storage medium for storing the aforementioned program. Here, the term "computer-readable storage medium" includes, but is not limited to, data storage repositories in the form of, for example, solid-state memory, optical media, and magnetic media.
[0182] Other configuration considerations
[0183] Unless otherwise specifically stated, it will be apparent from the following discussion that throughout this disclosure, the use of terms such as “processing,” “computing,” “operation,” “determining,” “analyzing,” etc., refers to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulate data represented as physical quantities (e.g., electronic quantities) and / or transform them into other data similarly represented as physical quantities.
[0184] In a similar manner, the term "processor" can refer to any device or part of a device that processes electronic data, for example, from registers and / or memory, to transform that electronic data into other electronic data, for example, that can be stored in registers and / or memory. "Computer," "computing machine," or "computing platform" may include one or more processors.
[0185] In one exemplary embodiment, the methods described herein can be executed by one or more processors that accept computer-readable (also known as machine-readable) code containing an instruction set that, when executed by the one or more processors, implements at least one of the methods described herein. This includes any processor capable of executing an instruction set (sequential or otherwise) specifying the operation to be performed. Thus, one example is a typical processing system including one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may also include a storage subsystem comprising main RAM and / or static RAM and / or ROM. A bus subsystem may be included for communication between components. The processing system may also be a distributed processing system with processors coupled via a network. If the processing system requires a display, it may include such a display, such as a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data input is required, the processing system also includes one or more input devices, such as an alphanumeric input unit (e.g., a keyboard), a pointing control device (e.g., a mouse), etc. The processing system may also include a storage system, such as a disk drive unit. In some configurations, the processing system may include a sound output device and a network interface device. Therefore, the storage subsystem includes a computer-readable carrier medium carrying computer-readable code (e.g., software) comprising an instruction set that, when executed by one or more processors, causes the execution of one or more of the methods described herein. It should be noted that when a method comprises several elements, such as several steps, the order of these elements is not implied unless specifically stated otherwise. The software may reside on a hard disk or may reside wholly or at least partially in RAM and / or a processor during execution by a computer system. Therefore, memory and processor also constitute a computer-readable carrier medium carrying computer-readable code. Furthermore, the computer-readable carrier medium may be formed or included in a computer program product.
[0186] In alternative example embodiments, one or more processors may operate as isolated devices or may be connected, for example, networked to one or more other processors. In a networked deployment, one or more processors may operate as a server or user machine in a server-user network environment, or as a peer-to-peer machine in a peer-to-peer or distributed network environment. One or more processors may form a personal computer (PC), tablet computer, personal digital assistant (PDA), cellular phone, network device, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) specifying the actions to be taken by the machine.
[0187] It should be noted that the term "machine" should also be considered as any collection of machines that individually or jointly execute one or more sets of instructions to perform any one or more of the methods discussed herein.
[0188] Therefore, an example embodiment of each method described herein is in the form of a computer-readable carrier medium carrying a set of instructions, such as a computer program for execution on one or more processors (e.g., one or more processors as part of a network server arrangement). Thus, as those skilled in the art will understand, example embodiments of this disclosure can be embodied as methods, apparatus such as dedicated devices, apparatus such as data processing systems, or computer-readable carrier media, such as computer program products. A computer-readable carrier medium carries computer-readable code comprising a set of instructions that, when executed on one or more processors, cause one or more processors to implement the method. Therefore, aspects of this disclosure can take the form of methods, entirely hardware example embodiments, entirely software example embodiments, or example embodiments combining software and hardware aspects. Furthermore, this disclosure can take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.
[0189] The software can also be transmitted or received over a network via a network interface device. While the carrier medium is a single medium in one example embodiment, the term "carrier medium" should be considered to include a single or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "carrier medium" should also be considered to include any medium capable of storing, encoding, or carrying a set of instructions for execution by one or more processors and causing one or more processors to perform one or more of the methods of this disclosure. The carrier medium can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical discs, magnetic disks, and magneto-optical discs. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wires, and optical fibers, including wiring that constitutes a bus subsystem. Transmission media can also take the form of acoustic or optical waves, such as those generated during radio wave and infrared data communication. For example, the term "carrier medium" should be understood accordingly to include, but is not limited to, solid-state storage; computer products embodied in optical and magnetic media; media carrying propagation signals that can be detected by at least one or more processors and represent a set of instructions that implement a method when executed; and transmission media in networks carrying propagation signals that can be detected by at least one of one or more processors and represent the set of instructions.
[0190] It should be understood that, in one example embodiment, the steps of the discussed method are performed by one or more suitable processors of a processing (e.g., a computer) system for executing instructions (computer-readable code) stored in memory. It will also be understood that this disclosure is not limited to any particular implementation or programming technique, and that this disclosure can be implemented using any suitable technique for implementing the functionality described herein. This disclosure is not limited to any particular programming language or operating system.
[0191] Throughout this disclosure, references to “one example embodiment,” “some example embodiments,” or “example embodiment” mean that a particular feature, structure, or characteristic described in connection with an example embodiment is included in at least one example embodiment of this disclosure. Therefore, the phrases “in one example embodiment,” “in some example embodiments,” or “in an example embodiment” appearing throughout this disclosure do not necessarily refer to the same example embodiment. Furthermore, as will be apparent to those skilled in the art from this disclosure, in one or more exemplary embodiments, a particular feature, structure, or characteristic may be combined in any suitable manner.
[0192] As used herein, unless otherwise stated, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to describe common objects merely indicates different instances of the similar objects being referred to, and is not intended to imply that the objects described in this way must be in a given sequence, whether in time, space, rank, or any other way.
[0193] In the appended claims and the description herein, the terms *including*, *consisting of*, or *comprise* are open-ended terms indicating at least the inclusion of the following element / feature, but not excluding other elements / features. Therefore, when used in the claims, the term *including* should not be construed as limiting oneself to the means, elements, or steps listed thereafter.
[0194] For example, the scope of the statement "the device includes A and B" should not be limited to a device consisting only of elements A and B. The terms "including" or "including" as used herein are also open-ended terms, meaning that at least the element / feature following the term is included, but other elements / features are not excluded. Therefore, "including" and "containing" are synonyms, meaning to include.
[0195] It should be understood that in the foregoing description of exemplary embodiments of this disclosure, various features of this disclosure are sometimes combined in a single example embodiment, figure, or description thereof for the purpose of simplifying this disclosure and aiding in the understanding of one or more of the various aspects of the invention. However, this approach to disclosure should not be construed as reflecting an intention that the claims require more features than are expressly recited in the claims. Rather, as reflected in the appended claims, the inventive aspect lies in fewer than all features of a single foregoing example embodiment. Therefore, the claims following this specification are hereby expressly incorporated, each claim serving as a separate example embodiment of this disclosure.
[0196] Furthermore, while some of the exemplary embodiments described herein include some features included in other exemplary embodiments but not others, combinations of features from different exemplary embodiments are intended to be within the scope of this disclosure and form different exemplary embodiments, as those skilled in the art should understand. For example, any claimed exemplary embodiment may be used in any combination as described in the appended claims.
[0197] Numerous specific details are set forth in the description provided herein. However, it should be understood that exemplary embodiments of this disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0198] Therefore, while the content considered to be the best mode of this disclosure has been described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of this disclosure, and it is intended that all such changes and modifications fall within the scope of this disclosure. For example, any formulas given above represent only usable processes. Functions can be added or removed from the block diagram, and operations can be interchanged between functional blocks. Steps can be added or removed from the methods described within the scope of this disclosure.
[0199] Various aspects of this disclosure can be understood from the following exemplary embodiments (EEE):
[0200] EEE 1. A method for training a deep learning-based system for sound source separation, wherein the system includes a separation level for extracting sound source representations frame-by-frame from a representation of an audio signal, and a clustering level for generating, for each frame, an estimate of the frame-to-frame assignment permutation of the extracted sound source representations in possible assignment permutations, wherein the audio signal representation is based on a waveform representation, the method comprising:
[0201] As input, a representation of a mixed audio signal and representations of at least two reference audio signals are obtained, wherein the representations are waveform-based, wherein the mixed audio signal includes at least two sound sources, and wherein the reference audio signals correspond to the respective sound sources included in the mixed audio signal;
[0202] A representation of a mixed audio signal and representations of at least two reference audio signals are input into a separation stage, and the separation stage is trained to extract the representation of the sound sources from the representation of the mixed audio signal in such a way that, for each frame of the representation of the mixed audio signal, a difference function is minimized, wherein the difference function is based on the difference between the frame of the extracted sound source representation and the frame of the reference audio signal representation, wherein the pairs of frames of the extracted sound source representation and frames of the reference audio signal representation are selected based on one of the possible assignment permutations to obtain this difference, and wherein, for each frame, in order to compute the difference function, this assignment permutation of the extracted sound source representation and the reference audio signal representation is selected to result in the minimum difference function.
[0203] EEE 2. Following the method of EEE 1, where the clustering level generates an estimated vector for each frame indicating the corresponding assigned permutation; and
[0204] This method also includes:
[0205] The input is a representation of the mixed audio signal, and for each frame of the mixed audio signal representation, the input is the frame of the extracted source representation and the indication of the assigned permutation that has been selected for the corresponding frame of the mixed audio signal representation to the cluster level, and the cluster level is trained to generate vectors indicating the assigned permutations of the extracted source representation to the corresponding source in such a way that the separation between the vector groups of the frames of the mixed audio signal is maximized, wherein the vectors of the frames are grouped according to the corresponding assigned permutations indicated by these vectors.
[0206] EEE 3. The method according to EEE 1 or 2, wherein the difference function indicates a combination of differences between frames of the extracted source representation and frames of the representation of the reference audio signal, wherein for each extracted source representation, the combination of differences includes the difference between frames of the extracted source representation and the corresponding frames of the representation of the reference audio signal.
[0207] EEE 4. According to EEE 2 or EEE 3 which depends on EEE 2,
[0208] The clustering level is trained such that the separation criteria are optimized for each frame of the mixed audio signal representation; and
[0209] The separation criterion is based on the Euclidean distance between vectors and / or groups of vectors.
[0210] EEE 5. Based on the method of EEE 2 or any of EEE 3 or 4 that depends on EEE 2,
[0211] The clustering level is trained such that the separation criteria are optimized for each frame of the mixed audio signal representation; and
[0212] For a given frame in the representation of the mixed audio signal, the separation criterion is based on the Euclidean distance between the vector indicating the assignment permutation of that frame and the vector group of other frames in the representation of the mixed audio signal.
[0213] EEE 6. According to the method of EEE 4 or 5, where the optimization separation criterion corresponds to maximizing the following for a given frame of the representation of the mixed audio signal:
[0214]
[0215] Where e i It is a vector of a given frame, c k Let be the centroid of the vector group of the k-th distribution permutation, P be the total number of distribution permutations, and d(-,-) be the squared Euclidean distance.
[0216] 7. According to the method described in any of the preceding EEE methods,
[0217] The system also includes a transform stage for converting the mixed audio signals into a representation of the mixed audio signals; and
[0218] The mixed audio signal is transformed into a signal space, which is a waveform-based signal space.
[0219] EEE 8. The method according to any one of EEE 1 to 6,
[0220] The system further includes a transformation stage for transforming the mixed audio signals into a representation of the mixed audio signals; and
[0221] The transformation mentioned above involves at least one of the following:
[0222] The mixed audio signal is divided into multiple frames in the time domain;
[0223] Deep learning-based coding is used to project mixed audio signals into a latent feature space optimized for sound source separation;
[0224] Mel space encoding; and
[0225] Problem-agnostic speech coding based on deep learning.
[0226] EEE 9. The method described according to any of the foregoing EEE,
[0227] The representation of the mixed audio signal involves dividing the mixed audio signal into waveform frames;
[0228] The separation stage is trained to determine the frames from which the sound sources are extracted from the frames of the mixed audio signal for each frame of the mixed audio signal in a manner that minimizes the following loss function.
[0229] ,
[0230] Where t indicates the frame, l indicates the number of samples within the frame, L is the total number of samples within the frame, n is the label of the extracted sound source, N is the total number of extracted sound sources, est represents the frame of the extracted sound source, and ref represents the frame of the reference audio signal. It is a permutation mapping for labels n=1, ..., N, and indicates the label of the reference audio signal; and
[0231] For each frame, the permutation map It was chosen as the function that produces the minimum loss.
[0232] EEE 10. The method according to any one of EEE 1 to 8,
[0233] The representation of the mixed audio signal involves the latent feature space representation of the mixed audio signal that can be generated by a pre-trained deep learning-based encoder;
[0234] The separation stage is trained to determine the extracted sound source frame from the frames of the mixed audio signal for each frame of the mixed audio signal in a manner that minimizes the following loss function.
[0235] ,
[0236] Where t indicates the frame, f indicates the features in the latent feature space, n indicates the labels of the extracted sound sources and the reference audio signal, and N is the total number of extracted sound sources. V indicates the frame representing the extracted sound source, and V indicates the frame representing the reference audio signal. It is a permutation mapping for labels n=1, ..., N, and indicates the label of the reference audio signal; and
[0237] For each frame, the permutation map It was chosen as the function that produces the minimum loss.
[0238] EEE 11. The method according to any one of EEE 1 to 8,
[0239] The representation of the mixed audio signal involves a Mel feature space representation that can be generated from the mixed audio signal by a Mel encoder; and
[0240] The corresponding Mel decoder and separator are jointly trained to minimize the joint loss function, which includes contributions from the Mel transform and contributions related to the difference function.
[0241] EEE 12. According to any one of EEE 1 to 8,
[0242] The representation of mixed audio signals involves a problem-agnostic speech encoder, PASE, which can be represented from the feature space generated by a deep learning-based PASE encoder; and
[0243] The separation stage and the PASE decoder corresponding to the PASE encoder are jointly trained to minimize the joint loss function, which includes contributions from the PASE transform and contributions related to the difference function.
[0244] EEE 13. The method according to any of the preceding EEEs, wherein the sound source relates to a speech source.
[0245] EEE 14. The method according to any of the preceding EEEs, wherein the system's separation level is based on one of the following:
[0246] Conv-TasNet architecture;
[0247] DPRNN architecture;
[0248] WaveNet architecture; and
[0249] Demucs architecture.
[0250] EEE 15. A sound source separation method using a deep learning-based system, wherein the system includes a separation level for extracting sound source representations frame-by-frame from a representation of an audio signal, and a clustering level for generating, for each frame, an assignment permutation of the extracted sound source representations to the corresponding sound sources, wherein the audio signal representation is based on a waveform representation, the method comprising:
[0251] A representation of the mixed audio signal is generated by applying a waveform-based transform to the mixed audio signal, wherein the representation is based on a waveform representation and wherein the mixed audio signal includes at least two sound sources;
[0252] The representation of the mixed audio signal is input into the separation stage to generate frames of the extracted sound source representation;
[0253] The extracted sound source representation frames are input into the clustering level to generate a vector for each frame of the mixed audio signal, which indicates the assignment permutation of the extracted sound source representation frames to the corresponding sound sources selected for the corresponding frames of the mixed audio signal.
[0254] The vectors generated at the cluster level are clustered according to the clustering algorithm, and the corresponding assignment permutation for the selected frames of the mixed audio signal is determined according to the vector's affiliation to different clusters. Each cluster represents a specific assignment permutation.
[0255] The extracted representation frames are assigned to the audio stream set according to the determined assignment permutation; and
[0256] The inverse of the waveform-based transformation is applied to the set of audio streams to generate a waveform audio stream of the extracted sound source.
[0257] EEE 16. A program comprising instructions that, when executed by a processor, cause the processor to perform all the steps of a method according to any of the preceding EEEs.
[0258] EE 17. A computer-readable storage medium for storing a program according to EE 16.
[0259] EEE 18. An apparatus comprising a processor and a memory coupled to the processor, wherein the processor is adapted to perform all steps of the method according to any one of EEE 1 to 15.
Claims
1. A computer-implemented method of training a deep learning based system for sound source separation, wherein the system comprises a deep learning based separation stage for frame-wise extracting representations of sound sources from a representation of an audio signal, and a clustering stage for, for each frame, generating an estimate of a permutation of the extracted representations of sound sources in a candidate permutation into a respective permutation of sound sources, wherein the representation of the audio signal is a waveform-based representation, and wherein the waveform-based representation is a time-domain representation suitable for a waveform-based model or a transformation of the time-domain representation different from a time-frequency transformation, the method comprising: obtaining, as input, a representation of a mixed audio signal and representations of at least two reference audio signals, wherein the representations are waveform-based representations, and wherein the waveform-based representations are time-domain representations suitable for a waveform-based model or a transformation of the time-domain representations different from a time-frequency transformation, wherein the mixed audio signal comprises at least two sound sources, and wherein the reference audio signals correspond to respective ones of the sound sources comprised in the mixed audio signal; and inputting the representation of the mixed audio signal and the representations of the at least two reference audio signals to the separation stage, and training the separation stage to extract representations of sound sources from the representation of the mixed audio signal in such a way that, for each frame of the representation of the mixed audio signal, a difference function is minimized, wherein the difference function is based on a difference between a frame of the extracted representations of sound sources and a frame of the representations of the reference audio signals, wherein the pair of the frame of the extracted representations of sound sources and the frame of the representations of the reference audio signals is selected based on one of the candidate permutations in order to take the difference, and wherein, for each frame, the permutation of the extracted representations of sound sources and the representations of the reference audio signals is selected that results in the smallest difference function for computing the difference function, wherein the clustering stage generates, for each frame, a vector indicative of the estimate of the respective permutation, and wherein the method further comprises: inputting the representation of the mixed audio signal, and, for each frame of the representation of the mixed audio signal, a frame of the extracted representations of sound sources and an indication of the permutation that has been selected for the respective frame of the representation of the mixed audio signal to the clustering stage, and training the clustering stage to generate the vector indicative of the permutation of the frame of the extracted representations of sound sources into the respective sound sources in such a way that a separation between groups of vectors of frames of the mixed audio signal is maximized, wherein the vectors of frames are grouped according to the respective permutations indicated by these vectors, wherein the clustering stage is trained such that a separation criterion is optimized for each frame of the representation of the mixed audio signal; and wherein the separation criterion is based on a Euclidean distance between the vectors and / or between groups of vectors.
2. The method according to claim 1, wherein the difference function is indicative of a combination of differences between the frames of the extracted representations of sound sources and the frames of the representations of the reference audio signals, wherein the combination of differences comprises, for each extracted representation of a sound source, a difference between the frame of the extracted representation of a sound source and the respective frame of the representations of the reference audio signals.
3. The method of claim 1, wherein, For a given frame of the mixed audio signal representation, the separation criterion is based on the Euclidean distance between a vector indicative of the separation configuration for that frame and a set of vectors of other frames of the representation of the mixed audio signal.
4. The method of claim 1, wherein the optimization separation criterion corresponds to maximizing, for a given frame of a representation of the mixed audio signal: where e i is the vector of a given frame, c k is the centroid of the vector group of the k-th partition, P is the total number of partitions, and d(-, -) is the squared Euclidean distance.
5. The method of claim 1, wherein the system further comprises a transformation stage for transforming the mixed audio signal into a representation of the mixed audio signal; and wherein the mixed audio signal is transformed into a signal space, which is a waveform-based signal space different from the frequency domain.
6. The method of claim 1, wherein the system further comprising a transformation stage for transforming the mixed audio signal into a representation of the mixed audio signal; and wherein the transformation involves at least one of: segmenting the mixed audio signal into a plurality of frames in the time domain; a deep learning based encoding for projecting the mixed audio signal into a latent feature space optimized for sound source separation; a Mel space encoding; and a deep learning based problem agnostic speech encoding.
7. The method of claim 1, wherein the representation of the mixed audio signal involves segmenting the mixed audio signal into waveform frames; wherein the separation stage is trained to determine, for each frame of the mixed audio signal, a frame of the extracted sound source from the frame of the mixed audio signal in a way that minimizes a loss function, 8. The method of claim 1, , where t indicates the frame, I the number of intra-frame samples, L the total number of intra-frame samples, n indicates the label of the extracted sound source, N the total number of extracted sound sources, est denotes the frame of the extracted sound source, ref denotes the frame of the reference audio signal, is a permutation mapping for labels n = 1,..., N and indicates the label of the reference audio signal; and wherein, For each frame, the permutation map is chosen to produce the smallest loss function. wherein the representation of the mixed audio signal involves a latent feature space representation of the mixed audio signal that can be generated by a pre-trained deep learning based encoder; wherein the separation stage is trained to determine, for each frame of the mixed audio signal, a frame of the extracted sound source from the frame of the mixed audio signal in a way that minimizes a loss function, 9. The method of any one of claims 1 to 7, , where t indicates frames, f indicates features within a latent feature space, n indicates labels of extracted sound sources and a reference audio signal, N is a total number of extracted sound sources, frames indicating representations of extracted sound sources, V indicates frames indicating representations of a reference audio signal, is a permutation mapping for labels n = 1,..., N and indicates labels of a reference audio signal; and wherein For each frame, the permutation map is chosen to produce the smallest loss function. wherein the representation of the mixed audio signal involves a Mel feature space representation that can be generated from the mixed audio signal by a Mel encoder; and a corresponding Mel decoder and the separation stage are jointly trained to minimize a joint loss function comprising a contribution from the Mel transformation and a contribution related to a difference function. wherein 10. The method of any one of claims 1 to 7, wherein the representation of the mixed audio signal involves a feature space representation that can be generated from the mixed audio signal by a deep learning based PASE encoder; and the separation stage and a corresponding PASE decoder of the PASE encoder are jointly trained to minimize a joint loss function comprising a contribution from the PASE transformation and a contribution related to a difference function. wherein 11. The method of any one of claims 1 to 7, wherein the sound source involves a speech source.
12. The method of any one of claims 1 to 7, wherein the separation stage of the system is based on one of: a Conv-TasNet architecture; a DPRNN architecture; a WaveNet architecture; and a Demucs architecture. 13. A computer-implemented method for sound source separation using a deep learning based system, wherein the system comprises a deep learning based separation stage for frame-wise extracting representations of sound sources from a representation of an audio signal, and a clustering stage for, for each frame, generating a vector indicating a frame-wise mapping of the extracted representations of sound sources to respective sound sources, wherein the representation of the audio signal is a waveform-based representation, and wherein the waveform-based representation is a time-domain representation or a transformation of the time-domain representation different from a time-frequency transformation, which is suitable for a waveform-based model, the method comprising: generating a representation of a mixed audio signal by applying a waveform-based transformation operation to the mixed audio signal, wherein the representation is a waveform-based representation, and wherein the waveform-based representation is a time-domain representation and the waveform-based transformation is a transformation different from a time-frequency transformation, which is suitable for a waveform-based model, and wherein the mixed audio signal comprises at least two sound sources; inputting the representation of the mixed audio signal to the separation stage to generate frames of extracted representations of sound sources; inputting the frames of extracted representations of sound sources to the clustering stage to generate, for each frame of the mixed audio signal, a vector indicating a frame-wise mapping of the extracted representations of sound sources to respective sound sources selected for the respective frame of the mixed audio signal; clustering the vectors generated by the clustering stage according to a clustering algorithm and determining the respective mappings selected for the frames of the mixed audio signal based on the attribution of the vectors to different clusters, each cluster representing a particular mapping; assigning the frames of extracted representations to a set of audio streams according to the determined mappings; and applying an inverse of the waveform-based transformation to the set of audio streams to generate waveform audio streams of the extracted sound sources.
14. A computer-readable storage medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 13.
15. A computer program product having instructions which, when executed by a computing device or system, cause the computing device or system to perform the method according to any one of claims 1 to 13.
16. An apparatus for training a deep learning based system for sound source separation, comprising means for performing the method according to any one of claims 1 to 12.
17. An apparatus for sound source separation using a deep learning based system, comprising means for performing the method according to claim 13.
18. An apparatus for training a deep learning based system for sound source separation, comprising a processor and a memory coupled to the processor, wherein the memory has stored instructions, wherein the processor is adapted to execute the instructions to carry out the method according to any one of claims 1 to 12.
19. An apparatus for sound source separation using a deep learning based system, comprising a processor and a memory coupled to the processor, wherein the memory has stored instructions, wherein the processor is adapted to execute the instructions to carry out the method according to claim 13.
20. An audio processing system comprising a deep learning based separation stage for frame-wise extracting representations of sound sources from a representation of an audio signal, and a clustering stage for generating, for each frame, an estimate of a frame-wise separation of the extracted representations of sound sources into respective sound sources among candidate separation configurations, wherein the representation of the audio signal is a waveform-based representation, and wherein the waveform-based representation is a time-domain representation suitable for a waveform-based model or a transformation of the time-domain representation different from a time-frequency transformation, wherein the separation stage and the clustering stage are trained by the method of any one of claims 1-12.
21. An audio processing system comprising a deep learning based separation stage for frame-wise extracting representations of sound sources from a representation of an audio signal, and a clustering stage for generating, for each frame, a vector indicating a frame-wise separation of the extracted representations of sound sources into respective sound sources, wherein the representation of the audio signal is a waveform-based representation, and wherein the waveform-based representation is a time-domain representation suitable for a waveform-based model or a transformation of the time-domain representation different from a time-frequency transformation, wherein the separation stage and the clustering stage implement the sound source separation of the method of claim 13.
Citation Information
Patent Citations
Permutation invariant training for talker-independent multi-talker speech separation
CN109313910A
Multi-speaker voice separation method based on convolutional neural network and depth clustering
CN110459240A