Method of adversarial training for universal sound separation

Through the adversarial training method, the separator is guided to separate multiple heterogeneous sound sources in any audio mix using context-based and instance-based discriminators, solving the problem of difficulty in adapting to multiple unknown sound sources in the prior art, and achieving efficient general sound separation.

CN120077436APending Publication Date: 2025-05-30DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073724.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-27
Filing Date
2023-09-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively separate multiple heterogeneous sound sources in any audio mix, especially in general sound separation tasks. Traditional methods often assume that a specific sound source exists and cannot adapt to the separation of multiple unknown sound sources.

Method used

Adversarial training methods are adopted to train context-based discriminators and instance-based discriminators, and the training of the separator is guided through adversarial loss prompts, allowing it to separate multiple heterogeneous sound sources in any audio mix.

Benefits of technology

The ability to efficiently separate multiple heterogeneous sound sources in any audio mix is ​​realized, the performance and flexibility of general sound separation is improved, and it can adapt to the separation tasks of multiple unknown sound sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077436A_ABST
    Figure CN120077436A_ABST
Patent Text Reader

Abstract

According to an aspect of the present disclosure, there is provided a method of adversarial training of a splitter (30) for generic sound separation of an audio mix m for any sound source sk = 1,..., K, the method comprising: training a context-based discriminator (34) configured to provide a context-based loss prompt based on consideration of a set of input separated sound sources; and training the separator (30) to minimize the loss according to the context-based loss cues provided by the context-based discriminator (34); wherein training the context-based discriminator (34) comprises maximizing loss based on a set of real sound sources and a false set of separated sound sources, where the false set of separated sound sources is ranked to match the order of the set of real sound sources, and where the false set of separated sound sources is ranked to match the order of the set of real sound sources, and where the false set of separated sound sources is ranked to match the order of the set of real sound sources. The false set of separated sound sources includes sources corresponding to the separated sound sources estimated by the separator (30) and further includes one or more real sound sources of the set of real sound sources.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This patent application claims the benefit of Spanish Patent Application No. P202230890, filed on October 17, 2022, US Provisional Application No. 63 / 440,568, filed on January 23, 2023, and US Provisional Application No. 63 / 498,794, filed on April 27, 2023, the entire contents of each of which are incorporated herein by reference. Technical Field

[0003] In one aspect, the present invention relates to a method for adversarially training a separator for general sound separation. Background Art

[0004] The problem of source separation is to separate the sources present in an audio mixture. For example, music source separation consists of extracting vocals, bass, and drums from a music mixture, while speech source separation consists of separating each speaker from a mixture produced by several speakers speaking simultaneously. An important characteristic of Western pop music mixtures is that certain musical instruments (e.g., vocals, bass, and drums) consistently appear in the song. For this reason, most music source separation methods assume that these instruments are always present in the mixture and separate vocals, bass, drums, and "other sources", where "other sources" refers to any other source in the mixture other than vocals, bass, or drums. Thus, most music source separation models are source-specific. This contrasts with speech source separation, where the speakers present in the mixture are not known in advance. If it cannot be assumed that it is known in advance which speakers are to be separated, then most speech source separation models are speaker-independent. Summary of the Invention

[0005] In a first aspect, there is provided a method for adversarially training a separator for general sound separation of an audio mixture m for any sound source s k=1,…,K The method includes: training a context-based discriminator configured to provide a context-based loss cue based on consideration of a set of separated sound sources of input; and training the separator to minimize a loss according to the context-based loss cue provided by the context-based discriminator. Training the context-based discriminator includes maximizing a loss based on a set of true sound sources and a false set of separated sound sources, where the false set of separated sound sources is sorted to match the order of the set of true sound sources, and where the false set of separated sound sources includes sources corresponding to the separated sound sources estimated by the separator and further includes one or more true sound sources from the set of true sound sources.

[0006] In some embodiments, a spurious set of separated sound sources is obtained by: obtaining a set of separated sound sources corresponding to a set of estimated separated sound sources estimated by the separator from the audio mix m; permuting the set of separated sound sources such that the order of the permuted set of separated sound sources matches the order of the set of true sound sources; and replacing one or more of the separated sound sources in the permuted set of separated sound sources with one or more true sound sources from the set of true sound sources to obtain a spurious set of separated sound sources. Corresponding to the set of separated sound sources; permuting the set of separated sound sources such that the order of the permuted set of separated sound sources matches the order of the set of true sound sources; and replacing one or more of the separated sound sources in the permuted set of separated sound sources with one or more true sound sources from the set of true sound sources to obtain a spurious set of separated sound sources.

[0007] In a second aspect, there is provided a computer program product comprising computer program code portions configured to perform the method according to the first aspect when executed on a computer processor.

[0008] In a third aspect, there is provided a neural network-based system for general sound separation of an audio mix m of any sound source s k=1,…,K wherein the system is configured to perform the method according to the first aspect.

[0009] In a fourth aspect, there is provided a method for general sound separation, the method comprising: estimating, by a separator, a set of separated sound sources from an audio mix m of any sound source s k=1,…,K wherein the separator is trained using the method according to the first aspect. wherein the separator is trained using the method according to the first aspect.

[0010] In a fifth aspect, there is provided a neural network-based system for general sound separation, the neural network-based system comprising a separator configured to estimate a set of separated sound sources from an audio mix m of any sound source s k=1,…,K wherein the separator is trained using the method according to the first aspect. wherein the separator is trained using the method according to the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The present invention will be described in more detail with reference to the accompanying drawings, which show presently preferred embodiments of the invention.

[0012] Figure 1 Examples of speech source separation are shown.

[0013] Figure 2 Examples of music source separation are shown.

[0014] Figure 3 Examples of general sound separation are shown.

[0015] Figure 4 Examples of instance-based adversarial loss for source separation are shown.

[0016] Figures 5a to 5b shows the S replacement for source separation of context-based adversarial loss, i.e., 2 replacement and 3 replacement examples ( Figure 5a and Figure 5b ). Detailed implementation mode

[0017] Figure 1 is a schematic depiction of the voice source separation 12 of two speakers 10 speaking simultaneously. The output of the voice source separation 12 is speakers 1 and 2 (reference numerals 10a, 10b). Figure 2 is a schematic depiction of the music source separation 16 of a music mix 14. The music source separation 16 outputs the separated music sources, including vocals 14a, "other" 14b, bass 14c, and drums 14d. At the same time, Figure 3 is a schematic depiction of the general sound separation 20 of a mix of any source, which in the example shown is a phone recording 18 (country, beside a highway). The separation 20 outputs the separated sound sources, including traffic noise 18a, wind noise 18b, dog 18c, and bird 18d.

[0018] Recently, deep learning-based general sound separation has been proposed. This technology lies in building source-independent models, which can separate any source given any audio mix. Different from music source separation, general separation is not source-specific and can separate any source given any audio mix. This means that general sound separation systems can not only separate music mixes but also separate user-generated phone recordings containing animal and traffic noises. Note that the task of general sound separation is similar to voice source separation because both rely on source-independent models. In short, general sound separation models are source-independent and not limited to a specific domain (such as the music or voice domain), enabling these general sound separation models to separate any source given any audio mix. The present invention is conceived in the context of deep learning-based general sound separation.

[0019] In this section, permutation invariant training (PIT) will be described. PIT is a technique commonly used to train deep learning-based general source separation models. Next, adversarial training is used to extend PIT, and finally, an example of adversarial PIT for general sound separation will be given.

[0020] Permutation - invariant training (PIT)

[0021] Consider the following audio mix m of length L composed of K' arbitrary sources s: In general sound separation, the separator model f θ predicts K estimated sources from this mix m The PIT optimizes the learnable parameters θ of the separator by minimizing the following permutation-invariant loss:

[0022]

[0023] where this minimization is performed over the set of all permutation matrices P, and L can be any regression loss. P * is represented as the optimal permutation matrix that minimizes Equation (1). Note that constructing a source-independent model requires a permutation-invariant loss because the output of f θ can be any source and in any order. Therefore, the model cannot focus on predicting one source type per output, and from the perspective of the loss, any possible permutation of the output sources must be equally correct. To accommodate this, PIT is adopted so that the model's predictions do not depend on any source type or any particular order. Finally, note that the separator outputs K sources, and in the case where the mixture contains K′ < K sources, this disclosure will set s k = 0 when k > K′. A commonly used L is the negative threshold signal-to-noise ratio (SNR):

[0024]

[0025] where determines the maximum SNR (signal-to-noise ratio) value, thus preventing the loss from being unreasonably amplified when is close to zero. Note that the numerator in Equation (2) does not depend on θ, so the loss is equivalent to the thresholded log mean squared error:

[0026]

[0027] Also note that when the target source is silent s k = 0, Equation (3) is unbounded. In that case, a different loss can be used based on thresholding with respect to the mixture (non-silent):

[0028]

[0029] In summary, the well-performing L for PIT that can be used in this disclosure for general sound separation is as follows:

[0030]

[0031] Adversarial permutation - invariant training

[0032] As described below, adversarial training is used to extend PIT. In the context of source separation, adversarial training consists of training two models simultaneously: the separator f θ , which produces a seemingly reasonable separation and the discriminator D, which judges the separation Is it generated by the separator f θ (fake), or is it the true separation s (real) from the dataset? In this setting, the goal of the separator is to estimate a separation (fake) that is as close as possible to the separation (real) from the dataset, such that D misclassifies the estimated source as real by exploiting the adversarial loss. Some aspects of the present invention combine variants of the following discriminators: an instance-based discriminator D inst and a novel S replacing the context-based discriminator D ctx,S . Each discriminator has a different role and is applicable to various domains, such as: waveforms, magnitude STFT, filter banks, and / or masks. Without loss of generality, the system and method are first presented in the waveform domain, and subsequently, it will be shown that multiple discriminators (D inst and D ctx,S ) that can operate (and be combined) in various domains can be used to adversarially train f θ , thereby generating an actual separation for general sound separation.

[0033] Instance - based adversarial loss

[0034] The role of the instance-based discriminator D inst is to provide, without any context, a loss hint regarding the authenticity of the separated source alone. To do so, D inst evaluates the authenticity of each source individually:

[0035]

[0036] where the square brackets [·] define the input to D inst , and left / right denote real / fake separation (not a division operation). In this case, the individual real / fake separations are fed into D inst that is trained to classify the separation as real or fake. D inst is trained to maximize the following term:

[0037]

[0038]

[0039] However, it should be noted that the adversarial setting can be unstable and challenging to train. Therefore, alternative adversarial training losses have been proposed, such as the least squares generative adversarial network (LSGAN) or the metric GAN. In some embodiments, the hinge (adversarial) loss can be utilized:

[0040]

[0041] Figure 4An example of an instance-based discriminator D of a neural network-based system is shown, and the system further includes a separator f inst (reference numeral 32 in the drawings). The separator 30 is configured (trained) to estimate a set of separated sound sources from the audio mix m θ (reference numeral 30 in the drawings). The separator 30 is configured (trained) to estimate a set of separated sound sources from the audio mix m In the example shown, K = 4, so the separator 30 estimates the sources from the audio mix m The instance-based discriminator 32 is configured to provide an instance-based loss cue based on consideration of the separately input separated sound sources (e.g., or ). Thus, the instance-based discriminator 32 of the example shown is configured to operate in the waveform domain. Thus, the sound sources input to the instance-based discriminator 32 are represented in the waveform domain. The instance-based loss cue (e.g., ) may also be referred to as an instance-based adversarial loss cue or an instance-based adversarial cue. As will be described in further detail herein, the separator 30 can be trained by minimizing the loss according to the instance-based loss cue provided by the instance-based discriminator 32

[0042] Note that in Figure 4 the discriminator 32 (i.e., D inst ) individually evaluates each source instance and classifies them as true or false, and this operation is performed for all estimated source and ground-truth source pairs. Previous studies have explored instance-based discriminators. However, these instance-based discriminators are typically employed in source-specific settings, where each D inst is dedicated to one source type. For example, in music source separation, source-specific discriminators are used for bass, guitar, drum, and vocal sources. Or in speech source separation, D inst is only trained to evaluate speech signals. However, according to the present disclosure, each D inst for general sound separation is not dedicated to any one source (they are source-independent) and evaluates the authenticity of any audio regardless of its source type

[0043] S replaces context - based adversarial loss

[0044] The context-based discriminator D using S replacement serves to provide a loss cue regarding the authenticity of the separated sources considering all the sources present in the mix (i.e., the context): ctx,S The context-based discriminator D using S replacement serves to provide a loss cue regarding the authenticity of the separated sources considering all the sources present in the mix (i.e., the context):

[0045]

[0046] Among them, the square brackets [·] define the input of D ctx,S and the left / right indicates true separation / false separation (not division operation). In this case, all separations are fed jointly to provide context for D that learns to classify such separations as true / false. In addition, D ctx,S can also be conditioned on the mixed sound m of the input: ctx,S

[0047]

[0048] The false examples contain entries which are obtained by randomly sampling S ∈ {1, …, K} indices k and replacing the estimated source with the true source s k :

[0049]

[0050] Among them, P * of is the optimal permutation matrix that minimizes formula (1). To further understand the role of P * , consider the following case as an example: K = 4, and among them, the estimated source is sorted as to match the order of the true source [s 1 , s 2 , s 3 , s 4 (because the estimation independent of the source does not necessarily need to match the same order as the true source). Therefore, for example, in the case of K = 4, the possible inputs to D ctx,S=2 are: For example, the permutation of the estimated source is such that the selected source can be replaced with its corresponding true source s k .

[0051] Figures 5a to 5b Additional examples in the case of K = 4 are also depicted in, which show examples of the context-based discriminator D ctx,S=2或者3 (reference numeral 34) of a neural network-based system, which also includes a separator f θ (reference numeral 30). Similar to Figure 4 , K = 4, so the separator 30 estimates the source from the audio mix m. The context-based discriminator 34 is configured to be based on the separated sound sources of a set of inputs (for example, ​) is considered to provide context-based loss cues. Therefore, the illustrated example context-based discriminator 34 is configured to operate in the waveform domain. Therefore, the separated sound sources of the set of inputs are represented in the waveform domain. The context-based loss cues (e.g., ) may also be referred to as a context-based adversarial loss cue or a context-based adversarial cue. As will be described in further detail herein, the separator 30 may be trained by minimizing the loss according to the context-based loss cue provided by the context-based discriminator 34. According to the previous discussion, training the context-based discriminator 34 may include a set s=[s 1 ,s 2 ,s 3 ,s 4 ] and the false set of separated sound sources To maximize the loss. Figure 5a and Figure 5b As shown, the false set of separated sound sources Obtained by doing the following:

[0052] - obtain the set of separated sound sources estimated by separator 30 from the audio mix m (ie represented in the waveform domain) (Step S1);

[0053] - Collection of separated sound sources Arrange so that the set of separated sound sources after arrangement The order of the real sound source [s 1 ,s 2 ,s 3 ,s 4 ] in order (step S2); and

[0054] -Use a collection of real sound sources [s 1 ,s 2 ,s 3 ,s 4 ] one or more real sound sources s k=1,2,3,4 The set of separated sound sources after replacement and arrangement One or more separated sound sources To obtain a pseudo set of separated sound sources (Step S3).

[0055] Thus, the pseudo set of separated sound sources can be obtained in, are sorted to match the order of the set s of real sound sources, and where the false set of separated sound sources including the separated sound source estimated by the separator 30 The corresponding sources (i.e., entries), and further includes one or more true sound sources in the set s of true sound sources. According to the previous discussion, D ctx,S The parameter S in represents the false set of separated sound sources The true sound source s in k The number of. Therefore, S represents the separated sound sources replaced by the true sound source s k The number of. In some embodiments, S is equal to or greater than 2, such as S = 3. More generally, S can be greater than 0. Note that when K = 4, D is used

[0056] means there will be the following input: ctx,S=0 This is equivalent to the standard context-based adversarial loss already used for voice source separation (without using S replacement). Therefore, the present disclosure lies in generalizing these systems and methods for general sound separation, and proposing an S replacement scheme to improve the separation quality. Formally, D

[0057] is trained to maximize the following loss: ctx,S This is equivalent to the standard context-based adversarial loss already used for voice source separation (without using S replacement). Therefore, the present disclosure lies in generalizing these systems and methods for general sound separation, and proposing an S replacement scheme to improve the separation quality.

[0058]

[0059] However, keep in mind that an alternative adversarial training loss can be used to make the adversarial training more stable. In some embodiments, the hinge (adversarial) loss can be utilized:

[0060]

[0061] Finally, note that since the present disclosure describes embodiments using formula (1) to estimate P * The S replacement context-based adversarial loss is also permutation invariant. The difference from the standard PIT is that the S replacement context-based adversarial loss does not rely on optimizing the parameters of f with respect to the regression loss in formula (1). θ Different from D that captures the local context related to a single source inst , the discriminator D ctx,S defines the loss related to the authenticity of the separation considering the global context between sources. Moreover, regarding D inst , note that the instance-based adversarial discriminator does not need to calculate P * to obtain a permutation invariant output, because D inst lacks the context required to evaluate the order of the sources.

[0062] Multi - discriminator training

[0063] The above embodiments present D in the waveform domain inst and D ctx,S , that is: and In the following, additional embodiments introduce discriminator D in the following domains inst and D ctx,S : the magnitude STFT domain ( and ) and the mask domain ( and ), and explains how to combine them to improve the separation quality. The motivation for combining multiple discriminators is to enable the loss of the guiding separator f θ to be based on a richer set of cues. Note that the instance-based and context-based discriminator D can provide different perspectives of the same signal in various domains: waveform, magnitude STFT, and mask. For example, the predicted mask is typically used to filter the input mixture, and the proposed discriminator can help evaluate both the authenticity of the mask and the authenticity of the waveform and magnitude STFT. This disclosure explores multiple discriminators D for general source separation for the first time. The short-time Fourier transform (STFT) of the mixture m is defined as M = STFT(m). The magnitude STFT is obtained by taking the absolute value of M in the complex domain (i.e., |M|). |S k | and represent the magnitude STFTs of the true source and the estimated source, respectively. The ratio mask R k (ranging from 0 to 1) is obtained from the magnitude STFT and is used to filter the source from the mixture:

[0064]

[0065] where: S k = M ⊙ R k and ⊙ represents the Hadamard (element-wise) product. Although for this example, the ratio mask R k is used, it should be noted that this process can be generalized to other types of masks (e.g., such as the binary mask). Using the same notation as above, the inputs to the instance-based and are defined as follows:

[0066]

[0067] And for example, for the context-based and inputs are defined as follows:

[0068]

[0069] where, and The entries follow the same S replacement process as in formula (6). Here, the optimal permutation matrix P required to reorder the false examples * is calculated considering the L1 loss between the magnitude STFT or the masks.

[0070] As pointed out above, multiple discriminators D can be combined to improve the separation quality. For example, at least a first context-based discriminator and a second context-based discriminator configured to operate in different domains can be combined. Additionally, at least a first instance-based discriminator and a second instance-based discriminator configured to operate in different domains can be combined. For example, these discriminators (i.e., context-based or instance-based discriminators) can be configured to operate in one of the corresponding domains of the waveform domain, the magnitude STFT domain, the filter bank domain, or the mask domain (e.g., can be combined with and / or while can be combined with and / or ). The instance-based and context-based discriminators D can be jointly trained to maximize formula (5) and formula (7) in different domains. For example, in addition to training each D separately, it can be trained in the following way: training together with , training together with , training together with , or all the proposed discriminators D can be combined together (at the cost of longer training time). Note that the more discriminators D are used, the higher the computational cost of running the loss and the longer the time spent on running the training. Also note that adding more discriminators D does not affect the inference time.

[0071] Separator loss

[0072] When using adversarial training, the separator model f θ is trained such that it produces a separation that is misclassified by the discriminator. So this means that the estimated separation (false) will be misclassified by the discriminator D as the true separation s (true). To do this, during each adversarial training step, first the discriminator is updated based on or any combination of losses in any of the above domains (without updating the separator). Then, minimize to train the separator (without updating the discriminator).

[0073] For example, when using ( When frozen), the following separator losses are minimized:

[0074]

[0075] Or when using the following losses are minimized by the present disclosure:

[0076]

[0077] Or, for another example, when training using two discriminators (such as ), the losses to be minimized are as follows:

[0078]

[0079] Also, remember that alternative adversarial training losses can be used to make adversarial training more stable. The hinge loss can also be used for exploration as follows:

[0080]

[0081] Although the present disclosure does not list all possible loss combinations for the sake of brevity, based on these examples, any possible combination (including hinge loss variants) described throughout the present disclosure can be easily inferred. Finally, the standard PIT regression loss in Equation (1) can also be used to extend adversarial PIT:

[0082]

[0083] where is a positive weighting factor. Interestingly, all previous studies using adversarial PIT for voice source separation (as opposed to general sound separation) have relied on However, given the powerful training losses provided by multiple discriminators and D ctx,s (with S substituted), our setup allows for the first time to discard the regression PIT loss and rely on a pure adversarial setup.

[0084] Architecture

[0085] This section describes exemplary embodiments of the architecture of the system for implementing the present disclosure. Thus, this is just one possible embodiment. The present disclosure does not depend on the separators, discriminators, or adversarial losses described herein.

[0086] Exemplary embodiments of the separator

[0087] The input mixed audio m is sampled at 16 kHz, where L = 160000 (10 s). First, the input is mapped to the STFT domain M = STFT(m), with a window of 32 ms and an overlap rate of 25%. The magnitude STFT|M| is obtained from M and fed as a tensor of shape [F, T] (F = 256 frequency bins, T = 1250 frames) into U-Netg θ . The output of U-Netg θ is designed to predict the ratio mask. To this end, a softmax layer K is applied on the source dimension: such that the sum of the predicted sources is the input mixed audio. Then, the estimated STFT is filtered out from the mixed audio: Finally, the inverse STFT is used to obtain the separated waveform: U-NETg θ consists of an encoder, a bottleneck, and a decoder.

[0088] The encoder is a sequence of 4 blocks, each block consisting of 2 ResNet blocks followed by a downsampler. Each ResNet block consists of the following layers: a group-norm layer, a SiLU non-linear layer, a 1D convolutional neural network (1D-CNN) layer, a group-norm layer, a SiLU layer, a dropout layer, and a 1D-CNN layer followed by a skip connection layer, where the 1D-CNN layer has a kernel size = 3 and a stride = 1. Using 1D-CNN makes the architecture more memory-efficient, and it is also the choice for other architectures for source separation such as TDCN++. The dropout layer probability is set to 0.1, and a group size of 32 is used in the group-norm layer. The downsampling layer is a 1D-CNN layer (kernel size = 3, stride = 2). The sequence of channels in the encoder is [256, 512, 512, 512].

[0089] The bottleneck block consists of the following layers: a ResNet layer (as defined above), a self-attention layer

[26] , and a ResNet layer, where all layers maintain 512 channels. The decoder block consists of 4 blocks, which is the opposite of the encoder's structure, replacing the downsampler with an upsampler, resulting in the following channel sequence: [512, 512, 512, 256]. The upsampler first performs linear interpolation, followed by a 1D-CNN layer (kernel size = 3, stride = 1). Following the standard Unet structure, the output of each encoder block is injected (at its corresponding block level) as input into the decoder block. The encoder features are concatenated along the channel dimension with the input of the decoder ResNet blocks (a linear layer adjusts the number of channels if necessary). Different from the encoder that uses 2 ResNet blocks, 3 ResNet blocks are used in each decoder block. The feature maps from 2 ResNet encoder blocks are concatenated with the input of the first 2 ResNet decoder blocks, and the output of the downsampling layer of the encoder is concatenated with the input of the third ResNet decoder block. The last linear layer adjusts the output to be able to predict the expected number of sources (K = 4), resulting in the following output:

[0090] Exemplary embodiments of the discriminator

[0091] The discriminator is a CNN and outputs a single scalar. and share the same architecture: x4 1D-CNN layers (kernel size = 4, stride = 3) interleaved by LeakyRe-Lus (negative slope of 0:2), resulting in the following channel sequence: [C, 128, 256, 256, 512], where, conditional on m and K = 4, for C = 1, and for C = 5. Then, the 512 channels are reduced to 1 channel using a 1D-CNN layer (kernel size = 4, stride = 1), and the output layer (mapping the resulting vector to a scalar) is linear. Follows the same architecture as and except that the 1D-CNN is 2D (kernel size = 44, stride = 33), and the number of channels is halved to: [C, 64, 128, 128, 256].

[0092] Exemplary embodiments of the adversarial loss

[0093] If there is evidence that training with the standard adversarial loss (also presented in this disclosure) can be challenging, the experiments of this disclosure rely on a hinge loss variant of adversarial training. That is, this disclosure does not rely on any specific adversarial loss, but can use any adversarial loss. Finally, this disclosure successfully combines the adversarial PIT scheme of this disclosure with PIT regression And different from previous studies, the results of this disclosure show that is not absolutely necessary and can be discarded. Due to the rich losses provided by the multiple Ds and the strong guidance of D ctx,S (which depends on S replacement), this can be possible.

[0094] Adversarial PIT for speech source separation

[0095] Adversarial PIT for speech source separation is relevant to this disclosure. As pointed out in the background art, the goal of speech source separation research is to develop speaker-independent models. A common technique for achieving this is to use PIT to develop speaker-independent models in a way similar to that for developing source-independent models. The field of speech source separation also uses adversarial training to extend PIT. Table 1 summarizes this disclosure compared to previous systems based on adversarial PIT for speech source separation:

[0096]

[0097] Table 1

[0098] In Table 1, [1] to [4] refer to:

[0099] [1] Chenxing Li, Lei Zhu, Shuang Xu, Peng Gao, and Bo Xu, “CBLDNN-based speaker-independent speech separation via generative adversarial training,” in ICASSP, 2018.

[0100] [2]Lianwu Chen, Meng Yu, Yanmin Qian, Dan Su, and Dong Yu, “Permutation invariant training of generative adversarial network for monaural speech separation,” in Interspeech, 2018.

[0101] [3]Ziqiang Shi, Huibin Lin, Liu Liu, Rujie Liu, Shoji Hayakawa, and Jiqing Han, “Furcax: End-to-end monaural speech separation based on deep gated (de) convolutional neural networks with adversarial example training,” in ICASSP, 2019.

[0102] [4]Chengyun Deng, Yi Zhang, Shiqian Ma, Yongtao Sha, Hui Song, and Xiangang Li, “Conv-TasSAN: Separative adversarial network based on conv-tasnet.,” in Interspeech, 2020, pp. 2647–2651.

[0103] Previous studies have found that using adversarial training to extend PIT improves their speech source separation system. Interestingly, the report of SSGAN-PIT states that all of their adversarial variants perform similarly, but variant (i) converges faster during training. The report of SSGAN-PIT also states that adversarial training alone (without using PIT) performs poorly. This disclosure differs from previous studies on adversarial PIT for speech source separation in many key aspects (see Table 1):

[0104] - This disclosure generalizes adversarial PIT for speech source separation to general sound separation. While previous systems have shown that it can be used to separate two speakers (K = 2), our experiments show that the general method and system proposed in this disclosure for general sound separation can be used to separate any type of four sources (K = 4).

[0105] - This disclosure proposes a new context-based discriminator using S substitution, which allows D ctx,S to provide better guidance when dealing with more than two heterogeneous sources (e.g., any type of four sources). Importantly, in adversarial PIT for speech source separation, the discriminator can rely on source-specific cues to judge the authenticity, but in adversarial PIT for general sound separation, the discriminator (or discriminators) cannot rely on source-specific cues because the model is source-independent. In the experiments, it was found that, as in previous studies using or the standard adversarial PIT without S substitution cannot obtain competitive results in general sound separation. However, when experimenting with the S substitution strategy, the model obtained better results. Therefore, embodiments of the present invention include D ctx,S using S substitution, which enables adversarial PIT to be generalized to general sound separation, in which multiple (e.g., 4) heterogeneous sources (e.g., any type of source) need to be separated from a mixture.

[0106] - Some embodiments of this disclosure improve the separation quality by using multiple discriminators operating on various domains - magnitude STFT, waveform, and mask. Therefore, different from previous studies, this disclosure can rely on multiple discriminators, which guide the separator based on a rich set of cues from different domains. Since the guidance provided by the multiple discriminators used is rich enough, this allows for the first time to discard the regression PIT loss and rely on a pure adversarial setting.

[0107] - The present disclosure has been successfully explored using HingeGAN in our experiments, but in principle, any other adversarial loss scheme (such as LSGAN or MetricGAN) should also be feasible.

[0108] Table 2 shows various D obtained using the example architectures presented above and using the reverberant FUSS dataset, ctx,S configured with 20k / 1k / 1k (training / validation / test) 10s mixes, each mix having one to four sources. The m column indicates whether D ctx,S is conditioned on m. The SI-SNR column indicates the SI-SNR I / SI-SNR S , where SI-SNR is the scale-invariant SNR:

[0109]

[0110] where

[0111]

[0112] and To account for inactive sources, estimated target pairs with a silent target source are discarded. For mixes with one source, this is equivalent to since for a one-source mix, the goal is to bypass the mix (the subscript S here represents single source). For mixes with two to four sources, the average of all sources is reported (subscript I represents improvement). To compare with a meaningful state-of-the-art baseline, the DCASE model, i.e., the TDCN++ model that predicts STFT masks, was used. This model was trained on the reverberant FUSS dataset and evaluated using a metric based on the standard SI-SNR. Finally, SI-SNR was reported for consistency S but SI-SNR I is more relevant for the comparison model since most SI-SNR SThe scores are already very close to the upper limit of 39.9 dB (see Table 3). These models are trained using the Adam optimizer until convergence (about 500k iterations), and the best model on the validation set is selected for evaluation. For training, the learning rate is adjusted to {10-5, 10-4, 10-3}, and the batch size is adjusted to {16, 32, 64, 96, 128} so that all experiments, including ablation experiments and baseline experiments, can achieve the best possible results. Finally, mixture-consistency projection is used during inference (instead of during training) because this operation systematically improves the SI-SNR I without reducing the SI-SNR S . The best model was trained for one month using 4 V100 GPUs with a learning rate of 10-4 and a batch size of 128.

[0113]

[0114] Table 2

[0115] Table 3 includes a comparison of adversarial PIT variants with the baseline. The SI-SNR column indicates the SI-SNR in dB I / SI-SNR S . All Ds in Table 3 ctx,S are conditioned on m, where S = 3 because this setting is superior to other settings (see Table 2). All adversarial PIT ablation experiments (rows 1 to 11 and Table 2) use the same f θ .

[0116]

[0117] Table 3

[0118] Those skilled in the art will recognize that the present invention is in no way limited to the preferred embodiments described above. On the contrary, many modifications and variations are possible within the scope of the appended claims.

[0119] Aspects of the systems described herein can be implemented in a suitable computer-based audio processing network environment for processing digital or digitized audio files. Portions of the adaptive audio system can include one or more networks that include any desired number of separate machines, the machines including one or more routers (not shown) for buffering and routing data transmitted between the computers. Such networks can be built on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0120] One or more components, blocks, processes, or other functional elements may be implemented by a computer program executed by a processor-based computing device of a control system. It should also be noted that any number of combinations of hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media may be used to describe the various functions disclosed herein in terms of behavior, register transfer, logic components, and / or other characteristics. Computer-readable media that may embody such formatted data and / or instructions include, but are not limited to, various forms of physical (non-transitory), non-volatile storage media, such as optical, magnetic, or semiconductor storage media.

[0121] Although one or more specific implementations have been described by way of example and in terms of specific embodiments, it should be understood that the one or more specific implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements that would be apparent to those skilled in the art. Accordingly, the scope of the appended claims should be given the broadest interpretation so as to include all such modifications and similar arrangements.

[0122] Explanation

[0123] A computing device implementing the techniques described above may have the following example architecture. Other architectures are possible, including architectures with more or fewer components. In some implementations, the example architecture includes one or more processors (e.g., dual-core processors), one or more output devices (e.g., LCD), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components may exchange communications and data via one or more communication channels (e.g., a bus), which may utilize various hardware and software to facilitate the transfer of data and control signals between the components.

[0124] The term "computer-readable medium" refers to a medium that participates in providing instructions to a processor for execution, including but not limited to non-volatile media (e.g., optical discs or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media includes, but is not limited to, coaxial cables, copper wire, and fiber optics.

[0125] Computer-readable media may further include an operating system (e.g., An operating system), a network communication module, an audio interface manager, an audio processing manager, and a real-time content distributor. The operating system can be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system performs basic tasks, including but not limited to: identifying inputs from network interfaces and / or devices and providing outputs thereto; tracking and managing files and directories on a computer-readable medium (e.g., a memory or a storage device); controlling peripheral devices; and managing traffic on one or more communication channels. The network communication module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.).

[0126] The architecture can be implemented in a parallel processing or peer-to-peer infrastructure, or on a single device having one or more processors. The software can include multiple software components or can be a single body of code.

[0127] The described features can be advantageously implemented in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device and to transmit data and instructions thereto. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform an activity or bring about a result. The computer program can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0128] By way of example, suitable processors for executing instruction programs include both general-purpose processors and special-purpose processors, as well as a single processor or one of multiple processors or cores of any kind of computer. Generally, the processor will receive instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data files or be operatively coupled to communicate with such mass storage devices; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, by way of example including: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by or incorporated in an ASIC (application-specific integrated circuit).

[0129] To provide interaction with a user, the features can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or a retinal display device. The computer can have a touch surface input device (e.g., a touch screen) or a keyboard and a pointing device such as a mouse or a trackball, by which the user can provide input to the computer. The computer can have a voice input device for receiving voice commands from the user.

[0130] The features can be implemented in a computer system that includes backend components such as data servers, or includes middleware components such as application servers or Internet servers, or includes frontend components such as client computers having a graphical user interface or an Internet browser, or any combination thereof. The components of the system can be connected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include, for example, LANs, WANs, and the computers and networks that form the Internet.

[0131] A computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship between the client and the server arises from computer programs running on the respective computers and having a client-server relationship with respect to each other. In some embodiments, the server transmits data (e.g., HTML pages) to the client device (e.g., to display data to a user interacting with the client device and to receive user input from the user). Data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.

[0132] A system of one or more computers can be configured to perform particular actions by virtue of software, firmware, hardware, or a combination thereof installed in the system that, in operation, causes the system to perform the actions. One or more computer programs can be configured to perform particular actions by virtue of including instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.

[0133] Although this specification contains many specific implementation details, these details should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. The specific features described in the context of separate embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may be described above as acting in a particular combination and even initially claimed as such, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0134] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0135] Unless otherwise specifically stated, it should be understood from the following discussion that throughout the discussion of this invention, terms such as "processing", "computing", "calculating", "determining", "analyzing", etc. are used to refer to actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or transform data represented as physical (such as electronic) quantities into other data similarly represented as physical quantities.

[0136] References throughout this invention to "an example embodiment", "some example embodiments", or "example embodiments" mean that a particular feature, structure, or characteristic described in connection with the example embodiment is included in at least one example embodiment of the invention. Thus, the phrases "in an example embodiment", "in some example embodiments", or "in example embodiments" that appear throughout this invention do not necessarily all refer to the same example embodiment. Additionally, in one or more example embodiments, the particular features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those of ordinary skill in the art from this invention.

[0137] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe common objects indicates merely that different instances of similar objects are referred to and is not intended to imply that the objects so described must be in a given order temporally, spatially, in ranking, or in any other manner.

[0138] Likewise, it should be understood that the phraseology and terminology used herein is for descriptive purposes and should not be regarded as limiting. The use of "including," "comprising," or "having" and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless otherwise specified or limited, the terms "mounted," "connected," "supported," and "coupled" and variations thereof are used broadly and encompass direct and indirect mountings, connections, supports, and couplings.

[0139] In the claims below and the description herein, any of the terms comprising, comprised of, or which comprises is an open term, which means including at least the elements / features that follow, but not excluding other elements / features. Therefore, when the term comprising is used in a claim, the term should not be interpreted as being limited to the devices or elements or steps listed thereafter. For example, the scope of the expression of a device comprising A and B should not be limited to being composed of only elements A and B. The term including, or any of which includes, or that includes, as used herein, is also an open term that also means including at least the elements / features that follow the term, but not excluding other elements / features. Therefore, including is synonymous with comprising and means comprising.

[0140] It should be understood that in the above description of example embodiments of the invention, various features of the invention are sometimes grouped together in a single example embodiment, figure, or description thereof in order to simplify the invention and aid in understanding one or more of the various inventive aspects. However, this approach to the invention should not be interpreted as reflecting an intention that the claims require more features than those expressly recited in each claim. On the contrary, as reflected in the following claims, the inventive aspects lie in less than all the features of a single, previously disclosed example embodiment. Therefore, the claims following the specification are hereby expressly incorporated into this specification, with each claim independently serving as a separate example embodiment of the invention.

[0141] In addition, although some example embodiments described herein include some but not other features included in other example embodiments, the combination of features of different example embodiments is meant to be within the scope of the present invention and forms different example embodiments, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed example embodiments can be used in any combination.

[0142] In the description provided herein, numerous specific details are set forth. However, it should be understood that example embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0143] Accordingly, while what has been described is considered to be the best mode of the present invention, those skilled in the art will recognize that other and further modifications can be made to the present invention without departing from the spirit thereof, and it is intended to claim all such changes and modifications that fall within the scope of the present invention. For example, any of the formulas given above merely represent procedures that can be used. Functions can be added or removed from the block diagrams, and operations can be interchanged among the functional blocks. Steps can be added or removed from the methods described within the scope of this disclosure.

[0144] Aspects of the example embodiments include the following enumerated example embodiments ("EEE"):

[0145] EEE 1. A method for adversarially training a separator for general sound separation of an audio mix m for an arbitrary sound source s k=1,…,K The method comprising:

[0146] Training a context-based discriminator configured to provide a context-based loss cue based on consideration of a set of separated sound sources of a set of inputs; and

[0147] Training the separator to minimize a loss according to the context-based loss cue provided by the context-based discriminator;

[0148] Wherein training the context-based discriminator includes maximizing a loss based on a set of true sound sources and a false set of separated sound sources, wherein the false set of separated sound sources is sorted to match the order of the set of true sound sources, and wherein the false set of separated sound sources includes sources corresponding to the separated sound sources estimated by the separator and further includes one or more true sound sources from the set of true sound sources.

[0149] EEE 2. The method according to EEE 1, wherein the false set of separated sound sources is obtained by:

[0150] Obtain a set of separated sound sources corresponding to a set of estimated separated sound sources estimated from the audio mix m by the separator ;

[0151] Permute the set of separated sound sources such that the order of the permuted set of separated sound sources matches the order of the set of true sound sources; and

[0152] Replace one or more of the separated sound sources in the permuted set of separated sound sources with one or more true sound sources from the set of true sound sources to obtain a spurious set of separated sound sources.

[0153] EEE 3. The method according to EEE 2, wherein the set of separated sound sources is permuted using a permutation matrix P * where P * is the permutation matrix among a set of all permutation matrices P that minimizes the loss between the set of true sources and the set of separated sound sources permuted using P, where optionally, the loss is a permutation-invariant loss.

[0154] EEE 4. The method according to any one of the foregoing EEEs, wherein the context-based discriminator is configured to operate in the waveform domain, the magnitude STFT domain, the filter bank domain, or the mask domain.

[0155] EEE 5. The method according to any one of the foregoing EEEs, wherein the context-based discriminator is a first context-based discriminator configured to operate in a first domain (e.g., and thus provide a first context-based loss cue based on consideration of a set of separated sound sources represented in the first domain), and the method further comprises training a second context-based discriminator configured to operate in a second domain different from the first domain and provide a second context-based loss cue based on consideration of a set of separated sound sources represented in the second domain,

[0156] wherein the loss minimized to train the separator is further based on the context-based loss cue provided by the second context-based discriminator (i.e., the loss is based on the first context-based loss cue and the second context-based loss cue).

[0157] EEE 6. The method according to EEE 5, wherein training the second context-based discriminator includes maximizing a loss based on a representation of a set of true sound sources in the second domain (i.e., the set of true sound sources represented in the second domain) and a second false set of separated sound sources represented in the second domain, wherein the second false set of separated sound sources is sorted to match the order of the set of true sound sources (e.g., in the second domain), and wherein the second false set of separated sound sources includes sources corresponding to the separated sound sources estimated by the separator and further includes one or more true sound sources from the set of true sound sources represented in the second domain.

[0158] EEE 7. The method according to any one of EEE 5 to EEE 6, wherein the first context-based discriminator and the second context-based discriminator are configured to operate in a respective one of a waveform domain, an amplitude STFT domain, a filter bank domain, or a mask domain.

[0159] EEE 8. The method according to any one of the foregoing EEEs, wherein the context-based discriminator is configured to operate in the waveform domain, and wherein the set of true sound sources [s 1 ,…,sK] and the false set of separated sound sources are represented in the waveform domain.

[0160] EEE 9. The method according to EEE 8, wherein the false set of separated sound sources is obtained by:

[0161] obtaining a set of separated sound sources represented in the waveform domain and estimated by the separator from the audio mix m

[0162] permuting the set of separated sound sources such that the order of the permuted set of separated sound sources matches the order of the set of true sound sources [s 1 ,…,sK]; and

[0163] replacing one or more of the separated sound sources in the permuted set of separated sound sources 1 ,…,s K with one or more true sound sources s k from the set of true sound sources [s to obtain the false set of separated sound sources

[0164] ​EEE 10. The method according to EEE 9, wherein the set of isolated sound sources is arranged using a permutation matrix P * to permute,

[0165] wherein P * is the permutation matrix that minimizes this.

[0166] EEE 11. The method according to any one of EEE 8 to EEE 10, wherein the context-based discriminator is represented as and is trained to maximize wherein, based on and based on

[0167] EEE 12. The method according to any one of EEE 1 to EEE 7, wherein the context-based discriminator is configured to operate in the magnitude STFT domain, and wherein the set of true sound sources [|S 1 |,…,|S K |] and the false set of isolated sound sources are represented in the magnitude STFT domain.

[0168] EEE 13. The method according to EEE 12, wherein the false set of isolated sound sources is obtained by:

[0169] obtaining the set of isolated sound sources represented in the magnitude STFT domain and corresponding to the set of estimated isolated sound sources estimated by the separator (30) from the audio mix m the set of isolated sound sources

[0170] permuting the set of isolated sound sources such that the order of the permuted set of isolated sound sources matches the order of the set of true sound sources [|S 1 |,…,|S K |]; and

[0171] replacing one or more of the isolated sound sources in the permuted set of isolated sound sources 1 |,…,|S K |] with one or more true sound sources |S k | in the set of true sound sources to obtain the false set of isolated sound sources ​

[0172] EEE 14. The method according to EEE 13, wherein |S k | and Represent these real sound sources s k and these estimated separated sound sources Amplitude STFT of .

[0173] EEE 15. The method according to any one of EEE 12 to EEE 14, wherein the context-based discriminator is represented as and is trained to maximize in, based on and based on Where M = STFT(m).

[0174] EEE 16. A method according to any one of EEE 1 to EEE 7, wherein the context-based discriminator is configured to operate in a ratio mask domain, and wherein the set of real sound sources [R 1 ,…,R K ] and the false set of separated sound sources is represented in the ratio mask field.

[0175] EEE 17. The method according to EEE 16, wherein the pseudo set of separated sound sources Obtained by doing the following:

[0176] Obtain separated sound sources represented in the ratio mask domain and compared to a set of estimates estimated by the separator from the audio mix m The corresponding set of separated sound sources

[0177] The collection of separated sound sources Arrange so that the set of separated sound sources after arrangement The sequence and the set of real sound sources [R 1 ,…,R K ] in sequence; and

[0178] Using a collection of real sound sources 1 ,…,R K ] one or more real sound sources R k The set of separated sound sources after replacement and arrangement One or more separated sound sources To obtain a pseudo set of separated sound sources

[0179] EEE 18. The method according to any one of EEE 16 to EEE 17, wherein the context-based discriminator is represented as and is trained to maximize wherein, based on and based on wherein M = STFT(m).

[0180] EEE 19. The method according to any one of the foregoing EEEs, further comprising training an instance-based discriminator configured to provide an instance-based loss cue based on consideration of individual isolated sound sources,

[0181] wherein the loss that is minimized to train the separator is further based on the instance-based loss cue provided by the instance-based discriminator.

[0182] EEE 20. The method according to EEE 19, further comprising training the instance-based discriminator by maximizing a loss based on individual ground truth sound sources in a set of ground truth sound sources and individual sound sources corresponding to individual isolated sound sources estimated by the separator.

[0183] EEE 21. The method according to any one of EEE 19 to EEE 20, wherein the instance-based discriminator is configured to operate in a waveform domain, an amplitude STFT domain, a filter bank domain, or a mask domain.

[0184] EEE 22. The method according to any one of EEE 19 to EEE 21, wherein the instance-based discriminator is a first instance-based discriminator configured to operate in a first domain (e.g., and thus provide a first instance-based loss cue based on consideration of individual isolated sound sources represented in the first domain), and the method further comprises training a second instance-based discriminator configured to operate in a second domain different from the first domain and provide a second instance-based loss cue based on consideration of individual isolated sound sources represented in the second domain,

[0185] wherein the loss that is minimized to train the separator is further based on the instance-based loss cue provided by the second instance-based discriminator (i.e., the loss is further based on the first instance-based loss cue and the second instance-based loss cue).

[0186] EEE 23. The method according to EEE 22, wherein the first instance-based discriminator and the second instance-based discriminator are configured to operate in a respective one of a waveform domain, an amplitude STFT domain, a filter bank domain, or a mask domain.

[0187] EEE 24. The method according to any one of EEE 19 to EEE 23, wherein the (one or more) instance-based discriminators and the (one or more) context-based discriminators are jointly trained.

[0188] EEE 25. The method according to any one of the foregoing EEEs, wherein the separated sound sources estimated by the separator are in the waveform domain.

[0189] EEE 26. A computer program product comprising computer program code portions configured to perform the method according to any one of EEE 1 to EEE 24 when executed on a computer processor.

[0190] EEE 27. A neural network-based system for general sound separation of an audio mix m of any sound source s k=1,…,K wherein the system is configured to perform the method according to any one of EEE 1 to EEE 25.

[0191] EEE 28. A method for general sound separation, the method comprising:

[0192] estimating, by a separator (30), a set of separated sound sources from an audio mix m of any sound source s k=1,…,K wherein the separator (30) is trained using the method according to any one of EEE 1 to EEE 25.

[0193] EEE 29. A neural network-based system for general sound separation, the neural network-based system comprising a separator (30) configured to estimate a set of separated sound sources from an audio mix m of any sound source s k=1,…,K wherein the separator (30) is trained using the method according to any one of EEE 1 to EEE 25.

[0194] EEE 30. A neural network-based system for general sound separation of an audio signal, the system comprising:

[0195] a separator configured to estimate a set of separated sound sources from the audio signal;

[0196] ​​An instance-based discriminator configured to provide an instance-based loss hint to the separator based on consideration of the separated sound sources in a set of separated sound sources;

[0197] A context-based discriminator configured to provide a context-based loss hint to the separator based on consideration of the set of separated sound sources; and

[0198] wherein the separator is configured to minimize a loss based on the instance-based loss hint and / or the context-based loss hint.

[0199] EEE 31. The system according to EEE 30, wherein the context-based loss hint is determined based on the authenticity of the set of separated sound sources, and / or wherein the instance-based loss hint is determined based on the authenticity of the separated sound sources in the set of separated sound sources.

[0200] EEE 32. The system according to any one of EEE 30 to EEE 31, wherein the instance-based discriminator is trained to maximize:

[0201] where

[0202]

[0203] EEE 33. The system according to any one of EEE 30 to EEE 32, wherein the instance-based discriminator and / or the context-based discriminator are trained based on an adversarial training loss including a least squares generative adversarial network, a metric generative adversarial network, or a hinge loss.

[0204] EEE 34. The system according to EEE 33, wherein training the context-based discriminator includes: receiving a set of true separated sound sources; receiving a false set of separated sound sources, wherein one or more false separated sound sources in the false set of separated sound sources are true sources determined based on a permutation matrix; and maximizing a loss based on the set of true separated sound sources and the false set of separated sound sources.

[0205] EEE 35. The system according to any one of EEE 30 to EEE 34, wherein the instance-based discriminator and the context-based discriminator are configured to operate in a waveform domain, an amplitude STFT domain, a filter bank domain, and / or a mask domain.

[0206] EEE 36. The system according to any one of EEE 30 to EEE 35, wherein the instance-based discriminator and the context-based discriminator are jointly trained.

[0207] EEE 37. The system according to any one of EEE 30 to EEE 36, wherein the separator is trained via adversarial training.

[0208] EEE 38. The system according to any one of EEE 30 to EEE 37, wherein the separator is trained via adversarial training and a permutation-invariant loss.

[0209] EEE 39. A method for adversarial training of a separator for general sound separation of an audio mix m for any sound source s k=1,…,K The method includes:

[0210] Training a context-based discriminator configured to provide a context-based loss cue based on consideration of a set of separated sound sources of a set of inputs;

[0211] Training an instance-based discriminator configured to provide an instance-based loss cue based on consideration of individual separated sound sources; and

[0212] Training the separator to minimize a loss according to the context-based loss cue provided by the context-based discriminator and the instance-based loss cue provided by the instance-based discriminator.

[0213] EEE 40. The method according to EEE 39, wherein the instance-based discriminator and the context-based discriminator are configured to operate in a waveform domain, an amplitude STFT domain, a filter bank domain, and / or a mask domain.

[0214] EEE 41. The method according to any one of EEE 39 to EEE 40, wherein the instance-based discriminator and the context-based discriminator are jointly trained.

Claims

1. A method for adversarially training a separator (30) for general sound separation of an audio mix m for an arbitrary sound source s k=1,…,K wherein the method Comprising: Training a context-based discriminator (34), the context-based discriminator being configured to provide context-based loss cues based on consideration of separated sound sources of a set of inputs; And Training the separator (30) to minimize loss based on the context-based loss cues provided by the context-based discriminator (34); Wherein training the context-based discriminator (34) includes maximizing loss based on a set of true sound sources and a false set of separated sound sources, wherein the false set of separated sound sources is sorted to match the order of the set of true sound sources, and wherein the false set of separated sound sources includes sources corresponding to the separated sound sources estimated by the separator (30) and further includes one or more true sound sources from the set of true sound sources.

2. The method according to claim 1, Wherein, The false set of separated sound sources is obtained by: Obtain a set of separated sound sources corresponding to the estimated separated sound sources estimated from the audio mix m by the separator (30) ; Permuting the set of separated sound sources such that the order of the permuted set of separated sound sources matches the order of the set of true sound sources; And Replacing one or more of the separated sound sources in the permuted set of separated sound sources with one or more true sound sources from the set of true sound sources to obtain the false set of separated sound sources.

3. The method according to claim 2, Wherein, The set of the separated sound sources is arranged using a permutation matrix P * where P * is the permutation matrix that minimizes the loss between the set of the true sources and the set of the separated sound sources arranged using P among all sets of permutation matrices P 4. The method according to any one of the preceding claims, Wherein, The context-based discriminator (34) is configured to operate in the waveform domain, the magnitude STFT domain, the filter bank domain, or the mask domain.

5. The method according to any one of the preceding claims, Wherein, The context-based discriminator (34) is a first context-based discriminator configured to operate in a first domain, and the method further includes training a second context-based discriminator, the second context-based discriminator being configured to operate in a second domain different from the first domain and to provide second context-based loss cues based on consideration of separated sound sources of a set of inputs represented in the second domain, Wherein the loss minimized to train the separator (30) is further based on the context-based loss cues provided by the second context-based discriminator.

6. The method according to claim 5, Wherein, Training the second context-based discriminator (34) includes maximizing loss based on the representation of the set of true sound sources in the second domain and a second false set of separated sound sources represented in the second domain, wherein the second false set of separated sound sources is sorted to match the order of the set of true sound sources, and wherein the second false set of separated sound sources includes sources corresponding to the separated sound sources estimated by the separator (30) and further includes one or more true sound sources from the set of true sound sources represented in the second domain.

7. The method according to any one of claims 5 to 6, Wherein, the first context-based discriminator and the second context-based discriminator are configured to operate in a respective one of a waveform domain, a magnitude STFT domain, a filter bank domain, or a mask domain.

8. The method according to any one of the preceding claims, wherein, The context-based discriminator (34) is configured to operate in the waveform domain, and wherein the set of true sound sources [s 1 , …, s K and the spurious set of separated sound sources are represented in the waveform domain.

9. The method according to claim 8, wherein, The false set of the isolated sound sources is obtained by the following operations: Obtain a set of separated sound sources represented in the waveform domain and estimated by the separator (30) from the audio mix m For the set of the isolated sound sources perform an arrangement such that the set of the isolated sound sources after the arrangement has an order matching the order of the set of the true sound sources [s 1 ,…, s K ; and using one or more real sound sources s 1 , …, s K in the set of real sound sources to replace one or more of the separated sound sources in the permuted set of separated sound sources k to obtain a fake set of the separated sound sources ​​ 10. The method according to claim 9, wherein, The set of separated sound sources using the permutation matrix P * to permute where P * is the minimized permutation matrix.

11. The method according to any one of claims 8 to 10, wherein, The context-based discriminator (34) is represented as and is trained to maximize wherein based on and based on 12. The method according to any one of claims 1 to 7, wherein, The context-based discriminator (34) is configured to operate in the magnitude STFT domain, and wherein the set of true sound sources [|S 1 |,…,|S K |] and the spurious set of separated sound sources are represented in the magnitude STFT domain.

13. The method according to claim 12, wherein, The false set of the isolated sound sources is obtained by the following operations: Obtain a set of separated sound sources represented in the magnitude STFT domain and corresponding to a set of estimated separated sound sources estimated by the separator (30) from the audio mix m of separated sound sources For the set of the separated sound sources perform an arrangement such that the set of the separated sound sources after the arrangement has an order matching the order of the set of the true sound sources [|S 1 |,…,|S K |]; and Replace one or more of the separated sound sources in the arranged set of separated sound sources with one or more real sound sources |S 1 |,…,|S K |] in the set of real sound sources[|S k to obtain a false set of the separated sound sources by replacing one or more of the separated sound sources in the arranged set of separated sound sources 14. The method according to claim 13, wherein, |S k | and respectively represent the magnitude STFT of the true sound source s k and the estimated separated sound source .

15. The method according to any one of claims 12 to 14, wherein, The context-based discriminator (34) is represented as and is trained to maximize where based on and based on where M = STFT(m).

16. The method according to any one of claims 1 to 7, wherein, The context-based discriminator (34) is configured to operate in the ratio mask domain, and wherein the set of true sound sources [R 1 , …, R K and the spurious set of separated sound sources are represented in the ratio mask domain.

17. The method according to claim 16, wherein, The false set of the separated sound sources is obtained by the following operations: Obtain a set of separated sound sources represented in the ratio mask domain and corresponding to a set of estimated separated sound sources estimated by the separator (30) from the audio mix m of the separated sound sources For the set of the separated sound sources perform an arrangement such that the set of the separated sound sources after the arrangement has an order matching the order of the set of the true sound sources [R 1 , …, R K ; and Replace one or more of the separated sound sources in the permuted set of separated sound sources with one or more real sound sources R 1 , …, R K in the set of real sound sources k to obtain a spurious set of the separated sound sources from the set of separated sound sources ​ 18. The method according to any one of claims 16 to 17, wherein, The context-based discriminator (34) is represented as and is trained to maximize where based on and based on where M = STFT(m).

19. The method according to any one of the preceding claims, further comprising training an instance-based discriminator (32), the instance-based discriminator being configured to provide an instance-based loss cue based on consideration of individual isolated sound sources, wherein, the loss to be minimized to train the separator (30) is further based on the instance-based loss cue provided by the instance-based discriminator (32).

20. The method according to claim 19, further comprising training the instance-based discriminator (32) by: maximizing a loss based on individual ground truth sound sources in the set of ground truth sound sources and individual sound sources corresponding to individual isolated sound sources estimated by the separator (30).

21. The method according to any one of claims 19 to 20, wherein, the instance-based discriminator (32) is configured to operate in a waveform domain, a magnitude STFT domain, a filter bank domain, or a mask domain.

22. The method according to any one of claims 19 to 21, wherein, the instance-based discriminator (32) is a first instance-based discriminator configured to operate in a first domain, and the method further comprises training a second instance-based discriminator, the second instance-based discriminator being configured to operate in a second domain different from the first domain and to provide a second instance-based loss cue based on consideration of individual isolated sound sources represented in the second domain, wherein the loss to be minimized to train the separator (30) is further based on the second instance-based loss cue provided by the second instance-based discriminator.

23. The method according to claim 22, wherein, the first instance-based discriminator and the second instance-based discriminator are configured to operate in a respective one of a waveform domain, a magnitude STFT domain, a filter bank domain, or a mask domain.

24. The method according to any one of claims 19 to 23, wherein, The instance-based discriminator (32) and the context-based discriminator (34) are jointly trained.

25. The method according to any one of the preceding claims, wherein, the separated sound sources estimated by the separator (30) are in the waveform domain.

26. A computer program product comprising a computer program code portion configured to perform the method according to any one of claims 1 to 25 when executed on a computer processor.

27. A neural network-based system for general sound separation of an audio mix m from any sound source s k=1,…,K ​ wherein, the system is configured to perform the method according to any one of claims 1 to 25.

28. A method for general sound separation, the method comprising: The set of sound sources separated from the audio mix m by a separator (30) from any sound source s k=1,…,K is estimated, wherein the separator (30) is trained using the method according to any one of claims 1 to 25.

29. A neural network-based system for general sound separation, the neural network-based system comprising a separator (30) configured to estimate a set of separated sound sources from an audio mix m of an arbitrary sound source s k=1,…,K of the separated sound sources wherein, the separator (30) is trained using the method according to any one of claims 1 to 25.