Methods, devices, and computer-readable media with trained user generated content source separation model

WO2026206648A1PCT designated stage Publication Date: 2026-10-01DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/019110
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-13
Publication Date
2026-10-01

Smart Images

  • Figure US2026019110_01102026_PF_FP_ABST
    Figure US2026019110_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Methods, devices, and non-transitory computer-readable media with a trained user generated content (UGC) source separation model are disclosed. A method may perform training data pair preparation and may performed a segmented training method to generate a trained UGC source separation model.
Need to check novelty before this filing date? Find Prior Art

Description

D25029W001 METHODS, DEVICES, AND COMPUTER-READABLE MEDIA WITH TRAINED USER GENERATED CONTENT SOURCE SEPARATION MODELCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from International Patent Application No. PCT / CN2025 / 084738, filed on March 25, 2025, which is incorporated by reference in its entirety.TECHNICAL FIELD

[0002] This application relates generally to audio processing, and more specifically to training, accessing, and implementing a source separation model that may be applied to user generated content (UGC).BACKGROUND

[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.

[0004] Music instrument source separation models have a wide range of applications. In the field of music production, source separation models are invaluable. Producers may use source separation models to isolate individual instruments from a mixed audio track. For example, when a guitarist wants to re-record a guitar part in a song that is already in a mixed audio track, a source separation model may be used to separate the guitar part from the rest of the mixed audio track, allowing for easy replacement or modification of the guitar part. The source separation model not only saves the producer’s time but the source separation model also offers more creative freedom to the producers and artists.

[0005] In audio restoration, the source separation models also play an important role. Old recordings often suffer from noise and interference. By separating different instruments with a source separation model, audio engineers may improve the audio quality of the separated instruments by reducing or removing the noise and interference from each individual instrument. For instance, classic vinyl recordings often suffer from crackling sounds which often are the result of static electricity, dust, debris, age, temperature and other conditions that may occur in recording and / or playback. An audio engineer may apply a source separation model to a vinyl recording to separate out the vocals, the individual instruments, and the background noise from each other, enabling the audio engineer to the improve the audio quality in a restorative manner.D25029W001

[0006] In music education, source separation models are also beneficial. Students may use these source separation models to aide in the study of playing techniques of different instruments. By isolating a particular instrument in a complex piece of music, the students may focus on the melody, rhythm, dynamics, and timbre of the music, enhancing the student’s understanding and learning efficiency with respect to the music.

[0007] It is with respect to these and other considerations that the disclosure made herein is presented.BRIEF SUMMARY OF THE DISCLOSURE

[0008] Techniques described herein generally relate to audio processing. Although many techniques are described herein in the context of user generated content and / or consumer products, these techniques may also be applied to professional generated content studio environment and / or professional products.

[0009] Music instrument source separation models are typically trained with Professional Generated Content (PGC) music data recorded in professional studios. Nevertheless, User Generated Content (UGC) music may be recorded in diverse scenes (both indoor and outdoor) and under various noise conditions. A source separation model trained using PGC music may lead to significant quality degradation when applied to UGC music because the UGC music may be recorded in diverse scenes and under various noise and environmental conditions that are not typically present in a controlled professional studio environment where PGC music is recorded. To address this issue, the present disclosure provides techniques for training a source separation model dedicated for user generated content (UGC) as well as techniques for accessing and implementing the trained UGC source separation model.

[0010] Briefly stated, methods, systems, devices, and non-transitory computer-readable media for training a UGC source separation model as well as accessing and implementing a trained UGC source separation model are disclosed. A method may perform training data pair preparation and may perform a segmented training method to generate a trained UGC source separation model.

[0011] In some described embodiments, the techniques described herein relate to a method for training a source separation model for user generated content, the method including: performing training data pair preparation; and performing a segmented training method to generate a trained user generated content (UGC) source separation model.D25029W001

[0012] In some described embodiments, the techniques described herein relate to a method for accessing a trained user generated content (UGC) source separation model, the method including: requesting access to a trained UGC source separation model: and receiving access to the trained UGC source separation model that is requested.

[0013] In some described embodiments, the techniques described herein relate to a source separation method, the method including: receiving user generated audio; and processing, with a trained user generated content (UGC) source separation model, the user generated audio to separate audio sources in the user generated audio.

[0014] Various aspects of the present disclosure provide for processing of audio signals, and effect improvements in at least the technical fields of audio processing, audio encoding, audio decoding, virtual reality, and the like.

[0015] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.

[0016] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims.DESCRIPTION OF THE DRAWINGS

[0017] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:

[0018] FIG. 1 is a flowchart illustrating various example training methods for training a source separation model for user generated content, in accordance with various aspects of the present disclosure.

[0019] FIG. 2 is a flowchart illustrating various examples of the segmented training method to generate trained user generated content (UGC) source separation model of FIG. 1, in accordance with various aspects of the present disclosure.D25029W001

[0020] FIG. 3 is a block diagram illustrating a first segment of the segmented training method of FIG. 2, in accordance with various aspects of the present disclosure.

[0021] FIG. 4 is a block diagram illustrating a second segment of the segmented training method of FIG. 2, in accordance with various aspects of the present disclosure.

[0022] FIG. 5 is a block diagram illustrating various preprocessing methods of one or more original songs, in accordance with various aspects of the present disclosure.

[0023] FIG. 6 is a block diagram illustrating various example preprocessing methods of one or more original songs, in accordance with various aspects of the present disclosure.

[0024] FIG. 7 illustrates a block diagram of an immersive voice and audio services (IVAS) coder / decoder ("codec") framework for encoding and decoding IVAS bitstreams, according to one or more embodiments.

[0025] FIG. 8 illustrates a schematic block diagram of an example device architecture (e.g., an apparatus) that may be used to implement various aspects of the present disclosure.

[0026] FIG. 9 illustrates a schematic block diagram of a first example CPU corresponding to the CPU implemented in the device architecture of FIG. 8 that may be used to implement various aspects of the present disclosure.

[0027] FIG. 10 illustrates a schematic block diagram of a second example CPU corresponding to the CPU implemented in the device architecture of FIG. 8 that may be used to implement various aspects of the present disclosure.

[0028] FIG. 11 is a flowchart illustrating various example methods for accessing a trained user generated content (UGC) source separation model, in accordance with various aspects of the present disclosure.

[0029] FIG. 12 is a flowchart illustrating various example source separation methods, in accordance with various aspects of the present disclosure.

[0030] FIG. 13 is a flowchart illustrating various examples for training a source separation model for user generated content, in accordance with various aspects of the present disclosure.DETAILED DESCRIPTION

[0031] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of various described embodiments with reference to the accompanyingD25029W001 drawings. The illustrative embodiments in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes made, without departing from the spirit or scope of the present disclosure. In light of the present disclosure, it will be apparent to one of ordinary skill in the art that the various described features and implementations may be practiced without many of these specific details. In some instances, well-known methods, procedures, components, and circuits, have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described hereafter that can each be used independently of one another or with any combination of other features. Thus, the features may be arranged, substituted, combined, separated, or designed into other configurations, which is contemplated in light of the present disclosure.

[0032] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terns are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0033] Various Acronyms that may appear throughout this disclosure and in the associated claims and / or drawings are listed below. Other commonly used acronyms and terms of art may be excluded from this list in the interest of brevity. Thus, a short list of acronyms is provided below as an easy reference for the reader.IVAS - Immersive Voice and Audio ServicesMD - Spatial MetadataBS - BitstreamEVS - Enhanced Voice ServicesD25029W001 IR - Impulse ResponseUGC - User Generated ContentPGC - Professionally Generated ContentSPAR - Spatial ReconstructionDirAC - Directional Audio CodingML - Machine LearningDNN - Deep Neural Network

[0034] The training data for music separation may include data for a clean (e.g., noise free) music recording that includes a set of separate audio tracks that can be combined together in a mix. An example set of audio tracks may include tracks for vocal, drum, guitar, bass, keyboard, as well as any other suitable tracks that may be required for a desired mix. The original training data may be organized by song. For example, a song may be stored as a folder named "songl ", and within this folder, there are individual files for each audio track, such as "vocal.wav", "drum.wav", “guitar.wav”, "bass.wav", “keyboard.wav”, and "others.wav". The set of audio tracks may be collectively mixed together to obtain a mix, such as "mixture.wav".

[0035] The mix can be utilized as input data to a training method, where the desired output corresponds to one of the individual audio tracks. For example, when training to separate the vocal track from the mix, the input of the training data is "mixture.wav", and the desired output of the training data is "vocal.wav". Similarly, when training to separate the drum track from the mix, the input of the training data is "mixture.wav", and the desired output of the training data is "drum.wav".

[0036] Since the number of original songs available for training may be limited, in order to increase the training data and make the trained UGC source separation model more robust across diverse usage scenarios, data augmentation may be used in training the UGC source separation model. Data augmentation involves manipulating one or more of the original songs with one or more augmentation operations, individually, sequentially, or a combination thereof as described in greater detail below.

[0037] FIG. 1 is a flowchart illustrating various example training methods 100 for training a source separation model for user generated content, in accordance with various aspects of the present disclosure. The methods 100 may be performed by one or more processors, which may be configured to perform methods 100 via machine-executable instructions. The methods 100 may be broken into various blocks or partitions, such as blocks 102 and 104. The variousD25029W001 process blocks illustrated in FIG. 1 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 102.

[0038] An example method for generating training data pairs is to use different tracks of the same song as a single set of training data and sum them up to obtain "mixture.wav". However, a problem with this example approach is that when the number of training songs is small, the combinations of different vocals and different instruments are limited by those covered by these songs. Such a lack of diversity in the training data may adversely impact the performance of the resulting source separation model in practical applications.

[0039] To address the above shortcomings, training data is prepared at block 102, “Performing Training Data Pair Preparation”, which may include randomly selecting different tracks from different songs as a set of training data. In one example implementation with randomly combined tracks in the data set, certain data augmentation methods (as described in greater detail below) may be excluded so that excessive differences within the same data set are avoided. For example, a data augmentation operation that applies convolution with different room impulse responses (IRs) may not be combined with a randomly adjusted energy operation. Processing may proceed from block 102 to block 104.

[0040] An additional consideration with respect to data augmentation is that adding all the data augmentations at once may result in overly diverse training data at the beginning of the model training process, which may make it difficult for the model parameters of the source separation model to achieve a successful convergence. To reduce or eliminate the likelihood of overly diverse training data at the beginning of the model training process, segmented training is employed at block 104, “Performing Segmented Training Method to Generate a Trained User Generated Content (UGC) Source Separation Model.”

[0041] FIG. 2 is a flowchart illustrating various examples 200 of the segmented training method to generate a trained user generated content (UGC) source separation model 104 of FIG. 1, in accordance with various aspects of the present disclosure. The methods 200 may be performed by one or more processors, which may be configured to perform methods 200 via machineexecutable instructions. The methods 200 may be broken into various blocks or partitions, such as blocks 202 and 204. The various process blocks illustrated in FIG. 2 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added,D25029W001 combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 202.

[0042] At block 202, “Training Basic Source Separation Model on Dataset with Limited Data Augmentation” may include training a basic source separation model using a dataset with limited data augmentation so that the basic source separation model may converge more easily.Processing may proceed from block 202 to block 204.

[0043] At block 204, “Adapting Basic Source Separation Model to Trained User Generated Content (UGC) Source Separation Model by Training Basic Source Separation Model with Second Training Dataset Including More Data Augmentation than Training Dataset,” may include adapting the basic source separation model to a trained UGC source separation model.

[0044] FIG. 3 is a block diagram illustrating a first segment 300 of the segmented training method of FIG. 2, in accordance with various aspects of the present disclosure. The first segment 300 corresponds to, block 202, “Training Basic Source Separation Model on Dataset with Limited Data Augmentation.”

[0045] The first segment 300 includes an electronic processor applying preprocessing 304 to raw audio data 302. For example, the preprocessing 304 may include a first augmentation operation (as described in greater detail below in FIGS. 5 and 6) to generate a dataset with limited augmentation 306. For an example where there are N possible data augmentations available, a limited number of data augmentations, n, may correspond to n < N or n <= N-l, such as for example n = 1, 2, ... N-l of the possible data augmentations.

[0046] In some examples, where there are a possibility of five different augmentation operations available, N = 5, and up to four different augmentations would correspond to the limited number of data augmentations, n = 1,2,3 or 4. When the training dataset used to train the basic source separation model is a single augmentation operation, the augmentation operation may be referred to as a first augmentation operation. In some other examples, the training dataset used to train the basic source separation model is augmented with two augmentation operations corresponding to a first augmentation operation and a second augmentation operation. In still other examples, the training dataset used to train the basic source separation model may be augmented with three augmentation operations corresponding to a first augmentation operation, a second augmentation operation, and a third augmentation operation. In yet other examples, the training dataset used to train the basic source separation model is augmented with four augmentation operationsD25029W001 corresponding to a first augmentation operation, a second augmentation operation, a third augmentation operation, and a fourth augmentation operation.

[0047] The first segment 300 further includes the electronic processor applying a learning algorithm 308 (e.g., a machine learning algorithm or neural network) to generate a basic user generated content (UGC) source separation model 312 by training the basic UGC source separation model 312 to recognize patterns in the dataset with limited augmentation 306.

[0048] The first segment 300 further includes a forward pass 310 in which the dataset with limited augmentation 306 is passed through the basic UGC source separation model 312, and the basic UGC source separation model 312 makes predictions based on the dataset with limited augmentation 306. The model’s predictions are comparted to actual values 314 using a loss function to calculate the error. The loss function is defined to measure how far a model’s predictions are from the actual values.

[0049] The first segment 300 also includes a backward pass 316 (also referred to as “backpropagation”) in which the error is propagated backward through the basic UGC source separation model 312 to adjust internal parameters (weights and / or biases) of the basic UGC source separation model 312. For example, the electronic processor may implement an optimizer that uses gradients from backpropagation to update the parameters of the basic UGC source separation model 312 and minimize the loss function (also called “convergence” herein). In some examples, the first segment 300 may further include the electronic processor processing the dataset with limited data augmentation as a whole. In other examples, the first segment 300 may further include the electronic processor performing batching on the dataset with limited data augmentation.

[0050] In some examples, the first segment 300 may further include the electronic processor performing training within a single epoch to generate the basic UGC separation model 312. In other examples, the first segment 300 may further include the electronic processor performing training over more than one epoch to generate the basic UGC separation model 312.

[0051] In some examples, the first segment 300 may further include the electronic processor performing regularization to minimize or eliminate overfitting. In other examples, the first segment 300 may further include the electronic processor performing performance monitoring by monitoring metrics (e.g., accuracy, loss, precision, or recall) to ensure the basic UGC source separation model 312 is continually improving during the first segment 300. In these other examples, the first segment 300 may also include the electronic processor performing earlyD25029W001 stopping when the monitored metrics indicate that the basic UGC source separation model 312 is no longer improving.

[0052] FIG. 4 is a block diagram illustrating a second segment 400 of the segmented training method of FIG. 2, in accordance with various aspects of the present disclosure. The second segment 400 corresponds to, block 204, “Adapting Basic Source Separation Model to Trained User Generated Content (UGC) Source Separation Model by Training Basic Source Separation Model with Second Training Dataset Including More Data Augmentation than Training Dataset.”

[0053] The second segment 400 includes an electronic processor applying preprocessing 404 to raw audio data 402. For example, the preprocessing 404 may include a plurality of augmentation operation (as described in greater detail below in FIG. 6) to generate a second training dataset 406 with more augmentation than the dataset with limited augmentation 306. For an example where there are N possible data augmentations available, a limited number of data augmentations, n, may correspond to n < N or n <= N-l, such as for example n = 1, 2, ... N-l of the possible data augmentations. In this example, the second training dataset 406 has m number of data augmentations, wherein n < m <= N.

[0054] The second segment 400 further includes the electronic processor applying a learning algorithm 408 (c.g., a machine learning algorithm or neural network) to adapt the basic user generated content (UGC) source separation model 312 of FIG. 3 to a trained user generated content (UGC) source separation model 412 by training the basic UGC source separation model 312 to recognize patterns in the second training dataset 406.

[0055] The second segment 400 further includes a forward pass 410 in which the second training dataset 406 is passed through the trained UGC source separation model 412, and the trained UGC source separation model 412 makes predictions based on the second training dataset 406. The model’s predictions are comparted to actual values 414 using a loss function to calculate the eiTor. The loss function is defined to measure how far a model’s predictions are from the actual values.

[0056] The second segment 400 also includes a backward pass 416 (also referred to as “backpropagation”) in which the error is propagated backward through the trained UGC source separation model 412 to adjust internal parameters (weights and / or biases) of the trained UGC source separation model 412. For example, the electronic processor may implement anD25029W001 optimizer that uses gradients from backpropagation to update the parameters of the trained UGC source separation model 412 and minimize the loss function (also called “convergence” herein).

[0057] In some examples, the second segment 400 may further include the electronic processor processing the second training dataset 406 as a whole. In other examples, the second segment 400 may further include the electronic processor performing batching on the second training dataset 406.

[0058] In some examples, the second segment 400 may further include the electronic processor adapting the basic UGC source separation model 312 within a single epoch to generate the trained UGC separation model 412. In other examples, the second segment 400 may further include the electronic processor adapting the basic UGC source separation model 312 over more than one epoch to generate the trained UGC separation model 412.

[0059] In some examples, the second segment 400 may further include the electronic processor performing regularization to minimize or eliminate overfitting. In other examples, the second segment 400 may further include the electronic processor performing performance monitoring by monitoring metrics (e.g., accuracy, loss, precision, or recall) to ensure the trained UGC source separation model 412 is continually improving during the second segment 400. In these other examples, the second segment 400 may also include the electronic processor performing early stopping when the monitored metrics indicate that the trained UGC source separation model 412 is no longer improving.

[0060] In this way, the trained UGC source separation model starts from a basic source separation model and the model training process is more likely to achieve convergence, while also having been exposed to more diverse data.

[0061] In some examples, the second training dataset 406 used to train the trained UGC source separation model is augmented with four of the five augmentation operations described in FIG. 6 for all audio tracks. In other examples, the second training dataset used to train the trained UGC source separation model is augmented with all five of the augmentation operations described in FIG. 6 for all audio tracks.

[0062] In yet other examples, the second training dataset 406 used to train the trained UGC source separation model is augmented with four of the five augmentation operations described in FIG. 6 for a single audio track (e.g., the vocal track). In these examples, the second training dataset used to train the trained UGC source separation model is augmented with all five of theD25029W001 augmentation operations described in FIG. 6 for all audio tracks other than the single audio track (e.g., the vocal track).

[0063] Referring back to FIG. 1, the segmented training method 104 has proven valuable in decreasing the size of a source separation model. In the absence of the segmented training method 104 as disclosed herein, an example source separation model may not converge during model training with the dataset augmented by all data augmentation techniques, and further may not achieve the decreased size realized through segmented training. However, by employing the segmented training method 104, the model size and the complexity of the trained UGC source separation model 412 of FIG. 4 may be reduced while maintaining comparable performance. In some examples, the complexity of the trained UGC source separation model 412 was reduced by over tenfold relative to the other source separation models by using the segmented training method 104.

[0064] FIG. 5 is a block diagram illustrating various preprocessing methods 500 of one or more original songs 502, in accordance with various aspects of the present disclosure. The methods 500 may be performed by one or more processors, which may be configured to perform methods 500 via machine-executable instructions. The methods 500 may be broken into various blocks or partitions, such as blocks 502 and 504. The various process blocks illustrated in FIG. 5 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 502.

[0065] At block 502, “One or More Original Songs” may correspond to a dataset (e.g., the raw audio data 302) that requires augmentation to be used in model training for a user generated content source separation model. At block 504, “First Augmentation” may be one of a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, or a noise addition operation.

[0066] In the example augmentation methods 500, the first augmentation operation 504 is the only augmentation operation applied to a first song of the one or more original songs 502.

[0067] In some examples, the pitch shift operation may be applied to all songs or less than all songs. In one example, the pitch shift operation may include randomly choosing a pitch shift amount based on a number of semitones from the original pitch, e.g., a random selection from a set of semitone pitch shift amounts such as [-2, -1, 0, +1, +2], In another example, the pitch shiftD25029W001 operation may include randomly choosing a pitch shift based on a number of cents from the original pitch, where 100 cents is equal to a semitone, e.g., a random selection from a set of cent-based pitch shift amounts such as [-500, -250, -110, -50, 0, +120, +270, +505], In still another example, the pitch shift operation may include randomly choosing a pitch shift based on a percentage or fractional tonal shift from the original pitch, e.g., a random selection from a set of percentage-based pitch shift amounts such as [-20%, -15%, -10%, -5%, 0, +4%, +12%, +18%].

[0068] In some examples, the time stretch operation may be applied to all songs or less than all songs. For example, the time stretch operations includes randomly choosing a time stretch ratio from [0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3],

[0069] In some examples, a room impulse response (IR) operation is performed by convolving one or more audio tracks of a song with a room impulse response (IR). The result of the convolution may be used to approximate the reverb and other characteristics of a particular environment (e.g., different room sizes with different echo and reflective characteristics). The IR model itself is created by collecting data (e.g., frequency response, magnitude, phase, timedomain characteristics, etc.) that is captured from a room when a test audio signal is output into the room. Collected reverberation data from various rooms can be used in different IR models to enhance the adaptability of the data to different environments. In a first implementation of the room impulse response operation, different tracks of a song may be convolved with different room impulse responses (IRs). However, this approach may cause significant differences among different tracks of the same song, which may bring some difficulties in model convergence during model training of the source separation model. In a second implementation of the room impulse response operation, instead of convolving the different tracks of a song with the reverberation of various rooms, all different tracks of the same song may be convolved with the same room impulse response. In the second implementation, when using this batch of data for training, convergence may be easier to achieve relative to the first implementation.

[0070] In the event that the room impulse response (IR) operation (e.g., data augmentation via convolving with the room impulse responses) is not used for augmentation, the trained source separation model may be incapable of separating the user generated contents that are recorded in reverberant environments. However, when this room impulse response (IR) operation is applied for augmentation, the trained source separation model has a higher likelihood of separating the contents recorded in such reverberant environments.

[0071] An example problem in user generated content (UGC) is that speech-based audio with extremely low energy may be regarded as noise. Therefore, an energy adjustment operation mayD25029W001 be performed to adjust the energy of the original data to enhance the diversity of energy. For example, a minimum energy adjustment coefficient, for example, -20dB may be set for the original data, and the energy may be randomly adjusted within the range of [-20dB, OdB],

[0072] The original PGC music rarely contains noise due to being recorded in a controlled professional environment. Therefore, a noise addition operation may be used to add noise to the original data to improve the robustness of the resulting trained source separation model with respect to noisy environments. In one example, when the goal is to extract clean vocal, bass, guitar, keyboard, and drums, noise may be added to the "others" track. In this way, the "mixture" will contain noise, while the output target vocal, bass, guitar, keyboard, and drums are still considered “clean” music (e.g., noise free music). In a different example, when the goal is to separate the drums, to avoid suppressing some of the separated percussion instruments, percussion noise should be limited when training the source separation model to improve separation for drums.

[0073] When data augmentation through the noise addition operation is not employed in preparing the data for model training, the resulting trained source separation model may be unable to effectively separate the user generated contents that are recorded in noisy environments. However, when the noise addition operation is utilized in preparing the data for model training, the resulting trained source separation model has a higher probability of successfully separating the user generated contents recorded in noisy environments.

[0074] Additionally, when the energy adjustment operation is not carried out in preparing the data for model training, the resulting trained source separation model may fail to separate low-energy user generated contents. On the other hand, with the energy adjustment operation, the resulting trained source separation model is more likely to separate low-energy user generated contents.

[0075] Moreover, while the above examples reference the first song of the one or more original songs 502, the above examples are applicable to any or all of the one or more original songs 502. Additionally, the one or more original songs 502 may be included in the training dataset or the second training dataset as described above in FIG. 2 after augmentation by any of the pitch shift operation, the time stretch operation, the room impulse response operation, the energy adjustment operation, or the noise addition operation.

[0076] FIG. 6 is a block diagram illustrating various example preprocessing methods 600 of one or more original songs 602, in accordance with various aspects of the present disclosure. TheD25029W001 methods 600 may be performed by one or more processors, which may be configured to perform methods 600 via machine-executable instructions. The methods 600 may be broken into various blocks or partitions, such as blocks 602 and 604. The various process blocks illustrated in FIG. 6 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 602.

[0077] Similar to FIG. 5, at block 602, “One or More Original Songs” may be a dataset (e.g., the raw audio data 402) that requires augmentation to be used in training a user generated content (UGC) source separation model. At block 604, “First Augmentation” may be one of a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, or a noise addition operation as described above in FIG. 5 applied to a first song of the original songs 602.

[0078] However, at block 606, “Second Augmentation” may be sequentially applied to a first song of the original songs 602 after the “First Augmentation." The “Second Augmentation” is different from the “First Augmentation” and may be one of the pitch shift operation, the time stretch operation, the room impulse response operation, the energy adjustment operation, or the noise addition operation as described above in FIG. 5.

[0079] At block 608, “Optional Third Augmentation” is optional and may be sequentially applied to a first song of the original songs 602 after the “First Augmentation” and the “Second Augmentation” are sequentially applied. The “Optional Third Augmentation” is different from the “First Augmentation” and the “Second Augmentation,” and may be one of the pitch shift operation, the time stretch operation, the room impulse response operation, the energy adjustment operation, or the noise addition operation as described above in FIG. 5.

[0080] At block 610, “Optional Fourth Augmentation” is optional and may be sequentially applied to a first song of the original songs 602 after the “First Augmentation,” the “Second Augmentation,” and the “Optional Third Augmentation” are sequentially applied. The “Optional Fourth Augmentation” is different from the “First Augmentation,” the “Second Augmentation,” and the “Optional Third Augmentation,” and may be one of the pitch shift operation, the time stretch operation, the room impulse response operation, the energy adjustment operation, or the noise addition operation as described above in FIG. 5.D25029W001

[0081] At block 612, “Optional Fifth Augmentation” is optional and may be sequentially applied to a first song of the original songs 602 after the “First Augmentation,” the “Second Augmentation,” the “Optional Third Augmentation,” and the “Optional Fourth Augmentation” are sequentially applied. The “Optional Fifth Augmentation” is different from the “First Augmentation,” the “Second Augmentation,” the “Optional Third Augmentation,” and the “Optional Fourth Augmentation,” and may be one of the pitch shift operation, the time stretch operation, the room impulse response operation, the energy adjustment operation, or the noise addition operation as described above in FIG. 5.

[0082] Moreover, while the above examples reference the first song of the one or more original songs 602, the above examples are applicable to any or all of the one or more original songs 602. Additionally, the one or more original songs 602 may be included in the training dataset or the second training dataset as described above in FIG. 2 after augmentation by any combination of the first through fifth augmentation operations.

[0083] FIG. 7 illustrates a block diagram of an immersive voice and audio services (IVAS) coder / decoder (“codec”) framework 700 for encoding and decoding IVAS bitstreams, according to one or more embodiments. IVAS is expected to support a range of audio service capabilities, including but not limited to mono to stereo upmixing and fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including but not limited to: mobile and smart phones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theatre devices, and other suitable devices.

[0084] The example IVAS codec 700 includes an IVAS encoder 701 and an IVAS decoder 704. The IVAS encoder 701 may be considered a first device that is located upstream from a second device (e.g., a mobile device) such as the IVAS decoder 704. Thus, the first device may also be referred to as an upstream device or an upstream encoder device, while the second device may be referred to as a downstream device or a downstream decoder device.

[0085] The IVAS encoder 701 includes a spatial encoder 702 and a core audio encoder 703. The input of the spatial encoder 702 corresponds to a first path 710. The spatial encoder 702 receives input audio (e.g., input audio content) via the first path 710. The spatial encoder 702 processes and encodes the received input audio. In some implementations, the spatial encoder 702 implements SPAR and DirAC for analyzing / downmixing N_dmx spatial audio channels, as described in further detail below. In some implementations, the spatial encoder 702 may also implement a trained user generated content (UGC) source separation model as described herein.D25029W001 The outputs of the spatial encoder 702 correspond to a second path 711 and a third path 712. The spatial encoder 702 is coupled to the core audio encoder 703 via the second path 711. The spatial encoder 702 is coupled to the IVAS decoder 704 via the third path 712. The output of the spatial encoder 702 includes a spatial metadata (MD) bitstream (BS) and N_dmx channels of spatial downmix. The N_dmx channels of spatial downmix are provided by the spatial encoder 702 to the core audio encoder 703 via the second path 711. The spatial MD BS is provided by the spatial encoder 702 to the IVAS decoder 704 via the third path 712. The spatial MD is quantized and entropy coded. In some implementations, quantization can include fine, moderate, coarse, and extra coarse quantization strategies and entropy coding can include Huffman or Arithmetic coding. The framework permits not more than three levels of quantization at a given operating mode; however, with decreasing bitrates, the three levels become increasingly coarser overall, to meet bitrate requirements.

[0086] The input of the core audio encoder 703 corresponds to the second path 711. The output of the core audio encoder 703 corresponds to a fourth path 713. The core audio encoder 703 (e.g., based on mono Enhanced Voice Services (EVS) encoding unit) encodes N_dmx channels (N_dmx = 1-16 channels) of the spatial downmix into an audio bitstream, which is combined (via the fourth path 713) with the spatial MD bitstream into an IVAS encoded bitstream transmitted to IVAS decoder 704 via the third path 712.

[0087] The IVAS decoder 704 includes a core audio decoder 705 (e.g., an EVS decoder) and a spatial decoder / renderer 706 (e.g., SPAR / DirAC). The input of the core audio decoder 705 corresponds to a fifth path 714. The core audio decoder 705 receives the audio bitstream via the fifth path 714. The core audio decoder 705 is configured to decode the audio bitstream extracted from the IVAS bitstream to recover the N_dmx audio channels. The output of the core audio decoder 705 corresponds to a sixth path 715. The core audio decoder 705 is coupled to the spatial decoder / renderer 706 via the sixth path 715. The core audio decoder 705 is configured to provide the N_dmx audio channels (e.g., the decoded spatial downmix) to the spatial decoder / renderer 706 via the sixth path 715.

[0088] The inputs of the spatial decoder / renderer 706 correspond to the third path 712 and the sixth path 715. The spatial decoder / renderer 706 receives the spatial MD bitstream from the IVAS encoder 701 via the third path 712 and receives the decoded spatial downmix from the core audio decoder 705 via the sixth path 715. The spatial decoder / renderer 706 decodes the spatial MD bitstream extracted from the IVAS bitstream to recover the spatial MD and synthesizes (e.g., renders) output audio channels using the spatial MD and a spatial upmix forD25029W001 playback on various audio systems with different speaker configurations and capabilities. The output of the spatial decoder / renderer 706 corresponds to a seventh path 716. The spatial decoder / renderer 706 provides the output audio (e.g., decoded audio) via the seventh path 716. In some implementations, the spatial decoder / renderer 706 may also implement a source combiner that combines sources separated by a trained user generated content (UGC) source separation model as described herein.

[0089] FIG. 8 illustrates a schematic block diagram of an example device architecture 800 (e.g., an apparatus 800) that may be used to implement various aspects of the present disclosure. Architecture 800 includes but is not limited to servers and client devices, systems, and methods as described in reference to FIGS. 1-7, 11, and 12. As shown, the architecture 800 includes central processing unit (CPU) 801 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 802 or a program loaded from, for example, storage unit 808 to random access memory (RAM) 803. The CPU 801 may be, for example, an electronic processor 801. In RAM 803, the data required when CPU 801 performs the various processes is also stored, as required. CPU 801, ROM 802, and RAM 803 are connected to one another via bus 804. Input / output interface 805 is also connected to bus 804.

[0090] The following components are connected to I / O interface 805: input unit 806, that may include a keyboard, a mouse, or the like; output unit 807 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 808 including a hard disk, or another suitable storage device; and communication unit 809 including a network interface card such as a network card (e.g., wired or wireless).

[0091] In some implementations, input unit 806 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0092] In some implementations, output unit 807 include systems with various number of speakers. Output unit 807 (depending on the capabilities of the hose device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0093] In some embodiments, communication unit 809 is configured to communicate with other devices (e.g., via a network). Drive 810 is also connected to I / O interface 805, as required. Removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or another suitable removable medium is mounted on drive 810, so that a computerD25029W001 program read therefrom is installed into storage unit 808, as required. A person skilled in the art would understand that although apparatus 800 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0094] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 809, and / or installed from the removable medium 811, as shown in FIG. 8.

[0095] FIG. 9 illustrates a schematic block diagram of a first example CPU 901 corresponding to the CPU 801 implemented in the device architecture 800 of FIG. 8 that may be used to implement various aspects of the present disclosure. The CPU 901 includes an electronic processor 920 and a memory 921. The electronic processor 920 is electrically and / or communicatively connected to the memory 921 for bidirectional communication. The memory 921 stores a user generated content (UGC) source separation model training software 922 and a trained UGC source separation model 923 that is generated by the UGC source separation model training software 922. In some examples, memory 921 may be located internal to the electronic processor 920, such as for an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 921 may be located external to the electronic processor 920, such as in a ROM 802, a RAM 803, flash memory or a removable medium 811, or another non-transitory computer readable medium that is contemplated for device architecture 800. In some instances, the electronic processor 920 may implement the user generated content (UGC) source separation model training software 922 stored in the memory 921 to perform, among other things, any of the method 100 of FIG. 1, the method 200 of FIG. 2, the first segment 300 of FIG. 3, the second segment 400 of FIG. 4, the method 500 of FIG. 5, the method 600 of FIG. 6, the method 1100 of FIG. 11, the method 1200 of FIG. 12, and / or the method 1300 of FIG. 13.

[0096] FIG. 10 illustrates a schematic block diagram of a second example CPU 1001 corresponding to the CPU 801 implemented in the device architecture 800 of FIG. 8 that may be used to implement various aspects of the present disclosure. The second example CPU 1001D25029W001 includes an electronic processor 1020 and a memory 1021. The electronic processor 1020 is electrically and / or communicatively connected to the memory 1021 for bidirectional communication. The memory 1021 stores a user generated content (UGC) source separation application software 1022 and the trained UGC source separation model 923 that is generated by the UGC source separation model training software 922 of FIG. 9. In some examples, memory 1021 may be located internal to the electronic processor 1020, such as for an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 1021 may be located external to the electronic processor 1020, such as in a ROM 802, a RAM 803 , flash memory or a removable medium 811 , or another non-transitory computer readable medium that is contemplated for device architecture 800. In some instances, the electronic processor 1020 may implement user generated content (UGC) source separation application software 1022 stored in the memory 921 to perform, among other things, any of the method 1100 of FIG. 11, and / or the method 1200 of FIG. 12.

[0097] FIG. 11 is a flowchart illustrating various example methods 1100 for accessing a trained user generated content (UGC) source separation model, in accordance with various aspects of the present disclosure. The methods 1100 may be performed by one or more processors, which may be configured to perform methods 1100 via machine-executable instructions. For ease of understanding, FIG. 11 is also described with respect to the device architecture 800 of FIG. 8, and in particular, the second example CPU 1001 of FIG. 10. The methods 1100 may be broken into various blocks or partitions, such as blocks 1102 and 1104. The various process blocks illustrated in FIG. 11 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 1102.

[0098] At block 1102, “Requesting Access To Trained UGC Source Separation Model,” an example method 1100 may include requesting access to a trained UGC source separation model. For example, the electronic processor 1020 requests access to the trained UGC source separation model 923 stored in the memory 1021. In another example, the electronic processor 1020 requests access to a trained UGC source separation model stored in a memory or a database that is separate from the second example CPU 1001 (e.g., the storage unit 808, the removable medium 811, the memory 921, cloud storage, and / or other suitable storage). Processing may proceed from block 1102 to block 1104.D25029W001

[0099] At block 1104, “Receiving Access To Trained UGC Source Separation Model That Is Requested,’’ an example method 1100 may include receiving access to the trained UGC source separation model that is requested. For example, the electronic processor 1020 receives access to the trained UGC source separation model 923 stored in the memory 1021. In another example, the electronic processor 1020 receives access to a trained UGC source separation model stored in a memory or a database that is separate from the second example CPU 1001 (e.g., the storage unit 808, the removable medium 811, the memory 921, cloud storage, and / or other suitable storage).

[0100] In some examples, receiving access to the trained UGC source separation model further includes receiving the trained UGC source separation model as part of an audio software package (e.g., an application for a mobile device that records user generated content). In other examples, receiving access to the trained UGC source separation model further includes receiving the trained UGC source separation model as part of a plugin to an audio software package.

[0101] FIG. 12 is a flowchart illustrating various example source separation methods 1200, in accordance with various aspects of the present disclosure. The methods 1200 may be performed by one or more processors, which may be configured to perform methods 1200 via machine-executable instructions. For ease of understanding, FIG. 12 is described with respect to the device architecture 800 of FIG. 8, and in particular, the second example CPU 1001 of FIG.10. However, the various example source separation methods 1200 of FIG. 12 are also applicable to other devices, for example, the IVAS compliant device of FIG. 7.

[0102] The methods 1200 may be broken into various blocks or partitions, such as blocks 1202 and 1204. The various process blocks illustrated in FIG. 12 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 1202.

[0103] At block 1202, “Receiving User Generated Audio,’’ an example method 1200 may include receiving user generated audio content. For example, the electronic processor 1020 receives the user generated audio from the memory 1021. In yet another example, the electronic processor 1020 receives the user generated audio from an electronic source that is external to the memory 1021 (e.g., the storage unit 808, the removable medium 811, the memory 921, cloudD25029W001 storage, and / or other suitable storage). With respect to the IVAS encoder 701, the spatial encoder 702 receives the input audio.

[0104] In some examples, receiving the user generated audio further includes recording audio to generate the user generated audio. For example, the electronic processor 1020 stores data generated by the input unit 806 in the storage unit 808 and / or the removable medium 811 as the user generated audio. Processing may proceed from block 1202 to block 1204.

[0105] At block 1204, “Processing, With Trained User Generated Content (UGC) Source Separation Model, User Generated Audio to Separate Audio Sources in User Generated Audio,” may include the example source separation method 1200 processing, with a trained user generated content (UGC) source separation model, the user generated audio to separate audio sources in the user generated audio. For example, the electronic processor 1020 processes, with the trained user generated content (UGC) source separation model 923, the user generated audio from the memory 1021 to separate audio sources in the user generated audio. In yet another example, the electronic processor 1020 processes, with the trained user generated content (UGC) source separation model 923, the user generated audio from an electronic source that is external to the memory 1021 to separate audio sources in the user generated audio (e.g., the storage unit 808, the removable medium 811, the memory 921, cloud storage, and / or other suitable storage). Additionally, with respect to the IVAS encoder 701, the spatial encoder 702 may process, with a trained user generated content (UGC) source separation model, the user generated audio to separate audio sources in the user generated audio. In some examples, the spatial metadata (MD) bitstream (BS), the N_dmx channels of spatial downmix, or a combination thereof may be based on the audio sources separated by the trained UGC source separation model.

[0106] FIG. 13 is a flowchart illustrating various examples for training a source separation model for user generated content 1300, in accordance with various aspects of the present disclosure. The methods 1300 may be performed by one or more processors, which may be configured to perform methods 1300 via machine-executable instructions. For ease of understanding, FIG. 13 is described with respect to the device architecture 800 of FIG. 8, and in particular, the second example CPU 1001 of FIG. 10. However, the various example source separation methods 1300 of FIG. 13 are also applicable to other devices, for example, the IVAS compliant device of FIG. 7.

[0107] The methods 1300 may be broken into various blocks or partitions, such as blocks 1302-1312. The various process blocks illustrated in FIG. 13 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added,D25029W001 combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 1302.

[0108] At block 1302, “Obtaining Raw Audio Data That Includes Plurality Of Randomly Selected Songs,” an example method 1300 may include obtaining raw audio data that includes a plurality of randomly selected songs, wherein each of the songs includes: a set of individual audio tracks for the song, each individual audio track corresponding to an isolated recording; and a mix of the set of individual audio tracks for the song. For example, the electronic processor 1020 receives raw audio data that includes a plurality of randomly selected songs, wherein each of the songs includes: a set of individual audio tracks for the song, each individual audio track corresponding to an isolated recording; and a mix of the set of individual audio tracks for the song from the memory 1021. With respect to the IV AS encoder 701, the spatial encoder 702 receives the input audio that includes a plurality of randomly selected songs, wherein each of the songs includes: a set of individual audio tracks for the song, each individual audio track corresponding to an isolated recording; and a mix of the set of individual audio tracks for the song. Processing may proceed from block 1302 to block 1304.

[0109] At block 1304, “Evaluating Raw Audio Data To Identify Limited Augmentation Operation,” may include the example source separation method 1300 evaluating the raw audio data to identify a limited augmentation operation, wherein the limited augmentation operation comprises a first portion of a full set of augmentation operations, the full set of augmentation operations including: a pitch shift operation; a time shift operation; a room impulse response operation; an energy adjustment operation; and a noise addition operation. For example, the electronic processor 1020 evaluates the raw audio data to identify a limited augmentation operation, wherein the limited augmentation operation comprises a first portion of a full set of augmentation operations, the full set of augmentation operations including: a pitch shift operation; a time shift operation; a room impulse response operation; an energy adjustment operation; and a noise addition operation. Additionally, with respect to the IVAS encoder 701, the spatial encoder 702 may evaluate the raw audio data to identify a limited augmentation operation, wherein the limited augmentation operation comprises a first portion of a full set of augmentation operations, the full set of augmentation operations including: a pitch shift operation; a time shift operation; a room impulse response operation; an energy adjustment operation; and a noise addition operation. Processing may proceed from block 1304 to block 1306.D25029W001

[0110] At block 1306, “Generating Initial Training Dataset With First Amount Of Data Augmentation,’’ may include the example source separation method 1300 generating an initial training dataset with a first amount of data augmentation by: applying the limited augmentation operation to the raw audio data such that data for each mix of the plurality of songs is adjusted by the first portion to generate a corresponding limited augmented mix; and preparing training data pairs that include each limited augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data. For example, the electronic processor 1020 generates an initial training dataset with a first amount of data augmentation by: applying the limited augmentation operation to the raw audio data such that data for each mix of the plurality of songs is adjusted by the first portion to generate a corresponding limited augmented mix; and preparing training data pairs that include each limited augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data. Additionally, with respect to the IVAS encoder 701 , the spatial encoder 702 may generate an initial training dataset with a first amount of data augmentation by: applying the limited augmentation operation to the raw audio data such that data for each mix of the plurality of songs is adjusted by the first portion to generate a corresponding limited augmented mix; and preparing training data pairs that include each limited augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data. Processing may proceed from block 1306 to block 1308.

[0111] At block 1308, “Generating Basic Source Separation Model,’’ may include the example source separation method 1300 generating a basic source separation model by iteratively applying a first learning algorithm to the initial training dataset with the first amount of data augmentation until a first convergence criteria is satisfied. For example, the electronic processor 1020 generates a basic source separation model by iteratively applying a first learning algorithm to the initial training dataset with the first amount of data augmentation until a first convergence criteria is satisfied. Additionally, with respect to the IVAS encoder 701 , the spatial encoder 702 may generate a basic source separation model by iteratively applying a first learning algorithm to the initial training dataset with the first amount of data augmentation until a first convergence criteria is satisfied. Processing may proceed from block 1308 to block 1310.

[0112] At block 1310, “Generating Subsequent Training Dataset With Second Amount Of Data Augmentation,” may include the example source separation method 1300 generating a subsequent training dataset with a second amount of data augmentation by: applying a second portion of the full set of augmentation operations to the raw audio data such that data for each mix of the plurality of songs is adjusted by the second portion to generate a correspondingD25029W001 second augmented mix; and preparing training data pairs that include each second augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data. For example, the electronic processor 1020 generates a subsequent training dataset with a second amount of data augmentation by: applying a second portion of the full set of augmentation operations to the raw audio data such that data for each mix of the plurality of songs is adjusted by the second portion to generate a corresponding second augmented mix; and preparing training data pairs that include each second augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data. Additionally, with respect to the IV AS encoder 701, the spatial encoder 702 may generate a subsequent training dataset with a second amount of data augmentation by: applying a second portion of the full set of augmentation operations to the raw audio data such that data for each mix of the plurality of songs is adjusted by the second portion to generate a corresponding second augmented mix; and preparing training data pairs that include each second augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data. Processing may proceed from block 1310 to block 1312.

[0113] At block 1312, “Generating Trained Source Separation Model,” may include the example source separation method 1300 generating a trained source separation model by iteratively applying a second learning algorithm to the subsequent training dataset with the second amount of data augmentation until a second convergence criteria is satisfied, where the trained source separation model is initially set to the basic source separation model, and where the second portion is different from, and includes more of the full set of augmentation operations relative to, the first portion. For example, the electronic processor 1020 generates a trained source separation model by iteratively applying a second learning algorithm to the subsequent training dataset with the second amount of data augmentation until a second convergence criteria is satisfied, where the trained source separation model is initially set to the basic source separation model, and where the second portion is different from, and includes more of the full set of augmentation operations relative to, the first portion. Additionally, with respect to the IVAS encoder 701, the spatial encoder 702 may generate a trained source separation model by iteratively applying a second learning algorithm to the subsequent training dataset with the second amount of data augmentation until a second convergence criteria is satisfied, where the trained source separation model is initially set to the basic source separation model, and where the second portion is different from, and includes more of the full set of augmentation operations relative to, the first portion.D25029W001

[0114] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units and modules discussed above can be executed by control circuitry (e.g., CPU 801 in combination with other components of FIG. 8), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0115] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0116] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer programD25029W001 codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0118] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.

[0119] EEE 1. A method (100) for training a source separation model for user generated content, the method comprising: performing (102) training data pair preparation; and performing (104) a segmented training method to generate a trained user generated content (UGC) source separation model.

[0120] EEE 2. The method of EEE 1, wherein performing (102) training data pair preparation further includes randomly selecting different tracks from different songs as a training dataset.

[0121] EEE 3. The method of EEE 2, wherein performing (104) the segmented training method to generate the trained user generated content (UGC) source separation model further includes training (202) a basic source separation model using the training dataset with limited data augmentation, and adapting (204) the basic source separation model to a trained user generated content (UGC) source separation model by training the basic source separation model with a second training dataset including more data augmentation than the training dataset.

[0122] EEE 4. The method of EEE 3, wherein training (202) the basic source separation model using the training dataset with limited data augmentation further includes generating the training dataset with limited data augmentation (306) by applying preprocessing (304) to raw audio data (302), iteratively generating the basic source separation model (312) by applying aD25029W001 learning algorithm (308) to the training dataset with limited data augmentation (306) until an error of the basic source separation model (312) is minimized or stable.

[0123] EEE 5. The method of EEE 4, wherein applying the preprocessing (304) to the raw audio data (302)further includes augmenting (504) a dataset with a first augmentation operation to generate the training dataset with the limited data augmentation.

[0124] EEE 6. The method of EEE 5, wherein the first augmentation operation (504) is one augmentation operation selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0125] EEE 7. The method of EEE 5 or EEE 6, wherein applying the preprocessing (404) to the raw audio data (402)further includes augmenting the dataset with the first augmentation operation (604) and a second augmentation operation (606) to generate the training dataset with the limited data augmentation, wherein the second augmentation operation is different from the first augmentation operation (604).

[0126] EEE 8. The method of EEE 7, wherein the second augmentation operation (606) is one augmentation operation selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0127] EEE 9. The method of EEE 7 or EEE 8, wherein applying the preprocessing (304) to the raw audio data (302)further includes augmenting the dataset with the first augmentation operation (604), the second augmentation operation (606), a third augmentation operation (608) to generate the training dataset with the limited data augmentation, wherein the third augmentation operation (608) is different from the first augmentation operation (604) and the second augmentation operation (606).

[0128] EEE 10. The method of EEE 9, wherein the third augmentation operation (608) is one augmentation operation selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0129] EEE 11. The method of EEE 9 or EEE 10, wherein applying the preprocessing (304) to the raw audio data (302)further includes augmenting the dataset with the first augmentation operation (604), the second augmentation operation (606), the third augmentationD25029W001 operation (608), and a fourth augmentation operation (610) to generate the training dataset with the limited data augmentation, wherein the fourth augmentation operation (610) is different from the first augmentation operation (604), the second augmentation operation (606), and the third augmentation operation (608).

[0130] EEE 12. The method of EEE 11, wherein the fourth augmentation operation (610) is one augmentation operation selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0131] EEE 13. The method of any of EEE 3 to EEE 12, wherein adapting (204) the basic source separation model to the trained user generated content (UGC) source separation model by training the basic source separation model with the second training dataset including more data augmentation than the training dataset further includes generating the second training dataset (406) by applying preprocessing (404) to raw audio data (402), iteratively generating the trained UGC source separation model (412) by applying a learning algorithm (408) to the second training dataset (406) until an error of the trained UGC source separation model (412) is minimized or stable.

[0132] EEE 14. The method of EEE 13, wherein the training dataset is augmented with a single augmentation operation (604), wherein applying preprocessing (404) to raw audio data (402)further includes augmenting a second dataset with two or more augmentation operations (604, 606) to generate the second training dataset, and wherein the two or more augmentation operations (604, 606) are different from each other.

[0133] EEE 15. The method of EEE 14, wherein the two or more augmentation operations (604, 606) are selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0134] EEE 16. The method of EEE 14 or EEE 15, wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

[0135] EEE 17. Ihe method of any of EEE 14 to EEE 16, wherein the second dataset is different from the dataset.

[0136] EEE 18. The method of any of EEE 4 to EEE 17, wherein the training dataset is augmented with two augmentation operations (604, 606), and wherein applying preprocessingD25029W001 (404) to raw audio data (402further includes augmenting a second dataset with three or more augmentation operations (604, 606, 608) to generate the second training dataset, and wherein the three or more augmentation operations (604, 606, 608) are different from each other.

[0137] EEE 19. The method of EEE 18, wherein the three or more augmentation operations (604, 606, 608) are selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0138] EEE 20. The method of EEE 18 or EEE 1 , wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

[0139] EEE 21. The method of any of EEE 18 to EEE 20, wherein the second dataset is different from the dataset.

[0140] EEE 22. The method of any of EEE 4 to EEE 21, wherein the training dataset is augmented with three augmentation operations (604, 606, 608), and wherein applying preprocessing (404) to raw audio data (402further includes augmenting a second dataset with four or more augmentation operations (604, 606, 608, 610) to generate the second training dataset, and wherein the four or more augmentation operations (604, 606, 608, 610) are different from each other.

[0141] EEE 23. The method of EEE 22, wherein the four or more augmentation operations (604, 606, 608, 610) are selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0142] EEE 24. The method of EEE 22 or EEE 23, wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

[0143] EEE 25. The method of any of EEE 22 to EEE 24, wherein the second dataset is different from the dataset.

[0144] EEE 26. The method of any of EEE 4 to EEE 25, wherein the training dataset is augmented with four augmentation operations (604, 606, 608, 610), and wherein applying preprocessing (404) to raw audio data (402 further includes augmenting a second dataset with five augmentation operations (604, 606, 608, 610, 612) to generate the second training dataset, and wherein the five augmentation operations (604, 606, 608, 610, 612) are different from each other.D25029W001

[0145] EEE 27. The method of EEE 26, wherein the five augmentation operations (604, 606, 608, 610, 612) include a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

[0146] EEE 28. The method of EEE 26 or EEE 27, wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

[0147] EEE 29. The method of any of EEE 26 to EEE 28, wherein the second dataset is different from the dataset.

[0148] EEE 30. A method (1100) for accessing a trained user generated content (UGC) source separation model, the method comprising: requesting access to a trained UGC source separation model (1102); and receiving access to the trained UGC source separation model that is requested (1104).

[0149] EEE 31. The method of EEE 30, wherein receiving access to the trained UGC source separation model that is requested further includes receiving the trained UGC source separation model as part of an audio software package.

[0150] EEE 32. Ihe method of EEE 30 or EEE 31, wherein receiving access to the trained UGC source separation model that is requested further includes receiving the trained UGC source separation model as part of a plugin to an audio software package.

[0151] EEE 33. A source separation method (1200), the method comprising: receiving user generated audio (1202); and processing, with a trained user generated content (UGC) source separation model, the user generated audio to separate audio sources in the user generated audio (1204).

[0152] EEE 34. A method (1300) for training a source separation model for user generated content, the method comprising: obtaining raw audio data (1302) that includes a plurality of randomly selected songs, wherein each of the songs includes: a set of individual audio tracks for the song, each individual audio track corresponding to an isolated recording; and a mix of the set of individual audio tracks for the song; evaluating (1304) the raw audio data to identify a limited augmentation operation, wherein the limited augmentation operation comprises a first portion of a full set of augmentation operations, the full set of augmentation operations including: a pitch shift operation; a time shift operation; a room impulse response operation; an energy adjustment operation; and a noise addition operation; generating (1306) an initial training dataset with a first amount of data augmentation by: applying the limited augmentation operationD25029W001 to the raw audio data such that data for each mix of the plurality of songs is adjusted by the first portion to generate a corresponding limited augmented mix; and preparing training data pairs that include each limited augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data; and generating (1308) a basic source separation model by iteratively applying a first learning algorithm to the initial training dataset with the first amount of data augmentation until a first convergence criteria is satisfied; generating (1310) a subsequent training dataset with a second amount of data augmentation by: applying a second portion of the full set of augmentation operations to the raw audio data such that data for each mix of the plurality of songs is adjusted by the second portion to generate a corresponding second augmented mix; and preparing training data pairs that include each second augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data; and generating (1312) a trained source separation model by iteratively applying a second learning algorithm to the subsequent training dataset with the second amount of data augmentation until a second convergence criteria is satisfied, wherein the trained source separation model is initially set to the basic source separation model, and wherein the second portion is different from, and includes more of the full set of augmentation operations relative to, the first portion.

[0153] EEE 35. An apparatus (800) comprising: an electronic processor (801) configured to perform operations including the method of any one of EEE 1 to EEE 34.

[0154] EEE 36. A non-transitory computer-readable storage medium (921, 1021) storing a program of instructions (922, 1022) that is executable by a device (800) to perform the method of any one of EEE 1 to EEE 34.

[0155] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0156] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should beD25029W001 determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0157] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0158] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

Claims

D25029W001CLAIMSWhat is claimed is:

1. A method (100) for training a source separation model for user generated content, the method comprising:performing (102) training data pair preparation; andperforming (104) a segmented training method to generate a trained user generated content (UGC) source separation model.

2. The method of claim 1, wherein performing (102) training data pair preparation further includes randomly selecting different tracks from different songs as a training dataset.

3. The method of claim 2, wherein performing (104) the segmented training method to generate the trained user generated content (UGC) source separation model further includes training (202) a basic source separation model using the training dataset with limited data augmentation, andadapting (204) the basic source separation model to a trained user generated content (UGC) source separation model by training the basic source separation model with a second training dataset including more data augmentation than the training dataset.

4. The method of claim 3, wherein training (202) the basic source separation model using the training dataset with limited data augmentation further includesgenerating the training dataset with limited data augmentation (306) by applying preprocessing (304) to raw audio data (302),iteratively generating the basic source separation model (312) by applying a learning algorithm (308) to the training dataset with limited data augmentation (306) until an error of the basic source separation model (312) is minimized or stable.

5. The method of claim 4, wherein applying the preprocessing (304) to the raw audio data (302) further includesaugmenting (504) a dataset with a first augmentation operation to generate the training dataset with the limited data augmentation.

6. The method of claim 5, wherein the first augmentation operation (504) is one augmentation operation selected from a group consisting of:D25029W001 a pitch shift operation,a time stretch operation,a room impulse response operation,an energy adjustment operation, anda noise addition operation.

7. The method of claim 5 or 6, wherein applying the preprocessing (404) to the raw audio data (402) further includesaugmenting the dataset with the first augmentation operation (604) and a second augmentation operation (606) to generate the training dataset with the limited data augmentation, wherein the second augmentation operation is different from the first augmentation operation (604).

8. The method of claim 7, wherein the second augmentation operation (606) is one augmentation operation selected from a group consisting of:a pitch shift operation,a time stretch operation,a room impulse response operation,an energy adjustment operation, anda noise addition operation.

9. The method of claim 7 or 8, wherein applying the preprocessing (304) to the raw audio data (302) further includesaugmenting the dataset with the first augmentation operation (604), the second augmentation operation (606), a third augmentation operation (608) to generate the training dataset with the limited data augmentation, wherein the third augmentation operation (608) is different from the first augmentation operation (604) and the second augmentation operation (606).

10. fhe method of claim 9, wherein the third augmentation operation (608) is one augmentation operation selected from a group consisting of:a pitch shift operation,a time stretch operation,a room impulse response operation,an energy adjustment operation, andD25029W001 a noise addition operation.

11. The method of claim 9 or 10, wherein applying the preprocessing (304) to the raw audio data (302)further includesaugmenting the dataset with the first augmentation operation (604), the second augmentation operation (606), the third augmentation operation (608), and a fourth augmentation operation (610) to generate the training dataset with the limited data augmentation, wherein the fourth augmentation operation (610) is different from the first augmentation operation (604), the second augmentation operation (606), and the third augmentation operation (608).

12. The method of claim 11, wherein the fourth augmentation operation (610) is one augmentation operation selected from a group consisting of: a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

13. The method of any one of claim 3 to claim 12, wherein adapting (204) the basic source separation model to the trained user generated content (UGC) source separation model by training the basic source separation model with the second training dataset including more data augmentation than the training dataset further includes generating the second training dataset (406) by applying preprocessing (404) to raw audio data (402), iteratively generating the trained UGC source separation model (412) by applying a learning algorithm (408) to the second training dataset (406) until an error of the trained UGC source separation model (412) is minimized or stable.

14. The method of claim 13, wherein the training dataset is augmented with a single augmentation operation (604), wherein applying preprocessing (404) to raw audio data (402) further includes augmenting a second dataset with two or more augmentation operations (604, 606) to generate the second training dataset, and wherein the two or more augmentation operations (604, 606) are different from each other.

15. The method of claim 14, wherein the two or more augmentation operations (604, 606) are selected from a group consisting of:a pitch shift operation,a time stretch operation,D25029W001 a room impulse response operation,an energy adjustment operation, anda noise addition operation.

16. The method of claim 14 or claim 15, wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

17. The method of any one of claim 14 to claim 16, wherein the second dataset is different from the dataset.

18. The method of any one of claim 4 to claim 17, wherein the training dataset is augmented with two augmentation operations (604, 606), and wherein applying preprocessing (404) to raw audio data (402) further includes augmenting a second dataset with three or more augmentation operations (604, 606, 608) to generate the second training dataset, and wherein the three or more augmentation operations (604, 606, 608) are different from each other.

19. The method of claim 18, wherein the three or more augmentation operations (604, 606, 608) are selected from a group consisting of:a pitch shift operation,a time stretch operation,a room impulse response operation,an energy adjustment operation, anda noise addition operation.

20. The method of claim 18 or claim 1 , wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

21. The method of any one of claim 18 to claim 20, wherein the second dataset is different from the dataset.

22. The method of any of claim 4 to claim 21, wherein the training dataset is augmented with three augmentation operations (604, 606, 608), and wherein applying preprocessing (404) to raw audio data (402) further includes augmenting a second dataset with four or more augmentation operations (604, 606, 608, 610) to generate the second training dataset, and wherein the four or more augmentation operations (604, 606, 608, 610) are different from each other.D25029W001 23. The method of claim 22, wherein the four or more augmentation operations (604, 606, 608, 610) are selected from a group consisting of:a pitch shift operation,a time stretch operation,a room impulse response operation,an energy adjustment operation, anda noise addition operation.

24. The method of claim 22 or claim 23, wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

25. The method of any one of claim 22 to claim 24, wherein the second dataset is different from the dataset.

26. The method of any of claim 4 to claim 25, wherein the training dataset is augmented with four augmentation operations (604, 606, 608, 610), and wherein applying preprocessing (404) to raw audio data (402 further includes augmenting a second dataset with five augmentation operations (604, 606, 608, 610, 612) to generate the second training dataset, and wherein the five augmentation operations (604, 606, 608, 610, 612) are different from each other.

27. The method of claim 26, wherein the five augmentation operations (604, 606, 608, 610, 612) include a pitch shift operation, a time stretch operation, a room impulse response operation, an energy adjustment operation, and a noise addition operation.

28. The method of claim 26 or claim 27, wherein the dataset includes one or more original songs (602), and wherein the second dataset is the dataset.

29. fhe method of any one of claim 26 to claim 28, wherein the second dataset is different from the dataset.

30. A method (1100) for accessing a trained user generated content (UGC) source separation model, the method comprising:requesting access to a trained UGC source separation model (1102); andreceiving access to the trained UGC source separation model that is requested (1104).D25029W001 31. The method of claim 30, wherein receiving access to the trained UGC source separation model that is requested further includes receiving the trained UGC source separation model as part of an audio software package.

32. The method of claim 30 or claim 31, wherein receiving access to the trained UGC source separation model that is requested further includes receiving the trained UGC source separation model as part of a plugin to an audio software package.

33. A source separation method (1200), the method comprising:receiving user generated audio (1202); andprocessing, with a trained user generated content (UGC) source separation model, the user generated audio to separate audio sources in the user generated audio (1204).

34. A method (1300) for training a source separation model for user generated content, the method comprising:obtaining raw audio data (1302) that includes a plurality of randomly selected songs, wherein each of the songs includes:a set of individual audio tracks for the song, each individual audio track corresponding to an isolated recording; anda mix of the set of individual audio tracks for the song;evaluating (1304) the raw audio data to identify a limited augmentation operation, wherein the limited augmentation operation comprises a first portion of a full set of augmentation operations, the full set of augmentation operations including:a pitch shift operation;a time shift operation;a room impulse response operation;an energy adjustment operation; anda noise addition operation;generating (1306) an initial training dataset with a first amount of data augmentation by:applying the limited augmentation operation to the raw audio data such that data for each mix of the plurality of songs is adjusted by the first portion to generate a corresponding limited augmented mix; andpreparing training data pairs that include each limited augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data; andD25029W001 generating (1308) a basic source separation model by iteratively applying a first learning algorithm to the initial training dataset with the first amount of data augmentation until a first convergence criteria is satisfied;generating (1310) a subsequent training dataset with a second amount of data augmentation by:applying a second portion of the full set of augmentation operations to the raw audio data such that data for each mix of the plurality of songs is adjusted by the second portion to generate a corresponding second augmented mix; andpreparing training data pairs that include each second augmented mix from the raw audio data, and each corresponding set of individual audio tracks from the raw audio data; andgenerating (1312) a trained source separation model by iteratively applying a second learning algorithm to the subsequent training dataset with the second amount of data augmentation until a second convergence criteria is satisfied,wherein the trained source separation model is initially set to the basic source separation model, andwherein the second portion is different from, and includes more of the full set of augmentation operations relative to, the first portion.

35. An apparatus (800) comprising:an electronic processor (801) configured to perform operations including the method of any one of claims 1 -34.

36. A non- transitory computer- readable storage medium (921. 1021) storing a program of instructions (922, 1022) that is executable by a device (800) to perform the method of any one of claims 1-34.