Training method of sound source separation model, sound source separation method, device, storage medium and program product

CN122511285APending Publication Date: 2026-08-04ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-06-15
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

在此类架构下,上述前级模型与后级模型需要各自预先进行独立训练,并在完成后直接串联使用,因此后级模型在训练时接收的是由未经处理的原始音源合成的理想“其他a”信号,然而,后级模型在实际推理时所实际接收的却是前级模型输出时已产生失真的信号,二者数据分布的不同,导致后级模型泛化能力不足,最终分离出的目标音源音质较差、性能指标偏低

Benefits of technology

[0010]The method provided in this specification proposes that the second separation model undergoes two different training processes: In the first training, the second separation model uses the enhanced mixed signal corresponding to the training samples as input to learn the basic ability to separate the target enhanced track signal from the ideal signal; in the second training, the input of the second separation model is replaced with the predicted enhanced mixed signal actually output by the first separation model, making this signal contain the distortion introduced by the first separation model during the separation process. Because the second separation model directly encounters the distorted signal with the same data distribution as the actual inference stage in the second training, and uses the corresponding enhanced track signal in the training samples as the learning target, the second separation model can actively adapt to the spectral impairments and separation errors in the previous stage output, thereby effectively bridging the data distribution differences between training and inference. After this training, the second separation model can more accurately reconstruct the target enhanced track signal from the damaged enhanced mixed signal output by the first separation model during inference, significantly improving the separation quality and performance indicators of the target audio source.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511285A_ABST
    Figure CN122511285A_ABST
Patent Text Reader

Abstract

This invention provides a training method, device, storage medium, and program product for an audio source separation model. The method includes: acquiring training samples, the training samples including a sample mixed audio signal composed of multiple sample reference track signals; training a first separation model and a second separation model respectively, such that: the first separation model learns to separate the preset track signal and the enhanced mixed signal from the sample mixed audio signal, and the second separation model learns to separate the target enhanced track signal from the enhanced mixed signal corresponding to the training samples; and retraining the second separation model, such that: the second separation model adjusts its parameters based on the predicted target enhanced track signal separated from the predicted enhanced mixed signal output by the first separation model and the corresponding enhanced track signal in the training samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio technology, and in particular to a training method, a method for separating audio sources, an apparatus, a storage medium, and a program product for a sound source separation model. Background Technology

[0002] Audio source separation technology, as an audio signal processing method, is used to recover the individual track signals of each audio source from mixed audio. Among these, audio source separation schemes, represented by separating vocals, bass, drums, and "other" tracks, are relatively mature and can meet basic application requirements such as accompaniment with noise reduction. Building on this, the industry hopes to further separate more audio sources, such as guitars and pianos, from the "other" tracks. However, the spectral characteristics of audio sources such as guitars and pianos are less significant and more difficult to distinguish compared to vocals and drums, making it difficult to achieve practically high separation accuracy for these audio sources. Therefore, this has become a pressing technical problem that needs to be solved.

[0003] In related technologies, a cascaded architecture is typically used to achieve multi-source audio separation. Taking the aforementioned mature audio source separation scheme as an example, such as... Figure 1 As shown, the pre-amplifier model in this scheme can separate vocals, bass, drums, and an "other a" track signal containing all remaining sound sources from the mixed audio. To separate specific sound sources, this cascaded architecture can further introduce n post-amplifier models. Each post-amplifier model receives the "other a" track as input and separates the target sound source track signals such as guitar, piano, etc. from it. If necessary, the guitar and piano track signals in the "other a" track can be eliminated to obtain another "other b" track output by the post-amplifier model. In this architecture, the pre-amplifier model and the post-amplifier model need to be trained independently beforehand and then directly cascaded after completion. Therefore, the post-amplifier model receives the ideal "other a" signal synthesized from the unprocessed original sound sources during training. However, the post-amplifier model actually receives the distorted signal output by the pre-amplifier model during actual inference. The difference in data distribution between the two leads to insufficient generalization ability of the post-amplifier model, resulting in poor sound quality and low performance of the separated target sound sources. Summary of the Invention

[0004] In view of this, the present invention provides a training method, a method for separating audio sources, an apparatus, a storage medium, and a program product for a sound source separation model, in order to address the shortcomings in related technologies.

[0005] Specifically, this specification is implemented through the following technical solution: According to a first aspect of this specification, a training method for a sound source separation model is provided, the method comprising: Acquire training samples, which include a mixed audio signal composed of multiple sample reference track signals. The multiple sample reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories. The first separation model and the second separation model are trained respectively, so that: the first separation model learns to separate the preset track-separated signal and the enhanced mixed signal from the sample mixed audio signal, and the second separation model learns to separate the target enhanced track-separated signal from the enhanced mixed signal corresponding to the training sample, wherein the enhanced mixed signal is composed of the superposition of each enhanced track-separated signal; The second separation model is retrained so that: the parameters of the second separation model are adjusted based on the predicted target enhanced track segmentation signal separated from the predicted enhanced mixed signal output from the first separation model, and the corresponding enhanced track segmentation signal in the training samples.

[0006] According to a second aspect of this specification, a sound source separation method is provided, the method comprising: A mixed audio signal composed of multiple reference track signals is acquired. The multiple reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories. The mixed audio signal is input into a pre-trained first separation model, so that: the first separation model separates the preset track segment signal and the predicted enhanced mixed signal from the mixed audio signal, and the predicted enhanced mixed signal is input into a second separation model, so that the second separation model separates the target enhanced track segment signal from the predicted enhanced mixed signal, wherein the predicted enhanced mixed signal is composed of the superposition of each enhanced track segment signal; The first separation model and the second separation model are trained according to the method described in the first aspect.

[0007] According to a third aspect of this specification, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the method described in either the first or second aspect by executing the executable instructions.

[0008] According to a fourth aspect of this specification, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in either the first or second aspect.

[0009] According to a fifth aspect of this specification, a computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] The method provided in this specification proposes that the second separation model undergoes two different training processes: In the first training, the second separation model uses the enhanced mixed signal corresponding to the training samples as input to learn the basic ability to separate the target enhanced track signal from the ideal signal; in the second training, the input of the second separation model is replaced with the predicted enhanced mixed signal actually output by the first separation model, making this signal contain the distortion introduced by the first separation model during the separation process. Because the second separation model directly encounters the distorted signal with the same data distribution as the actual inference stage in the second training, and uses the corresponding enhanced track signal in the training samples as the learning target, the second separation model can actively adapt to the spectral impairments and separation errors in the previous stage output, thereby effectively bridging the data distribution differences between training and inference. After this training, the second separation model can more accurately reconstruct the target enhanced track signal from the damaged enhanced mixed signal output by the first separation model during inference, significantly improving the separation quality and performance indicators of the target audio source. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0012] Figure 1 This is a schematic diagram of an audio source separation scheme based on a cascaded architecture, as shown in an embodiment of the present invention. Figure 2 This is a schematic flowchart illustrating a training method for a sound source separation model according to an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating a first-stage training of a separation model according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating a one-stage training of a second separation model according to an embodiment of the present invention; Figure 5 This is a schematic diagram illustrating a two-stage training of a second separation model according to an embodiment of the present invention; Figure 6 This is a schematic flowchart illustrating a sound source separation method according to an embodiment of the present invention; Figure 7 This is a schematic structural diagram of an electronic device according to an embodiment of the present invention; Figure 8 This is a block diagram of a training device for a sound source separation model, as shown in an embodiment of the present invention. Figure 9This is a block diagram of a sound source separation device shown in an embodiment of the present invention. Detailed Implementation

[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present invention.

[0014] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0015] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] Figure 2 This is a flowchart illustrating a training method for a sound source separation model according to an exemplary embodiment of the present invention. The method may specifically include the following steps: Step S202: Obtain training samples. The training samples include a mixed audio signal composed of multiple sample reference track signals. The multiple sample reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories.

[0017] The training samples mentioned above can include sample mixed audio signals composed of multiple sample reference track signals superimposed on each other. The so-called sample reference track signal refers to a track audio that is not mixed with other sound sources and corresponds to a single sound source category, such as a recording of a piano solo or a vocal solo. Each sample reference track signal corresponds to an independent sound source category. In the training samples, these sample reference track signals can be divided into two sets according to the sound source category: preset track signals and enhanced track signals. The preset track signals can correspond to relatively conventional sound source categories with significant acoustic characteristics, such as vocals, bass, and drums; these are collectively referred to as preset sound source categories. The enhanced track signals correspond to sound source categories such as guitar, piano, strings, wind instruments, synthesizers, and percussion instruments, where further fine separation is desired but acoustic characteristics are relatively less prominent; these are collectively referred to as enhanced sound source categories. Of course, enhanced track signals whose acoustic characteristics are weakened to the point that it is difficult to distinguish a specific sound source category can be uniformly classified as "other" sound source categories. The sample mixed audio signal is composed of the above-mentioned sample reference track signals and is used to simulate the mixed audio to be separated in a real-world scenario.

[0018] It is worth noting that the above distinction between "preset sound source category" and "enhanced sound source category" is not based on the fixed attributes of the sound source category itself, but rather on the division of labor between the first separation model and the second separation model in a specific task. In practical applications, the first separation model can be considered as... Figure 1 The front-end model in the architecture shown is responsible for directly separating the output audio source category, i.e., the corresponding preset track signal; the second separation model can be regarded as... Figure 1 In the architecture shown, any subsequent model that further refines the separation of audio sources corresponds to an enhanced track signal. For example, suppose in a certain application scenario, there exists a mature first separation model that only supports separating individual vocals and piano from mixed audio, with the remaining unseparable audio sources considered "others." A second separation model is expected to further separate the guitar from the other track signals output by the first separation model with higher precision. In this scenario, the track signals corresponding to vocals and piano are the preset track signals, and the track signal corresponding to guitar is the enhanced track signal. In other words, the same type of audio source may be classified into different sets under different model configurations; this specification does not impose specific restrictions on this.

[0019] Furthermore, the training samples mentioned above, in addition to containing a mixed audio signal composed of multiple superimposed sample reference track signals, may also directly contain the individual sample reference track signals that constitute the mixed audio signal. In other words, a complete training sample provides both the mixed signal used to simulate actual separated inputs and the individual reference track signals used to supervise model training. Of course, depending on the audio file format, if the mixed audio signal itself uses an audio encoding format that supports multi-track storage, the corresponding sample reference track signals for each track can be directly extracted from the mixed audio signal, thus eliminating the need for separate storage. In conclusion, this specification does not limit the specific organization of the training samples.

[0020] Having clarified the basic composition of the training samples, the following further explains the specific construction method of the sample mixed audio signal. In actual training, this manual can employ various strategies to combine the sample reference track signals to generate the sample mixed audio signal, thereby expanding the scale and diversity of the training data and improving the model's generalization ability.

[0021] In one embodiment, original track-segmented signals can be extracted from different original audio sources or from different time periods of the same original audio source, collectively referred to here as random original track-segmented signals; then, these random original track-segmented signals are used as sample reference track-segmented signals and superimposed to form a sample mixed audio signal.

[0022] In another embodiment, original track-segment signals can be extracted from the same time period of the same original audio, collectively referred to here as aligned original track-segment signals; then, these aligned original track-segment signals are used as sample reference track-segment signals and superimposed to form a sample mixed audio signal.

[0023] The former embodiment can significantly expand the combination space of training samples by shuffling the combination relationship of sound sources of different audio or different time periods, which helps the model learn a wider range of sound source co-occurrence patterns; the latter embodiment maintains the original consistency of melody and rhythm between each sound source, making the training signal closer to the acoustic characteristics of real mixed audio.

[0024] In actual training, the two construction methods described above can be used individually or in combination. For example, if the sample mixed audio signal in the training samples is determined to be composed of the random original track-splitting signal and the aligned original track-splitting signal, the first separation model and the second separation model can be trained respectively with the sample mixed audio signal constructed from the random original track-splitting signal in the early stage of training, i.e., step S204 below; in the later stage of training, i.e., step S206 below, the second separation model can be trained with the sample mixed audio signal constructed from the aligned original track-splitting signal, so that the second separation model first establishes basic separation capabilities in a wide and diverse range of sound source combination scenarios, and then further refines its learning in mixed signals with highly consistent melody and rhythm, ultimately improving the separation accuracy and generalization ability of the target enhanced track-splitting signal below; or, in each round of training, one of the two methods can be randomly selected with a certain probability to construct the training samples of the current batch. This specification does not limit the specific mixing strategy selection and combination method.

[0025] Step S204: Train the first separation model and the second separation model respectively, so that: the first separation model learns to separate the preset track-separated signal and the enhanced mixed signal from the sample mixed audio signal, and the second separation model learns to separate the target enhanced track-separated signal from the enhanced mixed signal corresponding to the training sample, wherein the enhanced mixed signal is composed of the superposition of each enhanced track-separated signal.

[0026] The first and second separation models mentioned above, as deep learning-based audio source separation models, can both adopt mask-based or mapping-based architectures. Taking a Convolutional Recurrent Neural Network (CRNN) as an example, the model can be designed as an encoder-splitter-decoder structure. For the input time-domain audio signal, it is first converted into a complex spectral tensor through a Short-Time Fourier Transform (STFT). This tensor is then used by the encoder to extract high-dimensional features in the time-frequency domain, followed by the splitter to independently model each audio source object, and then the decoder to reconstruct the complex spectrum of each target track signal. Finally, the complex spectrum is restored to the time-domain track-separated audio signal through an Inverse Short-Time Fourier Transform (ISTFT). Through this architecture, the model can learn the mapping relationship from mixed audio to each independent track signal end-to-end.

[0027] After acquiring training samples, the first and second separation models can be trained independently, without dependence on each other. Specifically, the first separation model takes the sample mixed audio signal as input and learns to separate each preset track signal, as well as an enhanced mixed signal composed of the superposition of all enhanced track signals. The second separation model takes the enhanced mixed signal corresponding to the training sample as input and learns to separate the target enhanced track signal. It is important to emphasize that the enhanced mixed signal used to train the second separation model at this stage is constructed by directly superimposing the enhanced track signals in the training sample, and is not the output of the first separation model, nor does it contain interference from other signals.

[0028] Through this training phase, the first separation model gains the ability to separate the preset track signals and the enhanced mixed signal from the mixed audio, while the second separation model gains the ability to extract the target enhanced track signal from the enhanced mixed signal constructed by directly superimposing the enhanced track signals.

[0029] The process of training the first separation model and the second separation model as described above is referred to as the first-stage training process. The training of the two models can be carried out using their own independent loss functions and parameter update mechanisms. The following will provide further explanation in conjunction with specific implementation methods.

[0030] For training the first separation model, the sample mixed audio signal is first input into the first separation model, which then outputs a predicted preset track segment signal and a predicted enhanced mixed signal. Next, the differences between the predicted preset track segment signal and the corresponding preset track segment signal in the training samples, as well as the differences between the predicted enhanced mixed signal and the corresponding enhanced mixed signal in the training samples, are calculated. The enhanced mixed signal in the training samples is constructed by directly superimposing the various enhanced track segment signals; its role at this stage is to provide a supervisory reference for the first separation model's output target—the enhanced mixed signal. Based on these two differences, the loss value of the first separation model is determined, and the parameters of the first separation model are updated accordingly.

[0031] by Figure 3For example, the pre-stage model in the figure can be used as the first separation model mentioned above. This model can receive and process the sample reference track signals, including those corresponding to vocals, drums, bass, guitar, piano, and other sound source categories (referred to as "Other 1" in the figure), which are combined by the audio processing device to form the sample mixed audio signal, i.e., the mixed audio in the figure. It then outputs the predicted preset track signals corresponding to vocals, bass, and drums, as well as the predicted enhanced mixed signals corresponding to other sound source categories (referred to as "Other 3" in the figure). Among them, the sample reference track signals corresponding to vocals, drums, and bass are preset track signals, and the sample reference track signals corresponding to guitar, piano, and Other 1 are enhanced track signals. These enhanced track signals can be further superimposed to form the enhanced mixed signal, which is equivalent to the other sound source categories (referred to as "Other 2" in the figure) used to directly form the above sample mixed audio signal. Based on the above architecture, the loss values ​​between each pair of preset track signals and predicted preset track signals belonging to the same sound source category, and between each pair of enhanced mixed signals and predicted enhanced mixed signals belonging to the same sound source category, can be calculated separately. These differences are then used to obtain the total loss value through methods such as weighted summation or summation averaging. The network parameters of the first separation model, such as the weights and biases of each convolutional layer and fully connected layer, are then updated based on the total loss value using the backpropagation algorithm.

[0032] For training the second separation model, the enhanced mixed signal corresponding to the training samples is input into the second separation model, which then outputs the predicted target enhanced track segmentation signal. Next, the difference between the predicted target enhanced track segmentation signal and the corresponding enhanced track segmentation signal in the training samples is calculated. Based on this difference, the loss value of the second separation model is determined, and the parameters of the second separation model are updated accordingly. It is important to emphasize again that the enhanced mixed signal received by the second separation model at this stage is constructed by directly superimposing the individual enhanced track segmentation signals from the training samples, and not from the output of the first separation model. Therefore, the second separation model learns at this stage its ability to extract the target enhanced track segmentation signal from the distortion-free enhanced mixed signal.

[0033] by Figure 4 For example, the piano model in the diagram can serve as the second separation model mentioned above. This model can receive and process the enhanced mixed signal "Other 2" in the training samples, which is composed of the superposition of preset track signals corresponding to guitars, pianos, and other sound source categories "Other 1". It then outputs the target enhanced track signal corresponding to the piano. Based on this architecture, the loss value between the target enhanced track signal and the predicted target enhanced track signal belonging to the same sound source category can be calculated separately. These differences can be used to obtain the total loss value through methods such as weighted summation or sum-average. The parameters of the second separation model are then updated based on this total loss value.

[0034] In this way, the first separation model and the second separation model each establish independent basic separation capabilities in the first stage, laying the foundation for the subsequent retraining stage.

[0035] Step S206: Retrain the second separation model so that: the second separation model adjusts the parameters of the second separation model based on the predicted target enhanced track segmentation signal separated from the predicted enhanced mixed signal output from the first separation model, and the corresponding enhanced track segmentation signal in the training samples.

[0036] After completing the separate training described above, the second separation model can be trained a second time, hereinafter referred to as the two-stage training process. This second training aims to adapt the second separation model to the signal distortion generated by the first separation model in actual inference, thereby improving the overall performance of the cascaded separation.

[0037] Specifically, the two-stage training process described above first inputs the sample mixed audio signal into a first separation model that has already completed training. The first separation model outputs a predicted enhanced mixed signal. It should be noted that the parameters of the first separation model are fixed during this stage; that is, the first separation model does not participate in subsequent parameter updates. Then, the predicted enhanced mixed signal output by the first separation model is used as input to a second separation model that has also completed preliminary training. The second separation model separates the predicted target enhanced track signal from this signal. Because the predicted enhanced mixed signal is distorted to some extent due to the separation error of the first separation model, there will inevitably be a deviation between the predicted target enhanced track signal output by the second separation model and the actual enhanced track signal. Based on the difference between the predicted target enhanced track signal and the corresponding enhanced track signal in the training samples, for example, by calculating the mean square error or mean absolute error, the loss value can be determined, and the parameters of the second separation model are adjusted only through the backpropagation algorithm. This process is repeated until the second separation model converges, thereby enabling the second separation model to gradually adapt to the signal distortion introduced by the first separation model in actual inference, improving the separation quality of the target enhanced track signal in real cascaded inference scenarios.

[0038] by Figure 5 For example, the piano model in the diagram can be used as the first separation model with fixed parameters after a one-stage training. Its output, the predicted enhanced mixed signal, and other sound source categories "Other 3" can be input into the second separation model after a one-stage training, to separate the predicted target enhanced track signal corresponding to the piano sound source category. Based on this architecture, the loss value between the target enhanced track signal and the predicted target enhanced track signal belonging to the piano sound source category can be calculated for each pair. These differences are then used to obtain the total loss value, which is used to update the parameters of the second separation model.

[0039] It should be noted that the second separation model described in this specification can be used in either of the following two scenarios, and the appropriate model can be flexibly selected according to actual needs: Scenario 1: In this case, the second separation model, acting as a single-output model, can be used to separate the target enhanced track signal corresponding to a certain type of enhanced audio source from the enhanced mixed signal. Figure 5 For example, in the embodiment, the "power piano model" in the figure belongs to this type of single-output model. It is only used to separate the predicted target enhanced track signal corresponding to the "piano" sound source category. If it is necessary to separate other enhanced sound source categories, such as guitar, another independent second separation model needs to be set up, such as the "power guitar model". This model also needs to undergo independent first-stage training and second-stage training, and receives "Other 3" output from the pre-amplifier model as input. In this scheme, each enhanced sound source category has its own dedicated second separation model. The models are independent of each other and do not interfere with each other, so the separation quality of each target track signal can be well guaranteed.

[0040] Scenario 2: In this scenario, the second separation model acts as a multi-output model, used to simultaneously separate the target enhanced track signals corresponding to multiple enhanced audio source categories from the enhanced mixed signal. (Corresponding to...) Figure 5 In this example, the "post-piano model" is designed as a multi-output model, whose inference results simultaneously include target enhanced track signals from multiple enhanced sound source categories, such as "piano" and "guitar." The advantage of this approach is that only one second separation model is needed to separate multiple enhanced sound source categories, resulting in a lower number of models and lower training costs. However, when training a sound source separation model that simultaneously outputs multiple tracks, the learning difficulty for different track signals varies. This leads to higher signal-to-distortion ratios (SDRs) for easier-to-learn track signals, while relatively lower SDRs for harder-to-learn track signals. Furthermore, this phenomenon does not significantly improve with increasing training iterations.

[0041] Therefore, from the perspective of separation quality, Scenario 1 is generally superior to Scenario 2; however, from the perspective of system cost and deployment efficiency, Scenario 2 is more advantageous. Nevertheless, those skilled in the art can freely choose between the two implementation methods based on the performance requirements and resource constraints of the actual application scenario, and this specification does not impose any restrictions on this.

[0042] After introducing the two implementation methods of the second separation model—single-output and multi-output—this specification also proposes a preferred training strategy: using a Generative Adversarial Network (GAN) to update the parameters of the second separation model, thereby improving the perceived quality and realism of its output prediction target enhancement track separation signal. Specifically, this strategy uses the second separation model as a generator and additionally introduces a discriminator model, with the two together forming a generative adversarial network.

[0043] In one embodiment, adjusting the parameters of the second separation model further includes: using the second separation model as a generator to form a generative adversarial network with the discriminator model. A total loss is determined based on regression loss, generative adversarial loss, and feature matching loss, and the parameters of the second separation model are updated according to the total loss. Wherein: The regression loss is determined based on the difference between the predicted target enhanced track splitting signal and the corresponding enhanced track splitting signal in the training samples. For example, it can be measured by mean squared error (MSE) or mean absolute error (MAE). The adversarial loss is determined based on the discriminator model's discrimination result of the predicted target enhanced track splitting signal. This loss is used to drive the generator to output a more realistic target enhanced track splitting signal that is difficult for the discriminator to distinguish. The feature matching loss is determined based on the difference between the features extracted by the discriminator model from the predicted target enhanced track splitting signal and the corresponding enhanced track splitting signal in the training samples. This loss helps to stabilize the training process and improve the time-frequency structure fidelity of the generated signal.

[0044] In summary, by weighted joint optimization of the above three losses, this specification can significantly improve the auditory quality and perceived naturalness of the separation results while ensuring the consistency between the separated signal and the real signal values, overcoming the fuzziness or smoothing problems that are easily generated when relying solely on regression loss.

[0045] The following uses the second separation model (hereinafter referred to as the "piano model") for separating piano sound sources as an example. The training process of the discriminator model described in this application is carried out alternately with the updating of the piano model. The specific steps are as follows: 1. Take the predicted enhanced mixed signal from a batch of training data as input, denoted as x, and input it into the piano model as the generator. The model outputs the predicted target enhanced track segmentation signal, denoted as y'. At the same time, obtain the real enhanced track segmentation signal corresponding to the input from the training samples, denoted as y.

[0046] 2. Calculate the regression loss =MSE(y',y).

[0047] 3. Input y' and y into the discriminator model respectively to obtain the discriminator's discrimination results for the two outputs. and .

[0048] 4. The loss function of Least Squares GAN (LSGAN) is adopted, based on... and Calculate the discriminator loss Then, backpropagation is performed to update the parameters of the discriminator model.

[0049] 5. Input y' and y into the discriminator model again to obtain the discriminator's discrimination results for the two and the discriminator's internal feature mapping: , and feature maps , .

[0050] 6. Based on and Difference calculation of feature matching loss ;based on Adversarial loss against all-one vector computation generator (It also uses the LSGAN form).

[0051] 7. Summate the regression loss, generative adversarial loss, and feature matching loss according to their pre-defined weights. , , The total loss of the generator is obtained. It then performs backpropagation to update the parameters of the piano model used as the generator.

[0052] 8. Repeat the above steps until the piano model converges.

[0053] It should be noted that the above example uses the piano as a single enhanced sound source category. For the second separation model of other enhanced sound source categories, the training process is exactly the same; simply replace the corresponding real enhanced track signals with signals from the target sound source category. Through this embodiment, the second separation model, within the framework of a generative adversarial network, can learn more realistic and natural target enhanced track signals, significantly improving the overall performance of the cascaded separation system in complex audio scenarios.

[0054] Figure 6 This is a flowchart illustrating an exemplary embodiment of the present invention for a sound source separation method, which may specifically include the following steps: Step 602: Obtain a mixed audio signal composed of multiple reference track signals, wherein the multiple reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories. Step 604: Input the mixed audio signal into a pre-trained first separation model so that: the first separation model separates the preset track segment signal and the predicted enhanced mixed signal from the mixed audio signal, and inputs the predicted enhanced mixed signal into a second separation model so that the second separation model separates the target enhanced track segment signal from the predicted enhanced mixed signal, wherein the predicted enhanced mixed signal is composed of the superposition of each enhanced track segment signal; The first separation model and the second separation model are trained according to the training method of the above-mentioned sound source separation model.

[0055] As previously stated, the second separation model is used to: separate a target enhanced track signal corresponding to one type of enhanced audio source from the enhanced mixed signal; or, simultaneously separate target enhanced track signals corresponding to multiple types of enhanced audio sources from the enhanced mixed signal.

[0056] As mentioned above, the parameters of the first separation model are fixed after training; during the training process, the parameters of the second separation model are adjusted based on the difference between the predicted target enhanced track segmentation signal separated from the predicted enhanced mixed signal output by the first separation model and the corresponding enhanced track segmentation signal in the training samples.

[0057] Figure 7 This is a schematic structural diagram of an electronic device according to an exemplary embodiment. Please refer to... Figure 7At the hardware level, the electronic device includes a processor 702, an internal bus 710, a network interface 704, memory 706, and non-volatile memory 708, and may also include other necessary hardware. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it, forming a training device or sound source separation device for the sound source separation model at the logical level. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.

[0058] Figure 8 This invention illustrates a block diagram of a training device for a sound source separation model. Please refer to... Figure 8 The device includes: The training sample acquisition unit 802 is used to acquire training samples, which include a sample mixed audio signal composed of multiple sample reference track signals. The multiple sample reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories. The first-stage acquisition unit 804 trains a first separation model and a second separation model respectively, so that: the first separation model learns to separate the preset track-separated signal and the enhanced mixed signal from the sample mixed audio signal, and the second separation model learns to separate the target enhanced track-separated signal from the enhanced mixed signal corresponding to the training sample, wherein the enhanced mixed signal is composed of the superposition of each enhanced track-separated signal; The second-stage acquisition unit 806 retrains the second separation model so that: the second separation model adjusts the parameters of the second separation model based on the predicted target enhanced track separation signal separated from the predicted enhanced mixed signal output from the first separation model, and the corresponding enhanced track separation signal in the training samples.

[0059] Optionally, the training sample acquisition unit 802 is specifically used for: Random original track-segmented signals are extracted from different original audio sources or different time periods of the same original audio source, and / or aligned original track-segmented signals are extracted from the same time period of the same original audio source. The multiple sample reference track signals are determined based on the extracted signals, and the sample mixed audio signal is constructed based on the multiple sample reference track signals as the training samples.

[0060] Optionally, the sample mixed audio signal in the training samples is composed of the random original track splitting signal and the aligned original track splitting signal; The training of the first separation model and the second separation model includes: training the first separation model and the second separation model respectively based on the sample mixed audio signal composed of the random original track-splitting signals; The second separation model is retrained, including training the second separation model based on the sample mixed audio signal composed of the aligned original track signals.

[0061] Optionally, the process of training the first separation model and the second separation model respectively includes: The sample mixed audio signal is input into the first separation model to obtain a predicted preset track segment signal and a predicted enhanced mixed signal; based on the difference between the predicted preset track segment signal and the corresponding preset track segment signal in the training sample, and the difference between the predicted enhanced mixed signal and the corresponding enhanced mixed signal in the training sample, the parameters of the first separation model are updated; The enhanced mixed signal corresponding to the training sample is input into the second separation model to obtain the predicted target enhanced track separation signal; based on the difference between the predicted target enhanced track separation signal and the corresponding enhanced track separation signal in the training sample, the parameters of the second separation model are updated.

[0062] Optionally, the process of retraining the second separation model includes: The sample mixed audio signal is input into the first separation model after training, and the predicted enhanced mixed signal output by the first separation model is input into the second separation model after training to obtain the predicted target enhanced track-separated signal; The parameters of the first separation model are fixed after training, and the parameters of the second separation model are adjusted based on the difference between the predicted target enhanced track separation signal and the corresponding enhanced track separation signal in the training samples.

[0063] Optionally, the second separation model is used for: Separate a target enhanced track signal corresponding to a certain category of enhanced audio source from the enhanced mixed signal; or... Target enhanced track signals corresponding to multiple enhanced audio source categories are simultaneously separated from the enhanced mixed signal.

[0064] Optionally, the two-stage acquisition unit 806 is specifically used for: The second separation model is used as a generator, and together with the discriminator model, they form a generative adversarial network. The total loss is determined by combining regression loss, generative adversarial loss, and feature matching loss, and the parameters of the second separation model are updated based on the total loss. The regression loss is determined based on the difference between the predicted target enhanced track splitting signal and the corresponding enhanced track splitting signal in the training samples; the adversarial loss is determined based on the discrimination result of the discriminator model on the predicted target enhanced track splitting signal; and the feature matching loss is determined based on the difference between the features extracted by the discriminator model from the predicted target enhanced track splitting signal and the corresponding enhanced track splitting signal in the training samples.

[0065] Figure 9 This invention illustrates a block diagram of a sound source separation device according to an embodiment of the present invention. Please refer to... Figure 9 The device includes: The audio acquisition unit 902 is used to acquire a mixed audio signal composed of multiple reference track signals, wherein the multiple reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories. The audio source separation unit 904 is used to input the mixed audio signal into a pre-trained first separation model, so that: the first separation model separates the preset track signal and the predicted enhanced mixed signal from the mixed audio signal, and inputs the predicted enhanced mixed signal into a second separation model, so that the second separation model separates the target enhanced track signal from the predicted enhanced mixed signal, wherein the predicted enhanced mixed signal is composed of the superposition of each enhanced track signal; The first separation model and the second separation model are trained according to the training method of the sound source separation model described above.

[0066] Optionally, the second separation model is used to: separate a target enhanced track signal corresponding to one type of enhanced audio source from the enhanced mixed signal; or, simultaneously separate target enhanced track signals corresponding to multiple types of enhanced audio sources from the enhanced mixed signal.

[0067] Optionally, the parameters of the first separation model are fixed after training; during the training process, the parameters of the second separation model are adjusted based on the difference between the predicted target enhanced track segmentation signal separated from the predicted enhanced mixed signal output by the first separation model and the corresponding enhanced track segmentation signal in the training samples.

[0068] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0069] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.

[0070] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0071] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0072] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.

[0073] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.

[0074] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a GPS receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0075] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0076] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0077] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0078] Therefore, specific embodiments of the subject matter have been described. Furthermore, the processes depicted in the figures are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0079] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

Claims

1. A training method for a sound source separation model, characterized in that, The method includes: Acquire training samples, which include a mixed audio signal composed of multiple sample reference track signals. The multiple sample reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories. The first separation model and the second separation model are trained respectively, so that: the first separation model learns to separate the preset track-separated signal and the enhanced mixed signal from the sample mixed audio signal, and the second separation model learns to separate the target enhanced track-separated signal from the enhanced mixed signal corresponding to the training sample, wherein the enhanced mixed signal is composed of the superposition of each enhanced track-separated signal; The second separation model is retrained so that: the parameters of the second separation model are adjusted based on the predicted target enhanced track segmentation signal separated from the predicted enhanced mixed signal output from the first separation model, and the corresponding enhanced track segmentation signal in the training samples.

2. The method according to claim 1, characterized in that, The acquisition of training samples includes: Random original track-segmented signals are extracted from different original audio sources or different time periods of the same original audio source, and / or aligned original track-segmented signals are extracted from the same time period of the same original audio source. The multiple sample reference track signals are determined based on the extracted signals, and the sample mixed audio signal is constructed based on the multiple sample reference track signals as the training samples.

3. The method according to claim 2, characterized in that, The sample mixed audio signal in the training samples is composed of the random original track splitting signal and the aligned original track splitting signal; The training of the first separation model and the second separation model includes: training the first separation model and the second separation model respectively based on the sample mixed audio signal composed of the random original track-splitting signals; The second separation model is retrained, including training the second separation model based on the sample mixed audio signal composed of the aligned original track signals.

4. The method according to claim 1, characterized in that, The process of training the first separation model and the second separation model respectively includes: The sample mixed audio signal is input into the first separation model to obtain a predicted preset track segment signal and a predicted enhanced mixed signal; based on the difference between the predicted preset track segment signal and the corresponding preset track segment signal in the training sample, and the difference between the predicted enhanced mixed signal and the corresponding enhanced mixed signal in the training sample, the parameters of the first separation model are updated; The enhanced mixed signal corresponding to the training sample is input into the second separation model to obtain the predicted target enhanced track separation signal; based on the difference between the predicted target enhanced track separation signal and the corresponding enhanced track separation signal in the training sample, the parameters of the second separation model are updated.

5. The method according to claim 1, characterized in that, The process of retraining the second separation model includes: The sample mixed audio signal is input into the first separation model after training, and the predicted enhanced mixed signal output by the first separation model is input into the second separation model after training to obtain the predicted target enhanced track-separated signal; The parameters of the first separation model are fixed after training, and the parameters of the second separation model are adjusted based on the difference between the predicted target enhanced track separation signal and the corresponding enhanced track separation signal in the training samples.

6. The method according to claim 1, characterized in that, The second separation model is used for: Separate a target enhanced track signal corresponding to a certain category of enhanced audio source from the enhanced mixed signal; or... Target enhanced track signals corresponding to multiple enhanced audio source categories are simultaneously separated from the enhanced mixed signal.

7. The method according to claim 1, characterized in that, The adjustment of the parameters of the second separation model includes: The second separation model is used as a generator, and together with the discriminator model, they form a generative adversarial network. The total loss is determined by combining regression loss, generative adversarial loss, and feature matching loss, and the parameters of the second separation model are updated based on the total loss. The regression loss is determined based on the difference between the predicted target enhanced track splitting signal and the corresponding enhanced track splitting signal in the training samples; the adversarial loss is determined based on the discrimination result of the discriminator model on the predicted target enhanced track splitting signal; and the feature matching loss is determined based on the difference between the features extracted by the discriminator model from the predicted target enhanced track splitting signal and the corresponding enhanced track splitting signal in the training samples.

8. A method for separating sound sources, characterized in that, The method includes: A mixed audio signal composed of multiple reference track signals is acquired. The multiple reference track signals are divided into preset track signals corresponding to preset sound source categories and enhanced track signals corresponding to enhanced sound source categories. The mixed audio signal is input into a pre-trained first separation model, so that: the first separation model separates the preset track segment signal and the predicted enhanced mixed signal from the mixed audio signal, and the predicted enhanced mixed signal is input into a second separation model, so that the second separation model separates the target enhanced track segment signal from the predicted enhanced mixed signal, wherein the predicted enhanced mixed signal is composed of the superposition of each enhanced track segment signal; The first separation model and the second separation model are trained according to the method described in claim 1.

9. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1 to 8 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 8.

11. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 8.