Music source separation method and device, network training method and device, equipment and storage medium

By using the knowledge distillation method in the music source separation technology, the music source separation network with larger parameters is used to guide smaller network training, solving the challenges in the computing power and delay in the existing technology, and real-time music source separation effect is achieved.

CN120048280APending Publication Date: 2025-05-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311602975.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing music source separation technology has great challenges in computing power and delay, and cannot meet the needs of real-time music source separation.

Method used

By using the first music source separation network with a larger parameter amount to knowledge distillation on the second music source separation network with a smaller parameter amount, a target music source separation network with better performance is obtained, thereby improving the effect of music source separation and reducing delay.

Benefits of technology

It realizes that in scenarios where real-time music source separation is required, the effect of music source separation is improved and the delay is reduced, which is suitable for the need to separate audio tracks in real-time on the end.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048280A_ABST
    Figure CN120048280A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a music source separation method and device, a network training method and device, equipment and a storage medium. The method comprises the steps that a to-be-processed first mixed music source is acquired; performing music source separation on the first mixed music source by adopting a target music source separation network to obtain prediction results of at least two sound components separated from the first mixed music source; wherein the target music source separation network is obtained by taking the first music source separation network as a teacher model, taking the second music source separation network as a student model, and performing knowledge distillation on the second music source separation network; a parameter amount of the first music source separation network is greater than a parameter amount of the second music source separation network. According to the method, knowledge distillation is carried out on the second music source separation network with smaller parameter quantity by using the first music source separation network with larger parameter quantity through a knowledge distillation technology, so that the target music source separation network with better performance and lower time delay can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular, to a music source separation method, a network training method, a device, a device, and a storage medium. Background Art

[0002] Music Source Separation (MSS) (or called music separation / sound separation) is the task of decomposing music (mixed music source) into its components. For example, separating music into tracks such as vocals, bass, drums, and others.

[0003] Currently, the mainstream MSS tasks use Artificial Intelligence (AI) technology to extract different tracks from conventional stereo music. Although good results can be achieved, the model is large, has high requirements for computing power, and the latency of the algorithm is also obvious. From the perspectives of space occupancy, computing power consumption, and algorithm latency, it cannot be applied to scenarios that require real-time music source separation. Summary of the Invention

[0004] The embodiments of the present application at least provide a music source separation method, a network training method, a device, a device, and a storage medium.

[0005] The technical solution of the present application is implemented as follows:

[0006] In a first aspect, the embodiments of the present application provide a music source separation method, which includes: obtaining a first mixed music source to be processed; performing music source separation on the first mixed music source by using a target music source separation network to obtain a prediction result of at least two sound components separated from the first mixed music source; wherein the target music source separation network is obtained by performing knowledge distillation on a second music source separation network with a first music source separation network as a teacher model; the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network.

[0007] Second aspect, an embodiment of the present application provides a training method for a music source separation network, the method comprising: separating a mixed music source in a training set by using a first music source separation network to obtain a first prediction result of at least two sound components separated from the mixed music source; separating the mixed music source by using a second music source separation network to obtain a second prediction result of at least two sound components separated from the mixed music source; wherein the number of parameters of the first music source separation network is greater than that of the second music source separation network; using the first music source separation network as a teacher model and the second music source separation network as a student model, and performing knowledge distillation on the second music source separation network based on the first prediction result and the second prediction result to obtain a target music source separation network.

[0008] Third aspect, an embodiment of the present application provides a music source separation device, the device comprising: an acquisition unit, configured to acquire a first mixed music source to be processed; a music source separation unit, configured to separate the first mixed music source by using a target music source separation network to obtain a prediction result of at least two sound components separated from the first mixed music source; wherein the target music source separation network is obtained by performing knowledge distillation on the second music source separation network by using the first music source separation network as a teacher model and the second music source separation network as a student model; and the number of parameters of the first music source separation network is greater than that of the second music source separation network.

[0009] Fourth aspect, an embodiment of the present application provides a training device for a music source separation network, the device comprising: a first music source separation unit, configured to separate a mixed music source in a training set by using a first music source separation network to obtain a first prediction result of at least two sound components separated from the mixed music source; a second music source separation unit, configured to separate the mixed music source by using a second music source separation network to obtain a second prediction result of at least two sound components separated from the mixed music source; wherein the number of parameters of the first music source separation network is greater than that of the second music source separation network; a knowledge distillation unit, configured to use the first music source separation network as a teacher model and the second music source separation network as a student model, and perform knowledge distillation on the second music source separation network based on the first prediction result and the second prediction result to obtain a target music source separation network.

[0010] Fifth aspect, an embodiment of the present application provides a music source separation device, the music source separation device comprising a memory and a processor; wherein, the memory is configured to store computer-executable instructions; the processor is connected to the memory and is configured to implement the method as described in the first aspect or the second aspect by executing the computer-executable instructions.

[0011] Sixth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by at least one processor, implements the method described in the first aspect or the second aspect.

[0012] Seventh aspect, an embodiment of the present application provides a chip. The chip includes: a processor configured to call and run a computer program from a memory, so that a device installed with the chip executes the method described in the first aspect or the second aspect.

[0013] In an embodiment of the present application, a target music source separation network may be used to perform music source separation on a first mixed music source to obtain a prediction result of at least two sound components separated from the first mixed music source; wherein, the target music source separation network is obtained by performing knowledge distillation on a second music source separation network with a first music source separation network as a teacher model and the second music source separation network as a student model, and the number of parameters of the first music source separation network is greater than that of the second music source separation network. This method can obtain a target music source separation network with better performance by using the first music source separation network with a larger number of parameters to perform knowledge distillation on the second music source separation network with a smaller number of parameters, thereby improving the effect of music source separation. At the same time, since the number of parameters of the second music source separation network is small, that is, the number of parameters of the obtained target music source separation network is small, therefore, this method can reduce the latency of music source separation, and thus can be applied to scenarios that require real-time music source separation.

[0014] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and do not limit the technical solution of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solution of the present application.

[0016] Figure 1 Schematic diagram of the basic principle of AI-based MSS;

[0017] Figure 2 Schematic diagram of the U-NET model architecture;

[0018] Figure 3 Schematic flow chart of a method for music source separation provided by an embodiment of the present application;

[0019] Figure 4 Structural schematic of the target music source separation network provided by an embodiment of the present application Figure 1 .

[0020] Figure 5Schematic diagram of the time-frequency feature extraction unit provided by the embodiment of the present application;

[0021] Figure 6 Schematic diagram of the structure of the unidirectional gated recurrent unit provided by the embodiment of the present application;

[0022] Figure 7 Schematic diagram of the structure of the unidirectional GRU layer provided by the embodiment of the present application;

[0023] Figure 8 Schematic diagram of the structure of the feature extraction module provided by the embodiment of the present application;

[0024] Figure 9 Schematic flowchart of a training method for a music source separation network provided by the embodiment of the present application;

[0025] Figure 10 Schematic flowchart of a possible implementation of the music source separation method provided by the embodiment of the present application;

[0026] Figure 11 Schematic flowchart of separating the mixed music source into music sources provided by the embodiment of the present application;

[0027] Figure 12 Schematic diagram of the structure of the target music source separation network provided by the embodiment of the present application Figure 2 ;

[0028] Figure 13 Schematic flowchart of knowledge distillation provided by the embodiment of the present application;

[0029] Figure 14 Schematic diagram of the composition structure of a music source separation device provided by the embodiment of the present application;

[0030] Figure 15 Schematic diagram of the composition structure of a training device for a music source separation network provided by the embodiment of the present application;

[0031] Figure 16 Schematic diagram of a hardware entity of the music source separation device in the embodiment of the present application. Detailed implementation manners

[0032] In order to understand the features and technical content of the embodiments of the present application in more detail, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and illustration purposes and are not used to limit the embodiments of the present application.

[0033] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0034] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. It should also be noted that the terms "first / second / third" in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific order for the objects. It is understood that "first / second / third" may be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0035] It should be understood that the term "and / or" in the embodiments of the present application is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0036] Music source separation (MSS) is the task of decomposing music (mixed music source) into its components. For example, separating music into tracks such as vocals, bass, drums, and other. Currently, mainstream MSS tasks use AI technology to extract different tracks from conventional stereo music. The separated tracks can be processed individually using track enhancement or spatial rendering algorithms and remixed again to build various applications. In the recreated music, the extracted tracks can be automatically placed at different positions in a virtual space to provide an immersive listening experience.

[0037] Figure 1 A schematic diagram of the basic principle of AI-based MSS is shown. As Figure 1 shown, the mixed music source includes four sound components: bass, drums, vocals, and other. Through AI technology, the mixed music source can be separated into the tracks of each of the four sound components. Further, the separated tracks can be processed individually using track enhancement or spatial rendering algorithms, and the extracted tracks can be placed at different positions in a virtual space.

[0038] With the continuous development of mobile terminal audio technology, spatial audio has broken through the two-channel limitation of stereo audio, opening up new dimensions of music and greatly enhancing the user's listening experience. In the music scenario, sound separation technology can adjust the sound field orientation and volume, or extract specific sound objects, providing users with a custom function that enables them to create their own sound effect styles, enhancing user participation and playability, and expanding the boundaries of audio experience. In the long run, separation technology will support a wider range of sound object separations, and its importance is analogous to image segmentation in the field of images, providing a prerequisite for subsequent immersive spatial rendering and various sound effects of sound object post-processing.

[0039] With the development of deep learning technology, the MSS technology based on Deep Neural Network (DNN) has achieved good results and has been widely applied to various music scenarios. Since the U-NET architecture was proposed in the field of image segmentation, it has been widely used in the MSS field and has made continuous progress. Figure 2 The schematic diagram of the U-NET model architecture is shown. As Figure 2 shown, this architecture mainly consists of an encoder layer, an intermediate layer, and a decoder layer.

[0040] In recent years, some studies have introduced the Recurrent Neural Network (RNN) and the transformer model into the MSS technology, achieving good results. However, although these MSS systems can achieve good results, the models are relatively large, have high requirements for computing power, and the latency of the algorithms is also obvious. They cannot be deployed on the edge side (such as the mobile phone side) and cannot be applied to scenarios that require real-time music separation. Therefore, in order to perform real-time separation of audio tracks on the edge side, it is necessary to design an MSS model with small parameter quantities, low latency, and high quality.

[0041] In view of this, the embodiments of the present application provide a music source separation method, a network training method, a device, a device, and a storage medium. This method can be executed by the processor of a computer device. Among them, the computer device can refer to devices with data processing capabilities such as servers, laptop computers, tablet computers, and desktop computers. In this method, a target music source separation network can be used to perform music source separation on the first mixed music source to obtain a prediction result of at least two sound components separated from the first mixed music source; among them, the target music source separation network is obtained by performing knowledge distillation on the second music source separation network with the first music source separation network as the teacher model and the second music source separation network as the student model, and the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network.

[0042] Since the number of parameters of the model is positively correlated with the representation ability of the model, generally, the larger the model, the greater the representation ability. Therefore, by using a first music source separation network with a larger number of parameters to perform knowledge distillation on a second music source separation network with a smaller number of parameters, a target music source separation network with better performance can be obtained, thereby improving the effect of music source separation. At the same time, since the number of parameters of the second music source separation network is small, that is, the number of parameters of the obtained target music source separation network is small. Therefore, this method can reduce the latency of music source separation, and thus can be applied to scenarios that require real-time music source separation.

[0043] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0044] Figure 3 A method for music source separation provided by an embodiment of the present application is shown. As Figure 3 shown, the method may include:

[0045] S301, obtaining a first mixed music source to be processed.

[0046] Exemplarily, the first mixed music source may be composed of at least one frame of sound data, and the first mixed music source may include multiple sound components. For example, the first mixed music source may include, but is not limited to, at least two of the following sound components: human voice, drumbeat, bass.

[0047] In some embodiments, when the first mixed music source is composed of multiple frames of sound data, there may be partial overlap between two adjacent frames of sound data in the multiple frames of sound data. For example, the length of each frame of sound data is 2048 sample points, and there are 1536 overlapping sample points between two adjacent frames of sound data.

[0048] S302, using the target music source separation network to perform music source separation on the first mixed music source, and obtaining prediction results of at least two sound components separated from the first mixed music source.

[0049] Wherein, the target music source separation network is obtained by performing knowledge distillation (KD) on the second music source separation network with the first music source separation network as the teacher model and the second music source separation network as the student model; the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network.

[0050] Since the number of parameters of the model is positively correlated with the representation ability of the model, generally, the larger the model, the stronger the representation ability. Therefore, by using a first music source separation network with a larger number of parameters to perform knowledge distillation on a second music source separation network with a smaller number of parameters, a target music source separation network with better performance can be obtained, thereby improving the effect of music source separation. At the same time, since the number of parameters of the second music source separation network is small, that is to say, the number of parameters of the obtained target music source separation network is small. Therefore, this method can reduce the latency of music source separation, and thus can be applied to scenarios that require real-time music source separation.

[0051] Taking the first music source separation network as the teacher model and the second music source separation network as the student model, the steps of performing knowledge distillation on the second music source separation network will be introduced in S903 and will not be elaborated here for the time being.

[0052] It should be understood that the prediction results of at least two sound components separated from the first mixed music source obtained through S302 can be the prediction results of any two or more sound components separated from the multiple sound components included in the first mixed music source. For example, the prediction results of vocals and bass separated from the first mixed music source can be obtained; for another example, the prediction results of vocals, drums, and bass separated from the first mixed music source can be obtained.

[0053] In some embodiments, the target music source separation network may include: an encoding sub-network and a decoding sub-network; the encoding sub-network includes a plurality of feature extraction modules and a plurality of downsampling modules; the decoding sub-network includes a plurality of feature extraction modules and a plurality of upsampling modules; the number of downsampling modules is the same as the number of upsampling modules.

[0054] Figure 4 shows a schematic structure of the target music source separation network provided by an embodiment of the present application Figure 1 . As Figure 4 shown, the target music source separation network 400 includes: an encoding sub-network 401 and a decoding sub-network 402. The encoding sub-network 401 includes 4 feature extraction modules (denoted as feature extraction module #1 to feature extraction module #4 respectively) and 3 downsampling modules (denoted as downsampling module #1 to downsampling module #3 respectively); the decoding sub-network 402 includes 3 feature extraction modules (denoted as feature extraction module #5 to feature extraction module #7 respectively) and 3 upsampling modules (denoted as upsampling module #1 to upsampling module #3 respectively).

[0055] In some embodiments, a target music source separation network is used to separate the first mixed music source, and a prediction result of at least two sound components separated from the first mixed music source is obtained, including: alternately using the feature extraction module and the downsampling module in the encoding sub-network to perform encoding processing on the first mixed music source to obtain the encoded first mixed music source; alternately using the upsampling module and the feature extraction module in the decoding sub-network to perform decoding processing on the encoded first mixed music source to obtain the prediction result of at least two sound components separated from the first mixed music source.

[0056] Taking Figure 4 as an example, the steps of using the target music source separation network 400 to separate the first mixed music source may include steps 3021) and 3022):

[0057] 3021) Alternately use the feature extraction module and the downsampling module in the encoding sub-network 401 to perform encoding processing on the first mixed music source to obtain the encoded first mixed music source.

[0058] For example, the feature extraction module #1, downsampling module #1, feature extraction module #2, downsampling module #2, feature extraction module #3, downsampling module #3, and feature extraction module #4 in the encoding sub-network 401 may be sequentially used to perform encoding processing on the first mixed music source. During the encoding process, each feature extraction module in the encoding sub-network 401 can be used to extract features from the data input to the feature extraction module; each downsampling module can be used to downsample the feature extraction result of the previous feature extraction module and use the downsampling result as the input data of the next feature extraction module. For example, taking the downsampling module #1 as an example, the downsampling module #1 can be used to downsample the feature extraction result of the feature extraction module #1 and use the downsampling result as the input data of the feature extraction module #2.

[0059] Among them, the data input to the first feature extraction module (i.e., feature extraction module #1) in the encoding sub-network 401 can be the first mixed music source, or it can be the convolution result obtained by performing convolution processing on the first mixed music source. The feature extraction result output by the last feature extraction module (i.e., feature extraction module #4) in the encoding sub-network 401 is the encoded first mixed music source.

[0060] 3022) Alternately use the upsampling module and the feature extraction module in the decoding sub-network to perform decoding processing on the encoded first mixed music source to obtain the prediction result of at least two sound components separated from the first mixed music source.

[0061] For example, the upsampling module #1, feature extraction module #5, upsampling module #2, feature extraction module #6, upsampling module #3, and feature extraction module #7 in the decoding sub-network 402 can be sequentially used to perform decoding processing on the encoded first mixed music source.

[0062] In some embodiments, in order to reduce the loss of spatial information caused by downsampling during the encoding process, and at the same time to make the feature maps restored by upsampling contain more low-level semantic information, a skip connection can be added between the encoding sub-network 401 and the decoding sub-network 402 (as Figure 4 shown by the dashed line in).

[0063] During the decoding process, each feature extraction module in the decoding sub-network 402 can be used to extract features from the data input to the feature extraction module; each upsampling module in the decoding sub-network 402 can be used to upsample the feature extraction result of the previous feature extraction module and use the upsampling result as the input data of the next feature extraction module. Taking the upsampling module #2 as an example, the upsampling module #2 can be used to upsample the feature extraction result of the feature extraction module #5 and use the upsampling result as the input data of the feature extraction module #6. It should be noted that the previous feature extraction module of the upsampling module #1 refers to the feature extraction module #4.

[0064] In the target music source separation network 400, each feature extraction module in the decoding sub-network 402 can correspond to a feature extraction module in the encoding sub-network 401. For example, the feature extraction module #7 corresponds to the feature extraction module #1, the feature extraction module #6 corresponds to the feature extraction module #2, and the feature extraction module #5 corresponds to the feature extraction module #3. For each feature extraction module in the decoding sub-network 402, the data input to the feature extraction module is the fusion result of two parts, one part is the output result of the previous upsampling module of the feature extraction module, and the other part is the output result of the feature extraction module in the decoding sub-network 401 corresponding to the feature extraction module. For example, taking the feature extraction module #7 as an example, the data input to the feature extraction module #7 is: the fusion result of the output result of the upsampling module #3 and the output result of the feature extraction module #1.

[0065] In Figure 4 the example of, the prediction result obtained by performing decoding processing on the encoded first mixed music source can be the output result of the last feature extraction module (i.e., the feature extraction module #7) in the decoding sub-network 402, or, it can be the convolution result obtained by performing convolution processing on the output result of the last feature extraction module.

[0066] It should be understood, Figure 4The numbers of the feature extraction module, the downsampling module, and the upsampling module are merely exemplary and should not impose any limitation on the implementation process of the embodiments of the present application.

[0067] In some embodiments, the feature extraction module (such as Figure 4 any one of the feature extraction modules in) includes: at least one time-frequency feature extraction unit and at least one one-way gated recurrent unit. Among them, the time-frequency feature extraction unit is used to perform causal convolution on the first data input to the time-frequency feature extraction unit in both the time domain and the frequency domain to obtain the time-frequency features of the first data; the one-way gated recurrent unit is used to extract the timing information of the second data input to the one-way gated recurrent unit to obtain the timing features of the second data.

[0068] As an example, the time-frequency feature extraction unit may, for example, include a concatenated batch normalization (BatchNorm) layer, a Gaussian Error Linear Unit (GeLU) layer, and a causal 2D convolution (Casual 2d Conv) layer. As Figure 5 shown, the time-frequency feature extraction unit 500 may include a concatenated batch normalization layer 501, a GeLU layer 502, and a causal 2D convolution layer 503. Among them, the batch normalization layer 501 can be used to perform batch normalization processing on the input first data, the GeLU layer 502 can be used as an activation function, and the causal 2D convolution layer 503 can be used to perform causal convolution in both the time domain and the frequency domain. Since causal convolution does not require the use of information from future times, it is not necessary to wait for the arrival of future information, thereby achieving the effect of reducing latency.

[0069] As an example, the one-way gated recurrent unit may, for example, include a concatenated batch normalization layer, a GeLU layer, and a one-way gated recurrent unit (Gate Recurrent Unit, GRU) layer. As Figure 6 shown, the one-way gated recurrent unit 600 may include a concatenated batch normalization layer 601, a GeLU layer 602, and a one-way GRU layer 603. Among them, the batch normalization layer 601 can be used to perform batch normalization processing on the input second data, the GeLU layer 602 can be used as an activation function, and the one-way GRU layer 603 can be used to extract the timing information. In addition, the one-way GRU layer 603 can ensure that the model does not utilize information from future times during the processing, thereby reducing the latency while utilizing the timing information of the music signal.

[0070] The network structure of the one-way GRU layer 603 is as Figure 7 shown. Among them, X t represents the input sound data of the t-th frame (such as the sound data of the t-th frame in the first mixed music source), and H t represents the one-way GRU layer's output for X tProcessing result of H t-1 Indicates the hidden state. H t It can be calculated by the following formula (1):

[0071]

[0072] Where, Z t Is the result of the update gate, Z t It can be calculated by the following formula (2):

[0073] Z t = σ(W z · [H t-1 , X t ) (2);

[0074] Indicates the candidate hidden state,[[]] It can be calculated by the following formula (3):

[0075]

[0076] R t Is the result of the reset gate, R t It can be calculated by the following formula (4):

[0077] R t = σ(W r · [H t-1 , X t ) (4);

[0078] Where, the above W z , W, and W r Are weight coefficients.[[]]

[0079] According to the above technical solution, the time-frequency feature extraction unit in the feature extraction module can perform convolution on time-frequency features (including time-domain features and frequency-domain features), or in other words, can perform convolution in two dimensions of the time domain and the frequency domain; the unidirectional gated recurrent unit can be used to extract temporal information, or in other words, can model the time domain dimension (temporal information). In this way, the time-domain features and frequency-domain features of the music signal are fully utilized, thereby improving the performance of the target music source separation network.[[]]

[0080] In some embodiments, the feature extraction module (such as Figure 4 Any feature extraction module in) includes: a plurality of serially connected feature extraction sub-modules, and each feature extraction sub-module includes two time-frequency feature extraction units and one unidirectional gated recurrent unit.[[]]

[0081] Figure 8 Shows the structural schematic diagram of the feature extraction module provided by the embodiment of the present application. As Figure 8As shown, the feature extraction module 800 includes: two cascaded feature extraction sub-modules (denoted as feature extraction sub-module 801 and feature extraction sub-module 802 respectively), and each feature extraction sub-module includes two time-frequency feature extraction units and a one-way gated recurrent unit.

[0082] Exemplarily, as Figure 8 shown, the processing flow of each feature extraction sub-module includes: using the first time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the data input to the feature extraction sub-module in both the time domain and the frequency domain to obtain the first time-frequency feature; using the one-way gated recurrent unit in the feature extraction sub-module to extract the timing information from the first time-frequency feature to obtain the first timing feature; using the second time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the result of adding the first time-frequency feature and the first timing feature in both the time domain and the frequency domain to obtain the second time-frequency feature; performing convolution processing (such as Shortcut) on the data input to the feature extraction sub-module, and adding the result of the convolution processing and the second time-frequency feature as the output result of the feature extraction sub-module.

[0083] It should be understood that Figure 8 the number of the feature extraction sub-modules is only exemplary and should not impose any limitation on the implementation process of the embodiments of the present application.

[0084] In some embodiments, the method may further include: saving the output results of multiple feature extraction modules in the target music source separation network, and the output results of the multiple feature extraction modules are used for the target music source separation network to perform music source separation on the second mixed music source to be processed; wherein, the second mixed music source is the next segment of the mixed music source of the first mixed music source.

[0085] For example, assuming that the complete mixed music source data includes multiple segments of mixed music sources, then, the first mixed music source and the second mixed music source may be two adjacent segments of the multiple segments of mixed music sources. Among them, the second mixed music source is after the first mixed music source. In one example, the first mixed music source and the second mixed music source may respectively include 8 frames of sound data, and among them, the 8 frames of sound data included in the second music source are after the 8 frames of sound data included in the first mixed music source.

[0086] In this embodiment, during the process of separating the first mixed music source using the target music source separation network, the output results of multiple feature extraction modules in the target music source separation network can be saved (or cached). In this way, when subsequently separating the second mixed music source using the target music source separation network, the saved information can be combined to process the second mixed music source, so as to ensure that at least two sound components separated from the second mixed music source can be correctly predicted. In addition, by dividing the complete mixed music source data into multiple segments of mixed music sources for processing one by one, the effect of real-time music source separation of the mixed music source data can be achieved, and thus it can be applied to scenarios that require real-time music source separation.

[0087] For example, in Figure 4 , the processing result of the feature extraction module #1 on the first mixed music source can be saved (cached). Thus, when the second mixed music is input, the feature extraction module #1 can process the second mixed music by combining the saved processing result of this feature extraction module #1; for another example, the processing result of the feature extraction module #2 on the first mixed music source can be saved (cached). Thus, when the second mixed music is input, the feature extraction module #2 can process the second mixed music by combining the saved processing result of this feature extraction module #2.

[0088] In some embodiments, before separating the first mixed music source using the target music source separation network, the method may further include: windowing the first mixed music source using an analysis window, and performing a short-time Fourier transform (STFT) on the result of the windowing process. In this way, the complex spectrum features of the first mixed music source can be extracted, and thus the target music source separation network can separate the first mixed music source based on the complex spectrum features of the first mixed music source. After obtaining the prediction result of at least two sound components separated from the first mixed music source, the method may further include: performing an inverse short-time Fourier transform (ISTFT) on the prediction result, and windowing the result of the inverse short-time Fourier transform using a synthesis window to reconstruct the time-domain signal of the prediction result. Among them, the analysis window and the synthesis window can be, for example, Hamming windows.

[0089] The embodiments of the present application also provide a training method for a music source separation network. As Figure 9 shown, the training method may include:

[0090] S901, separating the mixed music source in the training set using the first music source separation network to obtain a first prediction result of at least two sound components separated from the mixed music source.

[0091] In this step, a first music source separation network can be used to separate the mixed music source in the training set. Among them, the mixed music source can be composed of, for example, at least one frame of sound data, and the mixed music source can contain multiple sound components. For example, the mixed music source can contain, but is not limited to, at least two of the following sound components: vocals, drums, bass. In some embodiments, when the mixed music source is composed of multiple frames of sound data, there may be partial overlap between two adjacent frames of sound data in the multiple frames of sound data. For example, the length of each frame of sound data is 2048 sample points, and there are 1536 overlapping sample points between two adjacent frames of sound data.

[0092] Among them, the first music source separation network can be, for example, a pre-trained music source separation network, and the number of parameters of the first music source separation network can be set to be large enough to achieve a good music source separation effect, so that it can guide the training of the second music source separation network.

[0093] In the embodiments of the present application, the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network. The first music source separation network can be used as a teacher model to guide the training of the second music source separation network (student model), or in other words, to perform knowledge distillation on the second music source separation network to obtain a target music source separation network.

[0094] It should be noted that the network structures of the first music source separation network and the second music source separation network can be the same or different, and the embodiments of the present application do not limit this. When the network structures of the first music source separation network and the second music source separation network are the same, the fact that the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network can be manifested as follows: compared with the second music source separation network, the first music source separation network has more layers / modules. For example, the network structure of the second music source separation network is as Figure 4 shown, including 7 feature extraction modules, 3 downsampling modules and 3 upsampling modules. Then, the first music source separation network can be based on Figure 4 this, add more feature extraction modules, downsampling modules and upsampling modules to achieve better prediction performance.

[0095] S902, Use the second music source separation network to separate the mixed music source to obtain a second prediction result of at least two sound components separated from the mixed music source.

[0096] In some embodiments, the second music source separation network may include: an encoding sub-network and a decoding sub-network; the encoding sub-network includes a plurality of feature extraction modules and a plurality of downsampling modules; the decoding sub-network includes a plurality of feature extraction modules and a plurality of upsampling modules; the number of downsampling modules is the same as that of the upsampling modules. The structure of the second music source separation network can be seen in Figure 4 .

[0097] In some embodiments, the second music source separation network is used to perform music source separation on a mixed music source to obtain a second prediction result of at least two sound components separated from the mixed music source, including: alternately using the feature extraction module and the downsampling module in the encoding sub-network to perform encoding processing on the mixed music source to obtain an encoded mixed music source; alternately using the upsampling module and the feature extraction module in the decoding sub-network to perform decoding processing on the encoded mixed music source to obtain the second prediction result.

[0098] For the detailed steps of using the second music source separation network to perform music source separation on the mixed music source, reference can be made to the related description of using the target music source separation network to perform music source separation on the first mixed music source in S302, which will not be elaborated here.

[0099] In some embodiments, the feature extraction module includes: at least one time-frequency feature extraction unit and at least one one-way gated recurrent unit. Among them, the time-frequency feature extraction unit is used to perform causal convolution on the first data input to the time-frequency feature extraction unit in both the time domain and the frequency domain to obtain the time-frequency features of the first data; the one-way gated recurrent unit is used to extract the timing information of the second data input to the one-way gated recurrent unit to obtain the timing features of the second data.

[0100] In some embodiments, the feature extraction module includes: a plurality of cascaded feature extraction sub-modules, and each feature extraction sub-module includes two time-frequency feature extraction units and one one-way gated recurrent unit. The processing flow of each feature extraction sub-module (see Figure 8 ) includes: using the first time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the data input to the feature extraction sub-module in both the time domain and the frequency domain to obtain the first time-frequency features; using the one-way gated recurrent unit in the feature extraction sub-module to extract the timing information of the first time-frequency features to obtain the first timing features; using the second time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the result of adding the first time-frequency features and the first timing features in both the time domain and the frequency domain to obtain the second time-frequency features; performing convolution processing (such as Shortcut) on the data input to the feature extraction sub-module, and adding the result of the convolution processing and the second time-frequency features as the output result of the feature extraction sub-module.

[0101] In some embodiments, before separating the mixed music source using the second music source separation network, the method may further include: windowing the mixed music source using an analysis window and performing a short-time Fourier transform on the result of the windowing; after obtaining the second prediction result of at least two sound components separated from the mixed music source, the method may further include: performing an inverse short-time Fourier transform on the second prediction result and windowing the result of the inverse short-time Fourier transform using a synthesis window. The analysis window and the synthesis window may be, for example, Hamming windows.

[0102] S903. Using the first music source separation network as the teacher model and the second music source separation network as the student model, based on the first prediction result and the second prediction result, perform knowledge distillation on the second music source separation network to obtain the target music source separation network.

[0103] In this step, the target music source separation network can be obtained through knowledge distillation. The target music source separation network can be used to separate the music source of the mixed music source (such as the first mixed music source; or the second mixed music source) in S302.

[0104] In some embodiments, performing knowledge distillation on the second music source separation network based on the first prediction result and the second prediction result to obtain the target music source separation network includes: determining the target loss based on the loss between the first prediction result and the second prediction result and the loss between the second prediction result and the true label of the at least two sound components; adjusting the network parameters of the second music source separation network based on the target loss so that the loss output by the adjusted second music source separation network satisfies the convergence condition to obtain the target music source separation network.

[0105] For example, assume that the first prediction result is denoted as q, the second prediction result is denoted as p, and the true label of the at least two sound components is denoted as y. Then, the loss between the first prediction result q and the second prediction result p can be expressed as MSELoss(q, p), and the loss between the second prediction result p and the true label y can be expressed as MSELoss(y, p). The target loss Loss can be calculated, for example, by formula (5):

[0106] Loss = MSELoss(y, p) + w × MSELoss(q, p) (5);

[0107] Among them, MSELoss is the Mean Squared Error (MSE), and w is the weight coefficient. According to formula (5), the first prediction result q can be used as the target, so that the second prediction result p is close to the first prediction result q, thereby achieving the purpose of guiding the training of the second music source separation network based on the first music source separation network. After calculating the target loss Loss, the network parameters of the second music source separation network can be adjusted based on the target loss Loss, so that the loss output by the adjusted second music source separation network meets the convergence condition to obtain the required target music source separation network.

[0108] According to the method of this embodiment, knowledge distillation is performed on the second music source separation network with a smaller number of parameters using the first music source separation network with a larger number of parameters, and a target music source separation network with better performance can be obtained, thereby improving the effect of music source separation. At the same time, since the number of parameters of the second music source separation network is small, that is to say, the number of parameters of the obtained target music source separation network is small. Therefore, the time delay during music source separation using the target music source separation network can be low, and thus it can be applied to scenarios that require real-time music source separation.

[0109] As mentioned above in combination with Figures 1 to 9 This application embodiment provides a method for music source separation and a method for training a music source separation network. To facilitate understanding of this application embodiment, a possible implementation process of the music source separation method and the music source separation network training method provided by this application embodiment is introduced below with examples.

[0110] In the implementation process of music source separation, artificial intelligence technology can be used to extract audio tracks from stereo mixed music (mixed music sources). The audio tracks that can be extracted can include, for example: vocals, drums, bass, and others. As Figure 10 shown, for the input mixed music source, preprocessing, music source separation, and postprocessing can be performed in sequence to obtain 4 audio tracks (vocals, drums, bass, and others) extracted from the mixed music source. Among them, the input mixed music source can contain 2 channels for example, the output contains 8 channels, and each extracted audio track contains 2 channels.

[0111] In the preprocessing stage, the time-domain signal of the mixed music source can be converted into a frequency-domain signal through STFT, and this solution uses complex spectral features. These features are then sent to the deep learning model for separated-track music extraction (target music source separation network) to generate complex spectral features of 4 independent audio tracks. In the postprocessing stage, ISTFT can be performed on the complex spectral features of these 4 independent audio tracks to reconstruct the time-domain signal of each audio track.

[0112] Figure 11Schematic diagram of the process for separating music sources from a mixed music source provided by an embodiment of the present application.

[0113] Among them, the target music source separation network can be obtained through knowledge distillation. As Figure 11 shown, the steps for separating music sources from a mixed music source may include:

[0114] S111, Preprocessing.

[0115] For the Wav sample of the input mixed music source, it can be preprocessed through steps 1111) to 1113) first.

[0116] 1111) Framing: The input signal (Wav sample of the mixed music source) is divided into frames based on a window length of 2048 points and a hop length of 512 points.

[0117] 1112) Analysis Window: Each frame of the signal is windowed with a 2048-point Hamming window.

[0118] 1113) STFT: During the STFT process, a 2048-point FFT complex spectrum is applied to the windowed signal through the Fast Fourier Transform (FFT) to obtain the complex spectral characteristics of the input signal, which contain amplitude and phase information.

[0119] In this embodiment, every time 8 frames of sound data are obtained, preprocessing is performed once to obtain the complex spectral characteristics of the current 8 frames of sound data.

[0120] S112, Separation of the mixed music source.

[0121] After obtaining the complex spectral characteristics of the current 8 frames of sound data, if the characteristics of the previous 8 frames of sound data are cached in the feature buffer, the complex spectral characteristics of the current 8 frames of sound data and part of the characteristics of the previous 8 frames of sound data cached in the buffer can be concatenated, and the concatenated characteristics are input into the target music source separation network.

[0122] During the process of the target music source separation network processing the input characteristics, the characteristics of the current 8 frames of sound data obtained (such as the characteristics output by the TFC-GRU module in the target music source separation network) can be saved (cached) in the feature buffer, and the currently cached characteristics (such as part of the characteristics) can be used for concatenation with the characteristics of the next 8 frames of sound data.

[0123] After being processed by the target music source separation network, the complex spectral features of the 4 independent audio tracks separated from the current 8-frame audio data can be predicted (the current prediction result). At this time, if the complex spectral features of the 4 independent audio tracks separated from the previous 8-frame audio data (the previous prediction result) are cached in the prediction result buffer, the current prediction result and a partial prediction result of the previous time can be concatenated. After the concatenated result is post-processed, the Wav samples of the 4 independent audio tracks corresponding to the current 8-frame audio data can be obtained. At the same time, the target music source separation network also needs to save (cache) the current prediction result to the prediction result buffer, and the currently cached prediction result (such as a partial prediction result) can be used for concatenation with the next prediction result.

[0124] S113, Post-processing.

[0125] Among them, post-processing can also be called multi-track music reconstruction. The process of post-processing includes:

[0126] 1131) ISTFT: During the ISTFT process, an inverse fast Fourier transform (IFFT) operation of 2048 points is performed on the complex spectral features of each independent audio track.

[0127] 1132) Synthesis Window: A 2048-point Hamming window is used for windowing.

[0128] 1133) Overlap-Add: Overlap and add are performed on each frame of the signal, and the final audio track signal, that is, the Wav samples of the 4 independent audio tracks, is output.

[0129] According to the above technical means, during the inference stage, after every 8 frames of stereo audio data are obtained, a preprocessing can be performed to obtain the complex spectral features including amplitude and phase information; then, the target music source separation network can be used to process the complex spectral features to predict the complex spectral features of 4 independent audio tracks; subsequently, the 8-frame complex spectral features of the 4 independent audio tracks predicted by the current model and the 8-frame complex spectral features predicted last time are concatenated and processed using the post-processing module, so as to obtain accurate 8-frame multi-track music, and the 8-frame prediction result of this time is retained for the next prediction. By performing processing every 8 frames, real-time music source separation of the mixed audio can be achieved. In addition, during the current prediction process, the previous prediction result can be combined to ensure the correctness of the current prediction result.

[0130] Figure 12 Shows the structural schematic of the target music source separation network provided by the embodiment of the present application Figure 2Among them, the hybrid complex spectrum can be the complex spectral features of 8 frames of sound data obtained in S111; the separated-track complex spectrum can be the complex spectral features of 4 independent audio tracks separated from the current 8 frames of sound data obtained in S112.

[0131] Figure 12 The structure of the target music source separation network shown is designed based on the U-NET structure, and this structure can be composed of an encoder (corresponding to the encoding sub-network in the foregoing embodiment), an intermediate layer, and a decoder (corresponding to the decoding sub-network in the foregoing embodiment). In some scenarios, the intermediate layer can also be considered as a part of the encoder (encoding sub-network).

[0132] Among them, there are 3 TFC-GRU modules (corresponding to the feature extraction module in the foregoing embodiment) and 3 downsampling layers (DownSample) (corresponding to the downsampling module in the foregoing embodiment) at the encoder end, and 3 TFC-GRU modules and 3 upsampling layers (UpSample) (corresponding to the upsampling module in the foregoing embodiment) at the decoder end. There is also an intermediate layer between the encoder and the decoder. Among them, the encoder of U-NET has 3 times of downsampling to obtain high-level semantic information, and the decoder corresponds to 3 times of upsampling to perform resolution restoration. In order to reduce the loss of spatial information caused by the downsampling process, skip connections are introduced, so that the feature maps restored by upsampling contain more low-level semantic information, making the result finer.

[0133] In Figure 12 , the features output by the current TFC-GRU module can be saved (cached) in the feature buffer. When the next 8 frames of sound data are input, the TFC-GRU module can read and combine the cached features to process the 8 frames of sound data.

[0134] Exemplarily, the TFC-GRU module can include at least one time-frequency feature extraction unit and at least one one-way gated recurrent unit. Among them, the time-frequency feature extraction unit (see Figure 5 ) can include batch normalization, GeLU, and causal 2D convolution. Since the convolutional network structure in the encoder and the decoder includes causal convolution, it can be ensured that the information of future time is not used, thereby reducing the latency and achieving the effect of streaming inference.

[0135] The one-way gated recurrent unit (see Figure 6) It may include batch normalization, GeLU, and GRU (such as unidirectional GRU). Among them, GRU is a type of recurrent neural network. To solve the problem of gradient vanishing in standard RNNs, GRU uses so-called "update gates" and "reset gates". These two vectors determine which information should be passed to the output. What makes them special is that they can be trained to preserve information from a long time ago, rather than disappearing over time or deleting information irrelevant to the prediction.

[0136] In order not to destroy the streaming inference ability of the model, this embodiment may adopt a unidirectional GRU model, which can make full use of the temporal information of the music signal and can preserve information from a long time ago, thereby improving the effect of the streaming model.

[0137] In the training stage, this embodiment may adopt knowledge distillation technology to improve the robustness of the streaming model.

[0138] The process of the training stage includes: preprocessing each batch of stereo music (mixed music source) training data to obtain complex spectral features containing amplitude and phase information; training a teacher model with a larger number of training parameters and higher latency (corresponding to the first music source separation network in the foregoing embodiment); using knowledge distillation technology, using the trained teacher model above to guide the training of a student model (corresponding to the second music source separation network in the foregoing embodiment), and this student model has a smaller number of parameters and lower latency and can meet the requirements of end-side real-time music source separation.

[0139] Exemplarily, the model can be regarded as a black box, and knowledge can be regarded as the mapping relationship from input to output. Therefore, a teacher model can be trained first, and then the output result q of the teacher model can be used as the target of the student model to guide the training of the student model, so that the output result p of the student model is close to q. Therefore, the loss function can be represented by formula (6):

[0140] Loss = MSELoss(y, p) + w × MSELoss(q, p) (6);

[0141] Among them, MSELoss is the mean squared error loss function, y is the true label, q is the output result of the teacher model, p is the output result of the student model, and w is the weight coefficient.

[0142] The calculation formula of MSELoss is shown in formula (7):

[0143]

[0144] Among them, n is the number of data samples. Since the student model is used in the inference stage, the number of parameters of the teacher model can be set to be large enough to achieve a good effect and can guide the training of the student model.

[0145] Figure 13 This is a schematic diagram of the knowledge distillation process provided by the embodiments of the present application. As Figure 13 shown, the teacher model and the student model can respectively process the training data and output their respective prediction data. Further, formula (6) can be used to calculate the loss Loss based on the prediction data q of the teacher model, the prediction data p of the student model, and the label data (true label) y, and the network parameters of the student model can be adjusted based on this loss Loss so that the loss output by the adjusted student model meets the convergence condition to obtain the final music source separation network for inference (i.e., the target music source separation network).

[0146] It can be understood that the number of parameters of the model and the representation ability of the model are positively correlated. The smaller the model, the generally worse the representation ability, and vice versa. Therefore, in order to improve the separation effect of the streaming small model, this embodiment introduces the knowledge distillation technology, and uses the teacher model with a large number of parameters and high latency to guide the training of the student model with a small number of parameters and low latency, so as to effectively improve the separation effect of the student model, or rather, a target music source separation network with a better separation effect can be obtained.

[0147] It should be understood that separating the mixed music source into four tracks (vocals, drums, bass, others) in the above solution is only exemplary. In actual applications, separating more tracks is also supported.

[0148] It should also be understood that the complex spectral features of the mixed music source are used as the input in the above solution. In some other scenarios, the time-domain signal and frequency-domain signal of the mixed music source can also be used as the input simultaneously to further improve the separation effect.

[0149] The music source separation method proposed by the embodiments of the present application is applicable to the application scenarios of holographic audio / spatial audio. In this scenario, the MSS solution has the following technical problems to be solved: the solution needs to be able to process real-time audio streams to ensure that the separated tracks can be transmitted to the downstream path in real time; the solution effect needs to be good enough so that the music recreated after splitting the tracks can meet the needs of users; the number of model parameters needs to be designed very small because the storage space of terminal devices is limited.

[0150] The music source separation method provided by the embodiments of the present application can achieve the following functions:

[0151] 1) The network structure of the TFC-GRU module is proposed. The TFC module can perform convolution on time-frequency features (including time-domain features and frequency-domain features), and the GRU module can model the time-domain dimension (sequential information). This network structure makes full use of the time-domain features and frequency-domain features of music signals, thereby improving the performance of the model.

[0152] 2) The introduction of the unidirectional GRU model ensures that the model does not utilize information from future time, while leveraging the sequential information of music signals and ensuring the streaming structure of the model.

[0153] 3) The knowledge distillation technique is introduced into the music source separation task. A teacher model with a large amount of data and high latency is used to guide the training of a student model with a small amount of data and low latency, improving the separation effect of the streaming model.

[0154] The music source separation method proposed in the embodiments of this application can expand the track upmixing framework of holographic audio, and output high-quality audio track data of different instruments in real time to the downstream spatial rendering module; enabling holographic audio / spatial audio to support richer and more flexible audio spatialization effects, and enhancing the sense of participation and playability of user-defined sound effects.

[0155] Based on the foregoing embodiments, the embodiments of this application provide related devices for music source separation. The devices include the various modules included, as well as the various sub-modules included in each module, and can be implemented by a processor in a computer device with information processing capabilities; of course, they can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0156] Figure 14 The composition structure of a music source separation device provided by the embodiments of this application is shown. As Figure 14 shown, the music source separation device 1400 (hereinafter simply referred to as device 1400) may include:

[0157] An acquisition unit 1401, configured to acquire a first mixed music source to be processed; a music source separation unit 1402, configured to perform music source separation on the first mixed music source by using a target music source separation network to obtain a prediction result of at least two sound components separated from the first mixed music source; wherein, the target music source separation network is obtained by performing knowledge distillation on a second music source separation network with a first music source separation network as the teacher model; the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network.

[0158] In some embodiments, the target music source separation network includes: an encoding sub-network and a decoding sub-network; the encoding sub-network includes a plurality of feature extraction modules and a plurality of downsampling modules; the decoding sub-network includes a plurality of feature extraction modules and a plurality of upsampling modules; the number of downsampling modules is the same as that of the upsampling modules; the music source separation unit 1402 is specifically configured to: alternately use the feature extraction modules and the downsampling modules in the encoding sub-network to perform encoding processing on the first mixed music source to obtain the encoded first mixed music source; alternately use the upsampling modules and the feature extraction modules in the decoding sub-network to perform decoding processing on the encoded first mixed music source to obtain a prediction result.

[0159] In some embodiments, the feature extraction module includes: at least one time-frequency feature extraction unit and at least one one-way gated recurrent unit; the time-frequency feature extraction unit is configured to perform causal convolution on the first data input to the time-frequency feature extraction unit in both the time domain and the frequency domain to obtain the time-frequency feature of the first data; the one-way gated recurrent unit is configured to extract temporal information from the second data input to the one-way gated recurrent unit to obtain the temporal feature of the second data.

[0160] In some embodiments, the feature extraction module includes: a plurality of serially-connected feature extraction sub-modules, and each feature extraction sub-module includes two time-frequency feature extraction units and one one-way gated recurrent unit; the processing flow of each feature extraction sub-module includes: using the first time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the data input to the feature extraction sub-module in both the time domain and the frequency domain to obtain the first time-frequency feature; using the one-way gated recurrent unit in the feature extraction sub-module to extract temporal information from the first time-frequency feature to obtain the first temporal feature; using the second time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the result of adding the first time-frequency feature and the first temporal feature in both the time domain and the frequency domain to obtain the second time-frequency feature; performing convolution processing on the data input to the feature extraction sub-module and adding the result of the convolution processing and the second time-frequency feature as the output result of the feature extraction sub-module.

[0161] In some embodiments, the apparatus 1400 further includes: a storage unit configured to store the output results of a plurality of feature extraction modules, and the output results of the plurality of feature extraction modules are used for the target music source separation network to perform music source separation on a second mixed music source to be processed; wherein, the second mixed music source is the next-segment mixed music source of the first mixed music source.

[0162] In some embodiments, the apparatus 1400 further includes: a preprocessing unit, configured to perform windowing processing on the first mixed music source by using an analysis window and perform short-time Fourier transform on the result of the windowing processing before performing music source separation on the first mixed music source by using a target music source separation network; a postprocessing unit, configured to perform inverse short-time Fourier transform on the prediction result after obtaining the prediction result of at least two sound components separated from the first mixed music source, and perform windowing processing on the result of the inverse short-time Fourier transform by using a synthesis window.

[0163] Figure 15 FIG. shows the composition structure of a training apparatus for a music source separation network provided by an embodiment of the present application. As Figure 15 shown, the training apparatus 1500 for a music source separation network (hereinafter simply referred to as the apparatus 1500) may include:

[0164] A first music source separation unit 1501, configured to perform music source separation on the mixed music source in the training set by using a first music source separation network to obtain a first prediction result of at least two sound components separated from the mixed music source; a second music source separation unit 1502, configured to perform music source separation on the mixed music source by using a second music source separation network to obtain a second prediction result of at least two sound components separated from the mixed music source; wherein, the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network; a knowledge distillation unit 1503, configured to use the first music source separation network as a teacher model and the second music source separation network as a student model, and perform knowledge distillation on the second music source separation network based on the first prediction result and the second prediction result to obtain a target music source separation network.

[0165] In some embodiments, the knowledge distillation unit 1503 is specifically configured to: determine a target loss based on the loss between the first prediction result and the second prediction result and the loss between the second prediction result and the true labels of at least two sound components; adjust the network parameters of the second music source separation network based on the target loss so that the loss output by the adjusted second music source separation network satisfies a convergence condition to obtain a target music source separation network.

[0166] In some embodiments, the second music source separation network includes: an encoding sub-network and a decoding sub-network; the encoding sub-network includes a plurality of feature extraction modules and a plurality of downsampling modules; the decoding sub-network includes a plurality of feature extraction modules and a plurality of upsampling modules; the number of downsampling modules is the same as the number of upsampling modules; the second music source separation unit 1502 is specifically configured to: alternately use the feature extraction modules and the downsampling modules in the encoding sub-network to perform encoding processing on the mixed music source to obtain an encoded mixed music source; alternately use the upsampling modules and the feature extraction modules in the decoding sub-network to perform decoding processing on the encoded mixed music source to obtain a second prediction result.

[0167] In some embodiments, the feature extraction module includes: at least one time-frequency feature extraction unit and at least one one-way gated recurrent unit; the time-frequency feature extraction unit is configured to perform causal convolution on the first data input to the time-frequency feature extraction unit in both the time domain and the frequency domain to obtain the time-frequency features of the first data; the one-way gated recurrent unit is configured to extract the temporal information from the second data input to the one-way gated recurrent unit to obtain the temporal features of the second data.

[0168] In some embodiments, the feature extraction module includes: a plurality of serially connected feature extraction sub-modules, each feature extraction sub-module including two time-frequency feature extraction units and one one-way gated recurrent unit; the processing flow of each feature extraction sub-module includes: using the first time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the data input to the feature extraction sub-module in both the time domain and the frequency domain to obtain the first time-frequency features; using the one-way gated recurrent unit in the feature extraction sub-module to extract the temporal information from the first time-frequency features to obtain the first temporal features; using the second time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the result of adding the first time-frequency features and the first temporal features in both the time domain and the frequency domain to obtain the second time-frequency features; performing convolution processing on the data input to the feature extraction sub-module and adding the result of the convolution processing and the second time-frequency features as the output result of the feature extraction sub-module.

[0169] In some embodiments, the apparatus 1500 further includes: a preprocessing unit configured to perform windowing on the mixed music source using an analysis window and perform short-time Fourier transform on the result of the windowing before separating the mixed music source using the second music source separation network; a postprocessing unit configured to perform inverse short-time Fourier transform on the second prediction result after obtaining the second prediction result of at least two sound components separated from the mixed music source and perform windowing on the result of the inverse short-time Fourier transform using a synthesis window.

[0170] The description of the above apparatus embodiments is similar to the description of the above method embodiments and has similar beneficial effects to those of the method embodiments. In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the apparatus embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0171] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software functional modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, read-only memory (ROM), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination among hardware, software, and firmware.

[0172] The embodiments of the present application further provide a music source separation device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0173] The embodiments of the present application further provide a chip. The chip includes: a processor for calling and running a computer program from the memory, so that a device installed with the chip executes some or all of the steps in the above method.

[0174] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.

[0175] The embodiments of the present application further provide a computer program, including computer-readable code. When the computer-readable code runs in a device (such as a music source separation device), the processor in the device executes some or all of the steps in the above method.

[0176] The embodiments of the present application further provide a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method. The computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0177] It should be noted here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or resemblances can be referred to each other. The descriptions of the above embodiments of the device, chip, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, chip, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0178] Figure 16 This is a schematic diagram of a hardware entity of the music source separation device in an embodiment of the present application. Figure 16 The music source separation device 1600 shown includes a processor 1610. The processor 1610 can call and run a computer program from a memory to implement the method in the embodiments of the present application.

[0179] In some embodiments, as Figure 16 shown, the music source separation device 1600 may further include a memory 1620. Among them, the processor 1610 can call and run a computer program from the memory 1620 to implement the method in the embodiments of the present application. Among them, the memory 1620 can be a separate device independent of the processor 1610 or integrated in the processor 1610.

[0180] In some embodiments, as Figure 16 shown, the music source separation device 1600 may further include a transceiver 1630. The processor 1610 can control the transceiver 1630 to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices. Among them, the transceiver 1630 may include a transmitter and a receiver. The transceiver 1630 may further include an antenna, and the number of antennas may be one or more.

[0181] It should be understood that the "one embodiment", "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment", "in an embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above steps / processes do not mean the order of execution. The execution order of each step / process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0182] It should be noted that, in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element.

[0183] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units or modules is only a logical functional division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0184] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0185] In addition, each functional unit in the embodiments of this application can be all integrated in a processing unit, or each unit can be separately a unit alone, or two or more units can be integrated in a unit; the above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0186] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage media include: various media that can store program codes such as removable storage devices, read-only memory (ROM), magnetic disks or optical discs.

[0187] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as removable storage devices, ROMs, magnetic disks, or optical discs.

[0188] As described above, the above are only the implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A method for music source separation, characterized in that, the method comprises: obtaining a first mixed music source to be processed; performing music source separation on the first mixed music source by using a target music source separation network to obtain prediction results of at least two sound components separated from the first mixed music source; wherein, the target music source separation network is obtained by performing knowledge distillation on a second music source separation network with a first music source separation network as a teacher model; the number of parameters of the first music source separation network is greater than that of the second music source separation network.

2. The method according to claim 1, characterized in that, the target music source separation network comprises: an encoding sub-network and a decoding sub-network; the encoding sub-network comprises a plurality of feature extraction modules and a plurality of downsampling modules; the decoding sub-network comprises a plurality of feature extraction modules and a plurality of upsampling modules; the number of the downsampling modules is the same as that of the upsampling modules; the performing music source separation on the first mixed music source by using the target music source separation network to obtain prediction results of at least two sound components separated from the first mixed music source comprises: alternately using the feature extraction modules and the downsampling modules in the encoding sub-network to perform encoding processing on the first mixed music source to obtain an encoded first mixed music source; alternately using the upsampling modules and the feature extraction modules in the decoding sub-network to perform decoding processing on the encoded first mixed music source to obtain the prediction results.

3. The method according to claim 2, characterized in that, the feature extraction module comprises: at least one time-frequency feature extraction unit and at least one one-way gated recurrent unit; the time-frequency feature extraction unit is configured to perform causal convolution on first data input to the time-frequency feature extraction unit in both the time domain and the frequency domain to obtain time-frequency features of the first data; the one-way gated recurrent unit is configured to extract temporal information from second data input to the one-way gated recurrent unit to obtain temporal features of the second data.

4. The method according to claim 2 or 3, characterized in that, the feature extraction module comprises: a plurality of serially-connected feature extraction sub-modules, and each feature extraction sub-module comprises two time-frequency feature extraction units and one one-way gated recurrent unit; the processing flow of each feature extraction sub-module comprises: using the first time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on data input to the feature extraction sub-module in both the time domain and the frequency domain to obtain first time-frequency features; using the one-way gated recurrent unit in the feature extraction sub-module to extract temporal information from the first time-frequency features to obtain first temporal features; using the second time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the result of adding the first time-frequency features and the first temporal features in both the time domain and the frequency domain to obtain second time-frequency features; Perform convolution processing on the data input to the feature extraction sub-module, and add the result of the convolution processing to the second time-frequency feature as the output result of the feature extraction sub-module.

5. The method according to claim 2 or 3, wherein, the method further includes: saving the output results of the multiple feature extraction modules, and the output results of the multiple feature extraction modules are used for the target music source separation network to perform music source separation on the second mixed music source to be processed; wherein, the second mixed music source is the next segment of the mixed music source of the first mixed music source.

6. The method according to any one of claims 1 to 3, wherein, before using the target music source separation network to perform music source separation on the first mixed music source, the method further includes: windowing the first mixed music source using an analysis window, and performing short-time Fourier transform on the result of the windowing process; after obtaining the prediction results of at least two sound components separated from the first mixed music source, the method further includes: performing inverse short-time Fourier transform on the prediction results, and windowing the result of the inverse short-time Fourier transform using a synthesis window.

7. A training method for a music source separation network, wherein, the method includes: using a first music source separation network to perform music source separation on the mixed music source in the training set, and obtaining a first prediction result of at least two sound components separated from the mixed music source; using a second music source separation network to perform music source separation on the mixed music source, and obtaining a second prediction result of at least two sound components separated from the mixed music source; wherein, the number of parameters of the first music source separation network is greater than the number of parameters of the second music source separation network; using the first music source separation network as a teacher model and the second music source separation network as a student model, and performing knowledge distillation on the second music source separation network based on the first prediction result and the second prediction result to obtain a target music source separation network.

8. The method according to claim 7, wherein, performing knowledge distillation on the second music source separation network based on the first prediction result and the second prediction result to obtain a target music source separation network includes: determining a target loss based on the loss between the first prediction result and the second prediction result, and the loss between the second prediction result and the true labels of the at least two sound components; adjusting the network parameters of the second music source separation network based on the target loss, so that the loss output by the adjusted second music source separation network satisfies the convergence condition to obtain the target music source separation network.

9. The method according to claim 7 or 8, wherein, the second music source separation network includes: an encoding sub-network and a decoding sub-network; the encoding sub-network includes multiple feature extraction modules and multiple downsampling modules; the decoding sub-network includes multiple feature extraction modules and multiple upsampling modules; the number of downsampling modules is the same as the number of upsampling modules; Performing music source separation on the mixed music source using the second music source separation network to obtain a second prediction result of at least two sound components separated from the mixed music source, including: Alternately using the feature extraction module and the downsampling module in the encoding sub-network to perform encoding processing on the mixed music source to obtain an encoded mixed music source; Alternately using the upsampling module and the feature extraction module in the decoding sub-network to perform decoding processing on the encoded mixed music source to obtain the second prediction result.

10. The method according to claim 9, wherein, the feature extraction module includes: at least one time-frequency feature extraction unit and at least one unidirectional gated recurrent unit; the time-frequency feature extraction unit is configured to perform causal convolution on the first data input to the time-frequency feature extraction unit in both the time domain and the frequency domain to obtain the time-frequency feature of the first data; the unidirectional gated recurrent unit is configured to extract temporal information from the second data input to the unidirectional gated recurrent unit to obtain the temporal feature of the second data.

11. The method according to claim 9, wherein, the feature extraction module includes: a plurality of serially connected feature extraction sub-modules, and each feature extraction sub-module includes two time-frequency feature extraction units and one unidirectional gated recurrent unit; the processing flow of each feature extraction sub-module includes: using the first time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the data input to the feature extraction sub-module in both the time domain and the frequency domain to obtain a first time-frequency feature; using the unidirectional gated recurrent unit in the feature extraction sub-module to extract temporal information from the first time-frequency feature to obtain a first temporal feature; using the second time-frequency feature extraction unit in the feature extraction sub-module to perform causal convolution on the result of adding the first time-frequency feature and the first temporal feature in both the time domain and the frequency domain to obtain a second time-frequency feature; performing convolution processing on the data input to the feature extraction sub-module and adding the result of the convolution processing and the second time-frequency feature as the output result of the feature extraction sub-module.

12. The method according to claim 7 or 8, wherein, before performing music source separation on the mixed music source using the second music source separation network, the method further includes: performing windowing processing on the mixed music source using an analysis window and performing short-time Fourier transform on the result of the windowing processing; after obtaining the second prediction result of at least two sound components separated from the mixed music source, the method further includes: performing inverse short-time Fourier transform on the second prediction result and performing windowing processing on the result of the inverse short-time Fourier transform using a synthesis window.

13. A music source separation device, wherein, the device includes: an acquisition unit configured to acquire a first mixed music source to be processed; a music source separation unit configured to perform music source separation on the first mixed music source using a target music source separation network to obtain a prediction result of at least two sound components separated from the first mixed music source; Among them, the target music source separation network is obtained by performing knowledge distillation on the second music source separation network with the first music source separation network as the teacher model and the second music source separation network as the student model; the number of parameters of the first music source separation network is greater than that of the second music source separation network.

14. A training device for a music source separation network, characterized in that the device includes: A first music source separation unit, configured to perform music source separation on the mixed music source in the training set by using a first music source separation network, so as to obtain a first prediction result of at least two sound components separated from the mixed music source; A second music source separation unit, configured to perform music source separation on the mixed music source by using a second music source separation network, so as to obtain a second prediction result of at least two sound components separated from the mixed music source; wherein, the number of parameters of the first music source separation network is greater than that of the second music source separation network; A knowledge distillation unit, configured to perform knowledge distillation on the second music source separation network with the first music source separation network as the teacher model and the second music source separation network as the student model, based on the first prediction result and the second prediction result, to obtain a target music source separation network.

15. A music source separation device, characterized in that the device includes: A memory, configured to store computer-executable instructions; A processor, connected to the memory, configured to implement the method according to any one of claims 1 to 6, or implement the method according to any one of claims 7 to 12 by executing the computer-executable instructions.

16. A chip, characterized in that the chip includes: A processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the method according to any one of claims 1 to 6, or executes the method according to any one of claims 7 to 12.

17. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by at least one processor, it implements the method according to any one of claims 1 to 6, or implements the method according to any one of claims 7 to 12.

Citation Information

Cited By

  • Music source separation method and wearable device

    CN120913585A