Audio separation method, device, electronic device and computer-readable storage medium
By using a two-dimensional window self-attention network and a dual-strategy framework that first roughly divides and then finely adjusts in the audio separation method, the problem of insufficient multi-task audio separation performance in the existing technology is solved, and a more efficient audio separation effect is achieved.
Patent Information
- Application Number
- CN202211100895.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-09-09
AI Technical Summary
The existing multitasking audio separation method has shortcomings in extracting global dependencies, local dependencies and similarities between time and frequency units, resulting in the need to improve the separation performance.
An audio separation method is used to extract features using a two-dimensional window self-attention network (2D-WA), including a multi-head self-attention layer and a two-dimensional window self-attention layer, which is used to capture global and local context dependencies, and to estimate and adjust residuals to improve separation effect through a dual-strategy framework that first roughly divides and then finely adjusts.
By comprehensively capturing the information required for multitasking audio separation, the performance of audio separation is improved, especially in terms of coarse division and residual estimation capabilities.
Smart Images

Figure CN115641868B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio processing, and particularly to an audio separation method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] In real life, audio signals mainly include speech, music, and background noise, etc. When these signals are mixed, the intelligibility of the audio signal decreases, which will damage subsequent audio understanding tasks. For example, in a music information retrieval task, foreground speech and background noise will reduce the retrieval accuracy; in a speech recognition task, background noise will reduce the recognition accuracy, and at the same time, the music in the audio will cause the recognition result to be chaotic. At the same time, in a speech quality assessment task, the energy level of the extracted background noise is an important assessment basis. It can be seen that multi-task audio separation, as a technology that can extract the three tracks (speech, music, noise) of a mixed audio at one time, is applicable to more application scenarios compared with single-task audio separation technology that can only extract two tracks.
[0003] The goal of multi-task audio separation is to extract the three tracks of a mixed audio at one time through a single model. Speech, music, and noise have different acoustic characteristics. Speech has short-term stationarity; music has periodicity, rich harmonic structures, and high-frequency components; noise is relatively random and has no obvious structural characteristics. Therefore, the model not only needs to have the ability to extract global context dependencies to capture long-term features such as short-term stationarity and periodicity, but also needs to be able to extract local context dependencies and similarities between time-frequency units to capture features such as harmonic structures and time-frequency distributions, so as to better distinguish speech, music, and noise. However, existing multi-task audio separation methods have deficiencies in extracting global dependencies, local dependencies, and similarities between time-frequency units, and the separation performance needs to be improved. Summary of the Invention
[0004] The present disclosure provides an audio separation method, apparatus, electronic device, and computer-readable storage medium to at least solve the problem in related technologies of how to improve multi-task audio separation performance, or may not solve any of the above problems.
[0005] According to a first aspect of the present disclosure, an audio separation method is provided. The audio separation method includes: for the audio to be separated, based on the coarse separation network of the audio separation model, obtaining the mixed audio complex spectrum of the audio to be separated and the coarse separation audio complex spectra of at least two tracks; for the coarse separation audio complex spectra of the at least two tracks and the mixed audio complex spectrum, based on the residual compensation network of the audio separation model, obtaining the complex spectrum residuals of the at least two tracks; for each track, determining the audio complex spectrum according to the coarse separation audio complex spectrum and the complex spectrum residuals; converting the audio complex spectra of the at least two tracks into audio signals respectively; wherein, both the coarse separation network and the residual compensation network include a two-dimensional window self-attention network, the two-dimensional window self-attention network includes a serial multi-head self-attention layer and a two-dimensional window self-attention layer, the multi-head self-attention layer and the two-dimensional window self-attention layer are respectively used to extract first intermediate features and second intermediate features, both the first intermediate features and the second intermediate features are two-dimensional matrices composed of multiple time-frequency units, and the two-dimensional window self-attention layer is used to extract three-dimensional features from the first intermediate features and extract the second intermediate features from the three-dimensional features.
[0006] Optionally, the two-dimensional window self-attention layer is used to perform the following steps on the first intermediate features: for each time-frequency unit in the first intermediate features, extracting a set number of time-frequency units that are spaced apart from the current time-frequency unit by a set scale in the first intermediate features as the feature vector of the current time-frequency unit, and combining the feature vectors of each time-frequency unit in the first intermediate features to form the three-dimensional features; performing windowing processing on the three-dimensional features to obtain a plurality of three-dimensional feature blocks; for each of the three-dimensional feature blocks, extracting two-dimensional features based on the multi-head self-attention mechanism, and combining the corresponding two-dimensional features according to the positions of the three-dimensional feature blocks in the three-dimensional features to obtain the second intermediate features.
[0007] Optionally, the coarse separation network includes a plurality of serial two-dimensional window self-attention networks.
[0008] Optionally, the residual compensation network includes an encoding network, the two-dimensional window self-attention network, and a decoding network.
[0009] Optionally, for the audio to be separated, based on the coarse separation network of the audio separation model, obtaining the coarse separation audio complex spectra of at least two tracks and the mixed audio complex spectrum of the audio to be separated includes: separating the mixed audio amplitude spectrum and the mixed audio complex spectrum from the audio to be separated, and the mixed audio complex spectrum includes the mixed audio phase; inputting the mixed audio amplitude spectrum into the coarse separation network to obtain the coarse separation audio amplitude spectra of the at least two tracks; for each track, determining the coarse separation audio complex spectrum according to the coarse separation audio amplitude spectrum and the mixed audio phase.
[0010] Optionally, for the rough-separated audio complex spectra and the mixed audio complex spectra of the at least two tracks, based on the residual compensation network of the audio separation model, obtaining the complex spectrum residuals of the at least two tracks includes: for each track, determining the difference between the mixed audio complex spectrum and the rough-separated audio complex spectrum as the initial residual; inputting the initial residuals of the at least two tracks and the mixed audio complex spectrum into the residual compensation network to obtain the complex spectrum residuals of the at least two tracks.
[0011] Optionally, the audio separation model is trained through the following steps: obtaining sample audio, where the sample audio is composed of the superimposed pure audio signals of at least two tracks, and obtaining the pure audio complex spectra of each pure audio signal; for the sample audio, based on the audio separation model to be trained, obtaining the estimated audio complex spectra and estimated audio signals of the at least two tracks of the sample audio; determining a loss value according to the pure audio complex spectra, the estimated audio complex spectra, the pure audio signals, and the estimated audio signals of the at least two tracks; adjusting the parameters of the audio separation model to be trained based on the loss value to obtain the audio separation model.
[0012] Optionally, determining the loss value according to the pure audio complex spectra, the estimated audio complex spectra, the pure audio signals, and the estimated audio signals of the at least two tracks includes: determining a first loss value according to the pure audio complex spectrum and the estimated audio complex spectrum of each track; determining a second loss value according to the pure audio complex spectrum of each track and the estimated audio complex spectra of the other tracks except the current track among the at least two tracks; determining a third loss value according to the pure audio signal and the estimated audio signal of each track; determining a total loss value according to the first loss value, the second loss value, and the third loss value, where the total loss value is positively correlated with the first loss value and negatively correlated with the second loss value and the third loss value.
[0013] According to a second aspect of the present disclosure, there is provided an audio separation device, the audio separation device comprising: a rough separation unit configured to perform a rough separation of the audio to be separated based on a rough separation network of an audio separation model to obtain a mixed audio complex spectrum of the audio to be separated and rough separation audio complex spectra of at least two tracks; a residual unit configured to perform, on the rough separation audio complex spectra of the at least two tracks and the mixed audio complex spectrum, based on a residual compensation network of the audio separation model, to obtain complex spectrum residuals of the at least two tracks; a compensation unit configured to perform, for each track, to determine an audio complex spectrum according to the rough separation audio complex spectrum and the complex spectrum residuals; a conversion unit configured to perform to convert the audio complex spectra of the at least two tracks into audio signals respectively; wherein, both the rough separation network and the residual compensation network include a two-dimensional window self-attention network, the two-dimensional window self-attention network includes a serial multi-head self-attention layer and a two-dimensional window self-attention layer, the multi-head self-attention layer and the two-dimensional window self-attention layer are respectively used to extract a first intermediate feature and a second intermediate feature, both the first intermediate feature and the second intermediate feature are two-dimensional matrices composed of a plurality of time-frequency units, and the two-dimensional window self-attention layer is used to extract three-dimensional features from the first intermediate feature and extract the second intermediate feature from the three-dimensional features.
[0014] Optionally, the two-dimensional window self-attention layer is configured to perform the following steps on the first intermediate feature: for each time-frequency unit in the first intermediate feature, extract a set number of time-frequency units in the first intermediate feature that are spaced apart from the current time-frequency unit by a set scale as a feature vector of the current time-frequency unit, and combine the feature vectors of each time-frequency unit in the first intermediate feature to form the three-dimensional feature; perform windowing processing on the three-dimensional feature to obtain a plurality of three-dimensional feature blocks; for each of the three-dimensional feature blocks, extract two-dimensional features based on a multi-head self-attention mechanism, and combine the corresponding two-dimensional features according to the positions of the three-dimensional feature blocks in the three-dimensional feature to obtain the second intermediate feature.
[0015] Optionally, the rough separation network includes a plurality of serial two-dimensional window self-attention networks.
[0016] Optionally, the residual compensation network includes an encoding network, the two-dimensional window self-attention network, and a decoding network.
[0017] Optionally, the rough separation unit is further configured to perform to separate a mixed audio amplitude spectrum and the mixed audio complex spectrum from the audio to be separated, the mixed audio complex spectrum includes a mixed audio phase; input the mixed audio amplitude spectrum into the rough separation network to obtain rough separation audio amplitude spectra of the at least two tracks; for each track, determine the rough separation audio complex spectrum according to the rough separation audio amplitude spectrum and the mixed audio phase.
[0018] Optionally, the residual unit is further configured to, for each track, determine the difference between the complex spectrum of the mixed audio and the complex spectrum of the roughly separated audio as an initial residual; input the initial residuals of the at least two tracks and the complex spectrum of the mixed audio into the residual compensation network to obtain the complex spectrum residuals of the at least two tracks.
[0019] Optionally, the audio separation model is trained through the following steps: obtaining a sample audio, where the sample audio is composed of the superimposed pure audio signals of at least two tracks, and obtaining the complex spectrum of the pure audio for each pure audio signal; for the sample audio, based on the audio separation model to be trained, obtaining the estimated complex spectrum of the audio of the at least two tracks and the estimated audio signal of the sample audio; determining a loss value according to the complex spectrum of the pure audio of the at least two tracks, the estimated complex spectrum, the pure audio signal, and the estimated audio signal; based on the loss value, adjusting the parameters of the audio separation model to be trained to obtain the audio separation model.
[0020] Optionally, the determining the loss value according to the complex spectrum of the pure audio of the at least two tracks, the estimated complex spectrum, the pure audio signal, and the estimated audio signal includes: determining a first loss value according to the complex spectrum of the pure audio and the estimated complex spectrum of each track; determining a second loss value according to the complex spectrum of the pure audio of each track and the estimated complex spectrum of the other tracks except the current track in the at least two tracks; determining a third loss value according to the pure audio signal and the estimated audio signal of each track; determining a total loss value according to the first loss value, the second loss value, and the third loss value, where the total loss value is positively correlated with the first loss value and negatively correlated with the second loss value and the third loss value.
[0021] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; at least one memory storing computer-executable instructions, where when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the audio separation method according to the present disclosure.
[0022] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium, where when the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the audio separation method according to the present disclosure.
[0023] According to a fifth aspect of the present disclosure, there is provided a computer program product including computer instructions which, when executed by at least one processor, implement the audio separation method according to the present disclosure.
[0024] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0025] For the audio separation method and the audio separation device according to the embodiments of the present disclosure, the most basic network structure applied is the two-dimensional window self-attention network. By using the multi-head self-attention layer, this network can capture global context dependencies. By establishing a two-dimensional window self-attention layer, first, the two-dimensional first intermediate feature composed of multiple time-frequency units is converted into a three-dimensional feature, that is, by extracting more detailed information for each time-frequency unit to increase the dimension of the feature, it can be ensured that the obtained three-dimensional feature contains richer local information compared with the two-dimensional first intermediate feature, which helps to realize the similarity analysis between time-frequency units. At the same time, as the name implies, the two-dimensional window self-attention layer can provide a two-dimensional window to window the obtained three-dimensional feature and can perform feature extraction based on the self-attention mechanism, so it can comprehensively capture the information required for multi-task audio separation, which helps to improve the separation performance. In addition, the exemplary embodiments of the present disclosure also adopt a dual-strategy framework of first rough separation and then fine-tuning. After the initial audio separation is completed, this framework combines the data before and after rough separation, estimates the residual of the rough separation result, and adjusts the rough separation result according to the residual to obtain the final separation result, which helps to improve the audio separation performance. By using the two-dimensional window self-attention network in both rough separation and fine-tuning, the rough separation ability and the residual estimation ability during fine-tuning can be further improved, thereby further improving the performance of audio separation.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.
[0028] Figure 1 is a schematic structural diagram showing a two-dimensional window self-attention network according to an exemplary embodiment of the present disclosure;
[0029] Figure 2 is a schematic structural diagram showing a two-dimensional window self-attention layer according to an exemplary embodiment of the present disclosure;
[0030] Figure 3 is a schematic structural diagram showing an audio separation model according to an exemplary embodiment of the present disclosure;
[0031] Figure 4 is a flowchart showing an audio separation method according to an exemplary embodiment of the present disclosure;
[0032] Figure 5 is a schematic structural diagram showing a coarse separation network according to an exemplary embodiment of the present disclosure;
[0033] Figure 6 is a schematic structural diagram showing a residual compensation network according to an exemplary embodiment of the present disclosure;
[0034] Figure 7 is a block diagram showing an audio separation device according to an exemplary embodiment of the present disclosure;
[0035] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0036] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0038] It should be noted that "at least one of several items" in the present disclosure all represents the inclusion of the following three parallel situations: "any one of the several items", "any combination of multiple items of the several items", and "the whole of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0039] It should be noted that the user information involved in the present disclosure (including but not limited to user device information, user personal information, etc.) is all information authorized by the user or fully authorized by all parties.
[0040] In real life, audio signals mainly include speech, music, and background noise, etc. When these signals are mixed, the intelligibility of the audio signal decreases, which will damage subsequent audio understanding tasks. For example, in the music information retrieval task, foreground speech and background noise will reduce the retrieval accuracy; in the speech recognition task, background noise will reduce the recognition accuracy, and at the same time, the music in the audio will cause the recognition results to be chaotic. At the same time, in the speech quality assessment task, the energy level of the extracted background noise is an important assessment basis. It can be seen that multi-task audio separation, as a technology that can extract the three tracks (speech, music, noise) of the mixed audio at one time, is applicable to more application scenarios compared with the single-task audio separation technology that can only extract two tracks.
[0041] In recent years, research work has mainly focused on single-task audio separation, such as speech enhancement and separation, music separation, and singing voice separation. These separation algorithms can be divided into frequency-domain based separation algorithms, time-domain based separation algorithms, and complex-domain based separation algorithms. The frequency-domain based separation algorithm transforms the separation problem into a supervised classification problem. First, using the masking effect of the audio, a mask value is obtained as a label; then, a deep learning model is used to learn a mapping function from the mixed speech to the label; finally, the label information is used to extract the frequency-domain units where the target speech is located. The time-domain based separation algorithm adopts an end-to-end processing mode, inputting the mixed audio and outputting the estimated clean audio. The complex-domain based separation algorithm is to deal with the phase mismatch problem in the frequency-domain separation algorithm. The separation effects of various methods such as DCCRN (Deep Complex Convolution Recurrent Network), TSCN (Two-Stage Complex Network), and SDDNet (Simultaneous Denoising and Dereverberation Network) show that the complex-domain based models are more suitable for enhancing speech from background noise.
[0042] Since multi-task audio separation was proposed, it has received extensive attention from researchers. Currently, it mainly includes separation models such as Complex-MTASSNet (Multi-Task Audio Source Separation base on Complex Network), MRX (Multi-resolution Cross Network), and EAD-Conformer (Conformer-base Encoder Attention Decoder Network). Among them, the Complex-MTASSNet and MRX models verified the feasibility of multi-task audio separation and outperformed mainstream single-task audio separation models such as GCRN (Gated Convolutional Recurrent Network), Conv-TasNet (fully-Convolutional Time-domain Audio Separation Network), D3Net (Densely Connected Dilated DenseNet), and Demucs (Music Source Separation Network) in terms of performance. EAD-Conformer, on the other hand, achieved separation by borrowing Conformer (Convolution-augmented Transformer), which has achieved SOTA (State-Of-The-Art) in speech recognition tasks, to extract audio features, significantly improving the separation performance. At the same time, Swin-Transformer (Shifted Windows Transformer) has good global and local feature capture capabilities by introducing strategies such as window partitioning and has achieved SOTA results in audio event detection and classification tasks, indicating that the Swin-Transformer structure also has good performance in audio feature extraction.
[0043] The goal of multi-task audio separation is to extract the three tracks of the mixed audio through a single model. Speech, music, and noise have different acoustic characteristics. Speech has short-term stationarity; music has periodicity, rich harmonic structures, and high-frequency components; noise is relatively random and has no obvious structural characteristics. Therefore, the model not only needs to have the ability to extract global context dependencies to capture long-term features such as short-term stationarity and periodicity, but also be able to extract local context dependencies and the similarity between time-frequency units to capture features such as harmonic structures and time-frequency distributions, so as to better distinguish speech, music, and noise. Existing methods have deficiencies in extracting global dependencies, local dependencies, and the similarity between time-frequency units: Complex-MTASSNet uses a stacked convolutional structure to extract separation features, resulting in the global correlation features it captures being limited by the receptive field of the convolutional network; MRX uses a bidirectional long short-term memory network to capture global dependencies. This scheme not only has high complexity but also lacks features of local dependencies and the similarity between time-frequency units; EAD-Conformer can capture global and local dependencies, but lacks the ability to capture the similarity between time-frequency units and has a relatively high complexity.
[0044] According to the audio separation method and apparatus of the exemplary embodiments of the present disclosure, the most basic network structure applied is a two-dimensional window self-attention network. This network can capture global context dependencies by using a multi-head self-attention layer. By establishing a two-dimensional window self-attention layer, the two-dimensional first intermediate feature composed of multiple time-frequency units is first converted into a three-dimensional feature, that is, by extracting more detailed information for each time-frequency unit to increase the dimension of the feature, the resulting three-dimensional feature can contain richer local information compared to the two-dimensional first intermediate feature, which helps to realize the similarity analysis between time-frequency units. At the same time, as the name implies, the two-dimensional window self-attention layer can provide a two-dimensional window to window the obtained three-dimensional feature and can perform feature extraction based on the self-attention mechanism, thus being able to capture local context dependencies. Therefore, the two-dimensional window self-attention network proposed by the exemplary embodiments of the present disclosure can comprehensively capture the information required for multi-task audio separation, which helps to improve the separation performance. In addition, the exemplary embodiments of the present disclosure also adopt a dual-strategy framework of coarse separation first and then fine-tuning. After completing the preliminary audio separation, this framework combines the data before and after the coarse separation, estimates the residual of the coarse separation result, and adjusts the coarse separation result according to the residual to obtain the final separation result, which helps to improve the audio separation performance. By using the two-dimensional window self-attention network in both the coarse separation and fine-tuning, the coarse separation ability and the residual estimation ability during fine-tuning can be further improved, thereby further improving the performance of audio separation.
[0045] Next, reference will be made to Figures 1 to 8 Specifically describe the audio separation method and audio separation apparatus according to the exemplary embodiments of the present disclosure.
[0046] First, the two-dimensional window self-attention network is introduced.
[0047] Figure 1 It is a schematic structural diagram showing a two-dimensional window self-attention network according to an exemplary embodiment of the present disclosure. Refer to Figure 1 , the two-dimensional window self-attention network (WA-Transformer Block, Window Attention-based Transformer Block) includes a serial multi-head self-attention layer (MSA, Mulit-head Self-Attention) and a two-dimensional window self-attention layer (2D-WA). The multi-head self-attention layer and the two-dimensional window self-attention layer are respectively used to extract first intermediate features and second intermediate features. Both the first intermediate features and the second intermediate features are two-dimensional matrices composed of multiple time-frequency units. Each time-frequency unit represents the features of a pixel point on the spectrogram. The spectrogram is a spectrum that can simultaneously display time-domain and frequency-domain information. The two-dimensional matrix shows the information of the time-domain dimension and the frequency-domain dimension. The two-dimensional window self-attention layer is used to extract three-dimensional features from the first intermediate features and extract second intermediate features from the three-dimensional features. Optionally, refer to Figure 1 , before the multi-head self-attention layer and after the two-dimensional window self-attention layer, a feed-forward neural network (FFN, Feed Forward Neural Network) is connected, and residual connection and normalization processing (Add&Norm, Add and Normalize) are also required after each layer, that is, calculate the sum of the input features and output features of the corresponding layer and perform normalization as the input of the next layer.
[0048] As an example, the processing of the two-dimensional window self-attention network can be expressed by the formula:
[0049]
[0050] z″ = 2DWA(z′); output = LN(z″ + 0.5 × FFN(z″)).
[0051] Among them, FFN(), MSA(), 2DWA(), LN() and BN() respectively represent the feed-forward neural network, multi-head attention layer, two-dimensional window self-attention layer, layer normalization and batch normalization. For the residual connection, in the above formula, the input and output feature weights of the feed-forward neural network are 1 and 0.5 respectively, the input and output feature weights of the multi-head self-attention layer are both 1, and the input and output feature weights of the two-dimensional window self-attention layer are 0 and 1 respectively.
[0052] Figure 2It is a schematic structural diagram showing a two-dimensional window self-attention layer according to an exemplary embodiment of the present disclosure. Refer to Figure 2 , the two-dimensional window self-attention layer is used to perform the following steps on the first intermediate feature:
[0053] First, for each time-frequency unit x(t,f) in the first intermediate feature, using a scale time-frequency unit extractor (TFBand, Embeds Time Frequency Mulit-Band), a set number of time-frequency units with a set scale interval from this time-frequency unit in the first intermediate feature are extracted as the feature vector of this time-frequency feature, which can be expressed by the formula:
[0054] x(t,f) = [..., x(t + i×d t , f + j×d f ),...];
[0055] i = 0, 1, 2, 3,..., k t -1; j = 0, 1, 2, 3,..., k f -1.
[0056] Among them, t represents the time domain coordinate, k t represents the number extracted in the time domain, d t represents the extraction scale in the time domain, that is, 1 time-frequency unit is extracted from every d t time-frequency units in the time-frequency domain. d t = 1 means continuous extraction without interval, that is, time-frequency units with the same scale as the input intermediate feature are extracted. All three are integers, and k t and d t are both greater than 0. Optionally, in order to extract rich information, time-frequency units with different scales from the input intermediate feature can be extracted, that is, d t ≥2. f represents the frequency domain coordinate, k f represents the number extracted in the frequency domain, d f represents the extraction scale in the frequency domain, which is the same as in the time domain. A total of K = k t ×k f time-frequency units can be extracted.
[0057] As an example, refer to Figure 2 , for the first time-frequency unit, k t = k f = 3, d t = d f = 2, and 9 time-frequency units with an interval of 1 in the time domain and frequency domain are extracted starting from this time-frequency unit.
[0058] The feature vectors of each time-frequency unit in the first intermediate feature are pooled together to form a three-dimensional feature, that is, a single x ∈ RT×F The two-dimensional matrix is processed into a three-dimensional tensor. As an example, referring to Figure 2 , when T = F = 9, the scale time-frequency unit extractor processes the two-dimensional square array into a three-dimensional cuboid array.
[0059] Then, a window partitioner (WP) is used to perform window partitioning on the three-dimensional features, obtaining multiple three-dimensional feature blocks. Referring to Figure 2 , the three-dimensional features are evenly divided into 9 three-dimensional feature blocks of 3×3.
[0060] Finally, for each three-dimensional feature block, based on the multi-head self-attention mechanism, two-dimensional features are extracted, and according to the positions of the respective three-dimensional feature blocks in the three-dimensional features, the corresponding two-dimensional features are combined to obtain the second intermediate feature. Specifically, referring to Figure 2 , taking the third three-dimensional feature block in the second row as an example, a stretching operation (strech) can be first performed on this three-dimensional feature block to convert the 3×3 three-dimensional feature block into a 9×1 three-dimensional feature block, and then a window-based multi-head self-attention layer (W-MSA) is used for calculation to obtain another 9×1 three-dimensional feature block. This calculation belongs to the mature technology in this field and will not be elaborated here. It should be understood that the difference between the window-based multi-head self-attention layer here and the multi-head self-attention layer in the two-dimensional window self-attention network lies in the different dimensions of the input features, and their ways of processing features are similar. Figure 2 The three subsequent 1×9 three-dimensional feature blocks after the 9×1 three-dimensional feature block in represent Q (query), K (key), and V (value) in the multi-head self-attention mechanism respectively. After that, through the processing of the fully connected layer, a 9×1 two-dimensional feature is obtained and reconverted into a 3×3 two-dimensional feature. When combining the respective two-dimensional features, it is placed in the position of the third column in the second row, and finally a 9×9 two-dimensional feature is obtained as the output intermediate feature.
[0061] Generally speaking, the processing of the two-dimensional window self-attention layer can be expressed by the formula:
[0062]
[0063]
[0064] Among them, LN() represents normalization processing using the method of Layer Normalization.
[0065] Next, the dual-strategy framework, that is, the execution process of the entire audio separation model, is introduced.
[0066] Figure 3 It is a schematic structural diagram showing an audio separation model according to an exemplary embodiment of the present disclosure. Refer to Figure 3 , the execution of the audio separation model mainly includes a rough separation stage (Separation stage) and a residual compensation stage (Residual compensation stage). The former mainly uses a rough separation network (Separator), and the latter mainly uses a residual compensation network (Residual compensation). Both of these networks include the aforementioned two-dimensional window self-attention network.
[0067] Figure 4 It is a flowchart showing an audio separation method according to an exemplary embodiment of the present disclosure. It should be understood that the audio separation method according to an exemplary embodiment of the present disclosure can be implemented in a terminal device such as a smart phone, a tablet computer, or a personal computer (PC), or can also be implemented in a device such as a server.
[0068] Refer to Figure 4 , in step 401, for the audio to be separated, based on the rough separation network of the audio separation model, the mixed audio complex spectrum of the audio to be separated and the rough separation audio complex spectra of at least two tracks are obtained. That is, not only the rough separation audio complex spectra of each track need to be obtained, but also the mixed audio complex spectrum of the entire audio needs to be obtained. It should be understood that the exemplary embodiment of the present disclosure has the ability to separate speech, music, and background noise, but is limited by the actual content of the audio to be separated, and may only separate the data of two tracks, so it is expressed as "at least two" tracks here.
[0069] Figure 5 It is a schematic structural diagram showing a rough separation network according to an exemplary embodiment of the present disclosure. Refer to Figure 5 , optionally, the rough separation network includes a plurality of serially connected two-dimensional window self-attention networks, with a simple structure, and can obtain more accurate rough separation audio complex spectra by repeatedly extracting intermediate features, which helps to balance the structural complexity and the audio separation effect. As an example, holes can be injected into each two-dimensional window self-attention network to increase the receptive field, and the dilation rate is, for example, 2 n -1 , which means injecting 2 n-1 -1 holes between two adjacent convolutional kernels, where n represents the serial number of each two-dimensional window self-attention network in the rough separation network.
[0070] Optionally, refer to Figure 3 , step 401 includes the following steps:
[0071] First, the mixed audio amplitude spectrum Y mag-mix and the mixed audio complex spectrum Y RI-mix, which can be specifically achieved through the Short-Time Fourier Transform (STFT), and the complex spectrum of the mixed audio contains the phase Y of the mixed audio phase-mix = arctan(Y I-mix / Y R-mix ), where Y R-mix and Y I-mix respectively represent the real part and the imaginary part of the complex spectrum of the mixed audio.
[0072] Then, input the magnitude spectrum Y of the mixed audio mag-mix into the rough classification network to obtain the rough classification audio magnitude spectra of at least two tracks. Specifically, the rough classification network needs to repeatedly extract intermediate features first, and then convert the finally obtained intermediate features into the rough classification audio magnitude spectra of at least two tracks. As an example, the rough classification audio magnitude spectrum of the speech track can be determined the rough classification audio magnitude spectrum of the music track and the rough classification audio magnitude spectrum of the background noise track By inputting the magnitude spectrum of the mixed audio, which is a real spectrum, into the rough classification network, first perform the separation of the magnitude spectrum, and then convert it into a complex spectrum, rather than directly separating the complex spectrum, the calculation error can be reduced, which helps to improve the rough classification accuracy.
[0073] Finally, for each track, based on the rough classification audio magnitude spectrum and the phase Y of the mixed audio phase-mix , determine the rough classification audio complex spectrum. As an example, the rough classification audio complex spectrum of the speech track can be determined the rough classification audio complex spectrum of the music track and the rough classification audio complex spectrum of the background noise track Specifically, taking as an example, its real part imaginary part The same applies to other tracks.
[0074] Return to reference Figure 4 , in step 402, for the rough classification audio complex spectra and the mixed audio complex spectrum of at least two tracks, based on the residual compensation network of the audio separation model, obtain the complex spectrum residuals of at least two tracks. As an example, the complex spectrum residual Res of the speech track can be obtained speech , the complex spectrum residual Res of the music track music and the complex spectrum residual Res of the background noise track noiseBy combining the complex spectrum data before and after rough separation and a residual compensation network containing a two-dimensional window self-attention network, the residual of the rough separation result can be estimated more accurately, which is used as compensation for the rough separation result, so that the subsequent steps can adjust the rough separation result according to the residual to obtain the final separation result, helping to improve the audio separation performance.
[0075] Figure 6 is a schematic structural diagram showing a residual compensation network according to an exemplary embodiment of the present disclosure. Refer to Figure 6 , optionally, the residual compensation network includes an encoding network (Encoder), a two-dimensional window self-attention network (Stacked WA-Transformer blocks, representing the use of multiple serial two-dimensional window self-attention networks), and a decoding network (Decoder). Among them, reshape is a method for processing feature tensors, used to reshape the shape parameter of the feature tensor. The shape parameter represents the length of the feature tensor in each dimension, that is, the total number of elements in each dimension. As an example, the residual compensation network is a network obtained by replacing the long short-term memory network (LSTM, Long Short-Term Memory) in the gated convolutional recurrent network (GCRN) with a two-dimensional window self-attention network. The encoding network includes multiple serial convolutional gated linear units (ConvGLU, Convolutional Gated Linear Unit). The decoding network can be divided into two branches, respectively used to process the real part and the imaginary part. Each decoding network includes multiple serial deconvolutional gated linear units (DeconvGLU, Deconvolutional Gated Linear Unit). This single-encoding and double-decoding structure can improve the estimation accuracy of the real part and the imaginary part. The convolutional unit with a gating mechanism can further improve the network's modeling ability. Combined with the two-dimensional window self-attention network, it can ensure the relatively accurate estimation of the complex spectrum residual.
[0076] Optionally, step 402 includes: for each track, determining the difference between the complex spectrum of the mixed audio and the complex spectrum of the roughly separated audio as the initial residual; inputting the initial residuals of at least two tracks and the complex spectrum of the mixed audio into the residual compensation network to obtain the complex spectrum residuals of at least two tracks. Specifically, similar to the rough separation network, the residual compensation network needs to repeatedly extract intermediate features first, and then convert the finally obtained intermediate features into the complex spectrum residuals of at least two tracks. As an example, the initial residual can be expressed by the formula:
[0077]
[0078] That is, the difference between the complex spectrum of the mixed audio and the complex spectrum of the roughly separated audio for each track is used as the initial residual for the corresponding track.
[0079] Accordingly, the input of the residual compensation network can be expressed by the formula:
[0080] InRes 1,2,3 =[InRes speech ,InRes music ,InRes noise ,Y RI-mix .
[0081] Accordingly, the complex spectral residual of the output can be expressed by the formula:
[0082] OutRes R-1,2,3 ,OutRes I-1,2,3 =RCN(InRes 1,2,3 ).
[0083] Wherein, outRes R-1,2,3 includes outRes R-speech , outRes R-music , outRes R-noise , which respectively represent the real parts of the complex spectral residuals of speech, music, and background noise; outRes I-1,2,3 includes outRes I-speech , outRes I-music , outRes I-noise , which respectively represent the imaginary parts of the complex spectral residuals of speech, music, and background noise, and RCN() represents the residual compensation network.
[0084] Returning to reference Figure 4 , in step 403, for each track, the audio complex spectrum is determined according to the roughly segmented audio complex spectrum and the complex spectral residual. Referring to Figure 3 , as an example, the two can be summed as the audio complex spectrum of each track, which can be expressed by the formula:
[0085]
[0086] In step 404, the audio complex spectra of at least two tracks are respectively converted into audio signals. As an example, it can be achieved through inverse Fourier transform.
[0087] Optionally, the audio separation model is trained through the following steps: obtaining sample audio, where the sample audio is composed of the superimposition of pure audio signals of at least two tracks, and obtaining the pure audio complex spectrum of each pure audio signal; for the sample audio, based on the audio separation model to be trained, obtaining the estimated audio complex spectra and estimated audio signals of at least two tracks of the sample audio; determining a loss value according to the pure audio complex spectra of at least two tracks, the estimated audio complex spectra, the pure audio signals, and the estimated audio signals; adjusting the parameters of the audio separation model to be trained based on the loss value to obtain the audio separation model. The audio separation model can be trained in a supervised manner, ensuring the reliability of the training results. By combining the complex spectrum and the audio signal to determine the loss value, the separation effect of the audio separation model to be trained can be comprehensively evaluated, thereby improving the training efficiency.
[0088] Optionally, determining the loss value according to the pure audio complex spectra of at least two tracks, the estimated audio complex spectra, the pure audio signals, and the estimated audio signals includes: determining a first loss value according to the pure audio complex spectrum and the estimated audio complex spectrum of each track; determining a second loss value according to the pure audio complex spectrum of each track and the estimated audio complex spectra of other tracks except the current track among at least two tracks; determining a third loss value according to the pure audio signal and the estimated audio signal of each track; determining a total loss value according to the first loss value, the second loss value, and the third loss value, where the total loss value is positively correlated with the first loss value and negatively correlated with the second loss value and the third loss value. The first loss value and the third loss value can directly refer to the pure audio data to evaluate the separation effect of the audio separation model to be trained from the perspectives of the complex spectrum and the audio signal. In addition, the second loss value also compares the pure audio complex spectrum of the current track with the estimated audio complex spectra of other tracks, which can enhance the discrimination ability of the model and help further improve the training effect.
[0089] As an example, the first loss value is the mean absolute error, which can be expressed by the formula:
[0090]
[0091] The second loss value is the discrimination loss, which can be expressed by the formula:
[0092]
[0093] The third loss value is the signal-to-noise ratio (SNR), which can be expressed by the formula:
[0094]
[0095] The loss value can be expressed by the formula:
[0096]
[0097] Among them, the number of extraction tracks of N is, for example, 3. S and S' respectively represent the pure audio complex spectrum and the predicted audio complex spectrum features. s and s' respectively represent the pure audio signal and the predicted audio signal. λ and α are hyperparameters used to balance each loss term in the loss value so that these loss terms are on the same numerical scale. T represents the maximum number of training times, and t represents the iteration number of the current training.
[0098] To verify the performance of the audio separation model of the exemplary embodiments of the present disclosure, the number of operations per second (MAC / s, where MAC is Memory Access Cost, video memory / memory access volume) of the standard system on a GPU (Nvidia 2080Ti), the time consumption required to process audio per second (real-time rate), and the size of the model parameters (occupied space volume) were statistically analyzed. Under the condition of 8.62M parameters, the audio separation model of the exemplary embodiments of the present disclosure not only has better performance than other models, but also has the best real-time rate and small computational amount. Under the condition that the number of parameters is comparable to that of EAD-Conformer, the performance is significantly better than other models. Generally speaking, the present disclosure achieves the best signal-to-noise ratio improvement on all three audio tracks, and the signal-to-noise ratio improvements on the speech, music, and noise tracks are 13.86dB, 12.22dB, and 11.21dB respectively. This shows the effectiveness and advancement of the present disclosure.
[0099] Figure 7 It is a block diagram showing an audio separation device according to an exemplary embodiment of the present disclosure. It should be understood that the audio separation device according to the exemplary embodiment of the present disclosure can be implemented in a software, hardware, or software-hardware combination manner in terminal devices such as smart phones, tablet computers, and personal computers (PCs), or can also be implemented in devices such as servers.
[0100] Referring to Figure 7 , the audio separation device includes a rough separation unit 701, a residual unit 702, a compensation unit 703, and a conversion unit 704.
[0101] The rough separation unit 701 can obtain the mixed audio complex spectrum of the audio to be separated and the rough separation audio complex spectra of at least two tracks based on the rough separation network of the audio separation model for the audio to be separated.
[0102] Optionally, the rough separation network includes a plurality of serially connected two-dimensional window self-attention networks.
[0103] Optionally, the coarse separation unit 701 can also separate a mixed audio amplitude spectrum and a mixed audio complex spectrum from the audio to be separated, wherein the mixed audio complex spectrum includes a mixed audio phase; the mixed audio amplitude spectrum is input into the coarse separation network to obtain a coarse audio amplitude spectrum for determining at least two tracks; for each track, the coarse audio complex spectrum is determined according to the coarse audio amplitude spectrum and the mixed audio phase.
[0104] The residual unit 702 can obtain complex spectrum residuals of at least two tracks based on the residual compensation network of the audio separation model for the coarse audio complex spectrum and the mixed audio complex spectrum of at least two tracks.
[0105] Optionally, the residual compensation network includes an encoding network, a two-dimensional window self-attention network, and a decoding network.
[0106] Optionally, the residual unit 702 can also determine the difference between the mixed audio complex spectrum and the coarse audio complex spectrum for each track as the initial residual; input the initial residual and the mixed audio complex spectrum of at least two tracks into the residual compensation network to obtain the complex spectrum residuals of at least two tracks.
[0107] The compensation unit 703 may determine the audio complex spectrum for each track based on the roughly divided audio complex spectrum and the complex spectrum residual.
[0108] The conversion unit 704 may convert the audio complex spectra of at least two tracks into audio signals respectively.
[0109] Among them, both the coarse-separation network and the residual compensation network contain a two-dimensional window self-attention network, which includes a serial multi-head self-attention layer and a two-dimensional window self-attention layer. The multi-head self-attention layer and the two-dimensional window self-attention layer are used to extract the first intermediate feature and the second intermediate feature, respectively. The first intermediate feature and the second intermediate feature are both two-dimensional matrices composed of multiple time-frequency units. The two-dimensional window self-attention layer is used to extract three-dimensional features from the first intermediate feature and to extract the second intermediate feature from the three-dimensional feature.
[0110] Optionally, the two-dimensional window self-attention layer is used to perform the following steps on the first intermediate feature: for each time-frequency unit in the first intermediate feature, extract a set number of time-frequency units in the first intermediate feature that are separated from the current time-frequency unit by a set scale as the feature vector of the current time-frequency unit, and merge the feature vectors of each time-frequency unit in the first intermediate feature to form a three-dimensional feature; perform window processing on the three-dimensional feature to obtain multiple three-dimensional feature blocks; for each three-dimensional feature block, extract two-dimensional features based on the multi-head self-attention mechanism, and combine the corresponding two-dimensional features according to the position of each three-dimensional feature block in the three-dimensional feature to obtain a second intermediate feature.
[0111] Optionally, the audio separation model is trained through the following steps: obtaining sample audio, where the sample audio is composed of the superimposition of pure audio signals of at least two tracks, and obtaining the pure audio complex spectrum of each pure audio signal; for the sample audio, based on the audio separation model to be trained, obtaining the estimated audio complex spectra and estimated audio signals of at least two tracks of the sample audio; determining a loss value according to the pure audio complex spectra of at least two tracks, the estimated audio complex spectra, the pure audio signals, and the estimated audio signals; and adjusting the parameters of the audio separation model to be trained based on the loss value to obtain the audio separation model.
[0112] Optionally, determining a loss value according to the pure audio complex spectra of at least two tracks, the estimated audio complex spectra, the pure audio signals, and the estimated audio signals includes: determining a first loss value according to the pure audio complex spectrum and the estimated audio complex spectrum of each track; determining a second loss value according to the pure audio complex spectrum of each track and the estimated audio complex spectra of other tracks except the current track among at least two tracks; determining a third loss value according to the pure audio signal and the estimated audio signal of each track; and determining a total loss value according to the first loss value, the second loss value, and the third loss value, where the total loss value is positively correlated with the first loss value and negatively correlated with the second loss value and the third loss value.
[0113] Regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0114] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.
[0115] Referring to Figure 8 , the electronic device 800 includes at least one memory 801 and at least one processor 802. A set of computer-executable instructions is stored in the at least one memory 801. When the set of computer-executable instructions is executed by the at least one processor 802, an audio separation method according to an exemplary embodiment of the present disclosure is executed.
[0116] As an example, the electronic device 800 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 800 does not have to be a single electronic device, and may also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 800 may also be a part of an integrated control system or a system manager, or may be configured to be interconnected with a local or remote device (e.g., via wireless transmission) as an interface.
[0117] In the electronic device 800, the processor 802 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.
[0118] The processor 802 may execute instructions or code stored in the memory 801, where the memory 801 may also store data. The instructions and data may also be sent and received over a network via a network interface device, where the network interface device may employ any known transmission protocol.
[0119] The memory 801 may be integrated with the processor 802, for example, by arranging RAM or flash memory within an integrated circuit microprocessor or the like. Additionally, the memory 801 may include a separate device, such as an external disk drive, a storage array, or other storage devices that may be used by any database system. The memory 801 and the processor 802 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 802 can read files stored in the memory.
[0120] Furthermore, the electronic device 800 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 800 may be connected to each other via a bus and / or a network.
[0121] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the audio separation method according to the exemplary embodiment of the present disclosure. Examples of the computer-readable storage medium herein include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, the any other device being configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0122] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided. The computer program product includes computer instructions. When the computer instructions are run by at least one processor, the at least one processor is caused to execute the audio separation method according to the exemplary embodiment of the present disclosure.
[0123] An audio separation method, apparatus, electronic device, and computer-readable storage medium according to an exemplary embodiment of the present disclosure apply the most basic network structure, which is a two-dimensional window self-attention network. By using a multi-head self-attention layer, this network can capture global context dependencies. By establishing a two-dimensional window self-attention layer, a two-dimensional first intermediate feature composed of multiple time-frequency units is first converted into a three-dimensional feature, that is, by extracting more detailed information for each time-frequency unit to increase the dimension of the feature, the obtained three-dimensional feature can contain richer local information compared to the two-dimensional first intermediate feature, which helps to realize the similarity analysis between time-frequency units. At the same time, as the name implies, the two-dimensional window self-attention layer can provide a two-dimensional window to window the obtained three-dimensional feature and can perform feature extraction based on the self-attention mechanism, so it can comprehensively capture the information required for multi-task audio separation and helps to improve the separation performance. In addition, the exemplary embodiment of the present disclosure also adopts a dual-strategy framework of first rough separation and then fine-tuning. After completing the preliminary audio separation, this framework combines the data before and after the rough separation, estimates the residual of the rough separation result, and adjusts the rough separation result according to the residual to obtain the final separation result, which helps to improve the audio separation performance. By using the two-dimensional window self-attention network in both rough separation and fine-tuning, the rough separation ability and the residual estimation ability during fine-tuning can be further improved, thereby further improving the performance of audio separation.
[0124] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0125] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio separation method, characterized in that, the audio separation method includes: For the audio to be separated, based on the coarse separation network of the audio separation model, obtaining the complex spectrum of the mixed audio of the audio to be separated and the complex spectra of the coarsely separated audio of at least two tracks; For the complex spectra of the coarsely separated audio of the at least two tracks and the complex spectrum of the mixed audio, based on the residual compensation network of the audio separation model, obtaining the complex spectrum residuals of the at least two tracks; For each track, determining the complex spectrum of the audio according to the complex spectrum of the coarsely separated audio and the complex spectrum residuals; Converting the complex spectra of the audio determined according to the complex spectrum of the coarsely separated audio and the complex spectrum residuals of the at least two tracks into audio signals respectively; wherein, both the coarse separation network and the residual compensation network include a two-dimensional window self-attention network, the two-dimensional window self-attention network includes a serial multi-head self-attention layer and a two-dimensional window self-attention layer, the multi-head self-attention layer is used to extract first intermediate features, the two-dimensional window self-attention layer is used to extract three-dimensional features from the first intermediate features, and extract second intermediate features from the three-dimensional features, and both the first intermediate features and the second intermediate features are two-dimensional matrices composed of a plurality of time-frequency units.
2. The audio separation method according to claim 1, characterized in that, the two-dimensional window self-attention layer is used to perform the following steps on the first intermediate features: For each time-frequency unit in the first intermediate features, extracting a set number of time-frequency units in the first intermediate features that are spaced apart from the current time-frequency unit by a set scale as the feature vector of the current time-frequency unit, and converging the feature vectors of each time-frequency unit in the first intermediate features to form the three-dimensional features; Performing window division processing on the three-dimensional features to obtain a plurality of three-dimensional feature blocks; For each of the three-dimensional feature blocks, based on the multi-head self-attention mechanism, extracting two-dimensional features, and combining the corresponding two-dimensional features according to the positions of the three-dimensional feature blocks in the three-dimensional features to obtain the second intermediate features.
3. The audio separation method according to claim 1, characterized in that, the step of obtaining the complex spectrum of the mixed audio of the audio to be separated and the complex spectra of the coarsely separated audio of at least two tracks based on the coarse separation network of the audio separation model includes: Separating the mixed audio amplitude spectrum and the complex spectrum of the mixed audio from the audio to be separated, and the complex spectrum of the mixed audio includes the phase of the mixed audio; Inputting the mixed audio amplitude spectrum into the coarse separation network to obtain the coarsely separated audio amplitude spectra of the at least two tracks; For each track, determining the complex spectrum of the coarsely separated audio according to the coarsely separated audio amplitude spectrum and the phase of the mixed audio.
4. The audio separation method according to claim 1, characterized in that, the step of obtaining the complex spectrum residuals of the at least two tracks based on the residual compensation network of the audio separation model for the complex spectra of the coarsely separated audio of the at least two tracks and the complex spectrum of the mixed audio includes: For each track, determining the difference between the complex spectrum of the mixed audio and the complex spectrum of the coarsely separated audio as the initial residual; Input the initial residuals of the at least two tracks and the complex spectrum of the mixed audio into the residual compensation network to obtain the complex spectrum residuals of the at least two tracks.
5. The audio separation method according to any one of claims 1 to 4, wherein, the audio separation model is trained through the following steps: Obtain a sample audio, wherein the sample audio is composed of the pure audio signals of at least two tracks superimposed, and obtain the pure audio complex spectrum of each pure audio signal; For the sample audio, based on the audio separation model to be trained, obtain the estimated audio complex spectrum and the estimated audio signal of the at least two tracks of the sample audio; Determine a loss value according to the pure audio complex spectrum, the estimated audio complex spectrum, the pure audio signal, and the estimated audio signal of the at least two tracks; Based on the loss value, adjust the parameters of the audio separation model to be trained to obtain the audio separation model.
6. The audio separation method according to claim 5, wherein, the determining the loss value according to the pure audio complex spectrum, the estimated audio complex spectrum, the pure audio signal, and the estimated audio signal of the at least two tracks includes: Determine a first loss value according to the pure audio complex spectrum and the estimated audio complex spectrum of each track; Determine a second loss value according to the pure audio complex spectrum of each track and the estimated audio complex spectrum of the other tracks except the current track among the at least two tracks; Determine a third loss value according to the pure audio signal and the estimated audio signal of each track; Determine a total loss value according to the first loss value, the second loss value, and the third loss value, wherein the total loss value is positively correlated with the first loss value and negatively correlated with the second loss value and the third loss value.
7. An audio separation device, wherein, the audio separation device includes: A rough separation unit configured to perform, for the audio to be separated, based on the rough separation network of the audio separation model, to obtain the complex spectrum of the mixed audio and the complex spectrum of the roughly separated audio of at least two tracks; A residual unit configured to perform, for the complex spectrum of the roughly separated audio and the complex spectrum of the mixed audio of the at least two tracks, based on the residual compensation network of the audio separation model, to obtain the complex spectrum residuals of the at least two tracks; A compensation unit configured to perform, for each track, to determine the audio complex spectrum according to the complex spectrum of the roughly separated audio and the complex spectrum residuals; A conversion unit configured to perform to convert the audio complex spectra determined according to the complex spectrum of the roughly separated audio and the complex spectrum residuals of the at least two tracks into audio signals respectively. Among them, a two-dimensional window self-attention network is included in both the coarse segmentation network and the residual compensation network. The two-dimensional window self-attention network includes a multi-head self-attention layer and a two-dimensional window self-attention layer in series. The multi-head self-attention layer is used to extract first intermediate features, and the two-dimensional window self-attention layer is used to extract three-dimensional features from the first intermediate features and extract second intermediate features from the three-dimensional features. Both the first intermediate features and the second intermediate features are two-dimensional matrices composed of a plurality of time-frequency units.
8. An electronic device, characterized in that it includes: at least one processor; at least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the audio separation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that when the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the audio separation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Audio separation method and device, electronic equipment and storage medium
CN114171051A
Systems and methods facilitating selective removal of content from a mixed audio recording
US9373320B1