Audio separation method, device, electronic device and storage medium

By extracting and fusion of the characteristics of audio in different dimensions, the problem of limited receptive field of audio separation methods in the prior art is solved, and a more accurate audio separation effect is achieved.

CN114171051BActive Publication Date: 2025-05-13BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111447488.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-05-13
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

The existing CNN-based audio separation method is difficult to capture the spectrum distribution rules of different instruments in chord music due to the limited receptive field, resulting in inaccurate separation results.

Method used

By obtaining the frequency domain amplitude spectrum of the audio to be separated, the characteristics of different frequency dimensions and frequency domain dimensions at the same time are extracted respectively, the frequency domain feature map and the time feature map are formed, and attention fusion processing is performed, and the amplitude spectrum of the vocal and background accompaniment is finally obtained through decoding processing.

Benefits of technology

By enlarging the receptive field, capturing the complex modes of audio in the frequency domain time and frequency dimensions, the accuracy and effect of audio separation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114171051B_ABST
    Figure CN114171051B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio separation method, device, electronic device and storage medium, the method comprising: obtaining a frequency domain amplitude spectrum corresponding to the audio to be separated; performing feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated; the frequency domain feature graph is used to characterize the characteristics of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and the time feature graph is used to characterize the characteristics of the frequency domain amplitude spectrum in different frequency dimensions at different times; performing attention fusion processing on the frequency domain feature graph and the time feature graph to obtain a fusion feature graph; performing decoding processing on the fusion feature graph to obtain a human voice amplitude spectrum and a background accompaniment amplitude spectrum corresponding to the audio to be separated. This method can capture the distribution law of different musical instruments in the frequency spectrum and improve the separation effect of the audio to be separated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of audio processing, and in particular to an audio separation method, device, electronic device and storage medium. Background Art

[0002] Audio separation is a technique for separating singing voices and background accompaniment from mixed music. With the development of deep learning technology, many neural network-based methods have been proposed for audio separation, for example, audio separation based on convolutional neural networks (CNNs). Since polyphonic music usually has long-term dependencies and wide bandwidths, the large receptive field of CNNs (the size of the area mapped to the original image by the pixels on the feature map output by each layer of the convolutional neural network) is crucial for singing voice separation.

[0003] At present, most of the methods to increase the receptive field are to cover a larger area of ​​the input spectrum by stacking CNN layers, or to use dilated convolution to cover a large receptive field by using a nonlinearly growing expansion factor. However, since polyphonic music usually contains more than one instrument, each instrument has its own unique sound characteristics, resulting in their distribution in the spectrum having certain regularities. The current CNN-based audio separation method ignores this point, and the limited receptive field makes it difficult for the CNN-based model to capture this pattern. Therefore, the separation results of the current audio separation method are still not accurate enough. Summary of the invention

[0004] The present disclosure provides an audio separation method, device, electronic device and storage medium to at least solve the problem that the separation result of audio separation in the related art is still not accurate enough. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an audio separation method, comprising:

[0006] Obtain the frequency domain amplitude spectrum corresponding to the audio to be separated;

[0007] Performing feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated; the frequency domain feature graph is used to characterize the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and the time feature graph is used to characterize the features of the frequency domain amplitude spectrum in frequency domain dimensions at different times;

[0008] Performing attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map;

[0009] The fusion feature map is decoded to obtain the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0010] In an exemplary embodiment, the step of obtaining a frequency domain amplitude spectrum corresponding to the audio to be separated includes:

[0011] Get the audio to be separated;

[0012] The audio to be separated is subjected to time-frequency conversion processing to obtain a frequency domain amplitude spectrum corresponding to the audio to be separated; the time-frequency conversion is used to convert the time domain audio sampling points of the audio to be separated into frequency domain features.

[0013] In an exemplary embodiment, the performing feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated includes:

[0014] Inputting the frequency domain amplitude spectrum into a frequency domain conversion network to obtain a frequency domain feature map of the audio to be separated; the frequency domain conversion network is used to extract the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time;

[0015] The transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum is input into a time conversion network to obtain a time feature graph of the audio to be separated; the time conversion network is used to extract the features of the frequency domain amplitude spectrum in the frequency domain dimension at different times.

[0016] In an exemplary embodiment, the process of obtaining the frequency domain feature map through the frequency domain conversion network includes:

[0017] Recording the position of each audio feature in the input frequency domain amplitude spectrum;

[0018] The self-attention mechanism is used to calculate the weights of the audio features at each input position on the target audio features of the output as the attention result;

[0019] The attention result is mapped to the separable space of the frequency domain amplitude spectrum to obtain a frequency domain feature map of the audio to be separated; the dimension of the separable space is the same as the dimension of the frequency domain amplitude spectrum.

[0020] In an exemplary embodiment, performing attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map includes:

[0021] Performing attention extraction processing on the frequency domain feature graph and the time feature graph to obtain attention graphs corresponding to the frequency domain feature graph and the time feature graph respectively; the attention graphs contain weight values ​​of various audio features;

[0022] The frequency domain feature map and the time feature map are weighted by the attention maps corresponding to the frequency domain feature map and the time feature map, so as to obtain weighted feature maps corresponding to the frequency domain feature map and the time feature map;

[0023] The weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map are added together to obtain the fused feature map.

[0024] In an exemplary embodiment, the performing attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map further includes:

[0025] Obtaining a high-dimensional feature map corresponding to the frequency domain amplitude spectrum;

[0026] Attention fusion processing is performed on the frequency domain feature map, the time feature map and the high-dimensional feature map to obtain a fused feature map.

[0027] In an exemplary embodiment, the decoding process of the fused feature map to obtain a human voice amplitude spectrum and a background accompaniment amplitude spectrum corresponding to the audio to be separated includes:

[0028] The fused feature map is input into a decoding network to obtain a human voice frequency domain amplitude spectrum and a background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated; the decoding network includes two convolution branch networks, which are respectively used to decode the human voice and the background accompaniment in the fused feature map;

[0029] The human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum are subjected to inverse time-frequency conversion processing to convert the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum into time domain audio signals respectively as the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0030] According to a second aspect of an embodiment of the present disclosure, there is provided an audio separation device, including:

[0031] An acquisition unit is configured to acquire a frequency domain amplitude spectrum corresponding to the audio to be separated;

[0032] An extraction unit is configured to perform feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated; the frequency domain feature graph is used to characterize the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and the time feature graph is used to characterize the features of the frequency domain amplitude spectrum in frequency domain dimensions at different times;

[0033] A fusion unit is configured to perform attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map;

[0034] The decoding unit is configured to perform decoding processing on the fused feature map to obtain the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0035] In an exemplary embodiment, the acquisition unit is specifically configured to acquire the audio to be separated; perform time-frequency conversion on the audio to be separated to obtain a frequency domain amplitude spectrum corresponding to the audio to be separated; and the time-frequency conversion is used to convert the time domain audio sampling points of the audio to be separated into frequency domain features.

[0036] In an exemplary embodiment, the extraction unit is specifically configured to input the frequency domain amplitude spectrum into a frequency domain conversion network to obtain a frequency domain feature map of the audio to be separated; the frequency domain conversion network is used to extract the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time; the transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum is input into a time conversion network to obtain a time feature map of the audio to be separated; the time conversion network is used to extract the features of the frequency domain amplitude spectrum in frequency domain dimensions at different times.

[0037] In an exemplary embodiment, the extraction unit is further configured to record the positions of each audio feature in the input frequency domain amplitude spectrum; calculate the weights of the audio features at each position of the input to the output target audio features through a self-attention mechanism as attention results; map the attention results to the separable space of the frequency domain amplitude spectrum to obtain a frequency domain feature map of the audio to be separated; the dimension of the separable space is the same as the dimension of the frequency domain amplitude spectrum.

[0038] In an exemplary embodiment, the fusion unit is specifically configured to perform attention extraction processing on the frequency domain feature map and the time feature map to obtain attention maps corresponding to the frequency domain feature map and the time feature map respectively; the attention map contains weight values ​​of various audio features; the frequency domain feature map and the time feature map are weighted through the attention maps corresponding to the frequency domain feature map and the time feature map respectively to obtain weighted feature maps corresponding to the frequency domain feature map and the time feature map respectively; the weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map are added to obtain the fused feature map.

[0039] In an exemplary embodiment, the fusion unit is further configured to execute the acquisition of a high-dimensional feature map corresponding to the frequency domain amplitude spectrum; perform attention fusion processing on the frequency domain feature map, the time feature map and the high-dimensional feature map to obtain a fused feature map.

[0040] In an exemplary embodiment, the decoding unit is specifically configured to input the fused feature map into a decoding network to obtain a human voice frequency domain amplitude spectrum and a background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated; the decoding network includes two convolution branch networks, which are respectively used to decode the human voice and the background accompaniment in the fused feature map; the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum are inversely converted to time-frequency conversion processing to convert the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum into time domain audio signals, respectively, as the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0041] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0042] processor;

[0043] a memory for storing instructions executable by the processor;

[0044] The processor is configured to execute the instructions to implement the method as described in any one of the above items.

[0045] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any of the methods described above.

[0046] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, wherein the computer program product includes instructions, and when the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute any of the methods described above.

[0047] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0048] After obtaining the frequency domain amplitude spectrum of the audio to be separated, the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time and the features of the frequency domain dimensions at different times are extracted respectively, and the frequency domain feature map and time feature map of the audio to be separated are obtained. The frequency domain feature map and the time feature map are further subjected to attention fusion processing to obtain a fusion feature map, and the human voice amplitude spectrum and background accompaniment amplitude spectrum corresponding to the audio to be separated are obtained by decoding the fusion feature map. This method extracts the features of the audio to be separated from the frequency domain time dimension and the frequency dimension respectively, thereby increasing the coverage of the receptive field of the audio to be separated, better learning the complex audio patterns of the frequency domain time dimension and the frequency dimension, thereby better capturing the distribution law of different instruments in the frequency spectrum, and improving the separation effect of the audio to be separated.

[0049] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0051] Figure 1 The figure is a flowchart of an audio separation method according to an exemplary embodiment.

[0052] Figure 2 It is an overall architecture diagram of a time-frequency domain Transformer network according to an exemplary embodiment.

[0053] Figure 3 is a schematic diagram of a network structure of an attention fusion module according to an exemplary embodiment.

[0054] Figure 4 The figure is a schematic diagram of a network structure of a decoding network according to an exemplary embodiment.

[0055] Figure 5 It is a flowchart of an audio separation method according to another exemplary embodiment.

[0056] Figure 6 is a schematic diagram showing comparison of audio separation results according to an exemplary embodiment.

[0057] Figure 7 The figure is a structural block diagram of an audio separation device according to an exemplary embodiment.

[0058] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0059] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0060] It should be noted that the embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0061] In an exemplary embodiment, if Figure 1As shown, an audio separation method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0062] In step 110, a frequency domain amplitude spectrum corresponding to the audio to be separated is obtained.

[0063] The audio to be separated refers to mixed music including singing voice and background accompaniment.

[0064] The frequency domain amplitude spectrum may represent an amplitude spectrum obtained by converting a feature map of the audio to be separated into the frequency domain.

[0065] In a specific implementation, after obtaining the audio to be separated that includes singing vocals and background accompaniment, the audio to be separated can be subjected to video conversion processing, that is, the time domain audio sampling points of the audio to be separated are converted into frequency domain features to obtain the frequency domain amplitude spectrum corresponding to the audio to be separated.

[0066] In step 120, feature extraction processing is performed on the frequency domain amplitude spectrum to obtain a frequency domain feature map and a time feature map of the audio to be separated, wherein the frequency domain feature map is used to characterize the characteristics of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and the time feature map is used to characterize the characteristics of the frequency domain amplitude spectrum in the frequency domain dimensions at different times.

[0067] In a specific implementation, in order to increase the receptive field and cover a larger area of ​​the frequency domain amplitude spectrum, feature extraction can be performed on the frequency domain amplitude spectrum of the audio to be separated from the frequency dimension and the frequency domain time dimension respectively, to obtain a frequency domain feature map representing the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and a time feature map representing the features of the frequency domain amplitude spectrum in the frequency domain dimensions at different times, thereby covering the global receptive field of the audio to be separated and better learning the complex audio patterns on the frequency domain time and frequency axes. More specifically, the method of extracting features from the frequency domain amplitude spectrum of the audio to be separated from the frequency dimension and the frequency domain time dimension respectively can be to extract features from the frequency domain amplitude spectrum separately through two different networks, one network is used to extract the frequency domain features of the audio to be separated, and the other network is used to extract the time features of the audio to be separated.

[0068] In step 130, attention fusion processing is performed on the frequency domain feature map and the time feature map to obtain a fused feature map.

[0069] Among them, the attention fusion processing is used to fuse the frequency domain feature map and time feature map of the audio to be separated to obtain richer audio features.

[0070] In the specific implementation, before performing attention fusion processing on the frequency domain feature map and the time feature map, it is necessary to first perform attention extraction processing on the frequency domain feature map and the time feature map respectively to obtain the attention maps of the frequency domain feature map and the time feature map respectively, and the attention maps contain the weight values ​​of each audio feature of the frequency domain feature map and the time feature map. Further, the frequency domain feature map, the time feature map and their corresponding attention maps are weighted to obtain the weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map, and the weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map are added to obtain the fused feature map.

[0071] In step 140, the fused feature map is decoded to obtain the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0072] In a specific implementation, a decoding network can be pre-trained and the fused feature map can be decoded by the decoding network. Since the fused feature map is obtained based on the frequency domain amplitude spectrum after the audio to be separated is converted into the frequency domain, after the fused feature map is decoded, the obtained vocal amplitude spectrum and background accompaniment amplitude spectrum also correspond to the frequency domain amplitude spectrum. Therefore, it is necessary to perform inverse time-frequency conversion on the decoded vocal frequency domain amplitude spectrum and background accompaniment frequency domain amplitude spectrum to convert the vocal frequency domain amplitude spectrum and background accompaniment frequency domain amplitude spectrum into time domain audio signals respectively, so as to obtain the vocal amplitude spectrum and background accompaniment amplitude spectrum of the audio to be separated.

[0073] In the above audio separation method, after obtaining the frequency domain amplitude spectrum of the audio to be separated, the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time and the features of the frequency domain dimensions at different times are extracted respectively, and the frequency domain feature map and time feature map of the audio to be separated are obtained, and the frequency domain feature map and the time feature map are further subjected to attention fusion processing to obtain a fusion feature map, and the human voice amplitude spectrum and background accompaniment amplitude spectrum corresponding to the audio to be separated are obtained by decoding the fusion feature map. This method increases the coverage of the receptive field of the audio to be separated by extracting the features of the audio to be separated from the frequency domain time dimension and the frequency dimension respectively, and better learns the complex audio patterns of the frequency domain time dimension and the frequency dimension, so as to better capture the distribution law of different musical instruments in the frequency spectrum and improve the separation effect of the audio to be separated.

[0074] In an exemplary embodiment, in step 110, obtaining the frequency domain amplitude spectrum corresponding to the audio to be separated can be specifically achieved through the following steps: obtaining the audio to be separated; performing time-frequency conversion on the audio to be separated to obtain the frequency domain amplitude spectrum corresponding to the audio to be separated; the time-frequency conversion is used to convert the time domain audio sampling points of the audio to be separated into frequency domain features.

[0075] The time-frequency conversion means converting the time domain representation of the audio signal of the audio to be separated into the frequency domain representation.

[0076] In a specific implementation, after obtaining the audio to be separated, the time domain audio sampling points of the audio to be separated can be converted into frequency domain features through time-frequency conversion processing to obtain the frequency domain amplitude spectrum corresponding to the audio to be separated. More specifically, the time domain audio sampling points of the audio to be separated can be converted into frequency domain features through short-time Fourier transform (STFT), wherein the window length of the short-time Fourier transform can be set to 64ms and the window shift can be set to 32ms.

[0077] In this embodiment, the frequency domain amplitude spectrum of the audio to be separated is obtained by performing time-frequency conversion processing on the audio to be separated, so as to further perform feature extraction on the frequency domain amplitude spectrum to obtain audio features of the audio to be separated in different dimensions.

[0078] In an exemplary embodiment, the feature extraction processing of the frequency domain amplitude spectrum in the above step 120 to obtain the frequency domain feature graph and time feature graph of the audio to be separated can be achieved in the following manner: the frequency domain amplitude spectrum is input into the frequency domain conversion network to obtain the frequency domain feature graph of the audio to be separated; the frequency domain conversion network is used to extract the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time; the transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum is input into the time conversion network to obtain the time feature graph of the audio to be separated; the time conversion network is used to extract the features of the frequency domain amplitude spectrum in the frequency domain dimensions at different times.

[0079] In the specific implementation, both the time transformation network (time-Transformer) and the frequency domain transformation network (frequency-Transformer) can be Transformer networks. The time transformation network inputs the time axis of the spectrum as the baseline axis into the network, and the time transformation network is used to capture the feature correlation of the frequency domain dimensions at different times; the frequency domain transformation network inputs the frequency axis of the spectrum as the baseline axis into the network, and the frequency domain transformation network is used to capture the features of different frequency dimensions at the same time. Moreover, the time transformation network and the frequency domain transformation network can have the same model structure. For example, refer to Figure 2, is the overall architecture diagram of the time-frequency domain Transformer network, including the time-frequency domain Transformer network, the attention fusion module and the decoding / separation network. The branch network on the left side of the figure represents the frequency domain conversion network, and the branch network on the right side represents the time conversion network. It can be seen from the figure that the model structure of the time conversion network and the frequency domain conversion network is the same, both including the position feature layer, the multi-head attention mechanism, the layer norm and the feedforward neural network. Among them, the position feature layer is used to record the positions of different features, to ensure the unidirectionality of the audio features and to prevent the order from being disordered; the multi-head attention mechanism includes multiple point multiplication self-attention mechanisms, which are used to associate different positions of the input sequence and calculate the attention weights of the output of a certain position and the input of different positions. The larger the weight, the greater the contribution of the input at this position to the output of the corresponding position; the multi-head attention mechanism is used to utilize the results of attention at different levels; the fully connected network is used to map the attention results to a separable space.

[0080] In this embodiment, two independent frequency domain conversion networks and time conversion networks are used to respectively extract the features of the audio to be separated in the frequency domain time dimension and the frequency dimension, thereby improving the efficiency of obtaining the frequency domain feature map and the time feature map of the audio to be separated. Moreover, both the frequency domain conversion network and the time conversion network adopt the Transformer architecture, and the self-attention mechanism of the Transformer architecture can be used to encode the global input-output dependency, thereby effectively capturing the long-term dependency of the input audio.

[0081] In an exemplary embodiment, the process of obtaining a frequency domain feature map through a frequency domain conversion network includes: recording the position of each audio feature in the input frequency domain amplitude spectrum; calculating the weight of the audio features at each position of the input to the output target audio features through a self-attention mechanism as the attention result; mapping the attention result to the separable space of the frequency domain amplitude spectrum to obtain the frequency domain feature map of the audio to be separated; the dimension of the separable space is the same as the dimension of the frequency domain amplitude spectrum.

[0082] In the specific implementation, refer to Figure 2 As shown in the overall architecture diagram of the time-frequency domain Transformer network, a position feature layer can be set in the frequency domain conversion network to record the position of each audio feature in the input frequency domain amplitude spectrum. The point product self-attention mechanism is used to calculate the weight of the audio features at each input position to the output target audio features as the attention result, and the attention result is mapped to the separable space of the frequency domain amplitude spectrum to obtain the frequency domain feature map of the audio to be separated.

[0083] In this embodiment, the position of each audio feature in the input frequency domain amplitude spectrum is recorded to ensure the unidirectionality of the audio features and prevent order confusion. The global input-output dependency can be encoded through the self-attention mechanism to effectively capture the long-term dependency of the input audio.

[0084] In an exemplary embodiment, in step 130, attention fusion processing is performed on the frequency domain feature map and the time feature map to obtain a fused feature map, which can be achieved in the following manner: attention extraction processing is performed on the frequency domain feature map and the time feature map to obtain attention maps corresponding to the frequency domain feature map and the time feature map respectively; the attention map contains weight values ​​of each audio feature; the frequency domain feature map, the time feature map and the attention maps corresponding to the frequency domain feature map and the time feature map are fused to obtain a fused feature map.

[0085] In the specific implementation, the frequency domain feature map and the time feature map can be subjected to attention extraction processing through the trained attention extraction network. More specifically, the frequency domain feature map and the time feature map are first subjected to dimension reduction processing through the pooling layer (Global average Pooling, gap), and then the feature map after dimension reduction processing is further subjected to feature extraction through the nonlinear fully connected layers (Fully Connected layers, FC), and the extracted features are separated according to the level, and the importance of each channel of the feature map is learned again through a fully connected layer, and finally the attention map corresponding to the frequency domain feature map and the attention map corresponding to the time feature map are obtained through the softmax layer. The frequency domain feature map, the time feature map, and the attention map corresponding to the obtained frequency domain feature map and the attention map corresponding to the time feature map are fused to obtain a fused feature map.

[0086] Furthermore, in an exemplary embodiment, the frequency domain feature map, the time feature map, and the attention maps corresponding to the frequency domain feature map and the time feature map are fused to obtain a fused feature map, including: weighting the frequency domain feature map and the time feature map through the attention maps corresponding to the frequency domain feature map and the time feature map to obtain weighted feature maps corresponding to the frequency domain feature map and the time feature map; adding the weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map to obtain a fused feature map.

[0087] In the specific implementation, since the attention map contains the weight values ​​of each audio feature, that is, the attention map corresponding to the frequency domain feature map contains the weight values ​​of each audio feature in the frequency domain feature map, and the attention map corresponding to the time feature map contains the weight values ​​of each audio feature in the time feature map, therefore, when performing the fusion processing, the frequency domain feature map and the time feature map are first weighted with their respective corresponding attention maps, that is, the frequency domain feature map is multiplied with the attention map corresponding to the frequency domain feature map to obtain the weighted feature map of the frequency domain feature map, and the time feature map is multiplied with the attention map corresponding to the time feature map to obtain the weighted feature map of the time feature map, and then the weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map are added through element-by-element addition operation to obtain a fused feature map.

[0088] In this embodiment, the frequency domain feature map and the time feature map are first subjected to attention extraction processing to obtain the attention maps corresponding to the frequency domain feature map and the time feature map respectively, and then the frequency domain feature map and the time feature map are subjected to weighted summation processing with their respective corresponding attention maps to obtain rich feature information of the audio to be separated.

[0089] In an exemplary embodiment, in step 130, the frequency domain feature map and the time feature map are subjected to attention fusion processing to obtain a fused feature map, and the step also includes: performing feature extraction processing on the frequency domain amplitude spectrum to obtain a characteristic frequency domain amplitude spectrum; performing attention fusion processing on the frequency domain feature map, the time feature map and the characteristic frequency domain amplitude spectrum to obtain a fused feature map.

[0090] In the specific implementation, before performing attention fusion processing on the frequency domain feature map and the time feature map, the high-dimensional features of the frequency domain amplitude spectrum can also be extracted to obtain the high-dimensional feature map corresponding to the frequency domain amplitude spectrum. The high-dimensional feature map, the frequency domain feature map and the time feature map are respectively subjected to attention extraction processing to obtain the attention maps corresponding to the three feature maps. The three feature maps and the corresponding attention maps are then weighted to obtain three weighted feature maps. Finally, the three weighted feature maps are fused through element-by-element addition operation to obtain a fused feature map.

[0091] For example, refer to Figure 3 , is a schematic diagram of the network structure of the attention fusion module, where S', F t and F sThey respectively represent the high-dimensional feature map, time feature map and frequency domain feature map corresponding to the frequency domain amplitude spectrum. The three feature maps are taken as input. After addition, they are first reduced in dimension through the pooling layer (GAP) to retain the important features in the three feature maps. Then, they are further extracted through the fully connected layer (FC). After that, the extracted features are separated into three layers according to different levels. The importance of each channel of the feature map is learned through three fully connected layers. Three attention maps are obtained through the softmax layer (conversion layer), which are used as the attention maps corresponding to the high-dimensional feature map, time feature map and frequency domain feature map respectively. Finally, the three input feature maps and the three attention maps are multiplied according to the corresponding relationship to obtain three weighted feature maps. The three weighted feature maps are added element by element to obtain a fused feature map.

[0092] In this embodiment, when the frequency domain feature map and the time feature map are subjected to attention fusion processing, a high-dimensional feature map corresponding to the frequency domain amplitude spectrum is also added. On the one hand, it is convenient to better perform the regression task through the high-dimensional feature map, and on the other hand, richer feature information of the audio to be separated can be obtained.

[0093] In an exemplary embodiment, the fused feature map is decoded in the above step S140 to obtain the vocal amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated, which can be achieved in the following way: the fused feature map is input into the decoding network to obtain the vocal frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated; the decoding network includes two convolution branch networks, which are respectively used to decode the vocals and the background accompaniment in the fused feature map; the vocal frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum are inversely time-frequency converted to convert the vocal frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum into time domain audio signals, respectively, as the vocal amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0094] In a specific implementation, the decoding network may include two convolutional branch networks, respectively used to decode the human voice and background accompaniment in the fusion feature map. For example, Figure 4 , is a schematic diagram of the network structure of the decoding network. It can be seen from the figure that the structures and parameters of the two branch networks are the same. In practical applications, the obtained fusion feature map can be input into the decoding network. The two branch networks of the decoding network respectively decode / separate the fusion feature map to obtain the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated. Since the fusion feature map is obtained based on the frequency domain amplitude spectrum of the audio to be separated as input, it is necessary to perform inverse Fourier transform processing on the obtained human voice frequency domain amplitude spectrum and background accompaniment frequency domain amplitude spectrum, and restore the separated singing vocals and accompaniment to time domain audio signals and output them.

[0095] In this embodiment, by setting two convolution branch networks in the decoding network to separate the human voice and the background accompaniment in the fusion feature map respectively, the efficiency and accuracy of the separation can be improved.

[0096] In another exemplary embodiment, Figure 5 FIG. 2 is a flowchart of an audio separation method according to another exemplary embodiment. In this embodiment, the method includes the following steps:

[0097] Step 510, performing time-frequency conversion processing on the audio to be separated to obtain a frequency domain amplitude spectrum corresponding to the audio to be separated;

[0098] Step 520, input the frequency domain amplitude spectrum into the frequency domain conversion network to obtain the frequency domain feature map of the audio to be separated; the frequency domain conversion network is used to extract the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time;

[0099] Step 530, inputting the transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum into a time conversion network to obtain a time characteristic graph of the audio to be separated; the time conversion network is used to extract the characteristics of the frequency domain amplitude spectrum at different time points in the frequency domain dimension;

[0100] Step 540, obtaining a high-dimensional feature map corresponding to the frequency domain amplitude spectrum;

[0101] Step 550, performing attention extraction processing on the high-dimensional feature map, the frequency domain feature map and the time feature map to obtain respective corresponding attention maps; the attention map contains the weight value of each audio feature;

[0102] Step 560, multiplying the high-dimensional feature map, the frequency domain feature map, and the frequency domain feature map with their corresponding attention maps, respectively, to obtain weighted feature maps corresponding to the high-dimensional feature map, the frequency domain feature map, and the frequency domain feature map;

[0103] Step 570, adding the weighted feature map of the high-dimensional feature map, the weighted feature map of the frequency domain feature map, and the weighted feature map of the time feature map to obtain a fused feature map.

[0104] Step 580, input the fused feature map into a decoding network to obtain a human voice frequency domain amplitude spectrum and a background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated;

[0105] Step 590, perform inverse time-frequency conversion on the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum to convert the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum into time domain audio signals respectively, as the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0106] The audio separation method based on the time-frequency domain attention mechanism proposed in this embodiment uses the time-frequency domain Transformer network to access the global receptive field of the input audio and better learn the complex patterns of music on the frequency and time axes. Specifically, by constructing a Transformer based on the frequency domain time axis and a Transformer based on the frequency axis, the two Transformers encode the input audio and output features respectively. The time-frequency domain Transformer effectively covers the global receptive field of the audio and better learns the complex patterns of the audio on the frequency domain time and frequency axes. The self-attention mechanism in the Transformer can encode the global input-output dependency and effectively capture the long-term dependency of the input audio.

[0107] In an exemplary embodiment, in order to facilitate those skilled in the art to understand the embodiment of the present application, the audio separation method proposed in the present disclosure will be further described below with reference to the specific examples of the accompanying drawings. Figure 2 As shown, it is an overall architecture diagram of the time-frequency domain Transformer network for audio separation proposed in the present disclosure, including: a time-frequency domain Transformer network, an attention fusion module and a decoding / separation network, wherein the time-frequency domain Transformer network includes two independent Transformer networks, a time-Transformer network and a frequency-Transformer network, and the model structures of the two Transformer networks are the same.

[0108] (1) After obtaining the audio to be separated, performing a short-time Fourier transform on the audio to be separated to obtain a frequency domain amplitude spectrum corresponding to the audio to be separated, and transposing the frequency domain amplitude spectrum to obtain a transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum;

[0109] (2) Figure 2 As shown, the frequency domain amplitude spectrum is input into the frequency-Transformer network to obtain the frequency domain feature graphs of the audio to be separated in different frequency dimensions at the same time, and the transposed amplitude spectrum is input into the time-Transformer network to obtain the time feature graphs of the audio to be separated in the frequency domain dimensions at different times;

[0110] (3) Extract high-dimensional features from the frequency domain amplitude spectrum to obtain a high-dimensional feature map corresponding to the frequency domain amplitude spectrum, such as Figure 3 As shown, the high-dimensional feature map S' and the time feature map F t , frequency domain feature map F s Input the attention fusion module, through which the high-dimensional feature map S' and the temporal feature map F are first obtained. t , frequency domain feature map F sThe corresponding attention maps are then combined with the high-dimensional feature map S' and the time feature map F t , frequency domain feature map F s Multiply the corresponding attention maps respectively to obtain their own weighted feature maps, and add the weighted feature maps to obtain the fused feature map;

[0111] (4) Input the fused feature map into Figure 4 The decoding network shown in the figure obtains the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated through the decoding network, and uses the inverse Fourier transform to restore the separated singing human voice and background accompaniment to the time domain audio signal and output it.

[0112] By testing and comparing the present disclosure with four different audio separation models, the effectiveness and robustness of the model proposed in the present disclosure can be confirmed. Specifically, four commonly used and advanced music separation methods can be selected as baseline systems, including: UNet based on convolutional network, GRU-Dilation based on recurrent neural network, D3Net based on hole convolutional network and Sams-Net based on Transformer. The time-frequency domain Transformer model proposed in the present disclosure and the four baseline methods are trained on the same dataset. After training, compared with the baseline methods, the method proposed in the present disclosure obtained the highest score overall, and the results clearly confirmed the effectiveness and robustness of our proposed model. As shown in the following tables (a)-(c), compared with other baseline methods, when focusing on the SDR of the singing vocal part, the proposed method is 0.8% higher than the second best D3Net in iKala, 78.6% higher in MUDB18, and 72.3% higher in DSD100. When focusing on the SDR of the background accompaniment part, the proposed method outperforms the second best D3Net by 25.2% in iKala, 7.2% in MUDB18, and 20.6% in DSD100.

[0113]

[0114] (a) Kala

[0115]

[0116] (b)MUSDB18

[0117]

[0118] (c)DSD100

[0119] A case study is conducted on a popular song “All Souls Moon” in the DSD100 dataset. Figure 6As shown, it can be observed that the model proposed in the present disclosure can effectively separate the singing voice and the background accompaniment, which also illustrates the effectiveness of the method proposed in the present disclosure.

[0120] The audio separation method based on the time-frequency domain attention mechanism proposed in this embodiment, the time-frequency domain Transformer model is used to learn the time and spectrum features with a global receptive field, and the feature attention fusion module fuses the feature information and separates it. The experimental results show that the proposed model outperforms several existing state-of-the-art models on three data sets. The audio separation method disclosed in the present invention can be used in the production of short videos and multi-track subtitles to increase the gameplay and fun of the video.

[0121] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0122] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can refer to each other, and each embodiment focuses on the differences from other embodiments. For related points, please refer to the description of other method embodiments.

[0123] Figure 7 is a structural block diagram of an audio separation device according to an exemplary embodiment. Figure 7 The device includes: an acquisition unit 710, an extraction unit 720, a fusion unit 730 and a decoding unit 740, wherein:

[0124] An acquisition unit 710 is configured to acquire a frequency domain amplitude spectrum corresponding to the audio to be separated;

[0125] The extraction unit 720 is configured to perform feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated; the frequency domain feature graph is used to characterize the characteristics of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and the time feature graph is used to characterize the characteristics of the frequency domain amplitude spectrum in the frequency domain dimensions at different times;

[0126] A fusion unit 730 is configured to perform attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map;

[0127] The decoding unit 740 is configured to perform decoding processing on the fused feature map to obtain the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0128] In an exemplary embodiment, the acquisition unit is specifically configured to acquire the audio to be separated; perform time-frequency conversion on the audio to be separated to obtain a frequency domain amplitude spectrum corresponding to the audio to be separated; and the time-frequency conversion is used to convert the time domain audio sampling points of the audio to be separated into frequency domain features.

[0129] In an exemplary embodiment, the extraction unit is specifically configured to execute the input of the frequency domain amplitude spectrum into the frequency domain conversion network to obtain the frequency domain feature map of the audio to be separated; the frequency domain conversion network is used to extract the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time; the transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum is input into the time conversion network to obtain the time feature map of the audio to be separated; the time conversion network is used to extract the features of the frequency domain amplitude spectrum in the frequency domain dimensions at different times.

[0130] In an exemplary embodiment, the extraction unit is further configured to record the positions of each audio feature in the input frequency domain amplitude spectrum; calculate the weights of the audio features at each position of the input to the output target audio features through a self-attention mechanism as attention results; map the attention results to the separable space of the frequency domain amplitude spectrum to obtain a frequency domain feature map of the audio to be separated; the dimension of the separable space is the same as the dimension of the frequency domain amplitude spectrum.

[0131] In an exemplary embodiment, the fusion unit is specifically configured to perform attention extraction processing on the frequency domain feature map and the time feature map to obtain attention maps corresponding to the frequency domain feature map and the time feature map respectively; the attention map contains weight values ​​of each audio feature; the frequency domain feature map, the time feature map and the attention maps corresponding to the frequency domain feature map and the time feature map are fused to obtain a fused feature map.

[0132] In an exemplary embodiment, the fusion unit is further configured to perform a multiplication process on the frequency domain feature map and the attention map corresponding to the frequency domain feature map to obtain a weighted feature map of the frequency domain feature map; to perform a multiplication process on the time feature map and the attention map corresponding to the time feature map to obtain a weighted feature map of the time feature map; and to perform an addition process on the weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map to obtain a fused feature map.

[0133] In an exemplary embodiment, the fusion unit is further configured to execute the acquisition of a high-dimensional feature map corresponding to the frequency domain amplitude spectrum; perform attention fusion processing on the frequency domain feature map, the time feature map and the high-dimensional feature map to obtain a fused feature map.

[0134] In an exemplary embodiment, the decoding unit is specifically configured to execute inputting the fused feature map into a decoding network to obtain a human voice frequency domain amplitude spectrum and a background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated; the decoding network includes two convolution branch networks, which are respectively used to decode the human voice and the background accompaniment in the fused feature map; the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum are inversely converted to time-frequency conversion processing to convert the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum into time-domain audio signals, respectively, as the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

[0135] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0136] Figure 8 1 is a block diagram of an electronic device 800 for implementing an audio separation method according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0137] Reference Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0138] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0139] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, an optical disk, or a graphene memory.

[0140] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0141] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0142] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.

[0143] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0144] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of the components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or the components of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the device 800 and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0145] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0146] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0147] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0148] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions. The instructions can be executed by the processor 820 of the electronic device 800 to complete the above method.

[0149] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. may also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments, which will not be described one by one here.

[0150] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0151] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.

[0152] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An audio separation method, characterized in that: include: Obtaining a frequency domain amplitude spectrum corresponding to the audio to be separated; the frequency domain amplitude spectrum represents an amplitude spectrum obtained by converting a feature map of the audio to be separated into the frequency domain; Performing feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated; the frequency domain feature graph is used to characterize the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and the time feature graph is used to characterize the features of the frequency domain amplitude spectrum in frequency domain dimensions at different times; Performing attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map; The fusion feature map is decoded to obtain the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

2. The method according to claim 1, characterized in that The obtaining of the frequency domain amplitude spectrum corresponding to the audio to be separated includes: Get the audio to be separated; The audio to be separated is subjected to time-frequency conversion processing to obtain a frequency domain amplitude spectrum corresponding to the audio to be separated; the time-frequency conversion processing is used to convert the time domain audio sampling points of the audio to be separated into frequency domain features.

3. The method according to claim 1, characterized in that: The step of performing feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated includes: Inputting the frequency domain amplitude spectrum into a frequency domain conversion network to obtain a frequency domain feature map of the audio to be separated; the frequency domain conversion network is used to extract the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time; The transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum is input into a time conversion network to obtain a time feature graph of the audio to be separated; the time conversion network is used to extract the features of the frequency domain amplitude spectrum in the frequency domain dimension at different times.

4. The method according to claim 3, characterized in that The process of obtaining the frequency domain feature map through the frequency domain conversion network includes: Recording the position of each audio feature in the input frequency domain amplitude spectrum; The self-attention mechanism is used to calculate the weights of the audio features at each input position on the target audio features of the output as the attention result; The attention result is mapped to the separable space of the frequency domain amplitude spectrum to obtain a frequency domain feature map of the audio to be separated; the dimension of the separable space is the same as the dimension of the frequency domain amplitude spectrum.

5. The method according to claim 1, characterized in that The performing attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map includes: Performing attention extraction processing on the frequency domain feature graph and the time feature graph to obtain attention graphs corresponding to the frequency domain feature graph and the time feature graph respectively; the attention graphs contain weight values ​​of various audio features; The frequency domain feature map and the time feature map are weighted by the attention maps corresponding to the frequency domain feature map and the time feature map, so as to obtain weighted feature maps corresponding to the frequency domain feature map and the time feature map; The weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map are added together to obtain the fused feature map.

6. The method according to claim 1, characterized in that The performing attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map further includes: Obtaining a high-dimensional feature map corresponding to the frequency domain amplitude spectrum; Attention fusion processing is performed on the frequency domain feature map, the time feature map and the high-dimensional feature map to obtain a fused feature map.

7. The method according to claim 1, characterized in that The decoding process of the fused feature map to obtain a human voice amplitude spectrum and a background accompaniment amplitude spectrum corresponding to the audio to be separated includes: The fused feature map is input into a decoding network to obtain a human voice frequency domain amplitude spectrum and a background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated; the decoding network includes two convolution branch networks, which are respectively used to decode the human voice and the background accompaniment in the fused feature map; The human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum are subjected to inverse time-frequency conversion processing to convert the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum into time domain audio signals respectively as the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

8. An audio separation device, characterized in that: include: An acquisition unit is configured to acquire a frequency domain amplitude spectrum corresponding to the audio to be separated; The frequency domain amplitude spectrum represents the amplitude spectrum obtained by converting the feature map of the audio to be separated into the frequency domain; An extraction unit is configured to perform feature extraction processing on the frequency domain amplitude spectrum to obtain a frequency domain feature graph and a time feature graph of the audio to be separated; the frequency domain feature graph is used to characterize the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time, and the time feature graph is used to characterize the features of the frequency domain amplitude spectrum in frequency domain dimensions at different times; A fusion unit is configured to perform attention fusion processing on the frequency domain feature map and the time feature map to obtain a fused feature map; The decoding unit is configured to perform decoding processing on the fused feature map to obtain the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

9. The device according to claim 8, characterized in that The acquisition unit is specifically configured to acquire the audio to be separated; perform time-frequency conversion on the audio to be separated to obtain a frequency domain amplitude spectrum corresponding to the audio to be separated; and the time-frequency conversion is used to convert the time domain audio sampling points of the audio to be separated into frequency domain features.

10. The device according to claim 8, characterized in that The extraction unit is specifically configured to execute inputting the frequency domain amplitude spectrum into a frequency domain conversion network to obtain a frequency domain feature map of the audio to be separated; the frequency domain conversion network is used to extract the features of the frequency domain amplitude spectrum in different frequency dimensions at the same time; the transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum is input into a time conversion network to obtain a time feature map of the audio to be separated; the time conversion network is used to extract the features of the frequency domain amplitude spectrum in frequency domain dimensions at different times.

11. The device according to claim 10, characterized in that The extraction unit is further configured to record the positions of each audio feature in the input frequency domain amplitude spectrum; calculate the weights of the audio features at each position of the input to the output target audio features through a self-attention mechanism as attention results; map the attention results to the separable space of the frequency domain amplitude spectrum to obtain a frequency domain feature map of the audio to be separated; the dimension of the separable space is the same as the dimension of the frequency domain amplitude spectrum.

12. The device according to claim 8, characterized in that The fusion unit is specifically configured to perform attention extraction processing on the frequency domain feature map and the time feature map to obtain the attention maps corresponding to the frequency domain feature map and the time feature map respectively; the attention map contains the weight values ​​of each audio feature; the frequency domain feature map and the time feature map are weighted by the attention maps corresponding to the frequency domain feature map and the time feature map respectively to obtain the weighted feature maps corresponding to the frequency domain feature map and the time feature map respectively; the weighted feature map of the frequency domain feature map and the weighted feature map of the time feature map are added to obtain the fused feature map.

13. The device according to claim 8, characterized in that The fusion unit is further configured to execute acquisition of a high-dimensional feature map corresponding to the frequency domain amplitude spectrum; perform attention fusion processing on the frequency domain feature map, the time feature map and the high-dimensional feature map to obtain a fused feature map.

14. The device according to claim 8, characterized in that The decoding unit is specifically configured to input the fused feature map into a decoding network to obtain a human voice frequency domain amplitude spectrum and a background accompaniment frequency domain amplitude spectrum corresponding to the audio to be separated; the decoding network includes two convolution branch networks, which are respectively used to decode the human voice and the background accompaniment in the fused feature map; and perform inverse time-frequency conversion processing on the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum to convert the human voice frequency domain amplitude spectrum and the background accompaniment frequency domain amplitude spectrum into time domain audio signals, respectively, as the human voice amplitude spectrum and the background accompaniment amplitude spectrum corresponding to the audio to be separated.

15. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the audio separation method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the audio separation method according to any one of claims 1 to 7.

17. A computer program product, comprising instructions, characterized in that: When the instructions are executed by a processor of an electronic device, the electronic device is enabled to perform the audio separation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Reverberation elimination method and device

    CN112289334A

  • Speech enhancement method based on two-channel convolutional attention network and system

    CN113611323A