A lightweight sound source separation method based on an improved attention mechanism
By introducing improved attention mechanism and GAP layer into the sound source separation technology and replacing the original gated-point convolution module and attention module, the problems of large parameters and large calculations in the existing technology are solved, and more efficient sound source separation performance and better generalization are achieved.
Patent Information
- Application Number
- CN202211507544.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-11-29
AI Technical Summary
The existing audio source separation technology has the problems of large amount of parameters, large amount of computing and long running time, and it is difficult to operate effectively in a real-time environment, especially on edge computing devices.
A lightweight sound source separation method based on an improved attention mechanism is proposed. By replacing the gating-point convolution full-connection computing module with global pooling in the LaSAFT network, the attention module is replaced by the Hybrid-Voiceformer multi-head spectrum hybrid attention module.
Without adding a very large amount of computing and parameter volume, the audio source separation performance is improved, which is better than the original LaSAFT network model, and the parameter volume is controlled during the separation process, which has better generalization.
Smart Images

Figure CN115798505B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech or sound processing, and specifically to a lightweight sound source separation method based on an improved attention mechanism. Background Technique
[0002] Music separation technology can be described as: separating the audio signals of one or more music audios from an existing mixed music audio signal. Although humans can easily distinguish melodies from mixed music signals, and even those with good musical senses and professional training can recognize melodies as musical scores, it is complex and extremely challenging for machines to perform this task. The main reasons are as follows: (1) A polyphonic music signal is composed of the superimposed sound waveforms generated by all the instruments in the recording. It is extremely difficult to separate the spectra that come from different sound sources and are highly coupled and superimposed according to the harmonic structure into the corresponding individual notes; (2) Even if the fundamental frequency sequence of the notes has been obtained, it is still necessary to determine which pitches belong to the melody and which belong to the accompaniment. When the melody is a vocal but there is background harmony, the detection is even more difficult.
[0003] Music separation technology has many direct or indirect applications. Studying music separation technology helps to promote the further improvement of related tasks such as music score recognition, musicology, intonation analysis, or music retrieval. In addition, by separating music sound sources, similarity analysis can be performed on the obtained dry sounds. Detecting similar segments between songs has important practical significance for the copyright protection and piracy detection of music works.
[0004] From the perspective of separation methods, music separation technology can be mainly divided into methods based on signal processing and methods based on neural networks. The representative of the methods based on signal processing is the non-negative matrix factorization method. This type of method decomposes the spectrogram into two types of values, dictionary values and activation values, and obtains different separation sources by multiplying the dictionary values with different activation values. The methods based on neural networks mainly use fully connected networks, recurrent neural networks, and time-domain separation method networks. This type of method repeatedly trains by importing a training set into the network and finally obtains good results.
[0005] The above two types of methods each have their own advantages and disadvantages. The methods based on signal processing do not require training and adjust operation parameters through human prior knowledge. This method avoids the large consumption of computing resources, but at the same time is highly dependent on the experience of professionals, and the adjusted parameters cannot be widely applied to various types of music, and a large amount of time and effort is required to adapt to another type of music again. The methods based on neural networks can minimize human intervention and let the computer learn autonomously. However, the existing music separation neural networks with good effects usually require dozens or even hundreds of layers of neural networks to learn parameters, which poses a great challenge to computing resources.
[0006] In terms of the types of neural networks, the methods based on neural networks are mainly divided into model methods based on convolutional neural networks (CNNs) and model methods based on recurrent neural networks (RNNs).
[0007] Compared with traditional methods, although these methods have made great progress in terms of performance and generalization ability, there are still some deficiencies. Compared with the low efficiency of RNNs, the CNN network has largely solved the problem of information transfer between layers. By expanding the receptive field, it gradually obtains global information at deeper levels. However, it is also restricted by the local receptive field characteristics of convolution. When dealing with large features in shallower neural networks, it cannot cover global information. Therefore, it is not sensitive to long-range dependencies and is prone to losing global information in feature calculations. The model methods based on recurrent neural networks are restricted by the inherent disadvantages of the model itself and will have the problem of forgetting for longer time-series data. The deficiencies of both limit the further improvement of the model effect.
[0008] In recent years, the attention mechanism architecture has gradually demonstrated excellent performance in multiple tasks such as computer vision and natural language processing. This architecture avoids the problem of the Markov process of the recurrent neural network architecture depending on the initial information, and at the same time can make up for the long-range deficiencies of CNNs.
[0009] The original LaSAFT network adds a single-layer attention and conditional vectors to the basic U-Net convolutional neural network to further separate sound sources, achieving progress compared with other similar architecture networks. However, there are also problems of large parameter quantity and redundancy in the separation process of the original LaSAFT network, and the running time is relatively long, making it difficult to run the neural network on edge computing devices such as mobile phones in real-time environments.
[0010] Therefore, generating better-quality audio separation results and reducing the parameter quantity and computational amount during the separation process is an urgent problem to be solved. Summary of the Invention
[0011] The present invention proposes a lightweight sound source separation method based on an improved attention mechanism, which is used to achieve music audio separation while generating better-quality audio separation results, superior to the original LASAFT network model in objective evaluation indicators, and controlling the parameter quantity and having better generalization during the separation process.
[0012] The present invention provides a lightweight sound source separation method based on an improved attention mechanism, including the following steps:
[0013] Construct a LaSHAFT network for performing sound source separation; among them, the LaSHAFT network is an improved LaSAFT network: replace the gated-point convolution fully connected calculation module in the original LaSAFT network with a global average pooling (GAP) layer to replace the fully connected layer, and replace the attention module in the LaSAFT network with a Hybrid-Voiceformer multi-head spectral hybrid attention module;
[0014] Among them, the Hybrid-Voiceformer multi-head spectral hybrid attention module includes a multi-head self-attention branch and a convolutional neural network (CNN) branch; perform a Concat connection operation on the results obtained after passing the multi-head self-attention branch and the convolutional neural network (CNN) branch through their respective convolutional branches, that is, obtain the output of the Hybrid-Voiceformer multi-head spectral hybrid attention module;
[0015] Use the constructed LaSHAFT network to separate the audio file to be separated, and obtain the sound source separation result of the audio.
[0016] Furthermore, the construction method of the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer includes the following steps:
[0017] Divide the features of the current layer into M l and M s two parts, segment the M l part and import it into the multi-head self-attention branch, and segment the M s and import it into the convolutional neural network (CNN) layer branch;
[0018] Among them, the network structure of the multi-head self-attention branch includes:
[0019] Two regularization layers, a multi-head self-attention mechanism (MSA) layer, and a multi-layer perceptron (MLP) layer;
[0020] Among them, the network structure of the convolutional neural network (CNN) layer branch includes:
[0021] A regularization layer, a convolutional layer with a 1×1 convolutional kernel size, and a LeakyReLU layer;
[0022] Perform a Concat connection operation on the results obtained after passing the network structure of the multi-head self-attention branch and the network structure of the CNN layer branch through their respective convolutional branches, then obtain the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer.
[0023] Furthermore, the composition function of the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer is as follows:
[0024] M 0 = M l + p
[0025]
[0026]
[0027] y l = LN(M l )
[0028] y s = LN(CNN(M s ))
[0029] y = Concat(y l , y s ) (1)
[0030] Among them, M is the given sound source feature, and its subscript n represents the feature of the nth layer;
[0031] MSA represents the multi-head self-attention mechanism; LN represents the layer normalization mechanism;
[0032] MLP represents the multi-layer perception mechanism; M 0 represents the song token input value set in the initial state; M l represents the token value output by the long-range information Transformer branch;
[0033] p represents the initial offset value, generally taking a small random number to avoid the denominator being 0;
[0034] M n represents the token value of the nth layer; represents the intermediate value after the operation of the multi-head self-attention mechanism MSA;
[0035] M s represents the token value output by the short-range information CNN layer branch;
[0036] y l represents the output of the Transformer branch; y s represents the output of the CNN layer branch; y represents the final output of the current layer.
[0037] Furthermore, the LaSHAFT network includes: one layer of two-dimensional convolution, three layers of downsampling blocks; three layers of upsampling blocks, and a conditional vector implemented by the GAP layer;
[0038] The corresponding levels of the downsampling block and the upsampling block are connected by skip connections;
[0039] The output of the conditional vector of the GAP layer is weighted and summed with the output of the upsampling block.
[0040] Further, the corresponding levels of the downsampling block and the upsampling block are connected by skip connections, including:
[0041] The skip connections connect the first downsampling block and the third upsampling block, the second downsampling block and the second upsampling block, and the third downsampling block and the first upsampling block respectively;
[0042] The results of upsampling and downsampling are added and output to the decoder through skip connections.
[0043] Further, the GAP layer includes a conditional vector generator and three layers of lightweight gated - improved point convolution modules;
[0044] Constructing a conditional vector implemented by the GAP layer and weighted - summing the output of the conditional vector of the GAP layer with the output of the upsampling block includes:
[0045] The three layers of lightweight gated - improved point convolution modules respectively output separated sound source conditional vectors through weighted summation;
[0046] The three layers of lightweight gated - improved point convolution modules are weighted and added to the feature maps of the corresponding levels of the three layers of upsampling blocks to the first layer, the second layer, and the third layer of the decoder.
[0047] Further, the formula for weighted - summing the output of the conditional vector of the GAP layer with the output of the upsampling block is:
[0048]
[0049]
[0050]
[0051]
[0052] where I represents the selection of the instrument conditional vector; is the j - th instrument feature of the i - th layer;
[0053] ⊙ represents the Hadamard operation; σ represents the Sigmoid activation function;
[0054] Equation (2) represents the output G obtained through depth - separable convolution operation of the current decoding layer features under the embedding of instrument I through the depth - separable convolution operation;
[0055] Equation (3) represents the output G in Equation (2) under the input of this feature, which is the output G' obtained through a 1×1 convolution operation by the current decoding layer feature through the 1×1 convolution operation;
[0056] Equation (4) represents the result G″ obtained by compressing the output G′ in Equation (3) into a single-channel feature;
[0057] Equation (5) represents the result of separating the features of the current layer obtained by passing the gating index G″ calculated in Equation (4) through the improved gating output mechanism GAP.
[0058] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0059] The present invention combines the good short-distance information transmission ability of the convolutional neural network CNN with the long-range information transmission ability of the Transformer framework, and improves the performance without increasing a large amount of computational complexity and the number of parameters compared with the CNN framework; the present invention uses a deep learning model to estimate the target audio source signal, which only requires data training compared with traditional methods, without introducing assumptions and relying on auxiliary information, and has better generalization; the present invention applies a GAP layer to strip irrelevant information in the output layer, further improving the separation quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0061] Figure 1 is a schematic diagram of the network structure of a lightweight audio source separation method based on an improved attention mechanism according to the present invention;
[0062] Figure 2 is a schematic diagram of the downsampling block in the present invention;
[0063] Figure 3 is a schematic diagram of the upsampling block in the present invention;
[0064] Figure 4 is a schematic diagram of the Hybrid-Voiceformer multi-head spectrum hybrid attention structure layer in the present invention;
[0065] Figure 5 is a schematic diagram of the GAP layer in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. However, it should be understood that the protection scope of the present invention is not limited by the specific embodiments.
[0067] Example 1
[0068] As Figure 1 shown, the present invention provides a lightweight sound source separation method based on an improved attention mechanism, including the following steps:
[0069] Step S1: Construct a LaSHAFT network for sound source separation; wherein, the LaSHAFT network is an improved LaSAFT network: replace the gated-point convolution fully connected calculation module in the original LaSAFT network with a GAP layer that replaces the fully connected layer with global pooling, and replace the attention module in the LaSAFT network with a Hybrid-Voiceformer multi-head spectrum hybrid attention module;
[0070] Among them, the Hybrid-Voiceformer multi-head spectrum hybrid attention module includes a multi-head self-attention branch and a convolutional neural network CNN branch; perform a Concat connection operation on the results obtained after passing the multi-head self-attention branch and the convolutional neural network CNN branch through their respective convolutional branches, that is, obtain the output of the Hybrid-Voiceformer multi-head spectrum hybrid attention module;
[0071] Step S2: Use the constructed LaSHAFT network to separate the audio file to be separated, and obtain the sound source separation result of the audio.
[0072] The present invention constructs a Hybrid-Voiceformer multi-head spectrum hybrid attention structure layer to capture key features in the sound source features; constructs a GAP layer at the decoder end to further separate the sound source and improve the effect of position encoding; uses the Hybrid-Voiceformer multi-head spectrum hybrid attention structure layer and the GAP layer to construct a LaSHAFT sound source separation network; uses the LaSHAFT sound source separation network to perform audio separation and give the audio separation result.
[0073] In step S2, the construction method of the Hybrid-Voiceformer multi-head spectrum hybrid attention structure layer includes the following steps:
[0074] Divide the features of the current layer into M l and M s two parts, segment the M l part and import it into the multi-head self-attention branch, and segment the M s and import it into the convolutional neural network CNN layer branch;
[0075] Among them, the network structure of the multi-head self-attention branch includes:
[0076] Two regularization layers, one multi-head self-attention mechanism (MSA) layer, and one multi-layer perceptron (MLP) layer;
[0077] Among them, the network structure of the convolutional neural network (CNN) layer branch includes:
[0078] One regularization layer, one convolutional layer with a 1×1 convolution kernel size, and one LeakyReLU layer;
[0079] Perform a Concat connection operation on the results obtained after passing the network structures of the multi-head self-attention branch and the CNN layer branch through their respective convolutional branches, then the Hybrid-Voiceformer multi-head spectrum hybrid attention structure layer is obtained.
[0080] Among them, the composition function of the Hybrid-Voiceformer multi-head spectrum hybrid attention structure layer is:
[0081] M 0 = M l + p
[0082]
[0083]
[0084] y l = LN(M l )
[0085] y s = LN(CNN(M s ))
[0086] y = Concat(y l , y s ) (1)
[0087] Among them, M is the given sound source feature, and its subscript n represents the nth layer feature;
[0088] MSA represents the multi-head self-attention mechanism; LN represents the layer normalization mechanism;
[0089] MLP represents the multi-layer perceptron mechanism; M 0 represents the song token input value set in the initial state; M l represents the token value output by the long-range information Transformer branch;
[0090] p represents the initial offset value, generally taking a small random number to avoid the denominator being 0;
[0091] M n represents the token value of the nth layer; Represents the intermediate value after the operation of the multi-head self-attention mechanism MSA;
[0092] M s Represents the token value output by the short-range information CNN layer branch;
[0093] y l Represents the output of the Transformer branch; y s Represents the output of the CNN layer branch; y represents the final output of the current layer.
[0094] The LaSHAFT network includes: one layer of two-dimensional convolution, three layers of downsampling blocks; three layers of upsampling blocks, and a conditional vector implemented by a GAP layer;
[0095] The corresponding layers of the downsampling block and the upsampling block are connected by skip connections;
[0096] The output of the conditional vector of the GAP layer and the output of the upsampling block are weighted and summed.
[0097] Among them, the corresponding layers of the downsampling block and the upsampling block are connected by skip connections, including:
[0098] The skip connections connect the first downsampling block and the third upsampling block, the second downsampling block and the second upsampling block, and the third downsampling block and the first upsampling block respectively;
[0099] The results of upsampling and downsampling are added and output to the decoder through skip connections.
[0100] Among them, the GAP layer includes a conditional vector generator and three layers of lightweight gated-improved point convolution modules;
[0101] Construct a conditional vector implemented by the GAP layer, and perform weighted summation of the output of the conditional vector of the GAP layer and the output of the upsampling block, including:
[0102] The three layers of lightweight gated-improved point convolution modules are respectively weighted and summed to output the separated sound source conditional vector;
[0103] The three layers of lightweight gated-improved point convolution modules are weighted and added to the feature maps of the corresponding layers of the three layers of upsampling blocks to the first layer, the second layer, and the third layer of the decoder.
[0104] The formula for weighted summation of the output of the conditional vector of the GAP layer and the output of the upsampling block is:
[0105]
[0106]
[0107]
[0108]
[0109] Among them, I represents the selection of the musical instrument condition vector; is the j-th musical instrument feature of the i-th layer;
[0110] ⊙ represents the Hadamard operation; σ represents the Sigmoid activation function;
[0111] Equation (2) represents the output G obtained through the depthwise separable convolution operation by the current decoding layer features under the embedding of the musical instrument I obtained through the depthwise separable convolution operation;
[0112] Equation (3) represents the output G' obtained through the 1×1 convolution operation by the current decoding layer features under the input of the feature of the output G in Equation (2) obtained through the 1×1 convolution operation;
[0113] Equation (4) represents the result G″ obtained by compressing the output G' in Equation (3) into a single-channel feature;
[0114] Equation (5) represents the result of separating the features of the current layer obtained by passing the gating index G″ calculated in Equation (4) through the improved gating output mechanism GAP.
[0115] In the present invention:
[0116] (1) Since the convolutional neural network CNN has advantages in obtaining short-range features, while the Transformer has advantages in obtaining long-range features, the combination of the two can integrate the advantages of both and enhance the performance without adding too many parameters;
[0117] The original multi-head attention uses the Pre-Norm operation, which performs regularization before introducing the features, increasing the parallelism of the neural network and being unfavorable for fast operation. The Post-Norm operation performs the equivalent of increasing the depth of the network after the feature operation, and also avoids too many branches in the network, accelerating the operation.
[0118] (2) The result of separating the features of the current layer is obtained through the improved gating output mechanism GAP layer.
[0119] The original gating mechanism facilitates the network to dynamically learn the features of the audio. However, the audio presents different features at different stages of the network. At the stage of the decoder feature map with a higher dimension, using the depthwise separable convolution can enable the network to learn the features more precisely. Different levels of training methods contribute to improving the training quality of the audio separation network.
[0120] The following is a description of the specific implementation manners in combination with specific embodiments.
[0121] 1. Construct the training set and the test set, including:
[0122] Allocate the training set and the test set from the Musdb18 dataset. Musdb18 is a dataset with 150 full-length music tracks. The dataset contains multi-track files of drums, bass, vocals, and other tracks of different genres. Randomly select 120 multi-track songs from Musdb18 as the training set, and the remaining 30 multi-track songs as the test set.
[0123] Select the training set and the test set from the DSD100 dataset. DSD100 is a dataset with 100 full-length music tracks. The dataset contains multi-track files of drums, bass, vocals, and other tracks of different genres. Randomly select 75 multi-track songs from Musdb18 as the training set, and the remaining 25 multi-track songs as the test set.
[0124] Mix the 120 multi-track songs selected from Musdb18 with the 75 multi-track songs selected from the DSD100 dataset to construct a training set with a total of 195 multi-track songs.
[0125] Mix the 30 multi-track songs selected from Musdb18 with the 25 multi-track songs selected from the DSD100 dataset to construct a test set with a total of 55 multi-track songs.
[0126] Use the librosa Python library function to pre-mix the 55 songs in the test set to obtain the finally used test data.
[0127] 2. As Figure 1 shown, the sound source first passes through a 2D convolution with a size of 1×2 to extract the local features of the sound source and adjust the number of channels of the sound source;
[0128] Secondly, extract the features of the song through a LeakyReLU function with α = 0.05, and at the same time retain a part of the negative sampling part of the audio to retain the details of the audio;
[0129] Then, pass through three downsampling blocks and three upsampling blocks respectively. There are skip connections at the corresponding levels of the downsampling blocks and the upsampling blocks to ensure that no information is lost during the long transmission of network information;
[0130] Finally, output the final features of the sound source and adjust the number of channels through a 2D convolution with a size of 1×2;
[0131] 3. The composition of the downsampling block and the upsampling block is as follows:
[0132] 3.1. As Figure 2As shown, the construction of the downsampling block includes: a 3×3 Time-Frequency Convolution (TFC) for expanding the receptive field and extracting local features of the sound source;
[0133] a Hybrid-Voiceformer multi-head spectral hybrid attention structure layer for further extracting the decoding features of the TFC layer;
[0134] a hierarchical regularization layer for regularizing the sound source feature map of the current layer for subsequent processing,
[0135] a LeakyReLU layer with α = 0.025 for extracting features of the song while retaining part of the negative sampling of the audio to preserve the details of the audio;
[0136] a downsampling layer for reducing the input to half of the original information;
[0137] 3.2. As Figure 3 shown, the construction of the upsampling block includes:
[0138] a 3×3 TFC convolution for expanding the receptive field and extracting local features of the sound source,
[0139] a hierarchical regularization layer for regularizing the sound source feature map of the current layer for subsequent processing,
[0140] a LeakyReLU layer with α = 0.025 for extracting features of the song while retaining part of the negative sampling of the audio to preserve the details of the audio,
[0141] a Hybrid-Voiceformer multi-head spectral hybrid attention structure layer for extracting important features of the TFC layer;
[0142] an upsampling layer for doubling the input to the original information;
[0143] Skip connections connect the first downsampling block and the third upsampling block, the second downsampling block and the second upsampling block, and the third downsampling block and the first upsampling block. The results of upsampling and downsampling are added and output to the decoder through the skip connections.
[0144] Among them, the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer, as Figure 5 shown.
[0145] The composition function of the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer is:
[0146]
[0147] 4. As Figure 5 shown, constructing a GAP layer to add a conditional vector at the output end and removing the attached hidden connections in the fully connected layer can accelerate the operation speed, reduce the number of parameters, and reduce the interference of irrelevant sound sources.
[0148] Based on the above benefits, compared with the original LaSAFT gating mechanism, this design effectively reduces the number of parameters, as shown in Table 1, where the higher the value, the better the effect.
[0149] Table 1 Comparison of Sound Source Separation Performance
[0150] Model Parameter calculation method GC <![CDATA[H×W×C in ×C out > GAP-GC <![CDATA[2×H×W×C in +C in ×C out >
[0151] Among them, H is the frequency information that is actually sound in sound operations,
[0152] W is the time information that is actually sound in sound operations,
[0153] C in is the number of feature input channels at the current level,
[0154] C out is the number of feature output channels at the current level,
[0155] Due to the condition that C in = C out under the output of the conditional vector, therefore, the ratio of GAP-GC to the original GC is at most only 17.4% of the number of parameters of the original GC version.
[0156] 5. Train the LaSHAFT network, and then regularize the training of the network through L1 loss and MSE, specifically including:
[0157] 5.1. Use the mean square error MSE to evaluate the pros and cons of sound source separation, including:
[0158] Normalize the data to limit the data range within [0, 1], and the calculation result range of MSE is also within [0, 1]; calculate the mean square error MSE, and the calculation formula is:
[0159]
[0160] Among them, n represents the number of samples, y predict represents the encoded value of the generated separated sound source, and y real represents the encoded value of the original sample sound source.
[0161] 5.2. Use the L1 loss to evaluate the pros and cons of sound source separation, and the calculation formula for calculating the L1 loss is:
[0162] L 1 = |y predict - y real | 1 (8)
[0163] 6. Test the trained network. Input the mixed sound source to obtain the separation result.
[0164] In this embodiment, the separation of the corresponding sound sources of vocals, bass, drums, and other musical instruments in the mixed dataset of the publicly available DSD100 and MUSDB18 datasets is completed. The verified metrics are to calculate the average values of the sound source separation quality of vocals, bass, drums, and other musical instruments respectively.
[0165] Table 3 shows the performance comparison of the method proposed in the present invention with other existing methods on the validation set after training on the mixed dataset of the DSD100 and MUSDB18 datasets. All the mentioned control groups have conducted experiments with the same number of rounds, validation set, and training set as the LaSHAFT network proposed in this paper. In Table 2, the higher the value, the better the effect.
[0166] Table 2 Sound Source Separation Performance Comparison (The higher the value, the better the effect)
[0167] Model Bass Drum Vocal Others DeepNMF 1.88 2.11 2.75 2.64 MMDenseNet 3.91 5.37 6.00 3.81 MMDenseNet-LSTM 3.73 5.46 6.31 4.33 LaSAFT 5.24 5.84 6.96 4.54 LaSHAFT 5.79 5.98 7.31 4.66
[0168] In summary, the present invention selects the publicly available Musdb18 and DSD100 datasets for mixing to obtain a new dataset, constructs a music separation network LaSHAFT based on lightweight attention, and the objective SDR sound quality evaluation index also proves that the generated separated melody is very close to the original melody, demonstrating the feasibility and effectiveness of this model in sound signal separation.
[0169] Finally, it should be noted that the above disclosure is only a specific embodiment of the present invention. However, the embodiments of the present invention are not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A lightweight sound source separation method based on an improved attention mechanism, characterized in that, it includes the following steps: Construct a LaSHAFT network for sound source separation; wherein, the LaSHAFT network is an improved LaSAFT network: replace the gated-point convolution fully connected calculation module in the original LaSAFT network with a global average pooling (GAP) layer that replaces the fully connected layer, and replace the attention module in the LaSAFT network with a Hybrid-Voiceformer multi-head spectral hybrid attention module; Among them, the Hybrid-Voiceformer multi-head spectral hybrid attention module includes a multi-head self-attention branch and a convolutional neural network (CNN) branch; perform a Concat connection operation on the results obtained after passing the multi-head self-attention branch and the CNN branch through their respective convolutional branches, that is, obtain the output of the Hybrid-Voiceformer multi-head spectral hybrid attention module. Use the constructed LaSHAFT network to separate the audio file to be separated, and obtain the sound source separation result of the audio.
2. The lightweight sound source separation method based on an improved attention mechanism according to claim 1, characterized in that: The construction method of the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer includes the following steps: Divide the features at the current level into M l and M s into two parts. After segmenting the M l part, import it into the multi-head self-attention branch, and after segmenting the M s part, import it into the convolutional neural network CNN layer branch; Among them, the network structure of the multi-head self-attention branch includes: Two regularization layers, one multi-head self-attention mechanism (MSA) layer, and one multi-layer perceptron (MLP) layer; Among them, the network structure of the CNN layer branch includes: One regularization layer, one convolutional layer with a 1×1 convolution kernel size, and one LeakyReLU layer; Perform a Concat connection operation on the results obtained after passing the network structure of the multi-head self-attention branch and the network structure of the CNN layer branch through their respective convolutional branches, then obtain the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer.
3. The lightweight sound source separation method based on an improved attention mechanism according to claim 2, characterized in that: The composition function of the Hybrid-Voiceformer multi-head spectral hybrid attention structure layer is: M 0 = M l + p y l = LN(M l ) y s = LN(CNN(M s )) y = Concat(y l , y s ) (1) Among them, M is the given sound source feature, and its subscript n represents the nth layer feature; MSA represents the multi-head self-attention mechanism; LN represents the layer normalization mechanism; MLP represents a multi - layer perception mechanism; M 0 represents the song token input value set in the initial state; M l represents the token value output by the long - range information Transformer branch; p represents the initial offset value, generally taking a small random number to avoid the denominator being 0; M n represents the token value of the nth layer; represents the intermediate value after the multi-head self-attention mechanism MSA operation; M s represents the token value output by the short-range information CNN layer branch; y l represents the output of the Transformer branch; y s represents the output of the CNN layer branch; y represents the final output of the current layer.
4. The lightweight sound source separation method based on an improved attention mechanism according to claim 1, characterized in that: The LaSHAFT network includes: one layer of two-dimensional convolution, three downsampling blocks; three upsampling blocks, and a conditional vector implemented by a GAP layer; The corresponding levels of the downsampling block and the upsampling block are connected through skip connections; Perform a weighted sum of the output of the conditional vector of the GAP layer and the output of the upsampling block.
5. The lightweight sound source separation method based on an improved attention mechanism according to claim 4, characterized in that: The corresponding levels of the downsampling block and the upsampling block are connected through skip connections, including: The skip connections are respectively connected to the first downsampling block and the third upsampling block, the second downsampling block and the second upsampling block, and the third downsampling block and the first upsampling block; The results of upsampling and downsampling are added and output to the decoder through the skip connection.
6. A lightweight sound source separation method based on an improved attention mechanism according to claim 5, characterized in that: The GAP layer includes a conditional vector generator and three layers of lightweight gated-improved point convolution modules; Constructing a conditional vector implemented by the GAP layer and performing weighted summation of the output of the conditional vector of the GAP layer and the output of the upsampling block, including: The three layers of lightweight gated-improved point convolution modules respectively output separated sound source conditional vectors through weighted summation; The three layers of lightweight gated-improved point convolution modules and the feature maps of the corresponding levels of the three layers of upsampling blocks are weighted and added to the first layer, the second layer, and the third layer of the decoder.
7. A lightweight sound source separation method based on an improved attention mechanism according to claim 6, characterized in that: The formula for performing weighted summation of the output of the conditional vector of the GAP layer and the output of the upsampling block is: Among them, I represents the selection of the instrument condition vector; is the j-th instrument feature of the i-th layer; ⊙ represents the Hadamard operation; σ represents the Sigmoid activation function; Equation (2) represents the output G obtained through depthwise separable convolution operation of the features of the current decoding layer with the musical instrument I embedded therein. Equation (3) represents the output G' obtained by a 1×1 convolution operation on the output G of the current decoding layer under the feature input of the feature input G in Equation (2). Equation (4) represents the result G″ obtained by compressing the output G′ in Equation (3) into a single-channel feature; Equation (5) represents obtaining the feature separation result of the current layer through the improved gated output mechanism GAP for the gated index G″ calculated by Equation (4).
Citation Information
Patent Citations
Single-channel speech enhancement method based on multi-scale information perception convolutional neural network
CN113936680A
Missing audio automatic restoration method based on U-Net model
CN114373469A