A method for directional speech separation based on a deep neural network
Through a deep neural network-based method, the time spectrum features and voiceprint vectors of speech are extracted, combined with the Mish activation function and the MSE loss function, the problem of insufficient characterization of voiceprint vectors is solved, and the performance of directional speech separation is improved, especially the separation effect in Chinese scenarios.
Patent Information
- Application Number
- CN202211622172.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-16
AI Technical Summary
The existing directional speech separation method lacks the ability to characterize the voiceprint vector, resulting in poor performance, and the Relu activation function does not smoothly around the zero value affects the system performance.
Using a deep neural network-based method, by extracting the time spectrum characteristics of speech, using a voiceprint encoder to extract delicate voiceprint vectors, a vocal separation network is constructed and a Mish activation function and MSE loss function are used, and a convolutional neural network, a long and short-term memory network and a full connection layer are trained, and the output mask matrix is separated.
The processing performance of directional speech separation is improved, especially the separation effect on Chinese data sets. Through delicate voiceprint information preservation and smooth network training, the rationality and separation effect of the mask matrix are improved.
Smart Images

Figure CN116030824B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech processing, and particularly relates to a method for directional speech separation based on a deep neural network. Background Art
[0002] Directional speech separation originates from the problem of single-channel speech separation. Single-channel speech separation aims to separate the mixed speech obtained through a single microphone to obtain signal channels equal in number to the number of speakers. Currently, the best-performing single-channel speech separation methods are basically based on deep learning, such as deep clustering methods, deep attractor networks, permutation invariant training, etc.
[0003] There are problems in single-channel speech separation such as the inability to know the number of speakers, permutation, and output channel selection. Therefore, directional speech separation based on voiceprint, also known as target speaker extraction, has emerged, which aims to extract the speech of a target speaker from the mixed speech using the voiceprint information of the target speaker. The most representative method among them is VoiceFilter, which inputs the extracted d-vector embedding vector into the extraction network. The performance of the directional voice separation system is related to the characterization ability of the voiceprint vector, and the voiceprint information extracted by the d-vector is not detailed enough. In addition, VoiceFilter adopts the Relu activation function, which is not smooth near zero and does not allow negative values, which also affects the performance of the directional voice separation system. Summary of the Invention
[0004] The present invention provides a method for directional speech separation based on a deep neural network, which can be used to improve the processing performance of directional voice separation.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for directional speech separation based on a deep neural network, the method comprising:
[0007] Step 1, extracting the time-frequency spectrum features of the speech: extracting the time-frequency spectrum features of the mixed human voice and the pure human voice, using the time-frequency spectrum features of the mixed human voice as the input time-frequency spectrum, using the time-frequency spectrum features of the pure human voice as the target time-frequency spectrum, and the target time-frequency spectrum is used for the training of the human voice separation network;
[0008] Step 2, using a voiceprint encoder to extract the voiceprint vector of the speech to extract the voiceprint vector of a reference human voice different from the pure human voice;
[0009] Step 3, constructing and training a human voice separation network, wherein, the activation function of the human voice separation network adopts the Mish function;
[0010] Step 4, jointly inputting the time-frequency spectrum and the voiceprint vector into the human voice separation network, and outputting the target human voice extracted from the mixed human voice.
[0011] Further, in step 3, the voice separation network sequentially includes: a convolutional neural network, a long short-term memory network LSTM, and a fully connected layer; input the voiceprint vector of the reference voice different from the pure voice extracted by the voiceprint encoder into the LSTM network of the voice separation network; and input the extracted input time-frequency spectrum into the convolutional neural network of the voice separation network; the convolutional neural network then inputs the extracted feature map into the LSTM network, and the fully connected layer is used to output the mask matrix of the spectrogram;
[0012] Multiply the mask matrix output by the voice separation network by the input time-frequency spectrum to obtain the output time-frequency spectrum, and obtain the target voice extracted from the mixed voice through the inverse process of time-frequency spectrum feature extraction;
[0013] When training the voice separation network, the mean square loss is used to update the network parameters, that is, the mean square loss of the voice separation network is calculated based on the target time-frequency spectrum and the output time-frequency spectrum, and then the network parameters of the voice separation network are updated through backpropagation and the gradient descent method until the number of training times reaches the upper limit or the mean square loss meets the specified conditions and then stop.
[0014] The technical solution provided by the present invention at least brings the following beneficial effects:
[0015] By introducing the ecapa-vector voiceprint feature, the present invention extracts more delicate voiceprint information, so that more features of the reference voice are retained, which is beneficial to separating the voice containing the above features from the mixed voice; by introducing the Mish activation function to replace Relu, the training of the neural network of the present invention is smoother, so as to extract a more reasonable mask matrix; by introducing the MSE loss function, the difference between the output time-frequency spectrum and the target time-frequency spectrum is better expressed. In addition, the present invention is trained on a Chinese dataset, which is more suitable for processing Chinese voice separation scenarios. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic diagram of the processing process of a directional voice separation method based on a deep neural network provided by an embodiment of the present invention;
[0018] Figure 2 It is a schematic diagram of the structure of the voiceprint encoder used in the embodiment of the present invention. Detailed Embodiments
[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0020] A method for directional voice separation based on a deep neural network provided by an embodiment of the present invention includes the following steps:
[0021] Step 1: Extract the time-frequency spectrum features of the voice: Extract the time-frequency spectrum features of the mixed human voices and the pure human voices, where the time-frequency spectrum features of the pure human voices are used for the training of the voice separation network;
[0022] Step 2: Use a voiceprint encoder (preferably the EcapaTDNN network) to extract the voiceprint vector of the voice to obtain an ecapa-vector, that is, extract the voiceprint vector of a reference human voice different from the pure human voice;
[0023] Step 3: Construct and train a voice separation network, where the activation function of the voice separation network uses the Mish function;
[0024] Step 4: Input the time-frequency spectrum and the voiceprint vector into the voice separation network together, and output the target human voice extracted from the mixed human voices.
[0025] As a possible implementation, in the embodiment of the present invention, extracting the time-frequency spectrum features of the voice specifically is:
[0026] During the process of extracting the frequency-domain feature Fbank, the audio is divided into frames of equal length. Each frame obtains a spectrogram through the discrete Fourier transform. Representing the amplitude with gray values, the spectral representation can be converted from two-dimensional to one-dimensional. Concatenating the amplitude gray representations of all audio frames, with the time-sequential frames as the abscissa and the frequency as the ordinate, a time-frequency spectrogram is obtained. The above process is the short-time Fourier transform (STFT). If it is necessary to recover the original audio from the time-frequency spectrum, the inverse short-time Fourier transform (ISTFT) can be used. Among them, Fbank is a front-end processing method that processes the audio in a way similar to the human ear and can improve the performance of speech recognition.
[0027] In the embodiment of the present invention, the processing process of voice separation is based on Google's VoiceFilter architecture, and its improved structure is as Figure 1As shown in the figure, it includes: a voiceprint encoder, an extraction network (voice separation network), a short-time Fourier transform (time-frequency spectrum feature extraction unit), and an inverse short-time Fourier transform (time-frequency spectrum feature conversion unit). Among them, the extraction network sequentially includes a convolutional neural network, a long short-term memory network LSTM, and a fully connected layer, and is used to output a mask matrix of the spectrogram; the voiceprint encoder is used to extract the voiceprint vector of the reference voice different from the pure voice, and input the extracted voiceprint vector of the reference voice into the LSTM network of the extraction network. The input of the convolutional neural network of the extraction network is the frequency domain feature (input time-frequency spectrum) of the mixed voice obtained by the short-time Fourier transform. The extracted feature map is input into the LSTM network, and the fully connected layer is used to output the mask matrix of the spectrogram; multiplying the mask matrix by the frequency domain feature obtained by the short-time Fourier transform to obtain the output time-frequency spectrum, and after the inverse short-time Fourier transform, the target voice extracted from the mixed voice is obtained.
[0028] When training the voice separation network of the present invention, the MSE loss is used to update the network parameters. The specific calculation method of the MSE loss (mean square loss) is as follows: input the pure voice into the time-frequency spectrum feature extraction unit to obtain the target time-frequency spectrum, and calculate the MSE loss of the voice separation network based on the target time-frequency spectrum and the output time-frequency spectrum. For a matrix with dimensions of n×m, the MSE calculation is: where, and y ij represent the predicted value and the true value respectively.
[0029] That is, during training, the input of the system is the mixed voice and two different voices of a specific speaker. One of the voices is the pure voice of the speaker in the mixed voice, and the other voice is used as the reference voice to extract the voiceprint information. During inference, the system inputs the mixed voice and the reference voice and outputs the extracted voice.
[0030] The voiceprint information encoder used is the EcapaTDNN network in the voiceprint recognition system, and the extracted ecapa-vector is used as the input. Among them, the structure of the EcapaTDNN network is as Figure 2 shown, and sequentially includes: a convolutional block (one-dimensional convolution, activation function, and batch normalization (BN)), three layers of SE-Res2Block, a convolutional block, an attention statistics module (attention statistics pool + BN), a fully connected layer with BN (FC), and a Softmax layer (using the improved AAM-Softmax). The fully connected layer with BN and the Softmax layer constitute the output layer of the EcapaTDNN network. Figure 2Among them, T represents the number of audio frames, C represents the number of channels; S represents the number of speakers. The EcapaTDNN network introduces Res2Net (a multi-scale backbone network) in the network structure part to reduce the number of parameters; adds the SE (Squeeze-and-Excitation) network; introduces a feature aggregation and accumulation mechanism. The pooling function introduces a channel attention mechanism; the loss function uses an improved version of Softmax, AAM-Softmax.
[0031] The activation function of the neural network adopted in the present invention is the Mish activation function (in the convolutional neural network), and its definition is:
[0032] f(x) = x · tanh(ln(1e x ))
[0033] Among them, tanh represents the hyperbolic tangent function, and e represents the natural base.
[0034] The present invention extracts more delicate voiceprint information by introducing ecapa-vector voiceprint features, so that more features of the reference human voice are retained, which is beneficial to separating the human voice containing the above features from the mixed human voices; by introducing the Mish activation function to replace Relu, the training of the neural network of the present invention is smoother, so as to extract a more reasonable mask matrix; by introducing the MSE loss function, the difference between the output time-frequency spectrum and the target time-frequency spectrum is better expressed. In addition, the present invention is trained on a Chinese dataset, which is more suitable for processing Chinese human voice separation scenarios. Through these optimization means, the separation method constructed by the present invention has achieved better results in the evaluation indicators commonly used in sound source separation systems such as SDR and Si-SDR, specifically manifested as larger values when tested on datasets such as Aishell-1 and Aidatatang_200zh. This shows that the present invention effectively improves the benchmark separation system and is a better directional human voice separation method.
[0035] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0036] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the creative concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A directional speech separation method based on a deep neural network, characterized in that It includes the following steps: Step 1, extracting the time-frequency spectrum features of speech: Extract the time-frequency spectrum features of the mixed human voice and the pure human voice. Take the time-frequency spectrum features of the mixed human voice as the input time-frequency spectrum, and take the time-frequency spectrum features of the pure human voice as the target time-frequency spectrum, and the target time-frequency spectrum is used for the training of the human voice separation network; Step 2, using a voiceprint encoder to extract the voiceprint vector of the speech to extract the voiceprint vector of a reference human voice different from the pure human voice; wherein, the voiceprint encoder is an EcapaTDNN network; Step 3, constructing and training a human voice separation network, wherein, the activation function of the human voice separation network adopts the Mish function; Step 4, inputting the time-frequency spectrum and the voiceprint vector into the human voice separation network together, and outputting the target human voice extracted from the mixed human voice; In the said Step 3, the human voice separation network successively includes: a convolutional neural network, a long short-term memory network LSTM and a fully connected layer; input the voiceprint vector of the reference human voice different from the pure human voice extracted by the voiceprint encoder into the LSTM network of the human voice separation network; and input the extracted input time-frequency spectrum into the convolutional neural network of the human voice separation network; the convolutional neural network then inputs the extracted feature map into the LSTM network, and the fully connected layer is used to output the mask matrix of the spectrogram; Multiply the mask matrix output by the human voice separation network by the input time-frequency spectrum to obtain the output time-frequency spectrum, and obtain the target human voice extracted from the mixed human voice through the inverse process of time-frequency spectrum feature extraction; When training the human voice separation network, the mean square loss is used to update the network parameters, that is, calculate the mean square loss of the human voice separation network based on the target time-frequency spectrum and the output time-frequency spectrum, and then update the network parameters of the human voice separation network through backpropagation and the gradient descent method until the number of training times reaches the upper limit or the mean square loss meets the specified conditions and then stop.
2. The method according to claim 1, characterized in that, in Step 1, the specific way of extracting the time-frequency spectrum features of the speech is: Divide the audio into frames of equal length, and each frame obtains a spectrogram through discrete Fourier transform, and the amplitude is represented by gray value; Stitch the amplitude gray representations of all audio frames, with the frames in time sequence as the abscissa and the frequency as the ordinate, to obtain the time-frequency spectrogram of the audio.
Citation Information
Patent Citations
Human voice separation method, electronic equipment and readable storage medium
CN115132221A