Iterative Fundamental Frequency Estimation and Speech Separation Method and Device Based on a Bidirectional Cascade Framework

The dual-loop iterative framework enhances base frequency extraction and speech separation by iteratively updating these processes, addressing the limitations of existing methods in complex environments and improving accuracy and separation performance.

CN115862659BActive Publication Date: 2025-07-15PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211184250.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-07-15
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

The existing fundamental frequency extraction and speech separation methods are insufficient in complex acoustic environments, especially in mixed speech scenarios of multiple speakers, which are difficult to accurately extract fundamental frequency and separate speech. Traditional methods cannot effectively process the unvoiced part, and the existing tandem algorithms have poor performance when the speech of different speakers overlaps.

Method used

Using the iterative method of the bidirectional cascade framework, the closed-loop cycle process of fundamental frequency extraction and speech separation is established through the loop iteration of the fundamental frequency prediction module, the speech separation module and the fundamental frequency update module. The conditions are used to generate an adversarial network and a convolutional neural network for fundamental frequency prediction and speech separation, and the fundamental frequency prediction value is updated frame by frame to improve accuracy.

Benefits of technology

It significantly improves the performance of fundamental frequency extraction and speech separation, can accurately separate the voice of multiple speakers in complex acoustic environments, and improves the effects of technologies such as speech recognition and smart speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862659B_ABST
    Figure CN115862659B_ABST
Patent Text Reader

Abstract

The present invention discloses an iterative fundamental frequency extraction and speech separation method and device based on a bidirectional cascade framework, which iteratively performs "fundamental frequency prediction - speech separation - fundamental frequency update" frame by frame on the mixed speech, and improves the performance of both in the iteration. The fundamental frequency prediction module provides fundamental frequency clues for the subsequent modules, solving the permutation problem caused by multiple outputs and the problem of uncertain number of speakers. The speech separation module uses a conditional generative adversarial network for generative speech separation to improve the quality of the separated speech. The fundamental frequency update module reextracts the fundamental frequency from the separated clean speech and updates the prediction value of the fundamental frequency prediction, realizing the closed-loop of the "prediction - separation - update" process. Under the bidirectional cascade framework proposed by the present invention, the two tasks of speech separation and fundamental frequency extraction are alternately updated by an iterative method, depending on and promoting each other, and both tasks have achieved better performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech signal processing, and relates to fundamental frequency extraction technology and speech separation technology. Specifically, it relates to an iterative fundamental frequency estimation and speech separation method and device based on a bidirectional cascade framework. Background Art

[0002] When a person's vocal organs produce voiced sounds, the vocal cords vibrate periodically. The fundamental frequency is determined by the frequency of the vocal cord vibration and characterizes the pitch of the sound. The frequency components of a speech signal usually consist of a fundamental frequency and a series of harmonics, and the harmonics are integer multiples of the fundamental frequency. This characteristic is called "harmonicity". The fundamental frequency extraction task aims to extract the corresponding fundamental frequency trajectory from a person's speech. For the fundamental frequency extraction of a single speaker, traditional methods have achieved good performance. However, for the fundamental frequency extraction of multiple speakers in a complex acoustic environment, there is still no sufficiently accurate fundamental frequency extraction method. The speech separation task aims to enable a machine to extract the speech of a specific speaker in a complex acoustic scene and ignore background interference sounds. Similarly, the multi-speaker speech separation task in a complex environment is also a key issue of concern. Both are of great significance for the field of speech signal processing, such as speech recognition, keyword wake-up, smart speakers and other technologies.

[0003] The existing cascaded algorithm can sequentially implement fundamental frequency estimation and speech separation. This cascaded algorithm first estimates the fundamental frequency of the time-frequency segments with high inter-band correlation in the input mixed speech, then concatenates the frame-level fundamental frequencies belonging to the same speaker using the principle of time continuity, and then uses the estimated fundamental frequency for the separation of the corresponding speech. However, the part of fundamental frequency estimation is only based on traditional signal processing methods, such as autocorrelation calculation and envelope-based feature extraction, which are too simple among the methods for extracting signal periodicity. When there are many overlapping segments of speech from different speakers in the input mixed speech, the performance of this type of method is poor. In addition, the above algorithm can only separate the pronunciation segments because only these segments have fundamental frequency values. For the voiceless part, the algorithm cannot extract the fundamental frequency and thus cannot perform further separation.

[0004] There are also methods that use speech separation as the front end of multi-speaker fundamental frequency extraction to improve the performance of the latter. The results show that the cascaded method of "speech separation - fundamental frequency extraction" can improve the performance of fundamental frequency extraction compared to only using mixed speech as the input. However, most of the existing methods aim to solve the speech enhancement problem and are for the mixed signal of speech and non-speech noise, that is, the input only contains the fundamental frequency of a single speaker and environmental noise. Obviously, the fundamental frequency extraction and speech separation of multi-speaker mixed speech targeted by the present invention are more challenging tasks.

[0005] Existing research shows that a voice separation system can serve as the front end of a multi-speaker fundamental frequency extraction task, effectively improving the performance of the latter. Conversely, a multi-speaker fundamental frequency extraction system can also serve as the front end of a voice separation task, significantly enhancing the separation effect of the latter. Generally speaking, accurate fundamental frequency extraction depends on the prior separation of speech, while the performance improvement of voice separation benefits from the accurately extracted fundamental frequency. Therefore, the two are in a mutually dependent and mutually promoting relationship, and they are in a closed-loop cycle relationship. Summary of the Invention

[0006] Aiming at the existing problems and centering around the mutually dependent and promoting relationship between fundamental frequency extraction and voice separation, the present invention proposes a bidirectional cascaded iterative "predict fundamental frequency - separate speech - update fundamental frequency" framework, enabling these two tasks to be alternately updated in an iterative manner, playing a role of mutual dependence and promotion, modeling the closed-loop cycle process, and achieving better performance for both tasks.

[0007] The technical solution of the present invention is as follows:

[0008] An iterative fundamental frequency estimation and voice separation method based on a bidirectional cascaded framework, the steps of which include:

[0009] 1) For the given mixed speech, perform frame segmentation, windowing, and short-time Fourier transform operations in sequence to obtain the mixed speech spectrogram, and then iteratively execute steps 2) to 4) frame by frame until all frames are traversed;

[0010] 2) Use the fundamental frequency prediction module to predict the fundamental frequency value at the current moment based on the mixed speech at the current moment and the fundamental frequency prediction value at the historical moment, so as to extract the fundamental frequency sequence of a certain speaker in the mixed speech;

[0011] 3) Use the voice separation module to take the mixed speech and the fundamental frequency sequence obtained in step 2) as inputs, and generate the speaker speech corresponding to the fundamental frequency sequence through a conditional generative adversarial network;

[0012] 4) Use the fundamental frequency update module to take the separated speaker speech generated in step 3) as the input, extract the fundamental frequency trajectory, and use it to update the fundamental frequency prediction value of the current frame output in step 2);

[0013] 5) After the loop described in step 1), the complete fundamental frequency trajectory of a certain speaker in the mixed speech can be obtained. This speaker is determined by the leading speaker in the mixed speech. Taking this fundamental frequency trajectory as the fundamental frequency condition and combining it with the mixed speech spectrogram, inputting them into the conditional generative adversarial network in step 3), the speaker speech corresponding to the fundamental frequency condition can be separated;

[0014] 6) Subtract the speaker voice separated in step 5) from the mixed voice, and perform the iterative process of steps 1) to 5) on the residual voice again. Repeat this process until there is no voice in the residual voice, and then the above loop process stops, so as to separate the voices of each speaker in the mixed voice.

[0015] Further, the fundamental frequency prediction module includes an encoder, a prediction network, and a joint network. The encoder uses a two-layer two-dimensional convolutional neural network followed by a four-layer bidirectional long short-term memory network. The prediction network adopts a two-layer long short-term memory network. The joint network is a one-layer fully connected network. The entire framework is trained and jointly optimized simultaneously, and its optimization objective is the categorical cross-entropy loss.

[0016] Further, the conditional generative adversarial network consists of a generator and a discriminator. The generator aims to generate a time-domain signal corresponding to the fundamental frequency condition from the magnitude spectrum of the mixed voice. The discriminator consists of multiple discriminators acting on different frequency bands of the voice signal.

[0017] Further, the fundamental frequency update module aims to extract the fundamental frequency from the output result of the voice separation module and use it to update the output result of the fundamental frequency prediction module. The fundamental frequency update module uses a convolutional neural network to model the local characteristics of the input spectrum, capture the harmonic structure between frequency components, and then connects to a fully connected layer to model the mapping relationship between the harmonics and the fundamental frequency of each frame. The optimization objective is the categorical cross-entropy loss function.

[0018] Further, the present invention adopts a frame-by-frame iterative "fundamental frequency prediction - voice separation - fundamental frequency update" framework, which cascades the fundamental frequency extraction and voice separation tasks bidirectionally, and improves the performance of both simultaneously. The rules for its loop iteration process are as follows:

[0019] For a given mixed voice, each time the framework runs, the separated voice of a certain speaker will be output. This speaker is determined by the leading speaker in the mixed voice. Subtract the speaker voice separated in the previous round from the mixed voice, and perform the above iterative process on the residual voice again. Repeat this process until there is no voice in the residual voice, that is, once no fundamental frequency value is predicted in the remaining signal, the loop process stops; or, the loop stop condition is determined by the energy value of the remaining signal. If the energy is less than a certain threshold, the stop condition is met.

[0020] An iterative fundamental frequency estimation and voice separation device based on a bidirectional cascade framework, which includes a signal preprocessing module, a fundamental frequency prediction module, a voice separation module, a fundamental frequency update module, and a loop separation module;

[0021] The signal preprocessing module is used to perform frame division, windowing, and short-time Fourier transform operations on the given mixed voice in sequence to obtain the time-frequency spectrum of the mixed voice;

[0022] The fundamental frequency prediction module is used to predict the fundamental frequency value at the current moment based on the mixed speech frame at the current moment and the fundamental frequency prediction value at the historical moment, and extract the fundamental frequency sequence of a certain speaker in the mixed speech;

[0023] The speech separation module is used to take the mixed speech and the fundamental frequency sequence of a certain speaker obtained by the fundamental frequency prediction module as inputs, and use a conditional generative adversarial network to generate the speaker speech corresponding to the fundamental frequency sequence;

[0024] The fundamental frequency update module is used to take the generated speaker speech as an input, extract the fundamental frequency trajectory, and use it to update the fundamental frequency prediction value of the current frame output by the fundamental frequency prediction module;

[0025] The loop separation module is used to run the fundamental frequency prediction module, the speech separation module, and the fundamental frequency update module for a given mixed speech to obtain the separated speech of a certain speaker, which is determined by the leading speaker in the mixed speech. Subtract the speaker speech that has been separated in the previous round from the mixed speech, and perform the iterative process of the fundamental frequency prediction module, the speech separation module, and the fundamental frequency update module on the residual speech again. Repeat this process until there is no speech in the residual speech, and then the loop process stops, so as to separate the speech of each speaker in the mixed speech.

[0026] Compared with the prior art, the positive effects of the present invention are as follows:

[0027] This method integrates the two tasks of fundamental frequency extraction and speech separation into one framework in an iterative manner, aiming to characterize the interdependent and mutually promoting relationship between the two. Specifically, the present invention proposes a bidirectional cascade framework to iteratively jointly optimize the two tasks of fundamental frequency extraction and speech separation. Compared with the unidirectional cascaded framework in the prior art, this framework can significantly improve the performance of these two tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the process framework of the present invention;

[0029] Figure 2 It is a framework diagram of the fundamental frequency prediction module used in the present invention;

[0030] Figure 3 It is a schematic diagram of the loop separation process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0031] The specific embodiments of the present invention will be described in more detail below.

[0032] The present invention proposes an iterative framework for simultaneously solving fundamental frequency estimation and speech separation, which can be summarized as a process of "predicting fundamental frequency - separating speech - updating fundamental frequency", as Figure 1As shown in the figure, it includes three modules: fundamental frequency prediction, voice separation, and fundamental frequency update. Specifically, the fundamental frequency prediction module aims to predict the fundamental frequency at the current moment using the mixed speech at the current moment and the updated fundamental frequency at the previous moment. The voice separation module takes the mixed speech as input, and uses the predicted fundamental frequency at the current moment and the updated fundamental frequencies at all historical moments as conditions, and uses conditional generative adversarial networks to generate the separated speech up to the current moment. The fundamental frequency update module takes the output result of the voice separation module as input, extracts the fundamental frequency trajectory, and uses this fundamental frequency trajectory to update the predicted value of the fundamental frequency at the current moment output by the fundamental frequency update module. In this way, the fundamental frequency of the current frame is predicted, separated, and updated iteratively and frame by frame, and finally a sentence-level updated fundamental frequency trajectory is obtained, and the separated speech of the corresponding speaker is generated. In this framework, the two tasks form a bidirectional cascaded relationship, promoting and depending on each other, and achieving better performance.

[0033] An iterative fundamental frequency estimation and voice separation method based on a bidirectional cascaded framework according to an embodiment of the present invention, the steps of which include:

[0034] 1) For the given mixed speech, perform frame division, windowing, and short-time Fourier transform operations in sequence to obtain the mixed speech spectrogram, and then loop through steps 2) to 4) frame by frame until all frames are traversed. The loop framework is as Figure 1 shown.

[0035] 2) Use the fundamental frequency prediction module to predict the fundamental frequency value at the current moment based on the mixed speech at the current moment and the predicted value of the fundamental frequency at the historical moment, so as to extract the fundamental frequency sequence of a certain speaker in the mixed speech. The framework of the fundamental frequency prediction module is as Figure 2 shown.

[0036] 3) Use the voice separation module to take the mixed speech and the fundamental frequency sequence obtained in step 2) as input, and generate the speaker speech corresponding to the fundamental frequency sequence through conditional generative adversarial networks.

[0037] 4) Use the fundamental frequency update module to take the separated speaker speech generated in step 3) as input, extract the fundamental frequency trajectory, and use it to update the predicted value of the fundamental frequency of the current frame output in step 2).

[0038] 5) After the loop described in step 1), a complete fundamental frequency trajectory of a certain speaker in the mixed speech can be obtained, and this speaker is determined by the leading speaker in the mixed speech. Taking this fundamental frequency trajectory as the fundamental frequency condition and combining it with the mixed speech spectrogram, input it into the conditional generative adversarial network in step 3), and the speaker speech corresponding to the fundamental frequency condition can be separated.

[0039] 6) Subtract the speaker voice separated in step 5) from the mixed voice, and perform the iterative process described in steps 1)-5) on the residual voice again. Repeat this cycle until there is no voice in the residual voice, and the above cycle process stops, thereby separating the voices of each speaker in the mixed voice. The cycle separation process is as Figure 3 shown.

[0040] The specific implementation steps of the method of the present invention include several parts such as signal preprocessing, fundamental frequency extraction, voice separation, fundamental frequency update, and cycle separation. The specific implementation processes of each step are as follows:

[0041] 1. Voice signal preprocessing

[0042] This method first performs short-time Fourier transform (STFT) on the mixed voice as the subsequent input, and performs short-time Fourier transform on the signal using the analysis window w(n), window length N, and frame shift R. The transformation formula is as follows:

[0043]

[0044] Among them, X(t,f) represents the value corresponding to the f-th frequency band of the t-th frame of the spectrum X, x(n) represents the value corresponding to the n-th sampling point of the signal x, n represents the index of the signal sampling point, and t and f respectively represent the indices of the frame and the frequency band. After the transformation, the STFT spectrum is obtained. In the specific implementation, the frame length used is 32 ms, the frame shift is 16 ms, and the window function is the Hamming window. The following steps of "fundamental frequency prediction - voice separation - fundamental frequency update" will be iteratively performed on each frame of the transformed mixed voice spectrum.

[0045] 2. Fundamental frequency prediction

[0046] The fundamental frequency prediction module predicts the fundamental frequency value at the current moment based on the mixed voice at the current moment and the fundamental frequency prediction value at the historical moment. The function of this module can be expressed by the following formula (taking the current moment t as an example):

[0047]

[0048] Among them, represents the prediction result of the fundamental frequency at moment t, Prediction represents the fundamental frequency prediction module, x t represents the mixed voice of the t-th frame, and p t-1 represents the fundamental frequency prediction result obtained at moment t-1.

[0049] The framework of this module is as Figure 2As shown, it mainly consists of an encoder, a prediction network, and a joint network. Among them, the prediction network and the joint network can be regarded as a whole and act as a decoder. The encoder adopts the network structure of a recurrent neural network (RNN) to map the input mixed speech x of the t-th frame t to a higher-dimensional representation which not only depends on the input at the current moment but also on the hidden layer output at the previous moment, and can be said to depend on the entire historical input sequence x0, …, x t .

[0050]

[0051] Among them, represents the high-dimensional representation of the t-th frame obtained by encoding the input sequence, and f enc represents the encoder.

[0052] In the traditional framework based on recurrent networks, the output at the current moment is independent of the outputs at historical moments. In the present invention, by introducing a prediction network, the output at the current moment also depends on the outputs at historical moments. Specifically, the prediction network uses the prediction result at the previous moment as the input to decode the hidden layer output

[0053]

[0054] at the current moment, where f dec represents the decoder.

[0055] Next, a joint network is used to integrate the outputs of the above-mentioned encoder and prediction network to obtain the prediction result of the fundamental frequency at the current moment, which is a joint probability distribution conditional on the mixed speech at the current moment and the fundamental frequency prediction at the previous moment The joint network here consists of several feed forward layers.

[0056]

[0057] Among them, represents the hidden layer representation vector output by the joint network at time t, and f joint represents the joint network.

[0058] The final output probability distribution is obtained through a softmax layer:

[0059]

[0060] The entire framework including the encoder, prediction network, and joint network is trained simultaneously and jointly optimized, with the optimization objective being the categorical cross-entropy loss:

[0061]

[0062] where \(t\) is the frame index, \(s\) is the index of the frequencies corresponding to 68 fundamental frequencies, \(O\) is the linear output layer for 68 classifications, and \(O\) t (s) represents the probability of the \(m\)-th frame magnitude spectrum corresponding to the \(s\)-th frequency value. A softmax activation function is connected after this linear layer. For a given \(x\) t and input, it is the posterior probability of falling into the \(s\)-th frequency.

[0063] In specific implementation, first, the operation of framing the mixed speech is carried out. The STFT spectra of 7 consecutive frames (3 - 1 - 3) are concatenated and input into the encoding network to encode the output of the middle frame. This encoding network first uses a 2 - layer two - dimensional convolutional neural network (CNN) with a convolutional kernel size of 6×6, followed by a 4 - layer bidirectional long short - term memory network layer (BLSTM), where the number of hidden nodes in each cell is 256. In the prediction network at the input end of the historical fundamental frequency, a 2 - layer unidirectional LSTM layer is adopted, where the number of hidden nodes in each cell is 512. The joint network is a one - layer fully - connected layer FCN with 512 hidden nodes.

[0064] 3. Speech Separation

[0065] The speech separation module takes the mixed speech sequence and the fundamental frequency sequence as inputs and uses a conditional generative adversarial network to generate the speaker speech corresponding to the fundamental frequency. The function of this module can be expressed by the following formula (taking the current moment \(t\) as an example):

[0066]

[0067] where \(y\) 0~t represents the sequence of the target speaker's speech signal from time 0 to time \(t\), Separation represents the speech separation module, \(x\) 0~t represents the sequence of the mixed speech signal from time 0 to time \(t\), and \(p\) 0~(t-1) represents the sequence of the target speaker's fundamental frequency from time 0 to time \(t\).

[0068] Specifically, given the mixed speech sequence of the entire sentence \((x_0,\ldots,x\) N ) and the sequence of the speaker's fundamental frequency output by the "fundamental frequency prediction module", the generator outputs the speech corresponding to this speaker \((y_0,\ldots,y\) N), where the fundamental frequency values from the 0th frame to the tth frame at the current moment are replaced by the prediction results of the fundamental frequency prediction module, and N is the total number of frames.

[0069] This module mainly consists of two parts: a generator and a discriminator. The generator aims to generate a time-domain signal corresponding to the fundamental frequency condition from the magnitude spectrum of the mixed speech. Specifically, it includes two stages: the first stage aims to generate the magnitude spectrum of the speaker corresponding to the fundamental frequency condition from the magnitude spectrum of the input mixed speech; the second stage is to upsample the output magnitude spectrum of the first stage to the dimension of the corresponding time-domain signal using a set of stacked transposed convolution modules and one-dimensional convolutions. The upsampling factor is determined by the frame shift (hop size) of the input magnitude spectrum. After each transposed convolution module, there will be a residual block, which consists of three layers of one-dimensional convolutions with dilation. By setting different dilation coefficients (1, 3, 9), a receptive field of size 27 can be obtained to expand the ability to sense the signal in the time dimension and better model the long-range dependencies in the time dimension. The generator part finally uses a layer of one-dimensional convolution and the tanh activation function to output a representation that satisfies the numerical range of the time-domain signal. The output channel of this one-dimensional convolution is set to 1, corresponding to the time-domain signal of the full frequency band.

[0070] In the discriminator part, a multi-scale strategy is adopted, that is, multiple discriminators are used instead of the traditional single discriminator for true / false discrimination. Previous studies have shown that the speech generated using only a single discriminator will have metallic noise. Since speech signals have different spectral characteristics in different frequency ranges, therefore, for different frequency bands, different discriminators will be used. Here, the concept of "multi-scale" refers to different frequency bands. Specifically, the multi-scale discriminators will share the same set of network structures and parameters, but act on different frequency bands of the speech signal. When using K discriminators, the training objectives of the discriminator and the generator are respectively:

[0071] Discriminator:

[0072] Generator:

[0073] Among them, G is the generator, G(x|y) represents the time-domain signal generated by the generator G given the magnitude spectrum x of the mixed speech and the fundamental frequency condition y; D k is the kth discriminator, x is the magnitude spectrum of the input mixed speech, y is the given fundamental frequency condition, s is the speaker time-domain waveform corresponding to this fundamental frequency, represents the mathematical expectation of the formula (D k (s|y) - 1) 2 under the condition that the true time-domain waveform signal is s, and D k(s|y) represents the output value of the k-th discriminator corresponding to the given fundamental frequency condition y and the time-domain waveform s of the speaker corresponding to this fundamental frequency. It represents the mathematical expectation of the formula (D k (G(x|y))) 2 under the condition that the magnitude spectrum of the mixed speech is x, where D k (G(x|y)) represents the output result of the k-th discriminator for the time-domain signal G(x|y) generated by the generator.

[0074] Based on the loss function of the generative adversarial network (GAN), the present invention introduces a multi-resolution short-time Fourier transform (STFT) loss. Previous studies have shown that this loss function can not only effectively measure the differences between real and fake time-domain waveforms in the latent feature space, but also stabilize the training of the GAN and accelerate convergence. For the loss of a single STFT, the goal is to minimize the spectral convergence error L between the estimated real target signal s and the signal estimated by the generator sc and the log magnitude spectrum error

[0075]

[0076]

[0077] where ‖·‖ F and ‖·‖1 are the Frobenius and L1 normalizations respectively, and |STFT(·)| represents the magnitude spectrum of the signal after STFT transformation, and N is the total number of elements in the magnitude spectrum.

[0078] The multi-resolution STFT loss consists of M single STFT losses, where each STFT loss corresponds to different FFT sizes, window lengths, and frame shift parameters. Averaging these M losses gives the final multi-resolution STFT loss function:

[0079]

[0080] where represents the mathematical expectation of the result obtained from the formula under the condition that the real target signal is s and the signal estimated by the generator is ; represents the spectral convergence error between the real target signal s and the signal estimated by the generator ; represents the log magnitude spectrum error between the real target signal s and the signal estimated by the generator .

[0081] Therefore, the objective function of the entire adversarial-generative network can be expressed as:

[0082]

[0083] Among them, λ represents the network training weight parameter set manually according to experience.

[0084] In terms of the network structure, the generator first uses three layers of two-dimensional convolution with a convolution kernel size of 3 and a stride of 2. Each layer adopts a ResNet structure with residual connections, followed by 4 Transformer modules. The input feature dimension of each layer is 512, where the self-attention part uses 8 heads (M = 1, d model = 512, H = 8), and the dimension of the forward layer is 1024. A linear layer is added between the convolution module and the Transformer module to make the output dimension of the former match the input dimension of the latter. The part up to this point is denoted as G1, which is the basic module of the generator with the magnitude spectrum (spectrum) as the generation target. On this basis, the subsequent upsampling module uses three upsampling layers to gradually upsample the input dimension to 64 times the original (determined by the length of the frame shift). The upsampling coefficient of each layer is [4x, 4x, 4x], and the output channel numbers are 256, 128, and 64 respectively. Each upsampling layer consists of a transposed convolution and a residual dilation convolution module (ResStack). Among them, the kernel size of the transposed convolution is twice the stride. The ResStack module consists of 4 layers of one-dimensional convolution with dilation. Its kernel size is 3, and the dilation size increases as 1, 3, 9, 27 with the increase of the layer number. Finally, a receptive field corresponding to 81 frames can be obtained. Previous studies have shown that expanding the receptive field of convolution to a reasonable range can improve the quality of the generated speech. The above module is denoted as G2. As the follow-up of G1, it is concatenated with G1 to form the generator of the method with the time-domain signal as the generation target.

[0085] For the discriminator, first use 2 layers of two-dimensional convolution to decompose the input spectrum into 12 * 5 patches, and the convolution kernel size and stride are both 7 * 5. Then, through a linear flattening layer, the time dimension and frequency dimension are merged to obtain a one-dimensional sequence with a sequence length of 12 * 5 and a feature dimension of the channel dimension after the above two-dimensional convolution. A position encoding and a symbol [cls] for discriminative classification are added at the beginning of this sequence. The above result is input into the Transformer layer.

[0086] 4. Fundamental frequency update

[0087] The fundamental frequency update module aims to extract the fundamental frequency from the output result of the speech separation module and use it to update the output result of the fundamental frequency prediction module. The function of this module can be expressed by the following formula (taking the current moment t as an example):

[0088] p 0~t = UpdatePitch(y 0~t ), (16)

[0089]

[0090] where p 0~t represents the fundamental frequency sequence of the target speaker from time 0 to time t, UpdatePitch represents the fundamental frequency update module, and y 0~t represents the speech signal sequence of the target speaker from time 0 to time t.

[0091] Specifically, given the output (y0, …, y N ) of the speech separation module as the input of this module, the fundamental frequency results (p0, …, p t , …, p N ) are obtained by using the frame-level fundamental frequency extraction network. Since this fundamental frequency is extracted from the relatively clean separated speech, the output of this module can be used as a more accurate result to update the prediction value of the fundamental frequency prediction module at the current moment. As the "prediction-separation-update" process proceeds frame by frame, the fundamental frequencies of all historical moments will be updated and used as the conditions for the fundamental frequency prediction module, and finally a more accurate fundamental frequency trajectory can be obtained.

[0092] Specifically, given the speech of a single speaker, the amplitude spectrum y m of each corresponding frame is obtained through short-time Fourier transform and used as the input of the neural network to estimate the posterior probability of the fundamental frequency of each frame, that is, p(z m |y m ). The frequency range of 60 - 404 Hz is quantized into 67 frequency ranges at a logarithmic scale with 24 frequency points per octave. This process quantizes the frequency range where the fundamental frequency may fall from continuous frequency values to discrete frequency values, and this value is determined by the center frequencies of the 67 frequency ranges. In addition, silence and unvoiced sounds are regarded as an additional type of fundamental frequency range, for a total of 68 discrete frequency ranges. Then p(z m |y m ) represents the probability that the fundamental frequency of the m-th frame of the given input mixed speech amplitude spectrum corresponds to a certain value among these 68 frequencies. If the fundamental frequency label of the m-th frame corresponds to the s-th frequency value, then p(z m (s)|y m ) is equal to 1. In terms of network structure design, first, a convolutional neural network is used to model the local characteristics of the input spectrum and capture the harmonic structure between frequency components. Then a fully connected layer is connected to model the mapping relationship between the harmonics and the fundamental frequency of each frame. The categorical cross-entropy is used as the loss function, which is defined as follows:

[0093]

[0094] Among them, O m (s) represents the probability that the m-th frame amplitude spectrum corresponds to the s-th frequency value.

[0095] 5. Loop separation

[0096] The above three modules constitute the iterative fundamental frequency estimation and speech separation framework proposed by the present invention. Given a mixed speech, running the framework once will output the fundamental frequency trajectory of a certain speaker, which is determined by the leading speaker in the mixed speech. Using this fundamental frequency trajectory as a condition, jointly with the mixed speech spectrum, it is input into the conditional generative adversarial network of the speech separation module to obtain the speaker speech corresponding to the fundamental frequency condition. This speaker depends on the leading speaker at the starting moment of the mixed speech. For the remaining speakers to be separated in the mixed speech, subtract the speaker speech separated in the previous round from the mixed speech, and perform the above iterative process on the residual speech again. Repeat this process until there is no speech in the residual speech, that is, once no fundamental frequency value is predicted in the remaining signal, the loop process stops. The stopping condition of the loop can be determined by the energy value of the remaining signal. If the energy is less than a certain threshold, the stopping condition is satisfied. The entire loop separation process is as Figure 3 shown.

[0097] The advantages of the present invention will be described below in conjunction with specific embodiments. The speech separation performance was tested on an experimental dataset using this method. The results of this method will be compared with those of previous methods. In addition, it will also be compared with other methods that use speech separation or fundamental frequency extraction as the front-end and back-end of each other.

[0098] 1) Experimental setup

[0099] The experimental dataset is based on two-speaker speech (WSJ0-2mix) and three-speaker speech (WSJ0-3mix) mixed from Wall Street Journal speech data (WSJ0). Among them, each type of mixed speech contains approximately 30 hours of training data, 10 hours of validation data, and 5 hours of test data, with the signal-to-noise ratio ranging from 0 dB to 10 dB. In addition, the fundamental frequency is extracted from the speech of a single speaker using the Praat tool to obtain the fundamental frequency label. The relative improvement value (SDRi) with respect to the original speech, the objective speech quality index (PESQ), and the short-time objective intelligibility (STOI) are used as evaluation indicators for the speech separation task. Use E Total as the evaluation indicator for the fundamental frequency extraction task. This indicator can evaluate the accuracy of both fundamental frequency estimation and speaker assignment at the same time. It is a combination of pronunciation discrimination errors (frames without fundamental frequency are judged as frames with fundamental frequency, or vice versa), permutation errors (incorrect fundamental frequency assignment between different speakers), coarse-grained errors, and fine-grained errors. The smaller this indicator is, the better.

[0100] The settings of the voice separation module and the fundamental frequency update module are as described in the specific implementation manners above. Their parameters are trained and fixed using training data and are not jointly trained with the fundamental frequency prediction module. The entire framework of each component of the fundamental frequency prediction module, including the encoder, the prediction network, and the joint network, is trained and jointly optimized simultaneously.

[0101] 2) Experimental results

[0102] The method of the present invention is compared with previous methods, and the results are shown in Table 1. This method can achieve performance superior to traditional methods (uPIT and DPCL) and comparable to the current best method (Conv-TasNet). The indicators of SDRi, PESQ, and STOI have all improved. However, for the mixed speech of two speakers, the performance on SDRi is slightly worse than that of Conv-TasNet. The reason may be that in the fundamental frequency prediction module, for the encoder structure of the mixed speech input, BLSTM is used, while Conv-TasNet uses a more powerful time-domain convolution TCN in the encoding of the input mixed speech. Therefore, on the basis of the model of the present invention, the encoder in the fundamental frequency prediction module is changed to the TCN structure, and the comparison results are shown in Table 2. The performance of the present invention has been improved and is slightly higher than that of the Conv-TasNet method.

[0103] Compared with the IRM and IBM methods for known separated speech (labels), this method shows superiority in the SDRi index, but is slightly worse in the objective perception indexes of PESQ and STOI. The IRM / IBM methods estimate the amplitude spectrum of the separated speech more accurately and have more advantages in the PESQ and STOI indexes calculated based on the signal amplitude spectrum; however, this type of method reconstructs the time-domain signal using the phase of the mixed speech, resulting in a poor SDRi index based on the time-domain signal energy. The present invention directly outputs the time-domain signal of the separated speech, which is beneficial to improving the SDRi index.

[0104] The present invention is also compared with other methods with voice separation or fundamental frequency extraction as the front and back ends, as shown in Table 3. Among them, the systems for comparison are one-way processes with a "fundamental frequency - separation" framework (extracting the fundamental frequency trajectories of each speaker, splicing them with the mixed speech and inputting them into the voice separation system to obtain the separated speech of each speaker), and one-way processes with a "separation - fundamental frequency" framework (first separating the speech of each speaker from the mixed speech, then splicing it with the mixed speech and inputting it into the fundamental frequency extraction system to extract the fundamental frequency trajectories of each speaker). The above two systems can be regarded as a one-round iterative process for a certain task. The cyclic iterative framework in the present invention can depict the interdependent and mutually promoting relationship between the two tasks, and the superiority of this framework is also reflected in the experimental results, which can improve the performance of both simultaneously.

[0105] Table 1. Comparison of the present invention with other methods in terms of speech separation performance

[0106]

[0107] Table 2. Influence of different structures of the encoder in the fundamental frequency prediction module of the present invention on speech separation performance

[0108]

[0109] Table 3. Comparison results of the present invention with other methods using speech separation or fundamental frequency extraction as the front and back ends

[0110] Model SDRi(dB) <![CDATA[E Total (%)]]> Pitch-cGAN (our method) 16.1 18.7 Pitch-SS 12.0 - SS-Picth - 19.6

[0111] Another embodiment of the present invention provides an iterative fundamental frequency estimation and speech separation device based on a bidirectional cascade framework, which includes a signal preprocessing module, a fundamental frequency prediction module, a speech separation module, a fundamental frequency update module, and a cyclic separation module;

[0112] The signal preprocessing module is used to perform frame division, windowing, and short-time Fourier transform operations on the given mixed speech in sequence to obtain the time-frequency spectrum of the mixed speech;

[0113] The fundamental frequency prediction module is used to predict the fundamental frequency value at the current moment based on the mixed speech frame at the current moment and the fundamental frequency prediction value at the historical moment, and extract the fundamental frequency sequence of a certain speaker in the mixed speech;

[0114] The speech separation module is used to take the mixed speech and the fundamental frequency sequence of a certain speaker obtained by the fundamental frequency prediction module as inputs, and use a conditional generative adversarial network to generate the speaker speech corresponding to the fundamental frequency sequence;

[0115] The fundamental frequency update module is used to take the generated speaker speech as an input, extract the fundamental frequency trajectory, and use it to update the fundamental frequency prediction value of the current frame output by the fundamental frequency prediction module;

[0116] The cyclic separation module is used to run the fundamental frequency prediction module, the speech separation module, and the fundamental frequency update module for the given mixed speech to obtain the separated speech of a certain speaker, which is determined by the leading speaker in the mixed speech. Subtract the already separated speaker speech from the mixed speech in the previous round, and perform the iterative process of the fundamental frequency prediction module, the speech separation module, and the fundamental frequency update module on the residual speech again. Repeat this process until there is no speech in the residual speech, and then the cyclic process stops, so as to separate the speech of each speaker in the mixed speech.

[0117] For the specific implementation process of each module, please refer to the description of the method of the present invention above.

[0118] Based on the same inventive concept, another embodiment of the present invention provides an electronic device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for performing each step in the method of the present invention.

[0119] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a magnetic disk, an optical disk). When the computer program stored in the computer-readable storage medium is executed by a computer, each step of the method of the present invention is implemented.

[0120] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention shall be subject to the scope defined by the claims.

Claims

1. An iterative fundamental frequency estimation and speech separation method based on a bidirectional cascaded framework, the steps of which include: 1) For the given mixed speech, perform frame segmentation, windowing, and short-time Fourier transform operations in sequence to obtain the time-frequency spectrum of the mixed speech, and then loop through steps 2) to 4) frame by frame until all frames are traversed; 2) Use the fundamental frequency prediction module to predict the fundamental frequency value at the current moment based on the mixed speech frame at the current moment and the fundamental frequency prediction value at the historical moment, and extract the fundamental frequency sequence of a certain speaker in the mixed speech; 3) Use the speech separation module. With the mixed speech and the fundamental frequency sequence obtained in step 2) as inputs, use a conditional generative adversarial network to generate the speech of the speaker corresponding to the fundamental frequency sequence; 4) Use the fundamental frequency update module. With the speech of the speaker generated in step 3) as the input, extract the fundamental frequency trajectory and use it to update the fundamental frequency prediction value of the current frame output in step 2); 5) After the loop described in step 1), the complete fundamental frequency trajectory of a certain speaker in the mixed speech can be obtained. The speaker is determined by the leading speaker in the mixed speech. Use this fundamental frequency trajectory as the fundamental frequency condition and combine it with the mixed speech spectrum, and input it into the conditional generative adversarial network in step 3) to separate the speech of the speaker corresponding to the fundamental frequency condition; 6) Subtract the separated speech of the speaker in step 5) from the mixed speech, and perform the iterative process of steps 1) to 5) on the residual speech again. Loop like this until there is no speech in the residual speech, and the above loop process stops, so as to separate the speech of each speaker in the mixed speech.

2. The method according to claim 1, wherein The fundamental frequency prediction module includes an encoder, a prediction network, and a joint network. The encoder uses a two-layer two-dimensional convolutional neural network followed by a four-layer bidirectional long short-term memory network. The prediction network uses a two-layer long short-term memory network. The joint network is a one-layer fully connected network. The entire framework is trained and jointly optimized at the same time, and its optimization goal is the categorical cross-entropy loss.

3. The method according to claim 1, wherein The conditional generative adversarial network consists of a generator and a discriminator. The generator aims to generate a time-domain signal corresponding to the fundamental frequency condition from the magnitude spectrum of the mixed speech. The discriminator consists of multiple sub-discriminators acting on different frequency bands of the speech signal. The optimization goals of the two are respectively: Discriminator: Generator: Among them, G is the generator, and D k is the k-th discriminator, x is the magnitude spectrum of the input mixed speech, y is the given fundamental frequency condition, and s is the speaker's time-domain waveform corresponding to the fundamental frequency; G(x|y) represents the time-domain signal generated by the generator G given x and y; represents the mathematical expectation of the formula (D k (s|y) - 1) 2 under the condition that the true time-domain waveform signal is s, and D k (s|y) represents the output value of the k-th discriminator given y and s, represents the mathematical expectation of the formula (D k (G(x|y))) 2 under the condition that the magnitude spectrum of the mixed speech is x; D k (G(x|y)) represents the output result of the k-th discriminator for the time-domain signal G(x|y) generated by the generator; K represents the number of discriminators.

4. The method according to claim 1, wherein The fundamental frequency update module aims to extract the fundamental frequency from the output result of the speech separation module and use it to update the output result of the fundamental frequency prediction module; the fundamental frequency update module uses a convolutional neural network to model the local characteristics of the input spectrum, capture the harmonic structure between frequency components, and then connects a fully connected layer to model the mapping relationship between the harmonics and the fundamental frequency of each frame. The optimization goal is the categorical cross-entropy loss function.

5. The method according to claim 1, characterized in that, Adopt a frame-by-frame iterative "fundamental frequency prediction - speech separation - fundamental frequency update" framework, cascade the fundamental frequency extraction and speech separation tasks bidirectionally, and improve the performance of both at the same time. The rules for its loop iteration processing are: For a given mixed speech, each time the framework runs, the separated speech of a certain speaker will be output. The speaker is determined by the leading speaker in the mixed speech. The separated speech of the speaker from the previous round is subtracted from the mixed speech, and the above iterative process is performed again on the residual speech. This loop continues until there is no speech in the residual speech, that is, once no fundamental frequency value is predicted in the remaining signal, the loop process stops; alternatively, the condition for the loop to stop is determined by the energy value of the remaining signal. If the energy is less than a certain threshold, the stop condition is met.

6. An iterative fundamental frequency estimation and voice separation device based on a bidirectional cascade framework, characterized in that It includes a signal preprocessing module, a fundamental frequency prediction module, a speech separation module, a fundamental frequency update module, and a loop separation module; The signal preprocessing module is used to perform frame division, windowing, and short-time Fourier transform operations on the given mixed speech in sequence to obtain the time-frequency spectrum of the mixed speech; The fundamental frequency prediction module is used to predict the fundamental frequency value at the current moment based on the mixed speech frame at the current moment and the fundamental frequency prediction value at the historical moment, and extract the fundamental frequency sequence of a certain speaker in the mixed speech; The speech separation module is used to take the mixed speech and the fundamental frequency sequence of a certain speaker obtained by the fundamental frequency prediction module as inputs, and use a conditional generative adversarial network to generate the speaker speech corresponding to the fundamental frequency sequence; The fundamental frequency update module is used to take the generated speaker speech as an input, extract the fundamental frequency trajectory, and use it to update the fundamental frequency prediction value of the current frame output by the fundamental frequency prediction module; The loop separation module is used to run the fundamental frequency prediction module, the speech separation module, and the fundamental frequency update module for the given mixed speech to obtain the separated speech of a certain speaker. The speaker is determined by the leading speaker in the mixed speech. The separated speech of the speaker from the previous round is subtracted from the mixed speech, and the iterative process of the fundamental frequency prediction module, the speech separation module, and the fundamental frequency update module is performed again on the residual speech. This loop continues until there is no speech in the residual speech, and the loop process stops, thereby separating the speech of each speaker in the mixed speech.

7. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. When the computer program is executed by the computer, it implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Single-channel voice separation method and device for multiple speakers

    CN112331218A

  • Blind source extraction method using direction of arrival information and de-mixing system therefor

    KR1020140106823A