A two-stage multi-speaker fundamental frequency trajectory extraction method
Through the two-stage multi-speaker fundamental frequency extraction method, the fundamental frequency is estimated at the frame level using a convolutional neural network and a fully connected layer, and the fundamental frequency is connected at the sentence level through a conditional chain model, the problem of fundamental frequency extraction in the multi-speaker mixed speech environment is solved, and the efficient and accurate fundamental frequency extraction effect is achieved.
Patent Information
- Application Number
- CN202211084602.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-09-06
AI Technical Summary
In a multi-speaker hybrid voice environment, traditional fundamental frequency extraction methods make it difficult to accurately estimate the fundamental frequency of each speaker, especially in the presence of background noise or other voice interference.
The two-stage multi-speaker fundamental frequency extraction method is adopted. First, the fundamental frequency is estimated at the frame level through the convolutional neural network and the fully connected layer, and then the fundamental frequency is connected at the sentence level using the conditional chain model to construct a fundamental frequency extraction model that is independent of the speaker's characteristics.
It realizes the accurate extraction of each speaker's fundamental frequency in a mixed voice environment of multiple speakers, improves the performance of fundamental frequency extraction in complex environments, and avoids the problem that the results caused by the preset output number in traditional methods do not match the preset output.
Smart Images

Figure CN115631744B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech signal processing, relates to fundamental frequency extraction technology, and specifically relates to a two-stage multi-speaker fundamental frequency trajectory extraction method. Background Art
[0002] The fundamental frequency, corresponding to the auditory perception attribute of pitch, determines the pitch of a sound. When the vocal organs produce voiced sounds, the vocal cords vibrate periodically, and the fundamental frequency is determined by the frequency of the vocal cord vibration. The frequency components of a periodic sound signal consist of the fundamental frequency and a series of harmonics, where the harmonics are integer multiples of the fundamental frequency, and this property is called "harmonicity". The auditory peripheral system can decompose the lower-order harmonics, and the regular intervals of the lower-order harmonics can promote perceptual integration, that is, only when the fundamental frequencies corresponding to all harmonics are the same, will the listener perceive the stimulus signal as a single sound source; if the harmonics can be divided into several groups corresponding to several fundamental frequencies, the listener will perceive separate sound sources. Fundamental frequency extraction has wide applications in sound signal processing, such as identifying music melodies, and can also be used in speech processing, such as speech separation, speech recognition, speech emotion analysis and other technical fields.
[0003] Traditional fundamental frequency extraction methods include the autocorrelation function method, the average magnitude difference function method, the cepstrum method, the simplified inverse filtering method, etc. These methods are all based on traditional signal processing and have poor generalization and robustness in complex environments, and can only be applied to the scenario of a single speaker in a quiet situation. When the speech of the target speaker is affected by background noise or other speech interference, these traditional methods usually fail. The fundamental frequency extraction task in the present invention involves a multi-source scenario, and the goal is to extract the fundamental frequency of each speaker from the mixed speech of multiple people, and this process is also called "multi-fundamental frequency tracking".
[0004] The difficulties in multi-speaker fundamental frequency extraction are as follows: 1) The fundamental frequency trajectory is continuous and time-varying; 2) In the mixed speech of multiple speakers (such as two speakers), there will not always be two fundamental frequencies at every moment (it may be silent). Therefore, this task not only requires accurately estimating the fundamental frequency value of each speaker at each moment, but also needs to concatenate the fundamental frequencies belonging to the same speaker at different moments, that is, it is necessary to assign the fundamental frequency estimated at each moment to the corresponding speaker, so as to obtain the fundamental frequency trajectory at the sentence level of each speaker. Summary of the Invention
[0005] In view of the shortcomings of the existing methods, the present invention proposes a two-stage multi-speaker fundamental frequency extraction method. By analyzing the spectral pattern of speech, the present invention can see that it has regularly spaced frequency components - harmonics. The harmonic components in the spectrum are integer multiples of the corresponding fundamental frequency. Therefore, a neural network can be used to model the mapping relationship between harmonics and fundamental frequency. The present invention uses a neural network to mine the harmonic components in the input spectrum and learn the mapping relationship between harmonics and fundamental frequency, thereby constructing a fundamental frequency extraction model that is independent of speaker characteristics and has no limit on the number of speakers.
[0006] The method for multi-speaker fundamental frequency estimation proposed in the present invention is divided into two stages: the frame-level fundamental frequency estimation stage and the sentence-level fundamental frequency concatenation stage. The former aims to separate the fundamental frequency values of different speakers at the same time in mixed speech, and the latter aims to assign the frame-level estimation results to the corresponding speakers to obtain the fundamental frequency trajectory of a single speaker at the sentence level. The learning objectives of each stage of this method are clear, which is significantly different from the traditional end-to-end "black box" method that directly estimates the fundamental frequency of each speaker simultaneously from the mixed speech.
[0007] The technical solution of the present invention is:
[0008] A two-stage method for extracting fundamental frequency trajectories of multiple speakers, comprising the following steps:
[0009] 1) Processing a given multi-speaker mixed speech to obtain a frequency spectrum of each frame in the multi-speaker mixed speech; the frequency spectrum includes an amplitude spectrum and a phase spectrum;
[0010] 2) using a convolutional neural network to obtain local features of the amplitude spectrum;
[0011] 3) inputting the local features of the amplitude spectrum of each frame into a fully connected layer, obtaining the mapping relationship between the harmonics and the fundamental frequency of each frame in the multi-speaker mixed speech, and obtaining all the fundamental frequency estimation values corresponding to each frame;
[0012] 4) Taking the fundamental frequency estimation values of each frame obtained in step 3) as input, extracting the fundamental frequency sequence of one speaker in each iteration until the fundamental frequency sequence of the last speaker is predicted; wherein the processing method of the i-th iteration is:
[0013] a) inputting the base frequency sequence separated in the i-1th round into the encoder for encoding to obtain the feature representation of the base frequency sequence separated in the i-1th round;
[0014] b) inputting the fundamental frequency sequence feature representation obtained in step a) and the fundamental frequency estimation values of all frames obtained in step 3) into the conditional chain module to obtain the hidden layer output vector corresponding to the i-th iteration;
[0015] c) The decoder decodes the hidden layer output vector corresponding to the i-th round of iteration into the fundamental frequency sequence of the i-th speaker, that is, the fundamental frequency sequence separated in the i-th round of iteration is obtained.
[0016] Furthermore, all fundamental frequency estimation values corresponding to each frame are obtained by using the trained frame-level fundamental frequency estimation network; wherein, the frame-level fundamental frequency estimation network includes a convolutional neural network and a fully connected layer. First, the convolutional neural network is used to model the local characteristics of the input amplitude spectrum, capture the harmonic structure between frequency components in the amplitude spectrum and input it into the fully connected layer; the fully connected layer models the mapping relationship between the harmonics and the fundamental frequency of each frame in the multi-speaker mixed speech; the loss function used for training the frame-level fundamental frequency estimation network is m is the frame index, s is the index of the fundamental frequency value, y m is the amplitude spectrum of the m-th frame, z m (s) represents the s-th component of the fundamental frequency label z m of the m-th frame, S is the total number of components of the fundamental frequency label, p(z m (s)|y m ) represents the probability that the fundamental frequency label of the m-th frame corresponds to the s-th frequency value, O m (s) is the probability that the m-th frame amplitude spectrum corresponds to the s-th frequency value.
[0017] Furthermore, the encoder is composed of two layers of bidirectional LSTM, with 256 hidden layer nodes in each layer; the conditional chain module includes one layer of LSTM layer, with 512 hidden layer nodes; the decoder converts the encoding dimension of each frame of the hidden layer into the number of fundamental frequency categories required for the output sequence through a linear layer.
[0018] Furthermore, the given multi-speaker mixed speech is sequentially subjected to frame splitting, windowing, and short-time Fourier transform operations to obtain the spectrum of each frame in the multi-speaker mixed speech.
[0019] Furthermore, if i = 1, the fundamental frequency sequence input to the encoder is a silent sequence of all zeros.
[0020] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the above method.
[0021] A computer-readable storage medium, on which a computer program is stored, characterized in that the steps of the above method are implemented when the computer program is executed by a processor.
[0022] Compared with the prior art, the positive effects of the present invention are:
[0023] The present invention has the following advantages: (1) Since the model only learns the mapping relationship between the harmonic components and the fundamental frequency in the spectrum, it is independent of the speaker characteristics and the number of speakers. Even if it is trained on the speech of a single speaker, it is also applicable to the case of multi-speaker mixed speech; (2) Assuming that the input is the mixed speech of two speakers, it does not mean that there are two fundamental frequencies at every moment. There may be silent and voiceless cases, and there is no fundamental frequency at these moments. That is, the situations of no fundamental frequency, one fundamental frequency, or two fundamental frequencies will all exist. Existing models usually preset the number of matching output layers under the condition of knowing the number of speakers, which will lead to the problem that the real result does not match the preset output. However, the method proposed by the present invention does not need to preset the output number, thus avoiding this problem. (3) The present invention only uses a simple convolutional neural network and fully connected layers, and has achieved comparable performance with the current advanced methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a general flowchart of the present invention.
[0025] Figure 2 It is a framework diagram of the frame-level fundamental frequency extraction process of the present invention.
[0026] Figure 3 It is a structural diagram of the convolutional neural network in the frame-level fundamental frequency extraction sub-process of the present invention.
[0027] Figure 4 It is a framework diagram of the conditional chain model used in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0028] The specific implementation manner of the present invention will be described in more detail below. The specific implementation steps of the method of the present invention include signal preprocessing (step 1), frame-level fundamental frequency estimation network (corresponding to steps 2-3), conditional chain model (corresponding to steps 4-7), etc. The specific implementation processes of each step are as follows:
[0029] 1. Speech signal preprocessing
[0030] This method first performs short-time Fourier transform (STFT) on the mixed speech as the subsequent input, and uses the analysis window w(n), window length N, and frame shift R to perform short-time Fourier transform on the signal. The transformation formula is as follows:
[0031]
[0032] Among them, t and f respectively represent the indices of the frame and frequency band. After the transformation, the STFT spectrum is obtained. In the specific implementation, the frame length used is 32 ms, the frame shift is 16 ms, and the window function is a Hamming window. The STFT spectrum contains the amplitude information and phase information of this section of speech signal in the time dimension and frequency dimension (that is, the amplitude spectrum and phase spectrum can be derived from the STFT spectrum).
[0033] 2. Frame-level fundamental frequency estimation network
[0034] The frame-level fundamental frequency estimation network takes the STFT magnitude spectrum obtained through short-time Fourier transform as input and outputs the estimated multi-speaker fundamental frequency values at the frame level. Specifically, given the speech of a single speaker, the magnitude spectrum y of each corresponding frame is obtained through short-time Fourier transform m , which is used as the input of the neural network to estimate the posterior probability of the fundamental frequency of each frame, that is, p(z m |y m ). The frequency range from 60 to 404 Hz is quantized into 67 frequency ranges with a logarithmic scale of 24 frequency points per octave. This process quantizes the frequency range where the fundamental frequency may fall from continuous frequency values to discrete frequency values, and this value is determined by the center frequencies of the 67 frequency ranges. In addition, silence and voiceless sounds are regarded as an additional type of fundamental frequency range, for a total of 68 discrete frequency ranges. Then p(z m |y m ) represents the probability that the fundamental frequency of the m-th frame of the given input mixed speech magnitude spectrum corresponds to a certain value among these 68 frequencies. If the fundamental frequency label of the m-th frame corresponds to the s-th frequency value, then p(z m (s)|y m ) is equal to 1.
[0035] The frame-level fundamental frequency estimation network consists of two parts: a convolutional neural network and a fully connected layer. In terms of network structure design, as Figure 1 shown, first, the convolutional neural network is used to model the local characteristics of the input magnitude spectrum to capture the harmonic structure between frequency components. Then, a fully connected layer is connected to model the mapping relationship between the harmonics and the fundamental frequency of each frame.
[0036] The fundamental frequency estimation aims to obtain the posterior probability of the fundamental frequency of each frame. The categorical cross-entropy is used as the loss function, which is defined as follows:
[0037]
[0038] where m is the frame index, s is the index of the 68 fundamental frequency values, O is the linear output layer for 68-class classification, y m represents the magnitude spectrum of the m-th frame of the input, obtained through preprocessing, z m (s) represents the s-th component of the fundamental frequency label z m of the m-th frame (obtained from the training data labels), and p(z m (s)|y m ) represents the probability that the fundamental frequency label of the m-th frame corresponds to the s-th frequency value (if the fundamental frequency label of the m-th frame corresponds to the s-th frequency value, then p(z m (s)|y m) is equal to 1), O m (s) is the probability corresponding to the s-th frequency value of the amplitude spectrum of the m-th frame, which is output by the network after mapping through the final linear layer. For the multi-classification task with multi-speaker mixed speech as training data, the fully connected linear layer uses the sigmoid activation function to ensure that the network output probability is between 0 and 1.
[0039] This network is trained using the publicly released Wall Street Journal mixed speech dataset (WSJ0-2mix). Among them, the training set contains approximately 30 hours of mixed speech, randomly selects sentences of two speakers, and mixes them with a signal-to-noise ratio uniformly sampled between 0dB and 10dB. We use this dataset to train for the task of multi-speaker fundamental frequency extraction. The fundamental frequency labels are obtained by extracting on the clean speech of a single speaker using the existing publicly available Praat software tool, and then converting the fundamental frequency value of each frame into a vector label in the format described above. For the training task of a single speaker, this label can be directly used; for the multi-speaker task, its label is obtained by taking the union of the fundamental frequency vectors of individual speakers contained in the mixed speech. The sampling rate of all speech data is 16kHz. When extracting STFT features, the frame length used is 32ms, the frame shift is 16ms, and the window function is the Hamming window.
[0040] The training of this network, as shown in Equation 2, p(z m (s)|y m ) is given by the dataset label, O m (s) is the output obtained after the network inputs the amplitude spectrum. Thus, the loss function can be calculated according to Equation 2 to train the network parameters. After the training is completed, the network parameters are fixed, and then the trained neural network and fully connected layer are used to process in Steps 2-3 of the invention content.
[0041] 3. Conditional chain model
[0042] As Figure 4 shown, in the conditional chain model, each output sequence is not only determined by the input sequence, but also affected by the previous output sequence, that is, the previous output sequence will be used as a conditional input to the module that determines the current output sequence. Therefore, the conditional chain model can not only model the direct mapping relationship from the input sequence to the output sequence, but also model the relationship between output sequences. The output fundamental frequency sequences of each speaker seemingly have an independent parallel relationship, but in fact, there is a mutually exclusive relationship between them, and the conditional chain model can model this relationship, which is expressed by the formula:
[0043]
[0044] That is, given the input sequence O, which is the frame-level fundamental frequency estimation sequence, the formula models the joint probability of the fundamental frequency sequences s of N speakers. For each output fundamental frequency sequence, it is jointly determined by the original input sequence and the previously output fundamental frequency sequence, characterized by a conditional probability, which can be implemented using the structure of a conditional encoder-decoder. Specifically, the encoder-decoder part is used to encode the input sequence and decode the output sequence, and the conditional chain part is used to store the information from the previous output sequence as a conditional input to the decoding process of the current output sequence.
[0045] The purpose of multi-speaker fundamental frequency trajectory extraction is to obtain the fundamental frequency trajectory output of a single speaker from the frame-level fundamental frequency input of unassigned speakers. Figure 2 The structure for using the conditional chain model to solve this task is shown. Specifically, the input is the result obtained in the previous section, which can be regarded as a binary vector graph composed of 0 and 1 where F and T are the number of frequency bands and the number of frames respectively. The positions with the value of "1" indicate the existence of the fundamental frequency at this frame and this frequency. The encoder and decoder in the conditional chain module are shared at each step i, and the information transmission between the output sequences is achieved through a unidirectional LSTM. The fusion module uses a concatenation operation to concatenate the hidden layer representations of the input sequence and the output sequence of the previous step in the feature dimension. The decoder decodes the hidden layer output H of the LSTM at the current moment i into the fundamental frequency sequence of the target speaker The whole process can be expressed by the following formula:
[0046]
[0047]
[0048]
[0049] where the encoder and decoder at each step i are shared, and the output sequence has the same time dimension as the input mixed speech. In addition to being able to characterize the (mutually exclusive) relationship between the output fundamental frequency trajectory sequences, this conditional chain model can also characterize the (temporal continuity) relationship between the fundamental frequency values at each moment within a single fundamental frequency trajectory sequence.
[0050] In terms of the network structure, the input frame-level fundamental frequency sequence passes through three parts: an encoder, a conditional chain module, and a decoder, and then outputs the fundamental frequency trajectory sequence. Among them, the encoder consists of two layers of bidirectional LSTM, with 256 hidden nodes in each layer. The conditional chain module consists of only one layer of LSTM, with 512 hidden nodes. The decoder converts the hidden layer output vector into the fundamental frequency sequence through a linear layer. The input dimension of the linear layer is the dimension of the hidden layer output vector, and the output dimension is the number of fundamental frequency categories required by the fundamental frequency sequence.
[0051] This network is trained using the publicly released Wall Street Journal Mixed Speech Dataset (WSJ0-2mix). The training set contains approximately 30 hours of mixed speech, where sentences of two randomly selected speakers are mixed with a signal-to-noise ratio uniformly sampled between 0 dB and 10 dB. We use this dataset to train for the task of multi-speaker fundamental frequency extraction. The fundamental frequency sequence labels are obtained by extracting on the clean speech of a single speaker using the existing publicly available Praat software tool, and then converting the fundamental frequency value of each frame into a vector label in the format described above. All speech data has a sampling rate of 16 kHz. When extracting STFT features, the frame length used is 32 ms, the frame shift is 16 ms, and the window function is a Hamming window.
[0052] For the training of this network, the dataset labels give the true fundamental frequency sequence After formulas 4 - 6, the separated fundamental frequency sequence is output by the neural network The loss is calculated using formula 7 to train the network parameters. After training is completed, the network parameters are fixed, and then in the application phase, the trained units are directly used to process the data to be processed in the corresponding steps 5 - 7.
[0053]
[0054] Among them, the dataset labels give the true fundamental frequency sequence The fundamental frequency sequence of the neural network output result k represents that there are k speakers, that is, k fundamental frequency trajectory sequences.
[0055] To solve the problem of the variable number of output sequences and enable this model to be applied to scenarios where the number of speakers in the input mixed speech is variable, a termination sequence is added after the final output sequence to guide the termination of the training process. Specifically, in this method, when the fundamental frequency sequence of the last speaker is predicted, there will be no signal available for decoding, so a sequence with a fundamental frequency of 0 is used as the termination symbol. That is, when the result of a certain decoding no longer outputs a changing fundamental frequency trajectory, the decoding process is considered to end.
[0056] During the training process, the true label sequence is used as a condition (denoted as ), rather than the result estimated in the previous step (denoted as ). This is because it is considered that the error estimated in the previous step may be passed to the decoding process of the current step, causing the accumulation of errors, and avoiding the influence brought by systematic errors rather than the method itself.
[0057] In addition, the output of the model is multiple unordered sequences, which brings about the problem of the order (similar to the permutation problem) between the network output and the labels when calculating the loss. To address this issue, a greedy search strategy is adopted to solve the problem of the output sequence order. That is, for each step of the decoding process, the training objective is to minimize the difference between the current decoded output and each target sequence in the remaining label set, and the sequence corresponding to the minimum difference in the set is used as the target sequence for the current step.
[0058] The advantages of the present invention will be described below in conjunction with specific embodiments.
[0059] The fundamental frequency extraction performance test was carried out on the experimental dataset using this method, and the results of this method were compared with those of previous methods using general recognized evaluation metrics.
[0060] 1. Experimental Setup
[0061] The speech mixed from the WSJ0 database was used as the training and test dataset. Among them, the mixed speech of two speakers is WSJ0-2mix, which contains 30 hours of training data, 10 hours of validation data, and 5 hours of test data. In addition, the mixed speech dataset containing three speakers, WSJ0-3mix, was also used. These two datasets have currently become the general benchmark sets for the speech separation task, and can also be used to verify the performance of the fundamental frequency extraction task involved in the present invention. The fundamental frequency sequences of each individual speaker contained in the mixed speech constitute the true label set of the above-mentioned training dataset, and the labels were extracted using the Praat tool on the corresponding individual speaker's speech.
[0062] To compare with other methods for multi-speaker fundamental frequency extraction, we use E Total as the evaluation metric for this task. This metric can simultaneously evaluate the accuracy of fundamental frequency estimation and speaker assignment. It is a combination of pronunciation discrimination errors (frames without fundamental frequency are judged as frames with fundamental frequency, or vice versa), permutation errors (wrong fundamental frequency assignment between different speakers), coarse-grained errors, and fine-grained errors. The smaller this metric is, the better.
[0063] 2. Experimental Results
[0064] Table 1 shows the fundamental frequency estimation E Total values for the speech mixed with different signal-to-noise ratios for speakers of different gender combinations. The smaller the value, the better. The performance of the conditional chain model (Cond Chain) used in the present invention and the previous model (uPIT) was compared, and an attempt was made to use time-domain convolution as the basic network (encoder) to replace the traditional BLSTM. The method based on uPIT is the current mainstream method for multi-speaker fundamental frequency extraction and is the best-known method with the best performance.
[0065] Table 1 shows the comparison between the present invention and traditional methods in terms of fundamental frequency extraction performance
[0066]
[0067] An obvious result is that for mixed speech of different genders, the accuracy of the fundamental frequency trajectory at the sentence level is higher than that of mixed speech of the same gender, which is in line with expectations, that is, the fundamental frequency sequences of speakers of different genders have better distinguishing features. Under the condition of the same gender combination, as the signal-to-noise ratio increases, the fundamental frequency E Total value first decreases and then increases. A possible reason is that when mixing the voices of two speakers with a relatively high (9 dB) signal-to-noise ratio, the voice of a certain speaker (the one with higher energy) will dominate, and the estimation of the fundamental frequency trajectory of this speaker is relatively accurate, but the estimation result of the other masked speaker (the one with lower energy) is poor, thus affecting the overall result. However, this problem is alleviated in this method because the conditional chain model estimates each fundamental frequency sequence successively, and the previously estimated fundamental frequency sequence is used as a mutually exclusive condition to guide the estimation of the current fundamental frequency sequence, rather than estimating the fundamental frequency sequences of all speakers simultaneously in the uPIT framework, where the information used is only the mixed speech input.
[0068] Generally speaking, under various conditions, the performance of the method based on the conditional chain model is better than that of the method based on uPIT. A possible reason is that the uPIT method directly estimates the fundamental frequency trajectories of each speaker from the mixed speech. This process not only has to complete the separation task but also uses the uPIT training strategy to minimize the training error at the sentence level, that is, to complete the concatenation task at the sentence level. And this method splits the above process into two stages: frame-level fundamental frequency estimation and fundamental frequency concatenation based on the conditional chain model. Each stage is optimized separately and can achieve its best performance.
Claims
1. A two-stage multi-speaker fundamental frequency trajectory extraction method, the steps of which include: 1) Process the given multi-speaker mixed speech to obtain the spectrum of each frame in the multi-speaker mixed speech; the spectrum includes an amplitude spectrum and a phase spectrum; 2) Use a convolutional neural network to obtain the local features of the amplitude spectrum; 3) Input the local features of the amplitude spectrum of each frame into a fully connected layer to obtain the mapping relationship between the harmonics and the fundamental frequency of each frame in the multi-speaker mixed speech, and obtain all fundamental frequency estimation values corresponding to each frame; 4) Take the fundamental frequency estimation values of each frame obtained in step 3) as input, and extract the fundamental frequency sequence of one speaker in each round of iteration until the fundamental frequency sequences of the last speaker are predicted; the processing method for the i-th round of iteration is: a) Input the fundamental frequency sequence separated in the (i - 1)-th round into an encoder for encoding to obtain the feature representation of the fundamental frequency sequence separated in the (i - 1)-th round; b) Input the fundamental frequency sequence feature representation obtained in step a) and the fundamental frequency estimation values of all frames obtained in step 3) into a conditional chain module to obtain the hidden layer output vector corresponding to the i-th round of iteration; c) The decoder decodes the hidden layer output vector corresponding to the i-th round of iteration into the fundamental frequency sequence of the i-th speaker, that is, obtains the fundamental frequency sequence separated in the i-th round of iteration.
2. The method according to claim 1, wherein All fundamental frequency estimation values corresponding to each frame are obtained by using the trained frame-level fundamental frequency estimation network; wherein, the frame-level fundamental frequency estimation network includes a convolutional neural network and a fully connected layer. First, the convolutional neural network is used to model the local characteristics of the input amplitude spectrum, capture the harmonic structure between frequency components in the amplitude spectrum, and input it into the fully connected layer; the fully connected layer models the mapping relationship between the harmonics and the fundamental frequency of each frame in the multi-speaker mixed speech; the loss function used to train the frame-level fundamental frequency estimation network is where m is the frame index, s is the index of the fundamental frequency value, and y m is the amplitude spectrum of the m-th frame, and z m (s) represents the s-th component of the fundamental frequency label z m of the m-th frame, S is the total number of components of the fundamental frequency label, and p(z m (s)|y m ) represents the probability that the fundamental frequency label of the m-th frame corresponds to the s-th frequency value, and O m (s) is the probability that the m-th frame amplitude spectrum corresponds to the s-th frequency value.
3. The method according to claim 1, wherein The encoder is composed of two layers of bidirectional LSTM, and the number of hidden layer nodes in each layer is 256; the conditional chain module includes one layer of LSTM layer, and the number of hidden layer nodes is 512; the decoder converts the encoding dimension of each frame in the hidden layer into the number of fundamental frequency categories required by the output sequence through a linear layer.
4. The method according to claim 1, wherein Perform frame segmentation, windowing, and short-time Fourier transform operations on the given multi-speaker mixed speech in sequence to obtain the spectrum of each frame in the multi-speaker mixed speech.
5. The method according to claim 1, characterized in that, If i = 1, the fundamental frequency sequence input to the encoder is a silent sequence of all zeros.
6. A server, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Multi-speaker voice separation method based on convolutional neural network and depth clustering
CN110459240A
Fundamental frequency trajectory model parameter extraction device, fundamental frequency trajectory model parameter extraction method, program, and recording medium
JP2010020258A