Small sample voice timbre conversion method based on double-layer migration generative adversarial network
Through the small sample voice and tone conversion method of the two-layer migration generation adversarial network, the problem of poor voice conversion effect under small sample conditions is solved, and efficient and accurate voice and tone conversion is achieved, and the generated voice and tone similarity and naturalness are improved.
Patent Information
- Application Number
- CN202510369968.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-24
AI Technical Summary
Under small sample conditions, traditional voice tone conversion methods are difficult to generate voice signals that are highly consistent with the target voice, and the high algorithm complexity leads to high training difficulty and slow speed.
A small sample voice and voice conversion method based on a two-layer migration generation adversarial network is adopted. Through the two-layer migration network and module migration strategy, the target voice data is enhanced to improve the generation speed and effect.
This method can more accurately capture the underlying features and high-dimensional nonlinear features of speech, improve model training speed, reduce the risk of gradient disappearance, and the generated speech tone similarity and naturalness are better than traditional methods.
Smart Images

Figure CN120199237A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of speech signal processing and artificial intelligence technology, and specifically relates to a small sample speech timbre conversion method based on a double-layer migration generative adversarial network. Background Art
[0002] With the rapid development of speech synthesis and conversion technology, the demand for speech timbre conversion in personalized voice services, entertainment industry, and forensic identification is growing. However, in many cases, there are problems such as difficulty in collecting target timbre voice data or too few available fragments. For example, in the judicial process, when it is necessary to conduct voice identity identification on criminal suspects, there is often a problem of too little available evidence corpus. Traditional methods are mainly based on statistical modeling or signal processing technology (such as spectrum modification) of parallel speech data, but these methods have high requirements on data volume and quality, and it is difficult to generate speech signals that are highly consistent with the target timbre under small sample conditions. Most timbre conversion methods based on non-parallel speech data are also difficult to train and slow due to the high complexity of the algorithm, and it is difficult to overcome the problems caused by small sample conditions. Summary of the invention
[0003] In response to these shortcomings, the present invention proposes an innovative method for speech timbre conversion, which is based on deep learning and neural network algorithms. The method aims to enhance the target timbre speech data through strategies such as double-layer migration networks and module migration, and improve the speed and effect of target timbre speech generation under small sample conditions. This method can not only more accurately capture the underlying characteristics and high-dimensional nonlinear characteristics of speech, but also speed up model training and reduce the risk of gradient vanishing in generative adversarial networks. Specifically, a small sample speech timbre conversion method based on a double-layer migration generative adversarial network is provided, which specifically includes the following steps:
[0004] Step 1: Collect the speech data of the source speaker and the target speaker to form speech pair data, and construct a complete sample parallel training set, a small sample training set and a test set based on the speech pair data to simulate the timbre conversion scenarios under different data amounts; the complete sample parallel training set contains sufficient speech pairs, and the small sample training set and the test set contain a small number of speech pairs;
[0005] Step 2: Extract three types of features, namely, Mel cepstral coefficients MCEPs, fundamental frequency F0, and non-periodic feature APs, from the speech data of the small sample training set and the test set, clean and normalize the extracted feature data, eliminate noise, and unify the feature scale to obtain the pre-processed speech data features. Among them, the Mel cepstral coefficients MCEPs use neural networks for feature migration in subsequent operations, the fundamental frequency F0 is normalized by Gaussian logarithm and then participates in the final target timbre speech synthesis, and the non-periodic feature APs are directly retained to participate in the final target timbre speech synthesis;
[0006] Step 3: Use the long short-term memory network model (LSTM network) as the coarse transfer network. Input the MCEPs features of the preprocessed small-sample training set speech pairs, and obtain the intermediate timbre features after coarse transfer through temporal modeling transfer. The pre-trained model is used to approximately approach the target timbre distribution initially;
[0007] Step 4: Design a generative adversarial network (GAN) based on LSTM as the fine transfer network. The generative adversarial network consists of two parts: a generator and a discriminator. Here, both the generator and the discriminator in the generative adversarial network adopt a multi-layer LSTM structure, and the gating mechanism of the LSTM network is used to capture the long-term dependencies of the speech signal; introduce a module transfer strategy to synchronously transfer the weights of the first two layers of LSTM of the generator to the corresponding layers of the discriminator, and the weights of the corresponding layers of the discriminator are updated with the training of the generator to ensure the consistency of the underlying feature extraction of both; use the intermediate timbre features output by the coarse transfer network and the MCEPs features in the target speaker's voice timbre features in the small-sample training set speech pairs as inputs for supervised learning;
[0008] Step 5: Fix the parameters of the coarse transfer network, jointly train the fine transfer generator and discriminator, and optimize the timbre similarity of the generated speech through adversarial learning until the generative adversarial network model in the fine transfer network converges;
[0009] Step 6: Input the MCEPs features of the source speaker's voice data in the small-sample test set into the trained network to generate the MCEPs features of the target timbre voice. Combine the fundamental frequency F0 after Gaussian logarithmic normalization and the aperiodic features APs directly retained, and use a vocoder to reconstruct the waveform to obtain the converted target timbre voice data. Finally, the identity of the generated speech and the target timbre can be determined through Mel Cepstral Distortion (MCD), cosine similarity, and subjective listening tests, and the speech timbre conversion task is completed.
[0010] Furthermore, in Step 1, by systematically collecting and organizing speech data, a training set and a test set suitable for the small-sample timbre conversion task are constructed; the speech data is sourced from public speech databases (such as CMU-ARCTIC) or self-built recording libraries, and professional recording equipment is used to record the speech; the speech sampling rate is 16 kHz, the quantization precision is 16 bit, stored in mono, and the duration of each speech segment is 4 - 5 seconds, with the format being WAV;
[0011] The collected speech data set is divided into the following three categories:
[0012] 1) Complete sample parallel training set: It contains 1000 speech segments for each speaker, constructing speech pairs with a one-to-one correspondence between the source speaker and the target speaker, totaling 200 pairs;
[0013] 2) Small sample training set: Randomly select 20 pairs of voices from the above complete sample parallel training set to simulate the scenario of scarce data of the target speaker;
[0014] 3) Small sample test set: Select an additional 20 pairs of voices that have not participated in the training from the above complete sample parallel training set for model performance verification.
[0015] Furthermore, extract key voice features from all the original voice data in the small sample training set and test set. The extracted voice features include Mel Cepstral Coefficients (MCEPs), fundamental frequency (F0), and aperiodic features (APs). Use linear interpolation to fill in the missing frames and perform data cleaning. The formula is as follows:
[0016]
[0017] where x t-1 , x t+1 are the feature values of the adjacent frames before and after the missing frame; outliers are detected and removed by the standard deviation method, and those deviating from the mean by ±3σ are considered outliers; x fill is the voice feature data after interpolation.
[0018] Furthermore, use the Gaussian logarithmic normalization method to normalize the fundamental frequency F0. The formula is as follows:
[0019]
[0020] where μ is the mean of the logarithmic fundamental frequency and σ is the standard deviation to eliminate the pitch differences between speakers.
[0021] Furthermore, use the LSTM network as a rough transfer network to approximately approach the target timbre distribution. The LSTM network captures the long-term dependencies of the voice signal through the gating mechanism. The specific process is as follows:
[0022] The training set of the MCEPs feature dataset of the input voice data at time t of the LSTM is x t , and the output values are h t , c t is the memory state; the LSTM memory unit is updated as follows:
[0023] (1) Calculation of the forgetting gate f t :
[0024] f t =σ(w f ×[h t-1 , x t +b f )
[0025] where f is the weight matrix of the forgetting gate f t at time t, ht-1 is the hidden state of the previous time step, w f is the forget gate weight matrix, b f is the bias term, and σ is the Sigmoid activation function;
[0026] (2) Calculation of the input gate i t :
[0027] i t = σ(w i × [h t-1 , x t + b i )
[0028] where w i is the weight matrix of the input gate i at time t t , b i is the bias;
[0029] (3) Calculation of the candidate state of the memory cell :
[0030]
[0031] where w c is the weight matrix of the candidate state at time t, b c is the bias;
[0032] (4) Update calculation of the memory cell state value c t :
[0033]
[0034] (5) Calculation of the output gate o t :
[0035] o t = σ(w0 × [h t-1 , x t + b0)
[0036] where w0 is the weight matrix of the output gate o t at time t, and b0 is the bias;
[0037] (6) Calculation of the hidden layer output value h t :
[0038] h t = o t × tanh(c t )
[0039] To solve the possible overfitting problem, the Dropout technique is introduced.
[0040] During the backpropagation of the coarse transfer LSTM network, the mean square error (MSE) is used to measure the spectral difference between the intermediate features and the target speaker's MCEPs. The formula is as follows:
[0041]
[0042] where h mid is the output of the coarse transfer network, and m target is the MCEPs feature of the target speaker;
[0043] During the backpropagation process, the intermediate variables and gradient parameters of the neural network are calculated and stored, and the relevant parameters are updated on the neurons.
[0044] Furthermore, a generative adversarial network (GAN) based on LSTM is used as the fine transfer network to refine the target tone features, and a module transfer strategy is combined to balance the training process of the generator and the discriminator. The specific process is as follows:
[0045] The generator G is constructed by 4 layers of unidirectional LSTM. Taking the intermediate tone features output by the coarse transfer network and the MCEPs features in the target speaker's voice tone features in the small sample training set speech pair as inputs, the output of each layer is mapped to the MCEPs of the target speaker through a fully connected layer, and finally the MCEPs sequence of the target tone is generated. The formula is as follows:
[0046] m gen = G(h mid ; W g )
[0047] where W g are the weight parameters of the generator, h mid is the intermediate feature output by the coarse transfer network, and m gen is the MCEPs sequence generated by the generator;
[0048] The loss function of the generator includes the adversarial loss and the feature matching loss. The expression of the total loss function is as follows:
[0049]
[0050] where λ is the balancing weight, m target is the MCEPs feature sequence in the target speaker's voice tone features in the small sample training set speech pair. The specific expression of the feature matching loss (L MCD ) is as follows:
[0051]
[0052] The discriminator D is constructed, including 3 layers of LSTM, and the output is mapped to the authenticity probability value through a fully connected layer:
[0053] D(m; W d ) = σ(W d ·h lstm +b d )
[0054] where m is the sequence of MCEPs input to the discriminator, W d is the discriminator weight parameter, h lstm is the hidden state of the last layer of LSTM, σ is the Sigmoid function, and b d is the bias term of the output layer of the discriminator;
[0055] The loss function of the discriminator is as follows:
[0056]
[0057] When constructing the generator and the discriminator, the weights of the first two layers of LSTM of the generator are migrated to the corresponding layers of the discriminator and the parameter updates of the migration module are frozen, and only backpropagation is performed during the training of the generator:
[0058]
[0059] where η is the learning rate, is the weight of the network layer of the migration part in the generator, is the weight of the network layer of the migration part in the discriminator.
[0060] Furthermore, the model parameters are optimized through a staged training strategy. First, the coarse migration network is independently trained until convergence, and then its parameters are fixed and the fine migration generator and discriminator are jointly trained to ensure the stable convergence of the generative adversarial network model under small sample conditions. The specific process of training and updating the weight values and biases of the neural network is as follows:
[0061] During the training of the coarse migration network and the fine migration network, the Adam (Adaptive Moment Estimation) optimizer is used; the specific update rules of the Adam optimizer are as follows:
[0062] m t = β1m t-1 +(1 - β1)g t
[0063]
[0064] where g t is the gradient of the parameter, β1 and β2 are the decay coefficients of the two exponentially weighted moving averages, and are the corrected moving averages of the gradient bias, and θt+1 is the updated parameter, η is the learning rate, and ∈ is a very small constant to avoid zero as the divisor;
[0065] In the fine-tuning transfer network, the generator and discriminator are trained using an alternating update strategy, that is, in each round of iteration, the discriminator parameters are updated first and then the generator parameters are updated; and an early stopping strategy is introduced in the training. If the L of the generator MCD has a fluctuation amplitude less than 1% for 10 consecutive epochs, or the total number of training epochs reaches 500, then the training is terminated.
[0066] Furthermore, an identity comparison index and calculation are performed on the generated speech and the target speech. The specific process is as follows:
[0067] The objective evaluation metrics selected are Mel Cepstral Distortion (MCD) and cosine similarity. Mel Cepstral Distortion measures the difference between the generated speech and the target speech in the spectral envelope. The calculation formula is:
[0068]
[0069] where is the MCEPs sequence generated by the generator after training is completed, is the MCEPs feature sequence in the target speaker's voice timbre features in the small sample test set speech pair.
[0070] Cosine similarity calculates the cosine similarity between the generated MCEPs and the target MCEPs in the feature space. The calculation formula is:
[0071]
[0072] In order to subjectively verify the results of the previous MCD and cosine similarity objective metrics and measure the naturalness of the generated speech, the subjective Mean Opinion Score (MOS) is used for subjective evaluation and verification: the naturalness and timbre consistency of the generated speech are evaluated through artificial listening tests. Several subjects are invited to rate the generated speech on a 5-point scale (1: extremely poor, 5: excellent), and the average score is calculated. The scoring criteria include timbre matching similarity and naturalness.
[0073] A small-sample voice timbre conversion method and system based on a double-layer transfer generative adversarial network proposed by the present invention divides the key features of voice data into temporal dynamic features (Mel cepstral coefficients, MCEPs) and auxiliary acoustic parameters (fundamental frequency F0 and aperiodic features APs), combines the advantages of a long short-term memory network (LSTM) and a generative adversarial network (GAN), and realizes efficient and high-fidelity timbre transfer. The LSTM network captures the long-term temporal dependence relationship of the voice signal through a gating mechanism and extracts the timbre-related spectral envelope characteristics from the dynamic features; the auxiliary acoustic parameters (F0 and APs) directly retain the prosody and naturalness characteristics of the source voice during the synthesis stage, ensuring the intonation coherence and sound source authenticity of the generated voice. Through a double-layer transfer architecture (the coarse transfer network pre-trains the underlying features, and the fine transfer network makes fine adjustments) and a module transfer strategy (synchronizing the underlying weights of the generator and the discriminator to balance the training degrees of the two), the model significantly improves the training stability and generation quality under small-sample conditions. The continuous offline training and optimization process enables the model to continuously update the network weights, and the finally generated voice is superior to traditional methods in terms of timbre similarity and naturalness, providing reliable technical support for scenarios such as judicial voice identity identification and virtual character timbre cloning. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is a schematic flow chart of a small-sample voice timbre conversion method based on a double-layer transfer generative adversarial network according to an embodiment of the present invention;
[0075] Figure 2 It is a schematic diagram of the process of an LSTM neural network model according to an embodiment of the present invention;
[0076] Figure 3 It is a schematic diagram of a dropout method for avoiding data overfitting according to an embodiment of the present invention;
[0077] Figure 4 It is a schematic diagram of the process of a module transfer strategy according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0079] As Figure 1 shown, the present invention provides a small-sample voice timbre conversion method and system based on a double-layer transfer generative adversarial network, including the following steps:
[0080] Step 1: Collect the speech data of the source speaker and the target speaker to form speech pair data, and construct a complete sample parallel training set, a small sample training set, and a test set based on the speech pair data to simulate the voice conversion scenarios under different data volumes; the complete sample parallel training set contains sufficient speech pairs, and the small sample training set contains a small number of speech pairs;
[0081] Step 2: Extract three types of features, namely Mel Cepstral Coefficients (MCEPs), fundamental frequency (F0), and aperiodic features (APs) from the speech data of the small sample training set and the test set. Clean and normalize the extracted feature data to eliminate noise and unify the feature scale to obtain the preprocessed speech data features. Among them, the Mel Cepstral Coefficients (MCEPs) are used for feature migration using a neural network in subsequent operations, the fundamental frequency (F0) participates in the final target voice synthesis after Gaussian logarithmic normalization, and the aperiodic features (APs) are directly retained to participate in the final target voice synthesis;
[0082] Step 3: Use the Long Short-Term Memory Network (LSTM) as a rough migration network. Input the MCEPs features of the speech pairs in the preprocessed small sample training set, and obtain the intermediate voice features after rough migration through temporal modeling migration. The pre-trained model is used to initially approximate the target voice distribution;
[0083] Step 4: Design a Generative Adversarial Network (GAN) based on LSTM as a fine migration network. The Generative Adversarial Network consists of two parts: a generator and a discriminator. Here, both the generator and the discriminator in the Generative Adversarial Network adopt a multi-layer LSTM structure, and the gating mechanism of the LSTM network is used to capture the long-term dependencies of the speech signal; introduce a module migration strategy to synchronously migrate the weights of the first two layers of the LSTM in the generator to the corresponding layers of the discriminator, and the weights of the corresponding layers of the discriminator are updated with the training of the generator to ensure the consistency of the underlying feature extraction of the two; use the intermediate voice features output by the rough migration network and the MCEPs features in the target speaker voice features of the speech pairs in the small sample training set as inputs for supervised learning;
[0084] Step 5: Fix the parameters of the rough migration network, jointly train the fine migration generator and discriminator, and optimize the voice similarity of the generated speech through adversarial learning until the Generative Adversarial Network model in the fine migration network converges;
[0085] Step 6: Input the MCEPs features of the source speaker speech data in the small sample test set into the trained network to generate the MCEPs features of the target voice. Combine the fundamental frequency (F0) after Gaussian logarithmic normalization and the directly retained aperiodic features (APs), and use a vocoder to reconstruct the waveform to obtain the converted target voice data. Finally, the identity of the generated speech and the target voice can be determined through Mel Cepstral Distortion (MCD), cosine similarity, and subjective listening tests to complete the voice conversion task.
[0086] In the described Step 1, the specific processes of voice data collection, dataset construction, and data splitting into a training set and a test set are as follows:
[0087] In this step, a training set and a test set suitable for the small-sample voice conversion task are constructed by systematically collecting and organizing voice data. In specific implementation, the voice data can be sourced from public voice databases (such as CMU-ARCTIC) or self-built recording libraries, and professional recording equipment is used to record the voice. The voice sampling rate is 16 kHz, the quantization precision is 16 bit, it is stored in mono, and the duration of each voice segment is 4 - 5 seconds, with the format being WAV. The dataset splitting uses the train_test_split method to randomly divide the sample set into a training set and a test set, and return the divided training set and test set data. The specific dataset splitting is as follows: The complete sample parallel training set contains 1000 voice segments for each speaker, and voice pairs corresponding one-to-one between the source speaker and the target speaker are constructed (a total of 200 pairs); The small-sample training set randomly selects 20 voice pairs from the complete sample parallel training set to simulate the scenario of scarce data of the target speaker; The test set additionally selects 20 voice pairs that have not participated in the training from the complete sample parallel training set for model performance verification.
[0088] In the described Step 2, the process of extracting three key features, namely Mel Cepstral Coefficients (MCEPs), fundamental frequency (F0), and aperiodic features (APs) from the voice signal, and cleaning and normalizing the data to eliminate noise and unify the feature scale is as follows:
[0089] Extracting signal features: Mel Cepstral Coefficients reflect the spectral envelope characteristics of the voice signal and are used to characterize the timbre features. The short-time Fourier transform spectrum of the voice is smoothed through a Mel filter bank, and after dimensionality reduction by the discrete cosine transform (DCT), 24-dimensional coefficients are extracted as the main training features; The fundamental frequency represents the vocal cord vibration frequency and determines the pitch characteristics of the voice. It is extracted from the voice signal using the DIO algorithm and is logarithmically Gaussian-normalized to eliminate differences between speakers; The aperiodic features describe the non-periodic components of vocal cord vibration and are used to improve the naturalness of the synthesized voice. 5-dimensional parameters are extracted through the WORLD vocoder and directly normalized to the interval [-1, 1]. MCEPs are selected as the core training features. In subsequent voice conversion using a neural network, the converted MCEPs are used for waveform reconstruction, and F0 and APs are used as auxiliary features. After Gaussian logarithmic normalization conversion of F0, it is used together with APs in the waveform reconstruction stage of the vocoder.
[0090] Missing value processing: According to the characteristics of the dataset, the impact of missing values in the dataset on the analysis is reduced by deleting the features containing missing values, and then the missing values are replaced with appropriate estimated values of known data using linear interpolation to perform data cleaning. The formula is as follows:
[0091]
[0092] where x t-1 and x t+1 are the eigenvalue of the adjacent frames before and after the missing frame. Outliers are detected by the standard deviation method (deviating from the mean by ±3σ is considered an outlier) and removed. x fill is the speech feature data after interpolation is completed..
[0093] Outlier detection: Identify outliers that are inconsistent with most of the data in the dataset through standard deviation and scatter plots, and use outlier detection methods based on cluster analysis to detect and process these outliers to ensure the quality and reliability of the data.
[0094] Normalization: To make the features have the same measurement scale and the distribution of the fundamental frequency (F0) more conform to the assumptions of model training, we use the Gaussian logarithmic normalization method for normalization, and its formula is as follows:
[0095]
[0096] where μ is the mean of the logarithmic fundamental frequency, and σ is the standard deviation to eliminate the pitch difference between speakers. Through Gaussian logarithmic normalization, the transformed F0 can follow the standard normal distribution, and the influence of extremely high fundamental frequency values is suppressed through logarithmic transformation, enhancing the adaptability of the model to pitch fluctuations.
[0097] In step 3 described above, a coarse transfer network is built using a long short-term memory network (LSTM) to perform the first step of processing on the MCEPs features of the speech signal. By inputting the MCEPs features of the preprocessed small sample training set speech pairs into the LSTM model, the process of deeply modeling and learning the time series data to initially capture its dependency relationship and time dynamic characteristics is as follows:
[0098] Input the processed MCEPs feature data into the LSTM model to screen the past information of the individual, thus completing the long-term memory of the effective information.
[0099] Each LSTM memory cell contains 3 control gates, namely the forget gate f t , the input gate i t and the output gate o t , and its structure is as Figure 2 shown. The input data of LSTM at time t is x t , and the output value is h t , and c t is the memory state. The LSTM memory cell is updated as follows:
[0100] The forget gate f tIt mainly depends on how much information is forgotten from the memory cell state, which is determined by the input value x at time t t and the hidden layer output h at time t-1 t-1 jointly. The forget gate f t is calculated as follows:
[0101] f t =σ(w f ×[h t-1 ,x t +b f )
[0102] where f is the forget gate f at time t t 's weight matrix, b f is the bias, and σ uses the Sigmoid function.
[0103] The input gate i t mainly determines how much of the current input data is input into the memory cell, which is determined by the input value x at time t t and the hidden layer output h at time t-1 t-1 jointly. The input gate i t is calculated as follows:
[0104] i t =σ(w i ×[h t-1 ,x t +b i )
[0105] where w i is the weight matrix of the input gate i at time t t , b i is the bias.
[0106] The candidate state of the memory cell is determined by the input value x at time t t and the hidden layer output h at time t-1 t-1 jointly. The candidate state of the memory cell is calculated as follows:
[0107]
[0108] where w c is the weight matrix of the candidate state at time t , b c is the bias.
[0109] The memory cell state value c t passes through the input gate i t and the forget gate f t to its own state c t-1 and the current candidate memory state value is adjusted to update the memory cell state. The memory cell state value c t The update calculation formula is:
[0110]
[0111] The output gate o t is mainly used to control how much of the memory cell state value needs to be output, which is jointly determined by the input value x at time t t and the hidden layer output h at time t - 1 t-1 The output gate o t The calculation formula is:
[0112] o t = σ(w0 × [h t-1 , x t + b0)
[0113] where w0 is the weight matrix of the output gate o at time t t and b0 is the bias.
[0114] The hidden layer output value h t is jointly determined by the output value o at time t t and the memory cell state value c t The calculation formula of the hidden layer output value h t is:
[0115] h t = o t × tanh(c t )
[0116] In the process of temporal feature modeling, LSTM dynamically adjusts the information storage and update within the cell state by integrating the current input, historical hidden state, cell state, and gating mechanism, in order to depict the temporal evolution law of dynamic features. The gating mechanism precisely regulates the retention and new addition ratio of information through weighted calculation: the forget gate determines the attenuation degree of historical features, and the input gate filters out the effective information of the current input. The update of the cell state consists of two parts - the historical state attenuated by the forget gate and the newly added information weighted by the input gate. Finally, the current hidden state is generated based on the output gate as the temporal feature representation for the next moment.
[0117] The coarse transfer network adopts a stacked structure of four LSTM units, and dynamically optimizes the number of hidden units in each layer through training to explore the best parameter combination. To alleviate the overfitting risk caused by the increase in network depth, the Dropout technique is introduced. The schematic diagram of its principle is as Figure 3 shown, and the implementation process is as follows:
[0118] (1) Random masking setting: Set the masking probability P (usually taken as 0.2 - 0.5). Each layer of neurons is temporarily masked during training according to this probability. During the masking period, they do not participate in forward calculation and parameter update, but their input and output paths remain intact.
[0119] (2) Forward calculation and gradient backpropagation: The input data undergoes forward propagation through the masked LSTM network to calculate intermediate variables and prediction results; based on the cross-entropy loss function (formula as follows), perform backpropagation and only update the parameters of the unmasked neurons. The formula for the cross-entropy loss function is as follows:
[0120] l = -[ylogy^+(1 - y)log(1 - y^)]
[0121] where y is the true sample label, taking values of 0 or 1. y^ is the probability that the model predicts the sample as the positive class.
[0122] (3) Parameter recovery and iteration: After a single training is completed, restore the activity of all neurons, inherit the updated parameters, and loop through the above steps until convergence. Then, repeat steps 1 to 2 to continue training the model.
[0123] In this way, during training, the Dropout technique can be used to randomly discard a part of the neurons, thereby reducing the complexity of the model and the risk of overfitting. In the testing phase, the Dropout technique is not applied, but the outputs of all neurons are retained to maintain the stability of the model.
[0124] In step 4 described above, a generative adversarial network (GAN) based on LSTM is designed as the fine transfer network to further capture the temporal dependencies in the speech data and perform precise feature transfer; a module transfer strategy is introduced during training to synchronously transfer the weights of the first two LSTM network layers of the generator to the corresponding layers of the discriminator, and the weights of the corresponding layers of the discriminator are updated along with the training of the generator to balance the training degree of the two. The specific implementation process is as follows:
[0125] Generator: A 4-layer unidirectional LSTM constructs the generator G, with 128 hidden units in each layer and a time step of 40 frames. Using the intermediate timbre features output by the coarse transfer network as the input, the output of each layer is mapped to the MCEPs of the target speaker through a fully connected layer, and finally an MCEPs sequence of the target timbre is generated. The formula is:
[0126] m gen = G(h mid ; W g )
[0127] where W g are the weight parameters of the generator, h mid is the intermediate feature output by the coarse transfer network, and m genThe MCEPs sequence generated by the generator;
[0128] The loss function of the generator includes adversarial loss and feature matching loss. The expression of the total loss function is as follows:
[0129]
[0130] where λ is the balance weight, and m target is the MCEPs feature sequence in the voice timbre feature of the target speaker in the small sample training set voice pair. The specific expression of the feature matching loss (L MCD ) is as follows:
[0131]
[0132] Discriminator: Build a discriminator D, which includes 3 layers of LSTM, with 128 hidden units in each layer and a time step of 40 frames. The output is mapped to a authenticity probability value through a fully connected layer, and the output authenticity probability D(m) ∈ [0, 1]:
[0133] D(m; W d ) = σ(W d ·h lstm + b d )
[0134] where m is the MCEPs sequence input to the discriminator, W d is the weight parameter of the discriminator, h lstm is the hidden state of the last layer of LSTM, σ is the Sigmoid function, and b d is the bias term of the output layer of the discriminator;
[0135] The loss function of the discriminator is as follows:
[0136]
[0137] Module migration strategy: When building the generator and the discriminator, migrate the weights of the first two layers of LSTM of the generator to the corresponding layers of the discriminator and freeze the parameter update of the migrated module (only backpropagate with the training of the generator). The schematic diagram of its principle is as Figure 4 shown:
[0138]
[0139] where η is the learning rate, is the weight of the network layer of the module migration part in the generator, is the weight of the network layer of the module migration part in the discriminator.;
[0140] In step 5, the model parameters are optimized through a phased training strategy. First, the coarse transfer network is independently trained until convergence. Subsequently, its parameters are fixed and the fine transfer generator and discriminator are jointly trained to ensure the stable convergence of the generative adversarial network model under the condition of few samples. The specific process of training and updating the weight values and biases of the neural network is as follows:
[0141] During the training of the coarse transfer network and the fine transfer network, the Adam (Adaptive Moment Estimation) optimizer is used; the specific update rules of the Adam optimizer are as follows:
[0142] m t = β1m t-1 + (1 - β1)g t
[0143]
[0144] where g t is the gradient of the parameter, β1 and β2 are the decay coefficients of the two exponentially weighted moving averages, and are the corrected moving averages of the gradient bias, θ t+1 is the updated parameter, η is the learning rate, and ∈ is a very small constant to avoid zero as the divisor.
[0145] In the fine transfer network, the training of the generator and the discriminator adopts an alternating update strategy, that is, in each round of iteration, the discriminator parameters are updated first and then the generator parameters are updated. And an early stopping strategy is introduced in the training. If the fluctuation range of the L MCD of the generator is less than 1% for 10 consecutive epochs, or the total number of training epochs reaches 500, then the training is terminated.
[0146] In step 6, the MCEPs features of the source speaker speech data in the few-sample test set are input into the trained network to generate the MCEPs features of the target timbre speech. Combining the fundamental frequency F0 after Gaussian logarithmic normalization and the aperiodic features APs directly retained, the waveform is reconstructed using a vocoder to obtain the converted target timbre speech data. The specific process is as follows:
[0147] The MCEPs generated by the fine transfer network (m gen)Perform parameter fusion with the fundamental frequency (F0) after Gaussian normalization and the aperiodic features (APs) of the source speech to construct the complete acoustic parameters of the target timbre. Generate the spectral envelope of the target timbre according to the MCEPs, which determines the formant positions and energy distribution; generate a periodic pulse sequence through the fundamental frequency F0 to simulate vocal cord vibration; add a noise component based on the aperiodic features APs to enhance the naturalness of the speech. Subsequently, use the WORLD vocoder to convert the fused parameters into a time-domain speech waveform to obtain the speech data of the target timbre after timbre conversion. WORLD is based on the source-filter model, and its synthesis formula is:
[0148]
[0149] where s(t) is the synthesized speech signal, A k (t) is the amplitude of the k-th harmonic (calculated from MCEPs and APs), f0(t) is the fundamental frequency trajectory, and φ k (t) is the phase modulation term (controlled by the aperiodic parameters of APs).
[0150] Furthermore, perform identity comparison metrics and calculations on the generated target speech, and the specific process is as follows:
[0151] The objective evaluation metrics select the Mel Cepstral Distortion (MCD) and cosine similarity. The Mel Cepstral Distortion measures the difference in the spectral envelope between the generated speech and the target speech, and its calculation formula is:
[0152]
[0153] where is the MCEPs sequence generated by the generator after training, and is the MCEPs feature sequence in the target speaker's speech timbre features in the small sample test set speech pair.
[0154] The cosine similarity calculates the cosine similarity between the generated MCEPs and the target MCEPs in the feature space, and its calculation formula is:
[0155]
[0156] To verify the results of the above MCD and cosine similarity objective metrics through subjective evaluation and measure the naturalness of the generated speech, the subjective Mean Opinion Score (MOS) is used for subjective evaluation verification: evaluate the naturalness and timbre similarity of the generated speech through human listening tests. Each converted speech is randomly played to the listeners, and they evaluate it on a 5-point scale, namely 1: lowest, 2: poor, 3: average, 4: good, 5: excellent, so as to compare the differences in naturalness and similarity of each speech. The higher the MOS value score, the better the subjective performance of the model.
[0157] The small-sample voice timbre conversion method and system based on the double-layer transfer generative adversarial network proposed by the present invention have important application values in the fields of judicial voice identification, virtual character interaction, medical auxiliary voice generation, and film and television dubbing. By generating high-fidelity target timbre voices, it can provide supplementary and comparative bases for scarce voice samples for judicial institutions, significantly improve the accuracy and efficiency of voice identity identification, and reduce the risk of misjudgment caused by insufficient samples. In the fields of virtual characters and intelligent assistants, this method can quickly clone personalized timbres and adapt to multi-language scenarios, enhance the immersion of user interaction, reduce the dependence on real voice actors, and greatly reduce the timbre customization cost. For the medical auxiliary scenario, this technology can generate natural voices based on the timbres of relatives for aphasic patients to help them restore their communication abilities, and dynamically adjust voice parameters in combination with rehabilitation needs to improve the personalization and humanization level of rehabilitation treatment. In addition, in the field of film and television dubbing, by efficiently generating dubbing content with multiple languages and timbres, it can accelerate the localization process of film and television works and provide technical support for the restoration of the timbres of historical figures, contributing to the digital protection of cultural heritage.
[0158] To ensure the long-term effectiveness of the technology, it is recommended to continuously collect user feedback and optimize model parameters in actual applications, and establish a cross-domain voice database to enhance the adaptability to complex scenarios. In data usage, it is necessary to strictly follow ethical norms to ensure the authorization and anonymization of voice data. Especially in sensitive scenarios such as justice and medical care, a third-party review mechanism should be introduced to prevent the abuse of technology. In the future, by promoting the construction of a standardized evaluation system and industry norms for voice synthesis technology, this method is expected to be further extended to fields such as barrier-free communication and intelligent education, providing core support for the popularization and sustainable development of artificial intelligence voice technology, and ultimately achieving the long-term goal of empowering society with technology.
[0159] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A small sample speech timbre conversion method based on a double-layer transfer generative adversarial network, characterized in that: The following steps are involved: Step 1: Collect speech data of the source speaker and the target speaker to form speech pair data, and construct a complete sample parallel training set, a small sample training set and a test set based on the speech pair data; the complete sample parallel training set includes multiple groups of speech pairs, and the small sample training set and the test set are composed of some speech pairs selected from the complete sample parallel training set; Step 2: Extract three types of features, namely, Mel-frequency cepstral coefficients MCEPs, fundamental frequency F0, and aperiodic features APs, from the speech data of the small sample training set and the test set, clean and normalize the extracted feature data, eliminate noise, and unify the feature scale to obtain the preprocessed speech data features; The Mel-frequency cepstral coefficients MCEPs are used for feature migration in subsequent operations using a neural network, the fundamental frequency F0 is normalized by Gaussian logarithm and then participates in the final target timbre speech synthesis, and the non-periodic features APs are directly retained to participate in the final target timbre speech synthesis; Step 3: Use the long short-term memory network model LSTM network as the coarse transfer network, input the MCEPs features of the pre-processed small sample training set speech pairs, and obtain the intermediate timbre features after coarse transfer through time series modeling migration. The pre-training model is used to preliminarily approximate the target timbre distribution; Step 4: Design a LSTM-based generative adversarial network (GAN) as a fine-tuning transfer network; The generative adversarial network consists of two parts: the generator and the discriminator. Both the generator and the discriminator adopt a multi-layer LSTM structure and use the gating mechanism of the LSTM network to capture the long-term dependency of speech signals. A module migration strategy is introduced to synchronously migrate the weights of the first two layers of LSTM of the generator to the corresponding layers of the discriminator. The weights of the corresponding layers of the discriminator are updated with the training of the generator to ensure the consistency of the underlying feature extraction of the two. The intermediate timbre features output by the rough transfer network and the MCEPs features in the target speaker's speech timbre features in the speech pairs of the small sample training set are used as input for supervised learning; Step 5: Fix the parameters of the coarse transfer network, jointly train the generator and discriminator of the fine transfer network, and optimize the timbre similarity of the generated speech through adversarial learning until the generative adversarial network model in the fine transfer network converges; Step 6: Input the MCEPs features of the source speaker's speech data in the small sample test set into the generators in the trained coarse transfer network and fine transfer network to generate the MCEPs features of the target timbre speech. Combined with the fundamental frequency F0 after Gaussian logarithm normalization and the directly retained non-periodic features APs, the waveform is reconstructed using the vocoder to obtain the converted target timbre speech data. Finally, the Mel-cepstral distortion MCD, cosine similarity and subjective listening test are used to determine the identity of the generated speech with the target timbre, thus completing the speech timbre conversion task.
2. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: In step 1, by systematically collecting and organizing speech data, a training set and a test set suitable for the small sample timbre conversion task are constructed; the speech data comes from a public speech database or a self-built recording library, and the speech is recorded using a recording device; the speech sampling rate is 16kHz, the quantization accuracy is 16bit, the storage is monophonic, the duration of each speech segment is 4-5 seconds, and the format is WAV; the collected speech data sets are divided into the following three categories: 1) Complete sample parallel training set: contains 1,000 speech segments for each speaker, and constructs speech pairs with one-to-one correspondence between the source speaker and the target speaker, totaling 200 pairs; 2) Small sample training set: 20 pairs of speech are randomly selected from the above complete sample parallel training set to simulate the scenario where the target speaker data is scarce; 3) Small sample test set: 20 additional speech pairs that did not participate in the training were selected from the above-mentioned complete sample parallel training set for model performance verification.
3. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: Extract key speech features from all the original speech data in the small sample training set and test set. The extracted speech features include Mel-frequency cepstral coefficients MCEPs, fundamental frequency F0 and non-periodic features APs. Linear interpolation is used to fill in missing frames and perform data cleaning. The formula is as follows: Among them, x t-1 、x t+1 is the feature value of the adjacent frames before and after the missing frame; outliers are detected and eliminated by the standard deviation method, and deviations from the mean ±3σ are considered abnormal; x fill It is the speech feature data after interpolation.
4. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: The Gaussian logarithmic normalization method is used to normalize the fundamental frequency F0, and the formula is as follows: Where μ is the mean of the logarithmic fundamental frequency and σ is the standard deviation to eliminate the pitch differences between speakers.
5. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: The LSTM network captures the long-term dependencies of speech signals through a gating mechanism. The specific process is as follows: The training set of LSTM’s input speech data MCEPs feature data set at time t is x t , the output value is h t ,c t is the memory state; the LSTM memory unit is updated as follows: (1) Forget gate f t Calculation: f t =σ(w f ×[h t-1 ,x t ]+b f ) Among them, f is the forget gate f at time t t The weight matrix, h t-1 is the hidden state of the previous time step, w f is the forget gate weight matrix, b f is the bias term, σ is the Sigmoid activation function; (2) Input gate i t Calculation: i t =σ(w i ×[h t-1 ,x t ]+b i ) Among them, w i is the input gate i at time t t The weight matrix, b i is the offset; (3) Candidate states of memory cells Calculation: Among them, w c is the candidate state at time t The weight matrix, b c is the offset; (4) Memory unit state value c t Update calculation of: (5) Output gate o t Calculation: about t =σ(w0×[h t-1 ,x t ]+b0) Among them, w0 is the output gate o at time t t The weight matrix of , b0 is the bias; (6) Hidden layer output value h t calculate: h t =o t ×tanh(c t )。 6. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: In the back propagation process of the rough transfer LSTM network, the mean square error MSE is used to measure the spectral difference between the intermediate features and the target speaker's MCEPs. The formula is: Among them, h mid is the output of the coarse migration network, m target is the MCEPs feature of the target speaker; During the back propagation process, the intermediate variables and gradient parameters of the neural network are calculated and stored, and the relevant parameters are updated on the neurons.
7. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: The LSTM-based generative adversarial network GAN is used as a fine transfer network to finely transfer the target timbre features, and the module transfer strategy is combined to balance the training process of the generator and the discriminator. The specific process is as follows: The generator G is constructed by a 4-layer unidirectional LSTM. The intermediate timbre features output by the coarse transfer network and the MCEPs features in the target speaker's speech timbre features in the speech pairs of the small sample training set are used as input. The output of each layer is mapped to the MCEPs of the target speaker through a fully connected layer, and finally the MCEPs sequence of the target timbre is generated. The formula is: m gen =G(h mid ;W g ) Among them, W g is the generator weight parameter, h mid is the intermediate feature output by the coarse migration network, m gen The MCEPs sequence generated by the generator; The loss function of the generator includes adversarial loss and feature matching loss. The total loss function expression is as follows: Among them, λ is the balance weight, m target is the MCEPs feature sequence in the target speaker’s speech timbre feature in the speech pair of the small sample training set, and the feature matching loss (L MCD )The specific expression is as follows: Construct the discriminator D, including 3 layers of LSTM, and map the output to the authenticity probability value through the fully connected layer: D(m;W d )=σ(W d ·h lstm +b d ) Where m is the MCEPs sequence input to the discriminator, W d is the discriminator weight parameter, h lstm is the hidden state of the last layer of LSTM, σ is the Sigmoid function, b d is the bias term of the discriminator output layer; The loss function of the discriminator is as follows: When constructing the generator and discriminator, the weights of the first two layers of LSTM of the generator are Migrate to the corresponding layer of the discriminator And freeze the parameter updates of the migration module, and only backpropagate with the generator training: Where η is the learning rate, Migrate the weights of some network layers for the modules in the generator, Transfer the weights of some network layers for the discriminator module.
8. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: The model parameters are optimized through a phased training strategy. First, the coarse migration network is trained independently until convergence. Then its parameters are fixed and the fine migration generator and discriminator are jointly trained to ensure that the generative adversarial network model converges stably under small sample conditions. The specific process of training and updating the weights and biases of the neural network is as follows: During the training of the coarse migration network and the fine migration network, the optimizer uses the Adam optimizer; the specific update rules of the Adam optimizer are as follows: m t =β1m t-1 +(1-β1)g t Among them, g t is the gradient of the parameter, β1 and β2 are the decay coefficients of the two exponentially weighted averages, and is the bias-corrected moving average of the gradient, θ t+1 is the updated parameter, η is the learning rate, ∈ is a small constant to avoid zero as a divisor; In the thin transfer network, the training of the generator and the discriminator adopts an alternating update strategy, that is, in each iteration, the discriminator parameters are updated first and then the generator parameters; and an early stopping strategy is introduced in the training. MCD If the fluctuation range is less than 1% for 10 consecutive cycles or the total training cycles reaches 500, the training is terminated.
9. A small sample speech timbre conversion method based on a double-layer migration generative adversarial network as claimed in claim 1, characterized in that: The target speech of the generated speech is compared with the target speech and the identity is calculated. The specific process is as follows: Mel-cepstrum distortion (MCD) and cosine similarity are selected as objective evaluation indicators. Mel-cepstrum distortion measures the difference between the generated speech and the target speech in the spectral envelope. The calculation formula is: in It is the MCEPs sequence generated by the generator after training. It is the MCEPs feature sequence in the target speaker's speech timbre features in the small sample test set speech pair; The cosine similarity calculation generates the cosine similarity between the MCEPs and the target MCEPs in the feature space. The calculation formula is: In order to subjectively verify the previous MCD and cosine similarity objective indicator results and measure the naturalness of the generated speech, the subjective mean opinion score MOS is used for subjective evaluation verification: the naturalness and timbre consistency of the generated speech are evaluated through manual listening tests.