Zero sample voice cloning method and device

By using text encoder and speaker encoder to obtain speech features in zero-sample speech cloning, and generating acoustic feature differences in combination with detailed encoders, the problems of high resource consumption and difficult generation in the prior art are solved, and high-quality zero-sample speech synthesis is achieved.

CN120071891AActive Publication Date: 2025-05-30INST OF ACOUSTICS CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510203115.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The prior art has problems in zero-sample voice cloning that high resource consumption, increased generation difficulty due to information compression, and inaccurate generation of samples in the absence of text and voice pairing data.

Method used

Obtain high-level pronunciation information by entering the reference audio text and target audio text into the text encoder, and inputting the speaker encoder in combination with the acoustic feature context of the reference audio to obtain high-level pronunciation information. These features are then spliced ​​and entered into the detail encoder to generate the acoustic feature difference amount and converted to the target audio by predicting the Mel spectrum.

Benefits of technology

The high-quality generation of target audio without the need for large amounts of text and speech pairing data is achieved, improving the accuracy of generating samples, and taking into account multiple key factors in speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071891A_ABST
    Figure CN120071891A_ABST
Patent Text Reader

Abstract

The invention provides a zero-sample voice cloning method and device, and the method comprises the steps: obtaining a first acoustic feature and a second acoustic feature in a text encoder and a speaker encoder, training a detail encoder through employing a flow matching method through employing the second acoustic feature, the first acoustic feature, a target Mel spectrum, and a Mel spectrum of a training reference audio, and obtaining a zero-sample voice signal. And finally obtaining a zero sample voice cloning model, inputting the reference audio of the audio to be synthesized and the audio text to be synthesized into the zero sample voice cloning model, and finally obtaining the audio to be synthesized. The method does not need a large amount of text and voice pairing data, uses the features having a clear corresponding relationship with the real voice acoustic features as the training set to train the model, improves the accuracy of sample generation, also considers a plurality of key factors in voice synthesis, including text content, speaker features and voice rhythm information, and improves the accuracy of voice synthesis. Through an advanced neural network structure and a training strategy, high-quality zero-sample speech synthesis is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the technical field of speech synthesis, and in particular, to a zero-shot voice cloning method and apparatus. Background Art

[0002] Zero-shot voice cloning, as an important branch of multi-speaker speech generation, enables a model to clone or imitate a speaker's voice by referring to a reference speech without specific speaker samples. In the prior art, either a limited number of speaker lists are used to maintain the model, and then the model is fine-tuned through thousands of steps of training, which is labor-intensive and resource-consuming; or in order to make the generated samples more realistic in style and timbre, a high-dimensional global speaker embedding (GSE) that stores a large amount of speech feature data is used, which compresses the speaker information and increases the difficulty for the model to generate the target speaker's speech based on it; or an audio compression model is trained through a large amount of text and speech paired data, and this method requires a large amount of text and speech paired data; or when using an ordinary differential equation model as the generation model, random noise is used as the training sample, resulting in a high curvature for the difference amount predicted by the neural network, and this curvature will seriously affect the accuracy of the generated samples, resulting in inaccurate samples generated by the trained model. Summary of the Invention

[0003] This application describes a zero-shot voice cloning method and apparatus that can solve the above technical problems.

[0004] According to a first aspect, a zero-shot voice cloning method is provided. The method includes:

[0005] Input a reference audio text and a target audio text into a text encoder to obtain a first acoustic feature. The text encoder is used to predict low-level prosodic information related to the pronunciation content of the target audio text from the reference audio text, and add the low-level prosodic information to the phoneme string of the target audio text to obtain the first acoustic feature;

[0006] Input the acoustic feature context of the reference audio and the first acoustic feature into a speaker encoder to obtain a second acoustic feature. The speaker encoder is used to add high-level prosodic information related to the speaker to the first acoustic feature according to the acoustic feature context of the reference audio to obtain the second acoustic feature;

[0007] After concatenating the second acoustic feature added with noise, the first acoustic feature, the second acoustic feature, and the acoustic feature context of the reference audio, input them into a detail encoder to obtain an acoustic feature difference amount;

[0008] Based on the amount of acoustic feature difference and the second acoustic feature, a predicted Mel spectrogram is obtained, and the predicted Mel spectrogram is converted into audio to obtain the target audio corresponding to the target audio text.

[0009] Based on the above further embodiments, it further includes:

[0010] The reference audio text sample set and the target audio text sample set are input into the text encoder to obtain a first acoustic feature sample set;

[0011] The first acoustic feature sample set and the reference audio sample set are input into the speaker encoder to obtain a second acoustic feature sample set;

[0012] A noise term is added to the second acoustic feature sample set as initial data, and then the first acoustic feature sample set, the acoustic feature context of the reference audio, and the second acoustic feature sample set are concatenated as the input data of the detail encoder to train the detail encoder. The second acoustic feature sample set has a corresponding relationship with the true speech acoustic features.

[0013] Based on the above further embodiments, it further includes:

[0014] The true speech acoustic features are used as target data. The true speech acoustic features are the acoustic features extracted from the true speech signals of the target audio text sample set;

[0015] The initial data and the target data are weighted and mixed through a weight factor to obtain intermediate data, and during the training process, the weight factor is adjusted so that the intermediate data gradually reaches the target data from the initial data;

[0016] The intermediate data is concatenated with the first acoustic feature sample set, the acoustic feature context of the reference audio sample set, and the second acoustic feature sample set as the input data of the detail encoder to train the detail encoder, so that the acoustic feature difference amount output by the detail encoder, combined with the second acoustic feature sample set and the noise term, has the smallest difference from the true speech acoustic features.

[0017] Based on the above further embodiments, the step of concatenating the intermediate data with the first acoustic feature sample set, the acoustic feature context of the reference audio sample set, and the second acoustic feature sample set as the input data of the detail encoder to train the detail encoder, so that the acoustic feature difference amount output by the detail encoder, combined with the second acoustic feature sample set and the noise term, has the smallest difference from the true speech acoustic features, specifically includes:

[0018] Construct an objective function using the real speech acoustic features, the second acoustic feature sample set, and the acoustic feature difference amount. The objective function reflects the difference between the acoustic feature difference amount output by the detail encoder combined with the second acoustic feature sample set and the noise term, and the real speech acoustic features;

[0019] With the goal of minimizing the function value of the objective function, determine the parameters of the detail encoder through iterative updates.

[0020] Based on the above further embodiments, the objective function is expressed as:

[0021]

[0022] where Mel is the real speech acoustic features, F 2 is the second acoustic feature sample set, Θ(F″ t , t) is the detail encoder, F″ t is the input value of the detail encoder, t is the weight factor in the iterative process, T is the number of iterations, is the expected value;

[0023] where, F″ t = tMel+(1 - t)(F 2 + Noise)+F 1 +F 2 + Prompt

[0024] Noise is the noise term, F 1 is the first acoustic feature sample set, Prompt is the acoustic feature context of the reference audio sample set.

[0025] Based on the above further embodiments, the low-level prosodic information is low-level prosodic information, and the low-level prosodic information includes information on pitch, duration, and intensity in the text.

[0026] The high-level prosodic information is high-level prosodic features, and the high-level prosodic features include intonation, speech rhythm, syllable duration, and the offset of speech intensity affected by the speaker and emotional information.

[0027] According to a second aspect, there is provided a zero-shot voice cloning device, the device comprising:

[0028] A first processing module, configured to input a reference audio text and a target audio text into a text encoder to obtain a first acoustic feature. The text encoder is used to predict low-level prosodic information related to the pronunciation content of the target audio text from the reference audio text, and add the low-level prosodic information to the phoneme string of the target audio text to obtain the first acoustic feature;

[0029] A second processing module, configured to input the reference audio and the first acoustic feature into a speaker encoder to obtain a second acoustic feature, where the speaker encoder is configured to add high-level prosodic information related to the speaker to the first acoustic feature according to the acoustic feature context of the reference audio to obtain the second acoustic feature;

[0030] A third processing module, configured to splice the second acoustic feature added with noise, the first acoustic feature, the second acoustic feature, and the acoustic feature context of the reference audio, and then input the spliced result into a detail encoder to obtain an acoustic feature difference amount;

[0031] A fourth processing module, configured to obtain a predicted Mel spectrogram according to the acoustic feature difference amount and the second acoustic feature, and convert the predicted Mel spectrogram into audio to obtain the target audio corresponding to the target audio text;

[0032] Based on the above further embodiments, a fifth processing module is further included, specifically configured to input a reference audio text sample set and a target audio text sample set into the text encoder to obtain a first acoustic feature sample set;

[0033] Input the first acoustic feature sample set and the reference audio sample set into the speaker encoder to obtain a second acoustic feature sample set;

[0034] Add a noise term to the second acoustic feature sample set as initial data, and then splice the first acoustic feature sample set, the acoustic feature context of the reference audio, and the second acoustic feature sample set as the input data of the detail encoder to train the detail encoder, where the second acoustic feature sample set has a corresponding relationship with the real speech acoustic feature.

[0035] Based on the above further embodiments, the fifth processing module is specifically configured to use the real speech acoustic feature as target data, where the real speech acoustic feature is an acoustic feature extracted from the real speech signal of the target audio text sample set;

[0036] Perform weighted mixing on the initial data and the target data through a weight factor to obtain intermediate data, and during the training process, adjust the weight factor so that the intermediate data gradually reaches the target data from the initial data;

[0037] Splice the intermediate data with the first acoustic feature sample set, the acoustic feature context of the reference audio sample set, and the second acoustic feature sample set as the input data of the detail encoder, and train the detail encoder so that the acoustic feature difference amount output by the detail encoder, combined with the second acoustic feature sample set and the noise term, has the smallest difference from the real speech acoustic feature.

[0038] Based on the above further embodiments, the third processing module is specifically configured to construct an objective function by using the true speech acoustic features, the second acoustic feature sample set, and the acoustic feature difference amount, where the objective function reflects the difference between the acoustic feature difference amount output by the detail encoder in combination with the second acoustic feature sample set and the noise term, and the true speech acoustic features;

[0039] Taking the minimum value of the function value of the objective function as the goal, the parameters of the detail encoder are determined by updating and iterating.

[0040] Based on the above further embodiments, the objective function is expressed as:

[0041]

[0042] where Mel is the true speech acoustic features, F 2 is the second acoustic feature sample set, Θ(F″ t , t) is the detail encoder, F″ t is the input value of the detail encoder, t is the weight factor in the iterative process, T is the number of iterations, is the expected value;

[0043] where, F″ t = tMel+(1 - t)(F 2 + Noise)+F 1 +F 2 + Prompt

[0044] Noise is the noise term, F 1 is the first acoustic feature sample set, and Prompt is the acoustic feature context of the reference audio sample set.

[0045] Based on the above further embodiments, the low-level prosodic information is low-level prosodic information, and the low-level prosodic information includes information on pitch, length, and intensity in the text.

[0046] The high-level prosodic information is high-level prosodic features, and the high-level prosodic features include intonation, speech rhythm, syllable duration, and the offset of speech intensity affected by the speaker and emotional information.

[0047] According to a third aspect, a computer storage medium is provided. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by one or more processors, the zero-shot voice cloning method described in any one of the above technical solutions is implemented.

[0048] According to a fourth aspect, there is provided an electronic device including a memory and one or more processors, where a computer program is stored on the memory, and when the computer program is executed by the one or more processors, the zero-shot voice cloning method described in any one of the above technical solutions is implemented.

[0049] In the above systems and methods provided in the embodiments of the present specification, a large amount of text and speech paired data is not required. Features with a clear corresponding relationship with the acoustic features of real speech are used as the training set to train the model, improving the accuracy of the generated samples. Moreover, multiple key factors in speech synthesis are also considered, including text content, speaker characteristics, and prosody information of speech. Through an advanced neural network structure and training strategy, high-quality zero-shot speech synthesis is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0051] Figure 1 A schematic diagram showing the transmission path estimated by the neural network provided in the embodiments of the present specification;

[0052] Figure 2 A schematic diagram showing the framework of the zero-shot voice cloning model provided in the embodiments of the present specification;

[0053] Figure 3 A schematic diagram showing the performance of the neural ordinary differential equation model under different solution steps provided in the embodiments of the present specification;

[0054] Figure 4 A schematic diagram showing the flow of the zero-shot voice cloning method provided in the embodiments of the present specification;

[0055] Figure 5 A schematic diagram showing the zero-shot voice cloning device provided in the embodiments of the present specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following describes the solutions provided in the present specification with reference to the drawings.

[0057] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the drawings.

[0058] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for illustration" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration" is intended to present relevant concepts in a specific manner.

[0059] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, B exists alone, and A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more.

[0060] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise particularly emphasized in other ways.

[0061] Multi-speaker voice generation technology has been able to synthesize voices with a pronunciation quality comparable to that of humans, and this technology has been widely applied in fields such as voice assistants, audiobooks, and video dubbing. Zero-shot voice cloning, as an important branch of multi-speaker voice generation, enables the model to clone or imitate the voice of a speaker only by referring to a reference voice without specific speaker samples.

[0062] To achieve zero-shot voice cloning, the existing method one is to maintain the model using a list of a limited number of speakers and fine-tune the model through thousands of steps of training to achieve zero-shot voice cloning. However, for each generation of zero-shot voice cloning, the model needs to be fine-tuned, which is very labor-intensive and resource-consuming. To avoid fine-tuning the model, the existing method two is to extract the Global Speaker Embedding (GSE) from the reference speech. GSE can capture the unique attributes of the speaker, such as timbre, emotion, and accent. The extracted GSE is used to train the speaker recognition model to achieve zero-shot voice cloning. However, even if the dimension of GSE is higher, a whole reference audio is compressed into a vector, resulting in information compression and unable to provide sufficient speaker information for the model, making it more difficult for the model to generate speech that conforms to the style and timbre of the target speaker... The existing method three is to use the fine-grained unseen speaker information extracted from the reference speech as samples, and these features can be used as training samples to improve the performance of the model. This method can capture the unique attributes of the speaker, such as timbre, emotion, and accent, or use context learning to continue or fill in the reference speech as training samples. This method can help the model better understand and replicate the style and timbre of the speaker, but this method uses an audio compression model to process speech data and requires a large amount of paired data to learn the mapping from text to discrete speech representations.

[0063] In the case of insufficient text and speech paired data, the existing method four repairs the masked audio by learning to utilize the context and the text to be filled in the acoustic feature space of the speech. This method uses the FlowMatching (FM) technology to train an Ordinary Differential Equation (ODE) model to reconstruct the masked Mel spectrogram from random noise. It does not need to rely on a large amount of paired data to learn the mapping from text features to speech representations and can effectively handle the situation of insufficient text and speech paired data. The existing method four includes the following steps:

[0064] For a generation task, the ordinary differential equation model can be expressed as follows:

[0065] H t+1 =H t +dH t ,t∈[0,T]

[0066] The above equation simulates the continuous dynamic process from the initial state H 0 to the target state H T , and H t represents the intermediate state at time t. dH tRepresents the transmission direction from the initial state to the target state, that is, the difference amount from the initial state to the target state.

[0067]

[0068] The above equation indicates that the direction of state change is determined by the function and the function can be regarded as a direction estimator.

[0069] To fit the direction estimator, a neural network Θ(H t , t) is used to approximate The training objective is to minimize the following loss function:

[0070]

[0071] where represents the expected value, and the loss function is the integral of the square of the difference between and Θ(H t , t) over the time interval [0, T]. When the neural network Θ(H t , t) is trained, starting from the initial state H 0 , the target state H t+1 = H t + Θ(H t , t) can be obtained through iteration to get the target state H T .

[0072] Because is causal and only related to H t , it is difficult to give its specific expression. In some methods, if starting from a non-causal perspective, for the non-causal intermediate state H' t of can be clearly defined as:

[0073] H′ t = tH T + (1 - t)H 0

[0074]

[0075] Here H' t is the intermediate state at time t, is the transmission direction from H 0 to H T . Since the ODE is a causal model, the neural network with as the fitting target will give an estimated transmission direction distribution imitating the marginal probability of . This kind of The method of training for the fitting target is called FM. However, the marginal probabilities of the transmission direction estimated by the neural network and the true transmission path direction being the same does not mean that the two states are the same. This is because, in deep learning, neural networks are usually used to approximate the true distribution, but the predicted distribution of the neural network may differ from the true distribution. This difference is not only reflected in the marginal probabilities but also in the overall structure of the distribution. Sometimes, even if the marginal probabilities are the same, the overall structures of the two distributions can still be very different. Moreover, the predicted distribution of the neural network may be affected by various factors such as the model structure, training data, and optimization algorithm, and these factors can all lead to differences between the predicted distribution and the true distribution. Therefore, the fact that the marginal probabilities of the transmission direction estimated by the neural network and the true transmission path direction are the same only indicates that the probability distributions in each dimension are the same, but it does not guarantee that the two distributions are exactly the same.

[0076] To reduce the gap between the transmission path estimated by the neural network and the true transmission path, the differences between the neural network and the true path when dealing with different types of couplings are studied, such as Figure 1 Shows the differences of the neural network when dealing with different types of coupling changes. The four parts in the figure respectively correspond to H 0 and H T The coupling relationships between them, including independent coupling and repeated coupling. The upper left figure shows the independent coupling between H 0 and H T and the true transmission path dH t , and the brown line represents the independent coupling change dH 0 from H T to H t . It can be seen that the transmission paths in each dimension are relatively independent and have few intersections. The upper right figure shows the independent coupling between H 0 and H T , and the change dH 0 from H T to H t predicted by the neural network. The predicted transmission path of the neural network is also relatively straight, indicating that the neural network can learn the true transmission path more accurately. The lower left figure shows the repeated coupling between H 0 and H T and the true transmission path. It can be seen that there are more intersections between the transmission paths in different dimensions, indicating that the coupling relationship from H 0 to H T is more complex. The lower right figure shows the repeated coupling between H 0 and H T , and the change dH 0 from H T to H t。It can be seen that the transmission path predicted by the neural network appears curved, indicating that the neural network may produce errors when dealing with complex coupling relationships, resulting in a distribution gap between the predicted difference amount and the real path. In summary, the degree of intersection between the real transmission paths depends on H 0 and H T 's coupling relationship. When H 0 and H T repeatedly couple, the degree of intersection is large, and the corresponding transmission path learned by the neural network is less accurate. When the coupling relationship between H 0 and H T is more independent, the degree of intersection is smaller, and the corresponding transmission path learned by the neural network is more accurate. In the prior art, random noise is used as the initial state H 0 , and the target Mel spectrum is used as H T . However, there are many intersections between the paths of random noise and the target Mel spectrum. Therefore, random noise and the target Mel spectrum are repeatedly coupled. Using random noise as the initial state H 0 to train the neural network, the transmission paths predicted by the neural network basically all have a high curvature, and this curvature will seriously affect the accuracy of the generated samples. To solve this problem, it is necessary to design a neural network structure that can handle high-curvature paths, or adopt other technologies to reduce the curvature of the paths to improve the accuracy of the generated samples.

[0077] To solve the above problems, as Figure 2 shown, the present invention proposes a training method for a zero-shot voice cloning model, specifically including: First, input the reference audio text sample set and the target audio text sample set into the text encoder TextEncoder to obtain the first acoustic feature sample set. The first acoustic feature sample set includes the phoneme sequence and low-level prosody information of the target audio text. The low-level prosody information is the change in pitch, duration, and intensity related to the text content other than phoneme pronunciation. For example, for the pronunciation sequence X = {x 1 , x 2 , …, x i}, it can correspond to a set carrying multiple low-level prosodies where i represents time, a j represents the jth low-level prosody pattern, and Y is the first acoustic feature.

[0078] Subsequently, extract the unmasked Mel feature context Prompt from the reference audio sample set, and input Prompt and the first acoustic feature sample set into the speaker encoder SpeakerModule to obtain the second acoustic feature sample set F 2 . The second acoustic feature sample set F 2It includes the first acoustic feature and high-level prosodic information. High-level prosody refers to the prosodic features affected by paralinguistic information, such as the spectral energy distribution pattern affected by the speaker's timbre, the intonation change pattern affected by factors such as emotion and accent, the rhythm and syllable duration of speech affected by emotion and accent, the intensity and energy of speech affected by emotional state, and the speech style formed by different accents and emotional states. For example, for the pronunciation sequence Y combined with different paralinguistic information, the combined set is represented as where b n represents the nth paralinguistic information, and Z is the second acoustic feature sample set F 2 . Z does not encompass all speech components, but it is sufficient to uniquely identify a piece of speech.

[0079] Finally, during the training process of the Detial ODE model, random noise is added to the second acoustic feature sample set F 2 as the actual initial data H 0 = F 2 + Noise. The real speech acoustic feature Mel is used as the target data. The intermediate data tMel+(1 - t)(F 2 + Noise) is constructed using the weight factor t. The intermediate data, the first acoustic feature sample set F 1 , the second acoustic feature sample set F 2 , and Prompt are concatenated in the hidden dimension as the input value and input into the Detial ODE model to obtain the Mel spectrum change amount. The predicted Mel spectrum is obtained by combining the noise, the Mel spectrum change amount, and the second acoustic feature sample set F 2 . The difference between the predicted Mel spectrum and the target Mel spectrum is minimized as the fitting target to train the ODE model. Finally, the trained ODE model is obtained, and thus the zero-shot voice cloning model is obtained.

[0080] Subsequently, during the inference process using the zero-shot voice cloning model, the reference audio text, the reference audio, and the target audio text are input into the zero-shot voice cloning model to obtain the target audio.

[0081] In the above method provided in the embodiments of this specification, during the training process of the zero-shot voice cloning model, the second acoustic feature sample set with a clear correspondence to the real speech acoustic feature is used as the initial state H 0 , so that the coupling relationship between H 0 and H T is independent and the degree of intersection is small, thereby reducing the curvature of the neural network generation path and improving the accuracy of the generated samples. And the present invention also considers multiple key factors in speech synthesis, including text content, speaker characteristics, and prosodic information of speech, and realizes high-quality speech synthesis through an advanced neural network structure and training strategy.

[0082] The following will combine Figure 4 to introduce in detail a zero-shot voice cloning method of the present invention. The zero-shot voice cloning model includes a text encoder, a speaker encoder, and a detail encoder, and specifically includes the following steps:

[0083] 110. Input the reference audio text and the target audio text into the text encoder to obtain a first acoustic feature. The text encoder is used to predict low-level prosodic information related to the pronunciation content of the target audio text from the phoneme string of the reference audio text, and add the low-level prosodic information to the phoneme string of the target audio text to obtain the first acoustic feature.

[0084] 120. Input the acoustic feature context of the reference audio and the first acoustic feature into the speaker encoder to obtain a second acoustic feature. The speaker encoder is used to add high-level prosodic information related to the speaker to the first acoustic feature according to the acoustic feature context of the reference audio to obtain the second acoustic feature.

[0085] The low-level prosodic information includes low-level prosodic features, and the low-level prosodic features include information on pitch, duration, and intensity in the text.

[0086] The high-level prosodic information includes high-level prosodic features, and the high-level prosodic features include intonation, speech rhythm, syllable duration, and the offset of speech intensity affected by the speaker and emotional information.

[0087] The acoustic feature is a set of parameters that describe the physical properties and characteristics of sound, and may include frequency-related parameters, amplitude-related parameters, time-domain-related parameters, frequency-domain feature-related parameters, and energy-related features, etc. In this embodiment, it may be a Mel spectrum, and specific limitations are not made, and it can be determined according to the actual implementation situation.

[0088] In addition, in specific implementation, for example, a three-layer one-dimensional convolutional neural network with layer normalization (LayerNorm) and a layer of bidirectional recurrent memory neural network can be used as the network structure of the text encoder. Among them, the convolutional kernel size of the convolutional neural network is set to 5, and the hidden layer dimension of all intermediate layers is set to 512.

[0089] The network structure of the speaker encoder module can combine a convolutional neural network and a bidirectional recurrent memory neural network. Through such a network structure, context information can be captured from both global and local perspectives, adding speaker-related paralinguistic and related prosodic information to the first acoustic feature. For example, the speaker encoding module consists of a layer of bidirectional recurrent memory neural network and seven layers of convolutional neural network with layer normalization. Among them, the convolutional kernel size of the convolutional neural network is set to 3, the dimensions of the recurrent neural network and the first 4 consecutive convolutional layers are set to 512, the dimensions of the middle 2 convolutional layers are set to 256, and the last layer can be set to 80.

[0090] 130. After concatenating the second acoustic feature with added noise, the first acoustic feature, the second acoustic feature, and the acoustic feature context of the reference audio, input them into the detail encoder to obtain the acoustic feature difference amount.

[0091] 140. According to the acoustic feature difference amount and the second acoustic feature, obtain the predicted Mel spectrogram, convert the predicted Mel spectrogram into audio, and obtain the target audio corresponding to the target audio text.

[0092] Specifically, it further includes step 200, where step 200 is used to train the zero-shot voice cloning model:

[0093] 210. Input the reference audio text sample set and the target audio text sample set into the text encoder to obtain the first acoustic feature sample set;

[0094] 220. Input the first acoustic feature sample set and the reference audio sample set into the speaker encoder to obtain the second acoustic feature sample set;

[0095] 230. Add a noise term to the second acoustic feature sample set as the initial data, and then concatenate the first acoustic feature sample set, the acoustic feature context of the reference audio, and the second acoustic feature sample set as the input data of the detail encoder to train the detail encoder. The second acoustic feature sample set has a corresponding relationship with the real speech acoustic feature.

[0096] Specifically, it further includes step 240,

[0097] 240. Use the real speech acoustic feature as the target data. The real speech acoustic feature is the acoustic feature extracted from the real speech signal of the target audio text sample set;

[0098] 250. Perform weighted mixing on the initial data and the target data through a weight factor to obtain intermediate data, and during the training process, adjust the weight factor so that the intermediate data gradually reaches the target data from the initial data;

[0099] 260. Concatenate the intermediate data with the first acoustic feature sample set, the acoustic feature context of the reference audio sample set, and the second acoustic feature sample set as the input data of the detail encoder, and train the detail encoder so that the acoustic feature difference amount output by the detail encoder, combined with the second acoustic feature sample set and the noise term, has the smallest difference from the true speech acoustic features.

[0100] Specifically, step 260 specifically includes:

[0101] Construct an objective function using the true speech acoustic features, the second acoustic feature sample set, and the acoustic feature difference amount. The objective function reflects the difference between the acoustic feature difference amount output by the detail encoder, combined with the second acoustic feature sample set and the noise term, and the true speech acoustic features;

[0102] Aim to minimize the function value of the objective function, and determine the parameters of the detail encoder through iterative updates.

[0103] Among them, the objective function is expressed as:

[0104]

[0105] Among them, Mel is the true speech acoustic features, F 2 is the second acoustic feature sample set, Θ(F″ t , t) is the detail encoder, F″ t is the input value of the detail encoder, t is the weight factor in the iterative process, T is 1, is the expected value;

[0106] Among them, F″ t = tMel+(1 - t)(F 2 + Noise)+F 1 +F 2 + Prompt

[0107] Noise is the noise term, F 1 is the first acoustic feature sample set, and Prompt is the acoustic feature context of the reference audio sample set.

[0108] Specifically, in implementation, a Transformer with convolutional positional embeddings, ALiBi attention bias, root mean square normalization, and U - net style skip connections can be used to construct the detail encoder in the zero - shot voice cloning model. Among them, it can be set that the Transformer has 8 layers, each layer has 16 attention heads, a hidden layer dimension of 1024, and a fully - connected dimension of 4096. The entire zero - shot voice cloning model contains 117M trainable parameters.

[0109] After the detail encoder is trained, a zero-shot voice cloning model including a text encoder, a speaker encoder, and a detail encoder is obtained.

[0110] In the above method provided by the embodiments of the present specification, a large amount of text and speech paired data is used, and features with a clear corresponding relationship with the acoustic features of real speech are used as the training set to train the model, improving the accuracy of the generated samples. Moreover, multiple key factors in speech synthesis are considered, including text content, speaker characteristics, and prosody information of speech. Through an advanced neural network structure and training strategy, high-quality zero-shot speech synthesis is achieved.

[0111] The following combines Figure 5 to introduce a zero-shot voice cloning device proposed by the present invention. The zero-shot voice cloning model includes a text encoder, a speaker encoder, and a detail encoder. The device includes:

[0112] A first processing module, configured to input a reference audio text and a target audio text into the text encoder to obtain a first acoustic feature. The text encoder is used to predict low-level prosody information related to the pronunciation content of the target audio text from the reference audio text, and add the low-level prosody information to the phoneme string of the target audio text to obtain the first acoustic feature;

[0113] A second processing module, configured to input the reference audio and the first acoustic feature into the speaker encoder to obtain a second acoustic feature. The speaker encoder is used to add speaker-related high-level prosody information to the first acoustic feature according to the acoustic feature context of the reference audio to obtain the second acoustic feature;

[0114] A third processing module, configured to splice the second acoustic feature with added noise, the first acoustic feature, the second acoustic feature, and the acoustic feature context of the reference audio, and input the spliced result into the detail encoder to obtain an acoustic feature difference amount;

[0115] A fourth processing module, configured to obtain a predicted Mel spectrogram according to the acoustic feature difference amount and the second acoustic feature, and convert the predicted Mel spectrogram into audio to obtain the target audio corresponding to the target audio text.

[0116] Based on the above further embodiments, a fifth processing module is further included, specifically configured to input a reference audio text sample set and a target audio text sample set into the text encoder to obtain a first acoustic feature sample set;

[0117] Input the first acoustic feature sample set and the reference audio sample set into the speaker encoder to obtain a second acoustic feature sample set;

[0118] Add a noise term to the second set of acoustic feature samples as the initial data, and then concatenate the first set of acoustic feature samples, the acoustic feature context of the reference audio, and the second set of acoustic feature samples as the input data for the detail encoder, and train the detail encoder. The second set of acoustic feature samples has a corresponding relationship with the acoustic features of real speech.

[0119] Based on the above further embodiment, the fifth processing module is specifically configured to use the acoustic features of real speech as the target data, and the acoustic features of real speech are the acoustic features extracted from the real speech signals of the target audio text sample set.

[0120] Perform weighted mixing on the initial data and the target data through a weight factor to obtain intermediate data, and during the training process, adjust the weight factor so that the intermediate data gradually reaches the target data from the initial data.

[0121] Concatenate the intermediate data with the first set of acoustic feature samples, the acoustic feature context of the reference audio sample set, and the second set of acoustic feature samples as the input data for the detail encoder, and train the detail encoder so that the difference in acoustic features output by the detail encoder, combined with the second set of acoustic feature samples and the noise term, has the smallest difference from the acoustic features of real speech.

[0122] Based on the above further embodiment, the third processing module is specifically configured to construct an objective function using the acoustic features of real speech, the second set of acoustic feature samples, and the difference in acoustic features. The objective function reflects the difference between the difference in acoustic features output by the detail encoder, combined with the second set of acoustic feature samples and the noise term, and the acoustic features of real speech.

[0123] With the goal of minimizing the function value of the objective function, determine the parameters of the detail encoder through iterative updates.

[0124] Based on the above further embodiment, the objective function is expressed as:

[0125]

[0126] where Mel is the acoustic features of real speech, F 2 is the second set of acoustic feature samples, Θ(F″ t ,t) is the detail encoder, F″ t is the input value of the detail encoder, t is the weight factor in the iterative process, T is 1, is the expected value;

[0127] where F″ t = tMel+(1 - t)(F2 +(Noise)+F 1 +F 2 +Prompt

[0128] Noise is the noise term, F 1 is the first set of acoustic feature samples, and Prompt is the acoustic feature context of the reference audio sample set.

[0129] Based on the above further embodiments, the low-level prosodic information is low-level prosodic information, and the low-level prosodic information includes information on pitch, duration, and intensity in the text.

[0130] The high-level prosodic information is high-level prosodic features, and the high-level prosodic features include intonation, speech rhythm, syllable duration, and the offset of speech intensity affected by the speaker and emotional information.

[0131] The present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by one or more processors, it implements the zero-shot voice cloning method described in any one of the above technical solutions.

[0132] In addition, an electronic device is also provided, including a memory and one or more processors. A computer program is stored on the memory, and when the computer program is executed by the one or more processors, it implements the zero-shot voice cloning method described in any one of the above technical solutions.

[0133] The following introduces the data comparison in the experiment using the zero-shot voice cloning model SF-Speech of the present invention and the baseline models. The Magic Data dataset is used in the experiment, and the texts in the dataset cover various daily scenarios, including interactive Q&A, music search, spoken text messages, home instruction control, etc. The total audio in this corpus is 755 hours, recorded by 1000 speakers from different accent regions. The dataset is split into 712.1 hours for training, 14.8 hours for validation, and 28.1 hours for testing. Three baseline models are selected for comparison: Baseline Model 1 is YourTTS, which is based on VITS and a speaker recognition model H / ASP. YourTTS is the best-performing zero-shot voice cloning model based on GSE. Baseline Model 2 is VALL-E, which is an autoregressive voice cloning model based on LLM context learning and discrete acoustic representations. Baseline Model 3 is VoiceBox, which is the latest zero-shot voice cloning model based on ODE. VoiceBox-S is smaller in scale, with the same 8-layer Transformer structure as the detailed ODE module of SF-Speech, and has a total of 110M trainable parameters. VoiceBox follows its original configuration exactly, consisting of 24 layers of Transfomers and having a total of 330M trainable parameters.

[0134] In the experiment, the zero-shot voice cloning task 1 is to perform text-to-speech synthesis (Zero Shot Text to Speech, ZS-TTS), and task 2 is speech restoration (SR). In the ZS-TTS task, 256 voices of 32 unseen speakers (16 females and 16 males) are cloned for subjective evaluation, and 4000 voices are used for objective evaluation. In the SR task, the first 15% and the last 15% of the speech are retained, and the middle 70% of the speech is masked for reconstruction. Similarly, 256 voices are selected for subjective evaluation and 4000 voices are used for objective evaluation. For subjective metrics, the quality mean opinion score (QMOS) is used to evaluate the quality of the generated speech; the similarity mean opinion score (SMOS) is used to evaluate the timbre similarity between the generated speech and the reference speech of unseen speakers. For objective metrics, the word error rate (WER) of a speech recognition model is used to measure the intelligibility of the generated speech; the cosine similarity of speaker representations (SIM-o) extracted by a speaker recognition model is used to measure the timbre similarity between the generated speech and the reference speech; the reconstructed speaker representation cosine similarity (SMI-r) is used to measure the timbre similarity between the generated speech and the reference audio reconstructed by the vocoder to remove the influence brought by the vocoder difference. To ensure fairness, all ODE models use 8 steps of the solution steps.

[0135]

[0136] Table 1: Subjective and objective scoring results on two types of tasks

[0137] As shown in the comparison results of the SF-Speech model and the baseline model in Table 1 above, in the ZS-TTS task, the SF-Speech model achieved the best performance in all metrics, which means that SF-Speech has more advantages than all baseline models in the ZS-TTS task. In addition, in terms of speech quality and intelligibility, the performance of YourTTS and VALL-E is significantly worse than that of those ODE-based models. Even the worst-performing ODE-based model, VoiceBox-S, has a QMOS 0.43 / 0.48 higher than YourTTS / VALL-E; SMOS 0.09 / 0.46 higher; SIM-o 0.007 / 0.229 higher; and the WER score is 15.67 / % / 11.05 / % lower (a decrease of nearly 50 / % ). On the other hand, compared with the ODE-based method, the SF-Speech model has a similar parameter scale to VoiceBox-S, but reduces the WER by 3.37% compared to it, increases the QMOS by 0.19 compared to it, increases the SIM-o by 0.029, and increases the SMOS by 0.3. Even when the parameters of VoiceBox are increased to 330M (almost three times that of SF-Speech), SF-Speech still has advantages in QMOS / WER / SMOS / SIM-o (0.05 / 0.21% / 0.03 / 0.002). These results prove that using independent and coupled training data can significantly improve the accuracy of model modeling of ODEs trained with FM. In the SR task, since the two baseline models, YourTTS and VALL-E, are not applicable to this task, only the SF-Speech model and two scales of VoiceBox were compared. As can be seen from the bottom of Table 1, in terms of the quality evaluation of the generated speech, the SF-Speech model has obvious advantages, with a QMOS 0.28 / 0.04 higher than VoiceBox-S / VoiceBox, and a WER 2.07% / 2.01% lower than VoiceBox. It is worth noting that after tripling the number of parameters, VoiceBox did not show significant improvement in terms of WER. This indicates that VoiceBox may have limitations in the speech intelligibility of this short-term SR task, while SF-Speech broke this limitation by only increasing 7M parameters compared to VoiceBox-S. In addition, in terms of timbre similarity, the SF-Speech model scores slightly worse than VoiceBox, being 0.004 / 0.005 / 0.03 lower in SIM-o / SIM-r / SMOS. However, since the parameters of VoiceBox are 182% more than those of the SF-Speech model, the score of the SF-Speech model is still competitive.On the other hand, the SF-Speech model outperforms VoiceBox-S, improving by 0.011 / 0.009 / 0.26 on SIM-o / SIM-r / SMOS. These results on the SR task once again demonstrate the great advantage of the SF-Speech model in zero-shot voice cloning.

[0138] To further analyze the advantages of the proposed model over the state-of-the-art best-performing ODE model, the objective metric changes of the SF-Speech model and two VoiceBoxes of different scales were measured at different numbers of function evaluations (NFE) in ZS-TTS. As Figure 3 A schematic diagram of the performance of the ODE-based model at different NFEs, showing the fluctuating trend of the objective metric as the NFE of the ODE model increases. It can be seen that increasing the inference steps does not improve the performance of the ODE-based model. This is because as the NFE increases, the ODE-based model adds more harmful noises (reverberation, ambient sound, artificial noise, etc.) to the generated Mel spectrogram, and these noises are widely present in the dataset used for training. In addition, different objective metrics have different trends. As Figure 3 (a) shows, the SIM of these three models has the same trend and reaches the peak when NFE is equal to 8. The SF-Speech model outperforms VoiceBox-S in all NFE cases and outperforms VoiceBox when NFE is greater than 8. This result indicates that the SF-Speech model can alleviate the performance degradation caused by the increase in NFE. On the other hand, as Figure 3 (b) shows, the SF-Speech model can obtain the lowest WER under any NFE condition. Moreover, when NFE is greater than 8, the increase in the WER of the SF-Speech model is smaller than that of VoiceBox-S and VoiceBox, which once again proves the superiority of the SF-Speech model. Finally, since WER and SIM cannot fully represent the quality of the generated speech, the DNSMOS of the generated speech at different NFEs was also tested. DNSMOS is a model that uses a neural network to simulate the mean opinion score of human perception of speech quality. As Figure 3As shown in (c), the DNSMOS scores of the SF-Speech model under any NFE exceed those of VoiceBox and VoiceBox-S. Moreover, most importantly, VoiceBox and VoiceBox-S achieve the best DNSMOS when NFE is equal to 8, while the SF-Speech model already reaches the best generation quality when NFE is equal to 4. This means that it is beneficial to uniformly select the number of solution steps equal to 8 to compare all ODE-based models for VoiceBox and VoiceBox-S. Even so, the SF-Speech model still approaches or even surpasses the existing best models in multiple metrics, which once again proves the superiority of the proposed model in zero-resource voice cloning.

[0139] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only for the specific embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.

Claims

1. A zero-sample voice cloning method, characterized in that: The method comprises: Inputting a reference audio text and a target audio text into a text encoder to obtain a first acoustic feature, wherein the text encoder is used to predict low-level prosodic information related to the pronunciation content of the target audio text from the reference audio text, and add the low-level prosodic information to a phoneme string of the target audio text to obtain the first acoustic feature; Inputting the acoustic feature context of the reference audio and the first acoustic feature into a speaker encoder to obtain a second acoustic feature, wherein the speaker encoder is used to add high-level prosody information related to the speaker to the first acoustic feature according to the acoustic feature context of the reference audio to obtain the second acoustic feature; splicing the second acoustic feature with noise added, the first acoustic feature, the second acoustic feature and the acoustic feature context of the reference audio, and inputting the concatenated ... A predicted Mel spectrum is obtained according to the acoustic feature difference and the second acoustic feature, and the predicted Mel spectrum is converted into audio to obtain a target audio corresponding to the target audio text.

2. The method according to claim 1, characterized in that The method further comprises: Inputting a reference audio text sample set and a target audio text sample set into the text encoder to obtain a first acoustic feature sample set; Inputting the first acoustic feature sample set and the reference audio sample set into the speaker encoder to obtain a second acoustic feature sample set; A noise item is added to the second acoustic feature sample set as initial data, and then the first acoustic feature sample set, the acoustic feature context of the reference audio and the second acoustic feature sample set are concatenated as input data of the detail encoder to train the detail encoder, and the second acoustic feature sample set has a corresponding relationship with the acoustic features of the real speech.

3. The method according to claim 2, characterized in that The method further comprises: Using the real speech acoustic feature as target data, wherein the real speech acoustic feature is an acoustic feature extracted from a real speech signal of the target audio text sample set; Performing weighted mixing of the initial data and the target data by a weight factor to obtain intermediate data, and adjusting the weight factor during training so that the intermediate data gradually reaches the target data from the initial data; The intermediate data is concatenated with the first acoustic feature sample set, the acoustic feature context of the reference audio sample set, and the second acoustic feature sample set as input data of the detail encoder, and the detail encoder is trained so that the acoustic feature difference output by the detail encoder, combined with the second acoustic feature sample set and the noise term, has the minimum difference with the acoustic features of the real speech.

4. The method according to claim 3, characterized in that The step of splicing the intermediate data with the first acoustic feature sample set, the acoustic feature context of the reference audio sample set, and the second acoustic feature sample set as input data of the detail encoder, and training the detail encoder so that the acoustic feature difference amount output by the detail encoder combined with the second acoustic feature sample set and the noise term has the smallest difference with the acoustic feature of the real speech, specifically includes: Constructing an objective function using the real speech acoustic features, the second acoustic feature sample set and the acoustic feature difference, wherein the objective function reflects the difference between the acoustic feature difference output by the detail encoder combined with the second acoustic feature sample set and the noise term and the real speech acoustic features; Taking the minimum function value of the objective function as the goal, the parameters of the detail encoder are determined through update iteration.

5. The method according to claim 4, characterized in that The objective function is expressed as: Among them, Mel is the real speech acoustic feature, F2 is the second acoustic feature sample set, Θ(F″ t ,t) is the detail encoder, F″ t is the input value of the detail encoder, t is the weight factor of the iterative process, T is 1, is the expected value; Among them, F t =tMel+(1-t)(F2+Noise)+F1+F2+Prompt Noise is the noise term, F1 is the first acoustic feature sample set, and Prompt is the acoustic feature context of the reference audio sample set.

6. The method according to claim 1, characterized in that The low-level prosody information is low-level prosody information, which includes information on pitch, length and intensity of a tone in a text.

7. The method according to claim 1, characterized in that The high-level prosodic information is high-level prosodic features, which include intonation, speech rhythm, syllable duration, and offset of speech intensity that are affected by speaker and emotion information.

8. A zero-sample voice cloning device, characterized in that: The device comprises: A first processing module, configured to input a reference audio text and a target audio text into a text encoder to obtain a first acoustic feature, wherein the text encoder is configured to predict low-level prosodic information related to a pronunciation content of the target audio text from the reference audio text, and add the low-level prosodic information to a phoneme string of the target audio text to obtain the first acoustic feature; a second processing module, configured to input the reference audio and the first acoustic feature into a speaker encoder to obtain a second acoustic feature, wherein the speaker encoder is configured to add high-level prosody information related to the speaker to the first acoustic feature according to the acoustic feature context of the reference audio to obtain the second acoustic feature; A third processing module is used to splice the second acoustic feature with noise added, the first acoustic feature, the second acoustic feature and the acoustic feature context of the reference audio, and input the spliced ​​features into a detail encoder to obtain an acoustic feature difference; The fourth processing module is used to obtain a predicted Mel spectrum according to the acoustic feature difference and the second acoustic feature, convert the predicted Mel spectrum into audio, and obtain the target audio corresponding to the target audio text.

9. A computer storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the zero-sample voice cloning method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: The method comprises a memory and one or more processors, wherein a computer program is stored in the memory, and when the computer program is executed by the one or more processors, the zero-sample voice cloning method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Voice generation method and device based on artificial intelligence, computer equipment and medium

    CN119360818A

  • Voice generation method and device, equipment and medium

    CN119360819A

  • Zero sample voice cloning method based on dynamic neural network and feature modulation

    CN119360821A

  • Method and Device for Zero-Shot Speech Generation with Prosody Control and Random Speaker Generation

    US20240404509A1

  • Techniques for improved zero-shot voice conversion with a conditional disentangled sequential variational auto-encoder

    WO2023229626A1