Voice cloning system based on few samples
By optimizing the collaborative architecture of encoder, synthesizer and vocoder, combining multi-stage loss function and adaptive control technology, the problem of insufficient generalization capabilities of traditional voice cloning systems in data scarce scenarios is solved, and efficient and high-fidelity voice cloning effect is achieved.
Patent Information
- Application Number
- CN202510429398.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional voice cloning systems require a large amount of audio data of target speakers for training, which is difficult to apply in scenarios with scarce data. In addition, the generalization ability is insufficient in scenarios with multi-speaker scenarios, and there are limitations in tone fitting, rhythm control and real-time.
The collaborative architecture of encoder, synthesizer and vocoder is adopted, combined with GE2E, Softmax and Contrast loss functions, tone and pronunciation are adaptively controlled through deep learning technology, and the orthogonal mirror filter group and LPCNet technology are used to achieve high-fidelity voice cloning.
It realizes efficient and high-fidelity voice cloning under a small number of samples, which can accurately preserve the speaker's emotions, tone and timbre characteristics, and maintain high feature extraction ability and synthesis efficiency in complex environments.
Smart Images

Figure CN120452410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech synthesis technology, and in particular to a speech cloning system based on a small number of samples. Background Art
[0002] Traditional voice cloning systems typically require a large amount of target speaker audio data for training, making them difficult to apply in data-scarce scenarios. Existing deep learning-based speech synthesis systems (such as Tacotron and WaveNet) can generate high-quality speech, but they require high data volumes and computing resources, and their generalization capabilities are insufficient in multi-speaker scenarios. Furthermore, existing methods still have limitations in timbre matching, prosody control, and real-time performance. Summary of the Invention
[0003] The purpose of the present invention is to provide a voice cloning system based on a small number of samples to solve the problems raised in the above background technology.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a speech cloning system based on a small number of samples, comprising an encoder, a synthesizer, and a vocoder. The encoder comprises an audio encoder and a content encoder. The audio encoder is used to extract the unique features of the speaker from the original audio. The content encoder extracts content features from a pre-trained ASR model as input. The synthesizer is responsible for fusing the input text with the speech features extracted by the encoder to generate a mel-spectrogram.
[0005] Furthermore, the GE2E loss function is used for optimization.
[0006] Furthermore, the GE2E loss function is optimized by calculating the cosine similarity between speech features and speaker centroid to evaluate the fit of the timbre, initializing the model with the Softmax loss function, and fine-tuning the model using the Contrast loss function.
[0007] Furthermore, prosodic embeddings are generated by separating the fundamental frequency, explicit variables of pronunciation decision, and latent variables such as emotion.
[0008] Furthermore, deep learning technology is used to adaptively control tone and pronunciation.
[0009] Furthermore, the vocoder converts the mel-spectrogram generated by the synthesizer into high-fidelity waveform audio.
[0010] Furthermore, the orthogonal mirror filter bank technology is used to divide the signal frequency band into multiple sub-bands, combined with the LPCNet parallel processing technology.
[0011] Furthermore, an attention mechanism is used to weight text features during the decoding process, so that the model can focus on the information in the text that is most relevant to the current part of the generated speech.
[0012] Furthermore, the sub-band signals are combined by an orthogonal mirror filter combiner to output the final synthesized speech.
[0013] The present invention provides a speech cloning system based on a small number of samples, which has the following beneficial effects: the present invention achieves efficient and high-fidelity speech cloning by optimizing the collaborative architecture of the encoder, synthesizer, and vocoder, combining multi-stage loss functions and adaptive control technology; it uses deep learning technology to achieve high-quality speech cloning, and can accurately preserve the speaker's emotion, intonation, and timbre characteristics in cross-language scenarios, while maintaining high feature extraction capabilities and synthesis efficiency in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is an overall framework diagram of a voice cloning system based on a small number of samples according to the present invention; Figure 2 This is a mel-spectrogram of a speech cloning system based on a small number of samples in the present invention. DETAILED DESCRIPTION
[0015] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0016] like Figure 1-Figure 2 As shown in the figure, a voice cloning system based on a small number of samples includes the following core modules: Encoder: Consists of an audio encoder and a content encoder. The audio encoder extracts the speaker's timbre features from the original audio using a convolutional neural network (CNN) and a long short-term memory network (LSTM). The content encoder uses a pre-trained ASR model (such as Wav2Vec2) to extract text content features to ensure pronunciation accuracy.
[0017] Synthesizer: Fusion of text features with timbre features based on an attention mechanism (such as Transformer or GMM attention) to generate a mel-spectrogram. By explicitly separating the fundamental frequency, pronunciation decision, and emotional latent variables (such as VAE structure), it generates a prosodic embedding to enhance naturalness.
[0018] Vocoder: Uses quadrature mirror filter bank (QMF) technology to convert spectrograms into waveform audio, combined with LPCNet parallel processing to improve real-time performance.
[0019] GE2E loss function: Calculates the cosine similarity between speech features and speaker centroids, initializes the model with Softmax loss, and then fine-tunes with Contrast loss to improve timbre similarity.
[0020] Prosody control: Separate fundamental frequency and emotional variables through adversarial training, and use deep learning models (such as PitchNet) to adaptively adjust pitch and pronunciation rhythm.
[0021] Attention mechanism: Multi-head attention is used in the synthesizer to dynamically weight the input text features to ensure that the decoding focuses on the relevant content of the current speech segment.
[0022] Vocoder optimization: After the QMF technology divides the signal into subbands, LPCNet is used to synthesize the subband signals in parallel, and finally merge them into a high-fidelity waveform through a filter combiner.
[0023] Take 5 target speaker audio samples as examples: The audio encoder extracts timbre features, and the content encoder extracts text phoneme sequences; The synthesizer combines the two to generate a mel-spectrogram, and the rhythm embedding module injects the fundamental frequency and emotion parameters; The vocoder generates the final speech through QMF-LPCNet, with a cloning similarity of over 90% (MOS evaluation).
[0024] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.
Claims
1. A voice cloning system based on a small number of samples, characterized by: The system includes an encoder, a synthesizer and a vocoder. The encoder includes an audio encoder and a content encoder. The audio encoder is used to extract the unique features of the speaker from the original audio. The content encoder extracts content features from a pre-trained ASR model as input. The synthesizer is responsible for fusing the input text with the speech features extracted by the encoder to generate a mel-spectrogram.
2. A voice cloning system based on a small number of samples according to claim 1, characterized in that: The GE2E loss function is used for optimization.
3. The voice cloning system based on a small number of samples according to claim 2, characterized in that: The GE2E loss function is optimized by calculating the cosine similarity between speech features and speaker centroid to evaluate the fit of the timbre, initializing the model with the Softmax loss function, and fine-tuning the model using the Contrast loss function.
4. The voice cloning system based on a small number of samples according to claim 3, characterized in that: Prosodic embeddings are generated by separating the explicit variables of fundamental frequency, pronunciation decision, and latent variables such as emotion.
5. The voice cloning system based on a small number of samples according to claim 4, characterized in that: Deep learning technology is also used to adaptively control pitch and pronunciation.
6. The voice cloning system based on a small number of samples according to claim 5, characterized in that: The vocoder converts the mel-spectrogram generated by the synthesizer into high-fidelity waveform audio.
7. The voice cloning system based on a small number of samples according to claim 6, characterized in that: The orthogonal mirror filter bank technology is used to divide the signal frequency band into multiple sub-bands, combined with LPCNet parallel processing technology.
8. The voice cloning system based on a small number of samples according to claim 7, characterized in that: The attention mechanism is used to weight text features during the decoding process, so that the model can focus on the information in the text that is most relevant to the current part of the generated speech.
9. The voice cloning system based on a small number of samples according to claim 8, characterized in that: The sub-band signals are combined through the orthogonal mirror filter combiner to output the final synthesized speech.