Voice cloning system based on few samples

By optimizing the collaborative architecture of encoder, synthesizer and vocoder, combining multi-stage loss function and adaptive control technology, the problem of insufficient generalization capabilities of traditional voice cloning systems in data scarce scenarios is solved, and efficient and high-fidelity voice cloning effect is achieved.

CN120452410AInactive Publication Date: 2025-08-08ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510429398.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional voice cloning systems require a large amount of audio data of target speakers for training, which is difficult to apply in scenarios with scarce data. In addition, the generalization ability is insufficient in scenarios with multi-speaker scenarios, and there are limitations in tone fitting, rhythm control and real-time.

Method used

The collaborative architecture of encoder, synthesizer and vocoder is adopted, combined with GE2E, Softmax and Contrast loss functions, tone and pronunciation are adaptively controlled through deep learning technology, and the orthogonal mirror filter group and LPCNet technology are used to achieve high-fidelity voice cloning.

Benefits of technology

It realizes efficient and high-fidelity voice cloning under a small number of samples, which can accurately preserve the speaker's emotions, tone and timbre characteristics, and maintain high feature extraction ability and synthesis efficiency in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452410A_ABST
    Figure CN120452410A_ABST
Patent Text Reader

Abstract

The invention discloses a voice cloning system based on a small number of samples, and relates to the technical field of voice synthesis, the voice cloning system comprises an encoder, a synthesizer and a vocoder, the encoder comprises an audio encoder and a content encoder, the audio encoder is used for extracting unique features of a speaker from original audio, and the content encoder is used for extracting unique features of the speaker from the original audio; a content encoder extracts content features from a pre-trained ASR model as input, a synthesizer is responsible for fusing input texts and voice features extracted by the encoder to generate a Mel spectrogram, a GE2E loss function is adopted for optimization, and the Mel spectrogram is obtained. The GE2E loss function optimization is to evaluate the tone fitting degree by calculating the cosine similarity between the voice features and the centroid of a speaker, and to initialize the model in combination with the Softmax loss function. According to the speech cloning system based on a small number of samples, efficient and high-fidelity speech cloning is realized by optimizing a collaborative architecture of an encoder, a synthesizer and a vocoder and combining a multi-stage loss function and a self-adaptive control technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis technology, and in particular to a speech cloning system based on a small number of samples. Background Art

[0002] Traditional voice cloning systems typically require a large amount of target speaker audio data for training, making them difficult to apply in data-scarce scenarios. Existing deep learning-based speech synthesis systems (such as Tacotron and WaveNet) can generate high-quality speech, but they require high data volumes and computing resources, and their generalization capabilities are insufficient in multi-speaker scenarios. Furthermore, existing methods still have limitations in timbre matching, prosody control, and real-time performance. Summary of the Invention

[0003] The purpose of the present invention is to provide a voice cloning system based on a small number of samples to solve the problems raised in the above background technology.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a speech cloning system based on a small number of samples, comprising an encoder, a synthesizer, and a vocoder. The encoder comprises an audio encoder and a content encoder. The audio encoder is used to extract the unique features of the speaker from the original audio. The content encoder extracts content features from a pre-trained ASR model as input. The synthesizer is responsible for fusing the input text with the speech features extracted by the encoder to generate a mel-spectrogram.

[0005] Furthermore, the GE2E loss function is used for optimization.

[0006] Furthermore, the GE2E loss function is optimized by calculating the cosine similarity between speech features and speaker centroid to evaluate the fit of the timbre, initializing the model with the Softmax loss function, and fine-tuning the model using the Contrast loss function.

[0007] Furthermore, prosodic embeddings are generated by separating the fundamental frequency, explicit variables of pronunciation decision, and latent variables such as emotion.

[0008] Furthermore, deep learning technology is used to adaptively control tone and pronunciation.

[0009] Furthermore, the vocoder converts the mel-spectrogram generated by the synthesizer into high-fidelity waveform audio.

[0010] Furthermore, the orthogonal mirror filter bank technology is used to divide the signal frequency band into multiple sub-bands, combined with the LPCNet parallel processing technology.

[0011] Furthermore, an attention mechanism is used to weight text features during the decoding process, so that the model can focus on the information in the text that is most relevant to the current part of the generated speech.

[0012] Furthermore, the sub-band signals are combined by an orthogonal mirror filter combiner to output the final synthesized speech.

[0013] The present invention provides a speech cloning system based on a small number of samples, which has the following beneficial effects: the present invention achieves efficient and high-fidelity speech cloning by optimizing the collaborative architecture of the encoder, synthesizer, and vocoder, combining multi-stage loss functions and adaptive control technology; it uses deep learning technology to achieve high-quality speech cloning, and can accurately preserve the speaker's emotion, intonation, and timbre characteristics in cross-language scenarios, while maintaining high feature extraction capabilities and synthesis efficiency in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is an overall framework diagram of a voice cloning system based on a small number of samples according to the present invention; Figure 2 This is a mel-spectrogram of a speech cloning system based on a small number of samples in the present invention. DETAILED DESCRIPTION

[0015] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0016] like Figure 1-Figure 2 As shown in the figure, a voice cloning system based on a small number of samples includes the following core modules: Encoder: Consists of an audio encoder and a content encoder. The audio encoder extracts the speaker's timbre features from the original audio using a convolutional neural network (CNN) and a long short-term memory network (LSTM). The content encoder uses a pre-trained ASR model (such as Wav2Vec2) to extract text content features to ensure pronunciation accuracy.

[0017] Synthesizer: Fusion of text features with timbre features based on an attention mechanism (such as Transformer or GMM attention) to generate a mel-spectrogram. By explicitly separating the fundamental frequency, pronunciation decision, and emotional latent variables (such as VAE structure), it generates a prosodic embedding to enhance naturalness.

[0018] Vocoder: Uses quadrature mirror filter bank (QMF) technology to convert spectrograms into waveform audio, combined with LPCNet parallel processing to improve real-time performance.

[0019] GE2E loss function: Calculates the cosine similarity between speech features and speaker centroids, initializes the model with Softmax loss, and then fine-tunes with Contrast loss to improve timbre similarity.

[0020] Prosody control: Separate fundamental frequency and emotional variables through adversarial training, and use deep learning models (such as PitchNet) to adaptively adjust pitch and pronunciation rhythm.

[0021] Attention mechanism: Multi-head attention is used in the synthesizer to dynamically weight the input text features to ensure that the decoding focuses on the relevant content of the current speech segment.

[0022] Vocoder optimization: After the QMF technology divides the signal into subbands, LPCNet is used to synthesize the subband signals in parallel, and finally merge them into a high-fidelity waveform through a filter combiner.

[0023] Take 5 target speaker audio samples as examples: The audio encoder extracts timbre features, and the content encoder extracts text phoneme sequences; The synthesizer combines the two to generate a mel-spectrogram, and the rhythm embedding module injects the fundamental frequency and emotion parameters; The vocoder generates the final speech through QMF-LPCNet, with a cloning similarity of over 90% (MOS evaluation).

[0024] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.

Claims

1. A voice cloning system based on a small number of samples, characterized by: The system includes an encoder, a synthesizer and a vocoder. The encoder includes an audio encoder and a content encoder. The audio encoder is used to extract the unique features of the speaker from the original audio. The content encoder extracts content features from a pre-trained ASR model as input. The synthesizer is responsible for fusing the input text with the speech features extracted by the encoder to generate a mel-spectrogram.

2. A voice cloning system based on a small number of samples according to claim 1, characterized in that: The GE2E loss function is used for optimization.

3. The voice cloning system based on a small number of samples according to claim 2, characterized in that: The GE2E loss function is optimized by calculating the cosine similarity between speech features and speaker centroid to evaluate the fit of the timbre, initializing the model with the Softmax loss function, and fine-tuning the model using the Contrast loss function.

4. The voice cloning system based on a small number of samples according to claim 3, characterized in that: Prosodic embeddings are generated by separating the explicit variables of fundamental frequency, pronunciation decision, and latent variables such as emotion.

5. The voice cloning system based on a small number of samples according to claim 4, characterized in that: Deep learning technology is also used to adaptively control pitch and pronunciation.

6. The voice cloning system based on a small number of samples according to claim 5, characterized in that: The vocoder converts the mel-spectrogram generated by the synthesizer into high-fidelity waveform audio.

7. The voice cloning system based on a small number of samples according to claim 6, characterized in that: The orthogonal mirror filter bank technology is used to divide the signal frequency band into multiple sub-bands, combined with LPCNet parallel processing technology.

8. The voice cloning system based on a small number of samples according to claim 7, characterized in that: The attention mechanism is used to weight text features during the decoding process, so that the model can focus on the information in the text that is most relevant to the current part of the generated speech.

9. The voice cloning system based on a small number of samples according to claim 8, characterized in that: The sub-band signals are combined through the orthogonal mirror filter combiner to output the final synthesized speech.