一种语音合成方法、装置、电子设备及存储介质
By filtering similar timbre features of the target speaker in the pre-trained speech synthesis model and performing fine-tuning, the problem of time-consuming and labor-intensive personalized speech synthesis in the existing technology is solved, realizing efficient personalized speech synthesis, which is suitable for single and multi-person scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DINGFU NEW POWER (BEIJING) INTELLIGENT TECH CO LTD
- Filing Date
- 2022-11-29
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech synthesis algorithms require remodeling when faced with new personalized speech, which is time-consuming and labor-intensive, reducing the efficiency of speech synthesis tasks.
By acquiring training data of the target speaker, extracting its timbre features, and selecting similar speaker timbre features in the pre-trained speech synthesis model, the pre-trained model is fine-tuned using the fine-tune method to form a fine-tune speech synthesis model, thereby achieving personalized speech synthesis.
It improves the training efficiency and overall efficiency of personalized speech synthesis models, supports single-person and multi-person speech synthesis, is highly adaptable, and can quickly adapt to more scenarios.
Smart Images

Figure CN116312467B_ABST