一种语音合成方法、装置、电子设备及存储介质

By filtering similar timbre features of the target speaker in the pre-trained speech synthesis model and performing fine-tuning, the problem of time-consuming and labor-intensive personalized speech synthesis in the existing technology is solved, realizing efficient personalized speech synthesis, which is suitable for single and multi-person scenarios.

CN116312467BActive Publication Date: 2026-07-17DINGFU NEW POWER (BEIJING) INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DINGFU NEW POWER (BEIJING) INTELLIGENT TECH CO LTD
Filing Date
2022-11-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech synthesis algorithms require remodeling when faced with new personalized speech, which is time-consuming and labor-intensive, reducing the efficiency of speech synthesis tasks.

Method used

By acquiring training data of the target speaker, extracting its timbre features, and selecting similar speaker timbre features in the pre-trained speech synthesis model, the pre-trained model is fine-tuned using the fine-tune method to form a fine-tune speech synthesis model, thereby achieving personalized speech synthesis.

Benefits of technology

It improves the training efficiency and overall efficiency of personalized speech synthesis models, supports single-person and multi-person speech synthesis, is highly adaptable, and can quickly adapt to more scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312467B_ABST
    Figure CN116312467B_ABST
Patent Text Reader

Abstract

本申请提供一种语音合成方法、装置、电子设备及存储介质,其中语音合成方法包括:获取目标说话人训练数据,并提取目标说话人音色特征;在预训练数据中筛选与目标说话人的相似说话人音色特征;将训练好的预训练语音合成模型的模型参数加载至finetune语音合成模型;采用相似说话人音色特征与目标说话人音色特征共同训练finetune语音合成模型;将待合成文本输入训练好的finetune语音合成模型进行语音合成任务。通过预选构建的预训练模型,通过finetune的方式对预训练模型进行微调,以满足语音合成任务的及时性需求,极大提升了个性化语音合成模型的训练效率,进而提升了个性化语音合成任务的整体效率。
Need to check novelty before this filing date? Find Prior Art