Speech synthesis based corpus augmentation method, system, device and medium
CN121260141BActive Publication Date: 2026-08-28深圳市友杰智新科技有限公司
Patent Information
- Application Number
- CN202511070041.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Technical Problem
然而,这种方法存在显著的局限性
Benefits of technology
[0015]本申请通过相比传统依赖人工录音和标注的方式,显著提升语料生产的效率和质量。基于预训练模型的微调技术确保了个性化语音合成的质量,而多线程并行计算则大幅提升了处理能力,使得系统可以同时处理多个说话人的语音数据。这种端到端的解决方案不仅降低了语料获取的门槛和成本,还保证了生成语料的多样性和可用性。更重要的是,该方法建立的个性化声学模型可以持续复用,为语音技术的研发和应用提供了稳定可靠的数据支持,通过少量语音样本快速生成个性化语音语料,显著降低数据采集成本,提升语料库构建效率。
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN121260141B_ABST
Abstract
The application relates to the technical field of voice synthesis, in particular to a corpus expansion method and system based on voice synthesis, equipment and a medium, wherein the method comprises the following steps: based on collected audio and corresponding text of a target speaker, pre-processing including data labeling is carried out; acoustic characteristics are extracted from the pre-processed audio; based on a pre-trained acoustic model, individual fine-tuning is carried out by using labeled data and acoustic characteristics, and multiple speaker individual acoustic models are simultaneously trained through multi-thread parallel computing; the trained individual acoustic model is called, and a voice corpus of a target text is synthesized in combination with a vocoder; and based on the trained individual acoustic model, a corpus library is continuously expanded by replacing texts. The application can quickly generate individual voice corpus through a small amount of voice samples, significantly reduces data collection cost, and improves corpus construction efficiency.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Speech synthesis method and device, computer equipment and storage medium
CN117894293A
Speech synthesis model training method and device, speech synthesis method and device and storage medium
CN119516996A