语音合成前端处理方法及语音合成方法
By using a front-end processing model trained on multiple tasks, combined with pre-trained semantic feature extraction and prediction units, the problem of low efficiency and poor accuracy in vowel recovery and prosodic boundary prediction in speech synthesis of target Asian and African languages such as Arabic is solved, achieving efficient and accurate speech synthesis results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD
- Filing Date
- 2022-12-19
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies suffer from low efficiency and poor accuracy in speech synthesis for target Asian and African languages such as Arabic, particularly in the absence of large-scale training text resources with vowel annotations, making it difficult to achieve efficient and accurate speech synthesis.
A multi-task trained front-end processing model is adopted, which combines pre-trained semantic feature extraction and prediction units. Semantic features are extracted through a character-level BERT model, and vowel symbols and prosodic boundaries are predicted using BiLSTM or a fully connected network, so as to achieve simultaneous vowel recovery and prosodic boundary prediction.
It improves the pronunciation accuracy and prosodic boundary representation of speech synthesis, enhances the intelligibility and naturalness of speech synthesis, and improves the efficiency and accuracy of front-end processing. In particular, it significantly improves the accuracy of vowel symbols and prosodic boundaries under limited resource conditions.
Smart Images

Figure CN116312475B_ABST