语音合成前端处理方法及语音合成方法

By using a front-end processing model trained on multiple tasks, combined with pre-trained semantic feature extraction and prediction units, the problem of low efficiency and poor accuracy in vowel recovery and prosodic boundary prediction in speech synthesis of target Asian and African languages ​​such as Arabic is solved, achieving efficient and accurate speech synthesis results.

CN116312475BActive Publication Date: 2026-07-17XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD
Filing Date
2022-12-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency and poor accuracy in speech synthesis for target Asian and African languages ​​such as Arabic, particularly in the absence of large-scale training text resources with vowel annotations, making it difficult to achieve efficient and accurate speech synthesis.

Method used

A multi-task trained front-end processing model is adopted, which combines pre-trained semantic feature extraction and prediction units. Semantic features are extracted through a character-level BERT model, and vowel symbols and prosodic boundaries are predicted using BiLSTM or a fully connected network, so as to achieve simultaneous vowel recovery and prosodic boundary prediction.

Benefits of technology

It improves the pronunciation accuracy and prosodic boundary representation of speech synthesis, enhances the intelligibility and naturalness of speech synthesis, and improves the efficiency and accuracy of front-end processing. In particular, it significantly improves the accuracy of vowel symbols and prosodic boundaries under limited resource conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312475B_ABST
    Figure CN116312475B_ABST
Patent Text Reader

Abstract

本发明涉及自然语言处理技术领域,提供一种语音合成前端处理方法及语音合成方法,该语音合成前端处理方法获取目标亚非语系语种的待处理文本;将待处理文本输入至前端处理模型,得到待处理文本中各字符对应的目标元音标符以及待处理文本中的目标韵律边界。该方法引入待处理文本中各字符对应的目标元音标符的恢复,可以提升后续合成的语音的读音准确性,引入对待处理文本中目标韵律边界的预测,可以提升后续合成的语音的韵律边界表现,进而提高语音合成的可懂度和自然度。该方法采用多任务训练的方式得到前端处理模型,使该前端处理模型可以同时得到待处理文本中各字符对应的目标元音标符以及待处理文本中的目标韵律边界,可以提高前端处理效率。
Need to check novelty before this filing date? Find Prior Art