合成具有一种或多种速度的语音的方法和计算机系统
By using deep neural networks for unsupervised speech synthesis and generating Mel spectrogram features based on speaking rate, the problem of unnatural speech speed control in existing technologies is solved, resulting in more natural and robust speech synthesis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2021-02-18
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech synthesis methods lack an understanding of human speech speed control when controlling speech speed, resulting in unnatural generated speech and an inability to effectively control speech speed in end-to-end models.
Unsupervised speech synthesis is performed using deep neural networks. By receiving the speech rate as a conditional input, the encoder and decoder generate Mel spectrogram features to achieve speech synthesis at different speeds, thus avoiding dependence on parallel data.
The generated speech is more natural and robust, maintaining its naturalness and fluency at different speeds, thus solving the problem of unnatural speech speed control in existing technologies.
Smart Images

Figure CN115210808B_ABST