合成具有一种或多种速度的语音的方法和计算机系统

By using deep neural networks for unsupervised speech synthesis and generating Mel spectrogram features based on speaking rate, the problem of unnatural speech speed control in existing technologies is solved, resulting in more natural and robust speech synthesis.

CN115210808BActive Publication Date: 2026-07-17TENCENT AMERICA LLC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2021-02-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech synthesis methods lack an understanding of human speech speed control when controlling speech speed, resulting in unnatural generated speech and an inability to effectively control speech speed in end-to-end models.

Method used

Unsupervised speech synthesis is performed using deep neural networks. By receiving the speech rate as a conditional input, the encoder and decoder generate Mel spectrogram features to achieve speech synthesis at different speeds, thus avoiding dependence on parallel data.

Benefits of technology

The generated speech is more natural and robust, maintaining its naturalness and fluency at different speeds, thus solving the problem of unnatural speech speed control in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115210808B_ABST
    Figure CN115210808B_ABST
Patent Text Reader

Abstract

提供了一种合成具有一种或多种速度的语音的方法和计算机系统。对与一个或多个音位相关联的、与说话语音相对应的上下文进行编码。基于已编码的上下文将所述一个或多个音位与一个或多个目标声学帧对准。利用经过对准的音位和所述目标声学帧递归地生成一个或多个梅尔语谱图特征;及使用生成的梅尔语谱图特征合成与所述说话语音相对应的给定速度的语音样本。
Need to check novelty before this filing date? Find Prior Art