一种声码器的构建方法、语音合成方法及相关装置

By constructing a combination of amplitude spectrum predictor and phase spectrum predictor, the contradiction between the generation efficiency and computational complexity of neural network vocoders is resolved, achieving efficient speech synthesis, improving speech generation efficiency and reducing computational complexity.

CN116524894BActive Publication Date: 2026-07-17UNIV OF SCI & TECH OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2023-01-16
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing speech synthesis technologies, neural network vocoders present a trade-off between generation efficiency and computational complexity, making it difficult to simultaneously improve speech generation efficiency and reduce overall computational complexity.

Method used

A combination of amplitude spectrum predictor and phase spectrum predictor is adopted to predict the speech amplitude spectrum and phase spectrum directly in parallel at the full frame level. The vocoder is constructed by training with amplitude spectrum loss, phase spectrum loss, short-time spectrum loss and waveform loss.

Benefits of technology

It significantly improves speech generation efficiency, reduces overall computational complexity, and enhances the quality and efficiency of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524894B_ABST
    Figure CN116524894B_ABST
Patent Text Reader

Abstract

本申请实施例公开了一种声码器的构建方法、语音合成方法及相关装置,先获取目标声学特征,并将目标声学特征分别输入到幅度谱预测模型和相位谱预测模型中得到第一对数幅度谱和第一相位谱,第一对数幅度谱包括第一幅度谱。接着根据第一幅度谱和第一相位谱进行计算得到第一重构短时谱,并对第一重构短时谱预处理得到第一重构语音波形。计算幅度谱损失、相位谱损失、短时谱损失、波形损失,并根据以上损失计算修正参数。再根据修正参数修正幅度谱预测模型和相位谱预测模型得到幅度谱预测器和相位谱预测器。本申请的幅度谱预测器和相位谱预测器可以实现平行直接预测幅度谱和相位谱,提高了语音生成的效率,降低了整体运算的复杂度。
Need to check novelty before this filing date? Find Prior Art