Speech synthesis model training method and device, equipment and medium
By extracting features from speech signals and using a pre-trained automatic speech recognition model, combined with a phoneme duration alignment strategy to generate a target phoneme sequence, the problem of poor phoneme duration alignment accuracy in existing technologies is solved, achieving higher-quality speech synthesis and lower training complexity.
Patent Information
- Application Number
- CN202511053340.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
AI Technical Summary
The existing technology has a strong reliance on external alignment tools, and has failed to solve or effectively solve specific problems.
By acquiring speech signals for feature extraction, speech recognition is performed using a pre-trained automatic speech recognition model. The target phoneme sequence and phoneme duration sequence are generated by combining the phoneme possibility matrix and the preset phoneme duration alignment strategy. Based on the style feature sequence and the target phoneme sequence, the target acoustic feature sequence is obtained through the speech synthesis model, and the model parameters are adjusted using the speech synthesis loss function.
The phoneme duration alignment accuracy of the speech synthesis model is improved, the speech synthesis quality is enhanced, the dependence on external alignment tools is reduced, and the training efficiency and generalization ability of the model are improved.
Smart Images

Figure CN120708596A_ABST