Speech synthesis model training method and device, equipment and medium

By extracting features from speech signals and using a pre-trained automatic speech recognition model, combined with a phoneme duration alignment strategy to generate a target phoneme sequence, the problem of poor phoneme duration alignment accuracy in existing technologies is solved, achieving higher-quality speech synthesis and lower training complexity.

CN120708596APending Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511053340.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing technology has a strong reliance on external alignment tools, and has failed to solve or effectively solve specific problems.

Method used

By acquiring speech signals for feature extraction, speech recognition is performed using a pre-trained automatic speech recognition model. The target phoneme sequence and phoneme duration sequence are generated by combining the phoneme possibility matrix and the preset phoneme duration alignment strategy. Based on the style feature sequence and the target phoneme sequence, the target acoustic feature sequence is obtained through the speech synthesis model, and the model parameters are adjusted using the speech synthesis loss function.

Benefits of technology

The phoneme duration alignment accuracy of the speech synthesis model is improved, the speech synthesis quality is enhanced, the dependence on external alignment tools is reduced, and the training efficiency and generalization ability of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708596A_ABST
    Figure CN120708596A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, can be applied to the fields of finance and medical treatment, and discloses a speech synthesis model training method, device, equipment and medium, and the method comprises the steps: carrying out the speech recognition of a frame-level acoustic feature sequence through a pre-trained automatic speech recognition model, and obtaining a phoneme possibility matrix; obtaining a corresponding target phoneme sequence and a phoneme duration sequence according to a preset phoneme duration alignment strategy and the phoneme possibility matrix; based on the style feature sequence, the target phoneme sequence and the phoneme duration sequence, obtaining a target acoustic feature sequence through a preset speech synthesis model; and adjusting parameters of the speech synthesis model according to speech synthesis loss obtained by the frame-level acoustic feature sequence, the target acoustic feature sequence and a preset speech synthesis loss function. The speech synthesis model is trained based on the accurate phoneme duration, the phoneme duration alignment precision of the model is improved, and the speech synthesis quality is improved.
Need to check novelty before this filing date? Find Prior Art