This invention provides a method and apparatus for Manchu
speech synthesis based on character
component modeling, belonging to the field of
speech synthesis technology. The method includes: converting Manchu text into Latin characters according to transcription rules, generating a character embedding sequence, and generating feature templates containing dimensions such as
fundamental frequency and
pitch band based on the acoustic characteristics of Manchu pronunciation; concatenating the two templates into a high-dimensional input
tensor, which is then fed into a pre-trained duration prediction network; constructing a
mask, extracting prosodic features from the Manchu Latin-transcribed character embedding sequence to generate a prosodic vector, and adjusting the attention weights between characters and speech frames in the pre-trained feature network to obtain the prosodic space, Mel-spectral sequence, and
optimal alignment path; finally, inputting the vector into a lightweight multi-band inverse short-time
Fourier transform decoder for Manchu to obtain a complete synthesized speech waveform. This invention can significantly and effectively solve the core problem of text-speech frame matching deviation in low-resource Manchu scenarios, and its
inference speed is significantly improved compared to general models.