The invention provides a full-
language speech synthesis method and device based on character
component modeling, and relates to the technical field of
speech synthesis. The method comprises the following steps: converting Manchu into Latin characters according to a transcription rule, generating a character embedding sequence, generating a feature template containing dimensions such as a
fundamental frequency and a
pitch band based on full-language pronunciation acoustic characteristics, splicing the feature template and the feature template into a high-dimensional input
tensor, and sending the high-dimensional input
tensor into a pre-
training time length prediction network; constructing a
mask, extracting
rhythm-related features from a character embedding sequence of Manchu Latin transcription to generate a
rhythm vector, adjusting attention weights of characters and voice frames in a pre-training feature network to obtain a
rhythm space, a Mel spectrum sequence and an
optimal alignment path, and finally inputting the rhythm space, the Mel spectrum sequence and the
optimal alignment path into a full-language lightweight multi-band inverse short-time
Fourier transform decoder. And a complete synthetic voice waveform is obtained. According to the method, the core problem of text and voice frame matching deviation in a full-language low-resource scene can be obviously and effectively solved, and the reasoning speed is obviously improved compared with a general model.