The present invention discloses a multi-style personalized Tibetan
speech synthesis model for low-resource conditions, including a speaker style
encoder and a sequentially connected
grapheme-to-phoneme conversion module, a text
encoder, a variance adapter, a mel spectrum decoder and a HifiGAN vocoder; the synthesis model is trained using a "pre-training + meta-learning" model
algorithm; the
model parameters that can be learned by the
deep learning model are converted to i Divided into parameters related to
rhythm i p , parameters related to the speaker i s and other remaining parameters #imgabs0#, i.e. #imgabs1# i p and i s The model includes learnable
model parameters for the speaker style
encoder, and #imgabs2# includes learnable
model parameters for the text encoder, mel-
spectrogram decoder, and variance adapter. This synthetic model can be applied to intelligent bilingual teaching products, as well as Tibetan spoken dialogue,
multimedia information processing, and other human-computer interaction fields.