The invention relates to the technical field of
speech synthesis, and particularly discloses a
speech synthesis method and
system for controllable
latent variable modeling based on semantic
distillation, and the method comprises the steps: converting a Mel spectrum into continuous
latent variable distribution through a
speech coding module, generating continuous latent variables through re-parameterization sampling, introducing a self-supervised model for semantic
distillation, and carrying out the semantic
distillation. According to the method, alignment of latent variables and semantic features is constrained through marginal
cosine similarity and
distance matrix structure loss, a text
encoder maps a phoneme sequence into
latent variable distribution,
time sequence alignment of a text and the latent variables is achieved in combination with monotonic alignment search, and a decoder reconstructs the latent variables into a Mel spectrum. According to the method, waveform synthesis through a vocoder and total
loss function joint optimization reconstruction, KL
divergence, distillation,
text alignment and confrontation loss are carried out, discrete
information loss is avoided through continuous latent variable modeling,
semantic consistency and
text alignment efficiency are enhanced, the naturalness, coherence and real-time performance of synthesized voice are improved, and the method is suitable for scenes such as voice assistants and virtual anchors.