This application relates to a semantically enhanced Amdo
Tibetan language generation method and
system. The method extracts semantic features containing lexical boundaries through a pre-trained
encoder with full-word masking. After dimensional transformation and phoneme-level repetition expansion, these features are added element-wise to the phoneme features, so that each phoneme directly carries the complete
semantics of its corresponding word, fundamentally eliminating semantic fragmentation in prosodic modeling. Simultaneously, the sequence termination vector is independently input into the end-energy predictor and jointly trained with a variational
inference adversarial training framework to obtain the generation model and the
sentence-end energy decay coefficient. During
inference, after synthesizing audio segments according to
sentence-end
punctuation, the end of the segments is faded out based on the decay coefficient and a
silence buffer block is spliced. This ensures the naturalness and accuracy of the speech flow
rhythm and logical stress while completely eliminating
sentence-end elision and acoustic truncation
noise, significantly improving the prosodic naturalness and speech purity of the synthesized Amdo Tibetan speech.