The application relates to the technical field of text-to-
speech synthesis, in particular to a multi-layer
prosody sentence segmentation and tone intonation decoupling control method for Chinese TTS, which comprises the following steps: extracting features related to
prosody modeling from a Chinese text to be synthesized to establish a text feature sequence; performing boundary position prediction and correcting the boundary prediction sequence for level consistency; predicting the pause duration of each layer prediction sequence; generating a tone baseline corresponding to the text, calculating a tone residual component corresponding to the
sentence layer tone change; generating a residual tone control mark at the end of each tone
phrase boundary; and generating a
prosody planning sequence. The application cooperatively models the PW, PPH and IPH three-layer prosody boundaries and their pause durations, combines level consistency correction, guarantees the nested consistency of prosody boundaries of different levels, decouples the tone information and
sentence tone information in Putonghua, and realizes stable and fine-grained end-of-sentence tone control in the Chinese TTS
scenario.