The invention provides a voice synthesis method and
system based on VITS improvement, and the method comprises the steps: optimizing a text
encoder of a VITS model, introducing a large
language model, and enabling the emotion, intention and
speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic
backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a
time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional
transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of
speech synthesis can be effectively improved, a new solution is provided for cross-language diversified
speech synthesis, and a reference is provided for the more efficient and more intelligent development of the
speech synthesis technology.