The invention belongs to the technical field of text-to-speech conversion, and particularly relates to a solution for TTS (text-to-speech) in a high-
concurrency scene through tone
cloning. According to the method, the text can be generated and returned to the
client by segmenting and parallelizing the text and reducing the
delay of the first segment, compared with a traditional serial whole segment synthesis mode, the
waiting time of a user is greatly shortened, the audio is returned after being generated, the
waiting time of the user can be effectively shortened under high
concurrency, the network bandwidth and the
delay pressure are reduced, and the user experience is improved. The real-time performance is guaranteed, reconnection disasters are avoided, underlying model services can be fully utilized, the influence of the serial
processing characteristic of a CosyVoice2 model is avoided, a
client can immediately start to play a first segment of audio, even if subsequent segments are still synthesized, the continuous playing experience of previous content is not influenced, and when the
concurrency is increased, the continuous playing experience of the previous content is not influenced. The model gateway can dynamically disperse requests to a plurality of GPU instances, and long
queue waiting after a single-card
video memory is fully occupied is avoided.