Solution for TTS (Tone To Send) in high-concurrency scene by clone timbre
Through text segmentation parallel processing and dynamic load balancing, the memory saturation and network pressure problems of tone cloning TTS in high concurrent scenarios are solved, and the low-latency audio playback experience and resource utilization are achieved.
Patent Information
- Application Number
- CN202510741062.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-11
AI Technical Summary
现有音色克隆TTS技术在高并发场景下容易导致显存饱和、推理任务排队、网络带宽压力增大、CPU负载增加及网络中断后重连风暴,影响用户体验。
Through parallel text segmentation processing, dynamic load balancing, streaming fragment writing and elastic scheduling, text segmentation is sent to tone cloning model instances in parallel, audio files are stored in real time and URL list is returned, and the client can download and play them in sequence.
在高并发场景下减少用户等待时间,降低网络带宽与延迟压力,保证实时性,避免重连灾难,充分利用后端资源,确保连续播放体验。
Smart Images

Figure CN120299447A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text-to-speech conversion, and specifically relates to a solution method for TTS with cloned voices in high-concurrency scenarios. Background Art
[0002] The core goal of the Text-To-Speech (TTS) technology is to convert text into natural and fluent speech output. In recent years, the development of multimodal large language models (LLMs) has significantly promoted the innovation of speech synthesis technology. Among them, voice cloning, as a key branch, is committed to reproducing the voice characteristics of specific speakers through a small number of samples.
[0003] Currently, voice cloning and TTS (text-to-speech) technologies have shifted from traditional models that rely on large-scale labeled data (such as HMM and DNN) to fine-tuning large language model solutions based on few-shot learning. For example, GPT-SoVITS can generate cloned audio with 80%-95% similarity with only 5 seconds of speech. Llasa TTS achieves fine-grained control of emotions and styles through the LLaMA 8B large model, while CosyVoice2 reduces latency through lightweight design (0.5B parameters) and streaming output technology.
[0004] In the prior art, CosyVoice, as a streaming speech synthesis model technology, already has many excellent features. However, although the existing voice cloning TTS technology has made important breakthroughs in voice similarity and streaming low latency, the model itself requires a large amount of GPU video memory (for example, CosyVoice2-0.5B requires about 6GB for a single inference). When the concurrency increases, it is easy to cause video memory saturation and inference task queuing, increasing the hardware cost. Streaming synthesis requires continuous pushing of audio data, resulting in a sudden increase in bandwidth pressure and network fluctuations easily causing playback stuttering. At the same time, in order to maintain low-latency output, the client usually needs to maintain a large number of long connections. Once the number of concurrent connections surges, it will bring serious memory and CPU loads and cause a "reconnection storm" after a network interruption. Summary of the Invention
[0005] The purpose of the present invention is to provide a solution method for TTS with cloned voices in high-concurrency scenarios, which can ensure low-latency first-packet startup at the user end through means such as segmented parallelism, dynamic load balancing, streaming shard writing, and elastic scheduling, and can maximize the utilization of backend computing power and network resources and smoothly handle high-concurrency requests.
[0006] The technical solutions adopted by the present invention are specifically as follows: A solution for TTS with cloned voices in high-concurrency scenarios, including: Receiving the text to be synthesized and splitting it into multiple text segments according to punctuation rules; Parallelly sending the multiple text segments to the model gateway, and the model gateway allocates them to the corresponding voice cloning model instances through a preset load balancing strategy; Controlling the voice cloning model instances to generate audio streams corresponding to the text segments in a streaming manner and storing them in the object storage system in real time to generate independent audio files; When the independent audio file corresponding to the first text segment is generated, immediately return an ordered list containing the predefined URLs of all audio files to the client, so that the client can download and play them sequentially, while the background continuously generates the remaining audio files.
[0007] In a preferred solution, the step of receiving the text to be synthesized and splitting it into multiple text segments according to punctuation rules includes: The TTS API system receives the text to be synthesized and detects the set of sentence boundary punctuation marks in the text to be synthesized, where the set of punctuation marks includes full stops, question marks, exclamation marks, and semicolons; Performing primary segmentation at each detected boundary punctuation mark position to generate initial text segments; If the length of the text segment generated by the primary segmentation is less than the preset character condition, then merge it backward to the adjacent segment until the preset character condition is met; Attaching a sequential identifier to each initial text segment to obtain multiple text segments.
[0008] In a preferred solution, the preset character condition is ten characters.
[0009] In a preferred solution, the step of parallelly sending the multiple text segments to the model gateway and the model gateway allocating them to the corresponding voice cloning model instances through a preset load balancing strategy includes: Creating independent request packets based on each text segment; Batch-pushing the multiple independent request packets to the load balancing proxy entry of the model gateway through an asynchronous thread pool; Based on the model gateway, obtaining real-time resource monitoring data and executing the preset load balancing strategy to allocate the multiple independent request packets to the corresponding voice cloning model instances.
[0010] In a preferred solution, the independent request packet includes the text segment content, sequential identifier, hash value of the voice source file, and audio parameter configuration.
[0011] In a preferred solution, the step of controlling the voice cloning model instances to generate audio streams corresponding to the text segments in a streaming manner and storing them in the object storage system in real time to generate independent audio files includes: Receive independent request packets, parse the text fragment content, sequence identifier, hash value of the timbre source file, and audio parameter configuration; Load the corresponding timbre cloning model instance based on the hash value of the timbre source file and configure the audio parameters; Based on the loaded timbre cloning model instance, generate an audio data stream corresponding to the content of the text segment in a streaming manner; Split the audio data stream into consecutive data blocks; Upload the segmented data blocks to the object storage system in real time; In the object storage system, a unique audio file identifier is generated according to the sequence identifier and the hash value of the timbre source file, and the received data blocks are sequentially written into the independent audio file corresponding to the audio file identifier.
[0012] In a preferred solution, when the independent audio file corresponding to the first text segment is generated, an ordered list containing predefined URLs of all audio files is immediately returned to the client so that the client can download and play them in sequence, while the background continues to generate the remaining audio files, including: Real-time monitoring of the generation status of independent audio files corresponding to multiple text segments; If it is detected that the independent audio file corresponding to the first text segment has been written and closed in the object storage system, a list of all predefined URLs sorted by sequential identifiers is immediately generated and returned to the client; After returning the ordered URL list to the client, the background continues to send multiple text segments in parallel to the model gateway to process the remaining text segments until all independent audio files corresponding to the text segments are generated in the object storage system; When a client request does not generate a file, an intermediate response containing the waiting time is returned.
[0013] In a preferred solution, if it is detected that the independent audio file corresponding to the first text segment has been written and closed in the object storage system, the step of immediately generating and returning a list of all predefined URLs sorted by sequential identifiers to the client includes: If it is detected that the independent audio file corresponding to the first text segment is written and closed in the object storage system, a predefined URL is generated in the object storage system for each text segment according to the sequential identifiers of the multiple text segments; Sort the generated URLs according to the values of the sequence identifiers from small to large to form an ordered URL list; Immediately returns an ordered list of URLs to the requesting client.
[0014] In a preferred embodiment, when the client request has not generated a file, the step of returning an intermediate response including the waiting time includes: Real-time monitor the client's download requests for subsequent URLs in the ordered list; If it is detected that the audio file requested by the client for download has not been completely generated in the object storage system, then increase the processing priority of the corresponding text segment in the model gateway, return the status code and the estimated waiting time to the client, and at the same time trigger the write completion monitoring of the corresponding text segment in the object storage system. When it is monitored that the corresponding text segment is completely written, immediately push the final URL that can be redirected to the complete file to the client.
[0015] And, a solution terminal for cloning voices for TTS in a high-concurrency scenario includes: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the solution method for cloning voices for TTS in a high-concurrency scenario.
[0016] The technical effects achieved by the present invention are: In the present invention, through text segmentation in parallel, the first segment often has a lower delay and can be generated and returned to the client. Compared with the traditional serial and other whole-segment synthesis methods, the waiting time of users is greatly shortened, and it can be returned after an audio is generated. It can effectively reduce the waiting time of users under high concurrency, reduce the network bandwidth and delay pressure, ensure real-time performance, and there will be no reconnection disasters. Moreover, it can make full use of the underlying model service and will not be affected by the serial processing characteristics of the CosyVoice2 model. The client can immediately start playing the first segment of the audio. Even if the subsequent paragraphs are still being synthesized, it will not affect the continuous playback experience of the previous content. When the concurrency increases, the model gateway can dynamically distribute requests to multiple GPU instances, avoiding long queue waiting after a single card's video memory is full. Each instance loads or unloads the model as needed, effectively reducing video memory fragmentation and ensuring video memory utilization. Description of the Drawings
[0017] Figure 1 is the flowchart of the method provided by the present invention. Detailed Embodiments
[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be made in conjunction with the accompanying drawings of the specification.
[0019] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those specifically described herein. Those skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0020] Secondly, as used herein, an "embodiment" or "embodiments" refer to specific features, structures, or characteristics that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments.
[0021] Thirdly, the present invention is described in detail with reference to the schematic diagrams. When describing the embodiments of the present invention in detail, for the sake of convenience of explanation, the schematic diagrams are only examples and should not limit the scope of protection of the present invention.
[0022] Please refer to the attached Figure 1 As shown, a solution for TTS with cloned voices in a high-concurrency scenario is provided, including: S1. Receive the text to be synthesized and split it into multiple text segments according to punctuation rules; S2. Send the multiple text segments in parallel to the model gateway, and the model gateway distributes them to the corresponding voice cloning model instances through a preset load balancing strategy; S3. Control the voice cloning model instances to generate audio streams corresponding to the text segments in a streaming manner and store them in real time in the object storage system to generate independent audio files; S4. When the independent audio file corresponding to the first text segment is generated, immediately return an ordered list containing the predefined URLs of all audio files to the client, so that the client can download and play them sequentially, while the background continuously generates the remaining audio files.
[0023] In the above steps S1 to S4, when the TTS API system receives a long text to be synthesized, it first performs a preliminary split of the text based on punctuation marks (periods, question marks, exclamation marks, semicolons, etc.) to obtain several "initial text segments". To ensure that each segment is long enough for streaming synthesis, the system will perform a backward merge on short segments with a length less than a preset threshold (e.g., 10 characters) until the minimum length requirement is met. Subsequently, a sequential identifier (such as 1, 2, 3...) is attached to each paragraph for subsequent audio file generation, sorting, and corresponding to the client playback order. For each text segment with an identifier, the TTS API system constructs an "independent request packet", which contains: segment text, sequential identifier, timbre source file hash value, audio parameters (sampling rate, bit rate, etc.). These request packets are pushed to the model gateway in parallel through an asynchronous thread pool. The model gateway continuously monitors metrics such as the video memory usage, inference load, and concurrency of each timbre cloning model instance internally, and according to the preset load balancing strategy, distributes each request packet to the currently most idle or most suitable model instance. In this way, regardless of the number of concurrent requests, balanced distribution can be achieved, avoiding overloading of one GPU while other GPUs are idle. The assigned model instance will quickly load the corresponding cloning model according to the timbre hash value in the request packet and start streaming synthesis according to the required parameters. The specific process is as follows: The model first generates and outputs the first batch of voice frames (the first packet of data), and at the same time continuously outputs subsequent frames. The generated voice data is sliced into data blocks of a fixed size and uploaded to the object storage system block by block. The storage end generates a unique file identifier for each audio segment based on the "sequential identifier + hash value" and writes these data blocks in order. In this way, without waiting for the entire segment to be generated, the first batch of written audio shards can be seen. The system detects the writing status of each audio file in real time in the background. Once it detects that the first segment audio file with the "sequential identifier 1" is written and closed, it can trigger: Predefined URL generation: Before the synthesis starts, the system has already pre-generated and reserved all URL formats (based on the hash value and identifier) for each text segment, making it basically accessible even if the file has not been completely written; URL list assembly and return: Sort the predefined URLs of all segments in ascending order of segment sequence, immediately package them into a list, and return it to the client. After receiving the list, the client can immediately use the first URL to download and play the first segment of audio, and the remaining segments will continue to be generated and written in parallel in the background, achieving "generate while playing"; Intermediate Waiting and Priority Boost: If the client attempts to access a URL that has not been fully written, the system will return an "intermediate status", indicating the estimated waiting time, and at the same time automatically increase the priority of this segment in the model gateway. After the writing is completed, the final available link will be pushed. Through text segmentation in parallel, the first segment has a lower delay, and can be generated and returned to the client. Compared with the traditional serial method of waiting for the entire segment to be synthesized, the user waiting time is significantly shortened, and it can be returned after an audio is generated. It can effectively reduce the user waiting time under high concurrency, reduce the network bandwidth and latency pressure, ensure real-time performance, and there will be no reconnection disasters. Moreover, it can make full use of the underlying model services and will not be affected by the serial processing characteristics of the CosyVoice2 model. The client can immediately start playing the first segment of the audio. Even if the subsequent paragraphs are still being synthesized, it will not affect the continuous playback experience of the previous content. When the concurrency increases, the model gateway can dynamically distribute requests to multiple GPU instances, avoiding long queue waits after a single card's video memory is full. Each instance loads or unloads the model as needed, effectively reducing video memory fragmentation and ensuring video memory utilization rate.
[0024] In a preferred embodiment, the step of receiving the text to be synthesized and splitting it into multiple text segments according to punctuation rules includes: S101. The TTS API system receives the text to be synthesized and detects the set of sentence boundary punctuation marks in the text to be synthesized. Among them, the set of punctuation marks includes full stops, question marks, exclamation marks, and semicolons; S102. Perform primary segmentation at each detected boundary punctuation mark position to generate initial text segments; If the length of the text segment generated by the primary segmentation is less than the preset character condition, then merge it backward to the adjacent segment until the preset character condition is met; S103. Append sequential identifiers to each initial text segment to obtain multiple text segments.
[0025] It should be noted that the preset character condition is ten characters.
[0026] In the above steps S101 to S103, when the TTS API system receives the complete text to be synthesized, it first retrieves the four types of boundary punctuation marks, namely "period, question mark, exclamation mark, and semicolon" from the text. The system can use regular expressions or tokenizers to scan the entire string. Once any of the above punctuation marks is matched, it can be determined that this is a "sentence boundary position". The reason for selecting these punctuation marks is that they usually represent the end points of independent semantic units in the Chinese context, which can ensure the semantic integrity of each paragraph after segmentation to the greatest extent and reduce the sense of discontinuity in the synthesized spoken language caused by segmentation at non-natural pauses. At each recognized boundary symbol, "primary segmentation" is performed, that is, the position where the punctuation mark is located is used as the segmentation point, and the entire text is split into several "initial text segments". For example, "Today the weather is very good. Let's go for a walk together!" will be initially split into two paragraphs: "Today the weather is very good," and "Let's go for a walk together!". In some cases, the length of the text segments generated by the primary segmentation may be too short (less than 10 characters). If only these extremely short segments are synthesized into speech, the proportion of the time for starting and warming up the model will be very high, and the synthesized audio time segments will be too scattered and the tone will be discontinuous. Therefore, the system will automatically merge the segments with a length of less than 10 characters backward with the adjacent segments until the length of the merged text is greater than or equal to the threshold. In this way, it is ensured that each segment is neither too short nor can the integrity of the sentence itself be ignored. For each text segment after the length merging process, a "unique sequence identifier" (such as 1, 2, 3...) is assigned one by one according to their order in the original text. This identifier is not only used to ensure that the audio of each paragraph can be played in the correct order during subsequent parallel inference, but also facilitates the backend object storage system to quickly locate and manage when writing or reading the segmented audio files. Each finally generated "text segment" consists of three parts of information: the paragraph content, the paragraph length meets the requirements, and the sequence identifier corresponding to the paragraph. By performing primary segmentation only at the common sentence-ending punctuation marks, the natural pause positions of the segmented text are more in line with the Chinese semantic habits, and users can obtain better coherence when listening, without the unnatural sense of sentence breaking caused by premature cutting. Merging the short segments with less than 10 characters can avoid the generation of extremely short audio such as "synthesizing with only one or two characters", reduce the sense of stiffness in hearing, and also reduce the obvious discontinuity caused by the frequent splicing of "short-time audio blocks" during playback.
[0027] In a preferred embodiment, the step of parallelly sending multiple text segments to the model gateway and distributing them to the corresponding voice cloning model instances by the model gateway through a preset load balancing strategy includes: S201. Create independent request packets based on each text segment; S202. Batch push multiple independent request packets to the load balancing proxy entry of the model gateway through an asynchronous thread pool; S203. Obtain real-time resource monitoring data based on the model gateway, and execute a preset load balancing strategy to distribute multiple independent request packets to the corresponding voice cloning model instances.
[0028] It is worth mentioning that the independent request packet includes the text segment content, sequence identifier, hash value of the voice source file, and audio parameter configuration.
[0029] In the above steps S201 to S203, for each text segment, an "independent request packet" is constructed, which contains the text segment content (string), sequence identifier (to ensure subsequent audio splicing in the correct order), hash value of the voice source file (used to quickly locate the corresponding cloning model in the model gateway or computing power layer), and audio parameter configuration (such as sampling rate, bit rate, number of channels, etc.). The purpose of this is to encapsulate all the context information of each segment in a request packet, so that subsequent inferences can be completely independent without the need to query any additional information at the model end. The TTS API system internally maintains an asynchronous thread pool (or coroutine pool / asynchronous task queue), and batches and "parallelizes" all independent request packets to the load balancing proxy interface of the model gateway at the same moment. Each thread in the thread pool is responsible for an HTTP(S) or gRPC call, thus making full use of the CPU and network bandwidth, and not blocking other requests due to the waiting of a single request. When a large number of segments are initiated simultaneously, the asynchronous thread pool can start dozens or even hundreds of concurrent requests within milliseconds to dozens of milliseconds, ensuring that short text segments can quickly reach the model gateway without having to wait linearly segment by segment. The model gateway continuously collects real-time resource monitoring data of each downstream voice cloning model instance, including GPU video memory occupancy and remaining space, GPU computing utilization (i.e., current inference latency and throughput), number of concurrent requests or current queue length, and network bandwidth usage. Under the preset load balancing strategy (such as minimum video memory occupancy first, least latency first, or multi-dimensional weighted scheduling), the model gateway distributes each arriving independent request packet to the instance that is "currently the most idle" or "most suitable" for deploying the corresponding voice cloning model. Once the distribution is completed, the instance immediately starts loading the corresponding hash voice model and executes streaming inference, while returning shard results in real-time, and other requests are distributed to different instances for parallel completion. Encapsulating each text segment into a separate request packet and using the asynchronous thread pool for concurrent sending can push dozens to hundreds of inference requests to the model gateway simultaneously within milliseconds. Compared with serial pushing, the network throughput rate is higher and the queuing waiting time is shorter. The model gateway makes an immediate load balancing decision and distributes the requests to the idle or lowest-latency instance, avoiding the instantaneous congestion of a single GPU resource, thus significantly reducing the tail latency of single-segment inference and making the overall synthesis experience smoother.
[0030] In a preferred embodiment, the step of controlling the timbre cloning model instance to generate an audio stream corresponding to the text segment in a streaming manner and storing it in the object storage system in real time to generate an independent audio file includes: S301. Receive an independent request packet, and parse the text segment content, sequence identifier, hash value of the timbre source file, and audio parameter configuration therein; S302. Load the corresponding timbre cloning model instance based on the hash value of the timbre source file, and configure the audio parameters; S303. Based on the loaded timbre cloning model instance, generate an audio data stream corresponding to the text segment content in a streaming manner; S304. Split the audio data stream into continuous data blocks; S305. Upload the split data blocks to the object storage system in real time; S306. In the object storage system, generate a unique audio file identifier according to the sequence identifier and the hash value of the timbre source file, and sequentially write the received data blocks into the independent audio file corresponding to the audio file identifier.
[0031] In the above steps S301 to S306, after the voice cloning model instance (which can be deployed on a GPU node or in a container) receives the "independent request packet" forwarded by the model gateway, it first parses the key fields contained therein: the text fragment content, which is the specific text string to be synthesized; the sequence identifier indicating the order of the fragment in the entire original text; the voice source file hash value used to quickly locate the corresponding cloned voice model (i.e., which person / character / anchor voice sample); and the audio parameter configuration including sampling rate (such as 22 kHz, 24 kHz), bit rate (such as 16 kbps, 32 kbps), number of channels (mono / stereo), etc. After parsing, the model instance can obtain all the context, sequence information, and voice configuration required to synthesize this segment of speech, without the need to query or wait for other services additionally. According to the "voice source file hash value", the model instance quickly locates the corresponding voice cloning model file locally (for example, the weights of a neural network that has been pre-trained and slightly fine-tuned). If the model has been loaded into the video memory and kept in a hot state before, it can be directly reused. Otherwise, the instance will load the corresponding voice model weights from the disk or cache into the GPU video memory and perform necessary initializations (such as building an inference graph, loading a decoder, initializing a streaming decoding buffer). At the same time, according to the audio parameter configuration carried in the request packet, the synthesis pipeline (streaming decoder, post-processing module, etc.) is set to ensure that the final output audio meets the specified sampling rate, bit rate, or number of channels requirements. The voice cloning model instance starts "streaming inference". The model first inputs the text fragment into the front-end encoder (usually including character-to-phoneme or word-embedding-to-context feature extraction), and the back-end decoder generates the input frames for the vocoder in the way of "decoding while outputting", outputting immediately every time a small segment of waveform data is generated, avoiding waiting for the entire text to be processed and then outputting all at once. In each step of the generation process, the model will generate several "audio frames" or "voice feature vectors", and these feature vectors are converted into corresponding PCM waveform data in real time through a neural vocoder (such as WaveGlow, HiFi-GAN). To adapt to the optimal granularity for network transmission and storage writing, the model instance will internally split the generated continuous PCM data into fixed-size data blocks. For example, each 20 ms or 40 ms of audio corresponds to a data block. Each data block will carry the sequence identifier metadata of this audio segment, as well as the offset information (data block number) of this segment in this fragment. This can ensure that the subsequent storage end knows how to reassemble the blocks in order into a complete file. The split data blocks are immediately uploaded to the object storage system through an asynchronous I / O interface. When uploading, each data block carries a file identifier, which consists of a unique prefix composed of "voice hash value + sequence identifier" plus a suffix of the block number, and the data block content. Since asynchronous calls are used, the upload operation does not block the streaming inference, but instead "streamingly generates one block → asynchronously uploads one block", achieving parallelization between network latency and encoding latency and reducing the overall waiting time.After receiving the uploaded data chunks, object storage systems (such as MinIO and Ceph) will, according to the unique audio file identifier pre-generated based on the "sequence identifier + timbre hash value", aggregate and write the chunked data into the same object (file) in the order of "data chunk serial number". In fact, object storage will use the multi-part upload or append-write mechanism for each chunk as the same file, and each shard will be spliced into the same file logic in the background. During the writing process, if all the data chunks of the same file have not all arrived, the storage system can mark the existing chunks as "accessible" and return the partial content status when the URL is accessible. Each generated audio data chunk is immediately uploaded asynchronously, and the model does not need to wait for the entire audio to be generated before writing it all at once. Since inference and upload are carried out in parallel, the model instance will not be blocked by upload I / O and can start synthesizing subsequent data chunks or the next parallel task as soon as possible, improving the overall throughput. In case of abnormal situations such as network fluctuations or server restart, as long as the uploaded chunks exist, there is no need to retransmit the previously completed shards when resuming the upload later, reducing resource waste.
[0032] In a preferred embodiment, when the independent audio file corresponding to the first text segment is generated and completed, an ordered list containing all the predefined URLs of the audio files is immediately returned to the client, enabling the client to download and play them in sequence. Meanwhile, the background continues the steps of generating the remaining audio files, including: S401. Continuously monitor the generation status of the independent audio files corresponding to multiple text segments in real time; S402. If it is detected that the independent audio file corresponding to the first text segment is written and closed in the object storage system, immediately generate and return a list of all predefined URLs sorted by sequence identifier to the client; S403. After returning the ordered URL list to the client, the background continues to execute the process of parallelly sending multiple text segments to the model gateway to process the remaining text segments until all the independent audio files corresponding to the text segments are generated in the object storage system; S404. When the client requests a non-generated file, return an intermediate response containing the waiting time.
[0033] In steps S401 to S404 as described above, the system starts a "status monitoring module" in the background, continuously polling or subscribing to the writing and closing status of the audio files corresponding to each text segment of this synthesis task from the object storage (such as MinIO). Whenever a new chunk is written, the object storage will trigger a callback or update the metadata, and the monitoring module can learn the current writing progress of the file. Once the entire file has completed writing and closing all chunks, it marks this segment as "completed". The monitoring module first focuses on the first segment with a "sequence identifier of 1". When it detects that the audio file of this segment has been fully written and closed in the object storage, it immediately triggers the following operations: URL predefined. The system has previously generated and reserved "predefined URLs" for all segments according to "voiceprint hash value + sequence identifier". Even if the corresponding files have not been generated yet, these URLs have been mapped to the access paths of the object storage, allowing the client to download at any time; Ordered list assembly. Arrange the predefined URLs of all segments in ascending order according to the sequence identifier (1, 2, 3,...) to form a complete and ordered URL list (e.g., [url_1, url_2, url_3,...]); Immediately respond to the client. Through the HTTP response, return this ordered URL list to the client. At this time, the client can start downloading with the first URL and playing the first segment of the audio; Even if the first segment has been generated and returned to the client, the background still continues to execute the complete process of "segment packet sending → model gateway inference → streaming writing → storage splicing" in parallel until all independent audio files corresponding to all text segments are generated. Since the client has obtained all the URLs - even if the subsequent files have not been generated yet, the client will request or buffer the next URL in advance when playing the first segment. At this time, if the subsequent audio is being uploaded in chunks, the client will obtain the shard data from the object storage and start playing. When the client requests in advance an audio URL that has not been completely generated yet (or some chunks are missing), the storage layer will return a partial content status or encapsulate it into a custom "file not ready" intermediate response, prompting the client that the file is still being written and informing the approximate waiting time. At this time, the system will dynamically increase the priority of its inference and writing inside the model gateway for this segment, for example, marking this request as "urgent", so that the backend model instance can complete the generation and upload of the remaining chunks of this segment as soon as possible. Once the audio of this segment is written, the object storage system will push a "file ready" notification. After receiving it, the client can automatically retry the download and continue playing without manual intervention. As long as the first segment file with "sequence identifier 1" is written, all the URL lists of all segments can be returned immediately. The client will start playing the first segment immediately, avoiding the long-time lag when starting to play after the whole segment is synthesized. For the user, it only takes about 150 ms (the first packet delay of the CosyVoice2 model) to hear the first sentence, greatly improving the real-time experience. Since all the URL lists sorted by the sequence identifier are returned at one time, the client can know in advance the order of the files to be downloaded, and can cache the subsequent segments in parallel when playing the first segment, ensuring no lag during the switch and not needing to dynamically calculate which the next URL should be. Even if a certain segment has not been completely generated yet, the client can still play the downloaded content first and can seamlessly splice it when this segment is ready. The user can hardly feel the discontinuity.
[0034] In a preferred embodiment, if it is detected that the independent audio file corresponding to the first text segment has been written and closed in the object storage system, the step of immediately generating and returning a list of all predefined URLs sorted by sequential identifiers to the client includes: S4021, if it is detected that the independent audio file corresponding to the first text segment is written and closed in the object storage system, a predefined URL is generated in the object storage system for each text segment according to the sequence identifiers of the multiple text segments, wherein the predefined URL is generated in advance based on the sequence identifier and the hash value, and is accessible when the file is not completed; S4022, sorting the generated URLs from small to large according to the values of the sequence identifiers to form an ordered URL list; S4023. Immediately return an ordered URL list to the client that initiated the request.
[0035] As in the above steps S4021 to S4023, when the system detects that the audio file corresponding to the first text segment has been completely written and closed in the object storage, it will immediately start the generation process of predefined links. Before this process starts, the system has determined the unique path of each segment in the storage according to the sequence number and timbre model hash value of each text segment. However, these links only exist in a reserved state until the first segment is generated. Once the first segment of audio is confirmed to be readable, the system will generate formal access links for all segments according to the predefined naming rules. Each link maps to the file location in the storage system. And even if the corresponding file has not been completely uploaded, when the client accesses these links, the storage system will return responses such as "file is being generated" or "partially accessible" to prepare for direct downloading after the subsequent file is ready. After generating the predefined links corresponding to all segments, the system arranges these links into a complete list according to the sequence number. At this time, the links with smaller sequence numbers are in the front of the list, and the links with larger numbers are at the back, so that the client can clearly know the playing order. Almost instantaneously, this link list that has been arranged in the correct order will be packed into a unified response and sent back to the client through a single network request. After the client gets this set of links, it can immediately start the download request for the first audio segment. As for the subsequent links, even if the corresponding files are still being generated in the background, the client can call these links in advance and start downloading immediately once the files are generated, thus minimizing the waiting time during playback. Once the system completes the processing of the first segment of audio, the user can quickly obtain the access addresses of the entire audio sequence without having to wait until all paragraphs are generated to obtain the links. For the user, the playback experience becomes smoother. The first segment of audio starts playing almost immediately after receiving the links, and the subsequent paragraphs are continuously generated in the background, and the user will not feel long-term lags or waits. At the same time, returning all the links to the client at once also greatly simplifies the client's logic, eliminating the need for multiple polls to query the generation status of each segment and the need to frequently initiate requests when it is uncertain whether the file exists. In this way, the number of network requests of the client is significantly reduced, and the server side will not cause additional burden due to a large number of polls.
[0036] In a preferred embodiment, the step of returning an intermediate response including the waiting time when the client requests a non-generated file includes: S4041. Monitor in real time the client's download requests for subsequent URLs in the ordered list; S4042. If it is detected that the audio file requested by the client for download has not been generated completely in the object storage system, the processing priority of the corresponding text segment in the model gateway is increased, and the status code and the estimated waiting time are returned to the client. At the same time, the write completion monitoring of the corresponding text segment in the object storage system is triggered. When it is monitored that the corresponding text segment is completely written, the final URL that can be redirected to the complete file is immediately pushed to the client As in steps S4041 to S4042 above, when the client attempts to access an audio file that has not been fully generated, the system will continuously monitor in the background all the URL lists returned to the client. When the monitoring module detects that the download request for a subsequent paragraph in the list reaches the client, it will first initiate a query to the object storage to determine whether the corresponding file is "fully written and closed". If it is found that the audio file is still being generated, that is, a complete object has not yet been formed in the storage system, then the system will immediately dynamically adjust the task priority of this segment at the model gateway. Specifically, the monitoring module will send an internal instruction to the model gateway to promote the inference request where this segment is located from the normal queue to the high-priority queue, ensuring that the remaining synthesis work of this audio segment is preferentially processed during the model instance scheduling. At the same time, the system will not let the client wait idly, but will quickly return an "intermediate business execution response" to the client. In this response, in addition to telling the client that the requested file is not yet ready, it will also give an estimated waiting time (such as approximately how many milliseconds or seconds it will take to see the complete file). When this segment is finally written in the object storage and the file status becomes "readable", the storage system will trigger a write completion event. After the monitoring module captures this event, it will immediately notify the client or tell the client in a redirect manner that "you can now use this link to redownload the audio of the entire paragraph", thus completing the closed-loop from the entire intermediate response to the final available link. The client does not need to continuously poll whether the "file is fully generated". When the system detects the download request and finds that the file is not yet ready, it will actively tell it "please wait approximately how much time", allowing the user to clearly know how long they need to wait, rather than letting the player get stuck in an unresponsive state for a long time. By dynamically increasing the priority of the urgently needed paragraphs of the consumer, the model gateway can concentrate resources on the most urgent requests, enabling the remaining inference and writing of this segment to be completed as soon as possible, thereby reducing the overall waiting time. Even when there are multiple concurrent synthesis tasks in the background, once they click on a paragraph and are ready to play, the system will give priority to processing this request, significantly reducing the lag when searching for the next audio segment. At the same time, through the write completion event triggering mechanism, the client can immediately initiate a redirect download after receiving the "file is ready" notification, without having to initiate a new round of polling requests, thereby reducing unnecessary network overhead and also reducing the pressure on the server interface caused by a large number of polls in a high-concurrency scenario. Overall, this process not only ensures the real-time response experience but also makes the computing power scheduling in the background more flexible and efficient.
[0037] And, a solution terminal for TTS with cloned voices in a high-concurrency scenario, including: One or more processors; A storage device on which one or more programs are stored; When one or more programs are executed by one or more processors, such that the one or more processors implement a solution for cloning timbres for TTS in a high-concurrency scenario.
[0038] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. The structures, devices, and operation methods not specifically described and explained in the present invention are implemented according to the conventional means in the art without special instructions and limitations.
Claims
1. A solution for cloning tones in TTS in high-concurrency scenarios, characterized in that, Include: Receive the text to be synthesized and split it into multiple text segments according to punctuation rules; Parallelly send the multiple text segments to the model gateway, and the model gateway distributes them to the corresponding voice cloning model instances through a preset load balancing strategy; Control the voice cloning model instances to generate audio streams corresponding to the text segments in a streaming manner and store them in the object storage system in real time to generate independent audio files; When the independent audio file corresponding to the first text segment is generated, immediately return an ordered list containing the predefined URLs of all audio files to the client, enabling the client to download and play them sequentially, while the background continuously generates the remaining audio files.
2. The solution method for TTS of cloned timbre in a high-concurrency scenario according to claim 1, characterized in that The step of receiving the text to be synthesized and splitting it into multiple text segments according to punctuation rules includes: The TTS API system receives the text to be synthesized and detects the sentence boundary punctuation symbol set in the text to be synthesized. Among them, the punctuation symbol set includes full stops, question marks, exclamation marks, and semicolons; Perform primary splitting at each detected boundary punctuation symbol position to generate initial text segments; If the length of the text segment generated by the primary splitting is less than the preset character condition, then merge it backward to the adjacent segment until the preset character condition is met; Attach sequential identifiers to each initial text segment to obtain multiple text segments.
3. The solution method for TTS of cloned timbre according to claim 2 in a high-concurrency scenario, characterized in that, The preset character condition is ten characters.
4. The solution method for TTS of cloned timbre in a high-concurrency scenario according to claim 1, characterized in that, The step of parallelly sending the multiple text segments to the model gateway, and the model gateway distributes them to the corresponding voice cloning model instances through a preset load balancing strategy includes: Create independent request packets based on each text segment; Batch push the multiple independent request packets to the load balancing proxy entry of the model gateway through an asynchronous thread pool; Based on the model gateway, obtain real-time resource monitoring data and execute the preset load balancing strategy to distribute the multiple independent request packets to the corresponding voice cloning model instances.
5. The solution method for TTS using cloned timbres in a high-concurrency scenario according to claim 4, characterized in that, The independent request packet includes the text segment content, sequential identifier, hash value of the voice source file, and audio parameter configuration.
6. The solution method for TTS of cloned timbre in a high-concurrency scenario according to claim 5, characterized in that, The step of controlling the voice cloning model instances to generate audio streams corresponding to the text segments in a streaming manner and storing them in the object storage system in real time to generate independent audio files includes: Receive the independent request packet and parse the text segment content, sequential identifier, hash value of the voice source file, and audio parameter configuration therein; Load the corresponding voice cloning model instance based on the hash value of the voice source file and configure the audio parameters; Based on the loaded voice cloning model instance, generate an audio data stream corresponding to the text segment content in a streaming manner; Split the audio data stream into continuous data blocks; Upload the split data blocks to the object storage system in real time; In the object storage system, generate a unique audio file identifier according to the sequential identifier and the hash value of the voice source file, and sequentially write the received data blocks into the independent audio file corresponding to the audio file identifier.
7. The solution method for TTS of cloned timbre in a high-concurrency scenario according to claim 1, characterized in that, The step of, when the independent audio file corresponding to the first text segment is generated, immediately returning an ordered list containing the predefined URLs of all audio files to the client, enabling the client to download and play them sequentially, while the background continuously generates the remaining audio files includes: Real-time monitor the generation status of the independent audio files corresponding to the multiple text segments; If it is detected that the independent audio file corresponding to the first text segment has been written and closed in the object storage system, a list of all predefined URLs sorted by sequential identifiers is immediately generated and returned to the client; After returning the ordered URL list to the client, the background continues to send multiple text segments in parallel to the model gateway to process the remaining text segments until all independent audio files corresponding to the text segments are generated in the object storage system; When a client request does not generate a file, an intermediate response containing the waiting time is returned.
8. The solution method for TTS of cloned timbre in a high-concurrency scenario according to claim 7, characterized in that, If it is detected that the independent audio file corresponding to the first text segment is written and closed in the object storage system, the steps of immediately generating and returning a list of all predefined URLs sorted by sequential identifiers to the client include: If it is detected that the independent audio file corresponding to the first text segment is written and closed in the object storage system, a predefined URL is generated in the object storage system for each text segment according to the sequential identifiers of the multiple text segments; Sort the generated URLs according to the values of the sequence identifiers from small to large to form an ordered URL list; Immediately returns an ordered list of URLs to the requesting client.
9. The solution method for TTS of cloned timbres in a high-concurrency scenario according to claim 7, characterized in that, When the client request does not generate a file, the steps to return an intermediate response containing a waiting time include: Real-time monitoring of client download requests for subsequent URLs in the ordered list; If it is detected that the audio file requested to be downloaded by the client has not been generated in the object storage system, the processing priority of the corresponding text fragment in the model gateway is increased, and the status code and estimated waiting time are returned to the client. At the same time, the write completion monitoring of the corresponding text fragment in the object storage system is triggered. When the corresponding text fragment is detected to be written, the final URL that can be redirected to the complete file is immediately pushed to the client.
10. A solution terminal for TTS of cloned timbres in high-concurrency scenarios, characterized in that, include: one or more processors; a storage device having one or more programs stored thereon; When one or more programs are executed by one or more processors, the one or more processors implement the solution of cloning the tone and performing TTS in a high-concurrency scenario as described in any one of claims 1 to 9.
Citation Information
Cited By
TTS audio generation system and method based on sound cloning
CN121171202A
Streaming audio synthesis method and device, storage medium and electronic device
CN121600905A