Oral language understanding model training method and system

Through the dynamic blocked autoregressive synthesis (DCAR) method, the multi-head prediction mechanism and reinforcement learning training strategy are used to solve the problems of stability and efficiency in long speech sequence processing by traditional autoregressive speech synthesis model, and high-quality and efficient speech synthesis are achieved.

CN120496498APending Publication Date: 2025-08-15AISPEECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510827965.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-15

Smart Images

  • Figure CN120496498A_ABST
    Figure CN120496498A_ABST
Patent Text Reader

Abstract

The invention discloses a speech synthesis method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: pre-training a TTS basic model; training a dynamic block scheduling strategy according to the TTS basic model; and performing dynamic block voice marking according to the dynamic block scheduling strategy, and performing voice synthesis according to a dynamic block voice marking result. According to the speech synthesis method, the optimal block size can be adaptively determined in each decoding step, so that the speech synthesis quality and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech synthesis technology, and in particular to a speech synthesis method, device, electronic device, storage medium and program product. Background Art

[0002] Among related technologies, autoregressive (AR) language models have become a mainstream approach for speech synthesis, capable of generating expressive speech while also offering scalable training capabilities. However, traditional AR speech synthesis models, which rely on the "next token prediction" paradigm, often encounter significant challenges when processing long speech sequences. They struggle to establish stable inter-frame attention, resulting in increased latency and reduced synthesis quality, limiting their feasibility for real-time applications. Summary of the Invention

[0003] The embodiments of the present application provide a speech synthesis method, device, electronic device, storage medium and program product, which are used to solve at least one of the above technical problems.

[0004] In a first aspect, an embodiment of the present application provides a speech synthesis method, comprising: Pre-train the TTS basic model; Training a dynamic block scheduling strategy based on the TTS basic model; Dynamic block speech marking is performed according to the dynamic block scheduling strategy, and speech synthesis is performed according to the result of the dynamic block speech marking.

[0005] The speech synthesis method of this embodiment can adaptively determine the optimal block size in each decoding step, thereby improving the quality and efficiency of speech synthesis.

[0006] In some embodiments, the TTS basic model includes 1 basic head and n-1 additional heads, wherein the basic head and the additional heads are used to predict the next n tokens; pre-training the TTS basic model includes: The basic head is used to predict the next token; The additional head is used to predict the second to n subsequent tokens; The TTS basic model is obtained based on the loss training of the basic head and the additional head.

[0007] The solution of this embodiment utilizes a multi-head prediction mechanism for training, enabling the speech synthesis model to model the relatively holistic acoustic part, and uses reinforcement learning methods to learn to schedule the predicted number of tokens to achieve the best intelligibility performance.

[0008] In some embodiments, the dynamic scheduling strategy is executed by a scheduling module, which includes a linear layer and a causal transformation layer connected after the linear layer, and the causal transformation layer is used to take the historical hidden state sequence of the predicted speech sequence as input.

[0009] In some embodiments, the training objective of the dynamic block scheduling strategy is set to:

[0010] Among them, the advantages is defined as .

[0011] In some embodiments, the training of a dynamic block scheduling strategy based on the TTS basic model includes: Determine a preset range of block length based on CAR speech synthesis; A relatively negative process reward is set for actions that exceed the preset range of the block length. The total reward is expressed as follows: .

[0012] In some embodiments, the base header includes a linear layer and the additional header includes 4 residual blocks and 1 linear prediction layer.

[0013] In a second aspect, an embodiment of the present application provides a speech synthesis device, comprising: The first training module is used to pre-train the TTS basic model; A second training module is used to train a dynamic block scheduling strategy based on the TTS basic model; The speech synthesis module is used for dynamically marking speech in blocks according to the dynamic block scheduling strategy, and performing speech synthesis according to the result of the dynamic block speech marking.

[0014] In a fifth aspect, an embodiment of the present application provides a storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned speech synthesis methods of the present application.

[0015] In a sixth aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above-mentioned speech synthesis methods of the present application.

[0016] In a seventh aspect, an embodiment of the present application further provides a computer program product, which includes a computer program stored on a storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned speech synthesis methods.

[0017] The speech synthesis method of this embodiment can adaptively determine the optimal block size at each decoding step, thereby improving the quality and efficiency of speech synthesis. The speech synthesis method of this embodiment can be implemented as a Dynamic Chunk-wise Autoregressive Synthesis (DCAR) method. This DCAR method uses a multi-head pre-trained TTS model and a reinforcement learning training strategy model to achieve a dynamic number of speech tokens decoded at each step, enhancing the stability of speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A flowchart of an embodiment of the speech synthesis method of the present application; Figure 2 A flowchart of another embodiment of the speech synthesis method of the present application; Figure 3 A flowchart of another embodiment of the speech synthesis method of the present application; Figure 4 A comparative diagram of speech synthesis using FAR, CAR, and DCAR in this application; Figure 5a Schematic diagram of the first layer attention weight of the FAR model; Figure 5b Schematic diagram of the first layer attention weight of the CAR model; Figure 5c Schematic diagram of FAR attention amplification; Figure 5d Schematic diagram of CAR attention amplification; Figure 6 Schematic diagram of speech synthesis taking the case of block size 2 as an example; Figure 7 A schematic diagram of the inference action of the CAR model with 6 additional heads added to the GT sequence is shown; Figure 8Schematic diagram of WER performance under different frame rates and strategies; Figure 9 This is a diagram showing the evaluation results of word error rate when using HuBERT tagging in this application; Figure 10 This is a diagram showing the evaluation results of word error rate when using S3 tokenizer for tokenization in this application; Figure 11 Schematic diagram of the evaluation results of the synthetic performance of HuBERT tagging on the LibriHeavy dataset in this application; Figure 12 This is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0021] It should also be noted that, in this document, the terms "include" and "comprising" include not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, the elements defined by the phrase "include..." do not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the elements.

[0022] like Figure 1 FIG. 1 is a flow chart of an embodiment of a speech synthesis method of the present application. In this embodiment, the speech synthesis method includes: S10. Pre-train the TTS basic model.

[0023] For example, the TTS basic model can adopt a pure decoder TTS architecture, which consists of 12 layers of transformers, each layer has 8 attention heads, 1024 hidden dimensions, and 4096 feedforward dimensions. It should be noted that the specific structure of the TTS basic model is not limited in this application.

[0024] S20: Training a dynamic block scheduling strategy according to the TTS basic model.

[0025] In some embodiments, the dynamic scheduling strategy is executed by a scheduling module, which includes a linear layer and a causal transformation layer connected after the linear layer, wherein the causal transformation layer is configured to take as input the historical hidden state sequence of the predicted speech sequence. Exemplarily, the training objective of the dynamic block scheduling strategy is set to:

[0026] Among them, the advantages is defined as .

[0027] S30 , dynamically marking speech in blocks according to the dynamic block scheduling strategy, and performing speech synthesis based on the result of the dynamic block speech marking.

[0028] The speech synthesis method of this embodiment can adaptively determine the optimal block size in each decoding step, thereby improving the quality and efficiency of speech synthesis.

[0029] like Figure 2 The flowchart of one embodiment of the speech synthesis method of the present application is shown. In this embodiment, the TTS basic model includes 1 basic head and n-1 additional heads, wherein the basic head and the additional heads are used to predict the next n tokens; the pre-trained TTS basic model includes: S11, using the basic header to predict the next tag; S12, using the additional head to predict the second to n subsequent tags; S13. Obtain the TTS basic model based on the loss training of the basic head and the additional head. In some embodiments, the basic head includes a linear layer, and the additional head includes four residual blocks and one linear prediction layer.

[0030] The solution of this embodiment utilizes a multi-head prediction mechanism for training, enabling the speech synthesis model to model the relatively holistic acoustic part, and uses reinforcement learning methods to learn to schedule the predicted number of tokens to achieve the best intelligibility performance.

[0031] like Figure 3 FIG. 1 is a flow chart of an embodiment of the speech synthesis method of the present application. In this embodiment, the dynamic block scheduling strategy is trained based on the TTS basic model, including: S21, determining a preset range of block length based on CAR speech synthesis; S22. Set a relatively negative process reward for actions that exceed the preset range of the block length. The total reward is expressed as follows: ; Among them, WER gen and WERgt refers to the WER metric of the generated and real speech, and T refers to the length of the generated speech sequence. To avoid reward explosion, the value of ϵ is -10, and λ is usually set to 0.1, so that the absolute value of the total negative reward is less than the order of magnitude of the WER supervision reward.

[0032] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0033] An embodiment of the present application further provides a speech synthesis device, comprising: The first training module is used to pre-train the TTS basic model; A second training module is used to train a dynamic block scheduling strategy based on the TTS basic model; The speech synthesis module is used for dynamically marking speech in blocks according to the dynamic block scheduling strategy, and performing speech synthesis according to the result of the dynamic block speech marking.

[0034] The speech synthesis apparatus of this embodiment can adaptively determine the optimal block size in each decoding step, thereby improving the quality and efficiency of speech synthesis.

[0035] In some embodiments, the TTS basic model includes 1 basic head and n-1 additional heads, wherein the basic head and the additional heads are used to predict the next n tokens; pre-training the TTS basic model includes: The basic head is used to predict the next token; The additional head is used to predict the second to n subsequent tokens; The TTS basic model is obtained based on the loss training of the basic head and the additional head.

[0036] The solution of this embodiment utilizes a multi-head prediction mechanism for training, enabling the speech synthesis model to model the relatively holistic acoustic part, and uses reinforcement learning methods to learn to schedule the predicted number of tokens to achieve the best intelligibility performance.

[0037] In some embodiments, the dynamic scheduling strategy is executed by a scheduling module, which includes a linear layer and a causal transformation layer connected after the linear layer, and the causal transformation layer is used to take the historical hidden state sequence of the predicted speech sequence as input.

[0038] In some embodiments, the training objective of the dynamic block scheduling strategy is set to:

[0039] Among them, the advantages is defined as .

[0040] In some embodiments, the training of a dynamic block scheduling strategy based on the TTS basic model includes: Determine a preset range of block length based on CAR speech synthesis; A relatively negative process reward is set for actions that exceed the preset range of the block length. The total reward is expressed as follows: .

[0041] In some embodiments, the base header includes a linear layer and the additional header includes 4 residual blocks and 1 linear prediction layer.

[0042] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned speech synthesis methods of the present application.

[0043] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned speech synthesis methods.

[0044] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a speech synthesis method.

[0045] To make the technical solution and effects of this application clearer, the research, exploration, implementation and experimental demonstration process of the invention are presented as follows: Recently, autoregressive (AR) language models have become a mainstream approach for speech synthesis, capable of generating expressive speech while also offering scalable training capabilities. However, traditional AR speech synthesis models, which rely on the "next token prediction" paradigm, often face significant challenges when processing long speech sequences: they struggle to establish stable inter-frame attention, resulting in increased latency and reduced synthesis quality, limiting their feasibility for real-time applications. To address these issues, this application proposes a novel dynamic chunk-wise autoregressive synthesis framework, dubbed DCAR (Dynamic Chunk-wise Autoregressive Synthesis), designed to simultaneously improve the efficiency and robustness of AR speech generation. DCAR introduces a chunk-to-frame attention mechanism through multi-token prediction training, enabling a lightweight module to dynamically predict chunks based on different contexts under on-policy training. This framework dynamically adjusts the prediction span, significantly reducing sequence length dependence while maintaining high-quality synthesis. Comprehensive empirical evaluation demonstrates that, compared to traditional "next token prediction" models, DCAR achieves up to 72.27% improvement in intelligibility and 2.61× acceleration in inference speed on a test set. Furthermore, we provide an in-depth analysis demonstrating its potential as a versatile foundation for next-generation speech synthesis systems.

[0046] 1. Introduction Autoregressive (AR) architectures are widely used for both comprehension and generation tasks in natural language processing (NLP). These models typically operate within a next-token prediction paradigm, predicting each subsequent token based on the sequence of previous tokens. This approach has underpinned many recent advances in text generation, particularly with the rise of large language models (LLMs).

[0047] In the field of speech, adopting similar AR modeling strategies naturally requires the use of quantization techniques to convert continuous sound signals into discrete representations, effectively acting as speech taggers. To this end, researchers have actively explored various speech discretization methods. Notable approaches include applying vector quantization (VQ) to existing neural representations of speech and training speech codecs using the VQ-Variational Autoencoder (VQ-VAE) architecture. Unlike text or visual tags, speech tags typically correspond to short frames or time window segments of audio, capturing low-level acoustic patterns but not directly encoding semantic information. These characteristics pose challenges to the Transformer architecture, as lengthy speech sequences require autoregressive modeling of long-range dependencies across frames using a next-token prediction paradigm.

[0048] In this application, we investigate whether the frame-level autoregressive (FAR) prediction paradigm is essential for autoregressive speech synthesis models. Inspired by recent developments in LLMs, multi-token prediction, a method for fast decoding and rapid training convergence of LLMs, offers a promising approach for speech synthesis, reducing the computational burden associated with lengthy speech sequences in AR models. We experimentally investigate autoregressive text-to-speech (TTS) using various speech token configurations and find that, at the same sequence position, the token predicted by the first token head is not always superior to the tokens predicted by subsequent token heads. This phenomenon plausibly explains why robust synthesis is achieved by directly predicting a large chunk of speech tokens at each decoding step without any inference strategy. We term this approach "chunk-wise autoregressive synthesis," or CAR. While related approaches have applied CAR to autoregressive TTS and achieved fast or high-quality synthesis using decoding techniques, they still rely on header-based parallel inference verification or simply accept fixed-size token chunks. This leaves room for improvement in token selection strategies, as current methods may retain redundant or suboptimal predictions.

[0049] To this end, this application proposes DCAR, a dynamic block synthesis strategy with two-stage training, which can arrange the appropriate block size in each CAR decoding step. In the first stage, we let the TTS basic model learn the block-to-frame attention pattern through multi-label prediction. The TTS basic model predicts the next n tokens through a basic head and an additional n-1 heads. The basic head is responsible for predicting the next token, and the additional head is responsible for predicting the second to n subsequent tokens. The loss combination coefficients of each head are added together to form the final CAR training standard. Among them, the head is the output layer module in the neural network that processes a specific task; the TTS basic model can use the AR part language model in VALL-E plus CTX-vec2wav as the vocoder.

[0050] Exemplarily, obtaining the TTS basic model based on the loss training of the basic head and the additional head includes: adding the loss combination coefficients of the basic head and the additional head to form a final CAR training standard, and then training the TTS basic model based on the changed CAR training standard.

[0051] DCAR then trains a lightweight policy based on a pre-trained TTS base model to dynamically determine the next chunk length using an on-policy strategy adapted from Group Relative Preference Optimization (GRP). Related research suggests that preference optimization strategies tend to lead models to learn output preferences rather than novel knowledge, suggesting the need to provide guidance for chunk scheduling policies and reduce the space for learning preferences. Therefore, we propose a "chase-then-exceed" strategy: a well-performing chunk size range is identified in a fixed-size CAR and used as a sampling strategy for half of a group, serving as guidance for the lightweight module trained on the policy to sample the remaining chunks. Furthermore, a negative process reward is imposed as a tracking penalty for each out-of-range chunk size to further encourage the policy to act within the desired range. The DCAR strategy for chunk size scheduling demonstrates its effectiveness with only two selections of 980 samples from the LibriTT training set, achieving both stable and competitive speed for TTS.

[0052] The DCAR evaluation was based on four key aspects: robustness of intelligibility, naturalness of synthesis, speaker similarity for zero-shot TTS, and autoregressive decoding speedup. Leveraging HuBERT and S3-Tokenizer tokenization, DCAR simultaneously achieved synthesis intelligibility improvements of up to 72.27% (bit error rate from 9.99% to 2.77%) and 20.34% (bit error rate from 2.31% to 1.84%), as well as synthesis speedups of 2.61x and 2.89x.

[0053] Our main contributions can be summarized as follows: - We propose a novel autoregressive method for speech synthesis, DCAR, which autoregressively predicts dynamic chunked speech tokens, thereby improving synthesis speed and robustness.

[0054] - We use the DCPO algorithm for lightweight policy training, which adopts a catch-up-and-surpass strategy. It effectively speeds up the convergence process and makes full use of good mid-term results.

[0055] - We analyze the block-to-frame attention patterns and predicted actions in the CAR model, providing valuable insights into the nature of self-attention in speech sequences.

[0056] 2. Related Work 2.1 Autoregressive Speech Synthesis Despite their expressive power, AR TTS models have long been plagued by lengthy speech sequences. Compared to text modalities, speaking five words often takes one second, consisting of 50 speech tokens quantized at 50Hz. Predicting long sequences accumulates token loss due to inaccurate decoding, causing instability during inference. Generated speech often exhibits unnatural silences, repeated words, and word omissions. Therefore, robustness is a major issue for current AR TTS models. Previous studies have proposed various strategies to enhance the robustness of generation, including alignment-guided sequence reordering or training, applying transformer loss, and thought chaining prompts.

[0057] When using fine-grained speech representations, high latency and a large semantic gap with the text modality remain major challenges for autoregressive decoding. Existing approaches to address these issues can be broadly categorized into three types: (1) downsampling the representations or leveraging their coarse-grained counterparts; (2) hierarchical generation from coarse-grained representations to fine-grained representations; and (3) adopting block-wise autoregressive modeling with multi-label predictions and carefully designed decoding strategies. Representation Downsampling: In recent years, increasing work has focused on compressing discrete speech representations to extremely low frame rates. One promising approach is to integrate the downsampling module directly into the quantization stage of the self-supervised representation or into a neural speech codec model. However, this strategy often results in a significant loss of intrinsic information in the speech, and coarse-grained speech markers still struggle to match fine-grained markers in capturing acoustic details and other subtle features.

[0058] Coarse-to-fine-grained generation: Since fine-grained representations excel in modeling acoustic details, recent studies have explored autoregressive generation of coarse-grained tokens and then conditioning these tokens using autoregressive or non-autoregressive strategies to predict fine-grained representations.

[0059] Speech synthesis via block-wise multi-label prediction: Multi-label prediction has been proposed in recent years to accelerate the decoding of large language models (LLMs) and speed up convergence during training. Recent research in the speech field has begun to adopt this approach for efficient autoregressive speech recognition and fast, high-quality speech synthesis.

[0060] 2.2 Preference Strategy through Reinforcement Learning Since ChatGPT, preference optimization strategies have attracted extensive attention and exploration. The Proximal Policy Optimization (PPO) algorithm pioneered the use of human evaluation as supervision for the critic model, providing environmental rewards when training a basic LLM as the actor model. Subsequently, Direct Preference Optimization (DPO) simplified this process by making the LLM itself the reward model. With the recent shift to the inference era, DeepSeek-R1 brought the Group Relative Policy Optimization (GRPO) algorithm to the forefront of this trend. GRPO leverages the inter-group advantages derived from multiple rounds of LLM inference as a reward signal, effectively stimulating and enhancing the model's reasoning capabilities. Several studies in the speech field have adapted these preference alignment strategies to advance robust, high-fidelity speech synthesis.

[0061] Our method leverages multi-label prediction and reinforcement learning to further optimize the chunk scheduling strategy, providing a new approach for robust, high-quality, and efficient speech synthesis.

[0062] 3. Re-examining autoregressive TTS using block prediction In this section, we discuss a core question: what is the optimal prediction size for speech synthesis? To explore this question, we implement a basic TTS framework based on multi-label prediction, where each prediction head is trained using either a decreasing or uniform target weighting scheme. We then conduct an extensive analysis to investigate why block-wise prediction outperforms frame-wise prediction and the remaining limitations of fixed-length block-wise prediction.

[0063] 3.1. Implementation of Block Autoregressive Text-to-Speech (TTS) Tagger: Various types of speech tags are used when building AR TTS models. Tags with a codebook are considered semantic tags. We use k-means clustering in the last layer of HuBERT-Large and S3-Tokenizer for unsupervised and supervised semantic tagging, respectively. Specifically, for HuBERT tagging, a k-means model is trained with 2048 centers on a randomly sampled 73-hour subset of the LibriTTS training set. Previous research has shown that phonemes and subwords behave similarly as text types, so for convenience, we choose phonemes as the input text unit.

[0064] Training the CAR decoder with multiple heads: We designed additional lightweight heads on the original pure decoder architecture based on related techniques and jointly trained them from scratch, using bidirectional attention for textual tokens and causal attention for speech tokens. Each head is responsible for predicting the position of the next sound block in sequence.

[0065] The training objectives can be described as:

[0066] in, s Refers to the length of T The speech tag sequence, x 、 θ refers to the text conditions and model parameters, i is the prediction head index, γ is a hyperparameter with a range of (0, 1]. Assume N is the number of prediction heads, then in this process N ′=min{ N , T - t}. In a series of γ In value setting, γ = 1 performs best in terms of synthesis robustness.

[0067] Speaker-controlled speech tagging vocoder: semantic tagging (such as HuBERT clusters and S 3 Tokenizer tags) do not have a built-in decoder to reconstruct waveforms, so an additional vocoder needs to be trained. In this work, we use CTX-vec2wav to reconstruct semantic tokens into waveforms due to its speaker controllability and simple end-to-end procedure. We use WavLM-Large to extract speaker representations in order to best control speaker identity during inference. S 3 For the Tokenizer, this structure is faster than the conditional stream matching decoder, which is crucial in the RL training process in the following chapters.

[0068] 3.2 Advantages and limitations of block prediction We analyze the rationale of our approach from three key perspectives: (1) We revisit the motivation and role of multi-label prediction and discuss its unique action when applied to speech modalities; (2) We demonstrate the differences between the CAR speech model and the FAR model through an analysis of attention weights; and (3) We identify the limitations of fixed-length CAR and point out the necessity of adopting dynamic block sizes. In particular, we argue that the mismatch between training and inference in autoregressive speech generation hinders the learning of reliable dynamic scheduling policies, motivating our use of reinforcement learning.

[0069] like Figure 4 The figure shows a comparison of the speech synthesis methods used in this application, namely FAR, CAR, and DCAR. FAR uses a frame-by-frame processing approach, while CAR uses a fixed block processing approach. DCAR, on the other hand, uses a dynamic block size determination approach, adaptively determining the optimal block size at each decoding step, thereby improving the quality and efficiency of speech synthesis.

[0070] 3.2.1. Revisiting Block Label Prediction In related technologies, block-wise multi-label prediction is mainly used as an auxiliary task to better utilize training data and accelerate model convergence, while still using single-step prediction for optimal inference. However, when used to accelerate inference, it often leads to a decline in generation quality because cross-label prediction is inherently more challenging. MEDUSA attempts to balance inference efficiency and generation quality by timely verifying additional outputs. A common point of the above studies is that the additional prediction task is challenging and often produces unreliable results. However, due to the strong local continuity of speech signals (even in the form of discrete labels), we believe that the difficulty of predicting additional labels will be reduced. This weakens the dominance of the basic head in prediction, making the prediction effect of the additional head comparable or even better.

[0071] Figure 5a to Figure 5b This is a schematic diagram of the block mark prediction in this application. Figure 5a Schematic diagram of the first layer attention weight of the FAR model; Figure 5b Schematic diagram of the first layer attention weight of the CAR model; Figure 5c Schematic diagram of FAR attention amplification; Figure 5b Schematic diagram of CAR attention amplification.

[0072] 3.2.2 Block-to-frame attention model The chunk prediction model aims to capture the relationship between the upcoming chunk and the previous frame. Compared to a single frame, chunks typically contain more coherent and semantically rich content, which significantly changes the model's attention pattern. By visualizing the CAR model's attention weights for a given passage, we can gain insight into how attention is distributed across different locations during chunk prediction.

[0073] Figure 5a and Figure 5b The distribution of attention weights for the first layer of the FAR and CAR models is shown. Compared to the text-only model, the text-guided speech model exhibits two prominent attention branches: (1) the corresponding text frame; (2) the short-range speech context. Occasionally, the model also pays attention to frames dozens of steps away, which we assume is due to specific characteristics of the speaker or genuinely similar acoustic patterns.

[0074] Figure 5a to Figure 5b The comparison results in

[15] show a clear difference between the two models. The CAR model shows a clearer alignment between speech and text in the earliest layers, as well as a more stable attention to short-range speech context. In contrast, the FAR model shows signs of confusion in these areas. This difference can be reasonably attributed to the fact that block-level representations are inherently more informative and easier to interpret than individual frames.

[0075] In addition, it is worth noting that the attention distribution of short-range speech context is very uneven across different locations. Figure 5c Figure 5d and Figure 5d magnify the local attention structure near the diagonal line, showing that attention within a speech block is often distributed to the first few frames. This pattern is consistent with the nature of speech, where certain key frames in the sequence largely determine the trajectory of the speech content.

[0076] 3.2.3 Attention Preference Between Frames like Figure 6 The following is a diagram of speech synthesis using a block size of 2 as an example. For a given frame, the predictions from different segments may be different. The predictions for the same frame may be different in different blocks. Figure 6 For example, if the block size is 2, p θ ( s t , s t+1 | s <t )right s t The forecast is weaker than p θ ( s t-1 , st | s <t-1 )right s t predictions, we should not simply attribute this to the presence of additional s t-1 , or includes additional prediction targets s t+1 Instead, we view this as evidence that, given limited model capacity, fused blocks ( s t , s t+1 ) of information than the block ( s t-1 , s t ) is more suitable for prediction s t In other words, each frame has an inherent preference for different block-to-frame attention patterns.

[0077] Figure 7 The figure shows the inference of a CAR model with six additional heads added to a ground-truth sequence. We slide the seven prediction heads (one base head and six additional heads) of CAR over a frame interval, so that the predictions within each block are diagonally aligned. Under teacher forcing, we compute the cross-entropy loss at each position and sort it into columns. These losses correspond to the predictions for the same position when the base head is positioned at different preceding positions. Lower losses (brighter colors in the figure) indicate greater confidence in the model's prediction of the ground-truth landmark.

[0078] Figure 7 Ranking of prediction losses for different heads in the dataset (teacher-forcing mode). Brighter colors indicate better prediction performance at each position. Green dots represent dynamic prediction paths, selecting the best predictions whenever possible.

[0079] We note that the most confident predictions for a particular position often do not come from the base head. Experiments on the dev-all subset of LibriTTS show that the base head makes the best prediction in only 60.07% of cases, while producing the worst prediction in 9.27%.

[0080] When considering minimizing the total prediction loss under teacher forcing, the problem becomes a dynamic programming problem. Figure 7 As shown, the heuristic principle is to follow the path through the lighter cells as much as possible, with each diagonal line representing a prediction step.

[0081] However, this principle is only heuristic and difficult to use to train effective scheduling policies. This is partly because it fails to account for the varying importance of different positions. More importantly, the model's actions during actual reasoning differ from those under teacher forcing: during AR generation, deviations from the true sequence may lead to cumulative errors, but may also produce alternative, satisfactory outputs. Despite this, single-step FAR or fixed-length CAR clearly have greater room for optimization.

[0082] This unique property of AR poses challenges for learning policies under teacher forcing. We believe that a reasonable scheduling strategy must consider the entire trajectory across multiple future frames, thus differing from speculative sampling. Therefore, we would like to obtain feedback on the scheduling strategy directly from the generated results, which motivates our use of reinforcement learning. We believe reinforcement learning can address these difficulties, and that feedback directly from audio metrics, unconstrained by real-world data, holds the potential to further improve speech synthesis performance.

[0083] 4. Dynamic block prediction through lightweight strategy 4.1 Introduction GRPO is a variant of PPO that abandons the critic model and uses group-related advantage values instead. It can directly utilize rule-based rewards and is an efficient RL algorithm. During training, the update policy is π θ , the reference strategy is π ref , the old policy is π θold From the collection Q Given a problem q ,use π θold Generate a set G Candidate Output { o 1, o 2,…, o G}. Will use rewards to build advantages , the training objectives are:

[0084] in, ε and β is a hyperparameter, D KL [ π θ ∥ π ref ] can be described as:

[0085] 4.2 Dynamic Block Strategy Optimization We designed a task-specific RL algorithm, Dynamic Chunk-wise Policy Optimization (DCPO), to train a lightweight dynamic chunking scheduling policy as an adaptive modification of GRPO, using some specific training strategies. We selected a small sample subset from the LibriTTS training set, consisting of 980 utterances from 20 speakers, each lasting at least 6 seconds, with the first 3 seconds used as audio cues.

[0086] Training Objective: The scheduling module (used to implement the scheduling policy) is structured as a linear prediction head followed by a causal transformer layer, which takes as input the historical hidden state sequence of the predicted speech sequence. Therefore, our training objective should take into account the parameters of the base model. ϕ :

[0087] Among them, the output set A Includes lightweight policy action sequences, functions s Calculate the hidden state of the TTS basic model, a i Indicates that the model is currently i Specifically, the function s The previously determined policy actions are used to schedule the block decoding steps, and then the sampled tokens are used as input to the TTS base model and the hidden states are obtained accordingly. In the forward pass, we introduce a mask on the inner block position to ensure that the policy only learns from the positions where the block size is actually scheduled in the sampling. In practice, we β Using a warm-up strategy, it is initialized to 0 and gradually increased during the training period. In addition, the advantage is defined as ,in r i For sequence r Elements of species, reward vector r See below for details.

[0088] The chase-then-exceed strategy determines a preset range of block lengths based on CAR speech synthesis. Specifically, to stimulate the robust synthesis capabilities of the random initialization strategy, we set a range C of block-scheduled actions that excel in CAR speech synthesis and set half of the CAR scheduling sample set using a fixed block length to values within this range. This softly constrains the action search space, enabling the strategy to first learn from CAR's experience and then surpass CAR.

[0089] Reward function is implemented through action guidance: During training, a relatively negative process reward is set for actions that exceed the preset range of the block length. For example, we extract batches from a small dataset and decode G times in each group. The generated samples and real speech are transcribed using an efficient automatic speech recognition model, and then their WER indicators are calculated to obtain the result reward. We set a relatively negative process reward for actions that exceed the range as an additional constraint. Total reward r The calculation function is:

[0090] Among them, WER gen and WE Rgt Refers to the WER index of generated speech and real speech, T refers to the length of the generated speech sequence, C For the above range of action, i To avoid reward explosion, the value of ϵ is -10 and λ is usually set to 0.1, so that the absolute value of the total negative reward is smaller than the order of magnitude of the WER supervision reward.

[0091] 5. Experiment 5.1 Experimental Setup Datasets: We conduct primary experiments on LibriTTS's 585-hour training set and perform data expansion experiments using LibriHeavy's 50k-hour training set. For evaluation, we test TTS performance on the UniCATS test set-B. This test set contains 500 utterances from 37 unseen speakers from the LibriTTS clean test subset. Each speaker has approximately 3 seconds of speech prompts. Section 4.2 describes the policy training set.

[0092] Architecture: The basic decoder-only TTS architecture consists of 12 Transformer layers, each with 8 attention heads, 1024 hidden dimensions, and 4096 feedforward dimensions. The base prediction head is a linear layer, and the additional prediction head consists of four residual blocks (a linear layer with SiLU activation) and a linear prediction layer. For fair comparison, we conducted the main experiments with six additional heads. During training, we used NeMo ASR2 to compute the WER.

[0093]

[0094] Table 1: Performance of FAR, CAR, and DCAR in terms of speech synthesis robustness, quality, and speed. “Avg.Token” refers to the average number of speech tokens generated simultaneously at each step.

[0095] Baselines: We select three methods as baselines for DCAR: (1) frame-level AR TTS (using next token prediction); (2) block-wise AR TTS (selecting a fixed number of tokens as the next block); and (3) VADUSA, a speculative decoding strategy that selects draft tokens via verification based on the prediction head, which also performs dynamic effects in the decoding step. All baselines adopt the same architecture except for the presence or absence of an additional head.

[0096] Evaluation Metrics: To comprehensively evaluate speech synthesis performance, we used various metrics: Word Error Rate (WER) and UTMOS indicate synthesis intelligibility and quality; Speaker Encoder Cosine Similarity (SECS) indicates zero-shot TTS speaker similarity; and Real-Time Factor (RTF) indicates synthesis speed. We measured WER using Whisper-Large-V33 and SECS using Resemblyzer4. We also measured the speed improvement achieved by DCAR compared to the baseline.

[0097] Equipment: Our main experiments are conducted on an NVIDIA RTX 4090 24GB GPU, and data scaling training is performed on an NVIDIA A800 80GB GPU.

[0098] 5.2 Performance of Different AR Generation Paradigms In Table 1, we find that the FAR-based generation method performs poorly in terms of robustness for both token types. It also falls short of the best performance in terms of zero-shot speaker similarity (SECS) and naturalness as measured by UTMOS. Furthermore, this method incurs significant inference latency. The table also reports the top three most robust results for CAR-based synthesis methods. We find that simply training a speech synthesis model using the CAR architecture significantly improves robustness. Furthermore, its inference speed scales proportionally with the number of predefined decoding tokens. DCAR achieves the lowest WER in the robustness evaluation. Importantly, it maintains sound quality and zero-shot performance while providing significant acceleration. Notably, when the average number of sampling steps is comparable, DCAR's acceleration is slightly reduced due to the overhead introduced by the policy model.

[0099] 5.3 Comparison between DCAR and speculative decoding Speculative decoding methods like VADUSA aim to accelerate inference, while DCAR leverages WER as the primary supervisory signal to enhance the robustness of speech synthesis. Despite differing goals and underlying mechanisms, both methods employ similar dynamic decoding strategies, ultimately achieving high-quality and efficient speech generation. We ensure architectural consistency across the CAR TTS model and demonstrate the generation quality and speed of both methods using shared HuBERT markers in Table 3. VADUSA recommends applying a tolerance to multiple sampling results of the base header; here, we set the tolerance to 2 and 3.

[0100] Table 2: WER strategy decomposition results

[0101] Table 3: Comparison with VADUSA, τ refers to the tolerance value

[0102] Table 4: Impact of DCPO action guidance range on DCAR synthesis robustness and inference speed.

[0103]

[0104] 5.4. Discussion on Frame Rate Factors We compared and integrated the CAR framework with low frame rate tokenization methods, specifically leveraging the 50Hz and 25Hz variants of the S3-Tokenizer. Given that the temporal receptive field covered by 25Hz tokens is twice that of 50Hz tokens, we adapted the CAR model to reduce the number of prediction heads to half the number used in the 50Hz configuration.

[0105] like Figure 8 The following is a schematic diagram of WER performance under different frame rates and strategies. In order to evaluate the effect of CAR, we Figure 8 A visualization is provided in Figure 1, where the x-axis represents the time dimension. As shown, in the 25Hz setting, CAR generation performance degrades as the block size increases. Despite this, more than half of these configurations outperform the FAR baseline, while providing up to a 2x theoretical speedup in synthesis. On the other hand, when controlling for the temporal field of view, CAR synthesis using 50Hz markers consistently outperforms 25Hz CAR. Furthermore, we report the performance of the 50Hz DCAR model, which achieves the highest robustness across all evaluation settings.

[0106] 5.5 Ablation Study Variation of the policy action guidance range: Table 3 shows the impact of changing the action guidance range of the DCPO algorithm. A wider range of CAR block sizes shows that the policy’s decoding ability is faster, but its robust synthesis ability is reduced.

[0107] Method Decomposition: We decompose the DCPO strategy, the catch-up-surpass strategy, and the outrange negative rewards strategy to reflect the effectiveness of the designed strategies, as shown in Table 2. The “random” strategy in the table directly means that the block size of each decoding step is randomly selected from the range followed.

[0108] 6. Conclusion In summary, this work addresses a key challenge in speech synthesis: the limitations of frame-level autoregressive (FAR) models, which often suffer from instability and high inference latency when processing fine-grained speech tokens. To overcome these issues, we analyze the effectiveness of block-based autoregressive (CAR) synthesis as an alternative architecture. By leveraging its architectural properties, CAR demonstrates significant advantages over FAR in both robustness and efficiency. Building on this, we introduce DCAR, a dynamically scheduled variant of CAR that employs reinforcement learning to train a lightweight policy network. This policy adaptively determines the optimal block size at each decoding step, further improving synthesis quality and efficiency. The appendix in the supplementary material further discusses its limitations and broader implications.

[0109] B: DCPO algorithm details B.1: Action Guide Configuration File We leverage the NeMo ASR model for fast configuration and select an action guidance set C from the CAR block sizes that achieve top-k performance on the WER metric (where k is a hyperparameter representing the set size, i.e., |C|).

[0110] B.2: Description To explain the DCPO (Dynamic Chunk-wise Policy Optimization) algorithm in detail, we introduce the process of the algorithm in Algorithm 1.

[0111] C: Training details C.1: CAR TTS Configuration We use 12 causal transformer layers for TTS decoding, each layer consists of 1024 hidden dimensions and 4096 feedforward dimensions. For HuBERT tokens, the model has 512 text token embeddings and 2049 speech token embeddings; for S 3The model was trained on 4097 speech token embeddings using the Tokenizer (codebook + [EOS] tokens). The model was trained on 8 NVIDIA RTX4090 GPUs for 20 epochs. The LibriTTS training subset was filtered to have durations between 3 and 20 seconds. We used a batch size of 2 per GPU and gradient accumulation of 2 steps (effective batch size = 32). Training was performed using the ScaledAdam optimizer (initial gradient = 0.01, β = (0.9, 0.95), gradient clipping scale = 2.0) and the Eden scheduler (200 warmup steps).

[0112] We augment the decoder with additional heads for parallel CAR prediction branches. Each branch consists of four stacked residual blocks followed by an unbiased linear projection of the output space to the number of speech tokens. Each residual block applies a 1024×1024 linear transformation and a SiLU activation. The cross-entropy loss is the critical factor for all heads. The vocoder is trained for 1 million steps with hyperparameters following CTX-vec2wav.

[0113]

[0114] C.2: DCPO training configuration The DCPO lightweight policy is a causal transformer layer with a hidden size of 1024, followed by a 1024×C linear layer that predicts C actions. We use the Adam optimizer with a learning rate of 2e-6 and a step size of ε of 0.2. For the weights β of the KL loss, we set it so that policy training takes approximately 8 hours on a single NVIDIA RTX4090 GPU.

[0115] D: Evaluation Metrics WER: Word Error Rate (WER) is a standard metric for evaluating the intelligibility of generated speech. It is defined as:

[0116] Where S, D, and I represent the number of replaced, deleted, and inserted words, respectively, and N represents the total number of words in the reference text. A lower WER indicates higher intelligibility.

[0117] SECS: Speaker Embedding Cosine Similarity (SCES) measures the degree of match between synthesized speech and the target speaker's speech. It is calculated by encoding the generated and reference speech using a fixed speaker embedding model and then calculating the cosine similarity between their embeddings. Values closer to 1 indicate higher speaker similarity.

[0118] UTMOS: UTMOS (UTokyo-SaruLab MOS) is an objective mean opinion score predictor that uses a set of trained neural models to estimate human listening scores on a scale of 1-5. It provides a reliable and reproducible alternative to subjective audio quality assessment without the need for labor-intensive listening tests. We report the standard error of UTMOS along with the mean score to ensure the robustness of the evaluation.

[0119] RTF: Real-Time Factor (RTF) measures the synthesis speed by dividing the total inference time by the duration of the generated audio.

[0120]

[0121] Table 5: FAR, CAR, DCAR performance achieved on CosyVoice. “Average tokens” refers to the average number of speech tokens generated simultaneously in each step.

[0122] E: Extended Experiment E.1: TTS Performance Overview This section summarizes the main evaluation findings. Figure 9 The figure shows the evaluation results of word error rate when HuBERT tagging is used in this application. Figure 10 The figure shows the evaluation results of word error rate when using S3 tokenizer in this application. Figure 11 Shown is a schematic diagram of the evaluation results of the synthetic performance of HuBERT labeling on the LibriHeavy dataset in this application.

[0123] E.2: Scalability The experimental results demonstrate the scalability of our approach. Specifically, we trained the CAR TTS model for 2 rounds using HuBERT tagging on the large LibriHeavy dataset (50k hours). Figure 11 As shown, the robustness and speed of the synthesis process are highlighted. In this experiment, the action guidance range of DCPO is set to [3, 4, 5, 6].

[0124] E.3 Implementation on CosyVoice To validate the versatility of our approach, we implemented DCAR on the open-source TTS framework CosyVoice. Specifically, we added six additional voice heads to CosyVoice's base model, trained it on LibriTTS for 20 epochs, and trained DCPO for 5 epochs using a small set of 980 samples. The zero-shot TTS results on the test set shown in Table 5 demonstrate the effectiveness of our approach. CosyVoice is an open-source multilingual large-scale speech generation model released by Alibaba Tongyi Lab, providing full-stack capabilities for inference, training, and deployment.

[0125] Figure 12 FIG. 1 is a schematic diagram of the hardware structure of an electronic device for executing a speech synthesis method according to another embodiment of the present invention. Figure 12 As shown, the device includes: One or more processors 1210 and memory 1220, Figure 12 A processor 1210 is taken as an example.

[0126] The apparatus for executing the speech synthesis method may further include: an input device 1230 and an output device 1240 .

[0127] The processor 1210, the memory 1220, the input device 1230 and the output device 1240 may be connected via a bus or other means. Figure 12 The bus connection is taken as an example.

[0128] Memory 1220, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the speech synthesis method in the embodiments of the present application. Processor 1210 executes the non-volatile software programs, instructions, and modules stored in memory 1220 to execute various server functional applications and data processing, thereby implementing the speech synthesis method in the above-mentioned method embodiment.

[0129] The memory 1220 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the speech synthesis device, etc. Furthermore, the memory 1220 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 1220 may optionally include a memory remotely located relative to the processor 1210, and such remote memory may be connected to the speech synthesis device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0130] The input device 1230 may receive input digital or character information and generate signals related to user settings and function control of the speech synthesis device. The output device 1240 may include a display device such as a display screen.

[0131] The one or more modules are stored in the memory 1220 and, when executed by the one or more processors 1210 , perform the speech synthesis method in any of the above method embodiments.

[0132] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0133] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a general hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A speech synthesis method, comprising: Pre-train the TTS basic model; Training a dynamic block scheduling strategy based on the TTS basic model; Dynamic block speech marking is performed according to the dynamic block scheduling strategy, and speech synthesis is performed according to the result of the dynamic block speech marking.

2. The method according to claim 1, characterized in that The TTS basic model includes 1 basic head and n-1 additional heads, wherein the basic head and the additional heads are used to predict the next n tokens; the pre-trained TTS basic model includes: The basic head is used to predict the next token; The additional head is used to predict the second to n subsequent tokens; The TTS basic model is obtained based on the loss training of the basic head and the additional head.

3. The method according to claim 1, characterized in that The dynamic block scheduling strategy is executed by a scheduling module, which includes a linear layer and a causal transformation layer connected after the linear layer, and the causal transformation layer is used to take the historical hidden state sequence of the predicted speech sequence as input.

4. The method according to claim 3, characterized in that The training objective of the dynamic block scheduling strategy is set to: Among them, the advantages is defined as 5. The method according to claim 4, characterized in that: The dynamic block scheduling strategy is trained according to the TTS basic model, including: Determine a preset range of block length based on CAR speech synthesis; A relatively negative process reward is set for actions that exceed the preset range of the block length. The total reward is expressed as follows:

6. The method according to claim 2, characterized in that The basic header includes a linear layer, and the additional header includes 4 residual blocks and 1 linear prediction layer.

7. A speech synthesis device comprising: The first training module is used to pre-train the TTS basic model; A second training module is used to train a dynamic block scheduling strategy based on the TTS basic model; The speech synthesis module is used for dynamically marking speech in blocks according to the dynamic block scheduling strategy, and performing speech synthesis according to the result of the dynamic block speech marking.

8. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.