A multi-layer prosody sentence breaking and tone intonation decoupling control method for Chinese TTS
By using multi-level prosodic boundary prediction and tone baseline model decoupling control, the problem of coupling between lexical tone and sentence mood in Chinese TTS is solved. This achieves multi-level prosodic phrasing and tone modulation coupling in Chinese text, improving the naturalness and stability of synthesized speech. It is particularly suitable for long sentence broadcasting and live streaming scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-04
AI Technical Summary
Existing technologies in Chinese TTS have failed to effectively address the coupling problem between lexical tone changes and sentence mood changes sharing the F0 channel, resulting in unstable sentence-end mood control, which affects the rhythmic naturalness and structural stability of synthesized speech, especially in long sentence continuous speech scenarios.
By using a multi-level prosodic boundary prediction model to predict the prosodic boundary position and pause duration of Chinese text, and combining the tone baseline model and intonation residual components, stable modeling of different levels of boundaries and intonation coupling control are achieved, generating a multi-level prosodic planning sequence.
It improves the naturalness of sentence breaks, the stability of pauses and connections, and the consistency of sentence-end tone in scenarios involving long Chinese sentences, making it suitable for scenarios such as long Chinese sentences, live-streaming sales, and continuous broadcasting.
Smart Images

Figure CN122511230A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text-to-speech synthesis technology, specifically to a multi-layer prosodic phrasing and tone-level modulation coupling control method for Chinese TTS. Background Technology
[0002] In the field of text-to-speech (TTS) technology, TTS technology is showing a trend towards controllable generation and attribute decoupling modeling. Prosodic segmentation modeling has become an independent and important research branch in controllable TTS. ProsodyFM explicitly treats phrasing and intonation as core issues, introduces a Phrase Break Encoder to capture the initial phrasing position, uses a Duration Predictor to adjust the phrasing duration, and models the terminal intonation through a Terminal Intonation Encoder and intonation shape tokens.
[0003] In Chinese speech, the tone at the end of a sentence typically contains complex information. Language prosody includes variations in pitch and loudness, as well as changes in the length of syllables, words, and phrases. In Chinese TTS scenarios, when language prosody affects the final words of Chinese sentences, tone changes are coupled with mood changes. While existing technologies can control prosody or terminal intonation, their overall approach tends to model general prosodic patterns or absolute pitch morphology. However, directly learning and controlling the absolute F0 (Fundamental Frequency) or sentence-end intonation patterns in Chinese TTS can easily lead to the misintegration of local variations belonging to the tone category into the sentence intonation representation, or the reverse contamination of sentence-end mood changes into the implementation of lexical tone. This results in unstable sentence-end mood control and insufficient consistency of mood between different sentences. This problem is more pronounced in long, continuous spoken scenarios. For applications requiring stable pauses and connections, such as long sentences and spoken sentences, this deficiency directly affects the rhythmic naturalness and structural stability of the synthesized speech.
[0004] Existing technologies have not yet specifically addressed the coupling problem caused by the sharing of the F0 channel between "lexical tone changes" and "sentence intonation changes" in Mandarin. Summary of the Invention
[0005] In view of this, this application discloses a method to address the problems of existing technologies failing to uniformly predict and coordinately control the three-layer prosodic boundaries of Chinese PW (prosodic words), PPH (prosodic phrases), and IPH (intonation phrases), making it difficult to stably model the pause duration at different division levels, and easily confusing lexical tone changes with sentence mood changes when directly modeling intonation based on absolute pitch. This method achieves stable control of multi-layer prosodic organization and sentence-final intonation for spoken Chinese texts, including:
[0006] S1. Obtain the Chinese text to be synthesized and perform text normalization processing; extract features related to prosody modeling from the normalized text and establish a text feature sequence;
[0007] S2. Based on a pre-trained multi-layer prosodic boundary prediction model, boundary position prediction is performed on the text feature sequence to obtain boundary prediction sequences at different levels; the boundary position prediction includes: PW boundary prediction, PPH boundary prediction, and IPH boundary prediction.
[0008] S3. Perform hierarchical consistency correction on the boundary prediction sequence to obtain a multi-level prosodic boundary sequence;
[0009] S4. Based on the multi-layer prosodic boundary sequence, the pause duration of each layer prediction sequence is predicted to obtain the layered pause duration of each layer.
[0010] S5. Use the tone baseline model to generate the basic pitch plan corresponding to the Chinese text to be synthesized, and use the basic pitch plan as the tone baseline.
[0011] S6. Based on the tone baseline, calculate the intonation residual components corresponding to the sentence-level intonation changes;
[0012] S7. Generate a corresponding residual intonation control marker at the end of each IPH boundary to indicate the trend at the end of the current intonation phrase; the trend includes: rising, falling, continuation, convergence, and flattening.
[0013] S8. Based on the multi-level prosodic boundary sequence, the layered pause duration of each level, the tone baseline, and the residual intonation control markers, generate the prosodic planning sequence of the Chinese text to be synthesized, and complete the multi-level prosodic punctuation and intonation coupling control.
[0014] The beneficial effects of this application include:
[0015] By jointly predicting the three prosodic boundaries of Chinese text (PW, PPH, and IPH) and introducing hierarchical consistency constraints, a unified modeling of multi-layered stop-connection structures is achieved.
[0016] By separating the boundary position prediction from the pause duration prediction, the prosodic planning information is input into the text to the speech synthesis backend. The front-end prosodic planning and the back-end speech synthesis work in a relatively separate manner, thereby improving the adaptability to different types of text to the speech synthesis backend and realizing differentiated control of pause intensity at different levels.
[0017] By establishing a tone baseline and modeling intonation residuals, the separation and control of Chinese lexical tone and sentence intonation are realized, thereby improving the naturalness of sentence breaks, the stability of pauses and connections, and the consistency of sentence-end tone control in the context of long Chinese sentence broadcasting. By only marking or parameterizing the intonation residuals at this position, fine-grained adjustment of intonation patterns such as rising, falling, continuation, and termination at the end of sentences is achieved.
[0018] The method designed in this application is particularly suitable for scenarios such as long Chinese sentence broadcasting, live-streaming sales, and continuous broadcasting. It can improve the naturalness of sentence segmentation, the stability of pauses, and the consistency of sentence-end tone of synthesized speech without changing the text content, providing a design idea for those skilled in the art. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the multi-layer prosodic phrasing and tone-speech coupling control method for Chinese TTS in the embodiments of this application;
[0020] Figure 2 This is a schematic diagram of the prosody planning model in the embodiments of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, features, and advantages of this application clearer and to facilitate a better understanding of the technical solutions of this application by those skilled in the art, the following detailed description of this application is provided in conjunction with the accompanying drawings and embodiments.
[0022] Example 1:
[0023] This embodiment includes a multi-layer prosodic phrasing and tone-level decoupling control method for Chinese TTS. Before text-to-speech synthesis, prosodic planning information corresponding to the text to be synthesized is generated, and the target speech is output from the speech synthesis backend based on the prosodic planning information. Figure 1 As shown, it includes:
[0024] S1. Obtain the Chinese text to be synthesized and perform text normalization processing; extract features related to prosody modeling from the normalized text and establish a text feature sequence.
[0025] Text feature sequences typically include: character-level or word-level sequences, pinyin or syllable sequences, audio information, tone categories, punctuation categories, and modal particle markers.
[0026] S2. Based on a pre-trained multi-layer prosodic boundary prediction model, boundary position prediction is performed on the text feature sequence to obtain boundary prediction sequences at different levels; the boundary position prediction includes: PW boundary prediction, PPH boundary prediction, and IPH boundary prediction.
[0027] In the process of boundary location prediction, candidate boundary locations can be set between adjacent characters, adjacent words, or adjacent syllables.
[0028] S3. Perform hierarchical consistency correction on the boundary prediction sequence to obtain a multi-level prosodic boundary sequence.
[0029] Given that existing technologies rely on punctuation information or single-layer break / no-break methods for sentence segmentation prediction, they cannot distinguish between fine-grained connected speech positions, local meaning group pauses, and sentence-end or strong pauses. If PW, PPH, and IPH are predicted independently without rule constraints, problems such as inconsistent hierarchical boundaries, overly dense local boundaries, or lack of low-level support for high-level boundaries may occur.
[0030] To reduce omissions, misinterpretations, and hierarchical confusion in long Chinese sentences and spoken sentences, this application performs hierarchical consistency correction on the boundary prediction sequence, correcting, downgrading, or eliminating prediction results that do not meet the preset criteria; correcting, filtering, or reassigning prediction results that conflict between boundaries of different granularities; and comprehensively judging the distribution of smaller granular boundaries, boundary prediction probabilities, punctuation types, and adjacent segment lengths when determining larger granular boundaries. The preset prediction results include: IPH segments containing one or more PPH segments, and PPH segments containing one or more PW segments. Specifically, this includes:
[0031] First, for each candidate boundary location The boundary prediction model outputs three-layer cumulative boundary probabilities:
[0032]
[0033]
[0034]
[0035] This application first addresses the three-layer boundary probability. Perform monotonic projection to obtain the correction probability. At the same time, it meets certain constraints:
[0036]
[0037]
[0038] This constraint allows a location to be considered both an IPH and PW boundary if it is determined to be an IPH boundary; and a location to be considered both a PPH boundary and a PW boundary if it is determined to be a PPH boundary.
[0039] After obtaining the corrected probability that satisfies the monotonic relationship, it is transformed into the mutually exclusive boundary label probability: the no-boundary probability is... The boundary probability of PW is The PPH boundary probability is The IPH boundary probability is Then Decoding is performed under set constraints to obtain the final multi-layer prosodic boundary sequence.
[0040] Decoding is achieved through dynamic programming. This involves a comprehensive judgment based on information such as the distribution of smaller-granular boundaries, boundary prediction probabilities, punctuation types, and the lengths of adjacent segments. The objective function incorporates a boundary label probability term, a segment length penalty term, a boundary density penalty term, and a punctuation reward term. The segment length penalty term avoids excessively short or long PW, PPH, and IPH segments; the boundary density penalty term avoids overly dense boundaries within short windows; and the punctuation reward term increases the probability that strong punctuation positions such as commas, periods, question marks, and semicolons are decoded as PPH or IPH boundaries.
[0041] S4. Based on the multi-layer prosodic boundary sequence, the pause duration of each layer prediction sequence is predicted to obtain the layered pause duration.
[0042] Compared to schemes that only determine "whether a sentence is broken" without separately modeling the pause intensity, this application separates the boundary position prediction and pause duration prediction, and predicts the pause duration for different levels of boundaries. Therefore, it can achieve differentiated control of pause relationships at different granularities in the same text, making the rhythm organization of synthesized speech more stable.
[0043] In this embodiment, the pause duration prediction process is based on the following principles: PW boundaries correspond to zero or very short pauses, PPH boundaries correspond to short pauses, and IPH boundaries correspond to medium-long pauses or significant pauses between sentences. By separating pause duration prediction from boundary position prediction, the existence of boundaries and the intensity of boundary pauses are modeled separately. This addresses the following problems encountered in traditional solutions: Traditional methods typically use a "punctuation rule, single-layer break / no-break classification + fixed pause duration mapping" approach to handle sentence breaks. That is, once a position is determined to be a sentence break, a fixed pause duration is assigned according to the classification rules. This approach couples boundary position judgment with pause intensity control, making it difficult to distinguish pause differences between different levels of boundaries (PW, PPH, and IPH). This can easily lead to problems such as excessively heavy pauses at minor boundaries, insufficient pauses in local semantic groups, indistinct strong pauses at the end of sentences, or excessively dense pauses in long sentence broadcasts. Therefore, this application separates the boundary position prediction from the pause duration prediction. After obtaining the multi-level prosodic boundaries, the pause duration is predicted separately for PW, PPH and IPH, so as to achieve independent control of the pause intensity at different levels, which can alleviate the above-mentioned contradiction problem.
[0044] S5. Use the tone baseline model to generate the basic pitch plan corresponding to the Chinese text to be synthesized, and use the basic pitch plan as the tone baseline.
[0045] The basic pitch planning is used to characterize the pitch change trend dominated by word tone. It is generated based on audio information in the text feature sequence. The audio information specifically includes one or more of the following: pinyin, tone category, syllable position, context of adjacent syllables, and tone sandhi information.
[0046] The tone baseline model is used to encode the pinyin, tone category, syllable position, adjacent syllable context, and tone sandhi information in the text feature sequence to obtain the basic pitch curve of each syllable on the normalized time axis. The basic pitch curves of each syllable are then expanded into frame-level basic pitch planning based on the predicted or aligned syllable duration. This embodiment uses a tone template-constrained context sequence regression model to implement the tone baseline model. The specific operations are as follows:
[0047] Let the syllable sequence corresponding to the Chinese text to be synthesized be:
[0048]
[0049] Among them, the syllable The pinyin, tone category, tone sandhi mark, syllable position, and contextual features are represented as follows:
[0050]
[0051] in, Represents an embedding map. Indicates the category of Pinyin. Indicates tone category, Indicates pitch change information. Indicates the position of a syllable in a word, phrase, or sentence. Representing the contextual features of adjacent syllables. Inputting the syllable feature sequence into the context encoder yields the context-dependent tone representation:
[0052]
[0053] The context encoder can be implemented using a bidirectional recurrent neural network, a convolutional sequence network, or a Transformer encoder, which will not be elaborated here.
[0054] For the j-th syllable, let its start and end times be respectively... and Then the normalized time position of any frame t within this syllable is:
[0055]
[0056] In this embodiment, the tone baseline is determined jointly by the "tone category template" and the "context correction item". Let... This is the base tone template corresponding to the tone category of the j-th syllable. Using a pre-defined time basis function, the model represents tone based on context. Prediction correction coefficient :
[0057]
[0058] The sound baseline can then be represented as:
[0059]
[0060] in, This is the base tone template for the corresponding tone category. For time basis functions, To represent based on context The predicted correction coefficient, and This is the pitch normalization parameter. The result is... It primarily characterizes the basic pitch variation trend dominated by lexical tone and serves as a benchmark for subsequent intonation residual component calculation.
[0061] S6. Based on the tone baseline, calculate the intonation residual components corresponding to the tone changes at the sentence level.
[0062] The intonation residual component calculated based on the intonation baseline does not directly represent the absolute pitch value, but should represent additional variation information relative to the intonation baseline. Specifically, the intonation residual component is determined based on the difference between the intonation baseline and the frame-level pitch trajectory, and is used to represent the pitch shift caused by sentence-level mood, intonation phrase position, sentence-ending punctuation, sentence structure information, modal particle information, and contextual semantic structure, while maintaining basic lexical intonation stability.
[0063] Typically, when calculating intonation residuals, one or more of the following are considered: IPH boundary location, sentence-ending punctuation type, sentence structure information, modal particle information, and contextual features. The specific calculation method for intonation residuals is as follows:
[0064] During the training phase, let the logarithmic fundamental frequency trajectory of the reference speech after interpolation, smoothing, and normalization be:
[0065]
[0066] in, Indicates the first The fundamental frequency value of the frame. This represents the corresponding logarithmic pitch value. Let the frame-level pitch baseline output by the pitch baseline model be... Then the monitoring objective for intonation residuals can be expressed as:
[0067]
[0068] in, This represents the intonation residual target calculated during the training phase. Indicates the audio frame mask; when the first... When the frame is a frame with sound, ,otherwise Therefore, intonation residuals are not a direct result of learning absolute pitch. Instead, it learns the reference pitch relative to the tone baseline. The offset is used to reduce the interference of lexical tone changes on sentence mood modeling. To incorporate IPH boundary positions, let the set of IPH boundaries corresponding to the text be... ,in, Indicates the first Intonation phrase boundaries. Based on IPH boundaries, the frame-level sequence can be divided into several intonation phrase segments. .in, This represents the set of frames corresponding to the k-th IPH segment. Indicates the first The frame position corresponding to each IPH boundary. For a fragment For any frame t within the current IPH segment, its normalized position is defined as:
[0069]
[0070] in, This is used to indicate whether the current frame is at the beginning, middle, or end of an intonation phrase. Since sentence-level intonation changes are typically more pronounced at the end of an IPH, this embodiment further defines an IPH end window. :
[0071]
[0072] in, The preset ratio threshold is used to determine the time range within which the IPH end participates in intonation residual control.
[0073] Sentence-ending punctuation type, sentence structure information, modal particle information, and contextual features are used in the intonation residual calculation through feature vectors. Specifically, let the intonation contextual features of the k-th IPH segment be:
[0074]
[0075] in, Indicates the punctuation type corresponding to the end of the k-th IPH. Indicates sentence structure information, Indicates information from modal particles, This represents the contextual representation of the text context encoder at frame t or the corresponding text position. This indicates a pooling operation. This indicates the current IPH segment's position within the sentence. Using the above method, the IPH boundary position is used to determine the segment and end window affected by the intonation residual; punctuation type, sentence structure information, and modal particle information are used to indicate the intonation trend; and contextual features are used to adjust the residual amplitude and shape.
[0076] During the inference phase, the residual intonation model predicts the intonation residual based on the contextual representation of the current frame, the intonation contextual features of the current IPH segment, and the normalized position within the segment: ,in Representing the residual intonation model, This represents the predicted intonation residual component. To ensure that the residual intonation primarily affects the IPH terminology, this embodiment can also introduce positional weights: ,in, The base weights representing non-IPH end positions, This indicates the end-to-end augmentation weight of IPH. This is an indicator function. When t is located in the IPH end window... When t is within the range, the control weight of intonation residual increases; when t is located at a non-IPH end position, the control weight of intonation residual is small or zero.
[0077] Therefore, the target pitch planning used for synthesis can be expressed as:
[0078]
[0079] in, Indicates the tone baseline, The expression represents the intonation residual component after position weight adjustment. As can be seen from the construction method of the above formula, this application does not directly modify the absolute pitch, but rather superimposes the residual changes related to the tone of the sentence layer on the tone baseline.
[0080] During model training, the residual intonation model can be constrained using the following objective function:
[0081]
[0082] The first term is used to constrain the predicted residuals. Approaching the residual target calculated during the training phase The second term is used to constrain the smoothness of the intonation residual curve. Represents the first-order difference. Here, the weights are for the smoothing term. The first-order difference can be expressed as: .
[0083] In this way, the IPH boundary position, sentence-ending punctuation type, sentence structure information, modal particle information, and contextual features are all explicitly incorporated into the calculation process of the intonation residual component, so that the intonation residual mainly represents the intonation change at the sentence level, rather than directly destroying the lexical intonation trend represented by the intonation baseline.
[0084] S7. Generate a corresponding residual intonation control marker at the end of each IPH boundary to indicate the trend at the end of the current intonation phrase; the trend includes: rising, falling, continuation, convergence, and flattening.
[0085] Residual intonation control markers are generated only at the end of the IPH boundary. No residual intonation control is applied at non-IPH end positions, or a control weight or amplitude lower than that at the end of the IPH is applied. Sentence-end tone control is not directly applied to the absolute F0. By pre-establishing a tone baseline dominated by lexical tone, and then modeling the intonation residual components that exceed the baseline, this design does not directly rewrite the tone baseline. As a result, the residual intonation control markers can achieve sentence-end tone control while maintaining the basic stability of lexical tone when acting on intonation residual components.
[0086] S8. Based on the multi-layer prosodic boundary sequence, the layered pause duration of each layer, the tone baseline, and the residual intonation control marker, generate the prosodic planning sequence of the Chinese text to be synthesized, and complete the multi-layer prosodic punctuation and intonation coupling control.
[0087] The prosodic planning sequence includes: the boundary positions of three layers, PW, PPH, and IPH, the pause duration corresponding to each layer boundary, and the residual intonation control information corresponding to the end of each IPH.
[0088] The prosodic planning sequence superimposes the tone baseline with the intonation residual component adjusted by position weights, and sets the position weight at the end of the IPH position to be greater than that at the end of the non-IPH position, so that the residual intonation control is concentrated at the end of the IPH, thereby achieving fine-grained control of the sentence-end intonation while maintaining the stability of Chinese vocabulary intonation as much as possible.
[0089] Furthermore, prosodic planning information is input into the text to the speech synthesis backend, enabling the backend to control the speech synthesis output and generate the target speech in a decoupled manner from the frontend and backend, based on the prosodic planning sequence control information. The speech synthesis backend can employ any neural network speech synthesis model and is not limited by a specific acoustic model structure.
[0090] In some implementations, in order to reduce the impact of pitch tracking error on intonation residual modeling, the pitch trajectory of the reference speech can be smoothed or normalized during the training phase and then used for parameter learning of the pitch baseline and intonation residual components; during the synthesis phase, the pitch baseline and intonation residual components are predicted separately and combined to form the target pitch plan.
[0091] Example 2:
[0092] This embodiment includes a multi-layer prosodic phrasing and intonation coupling control method for Chinese TTS. The difference from Embodiment 1 is that the multi-layer prosodic phrasing and intonation coupling control method in this embodiment is based on... Figure 2 The prosody planning model shown is implemented. Figure 2 In the text structure representation, the text is fed into a multi-layer prosodic segmentation branch and a tone-speech coupling branch. The multi-layer prosodic segmentation branch includes joint prediction of three-layer prosodic boundaries, hierarchical consistency constraint decoding, and hierarchical pause duration prediction. The tone-speech coupling branch includes tone baseline modeling, intonation residual generation, and IPH end residual intonation control. The outputs of the two branches together form prosodic planning information, which is used to drive the TTS backend for speech synthesis, thereby realizing the collaborative modeling of Chinese sentence segmentation organization and sentence-end intonation control.
[0093] The following example, using a specific Chinese spoken text, illustrates the method designed in this application:
[0094] First, obtain the Chinese text to be synthesized: "Today I'd like to recommend a sweet and sticky corn from Northeast China. It's soft, sticky, and sweet when steamed. Order now for a better deal." Then, perform text normalization processing on the above text to obtain a standardized text sequence.
[0095] Text features such as character-level sequences, word-level sequences, pinyin or syllable sequences, tone categories, punctuation categories, and modal particle markers are extracted from the text sequence to form a structured representation of the text for subsequent prosodic modeling.
[0096] Based on the structured representation of text, the system performs joint prediction of prosodic boundaries using three layers: PW, PPH, and IPH. Specifically, the system can output the following set of possible multi-layer boundary results:
[0097] The PW level can be divided into: "Today / I recommend / a / sweet glutinous corn / from Northeast China / steamed / soft, glutinous and sweet / order now / for a better deal";
[0098] The PPH level can be divided into: "Today I'm recommending a sweet and sticky corn from Northeast China / It's soft, sticky, and sweet when steamed / Order now for a better deal";
[0099] The IPH (Illustrated Product) level can be divided into: "Today I'm recommending a sweet and sticky corn from Northeast China. It's soft, sticky, and sweet when steamed / Order now for a better deal."
[0100] A hierarchical consistency correction is performed on the three-layer boundary results to ensure that each IPH contains at least one or more PPHs, and each PPH contains at least one or more PWs. This eliminates conflicting boundary results between different levels, resulting in a final multi-layer prosodic boundary result that satisfies the nesting relationship.
[0101] After obtaining the final multi-layer prosodic boundary results, the pause duration of each layer boundary is predicted. The PW layer boundary corresponds to zero or very short pauses, the PPH layer boundary corresponds to short pauses, and the IPH layer boundary corresponds to relatively strong pauses. This makes the text form a local semantic group pause after "sweet corn" and a relatively strong pause after "soft and sweet", so that the subsequent "order now for a better deal" is output as a new intonation phrase.
[0102] A tone baseline is established based on the pinyin, syllables, and tone categories in the text. Taking "sweet and sticky corn" as an example, its corresponding syllables and tone categories can be represented as tian2, nuo4, yu4, mi3. The system generates the basic pitch trend dominated by the tone of the word according to the second tone, fourth tone, and third tone, respectively, to form the tone baseline representation of the word phrase.
[0103] Based on the tone baseline, the system extracts or predicts sentence-level tone variation information beyond the lexical tone baseline, obtaining intonation residual components. At the end of the second IPH (Intonation Per Hour), at the position "cost-effective", the system generates residual intonation control markers corresponding to the sentence-ending intonation, so that the sentence-ending intonation is mainly reflected in the intonation residuals, rather than directly rewriting the basic pitch trend determined by the lexical tone.
[0104] The final multi-layer prosodic boundary results, the pause duration results corresponding to each layer boundary, the tone baseline, and the residual intonation control marker at the end of IPH are combined to form the prosodic planning information corresponding to the text to be synthesized, including: the boundary positions of the three layers PW, PPH, and IPH, the pause information of each layer boundary, and the residual intonation control information at the end of IPH.
[0105] Finally, the prosodic planning information is input into the text and then into the speech synthesis backend for speech synthesis. Through this embodiment, the output speech can show a weak pause after "sweet and sticky corn", a strong pause after "soft, sticky and sweet", and a closing tone at the end of the sentence "more cost-effective", thus making the whole sentence more in line with the prosodic organization and tone expression requirements of Chinese spoken broadcasting scenarios.
[0106] Compared with existing technologies, this application achieves unified modeling of multi-layered pause structures by jointly predicting the three-layer prosodic boundaries (PW, PPH, and IPH) of Chinese text and introducing hierarchical consistency constraints. By separating the prediction of boundary positions from the prediction of pause duration, the prosodic planning information is input into the speech synthesis backend, working in a way that separates the front-end prosodic planning from the back-end speech synthesis, thereby improving the adaptability to different types of texts to the speech synthesis backend and achieving differentiated control of pause intensity at different levels. By establishing a tone baseline and modeling the intonation residual, the application achieves separate control of Chinese lexical tone and sentence intonation, thereby improving the naturalness of sentence breaks, the stability of pauses, and the consistency of sentence-end tone control in long Chinese sentence broadcasting scenarios. By only marking or parameterizing the intonation residual at this position, fine-grained adjustment of intonation patterns such as sentence-end rising, falling, continuation, and termination is achieved. The method designed in this application is particularly suitable for scenarios such as long Chinese sentence broadcasting, live-streaming sales, and continuous broadcasting. It can improve the naturalness of sentence segmentation, the stability of pauses, and the consistency of sentence-end tone of synthesized speech without changing the text content.
[0107] Finally, it should be noted that the above description only depicts some embodiments of this application. For those skilled in the art, various changes, modifications, substitutions, and variations can be conceived of these embodiments without departing from the principles and spirit of this application. The scope of protection of this application is defined by the appended claims and their equivalents, and all the above-mentioned behaviors should be covered within the scope of protection of this application.
Claims
1. A multi-layered prosodic phrasing and tone-based speech coupling control method for Chinese TTS, characterized in that, include: S1. Obtain the Chinese text to be synthesized and perform text normalization processing; extract features related to prosody modeling from the normalized text and establish a text feature sequence; S2. Based on a pre-trained multi-layer prosodic boundary prediction model, boundary position prediction is performed on the text feature sequence to obtain boundary prediction sequences at different levels; the boundary position prediction includes: PW boundary prediction, PPH boundary prediction, and IPH boundary prediction. S3. Perform hierarchical consistency correction on the boundary prediction sequence to obtain a multi-level prosodic boundary sequence; S4. Based on the multi-layer prosodic boundary sequence, the pause duration of each layer prediction sequence is predicted to obtain the layered pause duration of each layer. S5. Use the tone baseline model to generate the basic pitch plan corresponding to the Chinese text to be synthesized, and use the basic pitch plan as the tone baseline. S6. Based on the tone baseline, calculate the intonation residual components corresponding to the sentence-level intonation changes; S7. Generate a corresponding residual intonation control marker at the end of each IPH boundary to indicate the trend at the end of the current intonation phrase; the trend includes: rising, falling, continuation, convergence, and flattening. S8. Based on the multi-level prosodic boundary sequence, the layered pause duration of each level, the tone baseline, and the residual intonation control markers, generate the prosodic planning sequence of the Chinese text to be synthesized, and complete the multi-level prosodic punctuation and intonation coupling control.
2. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, After generating the prosodic planning sequence of the Chinese text to be synthesized, the prosodic planning information is input into the text to the speech synthesis backend. The speech synthesis backend then drives the speech synthesis output to generate the target speech in a decoupled manner from the frontend and backend, based on the prosodic planning sequence control information.
3. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, The hierarchical consistency correction includes: correcting, downgrading, or eliminating prediction results that do not meet the preset criteria; correcting, filtering, or reassigning prediction results that conflict between different granularity boundaries; and making a comprehensive judgment by combining information such as the distribution of smaller granularity boundaries, boundary prediction probability, punctuation type, and adjacent segment lengths when determining larger granularity boundaries. The preset prediction results include: IPH segments containing one or more PPH segments, and PPH segments containing one or more PW segments.
4. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, The pause duration is predicted for each layer of the predicted sequence. The PW boundary corresponds to zero or very short pauses, the PPH boundary corresponds to short pauses, and the IPH boundary corresponds to medium-long pauses or significant pauses between sentences.
5. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, The basic pitch planning is used to characterize the pitch variation trend dominated by word tone.
6. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, The trend indicated at the end of the current intonation phrase includes: rising, falling, continuation, termination, and flattening.
7. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, The prosodic planning sequence includes: the boundary positions of three layers, PW, PPH, and IPH, the pause duration corresponding to each layer boundary, and the residual intonation control information corresponding to the end of each IPH.
8. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, The intonation residual component is determined based on the difference between the intonation baseline and the frame-level pitch trajectory. It represents the pitch shift caused by sentence-level tone, intonation phrase position, sentence-ending punctuation, sentence structure information, modal particle information, and contextual semantic structure, while maintaining the basic stability of lexical intonation.
9. The multi-layer prosodic phrasing and tone-based speech coupling control method for Chinese TTS according to claim 1, characterized in that, The prosodic planning sequence superimposes the tone baseline with the intonation residual component and concentrates the residual intonation control at the end of the IPH.