A hierarchical rhythm mapping driven short video script tone automatic adjustment method

By constructing a hierarchical indexing system and a semantic anchoring mechanism, the problem of controlling tone and rhythm in speech synthesis technology in short videos has been solved, achieving efficient and natural dubbing generation, which is suitable for short video content of various styles.

CN121260143BActive Publication Date: 2026-03-03CLOUD ATTACK NETWORK TECH HEBEI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511374603.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-03-03
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing speech synthesis technologies struggle to flexibly control tone and rhythm in short video scenarios, resulting in unstable and mechanical outputs. Furthermore, they require significant computational resources, making them unsuitable for low-latency, lightweight applications.

Method used

A four-level indexing system of sentences, phrases, words, and syllables is constructed. Semantic anchors, frozen boundaries, and mirror backfilling mechanisms are introduced. Through the hierarchical propagation and conflict resolution of tone markers, accent markers, and duration markers, a target prosodic control sequence is generated to drive a rule-based formant synthesizer for audio generation.

Benefits of technology

It improves the accuracy and expressiveness of short video dubbing, reduces the time and cost of manual adjustments, ensures the emotional delivery and auditory coherence of the output, and is suitable for short video scripts of various styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260143B_ABST
    Figure CN121260143B_ABST
Patent Text Reader

Abstract

The application discloses a kind of short video script tone automatic adjustment methods of hierarchical prosody mapping drive, it is related to video processing technical field, method includes: step 1: receive text string and language category identification, and establish placeholder column for carrying tonal mark, stress mark and time mark in each level;Step 2: each sentence is divided into phrase segment based on hierarchical index table, freeze boundary as unit of phrase segment, terminate style is preset at the end of sentence according to punctuation, determine core phrase according to semantic anchor point and initialize its direction, and carry out time sequence elastic alignment and hierarchical backfilling, finally output the ternary sequence consisting of tonal mark, stress mark and time mark covering all syllables, as target prosody control sequence;Step 3: generate audio based on target prosody control sequence, obtain new dubbing.The application not only improves the tone accuracy and expressiveness of short video dubbing, but also significantly reduces the time and cost of manual adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, specifically to a method for automatically adjusting the tone of short video copy driven by hierarchical prosody mapping. Background Technology

[0002] With the widespread adoption of short video platforms and the accelerating pace of their dissemination, short video content has become a crucial form of information access, entertainment, and social interaction for internet users. In the dissemination of short videos, the tone and intonation of the script and voice-over play a vital role. When watching short videos, users often rely not only on the text and visuals to convey information but also on the tone, rhythm, and cadence of the voice-over to perceive the emotional tone and communicative intent of the content. Especially in scenarios such as advertising, knowledge dissemination, entertainment explanations, and storytelling, the naturalness of the tone, the matching of emotions, and the harmony with the visuals directly affect the audience's acceptance and the effectiveness of the message.

[0003] Existing speech synthesis technologies can be broadly categorized into two types: concatenation-based speech synthesis and parametric-based speech synthesis. The former relies on a recorded corpus, synthesizing complete speech by concatenating existing audio segments. While maintaining a high degree of naturalness in sound quality, it struggles to flexibly control tone and rhythm, appearing rigid, especially in short videos where rapid adaptation to different text styles is crucial, and failing to meet diverse needs. The latter generates speech through parametric modeling, with Hidden Markov Model (HMM)-based synthesis methods being a typical example. These methods offer superior control flexibility compared to concatenation-based synthesis, but the generated speech generally has a "mechanical" feel and insufficient naturalness in rhythm, particularly limiting its ability to express emotions.

[0004] In recent years, with the development of deep learning, neural network-based speech synthesis technology has gradually become mainstream. Typical methods include end-to-end sequence-to-sequence modeling (such as the Tacotron series) and speech synthesis based on generative adversarial networks (GANs) and variational autoencoders (VAEs). These technologies can generate synthesized results with near-natural speech quality under large-scale corpus training conditions, and the ability to control tone and prosody is also significantly improved. However, these methods have the following problems: First, the training and inference of the model require a lot of computing resources, making them unsuitable for deployment in low-latency, lightweight short video application scenarios; second, the model is highly dependent on large-scale labeled corpora, while the text in short videos is often open-domain and multi-style, making it difficult for existing models to accurately control tone and prosody; third, the generated results are unstable, and problems such as prosodic imbalance, misplaced stress, or abrupt pitch changes may occur, which contradicts the high consistency and high controllability goals required for short videos. Summary of the Invention

[0005] To address the aforementioned technical challenges, this invention provides a hierarchical prosodic mapping-driven method for automatically adjusting the tone of short video text. By constructing a four-level index system (sentences, phrases, words, and syllables) and introducing semantic anchors, frozen boundaries, and mirror backfilling mechanisms, it achieves hierarchical propagation and conflict resolution of tone markers, stress markers, and duration markers, thereby forming a target prosodic control sequence covering all syllables. This sequence drives a rule-based formant synthesizer, applying fluctuations, emphasis, and rhythmic adjustments syllable-by-syllable in the three control channels of tone, stress, and duration, generating a new voiceover that is semantically consistent with the text and flows naturally. Compared to existing technologies, this invention not only improves the accuracy and expressiveness of short video voiceovers but also significantly reduces the time and cost of manual adjustments, ensuring consistency in emotional delivery and auditory coherence, thus possessing strong application value.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A hierarchical prosodic mapping-driven method for automatically adjusting the tone of short video copy, the method comprising:

[0008] Step 1: Receive the text string and language type identifier, perform character normalization, sentence delimitation based on punctuation, word segmentation based on vocabulary and rules, and dependency relationship template recognition, generate a hierarchical index table and semantic anchor table containing four levels: sentence, phrase, word, and syllable, and establish placeholder columns at each level to carry tone markers, stress markers, and duration markers;

[0009] Step 2: Divide each sentence into phrase segments based on the hierarchical index table, freeze the boundaries of each phrase segment, pre-set the sentence termination style according to punctuation, determine the core phrase and initialize its direction based on semantic anchor points, propagate tone and stress between adjacent segments through mirror backfilling, enhance the stress at the anchor point position and set the local contour through co-position alignment, allocate the duration to words and their syllables according to word class, adjudicate the conflict between mirror and co-position differences according to preset priority, implement echo smoothing and stress spacing constraints at the segment connection, and perform time series elastic alignment and hierarchical backfilling. Finally, output a triplet sequence consisting of tone markers, stress markers and duration markers covering all syllables as the target prosodic control sequence.

[0010] Step 3: Generate audio based on the target prosodic control sequence to obtain a new dubbing.

[0011] Furthermore, step 3 specifically includes: inputting the target prosodic control sequence into a rule-based formant synthesizer, synthesizing speech in syllable order using a formant frequency table and an excitation waveform generator, and applying fluctuations, emphasis, and stretching to each syllable in the pitch control, accent control, and duration control channels according to the target prosodic control sequence, thereby generating a new dubbing.

[0012] Further, step 1 specifically includes: receiving the text string and language type identifier, performing character standardization and punctuation unification; dividing the sentence into segments according to periods, question marks, exclamation marks, and line breaks while preserving the original punctuation; segmenting the text within the sentence using a word segmenter based on a lexicon and rules, where the lexicon is a finite set containing function words, conjunctions, negation words, interrogative words, quantifiers, and modal particles; identifying the predicate headword and modifying components using a rule set based on dependency relation templates, and generating a semantic anchor table accordingly, which records interrogative words, negation words, modal particles, quantifiers, conjunctions, and modal particles. The position of the predicate headword in the word segmentation sequence; based on sentence boundaries, conjunctions and pauses, phrase segments are divided within a sentence to obtain a hierarchical index table, which includes a sentence index, a phrase index and a word index, and is expanded into a syllable index through a rule converter from character shape to pronunciation; placeholder columns are established for sentences, phrases, words and syllables according to the hierarchical index table. The placeholder columns are in a sequential list structure and include three items: tone mark, stress mark and duration mark. The initial values ​​of tone mark and stress mark are set to undetermined, the initial value of duration mark is set to baseline, and boundary markers are set at both ends of each phrase segment.

[0013] Furthermore, step 2 specifically includes:

[0014] Step 2.1: Divide each sentence into phrase segments based on the hierarchical index table, and set frozen boundaries for the start and end positions of each phrase segment in the phrase-level placeholder column; preset the sentence termination style according to punctuation, and set transition areas on both sides of the sentence-end syllable for subsequent smoothing;

[0015] Step 2.2: Determine the core phrase in each sentence based on the semantic anchor table, and initialize tone and stress markers in the placeholder column of the core phrase; using the core phrase as a reference, propagate tone and stress markers between adjacent phrase segments using mirror backfilling, maintaining relative positional symmetry during propagation and not crossing frozen boundaries; perform symmetry alignment at semantic anchor positions, raise stress markers at corresponding word-level positions and set local tone contours, and write them into their respective placeholder columns level by level along phrases, words and syllables;

[0016] Step 2.3: Assign duration markers at the word level based on word class and expand them to the syllable level to ensure that the cumulative duration markers within phrase segments are consistent with those at the phrase level; when mirror backfilling, co-position alignment, and sentence-end terminating style differ at the same position, execute conflict resolution according to preset priorities to cover conflicting positions, with sentence-end terminating style taking precedence over mirror backfilling and mirror backfilling taking precedence over co-position alignment; implement echo smoothing at the connection points of adjacent phrase segments, set fade-in and fade-out in the splicing area, and constrain the minimum spacing between adjacent strong accents; if the constraint is not met, downgrade the subsequent strong accent to a medium level;

[0017] Step 2.4: Perform time series flexible alignment on all duration markers, aligning the short, baseline, and long to the corresponding beat grids, while keeping the sentence termination style unchanged; write the tone markers and stress markers determined at the phrase level back to the word level and syllable level, and together with the duration markers, form a triplet sequence consisting of tone markers, stress markers, and duration markers covering all syllables, as the target prosodic control sequence.

[0018] Furthermore, in step 2.2, when propagating tone and accent marks between adjacent phrase segments using mirror backfill with the core phrase as a reference, the phrase segment containing the core phrase is located in the hierarchical index table and recorded as the mirror reference segment; the tone and accent mark fields in the phrase-level placeholder column of the mirror reference segment are read as the initial propagation source; with the mirror reference segment as the mirror center, an ordered list of phrase segments on the left and right sides is determined, and the list order is arranged sequentially according to the adjacency relationship with the mirror reference segment; the presence of frozen boundary markers between adjacent relationships is checked in the phrase-level placeholder column; propagation only occurs within continuous segments not separated by frozen boundaries; when any frozen boundary marker is encountered, propagation stops immediately at that boundary and cannot cross that boundary to continue setting the marks of subsequent segments; when the target segment already has a defined tone and accent mark field in the previous step, its existing values ​​are maintained and it is no longer overwritten by mirror backfill.

[0019] Furthermore, in step 2.2, when the tone marker field of the mirror reference segment is rising, the tone marker field of the first adjacent segment to its right is set to rising, the tone marker field of the first adjacent segment to its left is set to falling, the second adjacent segment to its right continues to be set to rising, the second adjacent segment to its left continues to be set to falling, and so on, until the freeze boundary or the end of the list is reached; when the tone marker field of the mirror reference segment is falling, the adjacent segments at all levels to its right are set to falling in sequence, the adjacent segments at all levels to its left are set to rising in sequence, and so on, until the stop condition is reached; when the tone marker field of the mirror reference segment is flat, the tone marker fields of the adjacent segments at all levels on both sides are set to flat, and so on, until the stop condition is reached; when the mirror reference segment is adjacent to the transition area where the sentence termination style is located, the phrase segments located in the transition area are not modified by mirror backfilling, and the preset result of the sentence termination style in step 2.1 is retained.

[0020] Furthermore, in step 2.2, when the accent mark field of the mirror reference segment is strong, the accent mark fields of its first adjacent segment to the right and the first adjacent segment to the left are both set to medium, the accent mark fields of its second adjacent segment to the right and the second adjacent segment to the left are both set to weak, and the accent mark fields of its third adjacent segment to the right and the third adjacent segment to the left remain indeterminate or weak, depending on whether the segment mainly functions as a conjunction or a function word; when the accent mark field of the mirror reference segment is medium, the accent mark fields of its first adjacent segments to the left and right are set to weak, and the second adjacent segments to the left and right remain indeterminate or weak; when the accent mark field of the mirror reference segment is weak, higher-level accent marks are not propagated outward, and its adjacent segments to the left and right remain indeterminate or weak; if the adjacent segments are located in the transition area of ​​the sentence termination pattern, the accent mark field of that segment is fixed as weak.

[0021] Further, in step 2.3, the phrase fragments of the current sentence are located in the hierarchical index table, and the corresponding word-level placeholder columns are read; a duration marker field is assigned to each word according to its part of speech: nouns, verbs, adjectives, proper nouns, and content adverbs are marked as long; function words, prepositions, conjunctions, and modal particles are marked as short; number strings and time expressions are marked as baseline; if there are no content words marked as long within the same phrase fragment and the phrase-level duration marker field of that phrase fragment is long, then the content word closest to the center of the phrase fragment is promoted from baseline to long; if the same phrase If a phrase contains too many words marked as long and the phrase-level duration marker field of that phrase phrase is short, then from right to left, the excess lengths are successively reduced to the baseline until only one length is retained or it matches the phrase-level duration marker field. The distribution of the number of word-level duration marker fields in each phrase phrase is counted. If the category with the most numbers does not match the phrase-level duration marker field of that phrase phrase, then, based on the principle of minimizing the number of replacements, the duration marker fields of words near the end of the phrase phrase are replaced with the same category as the phrase-level duration marker field, until most categories match.

[0022] Furthermore, in step 2.3, when the word-level duration marker field is long, the syllable carrying the main vowel is marked as long, and the remaining syllables are marked as baseline; when the word-level duration marker field is baseline, all syllables are marked as baseline; when the word-level duration marker field is short, all syllables are marked as short; in words containing both voiceless and voiced consonants, ensure that the first and last syllables are not both long: if the first and last syllables are both long, then only the syllable carrying the main vowel is retained as long, and the other end is changed to baseline; count the number distribution of syllable-level duration marker fields in each phrase segment, if the category with the most quantity is inconsistent with the phrase-level duration marker field, then from right to left in that phrase segment, change several baselines to the category consistent with the phrase-level duration marker field, until most categories are consistent.

[0023] Compared with existing technologies, the advantages of this invention are as follows: Based on existing speech synthesis and text-to-speech technologies, it introduces a multi-level indexing and semantic anchoring annotation mechanism, enabling sentences, phrases, words, and syllables to establish placeholder columns capable of carrying tone markers, stress markers, and duration markers within a unified framework. Compared with existing methods that rely solely on statistical models or single acoustic parameter control, this invention achieves a close integration of the semantic and prosodic layers during the text preprocessing stage, ensuring clear objectives and hierarchical structure for subsequent prosodic control. During sentence segmentation, phrase grouping, and semantic anchor recognition, the setting of frozen boundaries and transition zones avoids crossing structures that should not be altered during prosodic propagation, thereby improving the naturalness and coherence of the tone. Through a mirror backfilling mechanism, using the core phrase as a reference, tone and stress markers are symmetrically diffused between phrase segments, maintaining semantic focus while ensuring overall tone balance. When multiple rules conflict, a priority adjudication mechanism resolves contradictions between tone, stress, and sentence termination patterns, ensuring the uniqueness and stability of the prosodic output. In terms of duration marker allocation, this invention combines word class and syllable structure for refined mapping, matching the duration and semantic weight of different word categories to enhance the rationality of rhythm. Finally, this invention drives a formant synthesizer through a target prosodic control sequence, applying syllable-by-syllable adjustments in the pitch control, accent control, and duration control channels, ensuring that the generated new voice-over is highly consistent with the semantics of the original text in terms of intonation, emphasis, and rhythmic distribution. This solution is particularly valuable in the context of voice-over for short video scripts, not only improving expressive effect and emotional impact but also significantly reducing the cost and time of manual adjustments, providing higher-quality speech output for the automated production of short video content. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating the method for automatically adjusting the tone of short video copywriting driven by hierarchical prosody mapping proposed in this invention.

[0025] Figure 2 This is a schematic diagram illustrating the principle of echo smoothing at phrase segment connections and accent spacing constraints proposed in this invention.

[0026] Figure 3 This is a schematic diagram illustrating the complete structure and technical features of the target prosodic control sequence for the final output proposed in this invention. Detailed Implementation

[0027] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0028] Reference Figure 1As shown, a method for automatically adjusting the tone of short video copy driven by hierarchical prosody mapping is described, the method including:

[0029] Step 1: Receive the text string and language type identifier, perform character normalization, sentence delimitation based on punctuation, word segmentation based on vocabulary and rules, and dependency relationship template recognition, generate a hierarchical index table and semantic anchor table containing four levels: sentence, phrase, word, and syllable, and establish placeholder columns at each level to carry tone markers, stress markers, and duration markers;

[0030] Step 1 uses a text string and a language type identifier as the sole input. A rule base bound to the language type identifier performs character normalization, punctuation-based sentence delimitation, word segmentation based on a vocabulary and rules, and dependency relation template recognition. This ensures the text is deterministically structured before entering the subsequent processes of the hierarchical prosodic mapping-driven automatic tone adjustment method for short video copywriting. The principle of character normalization is to unify character morphology and encoding without altering semantics, folding full-width and half-width characters, unified quotation marks and dashes, whitespace and control characters into a single standardized form, thus ensuring the reproducibility of downstream boundary determination and position indexing. The language type identifier is used to select the corresponding punctuation set, vocabulary set, and dependency relation template set, avoiding ambiguity caused by mixing cross-language rules.

[0031] The principle of punctuation-based sentence delimitation is to treat terminating punctuation marks such as periods, question marks, exclamation marks, and line breaks as hard boundaries, and ellipses and bracket-type punctuation as soft boundaries, which only become hard boundaries when combined with terminating punctuation marks or co-occurring with semantic pauses. This process does not make lexical or semantic guesses, but only uses the normalized punctuation sequence and its order of occurrence in linear positions as the criteria, thereby obtaining stable sentence-level segmentation results and providing start and end ranges for subsequent phrase-level segmentation.

[0032] The word segmentation based on a lexicon and rules employs a deterministic priority matching principle: first, it matches multi-term function words, conjunctions, negation words, interrogative words, quantifiers, and modal particles; then, it matches general word forms and proper noun patterns; finally, it falls back to single characters or the smallest pronable unit. The purpose of priority matching is to reduce segmentation ambiguities without invoking statistical models, prioritizing terms with clear grammatical functions. The lexicon only contains a set of functional and triggering terms related to prosody, avoiding interference from external word frequencies. Segmentation results are recorded in a linear order in a word-level index, preserving the original start and end positions to ensure that the original text can be mapped back from any position at any time.

[0033] Dependency relation template recognition performs directed relation pairing on segmented sequences using a rule set. Templates are constructed around predicate headwords, subjects, objects, adverbs, negation markers, interrogative markers, quantity modifiers, and parallel connectors. When a rule matches multiple paths simultaneously, it decides based on template priority and the shortest span principle. This recognition does not generate probabilities; it only outputs relation triples that satisfy the templates, and extracts a semantic anchor table based on these. The semantic anchor table records the positions of interrogative words, negation words, modal words, quantity words, conjunctions, and predicate headwords in the segmented sequence and the indices of their respective sentences. This is used to trigger mirror backfilling, apposition alignment, and consistency control of termination styles in subsequent steps.

[0034] Based on the above, a hierarchical index table is constructed, expanding from top to bottom according to four levels: sentence, phrase, word, and syllable. The sentence level is directly given by sentence boundaries based on punctuation. The phrase level uses conjunctions, commas, semicolons, and dependency templates to generate phrase fragments within sentences and records the start and end positions and adjacency relationships of these fragments. The word level is directly given by the word segmentation results and retains the original offset. The syllable level uses a glyph-to-pronunciation rule converter to map the word-level sequence into a sequence of articulated syllables. When encountering multiple pronunciations or ambiguities in stress, the corresponding pronunciation priority table is used to resolve the issue. The conversion result establishes a stable one-to-many reference relationship with the word level, ensuring the refillability of the syllable and word levels in subsequent duration and stress propagation.

[0035] Placeholder columns are created at each level of the hierarchical index table to serve as containers for subsequent prosodic calculations. Each placeholder column corresponds one-to-one with the linear order of its level, and each position contains three fields: tone marker, stress marker, and duration marker, initialized to undetermined, undetermined, and baseline. Boundary markers are simultaneously written to both ends of each phrase segment within a sentence in the placeholder columns, serving as trigger criteria for subsequent boundary freezing and echo smoothing. Through this structure, tone markers, stress markers, and duration markers can perform top-down distribution and bottom-up consistency checks among phrases, words, and syllables. The trigger positions provided by the semantic anchor table can be directly indexed to the corresponding units in the word-level and syllable-level placeholder columns, achieving deterministic placement for alignment and priority determination.

[0036] Step 2: Divide each sentence into phrase segments based on the hierarchical index table, freeze the boundaries of each phrase segment, pre-set the sentence termination style according to punctuation, determine the core phrase and initialize its direction based on semantic anchor points, propagate tone and stress between adjacent segments through mirror backfilling, enhance the stress at the anchor point position and set the local contour through co-position alignment, allocate the duration to words and their syllables according to word class, adjudicate the conflict between mirror and co-position differences according to preset priority, implement echo smoothing and stress spacing constraints at the segment connection, and perform time series elastic alignment and hierarchical backfilling. Finally, output a triplet sequence consisting of tone markers, stress markers and duration markers covering all syllables as the target prosodic control sequence.

[0037] In step 2, each sentence is first segmented into phrase fragments based on the hierarchical index table, and frozen boundaries are written at the start and end positions of the phrase-level placeholder columns, so that any subsequent propagation and modification of tags are controlled within a fixed fragment range. The sentence termination style is pre-set to the transition area on both sides of the sentence's final syllable based on punctuation, forming a higher-level constraint on the overall sentence direction; this constraint takes precedence over all subsequent propagation and alignment, ensuring that the final output is consistent with the termination sense corresponding to periods, question marks, and exclamation marks. The determination of the core phrase relies on the semantic anchor table, identifying phrase fragments containing predicate headwords and with high anchor density as the main carrying fragments within the sentence, and initializing tone and stress marks in their phrase-level placeholder columns, thereby providing a unique direction and emphasis source for the entire sentence. Mirror backfilling uses the core phrase as the mirror center and expands outward along the left-right adjacency relationship, propagating tone and stress markers between adjacent segments while maintaining relative positional symmetry. It stops upon encountering a frozen boundary, without setting values ​​for subsequent segments across that boundary. Therefore, the prosodic skeleton within the sentence unfolds layer by layer along the adjacent sequence of segments without disrupting the initial structure. When performing co-alignment on semantic anchor points, the stress is enhanced and a local outline is set at the corresponding word level, then written into the respective placeholder columns level by phrase, word, and syllable, ensuring that the emphasis of the triggering word is coupled with the phrase-level direction without losing local discernibility. The allocation of duration markers adopts a word-class driven strategy, first distinguishing between long, baseline, and short at the word level, then expanding to syllables within the word, ensuring that the cumulative duration markers within phrase segments are consistent with the phrase level, providing a stable baseline for subsequent time mapping from a rhythmic perspective.

[0038] When mirror backfill, alignment, and sentence-ending patterns differ at the same position, a comprehensive conflict resolution is performed according to a preset priority, with sentence-ending patterns taking precedence over mirror backfill, and mirror backfill taking precedence over alignment. This priority arranges sentence convergence, skeleton propagation, and local emphasis from largest to smallest impact range, ensuring a monotonous and traceable decision-making process, and preventing re-rolling within the same round after resolution. Echo smoothing is implemented at the connection points of adjacent phrase segments, with fade-in and fade-out settings applied to both ends within the splicing area. The pitch markers in the splicing area are unified as sustained, and the accent markers are unified as weak, reducing the sense of boundary and energy abrupt changes between segments. To avoid auditory crowding caused by emphasis clustering, a constraint is imposed on the minimum spacing between adjacent strong accents. If the constraint is not met, the subsequent strong accent is downgraded to a medium level, and if necessary, its position is slightly shifted without touching the frozen boundary, maintaining a perceptible sparsity in the distribution of strong accents on the time axis. Temporal series flexible alignment maps short, reference, and long segments to corresponding beat grids while maintaining sentence termination patterns, thus transforming discrete duration markers into a synthesizable temporal layout. Finally, hierarchical backfilling is performed, writing the phrase-level determined tone and accent markers back to the word and syllable levels, forming a triplet sequence of tone, accent, and duration markers covering all syllables, which is output as the target prosodic control sequence. Through the orderly integration of frozen boundaries, termination patterns, mirror backfilling, syntagmatic alignment, duration allocation, conflict resolution, echo smoothing, accent spacing constraints, and temporal series flexible alignment, step 2 assembles and converges the prosodic intent from sentence to syllable level within the same indexing system from top to bottom, providing a definite, complete, and directly driveable control signal for audio generation in step 3.

[0039] Step 3: Generate audio based on the target prosodic control sequence to obtain a new dubbing.

[0040] Specifically, step 3 is driven syllable by syllable through the target prosody control sequence in the pitch control channel, accent control channel, and duration control channel: the sampling rate is set to 24000 Hz, the quantization bit depth to 16 bits, and the channel to mono. The formant frequency table is loaded, which provides the reference values ​​for the first, second, and third formant frequencies and their corresponding bandwidths based on the vowel and consonant combinations of the syllable. The excitation waveform generator is started, with excitation types including periodic pulses and broadband noise; periodic pulses are used when the syllable is predominantly voiced consonants or vowels, and broadband noise is used when the syllable is predominantly voiceless fricatives or affricates. When the syllable contains both voiced and voiceless segments, segments are called in either the voiceless-to-voiced or voiced-to-voiced order. A tandem formant filter bank is established, with the center frequency and bandwidth derived from the formant frequency table, allowing for slow variations within the syllable by the pitch control channel and accent control channel. An envelope generator is established to generate syllable-by-syllable amplitude envelopes and transition envelopes.

[0041] When the tone mark is rising, the tone control channel within the syllable will slowly scale the fundamental period of the excitation waveform from short to long over time, raising the pitch from the starting point to the ending point by 10% to 25%. When the tone mark is falling, the pitch will decrease from the starting point to the ending point by 10% to 25%. When the tone mark is level, the pitch change remains within ±2%. When the tone mark is partially rising or falling, the corresponding change is applied only in the middle of the syllable, while the beginning and end of the syllable remain level. When the tone mark is terminating rising or falling, the corresponding change is applied only in the latter half of the last syllable of the sentence. When the tone mark is continuing, the pitch ending point of the previous syllable is directly used as the starting point and main direction of this syllable.

[0042] When the accent mark is strong, the accent control channel increases the peak amplitude envelope to 1.4 to 1.6 times the base duration while narrowing the first formant bandwidth by 10% to 20% to enhance focus. When the accent mark is medium, the peak amplitude envelope increases to 1.15 to 1.3 times the base duration while maintaining the bandwidth. When the accent mark is weak, the peak amplitude envelope remains at the baseline or decreases to 0.9 to 1.0 times the base duration. When the duration mark is long, the duration control channel assigns a target duration of 1.2 to 1.4 times the base duration to the syllable. When the duration mark is baseline, the target duration is the baseline duration. When the duration mark is short, the target duration is 0.7 to 0.9 times the base duration. The baseline duration is given by a rule base of word class and syllable structure, with vowel-carrying syllables receiving a longer baseline duration.

[0043] Read the reference entries of the syllable in the formant frequency table to obtain the reference values ​​of the first, second, and third formant frequencies and their bandwidths. Based on the consonant and vowel order of the syllable, segment the excitation waveform generator; in the voiced segment, use periodic pulses and generate a sampled pitch trajectory according to the pitch control channel; in the unvoiced segment, use broadband noise and keep the pitch uninvolved in the synthesis. Pass the excitation signal sequentially through a series formant filter bank; when the accent is marked as strong or moderate, adjust the bandwidth or amplitude response according to the rules in Part Two; when there is a transition from unvoiced to voiced or vice versa within the syllable, use an envelope generator to set a fade-in or fade-out of no less than 5 milliseconds at the transition point. Use the duration control channel to stretch or compress the number of sampling points for the syllable; stretching and compression use a combination of resampling and overlay addition, and the stretching or compression ratio is given by the mapping in Part Three. Crossfade splicing zones are set between adjacent syllables, with a splicing zone length of 5 to 15 milliseconds. Linear fade-in and linear fade-out of amplitude are applied at the entry and exit points of the splicing zone, while zero crossover is aligned at the midpoint of the splicing zone to avoid popping. Average energy normalization is performed on the synthesized segments of all syllables within each sentence to maintain the relative elevation relationship of strong accents and ensure overall loudness consistency. When the sentence termination style is either terminating upward or downward, the second half of the last syllable is re-checked and corrected to ensure consistency with the target prosodic control sequence. A mono pulse-code modulated audio data file is output. The file header records the sampling rate and quantization bit depth, and the file body consists of waveform data arranged in syllable order, which constitutes the new dubbing.

[0044] When the target prosodic control sequence is missing a pitch marker at a certain position, the pitch control channel uses a flat rule; when an accent marker is missing, a weak rule is used; and when a duration marker is missing, a baseline rule is used. When the character corresponding to a syllable is an emoji, a hashtag, or an uncommon character, the three controls—flat, weak, and baseline—are directly generated and connected to the adjacent syllable with the shortest splicing area. When the interval between two adjacent strong accents is too small, causing the amplitude peaks to overlap, the amplitude envelope peak of the latter strong accent is lowered to the medium range, and the crossfading splicing area is extended to no more than 20 milliseconds. Through the above processes, the target prosodic control sequence is completely mapped within the pitch control channel, accent control channel, and duration control channel and applied to the rule-based formant synthesizer. The formant frequency table and the excitation waveform generator are stably driven in syllable order, ultimately resulting in a new dubbing consistent with the target prosodic control sequence.

[0045] Furthermore, in step 1, the text string and the language identifier are used as the sole inputs into the same deterministic processing chain. Character normalization and punctuation unification are performed first, folding full-width and half-width characters, irregular quotation marks and dashes, invisible control characters and redundant whitespace into a single standardized form, while retaining the original offset index without changing the character order, so that any subsequent position can be traced back to the source text. The language identifier is used to select the punctuation set, vocabulary set, and dependency relationship template set bound to that language, avoiding ambiguity and non-reproducibility caused by cross-language mixing. Placeholder columns are established for sentences, phrases, words, and syllables according to the hierarchical index table. The placeholder columns are in a sequential list structure, with each position corresponding to the linear order of the corresponding level. Each position contains three items: tone mark, accent mark, and duration mark, used to carry the control information generated in steps 2 and 3. During initialization, the initial values ​​of the tone mark and accent mark are set to undetermined, and the initial value of the duration mark is set to the baseline to ensure consistency of null values ​​before subsequent propagation and adjudication. Boundary markers are set at both ends of each phrase segment to serve as hard stop conditions and transition zone trigger conditions when freezing boundaries, echo smoothing, and stress spacing constraints are executed in step 2. Through the above structured output, step 1 transforms free text into an indexable, backfillable, and adjudicable data carrier under unified language rules. The hierarchical index table and semantic anchor table together provide a definite entry point and boundary for the subsequent assembly and synthesis of the hierarchical prosodic mapping-driven short video copywriting tone automatic adjustment method.

[0046] Furthermore, in step 2.1, each sentence is first segmented into phrase fragments based on the hierarchical index table, and frozen boundaries are written at the start and end positions of the phrase-level placeholder columns. This ensures that the generation, propagation, and modification of any subsequent tone markers, stress markers, and duration markers are confined to the fixed fragment range and do not cross the determined structural boundaries. The sentence-ending style is pre-set to the transition area on both sides of the sentence-ending syllable based on punctuation, serving as a superordinate constraint on the overall sentence direction. The transition area provides a unified buffer zone in terms of time and boundaries for subsequent mirror backfilling, co-position alignment, and echo smoothing, preventing the ending style from being covered by local adjustments.

[0047] In step 2.2, a core phrase is determined within each sentence based on the semantic anchor table, and tone and stress markers are initialized in the phrase-level placeholder column of the core phrase, ensuring that the entire sentence has a unique direction and emphasis source. Mirror backfilling is performed with the core phrase as a reference, propagating tone and stress markers outward in pairs along the left and right adjacent relationships, maintaining relative positional symmetry; it stops once a frozen boundary is encountered to avoid cross-segment interference with the established structure. Co-alignment is performed on the semantic anchor positions, raising the stress marker at the corresponding word-level position and setting the local tone outline, then writing it step by step along the phrase, word, and syllable into their respective placeholder columns, ensuring that local emphasis and segment-level direction coexist. Mirror backfilling is responsible for expanding the skeleton, while co-alignment is responsible for highlighting trigger words; both are jointly constrained by frozen boundaries and transition areas, and are also subject to priority control by the sentence-end termination style.

[0048] In step 2.3, duration markers are assigned at the word level according to word class and expanded to the syllable level to ensure that the cumulative duration markers within phrase segments are consistent with those at the phrase level, providing a stable time allocation for the entire sentence from a rhythmic perspective. If mirror backfilling, co-position alignment, and sentence-end terminating style differ at the same position, a coverage-style conflict adjudication is performed according to a preset priority, with sentence-end terminating style taking precedence over mirror backfilling, and mirror backfilling taking precedence over co-position alignment, ensuring that the decision-making process is monotonous and traceable. Echo smoothing is implemented at the connection points of adjacent phrase segments, with fade-in and fade-out settings within the splicing area, and the tone markers in the splicing area are uniformly set to sustained and the accent markers to weak, reducing the sense of segment boundaries and energy abrupt changes. A minimum spacing constraint is applied between adjacent strong accents. When the constraint is not met, the subsequent strong accent is downgraded to a medium level, and if necessary, its position is slightly shifted without touching the frozen boundary, so that the emphasis maintains a perceptible sparsity on the time axis without disrupting the continuity of the segment.

[0049] In step 2.4, time-series flexible alignment is performed on all duration markers, aligning short, reference, and long markers to their corresponding beat grids. This transforms the discrete duration markers into a directly driveable temporal layout under a unified time reference, while maintaining the sentence-ending style to ensure stable presentation of the ending. The phrase-level determined tone and accent markers are written back to the word and syllable levels and combined with the duration markers to form a triplet sequence covering all syllables, consisting of tone, accent, and duration markers, thus forming the target prosodic control sequence. Through the sequential connection of frozen boundaries, termination styles, mirror backfilling, syntagmatic alignment, duration marker allocation, conflict resolution, echo smoothing, minimum spacing constraints, and time-series flexible alignment, step 2 completes the top-down assembly and consistency convergence from sentence to syllable within the same indexing system, providing a definite, complete, and traceable control signal for subsequent audio generation.

[0050] Furthermore, in step 2.2, a mirror backfilling propagation mechanism is established with the core phrase as a reference, starting from the phrase segment containing the core phrase in the hierarchical index table. This segment already has initialized values ​​for the tone marker and accent marker fields in the phrase-level placeholder column, serving as the initial propagation source. To ensure the reproducibility of the update order and results, firstly, using this segment as the mirror center, an ordered list of phrase segments is generated in both the left and right directions according to the adjacency relationships recorded in the hierarchical index table. The lists are ordered by the adjacent levels from the mirror center, forming a sequence from level 1, level 2, down to the sentence boundary. While generating the lists, the frozen boundary markers at the boundaries of each segment in the phrase-level placeholder column are read and used as hard stopping conditions for propagation. In this way, the propagation range is limited to the set of continuous segments not separated by frozen boundaries. As soon as any frozen boundary marker is encountered, the propagation terminates at that boundary, without crossing into segments outside the boundary, thus avoiding intrusive modifications to the structure already determined by step 2.1 or the sentence-end termination pattern. If the target segment already has a defined pitch or accent mark field in the preceding steps, mirror backfilling does not overwrite its existing value, but leaves it unchanged. It is then handled uniformly by subsequent conflict resolution to ensure that the marks from different sources have clear priorities and boundaries of responsibility.

[0051] Mirror backfilling unfolds symmetrically around relative positions. Updates are performed using a "paired advancement" strategy, meaning that in each round, one adjacent segment with the same level number on the right and one on the left are processed simultaneously, keeping both sides synchronized in propagation depth. If one side reaches the freeze boundary and stops during advancement, the other side can continue, but is no longer forced to maintain synchronization with the stopped side, preventing the entire sentence propagation from being abandoned halfway due to unilateral limitation. Before each round of processing begins, the scheduler checks whether the target segments on both sides of the layer have already been written with definite values ​​in previous steps. If either segment already has a definite value, the mirror writing of that segment is skipped, and backfilling is only performed on undefined segments, thus ensuring that mirror backfilling only fills blanks and does not repeatedly overwrite results from collocation alignment or sentence termination patterns.

[0052] The propagation of tone markers follows a symmetrical outward expansion rule based on the core phrase. When the core phrase is rising, adjacent segments on the right are successively set to rising, and adjacent segments on the left are successively set to falling; when the core phrase is falling, the right side is successively set to falling, and the left side is successively set to rising; when the core phrase is level, both sides are successively set to level. This rule allows the overall direction of the sentence to expand skeletally along adjacent structures starting from the core phrase, maintaining consistency in the direction of tone closure while creating a natural contrast and echo through the left-right contrast. Segments located in the transition area of ​​the sentence-ending pattern are not subject to mirroring changes, retaining the results preset in step 2.1, ensuring that the sentence-ending closure is not altered midway. Through this exclusive retention strategy, the sentence-ending directive constraint becomes the highest priority control point above all propagation and fine-tuning, ensuring a stable and controllable sense of ending for the entire sentence.

[0053] The propagation of the accent marker field exhibits a decreasing emphasis distribution from the center outwards. When the core phrase is strong, the first layer on both sides is set to medium, the second layer to weak, and the third layer and beyond remain indeterminate or weak, depending on whether the segment primarily functions as a conjunction or function word. When the core phrase is medium, the first layer on both sides is set to weak, and the second layer and beyond remain indeterminate or weak. When the core phrase is weak, higher-level accent markers are not propagated outwards, and adjacent segments remain indeterminate or weak. This decreasing distribution creates a perceptible sparsity of emphasis on the time axis, preventing multiple strong accents from clustering together in a short period and causing perceptual crowding, while providing a natural initial layout for the minimum spacing constraint of subsequent adjacent strong accents. Segments located within the sentence-ending transition zone are fixed as weak in the accent marker field, consistent with the tonal direction preservation strategy, avoiding the introduction of additional energy peaks at the end.

[0054] To ensure that mirror backfill respects the existing structure and that subsequent processing is predictable, the update process employs a combination of strategies: writing only indeterminate values, stopping at boundaries, pairwise advancement, and unilateral release. Writing only indeterminate values ​​ensures that local emphasis from co-alignment and convergence from sentence-ending patterns are not overwritten; stopping at boundaries elevates the frozen boundary to a hard threshold for process control, making intra-segment propagation and inter-segment isolation clearly achievable; pairwise advancement establishes relative positional symmetry in propagation, meaning that both sides of each layer are evaluated and written simultaneously, ensuring balanced expansion of both sides of the skeleton; unilateral release avoids information loss on one side due to premature stopping on the other, thus enabling mirror backfill to output maximum coverage even under imperfect structural conditions. These strategies work together to give mirror backfill a definite advancement sequence and repeatable convergence state, meaning that multiple runs with the same input and the same frozen boundary layout yield the same phrase-level placeholder column writing results, meeting the engineering requirements for testability and verifiability.

[0055] The connection between alignment and conflict resolution is achieved through two interface constraints. First, mirror backfilling does not write directly at the word and syllable levels. Instead, it first completes the phrase-level skeleton construction, then alignment enhances the stress at specific word-level positions and sets the local tone contour, and finally writes step by step along the phrase, word, and syllable levels, avoiding repetition and reverse writing at the same level. Second, mirror backfilling keeps the determined positions read-only. When it encounters overlap with alignment or sentence-end termination patterns, it marks the position as a candidate conflict point and submits it to the preset priority in step 2.3 for overriding processing, ensuring that the decision chain is monotonous and traceable. Through these two constraints, mirror backfilling assumes the responsibility of "skeleton first, then refinement," avoiding overlap with the responsibility of alignment, which prioritizes "local emphasis."

[0056] At the implementation level, mirror backfill is built upon the linear order and adjacency relationships of the hierarchical index table, without introducing any additional implicit structure. Each phrase fragment write is accompanied by a source identifier, indicating which level of propagation the value comes from, or from initialization or alignment in a previous step. The source identifier allows subsequent conflict resolution to be overridden according to the established order: sentence termination style takes precedence over mirror backfill, and mirror backfill takes precedence over alignment. It also facilitates accurate location of recalculated intervals when backtracking is needed. Since mirror backfill only writes to indeterminate positions and stops immediately upon encountering a frozen boundary, the entire process maintains minimal intrusion into the input structure. For scenarios involving parallel processing of multiple sentences, mirror backfill is completely closed within the sentence, without accessing data outside the sentence, naturally supporting parallelization.

[0057] When the mirror reference fragment is adjacent to the transition zone where the sentence-ending pattern is located, the propagation start point is adjacent to the convergence zone. In this case, no mirroring is performed on the phrase fragments within the transition zone to avoid directly extending the core phrase's direction to the end, thus maintaining the purity of the termination pattern. Instead, propagation only proceeds within continuous fragments outside the transition zone, forming a "core-emerging, tail-self-protecting" structural configuration, allowing the upward or downward movement at the end of the sentence to be clearly identified in terms of energy and direction. This strategy is consistent with the time-series flexible alignment in step 2.4, which does not make alignment changes at the end of the sentence, further consolidating the stable presentation of the sentence-end from a temporal perspective.

[0058] From a holistic perspective, the core value of mirror backfilling lies in transforming the single-point direction and emphasis provided by the core phrase into an executable skeleton across segments. Simultaneously, by freezing boundaries, transition zones, and adhering to the discipline of writing only indeterminate values, this skeleton is stably embedded into the sentence structure. Once the skeleton is formed, appositional alignment can overlay local contours at specified words, and duration marker allocation and flexible time series alignment further solidify rhythmic information into a synthesizable temporal allocation. These three elements work in coordination within the same indexing system, ultimately converging at the end of step 2 into a triplet sequence covering all syllables, composed of tone markers, stress markers, and duration markers. The deterministic advancement, symmetrical expansion, boundary constraints, and the combination of writing only indeterminate values ​​in mirror backfilling ensure that this triplet sequence reflects both the overall sentence direction and emphasis gradient, while providing clear, stable, and traceable input for subsequent syllable-by-syllable driving in the tone control, stress control, and duration control channels.

[0059] Furthermore, in step 2.3, for all phrase fragments in the current sentence in the hierarchical index table, a one-time determination and hierarchical expansion are performed around the duration marker field of the word-level placeholder column. The goal is to provide a stable rhythmic baseline for subsequent flexible alignment of the time series without changing the tone marker field and accent marker field. The processing starts from phrase fragment location and sequentially reads the corresponding word-level placeholder columns, so that each word has a writable and traceable duration marker field position. To avoid ambiguity and repeated modifications, allocation and expansion are completed in the same round, and each write action only occurs in undetermined positions or positions inconsistent with the phrase-level duration marker field.

[0060] The word-level allocation is divided into three parts based on the differences in the rhythmic effect of word classes: nouns, verbs, adjectives, proper nouns, and content adverbs are marked as long; function words, prepositions, conjunctions, and modal particles are marked as short; and number strings and time expressions are marked as baseline. This division gives more time to information-carrying words, converges words that connect organizational structures or speech flow to shorter time values, and fixes readings and time expressions at the neutral length of the baseline, facilitating the formation of discernible strong and weak relationships within the same phrase segment. After the initial allocation, the consistency between the phrase segment and its phrase-level duration marker field is checked. When the phrase level is long but no content words in the segment receive a long marker, the selection rule closest to the center of the phrase segment is used, elevating the content word closest to the center of the phrase segment from the baseline to long. This rule concentrates the stretching on the carrying center of the segment, avoiding rhythmic bias caused by stretching only at the edges. When a phrase is short but there are too many long phrases within a segment, the excess lengths are reduced from right to left to the baseline until only one length is retained or the length matches the phrase-level duration marker field. The right-to-left reduction order aligns with the natural closing direction within a sentence, prioritizing the retention of later emphasis values ​​within the segment and reserving a small amount of closing space for the sentence-ending style.

[0061] After completing the above local corrections, the distribution of word-level duration marker fields is statistically analyzed for each phrase segment. If the most frequent category is inconsistent with the phrase-level duration marker field, the replacement is performed from the end of the segment backwards with the goal of minimizing the number of replacements. The duration marker fields of several words near the end of the phrase segment are replaced with categories consistent with the phrase-level category until most categories are consistent. Choosing the end of the segment as the replacement starting point reduces disturbance to the rhythm near the core phrase and ensures a smoother transition with the ending pattern of subsequent sentences. This majority consistency strategy ensures that the overall trend at the word level is aligned with the phrase-level tags, while controlling the extent of modification through the "minimum number of replacements" constraint, maintaining compatibility with the tone and stress arrangement generated in step 2.2.

[0062] The expansion from word level to syllable level follows a mapping of "prioritizing the main vowel." For words marked as long, the syllable carrying the main vowel is marked as long, and the remaining syllables are marked as baseline; for words marked as baseline, all syllables are marked as baseline; for words marked as short, all syllables are marked as short. This expansion concentrates the stretching within the word on the main vowel position, which best reflects the timbre and sound quality, avoiding unnecessary lengthening of weak syllables and wasting energy. For compound words containing both voiceless and voiced consonants, to avoid boundary stretching and splicing conflicts caused by simultaneous lengthening at both ends, a constraint is set that the first and last syllables cannot be long simultaneously; if the first and last syllables are already long simultaneously, only the end carrying the main vowel is retained as long, and the other end is changed to baseline. Through this constraint, the center of gravity of the word's duration is stabilized at the core of the sound quality, leaving the boundary for subsequent echo smoothing and crossfading.

[0063] After the initial expansion at the syllable level, the distribution of syllable-level duration markers is recalculated for each phrase segment to check if the majority of categories match the phrase-level duration markers. If they do not match, several baselines within the phrase segment are changed to categories consistent with the phrase level, from right to left, until the majority of categories match. The right-to-left order is still used to coordinate with the closing direction of the terminating pattern and to minimize disruption to the rhythmic structure of the preceding segment and its neighbors in the core phrase. This round only replaces between baselines and target categories, without altering syllables already set to long or short, thus avoiding further fragmentation of word duration distribution after expansion.

[0064] The entire allocation and expansion process is executed under four disciplines: no cross-level restrictions, no overwriting of established rules, minimal replacements, and end-priority. Restricting cross-level restrictions ensures the source of the duration marker field is clear: word-level is determined first, with only mapping and fine-tuning done at the syllable level; no overwriting of established rules ensures a stable priority order with step 2.2 and sentence-end termination patterns; minimal replacements ensure that each consistency check only makes necessary changes, maintaining coordination with existing tone and accent marker fields; and end-priority ensures that the time allocation within a sentence is consistent with the convergence direction, providing directly alignable category proportions for the flexible time series alignment in step 2.4.

[0065] The relationship between mirror backfilling and co-alignment is kept clearly defined. Mirror backfilling defines the segment-level direction and emphasis gradient, while co-alignment imparts locally perceptible stress and contour at anchor points. The allocation and expansion of duration marker fields do not attempt to alter these two aspects, but rather provide space for them in the time dimension. Even if long words appear near weak accent marker fields in certain positions, the accent marker fields are not directly altered. Instead, time yielding and energy yielding are handled separately, with subsequent echo smoothing and minimum spacing constraints jointly suppressing potential conflicts. The practice of fixing numerical strings and time expressions as the baseline ensures that they do not actively compete for time value and energy within any segment, thereby avoiding weakening the semantic focus of the segment.

[0066] The convergence of consistency between phrase-level, word-level, and syllable-level units is based on majority category consistency, rather than imposing rigid proportional constraints on all units. This aims to accommodate subtle fluctuations in natural speech flow while maintaining the stability of the overall rhythmic profile. In engineering implementation, this strategy relies only on statistical analysis and a finite number of substitutions, without requiring continuous parameter solving, thus possessing predictable complexity and a clear stopping condition. Each substitution is recorded in a placeholder column, which can be directly read and converted into a raster alignment target by the time-series flexible alignment in step 2.4, reducing intermediate inferences across steps.

[0067] When phrase segments are short or contain only a few words, the two rules of center selection and end replacement automatically degenerate into single adjustments to individual positions, still ensuring consistency with the phrase-level duration marker field. For longer segments, most consistency statistics are performed once at the word level and once at the syllable level, avoiding syllable-level imbalances caused by intra-word expansion after consistency is achieved only at the word level. The combination of two-level statistics and replacement ensures consistency at different granularities, thus eliminating the need for additional correction steps during the time series flexible alignment phase.

[0068] Through the above process, the duration marker field performs selective stretching and contraction at the word level, focuses on the main vowel at the syllable level, and achieves majority category consistency with the phrase-level duration marker field at both levels. This result complements step 2.2: the former provides a rhythmic framework, while the latter provides direction and emphasis. Subsequently, in step 2.4, long, baseline, and short are mapped to the corresponding beat grid while maintaining the sentence termination style. The three types of markers together form a triplet sequence covering all syllables, consisting of tone markers, accent markers, and duration markers, which is input as the target prosodic control sequence to step 3. The entire processing chain maintains the engineering characteristics of deterministic input, finite-step convergence, and minimal substitution, ensuring that the same duration allocation and expansion results can be reproduced under different text and language type identifiers, providing a clear and stable time baseline for subsequent syllable-by-syllable driven synthesis.

[0069] The following presents an implementation example that can be fully realized in engineering. The example text is the Chinese sentence: "Now place an order, and immediately participate?". The language type identifier is "Chinese". The input text string and the language type identifier enter the processing chain. After character normalization and punctuation unification, single sentences (ending with a question mark) are obtained based on punctuation-based sentence boundaries. The word segmentation result based on the word list and rules is the word sequence: [Now], [place an order], [immediately], [participate], [?]. The dependency relationship template identifies the predicate central word as [participate], the secondary predicate as [place an order], and marks the modal particle [?] and the adverb [immediately], and writes them into the semantic anchor table. Using the comma as the pause symbol, phrase segments are divided within the sentence, obtaining phrase segment P1 = [Now, place an order], phrase segment P2 = [immediately, participate,?], and freezing boundaries are set at both ends of the phrase-level occupancy column. The grapheme-to-phoneme rule expands the word level to the syllable level (represented by the main vowel as the representative syllable): [xian][zai][xia][dan][ma][shang][can][yu][ma], a total of 9 syllables. Occupancy columns are established for sentences, phrases, words, and syllables step by step, and each position contains three items: tone mark, stress mark, and duration mark, which are initialized as undetermined, undetermined, and baseline.

[0070] Based on the question mark, the end termination style is set to terminate with an upward pitch, and transition zones are set on both sides of the last syllable. According to the semantic anchor table, the core phrase is determined to be P2 (including the predicate central word [participate] and the modal particle [?]). In the phrase-level occupancy column of P2, the tone mark is initialized as upward and the stress mark as strong; P1 is temporarily undetermined at the phrase level. Using P2 as the mirror reference, it is propagated to the adjacent left and right segments: the tone mark of the left P1 is set to downward and the stress mark to medium; there is no segment on the right. Perform co-occurrence alignment on the semantic anchors: at the word level, the stress mark of [participate] is raised to strong and the local tone contour is set to local upward; the stress marks of [place an order] and [immediately] are raised to medium and the local tone contour is set to local upward; [?] is fixed as weak and follows the termination with an upward pitch.

[0071] Map the word classes to duration marks: [Now] is a time expression, marked as baseline; [place an order] is a verb, marked as long; [immediately] is a content adverb, marked as long; [participate] is a verb, marked as long; [?] is a modal particle, marked as short. The phrase-level duration mark of phrase segment P1 is set to baseline; there is already one long ([place an order]) in P1, so there is no need to backtrack; the phrase-level duration mark of phrase segment P2 is set to long; there are already longs ([immediately], [participate]) in P2, meeting the tendency of "most are long". Expand the word-level marks to the syllable level: when the word mark is long, only the syllable carrying the main vowel is marked as long, and the rest are baseline; when the word mark is baseline, all syllables are baseline; when the word mark is short, all syllables are short. Thus, the syllable-level duration mark sequence (corresponding to the above syllables) is obtained: baseline, baseline, long, baseline, long, long, long, long, short.

[0072] If any positional alignment conflicts with mirror backfill or sentence-ending style, the following order applies: sentence-ending style takes precedence over mirror backfill, and mirror backfill takes precedence over positional alignment. Echo smoothing is applied at the connection points of adjacent phrase segments: the splicing area length is 10 milliseconds, and the tone markers within the splicing area are set to sustained, while the accent markers are set to weak. The minimum spacing between adjacent strong accents is scanned; if it is less than one syllable, the subsequent strong accent is downgraded to medium (in this example, the strong accents are concentrated in [participate], and the spacing meets the requirements).

[0073] Align the short, reference, and long beats to the beat grid respectively. Let the reference beat duration be a parameter. ( The duration of a single reference syllable (in seconds). Let the duration ratios of the long, reference, and short syllables be parameters. , , In this embodiment, we take... The tone and stress markers determined at the phrase level are written back to the word and syllable levels, and together with the duration markers, they form the target prosodic control sequence.

[0074] Let the sampling rate be a parameter. (Unit: Hertz), in this example, we take... Let the target duration of a single syllable be a parameter. (No. The duration of each syllable (in seconds), and the number of sample points are parameters. (No. The number of sampling points per syllable (unit point), then Tone control uses the fundamental frequency as a parameter. (time The instantaneous fundamental frequency (in Hertz), and each syllable has an initial fundamental frequency parameter. With termination fundamental frequency parameters Accent control uses amplitude gain as a parameter. (No. The amplitude multiple of each syllable (dimensionless). Let the gain corresponding to strong, medium, and weak accents be . .

[0075] Let the fundamental frequency baseline be a parameter. (Unit: Hertz), in this example, we take... Let the upward ratio parameter be... (Dimensionless), the descent ratio parameter is (Dimensionless), in this example, we take... The first resonant filter The center frequency of each resonance peak is a parameter (Unit: Hertz), bandwidth is a parameter (Unit: Hertz), pole radius parameter is . No. The transfer function of a second-order resonator is . The total sound channels are in series product .

[0076] The target prosody control sequence to duration and fundamental frequency

[0077] According to the syllable-level duration markings obtained in step 2, calculate the for each syllable. The categories of the 9 syllables in this example are: reference, reference, long, reference, long, long, long, long, short. Thus ; corresponding number of sample points .

[0078] The pitch markings are given by the segment-level trend and local contour: P1 is descending, P2 is ascending; [xiadan], [mashang] are locally ascending, [canyu] and the final [ma] follow ascending and terminating ascending. To maintain cross-syllable continuity, take the starting fundamental frequency of the first syllable . For a "descending" syllable, set . For an "ascending" syllable, set .

[0079] For a "flat" syllable, set . For a "terminating ascending" syllable, apply the above formula only in the second half, and it is flat in the first half. For a "locally ascending", apply the above formula only in the interval, and it is flat at the beginning and end. A linear trajectory is adopted within each syllable .

[0080] The starting fundamental frequency of adjacent syllables is connected to the terminating fundamental frequency of the previous syllable: .

[0081] Taking the calculation of the first three syllables as an example:

[0082] Syllable 1 [xian] is descending, , .

[0083] Syllable 2 [zai] is descending, , .

[0084] Syllable 3 [xia] is descending and locally ascending, . Its start is flat until , the middle ascends to , the end linearly returns to the descending end . The ascending syllables within P2 are calculated in the same way, and the terminating ascending is applied only in the second half of the final [ma] .

[0085] Map the accent strength, medium, and weak to gain. To avoid the overlap of adjacent strong accent peaks, if the interval between adjacent strong accents is less than one syllable, the latter one will be replaced with a different accent. The envelope uses a linear rise and fall, with the crossfade duration set as a parameter. (Unit: seconds), in this example, we take... Monosyllabic envelopes are located before and after each syllable. Perform linear fade-in and linear fade-out.

[0086] For each syllable, select the main vowel and take three sets of formant parameters. With bandwidth This example uses the following approximation (in Hertz):

[0087] [a]: ;

[0088] [i]: ;

[0089] [u]: ;

[0090] [y]: .

[0091] For compound vowels (such as [ian], [ang]), the main vowel [a] is used as an approximation. When the stress is strong, the bandwidth of the first formant is contracted. : .

[0092] The sound section uses periodic pulse excitation. Let the instantaneous fundamental frequency be... The instantaneous fundamental period is Using phase parameters (No. The normalized phase of each sample (dimensionless) is recursively used to generate pulses: when Output unit pulse and command .

[0093] The unvoiced passage uses zero-mean broadband noise, with amplitude and Synchronous control. Let the excitation be... Three sets of second-order resonators connected in series Resulting in audio channel output Apply envelope With gain Syllable output is .

[0094] Between adjacent syllables, the length is Cross-depreciation zone according to Calculate the root mean square parameter for the entire waveform. And scaled to the target loudness parameter (In this example, take 's full width), and the scaling factor is .

[0095] Check again at the end of the sentence [ma] for the termination upward inflection: If the termination frequency in the second half is not , then recalculate according to this formula. Finally, write it in the mono pulse code modulation format, with a sampling rate of , and a quantization bit number of 16 bits.

[0096] Take the syllable sequence and its tags as an example: [xian] descends, weak, reference → , , , , . [zai] descends, weak, reference → , , , , . [xia] descends and has a local upward inflection, medium, long → , , , and the fundamental frequency is divided into three segments: at the beginning remains , in the middle rises to , and at the end descends to . [dan] descends, weak, reference → , , , , . [ma] rises, medium, long → , , , , . [shang] rises, weak, long → , , , , . [can] rises, strong, long → , , , , ; at the same time, narrow the bandwidth of the first formant according to . [yu] rises, medium, long → , , , , , and take the main vowel [y] as . [ma] terminates with an upward inflection, weak, short → , , , The first half is straight, and the second half rises to... .

[0097] If any syllable lacks a tone mark, accent mark, or duration mark, it reverts to level, weak, and baseline respectively. If the interval between adjacent strong accents is less than one syllable, the latter is treated as the baseline. And extend the crossfading zone of the syllable to ,correspond .

[0098] Figure 2 This paper details the implementation process of echo smoothing at phrase segment connections and stress spacing constraints in this invention. The experimental curve, with time on the x-axis and stress intensity on the y-axis, clearly depicts the complete trajectory from the original stress pattern to the smoothed version. The experimental data reveals a significant discontinuity in the original stress pattern, particularly abrupt changes at phrase segment connections. At connection point 1 (0.8-1.2 seconds) and connection point 2 (2.3-2.7 seconds), the system detects sharp jumps in stress intensity between adjacent phrase segments, transitioning directly from strong to weak stress. This discontinuity leads to noticeable splicing artifacts in the synthesized speech. The echo smoothing algorithm effectively addresses this issue by implementing a fade-in / fade-out mechanism at the connection points. Specifically, at connection point 1, the system implements a fade-in transition within the 220-260 millisecond range, gradually attenuating the strong stress level (level 3) to the medium stress level (level 2), and then implements a fade-out transition within the 260-300 millisecond range, smoothly transitioning to the stress level of the next segment. The stress spacing constraint mechanism also demonstrates significant effectiveness. The experiment identified two strong accent positions: strong accent 1 at 1.0 second and strong accent 2 originally planned for 2.0 seconds, with a 1.0-second interval between them. Since this interval was less than the preset minimum constraint threshold of 2.0 seconds, the system automatically triggered a constraint processing mechanism, downgrading the later-appearing strong accent 2 from its original strong accent level to a medium accent level, reducing the accent intensity from level 6 to level 4. This adjustment not only met the minimum interval constraint requirement but also ensured the naturalness of the overall rhythm and the comfort of the listening experience. After echo smoothing and accent interval constraint processing, the final generated accent curve exhibited a smooth continuity, effectively eliminating the unnaturalness of segment splicing.

[0099] Figure 3Fully demonstrates the complete structure and technical features of the target prosody control sequence finally output by the present invention. This experimental curve adopts a hierarchical display method, sequentially showing three control dimensions of tone marks, stress marks, and duration marks from top to bottom, and providing a typical triple structure example at the bottom. The tone mark layer shows the final tone contour after mirror backfilling,同位 alignment, and conflict adjudication. It can be seen from the curve changes that the syllable "jīn" at the beginning of the sequence is marked as flat, and then syllables such as "tiān", "tiān", "qì", "zhēn" show a continuous upward trend, reflecting the effect of the mirror backfilling algorithm spreading outward from the nuclear phrase. At the positions of "bù" and "cuò", the tone starts to decline after reaching the peak, which exactly matches the semantic features of the negative word and the predicate head word recognized by the dependency template. The tones of the three syllables "zěn", "me", and "yàng" at the end of the sentence show a downward trend, conforming to the typical prosody pattern of interrogative sentences. The stress mark layer intuitively shows the stress level distribution of each syllable in the form of a bar chart. The nuclear phrase "zhēn" is given the highest strong stress mark, and its adjacent syllables "tiān", "qì", "bù", "cuò" are sequentially marked as strong stress and medium stress, reflecting the gradually decreasing feature of mirror propagation. The pronoun "nǐ" and the interrogative phrase "zěn me yàng" are both marked as weak stress, conforming to the fixed stress rules for function words and the end-of-sentence transition area. The duration mark layer shows the duration control strategy allocated according to word classes. Content words such as the noun "weather", the adverb "really", the adjective "good", and the verb "feel" are all marked as long durations, while the pronoun "you" is marked as a short duration, reflecting the differential treatment of the semantic importance of word classes. After elastic alignment processing, each duration mark precisely corresponds to the beat grid of 0.2 seconds, 0.6 seconds, and 1.0 seconds, ensuring the rhythm of the overall prosody. The finally generated triple sequence such as (flat, medium, long), (upward, strong, baseline), etc. provides complete and accurate prosody control parameters for the subsequent formant synthesizer.

[0100] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A hierarchical rhythm mapping driven short video script tone automatic adjustment method, characterized in that, The method comprises: Step 1: receiving a text string and a language category identifier, performing character normalization, punctuation-based sentence boundary division, word table and rule-based word segmentation, and dependency relationship template identification, generating a hierarchical index table containing four levels of sentences, phrases, words, and syllables, and a semantic anchor point table, and establishing placeholder columns for carrying tone marks, stress marks, and duration marks at each level; Step 2: based on the hierarchical index table, each sentence is divided into phrase segments, the boundaries are frozen in phrase segment units, the end-of-sentence termination style is preset according to punctuation, the core phrase is determined according to the semantic anchor point and its direction is initialized, the tone and stress are propagated between adjacent segments through mirror backfilling, the stress is raised and the local contour is set at the anchor point position through homologue alignment, the duration is allocated to the word and its syllable according to the part of speech, the mirror and homologue differences are resolved according to the preset priority, the echo smoothing and stress interval constraint are implemented at the segment connection, and the time sequence elastic alignment and hierarchical backfilling are performed, finally outputting a triple sequence composed of tone marks, stress marks and duration marks covering all syllables as the target prosody control sequence; Step 3: based on the target prosody control sequence, audio generation is performed to obtain new dubbing.

2. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 1, wherein, Step 3 specifically comprises: inputting the target prosody control sequence into a rule-based formant synthesizer, synthesizing speech according to the order of syllables from the formant frequency table and the excitation waveform generator, and applying fluctuations, emphasis and stretching according to the target prosody control sequence in the tone control, stress control and duration control channels, thereby generating new dubbing.

3. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 2, wherein, Step 1 specifically comprises: receiving a text string and a language category identifier, performing character normalization and punctuation unification; performing sentence boundary division according to period, question mark, exclamation mark and line feed and retaining original punctuation; performing word segmentation on the text within the sentence using a word table and rule-based word segmenter, the word table being a limited set containing functional words, conjunctions, negative words, interrogative words, quantity words and mood words; identifying the predicate center word and the modifier component using a rule set based on the dependency relationship template, and generating a semantic anchor point table accordingly, the semantic anchor point table recording the positions of interrogative words, negative words, modal words, quantity words, conjunctions and predicate center words in the word segmentation sequence; dividing phrase segments within the sentence based on sentence boundaries, conjunctions and pauses to obtain a hierarchical index table, the hierarchical index table including sentence index, phrase index and word index, and expanding into syllable index through a rule converter from character to pronunciation; establishing placeholder columns for sentences, phrases, words and syllables at each level according to the hierarchical index table, the placeholder column being a sequential table structure containing tone marks, stress marks and duration marks, the initial values of the tone marks and stress marks being set as undefined respectively, and the initial value of the duration mark being set as a baseline, and setting boundary markers at both ends of each phrase segment.

4. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 3, wherein, Step 2 specifically comprises: Step 2.1: based on the hierarchical index table, each sentence is divided into phrase segments, and the start and end positions of each phrase segment are set as frozen boundaries in the phrase level placeholder column; the end-of-sentence termination style is preset according to punctuation, and transition zones are set on both sides of the last syllable of the sentence for subsequent smoothing processing; Step 2.2: Determine the kernel phrase in each sentence according to the semantic anchor table, and initialize the tone mark and stress mark in the placeholder column of the kernel phrase; propagate the tone mark and stress mark between adjacent phrase segments by mirror backfilling with reference to the kernel phrase, and the propagation process maintains relative position symmetry and cannot cross the frozen boundary; perform homophonic alignment at the semantic anchor position, promote the stress mark at the corresponding word level position and set the local tone contour, and write the respective placeholder column at the phrase, word and syllable levels; Step 2.3: Assign duration marks at the word level according to the part of speech and expand them to the syllable level, ensuring that the cumulative duration marks within the phrase segment are consistent with the phrase level; when differences occur between mirror backfilling, homophonic alignment and sentence end termination patterns at the same position, perform conflict resolution according to the preset priority, cover the conflict position, and the priority is that the sentence end termination pattern is prior to the mirror backfilling and the mirror backfilling is prior to the homophonic alignment; implement echo smoothing at the connection of adjacent phrase segments, set fade-in and fade-out in the splicing area, and constrain the minimum distance between adjacent strong stresses, and if the constraint is not met, the later-occurring strong stress is downgraded to the middle level; Step 2.4: Perform time series elastic alignment on all duration marks, align short, reference and long to the corresponding beat grid respectively, and the sentence end termination pattern remains unchanged; write the tone mark and stress mark determined at the phrase level back to the word level and the syllable level, and form a ternary sequence consisting of tone mark, stress mark and duration mark covering all syllables as the target prosody control sequence.

5. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 4, wherein, In step 2.2, when propagating the tone mark and stress mark between adjacent phrase segments by mirror backfilling with reference to the kernel phrase, locate the phrase segment where the kernel phrase is located in the hierarchical index table, denoted as the mirror reference segment; read the tone mark field and stress mark field in the phrase level placeholder column of the mirror reference segment as the initial propagation source; determine the ordered list of the left and right phrase segments with the mirror reference segment as the mirror center, and the list order is arranged in turn according to the adjacent relationship with the mirror reference segment; check in the phrase level placeholder column whether there is a frozen boundary marker between the adjacent relationships; propagation is only performed in continuous segments that are not separated by frozen boundaries; when encountering any frozen boundary marker, the propagation immediately stops at the boundary and cannot continue to set the mark of the subsequent segment across the boundary; when the target segment already has a determined tone mark field and stress mark field in the previous step, keep the existing values and do not be covered by mirror backfilling.

6. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 5, wherein, In Step 2.2, when the pitch mark field of the mirror-reference segment is up, the pitch mark field of the first right adjacent segment is set to up, the pitch mark field of the first left adjacent segment is set to down, the second right adjacent segment continues to be set to up, the second left adjacent segment continues to be set to down, and so on until the freezing boundary or the end of the list is reached; when the pitch mark field of the mirror-reference segment is down, the pitch mark field of each level of the right adjacent segments is set to down in turn, the pitch mark field of each level of the left adjacent segments is set to up in turn, and so on until the stop condition is reached; when the pitch mark field of the mirror-reference segment is flat, the pitch mark field of each level of the adjacent segments on both sides is set to flat, and so on until the stop condition is reached; when the mirror-reference segment is adjacent to the transition zone of the end-of-sentence termination pattern, the phrase segment located in the transition zone is not changed by the mirror backfill, and the preset result of the end-of-sentence termination pattern in Step 2.1 is retained.

7. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 6, wherein, In Step 2.2, when the stress mark field of the mirror-reference segment is strong, the stress mark fields of the first right and left adjacent segments are both set to medium, the stress mark fields of the second right and left adjacent segments are both set to weak, and the stress mark fields of the third right and left adjacent segments remain undefined or weak, depending on whether the segments mainly undertake the function of conjunctions or function words; when the stress mark field of the mirror-reference segment is medium, the stress mark fields of the first left and right adjacent segments are set to weak, and the stress mark fields of the second left and right adjacent segments remain undefined or weak; when the stress mark field of the mirror-reference segment is weak, higher-level stress mark fields are not propagated outward, and the stress mark fields of the adjacent segments remain undefined or weak; if the adjacent segment is located in the transition zone of the end-of-sentence termination pattern, the stress mark field of the segment is fixed to weak.

8. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 7, wherein, In Step 2.3, the phrase segments of the current sentence are located in the hierarchical index table, and the corresponding word-level placeholder columns are read; a duration mark field is assigned to each word according to its part of speech: nouns, verbs, adjectives, proper nouns, and substantive adverbs are marked as long; function words, prepositions, conjunctions, and mood words are marked as short; numerical strings and time expressions are marked as baseline; if there is no marked long substantive word in the same phrase segment and the phrase-level duration mark field of the phrase segment is long, the substantive word closest to the center of the phrase segment is promoted from baseline to long; if there are too many marked long words in the same phrase segment and the phrase-level duration mark field of the phrase segment is short, the excess longs are sequentially reduced from right to left to baseline until only one long or the phrase-level duration mark field is retained; The number distribution of the word-level duration mark field in each phrase segment is counted, and if the most numerous category is inconsistent with the phrase-level duration mark field of the phrase segment, the duration mark field of the word closest to the end of the phrase segment is replaced with the same category as the phrase-level duration mark field according to the principle of minimum replacement number until the majority of categories are consistent.

9. The hierarchical rhythm mapping driven short video script tone automatic adjustment method of claim 8, wherein, In step 2.3, when the word-level duration marking field is long, the syllable carrying the main vowel is marked as long, and the rest of the syllables are marked as reference; when the word-level duration marking field is reference, all the syllables are marked as reference; when the word-level duration marking field is short, all the syllables are marked as short; in a combination word containing a clear consonant and a dark consonant, the first and last syllables are ensured not to be long at the same time; if the first and last syllables are long at the same time, only the syllable carrying the main vowel is kept long, and the other end is changed to reference; the number distribution of the syllable-level duration marking field in each phrase segment is counted, if the most numerous category is inconsistent with the phrase-level duration marking field, then in the phrase segment, from right to left, a number of references are changed to the category consistent with the phrase-level duration marking field until the majority of the categories are consistent.

Citation Information

Patent Citations

  • Multilingual speech synthesis method, device and system

    CN113160792A

  • Multi-language cross-culture communication auxiliary method and system based on large model

    CN120636412A