This invention discloses a text pronunciation optimization method and
system for
speech synthesis, belonging to the field of
speech synthesis management technology. In this method, after emotional
processing, long sentences are split into segments based on a triple
rule system of semantic blocks, classical Chinese function words, and character length. The English portion of the text undergoes layered
processing, distinguishing between pure uppercase letter combinations and regular English words, adding splitting and prosodic markers to pure uppercase letter combinations, and generating optimized text pronunciation information. This invention breaks through the limitations of existing single-rule
adaptation in
speech synthesis text processing, pioneering a multi-dimensional layered optimization framework that integrates technologies such as semantic
parsing of numbers and operators, scene-based matching of polyphonic characters, emotional markers for literary texts, and semantic long
sentence splitting. It designs multiple exclusive optimization rules for the speech synthesis and reading needs of literary and everyday spoken texts, solving the core pain point of existing technologies that emphasize generality but
neglect specific scenarios.