Text-to-Speech Normalization for Peculiar Expression Mood Reproduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-to-speech devices fail to accurately analyze and reproduce the intended mood in texts containing 'leet-speak' or 'peculiar expressions,' leading to irrational readings.
Innovation Solution
A text-to-speech device with a normalization process that receives input text, applies normalization rules to convert peculiar expressions into normal forms, selects the most plausible normalized text through language processing, and modifies phonetic parameters to reflect the intended expression styles, enabling accurate mood representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional text-to-speech devices perform speech synthesis on texts containing peculiar expressions, then the synthesis process can be completed, but the intended mood cannot be reproduced and the reading becomes irrational
Solution Approach 1:
The patent introduces a normalization rule database as an intermediary component between the input text and the speech synthesis engine. This database stores normalization rules that map peculiar expressions to their standard forms, enabling the system to correctly interpret and synthesize texts with peculiar expressions while preserving the intended mood and meaning
Solution Approach 2:
The system performs preliminary normalization processing before speech synthesis by converting peculiar expressions into standard expressions using pre-defined normalization rules. This preliminary action ensures that the subsequent speech synthesis can accurately reproduce the intended mood without being confused by non-standard expressions
2Reliability
If normalization rules are added to handle peculiar expressions, then accuracy of mood reproduction improves, but system complexity increases
Solution Approach 1:
The patent pre-defines normalization rules in a database that map peculiar expressions to standard expressions. These rules are prepared in advance and stored for efficient retrieval during text processing, avoiding the need for complex real-time analysis and reducing computational complexity while maintaining high accuracy
Solution Approach 2:
The normalization rule database serves itself by automatically matching peculiar expressions in the input text with corresponding normalization rules. The system retrieves applicable rules based on the input text characteristics without requiring complex external intervention or manual configuration, simplifying the overall system architecture
Data Source
AI summary
According to an embodiment, a text-to-speech device includes a receiver to receive an input text containing a peculiar expression; a normalizer to normalize the input text based on a normalization rule in which the peculiar expression, a normal expression of the peculiar expression, and an expression style of the peculiar expression are associated, to generate normalized texts; a selector to perform language processing of each normalized text, and select a normalized text based on result of the language processing; a generator generate a series of phonetic parameters representing phonetic expression of the selected normalized text; a modifier modifies a phonetic parameter in the normalized text corresponding to the peculiar expression in the input text based on a phonetic parameter modification method according to the normalization rule of the peculiar expression; and a output unit to output a phonetic sound synthesized using the series of phonetic parameters including the modified phonetic parameter.


