Text-to-Speech Normalization for Peculiar Expression Mood Reproduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech devices fail to accurately analyze and reproduce the intended mood in texts containing 'leet-speak' or 'peculiar expressions,' leading to irrational readings.

Innovation Solution

A text-to-speech device with a normalization process that receives input text, applies normalization rules to convert peculiar expressions into normal forms, selects the most plausible normalized text through language processing, and modifies phonetic parameters to reflect the intended expression styles, enabling accurate mood representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional text-to-speech devices perform speech synthesis on texts containing peculiar expressions, then the synthesis process can be completed, but the intended mood cannot be reproduced and the reading becomes irrational

Engineering Contradiction:
Improveaccuracy of mood reproductionVSAvoidability to handle peculiar expressions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces a normalization rule database as an intermediary component between the input text and the speech synthesis engine. This database stores normalization rules that map peculiar expressions to their standard forms, enabling the system to correctly interpret and synthesize texts with peculiar expressions while preserving the intended mood and meaning

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary normalization processing before speech synthesis by converting peculiar expressions into standard expressions using pre-defined normalization rules. This preliminary action ensures that the subsequent speech synthesis can accurately reproduce the intended mood without being confused by non-standard expressions

Inventive Principle:
Principle #10Preliminary action

2Reliability

If normalization rules are added to handle peculiar expressions, then accuracy of mood reproduction improves, but system complexity increases

Engineering Contradiction:
Improveaccuracy of expression analysisVSAvoidcomplexity of normalization process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent pre-defines normalization rules in a database that map peculiar expressions to standard expressions. These rules are prepared in advance and stored for efficient retrieval during text processing, avoiding the need for complex real-time analysis and reducing computational complexity while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The normalization rule database serves itself by automatically matching peculiar expressions in the input text with corresponding normalization rules. The system retrieves applicable rules based on the input text characteristics without requiring complex external intervention or manual configuration, simplifying the overall system architecture

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9570067B2Text-to-speech system, text-to-speech method, and computer program product for synthesis modification based upon peculiar expressions
Publication Date: 2017.02.14 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9570067B2 patent drawing
  • US9570067B2 patent drawing
  • US9570067B2 patent drawing

AI summary

According to an embodiment, a text-to-speech device includes a receiver to receive an input text containing a peculiar expression; a normalizer to normalize the input text based on a normalization rule in which the peculiar expression, a normal expression of the peculiar expression, and an expression style of the peculiar expression are associated, to generate normalized texts; a selector to perform language processing of each normalized text, and select a normalized text based on result of the language processing; a generator generate a series of phonetic parameters representing phonetic expression of the selected normalized text; a modifier modifies a phonetic parameter in the normalized text corresponding to the peculiar expression in the input text based on a phonetic parameter modification method according to the normalization rule of the peculiar expression; and a output unit to output a phonetic sound synthesized using the series of phonetic parameters including the modified phonetic parameter.