Dynamic Voice Configuration for Natural Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech systems produce mechanical and unpleasing synthetic speech that lacks personality, as they fail to differentiate and accurately render spoken passages with distinct speaker identities and genders, resulting in a lack of natural inflection and emotion.

Innovation Solution

A method and apparatus for generating natural sounding synthetic speech by automatically identifying spoken and non-spoken passages within a text source, determining speaker identity and gender, and dynamically applying voice configurations to each portion of text based on these attributes, allowing for the selective application of voice configurations to create a more natural audio rendition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If text-to-speech systems use a single uniform voice configuration for all text passages, then the system is simple and fast to implement, but the resulting speech sounds mechanical and lacks personality

Engineering Contradiction:
Improvenaturalness of speechVSAvoidvoice configuration management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the text source into spoken passages and non-spoken passages, and further segments spoken passages by speaker identity. Different voice configurations are applied to each segment type, allowing natural speech rendering without requiring a single complex configuration for all text.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different voice configurations to different portions of the text based on local characteristics (spoken vs. non-spoken, different speakers). This local differentiation enables natural speech output while keeping the overall system manageable through context-specific configuration application.

Inventive Principle:
Principle #3Local quality

2Reliability

If text-to-speech systems differentiate voice configurations for different speakers and passages, then the speech becomes more natural and expressive, but the processing complexity and time increase

Engineering Contradiction:
Improveexpression and personality in speechVSAvoidtext processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary analysis of the text source to identify spoken passages and determine speaker identities before speech synthesis. This advance segmentation and tagging enables efficient processing during synthesis, as the system already has organized information about which voice configurations to apply.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically selects voice configurations based on the analyzed text characteristics. The system adapts its voice output in real-time based on the identified speaker and passage type, enabling expressive speech while maintaining processing efficiency through rule-based dynamic selection.

Inventive Principle:
Principle #15Dynamics

3Reliability

If text-to-speech systems apply multiple voice configurations selectively, then the audio quality improves with personality and emotion, but the system complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidvoice configuration management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes voice configuration parameters (such as pitch, tone, and other acoustic properties) based on the identified speaker identity and passage type. This parameter differentiation enables high-quality audio with distinct speaker characteristics while managing complexity through systematic parameter selection based on text analysis results.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8326629B2Dynamically changing voice attributes during speech synthesis based upon parameter differentiation for dialog contexts
Publication Date: 2012.12.04 CERENCE OPERATING CO
  • US8326629B2 patent drawing
  • US8326629B2 patent drawing
  • US8326629B2 patent drawing

AI summary

A method of speech synthesis can include automatically identifying spoken passages and non-spoken passages within a text source and converting the text source to speech by applying different voice configurations to different portions of text within the text source according to whether each portion of text was identified as a spoken passage or a non-spoken passage. The method further can include identifying the speaker and/or the gender of the speaker and applying different voice configurations according to the speaker identity and/or speaker gender.