Blending Recorded Speech with TTS Output for Domain-Specific Prosody

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Combining high-quality recorded speech with text-to-speech (TTS) synthesizer output in various applications, such as navigation and voice-activated systems, is time-consuming and difficult due to the need for manual selection and alignment of specific prompts.

Innovation Solution

A text-to-speech engine that identifies the domain of input text and selects domain-specific recorded speech to blend with TTS output, using a domain detector and blending unit to match and refine the prosody of static phrases with synthesized speech, thereby smoothing the acoustic trajectory.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual selection and alignment of recorded speech prompts is performed, then high-quality domain-specific speech output is achieved, but the process becomes time-consuming and difficult

Engineering Contradiction:
Improvespeech qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system automatically identifies the domain from input text and selects appropriate recorded speech prompts without manual intervention. The domain detector and blending unit work autonomously to match recorded speech with TTS output, eliminating the need for manual selection and alignment while maintaining high speech quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameter of domain identification from manual to automatic based on text analysis. By detecting the domain from input text characteristics and using this information to select recorded speech prompts, the system achieves both high quality and time efficiency.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If manual selection and alignment of recorded speech prompts is performed, then accurate domain-specific speech is achieved, but the operation becomes difficult and complex

Engineering Contradiction:
Improvespeech accuracyVSAvoidoperation simplicity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system performs self-service by automatically detecting the domain from input text and selecting the appropriate recorded speech prompts. The blending unit automatically aligns and combines recorded speech with TTS output, eliminating complex manual operations while maintaining accurate domain-specific speech generation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system segments the speech generation process into distinct components: domain detection, recorded speech selection, and blending with TTS output. This segmentation allows each component to operate autonomously and simplifies the overall operation while maintaining precision.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If recorded speech is combined with TTS synthesizer output, then high-quality domain-specific speech is produced, but the process becomes time-consuming

Engineering Contradiction:
Improvespeech qualityVSAvoidprocessing efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system achieves self-service automation where the domain detector automatically identifies the domain from input text, selects appropriate recorded speech prompts, and the blending unit automatically combines them with TTS output. This eliminates manual processing steps while maintaining high speech quality, significantly improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-identifying the domain from input text before selecting and blending recorded speech. This advance domain detection enables efficient selection of appropriate recorded prompts, reducing overall processing time while maintaining quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8996377B2Blending recorded speech with text-to-speech output for specific domains
Publication Date: 2015.03.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8996377B2 patent drawing
  • US8996377B2 patent drawing
  • US8996377B2 patent drawing

AI summary

A text-to-speech (TTS) engine combines recorded speech with synthesized speech from a TTS synthesizer based on text input. The TTS engine receives the text input and identifies the domain for the speech (e.g. navigation, dialing, . . . ). The identified domain is used in selecting domain specific speech recordings (e.g. pre-recorded static phrases such as “turn left”, “turn right” . . . ) from the input text. The speech recordings are obtained based on the static phrases for the domain that are identified from the input text. The TTS engine blends the static phrases with the TTS output to smooth the acoustic trajectory of the input text. The prosody of the static phrases is used to create similar prosody in the TTS output.