Dynamic Vocal Characteristic Sets for Speech Output Emphasis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice interfaces for computing devices often generate flat or monotone speech output, which can lead to reduced attentiveness and information retention, especially in 'eyes-busy' and 'hands-busy' environments, as they fail to effectively emphasize relevant information in search results or responses.

Innovation Solution

The method involves generating spoken output using a plurality of vocal characteristic sets, where salient portions of the text are emphasized with distinct vocal characteristics, such as different voices or emphasis levels, to differentiate them from standard portions, thereby increasing attentiveness and persuasiveness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If monotone speech output is used for voice interfaces, then device complexity is reduced, but information retention and user attentiveness deteriorate

Engineering Contradiction:
Improvespeech generation complexityVSAvoidinformation retention
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies local quality by varying vocal characteristics (pitch, tone, emphasis) specifically for salient portions of speech output while maintaining standard characteristics for non-salient portions. This creates localized quality differences that highlight important information without requiring complete system complexity overhaul.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts vocal characteristics based on the importance of information being conveyed. By making the speech output dynamic rather than static/monotone, the system adapts its delivery to emphasize key information, improving retention without permanent complexity increase.

Inventive Principle:
Principle #15Dynamics

2Loss of information

If multiple vocal characteristic sets are used to emphasize salient portions, then information retention improves, but device complexity increases

Engineering Contradiction:
Improveinformation retentionVSAvoidspeech generation complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The speech output is segmented into salient and non-salient portions, with different vocal characteristics applied to each segment. This segmentation allows targeted enhancement of important information while keeping the overall system manageable through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes vocal parameters (pitch, tone, emphasis level) to differentiate salient from non-salient information. By manipulating existing parameters rather than introducing entirely new systems, the patent achieves enhanced information retention with controlled complexity increase.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If vocal emphasis is applied to salient portions, then user attentiveness improves, but processing time increases

Engineering Contradiction:
Improveuser attentivenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

Instead of applying uniform emphasis to all speech output, the system applies vocal emphasis partially—only to salient portions that require user attention. This selective approach improves attentiveness while minimizing the additional processing time required compared to comprehensive processing.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8856007B1Use text to speech techniques to improve understanding when announcing search results
Publication Date: 2014.10.07 GOOGLE LLC
  • US8856007B1 patent drawing
  • US8856007B1 patent drawing
  • US8856007B1 patent drawing

AI summary

Disclosed are apparatus and methods for generating synthesized utterances related to output of commands. A command is received at a computing device. A textual output for the command is determined using the computing device. A spoken output of the computing device is generated that utilizes a plurality of vocal characteristic sets. At least a portion of the spoken output corresponds to the textual output. At least a first part of the spoken output utilizes vocal characteristics of a first vocal characteristic set. At least a second part of the spoken output utilizes vocal characteristics of a second vocal characteristic set, where at least some of the vocal characteristics of the first vocal characteristic set differ from the vocal characteristics of the second vocal characteristic set.