Dynamic TTS Output Adaptation for Speech-Controlled Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech processing systems face challenges in providing a dynamic and user-centric experience, as they often output all default TTS content and speech at a constant speed, which may not align with the user's current situation or preferences, potentially leading to an undesirable experience, especially when the user is in a hurry or has limited time.

Innovation Solution

The system dynamically alters TTS output based on the number of words spoken, speech characteristics, electronic calendar data, and geographic location, allowing for truncated responses, faster speech synthesis, and tailored content delivery to enhance user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the system outputs all default TTS content at a constant speed, then the information completeness is maintained, but the user experience deteriorates when the user is in a hurry or has limited time

Engineering Contradiction:
Improveinformation completenessVSAvoiduser context adaptability
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The TTS output is made dynamic by adjusting both the content length and speech speed based on real-time analysis of user speech characteristics. The system analyzes the rate of speech, number of words, and speech characteristics of the input command, then dynamically modifies the output accordingly - truncating content and increasing speed when the user speaks quickly, and providing complete content at normal speed when the user speaks slowly.

Inventive Principle:
Principle #15Dynamics

2Loss of information

If the system provides complete TTS content, then the information accuracy is improved, but the response time increases which is undesirable when the user is in a hurry

Engineering Contradiction:
Improveinformation accuracyVSAvoidresponse time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system changes multiple parameters of the TTS output based on user speech analysis: (1) content length parameter - truncating or providing complete responses; (2) speech speed parameter - adjusting the rate of speech synthesis. These parameter changes are determined by analyzing the input speech characteristics, allowing the system to optimize the balance between information accuracy and response time for each user interaction.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If the system uses constant speech speed for TTS output, then the system complexity is reduced, but the user experience deteriorates when users have different speech patterns and time constraints

Engineering Contradiction:
Improvesystem complexityVSAvoiduser experience
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The system implements feedback by analyzing the user's input speech characteristics (rate of speech, number of words, speech characteristics) and using this feedback to adjust the TTS output parameters. This closed-loop approach allows the system to adapt to individual user preferences and contexts, significantly improving user experience while maintaining manageable system complexity through automated analysis and adjustment.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10276149B1Dynamic text-to-speech output
Publication Date: 2019.04.30 AMAZON TECH INC
  • US10276149B1 patent drawing
  • US10276149B1 patent drawing
  • US10276149B1 patent drawing

AI summary

Systems, methods, and devices for dynamically outputting TTS content are disclosed. A speech-controlled device captures a spoken command, and sends audio data corresponding thereto to a server(s). The server(s) determines output content responsive to the spoken command. The server(s) may also determine a user that spoke the command and determine an average speech characteristic (e.g., tone, pitch, speed, number of words, etc.) used by the user when speaking commands. The server(s) may also determine a speech characteristic of the presently spoken command, as well as determine a difference between the speech characteristic of the presently spoken command and the average speech characteristic of the user. The server(s) may then cause the speech-controlled device to output audio based on the difference.