Dynamic TTS Rate Control via User Speaking Speed Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice assistants maintain a uniformly set text-to-speech (TTS) rate, failing to adapt to individual users' speaking rates, which can lead to suboptimal user experience and understanding, especially for users with different speaking habits or those unfamiliar with voice assistants.

Innovation Solution

An electronic device with a processor that receives a voice signal, calculates the speaking rate, and adjusts the TTS rate accordingly, allowing for dynamic conversion of text to speech based on the user's speaking rate, thereby improving the intimacy and clarity of the interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a uniformly set TTS rate is used for all users, then the system is simple and easy to implement, but the user experience deteriorates for users with different speaking rates

Engineering Contradiction:
Improvesimplicity of TTS rate settingVSAvoidadaptability to individual user speaking rates
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The TTS rate is changed from a static uniform value to a dynamic value that automatically adjusts based on the user's speaking rate. The system calculates the user's speaking rate in real-time and adapts the TTS rate accordingly, making the system flexible and adaptive to individual users while maintaining ease of operation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system introduces feedback by calculating the user's speaking rate from their voice input and using this information to adjust the TTS rate. This closed-loop feedback mechanism allows the system to adapt to user preferences automatically, improving user experience without requiring manual configuration.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If the TTS rate is adjusted to match each user's speaking rate, then user experience and understanding improve, but the system complexity increases

Engineering Contradiction:
Improveadaptability to user speaking ratesVSAvoidcomplexity of TTS rate control system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically calculating the user's speaking rate and adjusting the TTS rate without requiring external configuration or complex control mechanisms. The voice assistant itself gathers the necessary information from user interaction and makes the adjustments autonomously, minimizing system complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The speaking rate calculation mechanism serves multiple functions: it not only determines the TTS rate but also provides insights into user behavior patterns. This multi-functionality reduces the need for separate systems and components, thereby limiting the increase in overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If the system calculates and adapts TTS rate for each user, then communication effectiveness improves, but the processing time and computational resources increase

Engineering Contradiction:
Improvecommunication effectivenessVSAvoidtime for speaking rate calculation
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary action by calculating the user's speaking rate during the initial interaction phases and using this information for subsequent TTS operations. This approach avoids repeated calculations for each TTS operation, reducing computational overhead and time loss while maintaining communication effectiveness.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240071363A1Electronic device and method of controlling text-to-speech (TTS) rate
Publication Date: 2024.02.29 SAMSUNG ELECTRONICS CO LTD
  • US20240071363A1 patent drawing
  • US20240071363A1 patent drawing
  • US20240071363A1 patent drawing

AI summary

Disclosed are an electronic device and a method of controlling a text-to-speech (TTS) rate. An electronic device may include a processor, and a memory configured to store instructions to be executed by the processor. The processor may receive a voice signal of a user. The processor may calculate a speaking rate of the voice signal based on the voice signal. The processor may generate an output text to be output to the user based on the voice signal. The processor may determine a TTS rate of the output text based on the speaking rate. The processor may convert the output text into voice data based on the TTS rate and output the voice data.