Dynamic Voice Profile TTS for Personalized Speech Output

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional speech output systems in IoT devices and smart appliances use generic machine-generated voices that fail to accurately represent the voice of the message sender, leading to confusion and inefficiencies, particularly in distinguishing between messages from different users.

Innovation Solution

A dynamic text-to-speech system that uses machine learning to develop and employ voice profiles approximating the voice of the message sender, allowing for personalized speech output based on attributes like tone, pitch, and timbre, and adjusts speech output based on user feedback to enhance user experience and interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If pre-recorded speech segments from celebrities or other persons are used for speech output, then the speech output quality is improved, but the system cannot accommodate situations in which the text to be output is not predetermined

Engineering Contradiction:
Improvespeech output qualityVSAvoidaccommodation of non-predetermined text
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically selects and switches between different voice profiles based on the sender's identity and message context. The TTS engine can adaptively choose between pre-recorded segments for predetermined text and synthesized speech for non-predetermined text, making the system both high-quality and versatile.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The speech output system is designed to handle multiple types of input text (predetermined and non-predetermined) using a unified architecture that supports both pre-recorded segments and real-time text-to-speech synthesis, allowing it to accommodate various messaging scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If generic machine-generated voices are used for speech output, then the system complexity is reduced, but the ability to accurately represent the voice of the message sender is lost

Engineering Contradiction:
Improvesystem complexityVSAvoidsender voice identity
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The system creates voice profiles that are copies or approximations of the sender's actual voice characteristics. These profiles capture tone, pitch, and timbre attributes, allowing the TTS engine to generate speech that sounds like the sender without requiring the sender's physical presence or complex real-time processing.

Inventive Principle:
Principle #26Copying

3Ease of operation

If multiple voice profiles are maintained for different senders, then the personalization of speech output is improved, but the data storage requirements and processing overhead increase

Engineering Contradiction:
Improvepersonalization of speech outputVSAvoiddata storage requirements
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

Instead of storing complete voice recordings or large audio files for each sender, the system extracts and stores only the essential local characteristics (tone, pitch, timbre parameters) that define each voice profile. This allows for personalized speech output while minimizing storage requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11398218B1Dynamic speech output configuration
Publication Date: 2022.07.26 UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)
  • US11398218B1 patent drawing
  • US11398218B1 patent drawing
  • US11398218B1 patent drawing

AI summary

Techniques are described for providing dynamically configured speech output, through which text data from a message is presented as speech output through a text-to-speech (TTS) engine that employs a voice profile to provide a machine-generated voice that approximates that of the sender of the message. The sender can also indicate the type of voice they would prefer the TTS engine use to render their text to a recipient, and the voice to be used can be specified in a sender's user profile, as a preference or attribute of the sending user. In some examples, the voice profile to be used can be indicated as metadata included in the message. A voice profile can specify voice attributes such as the tone, pitch, register, timbre, pacing, gender, accent, and so forth. A voice profile can be generated through a machine learning (ML) process.