Voice Synthesis Adaptation via Context Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice synthesis technologies lack the ability to provide personalized and engaging voice interactions, as they typically produce monotonous voices that fail to build familiarity or capture user attention due to uniform voice quality and tone.

Innovation Solution

A learning device and method that utilizes voice recognition, image recognition, and context estimation to generate voice synthesis data tailored to specific users, taking into account their emotions, noise levels, and relationships, allowing for dynamic voice synthesis that mimics the voice quality and tone of individual users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If voice synthesis uses uniform voice quality and tone for all users, then the system is simple and easy to implement, but user familiarity and attention are not built

Engineering Contradiction:
Improvevoice personalizationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments voice synthesis into multiple distinct voice types (e.g., friendly, professional, energetic) and assigns them to different users based on their preferences and contexts. This allows personalized voice experiences without requiring a completely new synthesis system for each user, thus improving adaptability while controlling complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic voice selection that adapts to changing contexts, user states, and interactions. The system can switch between different voice qualities and tones based on real-time conditions, enabling personalized adaptation while using a finite set of pre-defined voice characteristics to manage system complexity.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If voice synthesis adapts to individual user contexts and emotions, then user engagement and understanding improve, but the processing complexity and data requirements increase

Engineering Contradiction:
Improveuser engagementVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent creates simplified representations or models of user preferences and contextual patterns based on observed interactions. These models allow the system to adapt to individual users by copying and applying learned patterns rather than performing complex real-time analysis, thus improving engagement while controlling processing complexity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent adjusts specific parameters of voice synthesis (such as pitch, speed, tone, and volume) based on user context and emotions rather than redesigning the entire synthesis process. This selective parameter adjustment enables personalized adaptation with manageable processing requirements.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If the system learns from multiple users' speech patterns and contexts, then voice personalization accuracy improves, but the learning time and computational resources increase

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidlearning time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary analysis and categorization of user speech patterns during initial interactions to establish baseline profiles quickly. This preliminary action allows the system to achieve functional personalization faster, with continued refinement occurring in the background without significantly increasing perceived learning time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements partial personalization by focusing on the most impactful speech characteristics and contextual factors first, achieving satisfactory recognition accuracy without analyzing every possible parameter. This selective approach reduces learning time while maintaining useful personalization levels.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11335322B2Learning device, learning method, voice synthesis device, and voice synthesis method
Publication Date: 2022.05.17 SONY GROUP CORP
  • US11335322B2 patent drawing
  • US11335322B2 patent drawing
  • US11335322B2 patent drawing

AI summary

The present technology relates to a learning device, a learning method, a voice synthesis device, and a voice synthesis method configured so that information can be provided via voice allowing easy understanding of contents by a user as a speech destination. A learning device according to one embodiment of the present technology performs voice recognition of speech voice of a plurality of users, estimates statuses when a speech is made, and learns, on the basis of speech voice data, a voice recognition result, and the statuses when the speech is made, voice synthesis data to be used for generation of synthesized voice according to statuses upon voice synthesis. Moreover, a voice synthesis device estimates statuses, and uses the voice synthesis data to generate synthesized voice indicating the contents of predetermined text data and obtained according to the estimated statuses. The present technology can be applied to an agent device.