Artificial Speech Generation with Emotional Indicators

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current artificial speech generation techniques lack the ability to effectively incorporate emotional features, resulting in synthesized voices that often sound robotic and unrealistic, failing to convey the emotional content present in human speech.

Innovation Solution

An information processing device and method that obtain speech emotional indicators and associated timing data from speech data, combining them with text data to generate artificial speech data, using machine learning algorithms and multimodal AI models to embed emotional information, allowing for realistic voice synthesis and voice cloning across languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional text-to-speech conversion is used to generate artificial speech, then the speech content can be produced efficiently, but the emotional expression and naturalness are lost resulting in robotic speech

Engineering Contradiction:
Improveemotional expression accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the speech generation process into distinct modules: emotional indicator extraction from speech data, text data processing, and artificial speech generation. Each module handles specific aspects (emotional features, linguistic content, synthesis), allowing complex emotional expression to be achieved through coordinated simple components rather than a monolithic complex system

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds an emotional dimension to traditional text-to-speech by extracting and incorporating speech emotional indicators as a new feature dimension. This transforms the generation process from handling only textual information to processing both text and emotional features in parallel, enabling realistic emotional expression without proportionally increasing overall system complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If speech emotional indicators and timing data are extracted and integrated with text data to generate artificial speech, then emotional realism is improved, but processing complexity and computational requirements increase

Engineering Contradiction:
Improveemotional feature precisionVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary extraction of speech emotional indicators and their timing data from speech recordings before the actual artificial speech generation process. This pre-processing step separates emotional feature extraction from the generation process, allowing the main generation system to receive pre-packaged emotional data along with text, thereby reducing real-time processing complexity while maintaining high emotional feature precision

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces speech emotional indicators as an intermediary element that bridges the gap between raw speech data and artificial speech generation. These indicators serve as mediators that carry emotional information from the source speech to the generation process, enabling precise emotional feature transfer without requiring direct complex processing between all system components

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If emotional indicators are accurately matched with timestamps and integrated into speech generation, then naturalness and human-like quality improve, but data processing requirements increase

Engineering Contradiction:
Improvetiming accuracyVSAvoiddata processing volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential timing information (timestamps) associated with emotional indicators from the complete speech data, separating this critical temporal information from the full audio signal. This extraction approach focuses processing on necessary timing data rather than analyzing entire speech waveforms, achieving accurate temporal alignment with reduced data processing volume

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing qualities to different aspects of the data: high precision is applied to timing data and emotional indicator matching where accuracy is critical, while less intensive processing is applied to other aspects of speech generation. This localized quality approach ensures timing accuracy where needed without uniformly increasing processing requirements across all data

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240242703A1Information processing device and information processing method for artificial speech generation
Publication Date: 2024.07.18 SONY GROUP CORP
  • US20240242703A1 patent drawing
  • US20240242703A1 patent drawing
  • US20240242703A1 patent drawing

AI summary

An information processing device for generating artificial speech data has circuitry which is configured to obtain, based on speech data, speech emotional indicators and associated timing data of the emotional indicators; and to obtain, based on the speech data, text data; and to generate artificial speech data based on the text data, the speech emotional indicators and the associated timing data.