Artificial Speech Generation with Emotional Indicators
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current artificial speech generation techniques lack the ability to effectively incorporate emotional features, resulting in synthesized voices that often sound robotic and unrealistic, failing to convey the emotional content present in human speech.
Innovation Solution
An information processing device and method that obtain speech emotional indicators and associated timing data from speech data, combining them with text data to generate artificial speech data, using machine learning algorithms and multimodal AI models to embed emotional information, allowing for realistic voice synthesis and voice cloning across languages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional text-to-speech conversion is used to generate artificial speech, then the speech content can be produced efficiently, but the emotional expression and naturalness are lost resulting in robotic speech
Solution Approach 1:
The system segments the speech generation process into distinct modules: emotional indicator extraction from speech data, text data processing, and artificial speech generation. Each module handles specific aspects (emotional features, linguistic content, synthesis), allowing complex emotional expression to be achieved through coordinated simple components rather than a monolithic complex system
Solution Approach 2:
The patent adds an emotional dimension to traditional text-to-speech by extracting and incorporating speech emotional indicators as a new feature dimension. This transforms the generation process from handling only textual information to processing both text and emotional features in parallel, enabling realistic emotional expression without proportionally increasing overall system complexity
2Manufacturing precision
If speech emotional indicators and timing data are extracted and integrated with text data to generate artificial speech, then emotional realism is improved, but processing complexity and computational requirements increase
Solution Approach 1:
The system performs preliminary extraction of speech emotional indicators and their timing data from speech recordings before the actual artificial speech generation process. This pre-processing step separates emotional feature extraction from the generation process, allowing the main generation system to receive pre-packaged emotional data along with text, thereby reducing real-time processing complexity while maintaining high emotional feature precision
Solution Approach 2:
The patent introduces speech emotional indicators as an intermediary element that bridges the gap between raw speech data and artificial speech generation. These indicators serve as mediators that carry emotional information from the source speech to the generation process, enabling precise emotional feature transfer without requiring direct complex processing between all system components
3Measurement precision
If emotional indicators are accurately matched with timestamps and integrated into speech generation, then naturalness and human-like quality improve, but data processing requirements increase
Solution Approach 1:
The system extracts only the essential timing information (timestamps) associated with emotional indicators from the complete speech data, separating this critical temporal information from the full audio signal. This extraction approach focuses processing on necessary timing data rather than analyzing entire speech waveforms, achieving accurate temporal alignment with reduced data processing volume
Solution Approach 2:
The patent applies different processing qualities to different aspects of the data: high precision is applied to timing data and emotional indicator matching where accuracy is critical, while less intensive processing is applied to other aspects of speech generation. This localized quality approach ensures timing accuracy where needed without uniformly increasing processing requirements across all data
Data Source
AI summary
An information processing device for generating artificial speech data has circuitry which is configured to obtain, based on speech data, speech emotional indicators and associated timing data of the emotional indicators; and to obtain, based on the speech data, text data; and to generate artificial speech data based on the text data, the speech emotional indicators and the associated timing data.


