AI Speech Style Synthesis via Vector Embedding and Sparse Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies struggle to replicate various speech styles like humans, limiting the diversity and accuracy of artificial intelligence devices in speech communication.
Innovation Solution
An artificial intelligence device and method that control speech style by generating a condition vector, reducing its dimension, acquiring a sparse code vector through sparse dictionary coding, and changing vector element values to synthesize speech with specific styles from text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If synthesized speech is generated using traditional methods, then speech production is achieved, but the ability to reproduce various speech styles like human speech is limited
Solution Approach 1:
The patent changes the parameter representation from raw audio signals to vector embeddings in a latent space. By representing speech styles as vectors and manipulating their elements, the system can continuously transform between different speech styles (e.g., from formal to casual, from male-like to female-like) while maintaining high fidelity to human speech characteristics.
Solution Approach 2:
The patent introduces a new dimensional space (vector embedding space) where speech styles are represented as points or regions. This allows the system to navigate between different speech styles by moving through this abstract space, enabling diverse speech style reproduction without requiring separate synthesis models for each style.
2Measurement precision
If more speech data is accumulated to improve recognition accuracy, then speech recognition accuracy approaches human parity, but the complexity of processing and managing large amounts of speech data increases
Solution Approach 1:
The patent extracts the essential characteristics of speech styles by transforming large amounts of speech data into compact vector embeddings. This extraction process captures the core features needed for speech style reproduction while discarding redundant information, thereby reducing processing complexity while maintaining accuracy.
Solution Approach 2:
The patent applies local quality by selectively modifying specific elements within the speech style vectors to achieve desired speech style transformations. Instead of processing entire speech signals or large datasets, the system makes targeted adjustments to specific vector elements that correspond to particular speech style characteristics.
3Ease of operation
If traditional speech synthesis methods are used, then speech generation is achieved, but fine control over prosody and speech style variations is difficult
Solution Approach 1:
The patent segments speech style control into discrete vector elements, where each element or group of elements corresponds to specific speech style characteristics (e.g., pitch, rhythm, timbre). This segmentation allows independent control of different prosodic features by modifying specific vector elements, enabling precise and easy speech style manipulation.
Data Source
AI summary
An artificial intelligence device includes a memory and a processor. The memory is configured to store audio data having a predetermined speech style. The processor is configured to generate a condition vector relating to a condition for determining the speech style of the audio data, reduce a dimension of the condition vector to a predetermined reduction dimension, acquire a sparse code vector based on a dictionary vector acquired through sparse dictionary coding with respect to the condition vector having the predetermined reduction dimension, and change a vector element value included in the sparse code vector.


