AI Speech Style Synthesis via Vector Embedding and Sparse Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies struggle to replicate various speech styles like humans, limiting the diversity and accuracy of artificial intelligence devices in speech communication.

Innovation Solution

An artificial intelligence device and method that control speech style by generating a condition vector, reducing its dimension, acquiring a sparse code vector through sparse dictionary coding, and changing vector element values to synthesize speech with specific styles from text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If synthesized speech is generated using traditional methods, then speech production is achieved, but the ability to reproduce various speech styles like human speech is limited

Engineering Contradiction:
Improvespeech style diversityVSAvoidhuman-like speech reproduction accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent changes the parameter representation from raw audio signals to vector embeddings in a latent space. By representing speech styles as vectors and manipulating their elements, the system can continuously transform between different speech styles (e.g., from formal to casual, from male-like to female-like) while maintaining high fidelity to human speech characteristics.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a new dimensional space (vector embedding space) where speech styles are represented as points or regions. This allows the system to navigate between different speech styles by moving through this abstract space, enabling diverse speech style reproduction without requiring separate synthesis models for each style.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If more speech data is accumulated to improve recognition accuracy, then speech recognition accuracy approaches human parity, but the complexity of processing and managing large amounts of speech data increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential characteristics of speech styles by transforming large amounts of speech data into compact vector embeddings. This extraction process captures the core features needed for speech style reproduction while discarding redundant information, thereby reducing processing complexity while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by selectively modifying specific elements within the speech style vectors to achieve desired speech style transformations. Instead of processing entire speech signals or large datasets, the system makes targeted adjustments to specific vector elements that correspond to particular speech style characteristics.

Inventive Principle:
Principle #3Local quality

3Ease of operation

If traditional speech synthesis methods are used, then speech generation is achieved, but fine control over prosody and speech style variations is difficult

Engineering Contradiction:
Improvespeech style controlVSAvoidprosody control precision
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent segments speech style control into discrete vector elements, where each element or group of elements corresponds to specific speech style characteristics (e.g., pitch, rhythm, timbre). This segmentation allows independent control of different prosodic features by modifying specific vector elements, enabling precise and easy speech style manipulation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11721319B2Artificial intelligence device and method for generating speech having a different speech style
Publication Date: 2023.08.08 LG ELECTRONICS INC
  • US11721319B2 patent drawing
  • US11721319B2 patent drawing
  • US11721319B2 patent drawing

AI summary

An artificial intelligence device includes a memory and a processor. The memory is configured to store audio data having a predetermined speech style. The processor is configured to generate a condition vector relating to a condition for determining the speech style of the audio data, reduce a dimension of the condition vector to a predetermined reduction dimension, acquire a sparse code vector based on a dictionary vector acquired through sparse dictionary coding with respect to the condition vector having the predetermined reduction dimension, and change a vector element value included in the sparse code vector.