AI Voice Sampling for Heterogeneous Speech Style Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis systems using AI struggle to provide natural and user-friendly dialogue, as they cannot analyze or replicate the user's speech style effectively, especially when dealing with diverse or shared speech styles, leading to limitations in providing intuitive voice responses.

Innovation Solution

An AI-based voice sampling apparatus and method that employs a rhyme encoder to analyze vocal features, a text encoder to reflect these features, and a processor to classify and generate a speech style, applying it to text through a rhyme decoder, thereby producing synthesized voice data that mimics the user's speech style.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If voice samples from different speech styles are shared or crossed, then the system can handle diverse speech styles, but error occurs in selecting representative samples

Engineering Contradiction:
Improvehandling diverse speech stylesVSAvoidaccuracy in selecting representative samples
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments voice samples into distinct speech style categories using style labels, and processes each segment separately through dedicated encoders. This segmentation allows the system to handle diverse speech styles while maintaining accurate representative sample selection within each style category, resolving the contradiction between versatility and reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces style embedding vectors as intermediary representations that bridge different speech styles. These embedding vectors serve as mediators that capture the essential characteristics of each speech style, enabling the system to accurately select representative samples across diverse styles without direct comparison errors.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the system generates only response signals without analyzing rhyme or speech style, then the system is simpler, but natural interactive service cannot be provided

Engineering Contradiction:
Improvesimplicity of system operationVSAvoidcapability to provide natural interactive service
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary analysis of speech style and rhyme characteristics before generating response signals. By extracting style embedding vectors and analyzing vocal features in advance, the system prepares natural speech style representations that are then applied to generate responses, enabling natural interactive service while maintaining operational simplicity through automated preprocessing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent copies and replicates the user's speech style characteristics through embedding vectors and applies them to the synthesized response. This copying mechanism allows the system to mimic natural speech patterns without complex manual intervention, providing natural interactive service while keeping the system operation simple through automated style transfer.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11056096B2Artificial intelligence (AI)-based voice sampling apparatus and method for providing speech style in heterogeneous label
Publication Date: 2021.07.06 LG ELECTRONICS INC
  • US11056096B2 patent drawing
  • US11056096B2 patent drawing
  • US11056096B2 patent drawing

AI summary

Disclosed is an artificial intelligence (AI)-based voice sampling apparatus for providing a speech style in a heterogeneous label, including a rhyme encoder configured to receive a user's voice, extract a voice sample, and analyze a vocal feature included in the voice sample, a text encoder configured to receive text for reflecting the vocal feature, a processor configured to classify the voice sample input to the rhythm encoder into a label according to the vocal feature, provide a weight by measuring a distance between a voice sample corresponding to the label and a voice sample corresponding to a heterogeneous label as a label other than the label and provide a weight by measuring similarity between the label and the heterogeneous label, extract an embedding vector representing the vocal feature, generate a speech style from the embedding vector, and apply the generated speech style to the text, and a rhyme decoder configured to output synthesized voice data in which the speech style is applied to the text by the processor.