AI Voice Sampling Apparatus for Natural Speech Style Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice user interfaces are limited in providing natural and interactive services as they cannot analyze a user's speech style or rhyme, resulting in a lack of user-friendly and natural dialogue, especially when dealing with unrecorded voices.

Innovation Solution

An AI-based voice sampling apparatus and method that utilizes a rhyme encoder to analyze vocal features from a user's voice sample, extracts an embedding vector, and applies the speech style to synthesized voice data through a processor and rhyme decoder, enabling the generation of natural and interactive voice responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio recording is used to implement voice response in a voice user interface, then the system can provide recorded voice responses, but it cannot provide responses for unrecorded voices and lacks flexibility

Engineering Contradiction:
Improveflexibility in providing voice responsesVSAvoidsystem complexity for handling unrecorded voices
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses text-to-speech synthesis to generate speech from text, creating a copy of human speech patterns without requiring actual human voice recording. This allows the system to provide voice responses for any text input, not just pre-recorded phrases, thereby improving adaptability while maintaining manageable system complexity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical approach of audio recording and playback with an AI-based text-to-speech synthesis system. This substitution enables the system to generate natural-sounding speech from text dynamically, providing flexibility for unrecorded voices without the constraints of pre-recorded content

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If the related art voice recognition system generates only a response signal without analyzing rhyme or speech style, then the system can provide basic voice responses, but it outputs text-type voice lacking naturalness and user-friendliness

Engineering Contradiction:
Improveanalysis precision of vocal featuresVSAvoidcomplexity of vocal feature analysis system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech synthesis process into distinct components: a rhyme encoder for analyzing vocal features, a text encoder for processing text input, and a decoder for generating synthesized speech. This segmentation allows the system to perform comprehensive vocal feature analysis while organizing the complexity into manageable, specialized modules

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a multi-functional system where the rhyme encoder analyzes various vocal features (pitch, rhythm, timbre), the text encoder processes text input, and the integrated system generates natural-sounding speech with appropriate speech style. This universal approach handles multiple aspects of speech synthesis within a unified framework, improving naturalness without proportionally increasing overall system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11107456B2Artificial intelligence (AI)-based voice sampling apparatus and method for providing speech style
Publication Date: 2021.08.31 LG ELECTRONICS INC
  • US11107456B2 patent drawing
  • US11107456B2 patent drawing
  • US11107456B2 patent drawing

AI summary

Discussed is an artificial intelligence (AI)-based voice sampling apparatus for providing a speech style, including a rhyme encoder configured to receive a user's voice, extract a voice sample, and analyze a vocal feature included in the voice sample, a text encoder configured to receive text for reflecting the vocal feature, a processor configured to classify the vocal feature of the voice sample input to the rhyme encoder according to a label, extract an embedding vector representing the vocal feature from the label, and generate a speech style from the embedding vector and apply the generated speech style to the text, and a rhyme decoder configured to output synthesized voice data in which the speech style is applied to the text by the processor.