AI Voice Sampling for Heterogeneous Speech Style Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis systems using AI struggle to provide natural and user-friendly dialogue, as they cannot analyze or replicate the user's speech style effectively, especially when dealing with diverse or shared speech styles, leading to limitations in providing intuitive voice responses.
Innovation Solution
An AI-based voice sampling apparatus and method that employs a rhyme encoder to analyze vocal features, a text encoder to reflect these features, and a processor to classify and generate a speech style, applying it to text through a rhyme decoder, thereby producing synthesized voice data that mimics the user's speech style.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice samples from different speech styles are shared or crossed, then the system can handle diverse speech styles, but error occurs in selecting representative samples
Solution Approach 1:
The patent segments voice samples into distinct speech style categories using style labels, and processes each segment separately through dedicated encoders. This segmentation allows the system to handle diverse speech styles while maintaining accurate representative sample selection within each style category, resolving the contradiction between versatility and reliability.
Solution Approach 2:
The patent introduces style embedding vectors as intermediary representations that bridge different speech styles. These embedding vectors serve as mediators that capture the essential characteristics of each speech style, enabling the system to accurately select representative samples across diverse styles without direct comparison errors.
2Ease of operation
If the system generates only response signals without analyzing rhyme or speech style, then the system is simpler, but natural interactive service cannot be provided
Solution Approach 1:
The patent performs preliminary analysis of speech style and rhyme characteristics before generating response signals. By extracting style embedding vectors and analyzing vocal features in advance, the system prepares natural speech style representations that are then applied to generate responses, enabling natural interactive service while maintaining operational simplicity through automated preprocessing.
Solution Approach 2:
The patent copies and replicates the user's speech style characteristics through embedding vectors and applies them to the synthesized response. This copying mechanism allows the system to mimic natural speech patterns without complex manual intervention, providing natural interactive service while keeping the system operation simple through automated style transfer.
Data Source
AI summary
Disclosed is an artificial intelligence (AI)-based voice sampling apparatus for providing a speech style in a heterogeneous label, including a rhyme encoder configured to receive a user's voice, extract a voice sample, and analyze a vocal feature included in the voice sample, a text encoder configured to receive text for reflecting the vocal feature, a processor configured to classify the voice sample input to the rhythm encoder into a label according to the vocal feature, provide a weight by measuring a distance between a voice sample corresponding to the label and a voice sample corresponding to a heterogeneous label as a label other than the label and provide a weight by measuring similarity between the label and the heterogeneous label, extract an embedding vector representing the vocal feature, generate a speech style from the embedding vector, and apply the generated speech style to the text, and a rhyme decoder configured to output synthesized voice data in which the speech style is applied to the text by the processor.


