Self-Supervised Voice Synthesis Using Segmented Neural Modules
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice synthesis technologies, including Text to Speech (TTS) and singing voice synthesis, struggle to generate natural and coherent voice signals due to the lack of representation of intonation and prosody, and require large datasets for training, limiting their ability to produce voices that mimic specific singers or styles.
Innovation Solution
A self-supervised learning-based voice synthesis method and apparatus that trains voice analysis and synthesis modules to extract and generate voice features, such as fundamental frequency, periodic and aperiodic amplitudes, linguistic, and timbre features, allowing for the synthesis of voices that closely resemble actual voices without the need for extensive training data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional parametric TTS method using artificial neural networks is used to generate natural voice signals with intonation and prosody, then voice naturalness is improved, but the quantity of training data required increases significantly
Solution Approach 1:
The patent segments the voice synthesis task into two distinct modules: a voice analysis module that extracts voice features from input voice signals, and a voice synthesis module that generates target voice signals from those features. This segmentation allows each module to be trained independently with smaller datasets, avoiding the need for large-scale joint training while maintaining naturalness.
Solution Approach 2:
The patent introduces voice features as an intermediary between the input voice signal and the synthesized target voice signal. The voice analysis module extracts these features (including fundamental frequency, periodic amplitude, aperiodic amplitude, and timbre), which then serve as input to the synthesis module. This intermediary representation enables more efficient learning with reduced data requirements.
2Device complexity
If concatenative TTS method is used to combine pre-recorded voice signals, then implementation complexity is reduced, but voice coherence and naturalness deteriorate due to lack of intonation and prosody representation
Solution Approach 1:
The patent replaces the mechanical concatenation of pre-recorded voice segments with a learning-based system. Instead of simply joining fixed phoneme or syllable recordings, the system uses trained neural network modules to analyze and synthesize voice signals, dynamically generating intonation and prosody patterns that maintain coherence and naturalness.
Solution Approach 2:
The patent transforms the voice synthesis approach from fixed parameter concatenation to dynamic parameter generation. The voice analysis module extracts parameters such as fundamental frequency, periodic amplitude, and aperiodic amplitude, which the synthesis module then uses to generate target voice signals with natural intonation and prosody variations, rather than relying on pre-recorded segments with fixed characteristics.
3Quantity of substance
If self-supervised learning is used to train voice synthesis models, then training data requirements are reduced, but the ability to learn complex intonation and prosody patterns may be compromised
Solution Approach 1:
The patent implements a self-supervised learning framework where the voice synthesis module generates target voice signals that are fed back to the voice analysis module for feature extraction. This creates a closed-loop training system where the model learns to produce features that, when synthesized, reconstruct the original voice characteristics, enabling effective learning of intonation and prosody patterns even with limited external supervision.
Solution Approach 2:
The patent enables the voice synthesis system to train itself by using its own output as training input. The synthesized target voice signals generated by the synthesis module are used to further train the voice analysis module, creating a self-improving system that can develop complex intonation and prosody representation capabilities without requiring large amounts of externally labeled training data.
Data Source
AI summary
Disclosed herein are a self-supervised learning-based unified voice synthesis method and apparatus. The self-supervised learning-based unified voice synthesis method and apparatus: train a voice analysis module to output voice features for training voice signals by using the training voice signals representing training voices, and output voice features for the training voices; and train a voice synthesis module to synthesize voice signals from the voice features for the training voices by using the output voice features, and synthesize synthesized voice signals, representing synthesized voices, from the output voice features. The self-supervised learning-based unified voice synthesis method and apparatus can synthesize voices similar to actual voices by using artificial neural networks that are trained by themselves through self-supervised learning, without the need to train the artificial neural networks on a large quantity of voice and text datasets.


