Self-Supervised Voice Synthesis Using Segmented Neural Modules

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice synthesis technologies, including Text to Speech (TTS) and singing voice synthesis, struggle to generate natural and coherent voice signals due to the lack of representation of intonation and prosody, and require large datasets for training, limiting their ability to produce voices that mimic specific singers or styles.

Innovation Solution

A self-supervised learning-based voice synthesis method and apparatus that trains voice analysis and synthesis modules to extract and generate voice features, such as fundamental frequency, periodic and aperiodic amplitudes, linguistic, and timbre features, allowing for the synthesis of voices that closely resemble actual voices without the need for extensive training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional parametric TTS method using artificial neural networks is used to generate natural voice signals with intonation and prosody, then voice naturalness is improved, but the quantity of training data required increases significantly

Engineering Contradiction:
Improvevoice naturalnessVSAvoidtraining data quantity
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments the voice synthesis task into two distinct modules: a voice analysis module that extracts voice features from input voice signals, and a voice synthesis module that generates target voice signals from those features. This segmentation allows each module to be trained independently with smaller datasets, avoiding the need for large-scale joint training while maintaining naturalness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces voice features as an intermediary between the input voice signal and the synthesized target voice signal. The voice analysis module extracts these features (including fundamental frequency, periodic amplitude, aperiodic amplitude, and timbre), which then serve as input to the synthesis module. This intermediary representation enables more efficient learning with reduced data requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If concatenative TTS method is used to combine pre-recorded voice signals, then implementation complexity is reduced, but voice coherence and naturalness deteriorate due to lack of intonation and prosody representation

Engineering Contradiction:
Improveimplementation complexityVSAvoidvoice coherence
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent replaces the mechanical concatenation of pre-recorded voice segments with a learning-based system. Instead of simply joining fixed phoneme or syllable recordings, the system uses trained neural network modules to analyze and synthesize voice signals, dynamically generating intonation and prosody patterns that maintain coherence and naturalness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the voice synthesis approach from fixed parameter concatenation to dynamic parameter generation. The voice analysis module extracts parameters such as fundamental frequency, periodic amplitude, and aperiodic amplitude, which the synthesis module then uses to generate target voice signals with natural intonation and prosody variations, rather than relying on pre-recorded segments with fixed characteristics.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If self-supervised learning is used to train voice synthesis models, then training data requirements are reduced, but the ability to learn complex intonation and prosody patterns may be compromised

Engineering Contradiction:
Improvetraining data quantityVSAvoidintonation and prosody representation
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent implements a self-supervised learning framework where the voice synthesis module generates target voice signals that are fed back to the voice analysis module for feature extraction. This creates a closed-loop training system where the model learns to produce features that, when synthesized, reconstruct the original voice characteristics, enabling effective learning of intonation and prosody patterns even with limited external supervision.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent enables the voice synthesis system to train itself by using its own output as training input. The synthesized target voice signals generated by the synthesis module are used to further train the voice analysis module, creating a self-improving system that can develop complex intonation and prosody representation capabilities without requiring large amounts of externally labeled training data.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240347037A1Method and apparatus for synthesizing unified voice wave based on self-supervised learning
Publication Date: 2024.10.17 SUPERTONE INC
  • US20240347037A1 patent drawing
  • US20240347037A1 patent drawing
  • US20240347037A1 patent drawing

AI summary

Disclosed herein are a self-supervised learning-based unified voice synthesis method and apparatus. The self-supervised learning-based unified voice synthesis method and apparatus: train a voice analysis module to output voice features for training voice signals by using the training voice signals representing training voices, and output voice features for the training voices; and train a voice synthesis module to synthesize voice signals from the voice features for the training voices by using the output voice features, and synthesize synthesized voice signals, representing synthesized voices, from the output voice features. The self-supervised learning-based unified voice synthesis method and apparatus can synthesize voices similar to actual voices by using artificial neural networks that are trained by themselves through self-supervised learning, without the need to train the artificial neural networks on a large quantity of voice and text datasets.