Audio Generation via Phoneme-Semantic Fusion for Zero-Shot Voice Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio synthesis technologies struggle to generate high-quality speech with consistent timbre and emotion, particularly in Chinese, and require extensive training datasets or user-generated corpora, limiting their versatility and effectiveness.
Innovation Solution
An audio generation method that fuses phoneme and semantic features of a target text with reference audio features to decode a target audio, utilizing encoding and decoding processes to achieve high-quality audio synthesis with consistent timbre and emotion, even in zero-shot or few-shot scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing audio synthesis technologies are used, then audio generation is possible, but the timbre and emotion consistency is poor and extensive training datasets are required
Solution Approach 1:
The patent segments the audio synthesis process into distinct feature extraction modules: phoneme feature extraction, semantic feature extraction, and reference audio feature extraction. Each module processes specific aspects of the input data independently before fusion, enabling targeted optimization without requiring extensive full-dataset training
Solution Approach 2:
The patent introduces an intermediary encoding process that transforms multiple input features (phoneme features, semantic features, reference audio features) into a unified encoding feature representation. This intermediary step acts as a mediator that integrates diverse feature types while maintaining their distinctive characteristics, achieving reliable timbre and emotion consistency without requiring large training datasets
2Manufacturing precision
If phoneme and semantic features are fused with reference audio features, then audio quality and naturalness improve, but the processing complexity increases
Solution Approach 1:
The patent divides the complex processing into segmented stages: first extracting phoneme features from target text, then extracting semantic features,接着 extracting reference audio features, and finally fusing them through encoding. This segmentation reduces processing complexity by handling each feature type separately before integration
Solution Approach 2:
The patent merges multiple feature representations (phoneme features, semantic features, reference audio features) into a unified encoding feature through a fusion process. This combining approach achieves high audio quality by integrating complementary information from different sources while maintaining manageable processing complexity through systematic integration
Data Source
AI summary
An audio generation method, a method of training an audio generation model, an electronic device, and a storage medium, which relate to a field of an artificial intelligence technology, in particular to fields of deep learning, large model and audio synthesis technologies. The audio generation method includes: fusing a target phoneme feature of a target text and a target semantic feature of the target text to obtain a target fusion feature; obtaining an encoding feature according to the target fusion feature, a reference fusion feature and a reference audio feature, where the reference fusion feature is obtained by fusing a reference phoneme feature of a reference text and a reference semantic feature of the reference text, and the reference audio feature is determined according to a reference audio corresponding to the reference text; and decoding the encoding feature to obtain a target audio corresponding to the target text.


