Audio Generation via Phoneme-Semantic Fusion for Zero-Shot Voice Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio synthesis technologies struggle to generate high-quality speech with consistent timbre and emotion, particularly in Chinese, and require extensive training datasets or user-generated corpora, limiting their versatility and effectiveness.

Innovation Solution

An audio generation method that fuses phoneme and semantic features of a target text with reference audio features to decode a target audio, utilizing encoding and decoding processes to achieve high-quality audio synthesis with consistent timbre and emotion, even in zero-shot or few-shot scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing audio synthesis technologies are used, then audio generation is possible, but the timbre and emotion consistency is poor and extensive training datasets are required

Engineering Contradiction:
Improvetimbre and emotion consistencyVSAvoidtraining dataset size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the audio synthesis process into distinct feature extraction modules: phoneme feature extraction, semantic feature extraction, and reference audio feature extraction. Each module processes specific aspects of the input data independently before fusion, enabling targeted optimization without requiring extensive full-dataset training

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary encoding process that transforms multiple input features (phoneme features, semantic features, reference audio features) into a unified encoding feature representation. This intermediary step acts as a mediator that integrates diverse feature types while maintaining their distinctive characteristics, achieving reliable timbre and emotion consistency without requiring large training datasets

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If phoneme and semantic features are fused with reference audio features, then audio quality and naturalness improve, but the processing complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the complex processing into segmented stages: first extracting phoneme features from target text, then extracting semantic features,接着 extracting reference audio features, and finally fusing them through encoding. This segmentation reduces processing complexity by handling each feature type separately before integration

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges multiple feature representations (phoneme features, semantic features, reference audio features) into a unified encoding feature through a fusion process. This combining approach achieves high audio quality by integrating complementary information from different sources while maintaining manageable processing complexity through systematic integration

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250363977A1Audio generation method, method of training model, device, and storage medium
Publication Date: 2025.11.27 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20250363977A1 patent drawing
  • US20250363977A1 patent drawing
  • US20250363977A1 patent drawing

AI summary

An audio generation method, a method of training an audio generation model, an electronic device, and a storage medium, which relate to a field of an artificial intelligence technology, in particular to fields of deep learning, large model and audio synthesis technologies. The audio generation method includes: fusing a target phoneme feature of a target text and a target semantic feature of the target text to obtain a target fusion feature; obtaining an encoding feature according to the target fusion feature, a reference fusion feature and a reference audio feature, where the reference fusion feature is obtained by fusing a reference phoneme feature of a reference text and a reference semantic feature of the reference text, and the reference audio feature is determined according to a reference audio corresponding to the reference text; and decoding the encoding feature to obtain a target audio corresponding to the target text.