Two-Stage Timbre Conversion Model With Diffusion Feature Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods fail to accurately and efficiently convert the timbre of audio within a predetermined time and require extensive training with large amounts of audio data.

Innovation Solution

A method involving the processing of first and second audio content with a first generation model to generate third audio content, providing this content to a second generation model to generate audio features, and training the second model based on these features, utilizing diffusion models for improved timbre conversion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods are used for timbre conversion, then the model can process audio content, but the conversion accuracy is insufficient and extensive training with large amounts of audio data is required

Engineering Contradiction:
Improvetimbre conversion accuracyVSAvoidamount of audio data for training
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the timbre conversion process into two distinct models: a first generation model for initial timbre conversion and a second generation model for refining and reconstructing audio features. This segmentation allows each model to specialize in specific aspects of timbre conversion, improving overall accuracy without requiring extensive training data for a single comprehensive model

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces audio features as an intermediary representation between the input audio content and the final converted output. The second generation model processes these intermediate features to reconstruct the audio, enabling more precise timbre control and better conversion accuracy with reduced training data requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If conventional methods are used for timbre conversion, then the model can generate converted audio, but the conversion efficiency is low and cannot complete within a predetermined time

Engineering Contradiction:
Improvetimbre conversion efficiencyVSAvoidtraining time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary timbre conversion using the first generation model to generate intermediate audio content and features before the second generation model refines the output. This preliminary action allows the system to quickly generate a baseline conversion that can be subsequently improved, reducing overall processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By dividing the conversion process into two sequential stages with different computational complexities, the system can optimize each stage for speed and accuracy respectively, improving overall conversion efficiency while meeting time constraints

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If a single generation model is used, then the system structure is simple, but the generation effect and timbre conversion accuracy are insufficient

Engineering Contradiction:
Improvegeneration effect accuracyVSAvoidmodel structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the single model into two specialized generation models: the first model handles initial timbre conversion while the second model focuses on audio feature reconstruction. This functional segmentation improves generation accuracy by allowing each model to specialize in specific tasks

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a second generation model that copies and builds upon the architecture and learning of the first model, then specializes it for feature reconstruction. This copying approach allows the system to leverage the first model's learned representations while adding specialized capabilities without completely redesigning the system

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260073897A1Method, apparatus, device, and storage medium for training generation model
Publication Date: 2026.03.12 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260073897A1 patent drawing
  • US20260073897A1 patent drawing
  • US20260073897A1 patent drawing

AI summary

A method, an apparatus, a device and a storage medium for training a generation model are provided. The method provided by the disclosure includes: obtaining first audio content corresponding to a first timbre and second audio content corresponding to a second timbre; processing the first audio content and the second audio content with a first generation model to generate third audio content; providing the third audio content and a first portion of the second audio content to a second generation model to generate a first audio feature; and training the second generation model based on the first audio feature and a second audio feature corresponding to the second audio content.