Tone Conversion Networks for Content-Preserving Audio Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing tone conversion methods struggle with maintaining the content information of the original audio unchanged while achieving a stable tone conversion effect, often degrading the naturalness and accuracy of synthesized audio due to reliance on pre-trained models that ignore non-semantic tone information and require costly labeling.

Innovation Solution

An unsupervised training method for tone conversion models using a tone extraction network, tone removal network, and vocoder to accurately extract and maintain tone and semantic features, reducing the need for costly labeling and improving the accuracy and reliability of tone conversion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional tone conversion models are used, then tone conversion can be performed, but the content information of the original audio degrades and the naturalness of synthesized audio deteriorates

Engineering Contradiction:
Improvetone conversion reliabilityVSAvoidcontent information loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments the audio signal into tone features and semantic features by using a tone extraction network to separate tonal information from content information. This allows the tone conversion model to process tone and semantic components independently, preventing degradation of content information during tone conversion while maintaining conversion reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The tone extraction network extracts tone features from the original audio signal, separating them from semantic content. This extraction process enables the model to perform tone conversion without losing content information, as the semantic features are preserved while only the tone aspects are transformed.

Inventive Principle:
Principle #2Taking out (Extraction)

2Ease of manufacture

If pre-trained models are used for tone conversion, then training can be avoided, but costly labeling is required and model accuracy decreases

Engineering Contradiction:
Improvemodel training easeVSAvoidtone conversion accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The model performs self-service by automatically extracting tone features and semantic features without requiring external labeling. The tone extraction network and semantic feature extraction mechanism enable the system to generate its own training data, eliminating the need for costly manual labeling while maintaining high accuracy in tone conversion.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates synthetic training data by copying and transforming existing audio data through the tone extraction network and semantic feature extraction. This allows the model to be trained on generated data that mirrors real audio characteristics, achieving high accuracy without requiring labeled real audio data for training.

Inventive Principle:
Principle #26Copying

3Reliability

If comprehensive feature extraction is performed, then tone and semantic information are preserved, but model complexity increases

Engineering Contradiction:
Improvetone conversion reliabilityVSAvoidmodel structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the feature extraction process into distinct modules: a tone extraction network for tone features and a semantic feature extraction mechanism for content features. This segmentation allows comprehensive feature extraction while managing model complexity through modular architecture, where each module handles specific aspects of feature extraction independently.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4425482B1Model training and tone conversion method and apparatus, device, and medium
Publication Date: 2025.09.03 BIGO TECH PTE LTD
  • EP4425482B1 patent drawingFigure 1
  • EP4425482B1 patent drawingFigure 2~3
  • EP4425482B1 patent drawingFigure 4~5

AI summary

The present application provides a model training and tone conversion method and apparatus, a device and a medium. By means of a tone extraction network, a first tone feature of input sample audio data can be obtained, so as to obtain tone information of the input sample audio data, which facilitates subsequently obtaining synthesized audio data according to the tone feature, thereby improving the accuracy of the tone of the synthesized audio data. By means of a tone-removing network, and on the basis of the first tone feature, a first semantic feature of the sample audio data can be obtained, thereby accurately obtaining a feature of the sample audio data that is not-related to the tone of the speaker but is related to the spoken content, which facilitates subsequently obtaining synthesized audio data according to the first semantic feature, and ensures the accuracy of the spoken content of the synthesized audio data. After obtaining the trained tone conversion model, tone conversion is carried out by means of the tone conversion model, so that the conversion effect and reliability of tone conversion can be improved.