Tone Conversion Networks for Content-Preserving Audio Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing tone conversion methods struggle with maintaining the content information of the original audio unchanged while achieving a stable tone conversion effect, often degrading the naturalness and accuracy of synthesized audio due to reliance on pre-trained models that ignore non-semantic tone information and require costly labeling.
Innovation Solution
An unsupervised training method for tone conversion models using a tone extraction network, tone removal network, and vocoder to accurately extract and maintain tone and semantic features, reducing the need for costly labeling and improving the accuracy and reliability of tone conversion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional tone conversion models are used, then tone conversion can be performed, but the content information of the original audio degrades and the naturalness of synthesized audio deteriorates
Solution Approach 1:
The patent segments the audio signal into tone features and semantic features by using a tone extraction network to separate tonal information from content information. This allows the tone conversion model to process tone and semantic components independently, preventing degradation of content information during tone conversion while maintaining conversion reliability.
Solution Approach 2:
The tone extraction network extracts tone features from the original audio signal, separating them from semantic content. This extraction process enables the model to perform tone conversion without losing content information, as the semantic features are preserved while only the tone aspects are transformed.
2Ease of manufacture
If pre-trained models are used for tone conversion, then training can be avoided, but costly labeling is required and model accuracy decreases
Solution Approach 1:
The model performs self-service by automatically extracting tone features and semantic features without requiring external labeling. The tone extraction network and semantic feature extraction mechanism enable the system to generate its own training data, eliminating the need for costly manual labeling while maintaining high accuracy in tone conversion.
Solution Approach 2:
The patent creates synthetic training data by copying and transforming existing audio data through the tone extraction network and semantic feature extraction. This allows the model to be trained on generated data that mirrors real audio characteristics, achieving high accuracy without requiring labeled real audio data for training.
3Reliability
If comprehensive feature extraction is performed, then tone and semantic information are preserved, but model complexity increases
Solution Approach 1:
The patent segments the feature extraction process into distinct modules: a tone extraction network for tone features and a semantic feature extraction mechanism for content features. This segmentation allows comprehensive feature extraction while managing model complexity through modular architecture, where each module handles specific aspects of feature extraction independently.
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
The present application provides a model training and tone conversion method and apparatus, a device and a medium. By means of a tone extraction network, a first tone feature of input sample audio data can be obtained, so as to obtain tone information of the input sample audio data, which facilitates subsequently obtaining synthesized audio data according to the tone feature, thereby improving the accuracy of the tone of the synthesized audio data. By means of a tone-removing network, and on the basis of the first tone feature, a first semantic feature of the sample audio data can be obtained, thereby accurately obtaining a feature of the sample audio data that is not-related to the tone of the speaker but is related to the spoken content, which facilitates subsequently obtaining synthesized audio data according to the first semantic feature, and ensures the accuracy of the spoken content of the synthesized audio data. After obtaining the trained tone conversion model, tone conversion is carried out by means of the tone conversion model, so that the conversion effect and reliability of tone conversion can be improved.