Mixed Language Audio Synthesis Tone Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice synthesis technologies face challenges in seamlessly integrating multiple language types, leading to tone differences and 'tone jumps' in synthesized speech, affecting playback quality.
Innovation Solution
An audio synthesis method that acquires mixed language text, performs text coding and decoding processing using encoders and decoders corresponding to multiple language types, and applies acoustic coding to generate unified tone audio, utilizing neural networks and attention mechanisms for natural and smooth output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple language type speeches are switched in voice synthesis, then language diversity is improved, but tone consistency deteriorates causing tone jumps
Solution Approach 1:
The patent applies parameter changes by transforming the input text into acoustic features through neural network encoding and decoding processes. The system adjusts acoustic parameters (such as pitch, duration, and spectral characteristics) during the synthesis process to maintain consistent tone across different language types, thereby resolving the tone jump issue while preserving language diversity
Solution Approach 2:
The patent introduces an intermediary neural network processing stage between text input and audio output. The encoder-decoder architecture acts as a mediator that processes text from multiple language types and transforms it into unified acoustic features, which are then synthesized into consistent speech output. This intermediary processing ensures tone consistency regardless of the input language type
2Stability of the object's composition
If text coding and decoding processing is performed for mixed language types, then tone consistency is improved, but processing complexity increases
Solution Approach 1:
The patent implements a universal encoder-decoder neural network architecture that can process multiple language types through a single unified system. The model is trained to handle mixed language inputs and produces consistent acoustic outputs, eliminating the need for separate processing pipelines for each language type. This multi-functional approach maintains tone consistency while avoiding the complexity of multiple specialized systems
Solution Approach 2:
The patent uses neural network models that learn and copy acoustic patterns from training data to generate synthesized speech. The encoder-decoder architecture copies linguistic and acoustic features from the input text through learned representations, enabling consistent tone generation across different languages without requiring explicit rule-based transformations for each language pair
Data Source
AI summary
This application discloses a method, an apparatus, a computer readable medium, and an electronic device for audio synthesis. The method includes: acquiring mixed language text information comprising text characters corresponding to at least two language types; performing text coding processing on the mixed language text information based on the at least two language types, to obtain an intermediate semantic coding feature of the mixed language text information; acquiring a target tone feature corresponding to a target tone subject, and performing decoding processing on the intermediate semantic coding feature based on the target tone feature to obtain an acoustic feature; and performing acoustic coding processing on the acoustic feature to obtain an audio corresponding to the mixed language text information.


