Mixed Language Audio Synthesis Tone Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis technologies face challenges in seamlessly integrating multiple language types, leading to tone differences and 'tone jumps' in synthesized speech, affecting playback quality.

Innovation Solution

An audio synthesis method that acquires mixed language text, performs text coding and decoding processing using encoders and decoders corresponding to multiple language types, and applies acoustic coding to generate unified tone audio, utilizing neural networks and attention mechanisms for natural and smooth output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple language type speeches are switched in voice synthesis, then language diversity is improved, but tone consistency deteriorates causing tone jumps

Engineering Contradiction:
Improvelanguage diversityVSAvoidtone consistency
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent applies parameter changes by transforming the input text into acoustic features through neural network encoding and decoding processes. The system adjusts acoustic parameters (such as pitch, duration, and spectral characteristics) during the synthesis process to maintain consistent tone across different language types, thereby resolving the tone jump issue while preserving language diversity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary neural network processing stage between text input and audio output. The encoder-decoder architecture acts as a mediator that processes text from multiple language types and transforms it into unified acoustic features, which are then synthesized into consistent speech output. This intermediary processing ensures tone consistency regardless of the input language type

Inventive Principle:
Principle #24Intermediary (Mediator)

2Stability of the object's composition

If text coding and decoding processing is performed for mixed language types, then tone consistency is improved, but processing complexity increases

Engineering Contradiction:
Improvetone consistencyVSAvoidprocessing complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent implements a universal encoder-decoder neural network architecture that can process multiple language types through a single unified system. The model is trained to handle mixed language inputs and produces consistent acoustic outputs, eliminating the need for separate processing pipelines for each language type. This multi-functional approach maintains tone consistency while avoiding the complexity of multiple specialized systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses neural network models that learn and copy acoustic patterns from training data to generate synthesized speech. The encoder-decoder architecture copies linguistic and acoustic features from the input text through learned representations, enabling consistent tone generation across different languages without requiring explicit rule-based transformations for each language pair

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12106746B2Audio synthesis method and apparatus, computer readable medium, and electronic device
Publication Date: 2024.10.01 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12106746B2 patent drawing
  • US12106746B2 patent drawing
  • US12106746B2 patent drawing

AI summary

This application discloses a method, an apparatus, a computer readable medium, and an electronic device for audio synthesis. The method includes: acquiring mixed language text information comprising text characters corresponding to at least two language types; performing text coding processing on the mixed language text information based on the at least two language types, to obtain an intermediate semantic coding feature of the mixed language text information; acquiring a target tone feature corresponding to a target tone subject, and performing decoding processing on the intermediate semantic coding feature based on the target tone feature to obtain an acoustic feature; and performing acoustic coding processing on the acoustic feature to obtain an audio corresponding to the mixed language text information.