Multi-Character Voice Synthesis via Character Attribute Speaker Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice synthesis technologies can only produce a single speaker voice, failing to generate distinct voices for multiple characters in scenarios like audio reading and human-machine dialogue.
Innovation Solution
A method that analyzes text information to identify characters, recognizes character attributes, and matches them with pre-stored speakers to generate multi-character synthesized voices, incorporating attributes such as gender, age, regional information, and pronunciation style to differentiate voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single speaker is used for voice synthesis, then the device complexity is reduced, but the ability to distinguish different characters is lost
Solution Approach 1:
The text is segmented into different character portions, and each character is assigned a corresponding speaker based on character attributes. This allows the system to maintain simplicity while preserving character distinction by processing different text segments with different speakers.
Solution Approach 2:
Different speakers with distinct voice characteristics are assigned to different characters based on their specific attributes (gender, age, region). This local differentiation enables character distinction while the overall system remains manageable through automated speaker selection.
2Loss of information
If multiple speakers are used for different characters, then character distinction is improved, but the device complexity increases
Solution Approach 1:
Speaker information is pre-stored in the system with associated character attributes. During voice synthesis, the system automatically matches characters with corresponding speakers based on pre-established attribute relationships, avoiding the need for complex real-time decision-making mechanisms.
Solution Approach 2:
The system automatically performs character recognition, attribute extraction, and speaker selection without requiring manual intervention. This self-service capability manages the complexity of multi-speaker systems through automated processing pipelines.
3Measurement precision
If character attributes are analyzed in detail, then speaker matching accuracy is improved, but the processing time increases
Solution Approach 1:
The system extracts essential character attributes (gender, age, region) that are sufficient for accurate speaker matching, rather than analyzing every possible character feature. This partial action approach achieves adequate precision while maintaining efficient processing speed.
Data Source
AI summary
Provided are a voice synthesis method, an apparatus, a device, and a storage medium, involving obtaining text information and determining characters in the text information and a text content of each of the characters; performing a character recognition on the text content of each of the characters, to determine character attribute information of each of the characters; obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, where the speakers are pre-stored pronunciation object having the character attribute information; and generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information. These improve pronunciation diversities of different characters in the synthesized voices, improve an audience's discrimination between different characters in the synthesized voices, and thereby improve experience of a user.


