Multi-Character Voice Synthesis via Character Attribute Speaker Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis technologies can only produce a single speaker voice, failing to generate distinct voices for multiple characters in scenarios like audio reading and human-machine dialogue.

Innovation Solution

A method that analyzes text information to identify characters, recognizes character attributes, and matches them with pre-stored speakers to generate multi-character synthesized voices, incorporating attributes such as gender, age, regional information, and pronunciation style to differentiate voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single speaker is used for voice synthesis, then the device complexity is reduced, but the ability to distinguish different characters is lost

Engineering Contradiction:
Improvevoice synthesis system complexityVSAvoidcharacter distinction information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The text is segmented into different character portions, and each character is assigned a corresponding speaker based on character attributes. This allows the system to maintain simplicity while preserving character distinction by processing different text segments with different speakers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different speakers with distinct voice characteristics are assigned to different characters based on their specific attributes (gender, age, region). This local differentiation enables character distinction while the overall system remains manageable through automated speaker selection.

Inventive Principle:
Principle #3Local quality

2Loss of information

If multiple speakers are used for different characters, then character distinction is improved, but the device complexity increases

Engineering Contradiction:
Improvecharacter distinction informationVSAvoidvoice synthesis system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

Speaker information is pre-stored in the system with associated character attributes. During voice synthesis, the system automatically matches characters with corresponding speakers based on pre-established attribute relationships, avoiding the need for complex real-time decision-making mechanisms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically performs character recognition, attribute extraction, and speaker selection without requiring manual intervention. This self-service capability manages the complexity of multi-speaker systems through automated processing pipelines.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If character attributes are analyzed in detail, then speaker matching accuracy is improved, but the processing time increases

Engineering Contradiction:
Improvespeaker matching accuracyVSAvoidvoice synthesis processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts essential character attributes (gender, age, region) that are sufficient for accurate speaker matching, rather than analyzing every possible character feature. This partial action approach achieves adequate precision while maintaining efficient processing speed.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11600259B2Voice synthesis method, apparatus, device and storage medium
Publication Date: 2023.03.07 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11600259B2 patent drawing
  • US11600259B2 patent drawing
  • US11600259B2 patent drawing

AI summary

Provided are a voice synthesis method, an apparatus, a device, and a storage medium, involving obtaining text information and determining characters in the text information and a text content of each of the characters; performing a character recognition on the text content of each of the characters, to determine character attribute information of each of the characters; obtaining speakers in one-to-one correspondence with the characters according to the character attribute information of each of the characters, where the speakers are pre-stored pronunciation object having the character attribute information; and generating multi-character synthesized voices according to the text information and the speakers corresponding to the characters of the text information. These improve pronunciation diversities of different characters in the synthesized voices, improve an audience's discrimination between different characters in the synthesized voices, and thereby improve experience of a user.