Voice Output System Using Label Data for Speaker Attribute Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis technologies cannot accurately switch between voices based on the attributes of speakers in text content, such as age and sex, leading to mismatched voices during reading aloud, as they rely on fixed system or application settings.

Innovation Solution

A speech output system that assigns labels to substrings in content using human computing technology, allowing for the selection of appropriate synthetic voices based on identified attributes, enabling voice switching during speech synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a fixed voice is used for speech synthesis, then the system is simple and easy to operate, but the voice does not match the intended speaker attributes in the content

Engineering Contradiction:
Improveease of operationVSAvoidvoice matching accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system performs voice attribute labeling in advance during content creation or preprocessing. Label data containing speaker attributes (age, sex, etc.) is assigned to substrings before speech synthesis occurs, allowing the TTS system to automatically select appropriate voices without requiring real-time manual intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Label data serves as an intermediary between the content text and the speech synthesis system. This intermediate layer contains structured attribute information that bridges the gap between static text and dynamic voice selection, enabling automated voice matching based on speaker attributes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If voice switching based on speaker attributes is implemented, then the voice matching accuracy is improved, but the system complexity increases

Engineering Contradiction:
Improvevoice matching accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

Voice attribute labeling is performed in advance during content creation or preprocessing. Label data containing speaker attributes (age, sex, etc.) is assigned to substrings before speech synthesis occurs, allowing the TTS system to automatically select appropriate voices without requiring real-time manual intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically performs voice selection based on pre-assigned label data without requiring manual intervention. The TTS system reads the label data, extracts speaker attributes, and autonomously selects matching voices from available voice data, making the complex process transparent to the user.

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If manual voice selection is required for each substring, then the voice matching accuracy is improved, but the operation time and complexity increase significantly

Engineering Contradiction:
Improvevoice matching accuracyVSAvoidoperation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

Voice attribute labeling is performed in advance during content creation or preprocessing. Label data containing speaker attributes (age, sex, etc.) is assigned to substrings before speech synthesis occurs, allowing the TTS system to automatically select appropriate voices without requiring real-time manual intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically performs voice selection based on pre-assigned label data without requiring manual intervention. The TTS system reads the label data, extracts speaker attributes, and autonomously selects matching voices from available voice data, making the complex process transparent to the user.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12125470B2Voice output method, voice output system and program
Publication Date: 2024.10.22 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12125470B2 patent drawing
  • US12125470B2 patent drawing
  • US12125470B2 patent drawing

AI summary

A speech output method carried out by a speech output system that includes a first terminal, a server, and a second terminal, wherein the first terminal carries out: a first label assignment step of assigning label data to character strings that are included in content, the label data representing attributes of speakers in a case where the character strings are to be read aloud by using synthetic speech; and a transmission step of transmitting the label data to the server, the server carries out a saving step of saving the label data transmitted from the first terminal, in a database, in association with content identification information that identifies the content, and the second terminal carries out: an acquisition step of acquiring label data that corresponds to the content identification information regarding the content, from the server.