Voice Conversion Rule Learning Using Attribute Mismatch

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice conversion techniques struggle to learn comprehensive voice conversion rules when using mass speech data from a conversion-source speaker and limited speech data from a conversion-target speaker, as the limited speech content restricts the learning of rules reflecting the information in the source speaker's speech unit database.

Innovation Solution

A speech processing apparatus that extracts and selects speech units from a conversion-source speaker's database based on attribute mismatch costs, generating target-speaker attribute information and selecting source-speaker speech units to create voice conversion rules for converting speech units from the conversion-source speaker to the conversion-target speaker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech data of the same content from conversion-source speaker and conversion-target speaker are associated to learn voice conversion rules, then the conversion rules can be learned from paired data, but the speech content is limited and cannot reflect the information in the mass speech unit database of the conversion-source speaker

Engineering Contradiction:
Improveaccuracy of voice conversion rulesVSAvoidcoverage of speech content
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the speech units into smaller granular units (e.g., phonemes, diphones, triphones) and associates them based on content similarity rather than requiring identical sentences. This allows comprehensive use of mass speech data while maintaining accurate conversion rule learning through segmented unit matching

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an automatic speech recognition (ASR) system as an intermediary to convert speech units into text, enabling content-based association between conversion-source and conversion-target speech data. This mediator allows matching speech units with similar semantic content even when the exact sentences differ, expanding the usable speech data coverage

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If mass speech data of conversion-source speaker and low-volume speech data of conversion-target speaker are used, then voice conversion can be achieved with limited target speaker data, but the speech contents are limited and conversion rules cannot reflect information in the mass speech unit database

Engineering Contradiction:
Improveefficiency of voice conversion rule learningVSAvoidinformation from mass speech unit database
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

By segmenting speech into fine-grained units and using ASR-based content matching, the patent enables efficient association of mass source speaker data with limited target speaker data. This segmentation approach maintains high productivity while preventing information loss by充分利用 the available mass speech unit database

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary ASR processing and speech unit segmentation before association, preparing the data in advance to enable comprehensive utilization of mass speech data. This preliminary action allows the system to efficiently match and associate speech units based on semantic content rather than requiring extensive manual pairing

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7580839B2Apparatus and method for voice conversion using attribute information
Publication Date: 2009.08.25 TOSHIBA DIGITAL SOLUTIONS CORP
  • US7580839B2 patent drawing
  • US7580839B2 patent drawing
  • US7580839B2 patent drawing

AI summary

A speech processing apparatus according to an embodiment of the invention includes a conversion-source-speaker speech-unit database; a voice-conversion-rule-learning-data generating means; and a voice-conversion-rule learning means, with which it makes voice conversion rules. The voice-conversion-rule-learning-data generating means includes a conversion-target-speaker speech-unit extracting means; an attribute-information generating means; a conversion-source-speaker speech-unit database; and a conversion-source-speaker speech-unit selection means. The conversion-source-speaker speech-unit selection means selects conversion-source-speaker speech units corresponding to conversion-target-speaker speech units based on the mismatch between the attribute information of the conversion-target-speaker speech units and that of the conversion-source-speaker speech units, whereby the voice conversion rules are made from the selected pair of the conversion-target-speaker speech units and the conversion-source-speaker speech units.