Voice Conversion Learning Device Multi-Attribution Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice conversion methods, such as CVAE and CycleGAN, face limitations in converting to multiple attributions due to exploding parameter numbers and do not directly consider the degree of target attribution, leading to limited effectiveness in attribution conversion.

Innovation Solution

A voice conversion system that learns a converter using a learning criterion incorporating real voice similarity, attribution code similarity, reconstruction error, and distance between converted and original sound feature values, with identifiers to minimize these criteria, allowing for conversion to desired attributions without excessive parameter growth and improved quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If CycleGAN is used to convert between multiple attributions, then conversion capability to multiple attributions is achieved, but the number of parameters explodes

Engineering Contradiction:
Improveconversion capability to multiple attributionsVSAvoidnumber of parameters
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies a single converter that can handle conversions between multiple attributions (e.g., speaker gender, age, ethnicity) by incorporating attribution codes as auxiliary inputs. This universal converter replaces the need for separate converters for each attribution pair, thereby achieving multi-attribution conversion capability without exploding the number of parameters.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces attribution codes as an additional dimension of input to the converter. By encoding attribution information as auxiliary inputs rather than creating separate conversion functions for each attribution, the system adds a dimensional layer that enables multi-attribution conversion without proportionally increasing the converter's parameter count.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If conventional voice conversion methods are used, then parallel data is required for learning, but it is difficult to provide pair data of conversion-source voice and target voice of the same utterance content

Engineering Contradiction:
Improvelearning accuracyVSAvoiddata preparation difficulty
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent introduces attribution codes as intermediary representations that bridge the conversion-source voice and target voice without requiring direct parallel pairing. The converter learns to map sound feature values between different attributions using these codes as mediators, enabling conversion learning without the need for difficult-to-obtain parallel utterance data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates synthetic parallel data by combining sound feature values from conversion-source voices with attribution codes and target attribution information. This copying approach generates virtual parallel pairs that train the converter without requiring actual recorded parallel data of the same utterance content, significantly easing data preparation requirements.

Inventive Principle:
Principle #26Copying

3Ease of manufacture

If CVAE is used for non-parallel voice conversion, then parallel data is not necessary, but the feature amount of the generated voice is excessively smoothed

Engineering Contradiction:
Improvedata requirementVSAvoidquality of converted voice
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent incorporates adversarial learning feedback through a discriminator that evaluates the realism of generated voice features. This feedback mechanism prevents excessive smoothing by penalizing overly generic or blurred features, while still maintaining the advantage of not requiring parallel data. The converter learns to generate natural-looking voice features that preserve speaker individuality and audio quality.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the learning objective by incorporating adversarial loss alongside the reconstruction loss. This parameter change in the learning criterion allows the system to avoid the excessive smoothing problem of CVAE while maintaining the non-parallel data advantage. The adversarial component pushes the generated features to be more precise and natural.

Inventive Principle:
Principle #35Parameter changes

4Ease of manufacture

If voice recognition is used to construct parallel data, then non-parallel data can be converted, but performance is limited when voice recognition accuracy is poor

Engineering Contradiction:
Improvedata construction capabilityVSAvoidconversion performance
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent replaces voice recognition as an intermediary step with direct converter learning. Instead of relying on voice recognition to construct parallel data, the converter directly learns the mapping between sound feature values of different attributions using attribution codes, eliminating the performance bottleneck introduced by inaccurate voice recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates synthetic parallel data pairs by copying and combining sound feature values with attribution codes, bypassing the need for voice recognition to identify and pair corresponding segments. This copying approach generates training data that does not depend on the accuracy of voice recognition algorithms.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11869486B2Voice conversion learning device, voice conversion device, method, and program
Publication Date: 2024.01.09 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11869486B2 patent drawing
  • US11869486B2 patent drawing
  • US11869486B2 patent drawing

AI summary

To be able to convert to a voice of the desired attribution. A learning unit learns a converter to minimize a value of a learning criterion of the converter, learns a voice identifier to minimize a value of a learning criterion of the voice identifier, and learns an attribution identifier to minimize a value of a learning criterion of the attribution identifier.