Voice Conversion Model Learning with Discriminator Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice conversion technologies face challenges in maintaining an appropriate experience distribution when there are many candidates for both the conversion source and destination attributes, leading to deviations in conversion results.

Innovation Solution

A voice signal conversion model learning device that includes a generation unit for generating conversion destination voice signals based on input voice signals, conversion source attribute information, and conversion destination attribute information, along with an identification unit for estimating whether the voice signal represents actual human vocal sound, allowing the generation and identification units to perform learning based on this estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a generator and discriminator system is used for voice conversion, then voice quality conversion can be achieved, but the experience distribution deviates when there are many attribute candidates

Engineering Contradiction:
Improvevoice conversion accuracyVSAvoidexperience distribution
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent introduces a feedback mechanism where the discriminator's estimation result about whether the voice signal represents actual human vocal sound is fed back to the generator. This feedback loop allows the generator to adjust its conversion process to maintain proper experience distribution, resolving the contradiction between conversion accuracy and distribution preservation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The discriminator serves as an intermediary between the generator and the final output. It evaluates the generated voice signals and provides estimation results that mediate the learning process, ensuring that the experience distribution is maintained while still achieving accurate voice conversion.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conversions between multiple attributes are learned simultaneously, then many-to-many conversion capability is achieved, but uniform learning of all combinations becomes impossible

Engineering Contradiction:
Improvemany-to-many conversion capabilityVSAvoidlearning uniformity
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the learning process by introducing the discriminator that evaluates each conversion independently. Instead of learning all attribute combinations uniformly, the system segments the evaluation into individual voice signal assessments, allowing each conversion to be optimized separately while maintaining overall versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the learning parameter from uniform combination learning to experience distribution-based learning. By using the discriminator's estimation results as a feedback parameter, the system adapts the learning process to maintain proper experience distribution across all attribute combinations, enabling versatile many-to-many conversion with improved learning uniformity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12254890B2Audio signal conversion model learning apparatus, audio signal conversion apparatus, audio signal conversion model learning method and program
Publication Date: 2025.03.18 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12254890B2 patent drawing
  • US12254890B2 patent drawing
  • US12254890B2 patent drawing

AI summary

A voice signal conversion model learning device includes: a generation unit configured to execute generation processing of generating a conversion destination voice signal on the basis of an input voice signal that is a voice signal of an input voice, conversion source attribute information that is information indicating an attribute of an input voice that is a voice represented by the input voice signal, and conversion destination attribute information indicating an attribute of a voice represented by the conversion destination voice signal that is a voice signal of a conversion destination of the input voice signal; and an identification unit configured to execute voice estimation processing of estimating whether or not a voice signal that is a processing target is a voice signal representing a vocal sound actually uttered by a person on the basis of the conversion source attribute information and the conversion destination attribute intonation, wherein the conversion destination voice signal is input to the identification unit, the processing target is a voice signal input to the identification unit, and the generation unit and the identification unit pertain learning on the basis of an estimation result of the voice estimation processing.