Voice Conversion Model Learning with Discriminator Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice conversion technologies face challenges in maintaining an appropriate experience distribution when there are many candidates for both the conversion source and destination attributes, leading to deviations in conversion results.
Innovation Solution
A voice signal conversion model learning device that includes a generation unit for generating conversion destination voice signals based on input voice signals, conversion source attribute information, and conversion destination attribute information, along with an identification unit for estimating whether the voice signal represents actual human vocal sound, allowing the generation and identification units to perform learning based on this estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a generator and discriminator system is used for voice conversion, then voice quality conversion can be achieved, but the experience distribution deviates when there are many attribute candidates
Solution Approach 1:
The patent introduces a feedback mechanism where the discriminator's estimation result about whether the voice signal represents actual human vocal sound is fed back to the generator. This feedback loop allows the generator to adjust its conversion process to maintain proper experience distribution, resolving the contradiction between conversion accuracy and distribution preservation.
Solution Approach 2:
The discriminator serves as an intermediary between the generator and the final output. It evaluates the generated voice signals and provides estimation results that mediate the learning process, ensuring that the experience distribution is maintained while still achieving accurate voice conversion.
2Adaptability or versatility
If conversions between multiple attributes are learned simultaneously, then many-to-many conversion capability is achieved, but uniform learning of all combinations becomes impossible
Solution Approach 1:
The patent segments the learning process by introducing the discriminator that evaluates each conversion independently. Instead of learning all attribute combinations uniformly, the system segments the evaluation into individual voice signal assessments, allowing each conversion to be optimized separately while maintaining overall versatility.
Solution Approach 2:
The system changes the learning parameter from uniform combination learning to experience distribution-based learning. By using the discriminator's estimation results as a feedback parameter, the system adapts the learning process to maintain proper experience distribution across all attribute combinations, enabling versatile many-to-many conversion with improved learning uniformity.
Data Source
AI summary
A voice signal conversion model learning device includes: a generation unit configured to execute generation processing of generating a conversion destination voice signal on the basis of an input voice signal that is a voice signal of an input voice, conversion source attribute information that is information indicating an attribute of an input voice that is a voice represented by the input voice signal, and conversion destination attribute information indicating an attribute of a voice represented by the conversion destination voice signal that is a voice signal of a conversion destination of the input voice signal; and an identification unit configured to execute voice estimation processing of estimating whether or not a voice signal that is a processing target is a voice signal representing a vocal sound actually uttered by a person on the basis of the conversion source attribute information and the conversion destination attribute intonation, wherein the conversion destination voice signal is input to the identification unit, the processing target is a voice signal input to the identification unit, and the generation unit and the identification unit pertain learning on the basis of an estimation result of the voice estimation processing.


