Voice Conversion Learning Device Multi-Attribution Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice conversion methods, such as CVAE and CycleGAN, face limitations in converting to multiple attributions due to exploding parameter numbers and do not directly consider the degree of target attribution, leading to limited effectiveness in attribution conversion.
Innovation Solution
A voice conversion system that learns a converter using a learning criterion incorporating real voice similarity, attribution code similarity, reconstruction error, and distance between converted and original sound feature values, with identifiers to minimize these criteria, allowing for conversion to desired attributions without excessive parameter growth and improved quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If CycleGAN is used to convert between multiple attributions, then conversion capability to multiple attributions is achieved, but the number of parameters explodes
Solution Approach 1:
The patent applies a single converter that can handle conversions between multiple attributions (e.g., speaker gender, age, ethnicity) by incorporating attribution codes as auxiliary inputs. This universal converter replaces the need for separate converters for each attribution pair, thereby achieving multi-attribution conversion capability without exploding the number of parameters.
Solution Approach 2:
The patent introduces attribution codes as an additional dimension of input to the converter. By encoding attribution information as auxiliary inputs rather than creating separate conversion functions for each attribution, the system adds a dimensional layer that enables multi-attribution conversion without proportionally increasing the converter's parameter count.
2Reliability
If conventional voice conversion methods are used, then parallel data is required for learning, but it is difficult to provide pair data of conversion-source voice and target voice of the same utterance content
Solution Approach 1:
The patent introduces attribution codes as intermediary representations that bridge the conversion-source voice and target voice without requiring direct parallel pairing. The converter learns to map sound feature values between different attributions using these codes as mediators, enabling conversion learning without the need for difficult-to-obtain parallel utterance data.
Solution Approach 2:
The patent creates synthetic parallel data by combining sound feature values from conversion-source voices with attribution codes and target attribution information. This copying approach generates virtual parallel pairs that train the converter without requiring actual recorded parallel data of the same utterance content, significantly easing data preparation requirements.
3Ease of manufacture
If CVAE is used for non-parallel voice conversion, then parallel data is not necessary, but the feature amount of the generated voice is excessively smoothed
Solution Approach 1:
The patent incorporates adversarial learning feedback through a discriminator that evaluates the realism of generated voice features. This feedback mechanism prevents excessive smoothing by penalizing overly generic or blurred features, while still maintaining the advantage of not requiring parallel data. The converter learns to generate natural-looking voice features that preserve speaker individuality and audio quality.
Solution Approach 2:
The patent changes the learning objective by incorporating adversarial loss alongside the reconstruction loss. This parameter change in the learning criterion allows the system to avoid the excessive smoothing problem of CVAE while maintaining the non-parallel data advantage. The adversarial component pushes the generated features to be more precise and natural.
4Ease of manufacture
If voice recognition is used to construct parallel data, then non-parallel data can be converted, but performance is limited when voice recognition accuracy is poor
Solution Approach 1:
The patent replaces voice recognition as an intermediary step with direct converter learning. Instead of relying on voice recognition to construct parallel data, the converter directly learns the mapping between sound feature values of different attributions using attribution codes, eliminating the performance bottleneck introduced by inaccurate voice recognition.
Solution Approach 2:
The patent creates synthetic parallel data pairs by copying and combining sound feature values with attribution codes, bypassing the need for voice recognition to identify and pair corresponding segments. This copying approach generates training data that does not depend on the accuracy of voice recognition algorithms.
Data Source
AI summary
To be able to convert to a voice of the desired attribution. A learning unit learns a converter to minimize a value of a learning criterion of the converter, learns a voice identifier to minimize a value of a learning criterion of the voice identifier, and learns an attribution identifier to minimize a value of a learning criterion of the attribution identifier.


