Masked Voice Conversion Learning for Time-Frequency Preservation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice quality conversion techniques struggle to accurately reproduce the time-frequency structure of voice signals while maintaining language information, as the conversion of nonverbal and paralanguage information often disrupts the arrangement of vowels and consonants.
Innovation Solution
A conversion model learning apparatus that employs a mask unit to generate a missing primary feature quantity sequence, a conversion unit to simulate a secondary feature quantity sequence, a calculation unit to calculate a learning reference value, and an update unit to update parameters based on the reference value, using a machine learning model to accurately reproduce the time-frequency structure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If machine learning is used for voice quality conversion to convert nonverbal information and paralanguage information, then voice quality conversion capability is improved, but accurate reproduction of time-frequency structure and language information deteriorates
Solution Approach 1:
The patent segments the voice signal processing into distinct components: primary voice signal processing and secondary voice signal processing. By separating the time-frequency structure extraction and reconstruction processes, the system can independently optimize each component to maintain language information while converting voice quality characteristics.
Solution Approach 2:
The patent applies different processing qualities to different parts of the voice signal. The time-frequency structure related to language information (vowels and consonants) is preserved with high fidelity, while the nonverbal information and paralanguage information are converted. This local differentiation allows simultaneous achievement of accurate language reproduction and voice quality transformation.
2Loss of information
If the arrangement of vowels and consonants is maintained to keep language information, then language information preservation is improved, but conversion of nonverbal information and paralanguage information deteriorates
Solution Approach 1:
The patent introduces a time-frequency structure as an intermediary representation between the primary and secondary voice signals. This intermediary structure serves as a bridge that preserves language information (vowels and consonants arrangement) while allowing transformation of nonverbal characteristics. The conversion model learns to map between different voice qualities through this intermediate time-frequency representation.
3Adaptability or versatility
If conventional voice conversion methods are used to convert nonverbal information, then voice quality transformation is improved, but accuracy of time-frequency structure reproduction deteriorates
Solution Approach 1:
The patent performs preliminary extraction and analysis of the time-frequency structure from the primary voice signal before conducting the conversion process. By pre-identifying and preserving the critical time-frequency components related to language information, the system ensures that subsequent voice quality transformation does not degrade the accuracy of time-frequency structure reproduction.
Data Source
AI summary
A mask unit generates a missing primary feature quantity sequence in which a part of a primary feature quantity sequence, which is an acoustic feature quantity sequence of a primary voice signal, on a time axis is masked. A conversion unit generates a simulated secondary feature quantity sequence in which a secondary feature quantity sequence which is an acoustic feature quantity sequence of a secondary voice signal having a time-frequency structure corresponding to a primary voice signal by inputting a missing primary feature quantity sequence to a conversion model that is a machine learning model. A calculation unit calculates a learning reference value which becomes higher as a time frequency structure of a simulated secondary feature quantity sequence is closer to a time frequency structure of a secondary feature quantity sequence. An update unit updates parameters of a conversion model on the basis of a learning reference value.


