Masked Voice Conversion Learning for Time-Frequency Preservation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice quality conversion techniques struggle to accurately reproduce the time-frequency structure of voice signals while maintaining language information, as the conversion of nonverbal and paralanguage information often disrupts the arrangement of vowels and consonants.

Innovation Solution

A conversion model learning apparatus that employs a mask unit to generate a missing primary feature quantity sequence, a conversion unit to simulate a secondary feature quantity sequence, a calculation unit to calculate a learning reference value, and an update unit to update parameters based on the reference value, using a machine learning model to accurately reproduce the time-frequency structure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If machine learning is used for voice quality conversion to convert nonverbal information and paralanguage information, then voice quality conversion capability is improved, but accurate reproduction of time-frequency structure and language information deteriorates

Engineering Contradiction:
Improvevoice quality conversion capabilityVSAvoidtime-frequency structure reproduction accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the voice signal processing into distinct components: primary voice signal processing and secondary voice signal processing. By separating the time-frequency structure extraction and reconstruction processes, the system can independently optimize each component to maintain language information while converting voice quality characteristics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different parts of the voice signal. The time-frequency structure related to language information (vowels and consonants) is preserved with high fidelity, while the nonverbal information and paralanguage information are converted. This local differentiation allows simultaneous achievement of accurate language reproduction and voice quality transformation.

Inventive Principle:
Principle #3Local quality

2Loss of information

If the arrangement of vowels and consonants is maintained to keep language information, then language information preservation is improved, but conversion of nonverbal information and paralanguage information deteriorates

Engineering Contradiction:
Improvelanguage information preservationVSAvoidnonverbal information conversion capability
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent introduces a time-frequency structure as an intermediary representation between the primary and secondary voice signals. This intermediary structure serves as a bridge that preserves language information (vowels and consonants arrangement) while allowing transformation of nonverbal characteristics. The conversion model learns to map between different voice qualities through this intermediate time-frequency representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If conventional voice conversion methods are used to convert nonverbal information, then voice quality transformation is improved, but accuracy of time-frequency structure reproduction deteriorates

Engineering Contradiction:
Improvevoice quality transformation capabilityVSAvoidtime-frequency structure reproduction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary extraction and analysis of the time-frequency structure from the primary voice signal before conducting the conversion process. By pre-identifying and preserving the critical time-frequency components related to language information, the system ensures that subsequent voice quality transformation does not degrade the accuracy of time-frequency structure reproduction.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12573416B2Conversion model learning apparatus, conversion model generation apparatus, conversion apparatus, conversion method and program
Publication Date: 2026.03.10 NT T INC
  • US12573416B2 patent drawing
  • US12573416B2 patent drawing
  • US12573416B2 patent drawing

AI summary

A mask unit generates a missing primary feature quantity sequence in which a part of a primary feature quantity sequence, which is an acoustic feature quantity sequence of a primary voice signal, on a time axis is masked. A conversion unit generates a simulated secondary feature quantity sequence in which a secondary feature quantity sequence which is an acoustic feature quantity sequence of a secondary voice signal having a time-frequency structure corresponding to a primary voice signal by inputting a missing primary feature quantity sequence to a conversion model that is a machine learning model. A calculation unit calculates a learning reference value which becomes higher as a time frequency structure of a simulated secondary feature quantity sequence is closer to a time frequency structure of a secondary feature quantity sequence. An update unit updates parameters of a conversion model on the basis of a learning reference value.