Voice Conversion Model Using Local Features to Preserve Correlations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice quality conversion technologies using machine learning lose important information during feature extraction, particularly the structural correlations between parts of the input data, leading to inadequate learning and conversion.

Innovation Solution

A voice signal conversion model that includes local feature quantity acquisition, adjustment parameter value acquisition, and mapping conversion processing to minimize information loss by retaining the structural correlations of input data through a conversion learning model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If convolutional neural network processing is used to extract feature quantities from voice signals, then voice quality conversion can be performed, but information indicating the structure of the input data is lost due to contraction processing

Engineering Contradiction:
Improvevoice quality conversion capabilityVSAvoidstructural correlation information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent divides the voice signal into multiple segments and extracts local feature quantities for each segment using convolutional neural networks. This segmentation allows the system to process and preserve structural correlations within local regions while still achieving voice quality conversion, thereby reducing the information loss that would occur in global contraction processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension by extracting local feature quantities at multiple levels (e.g., frame level, phoneme level) and combining them with global feature quantities. This multi-dimensional approach preserves structural correlation information that would otherwise be lost in conventional single-level contraction processing, while maintaining voice quality conversion capability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If feature quantity extraction is performed using contraction processing, then voice signal conversion can be achieved, but learning is not appropriately performed due to information loss

Engineering Contradiction:
Improveconversion processing efficiencyVSAvoidlearning accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by extracting local feature quantities for specific regions of the voice signal and combining them with global features. This allows the system to maintain high conversion processing efficiency through localized processing while improving learning accuracy by preserving structural correlation information that conventional global contraction processing loses.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameters of feature extraction by introducing local feature quantity extraction as an additional component to the conventional contraction processing. This parameter change enables the system to maintain both processing efficiency and learning accuracy by preserving information about structural correlations between parts of the input data.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If conventional machine learning processing is used for voice conversion, then conversion can be performed, but information on correlation between parts of input data is lost

Engineering Contradiction:
Improvevoice conversion capabilityVSAvoidcorrelation information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent implements a nested structure by extracting local feature quantities at multiple levels (frame level, phoneme level) and nesting them within the global feature quantity extraction framework. This nested approach allows the system to maintain voice conversion capability while preserving correlation information between parts of the input data through the hierarchical structure.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS12475904B2Audio signal conversion model learning apparatus, audio signal conversion apparatus, audio signal conversion model learning method and program
Publication Date: 2025.11.18 NT T INC
  • US12475904B2 patent drawing
  • US12475904B2 patent drawing
  • US12475904B2 patent drawing

AI summary

The present invention is a voice signal conversion model learning device provided with: a learning data acquisition unit for acquiring learning input data, which is an input voice signal; and a learning stage conversion unit for executing a conversion learning model, which is a machine learning model, including learning stage conversion processing for converting the learning input data into learning stage conversion-destination data, which is a conversion-destination voice signal. The learning stage conversion processing includes local feature acquisition processing for acquiring, on the basis of input data to be processed, a feature for each learning input-side subset. The conversion learning model further includes adjustment-parameter-value acquisition processing for acquiring, on the basis of the learning input data, an adjustment parameter value. The learning stage conversion processing converts the learning input data into the learning stage conversion-destination data by using a result of a predetermined computation based on the adjustment parameter value.