Voice Conversion Model Using Local Features to Preserve Correlations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice quality conversion technologies using machine learning lose important information during feature extraction, particularly the structural correlations between parts of the input data, leading to inadequate learning and conversion.
Innovation Solution
A voice signal conversion model that includes local feature quantity acquisition, adjustment parameter value acquisition, and mapping conversion processing to minimize information loss by retaining the structural correlations of input data through a conversion learning model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If convolutional neural network processing is used to extract feature quantities from voice signals, then voice quality conversion can be performed, but information indicating the structure of the input data is lost due to contraction processing
Solution Approach 1:
The patent divides the voice signal into multiple segments and extracts local feature quantities for each segment using convolutional neural networks. This segmentation allows the system to process and preserve structural correlations within local regions while still achieving voice quality conversion, thereby reducing the information loss that would occur in global contraction processing.
Solution Approach 2:
The patent introduces a new dimension by extracting local feature quantities at multiple levels (e.g., frame level, phoneme level) and combining them with global feature quantities. This multi-dimensional approach preserves structural correlation information that would otherwise be lost in conventional single-level contraction processing, while maintaining voice quality conversion capability.
2Productivity
If feature quantity extraction is performed using contraction processing, then voice signal conversion can be achieved, but learning is not appropriately performed due to information loss
Solution Approach 1:
The patent applies local quality by extracting local feature quantities for specific regions of the voice signal and combining them with global features. This allows the system to maintain high conversion processing efficiency through localized processing while improving learning accuracy by preserving structural correlation information that conventional global contraction processing loses.
Solution Approach 2:
The patent changes the parameters of feature extraction by introducing local feature quantity extraction as an additional component to the conventional contraction processing. This parameter change enables the system to maintain both processing efficiency and learning accuracy by preserving information about structural correlations between parts of the input data.
3Adaptability or versatility
If conventional machine learning processing is used for voice conversion, then conversion can be performed, but information on correlation between parts of input data is lost
Solution Approach 1:
The patent implements a nested structure by extracting local feature quantities at multiple levels (frame level, phoneme level) and nesting them within the global feature quantity extraction framework. This nested approach allows the system to maintain voice conversion capability while preserving correlation information between parts of the input data through the hierarchical structure.
Data Source
AI summary
The present invention is a voice signal conversion model learning device provided with: a learning data acquisition unit for acquiring learning input data, which is an input voice signal; and a learning stage conversion unit for executing a conversion learning model, which is a machine learning model, including learning stage conversion processing for converting the learning input data into learning stage conversion-destination data, which is a conversion-destination voice signal. The learning stage conversion processing includes local feature acquisition processing for acquiring, on the basis of input data to be processed, a feature for each learning input-side subset. The conversion learning model further includes adjustment-parameter-value acquisition processing for acquiring, on the basis of the learning input data, an adjustment parameter value. The learning stage conversion processing converts the learning input data into the learning stage conversion-destination data by using a result of a predetermined computation based on the adjustment parameter value.


