Conversion Learning Apparatus Using Attention Matrix for Voice Prosody
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice conversion technologies are inadequate in effectively converting prosodic features, which are crucial for characterizing speaker identities and speaking modes, and also lack applicability to image and video conversion.
Innovation Solution
A conversion learning device and method that employs a multi-stage machine learning approach, including source encoding, target encoding, attention matrix calculation, and target decoding units, to transform feature sequences between domains, with a focus on minimizing distance between submatrices to achieve effective conversion processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice conversion schemes (GMM, NN, NMF) are used, then voice quality conversion is achieved, but prosodic feature conversion is inadequate
Solution Approach 1:
The patent segments the voice conversion task into distinct components: linguistic information processing and prosodic feature processing. By separating these functions and applying specialized processing to each, the system achieves accurate prosodic conversion while maintaining voice quality conversion capabilities.
Solution Approach 2:
The patent creates a unified voice conversion framework that handles multiple types of conversion simultaneously - voice quality, prosodic features, speaker identity, and speaking mode. This multi-functional approach allows a single system to address all conversion needs effectively.
2Adaptability or versatility
If sequence-to-sequence conversion learning is applied, then effectiveness in machine translation and speech recognition is improved, but applicability to image and video conversion is limited
Solution Approach 1:
The patent introduces an attention mechanism as an intermediary component that bridges the source and target feature sequences. This attention mechanism enables precise alignment and correspondence identification between features in different domains, facilitating accurate conversion while maintaining the ability to handle diverse application types including image and video conversion.
3Measurement precision
If attention mechanisms are used to align feature sequences, then conversion accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent applies preliminary linear conversion to the source feature sequences before applying the attention mechanism. This preprocessing step simplifies the subsequent attention computation by transforming the data into a more favorable representation space, thereby reducing the overall computational complexity while maintaining alignment accuracy.
Data Source
AI summary
A conversion learning device includes: a source encoding unit that converts, by using a first machine learning model, a feature amount sequence of a source domain that is a characteristic of conversion-source content data, into a first internal representation vector sequence that is a matrix in which internal representation vectors at individual locations of the feature amount sequence of the source domain are arranged; a target encoding unit that converts, by using a second machine learning model, a feature amount sequence of a target domain that is a characteristic of conversion-target content data, into a second internal representation vector sequence that is a matrix in which internal representation vectors at individual locations of the feature amount sequence of the target domain are arranged; an attention matrix calculation unit that calculates, by using the first internal representation vector sequence and the second internal representation vector sequence, an attention matrix that is a matrix mapping the individual locations of the feature amount sequence of the source domain to the individual locations of the feature amount sequence of the target domain, and calculates a third internal representation vector sequence that is a product of an internal representation vector sequence calculated by linear conversion of the first internal representation vector sequence and the attention matrix; a target decoding unit that calculates, by using the third internal representation vector sequence, a feature amount sequence of a conversion domain that is used to convert the source domain into the conversion domain, by using a third machine learning model; and a learning execution unit that causes at least one of the target encoding unit and the target decoding unit to learn such that a distance between a submatrix of the feature amount sequence of the target domain and a submatrix of the feature amount sequence of the conversion domain becomes shorter.


