Conversion Learning Apparatus Using Attention Matrix for Voice Prosody

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice conversion technologies are inadequate in effectively converting prosodic features, which are crucial for characterizing speaker identities and speaking modes, and also lack applicability to image and video conversion.

Innovation Solution

A conversion learning device and method that employs a multi-stage machine learning approach, including source encoding, target encoding, attention matrix calculation, and target decoding units, to transform feature sequences between domains, with a focus on minimizing distance between submatrices to achieve effective conversion processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional voice conversion schemes (GMM, NN, NMF) are used, then voice quality conversion is achieved, but prosodic feature conversion is inadequate

Engineering Contradiction:
Improveprosodic feature conversion accuracyVSAvoidvoice conversion effectiveness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the voice conversion task into distinct components: linguistic information processing and prosodic feature processing. By separating these functions and applying specialized processing to each, the system achieves accurate prosodic conversion while maintaining voice quality conversion capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a unified voice conversion framework that handles multiple types of conversion simultaneously - voice quality, prosodic features, speaker identity, and speaking mode. This multi-functional approach allows a single system to address all conversion needs effectively.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If sequence-to-sequence conversion learning is applied, then effectiveness in machine translation and speech recognition is improved, but applicability to image and video conversion is limited

Engineering Contradiction:
Improveconversion application scopeVSAvoidfeature sequence alignment accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces an attention mechanism as an intermediary component that bridges the source and target feature sequences. This attention mechanism enables precise alignment and correspondence identification between features in different domains, facilitating accurate conversion while maintaining the ability to handle diverse application types including image and video conversion.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If attention mechanisms are used to align feature sequences, then conversion accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvefeature sequence alignment accuracyVSAvoidmachine learning model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary linear conversion to the source feature sequences before applying the attention mechanism. This preprocessing step simplifies the subsequent attention computation by transforming the data into a more favorable representation space, thereby reducing the overall computational complexity while maintaining alignment accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12051433B2Conversion learning apparatus, conversion learning method, conversion learning program and conversion apparatus
Publication Date: 2024.07.30 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12051433B2 patent drawing
  • US12051433B2 patent drawing
  • US12051433B2 patent drawing

AI summary

A conversion learning device includes: a source encoding unit that converts, by using a first machine learning model, a feature amount sequence of a source domain that is a characteristic of conversion-source content data, into a first internal representation vector sequence that is a matrix in which internal representation vectors at individual locations of the feature amount sequence of the source domain are arranged; a target encoding unit that converts, by using a second machine learning model, a feature amount sequence of a target domain that is a characteristic of conversion-target content data, into a second internal representation vector sequence that is a matrix in which internal representation vectors at individual locations of the feature amount sequence of the target domain are arranged; an attention matrix calculation unit that calculates, by using the first internal representation vector sequence and the second internal representation vector sequence, an attention matrix that is a matrix mapping the individual locations of the feature amount sequence of the source domain to the individual locations of the feature amount sequence of the target domain, and calculates a third internal representation vector sequence that is a product of an internal representation vector sequence calculated by linear conversion of the first internal representation vector sequence and the attention matrix; a target decoding unit that calculates, by using the third internal representation vector sequence, a feature amount sequence of a conversion domain that is used to convert the source domain into the conversion domain, by using a third machine learning model; and a learning execution unit that causes at least one of the target encoding unit and the target decoding unit to learn such that a distance between a submatrix of the feature amount sequence of the target domain and a submatrix of the feature amount sequence of the conversion domain becomes shorter.