Neural Network Speech Rhythm Conversion Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech rhythm conversion techniques are insufficient in increasing conversion accuracy when no speech signal from the same text is available, due to the non-linear relationship between speech rhythms.

Innovation Solution

A neural network-based speech rhythm conversion device and model learning device that extract feature values from speech signals, including vocal tract spectrum and rhythm information, to convert speech rhythms from one language group to another, even without matching text-based speech signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a Gaussian mixture model is used to convert speech rhythm from speeches of the same text, then conversion can be performed even without native speaker speech, but conversion accuracy is insufficient due to non-linear relationship

Engineering Contradiction:
Improveconversion capability without matching text speechVSAvoidconversion accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the modeling approach from Gaussian mixture model to neural network, transforming the parameter representation from linear weighted addition to non-linear neural network computation. This enables accurate capture of non-linear speech rhythm relationships while maintaining the capability to convert speech without requiring matching text speech from native speakers.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the traditional Gaussian mixture model mechanism with a neural network mechanism. The neural network automatically learns complex non-linear mappings between speech rhythms of different languages, substituting the insufficient linear interpolation approach with a powerful non-linear function approximator that achieves high conversion accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If traditional speech rhythm conversion methods are used, then simple processing is possible, but accuracy is insufficient for non-linear speech rhythm relationships

Engineering Contradiction:
Improveprocessing simplicityVSAvoidconversion accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent substitutes traditional simple processing methods with neural network-based processing. The neural network automatically learns complex non-linear relationships from training data, providing high conversion accuracy while maintaining ease of use through automated feature extraction and rhythm transformation without requiring manual rule creation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the processing approach by changing from fixed conversion rules to learned parameters in a neural network. The model learns optimal conversion parameters from parallel speech corpora, achieving high accuracy while keeping the system simple to operate through automatic parameter adaptation to different language pairs.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11869529B2Speaking rhythm transformation apparatus, model learning apparatus, methods therefor, and program
Publication Date: 2024.01.09 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11869529B2 patent drawing
  • US11869529B2 patent drawing
  • US11869529B2 patent drawing

AI summary

It is intended to accurately convert a speech rhythm. A model storage unit (10) stores a speech rhythm conversion model which is a neural network that receives, as an input thereto, a first feature value vector including information related to a speech rhythm of at least a phoneme extracted from a first speech signal resulting from a speech uttered by a speaker in a first group, converts the speech rhythm of the first speech signal to a speech rhythm of a speaker in a second group, and outputs the speech rhythm of the speaker in the second group. A feature value extraction unit (11) extracts, from the input speech signal resulting from the speech uttered by the speaker in the first group, information related to a vocal tract spectrum and information related to the speech rhythm. A conversion unit (12) inputs the first feature value vector including the information related to the speech rhythm extracted from the input speech signal to the speech rhythm conversion model and obtains the post-conversion speech rhythm. A speech synthesis unit (13) uses the post-conversion speech rhythm and the information related to the vocal tract spectrum extracted from the input speech signal to generate an output speech signal.