Neural Network Speech Rhythm Conversion Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech rhythm conversion techniques are insufficient in increasing conversion accuracy when no speech signal from the same text is available, due to the non-linear relationship between speech rhythms.
Innovation Solution
A neural network-based speech rhythm conversion device and model learning device that extract feature values from speech signals, including vocal tract spectrum and rhythm information, to convert speech rhythms from one language group to another, even without matching text-based speech signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a Gaussian mixture model is used to convert speech rhythm from speeches of the same text, then conversion can be performed even without native speaker speech, but conversion accuracy is insufficient due to non-linear relationship
Solution Approach 1:
The patent changes the modeling approach from Gaussian mixture model to neural network, transforming the parameter representation from linear weighted addition to non-linear neural network computation. This enables accurate capture of non-linear speech rhythm relationships while maintaining the capability to convert speech without requiring matching text speech from native speakers.
Solution Approach 2:
The patent replaces the traditional Gaussian mixture model mechanism with a neural network mechanism. The neural network automatically learns complex non-linear mappings between speech rhythms of different languages, substituting the insufficient linear interpolation approach with a powerful non-linear function approximator that achieves high conversion accuracy.
2Ease of manufacture
If traditional speech rhythm conversion methods are used, then simple processing is possible, but accuracy is insufficient for non-linear speech rhythm relationships
Solution Approach 1:
The patent substitutes traditional simple processing methods with neural network-based processing. The neural network automatically learns complex non-linear relationships from training data, providing high conversion accuracy while maintaining ease of use through automated feature extraction and rhythm transformation without requiring manual rule creation.
Solution Approach 2:
The patent transforms the processing approach by changing from fixed conversion rules to learned parameters in a neural network. The model learns optimal conversion parameters from parallel speech corpora, achieving high accuracy while keeping the system simple to operate through automatic parameter adaptation to different language pairs.
Data Source
AI summary
It is intended to accurately convert a speech rhythm. A model storage unit (10) stores a speech rhythm conversion model which is a neural network that receives, as an input thereto, a first feature value vector including information related to a speech rhythm of at least a phoneme extracted from a first speech signal resulting from a speech uttered by a speaker in a first group, converts the speech rhythm of the first speech signal to a speech rhythm of a speaker in a second group, and outputs the speech rhythm of the speaker in the second group. A feature value extraction unit (11) extracts, from the input speech signal resulting from the speech uttered by the speaker in the first group, information related to a vocal tract spectrum and information related to the speech rhythm. A conversion unit (12) inputs the first feature value vector including the information related to the speech rhythm extracted from the input speech signal to the speech rhythm conversion model and obtains the post-conversion speech rhythm. A speech synthesis unit (13) uses the post-conversion speech rhythm and the information related to the vocal tract spectrum extracted from the input speech signal to generate an output speech signal.


