Parametric Speech Synthesis Using Deep Recurrent Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies face challenges in producing high-quality speech in multiple languages with the same voice, as existing concatenative systems require extensive recording and are limited to a single language, while parametric systems struggle with quality and flexibility.

Innovation Solution

A deep recurrent neural network (DRNN) based parametric speech synthesis system that uses a universal phoneme set and embedding techniques to generate speech in any language and accent, allowing for flexible speaker conversion and improved speech quality by learning voice patterns from a small dataset.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If concatenative speech synthesis is used to produce high-quality speech, then speech quality is improved, but the system requires extensive recording investment and is limited to a single language

Engineering Contradiction:
Improvespeech qualityVSAvoidrecording investment
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical concatenative approach (assembling pre-recorded speech segments) with a neural network-based parametric synthesis system. The neural network learns voice patterns from minimal recordings and generates speech signals dynamically, eliminating the need for extensive pre-recorded speech libraries while maintaining high quality and enabling multi-language support

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the fundamental parameters of speech synthesis by using neural networks to model voice characteristics and generate acoustic features dynamically. This allows the system to adapt to different languages and speakers without requiring re-recording, as the neural network can transform learned voice patterns across different linguistic domains

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If concatenative speech synthesis is used to produce high-quality speech, then speech quality is improved, but the system cannot produce speech in languages other than the recorded language

Engineering Contradiction:
Improvespeech qualityVSAvoidlanguage flexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The neural network model is trained to be universal across multiple languages and speakers. By learning the underlying voice patterns and acoustic characteristics from diverse training data, the same model can generate speech in different languages and accents, making the system multi-functional rather than language-specific

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The neural network acts as an intermediary that translates text into speech signals by first converting text to phonemes and then generating acoustic features based on learned voice patterns. This intermediary processing layer enables the system to handle multiple languages without requiring separate recording sessions for each language

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If parametric speech synthesis is used to reduce recording requirements, then recording investment is reduced, but speech quality is lower than concatenative systems

Engineering Contradiction:
Improverecording investmentVSAvoidspeech quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The system uses feedback mechanisms where the neural network continuously refines its speech generation based on the input text, phoneme sequences, and learned acoustic patterns. This iterative process allows the system to compensate for the simplified recording requirements while maintaining high speech quality through intelligent feature synthesis

Inventive Principle:
Principle #23Feedback

4Stability of the object's composition

If a single speaker records speech for all languages, then voice consistency is maintained, but it is impractical to require a single speaker to speak in all languages

Engineering Contradiction:
Improvevoice consistencyVSAvoidlanguage coverage
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

Instead of requiring a single speaker to physically perform in all languages, the neural network creates virtual copies of speaker voices by learning from minimal recordings. These digital voice models can be transformed to speak different languages while maintaining the original speaker's voice characteristics, effectively copying the voice across linguistic boundaries

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240428778A1Method and System for a Parametric Speech Synthesis
Publication Date: 2024.12.26 GEORGETOWN UNIV
  • US20240428778A1 patent drawing
  • US20240428778A1 patent drawing
  • US20240428778A1 patent drawing

AI summary

Embodiments of the present systems and methods may provide techniques for synthesizing speech in any voice in any language in any accent. For example, in an embodiment, a text-to-speech conversion system may comprise a text converter adapted to convert input text to at least one phoneme selected from a plurality of phonemes stored in memory, a machine-learning model storing voice patterns for a plurality of individuals and adapted to receive the at least one phoneme and an identity of a speaker and to generate acoustic features for each phoneme, and a decoder adapted to receive the generated acoustic features and to generate a speech signal simulating a voice of the identified speaker in a language.