Neural Speech Model for Natural Audio Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech (TTS) systems face limitations in generating natural-sounding speech, particularly in producing diverse vocal attributes and qualities, as they require extensive recorded speech datasets and rely on either unit selection or parametric synthesis, which are time-consuming and result in less natural output.

Innovation Solution

A machine-learning-based speech model is trained to directly generate audio data, using a sample model, conditioning model, and output model that can produce tens of thousands of audio samples per second, leveraging causal convolutions and linguistic context features to create high-quality audio with specific vocal attributes, tones, and languages, improving output quality beyond traditional methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional unit selection or parametric synthesis methods are used, then TTS systems can generate speech output, but the output lacks naturalness and diverse vocal attributes

Engineering Contradiction:
Improvespeech qualityVSAvoidvocal attribute diversity
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the TTS system from selecting or synthesizing speech parameters to directly generating audio waveforms. By changing the fundamental parameter representation from discrete speech parameters to continuous audio samples, the system achieves both high naturalness and diverse vocal attributes simultaneously.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical TTS systems (unit selection and parametric synthesis) with a neural network-based audio generation system. This substitution enables direct waveform generation that naturally captures human speech characteristics without relying on pre-recorded units or parameter-based synthesis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If extensive recorded speech datasets are used to improve speech quality, then output quality improves, but the system becomes less flexible and more time-consuming

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent implements continuous audio waveform generation instead of discrete unit selection or parameter synthesis. The neural network generates audio samples continuously at high rates (tens of thousands per second), eliminating the time-consuming selection and concatenation processes of traditional systems while maintaining high audio quality.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary training with extensive datasets to create a pre-trained neural network model. Once trained, the model can generate high-quality speech rapidly without requiring access to the extensive training data during operation, thus achieving both high quality and fast processing.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If traditional TTS methods are used, then the system structure is simpler, but the ability to produce natural-sounding speech with diverse qualities is limited

Engineering Contradiction:
Improvespeech quality variationVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal audio generation model that can produce diverse vocal attributes and speech qualities through a single neural network architecture. The model handles multiple tasks (different languages, tones, and vocal characteristics) without requiring separate systems, achieving versatility through unified design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10692484B1Text-to-speech (TTS) processing
Publication Date: 2020.06.23 AMAZON TECH INC
  • US10692484B1 patent drawing
  • US10692484B1 patent drawing
  • US10692484B1 patent drawing

AI summary

A speech model is trained using multi-task learning. A first task may correspond to how well predicted audio matches training audio; a second task may correspond to a metric of perceived audio quality. The speech model may include, during training, layers related to the second task that are discarded at runtime.