Text-to-Speech Intonation Modification via Human Audio Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech generated audio signals often sound robotic and lack emotional inflection, making it difficult to achieve a natural, human-like audio output, especially in applications requiring professional-sounding voices.

Innovation Solution

A method and system that dynamically modify intonations in a text-to-speech generated audio waveform to match those detected in a human-produced audio waveform, using AI and machine learning to analyze and adjust frequency characteristics, pitch, volume, and emotional cues, allowing for the creation of more lifelike and professional-sounding audio signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If text-to-speech generates audio signals, then automated speech production is achieved, but the audio sounds robotic and lacks natural intonation

Engineering Contradiction:
Improveautomated speech productionVSAvoidnatural intonation accuracy
Core Design Contradiction:
Extent of automationVSManufacturing precision

Solution Approach 1:

The patent extracts intonation patterns from human-recorded audio waveforms and applies them to text-to-speech generated audio. By copying the intonation contours, pitch variations, and rhythmic patterns from human speech examples, the system transforms robotic synthetic speech into natural-sounding audio while maintaining automated production capabilities

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system modifies audio parameters including pitch, volume, and frequency characteristics dynamically. By adjusting these parameters based on detected human intonation patterns, the synthetic speech transitions from robotic to natural-sounding, resolving the contradiction between automation and naturalness

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If traditional audio editing techniques are used, then audio signals can be modified, but advanced knowledge and manual expertise are required

Engineering Contradiction:
Improveaudio editing accessibilityVSAvoidediting skill requirement
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The system performs automated intonation extraction and application without requiring manual editing. The machine learning model automatically analyzes human audio examples, identifies intonation patterns, and applies them to synthetic speech, eliminating the need for skilled manual editing while maintaining high-quality results

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual audio editing operations with automated machine learning-based processing. Instead of requiring editors to manually adjust audio parameters, the system uses AI algorithms to automatically extract and apply intonation patterns, substituting mechanical manual editing with automated intelligent processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20230386446A1Modifying an audio signal to incorporate a natural-sounding intonation
Publication Date: 2023.11.30 AUTHENTICVOICE INC
  • US20230386446A1 patent drawing
  • US20230386446A1 patent drawing
  • US20230386446A1 patent drawing

AI summary

Techniques for modifying intonations in audio files are disclosed. A first audio waveform and a second audio waveform are accessed. A first intonation in the first audio waveform is identified, and a second intonation in the second audio waveform is identified. The first and second intonations correspond to the same unit of pronunciation. The first intonation is modified until it sufficiently matches the second intonation.