Audio Signal Phoneme Mapping for Speech Mimicry

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for processing audio signals in computing devices are limited in replicating speech in a way that mimics the original speaker, often resulting in computer-generated sounds that fail to convincingly portray speech as if it were produced by a character or entity.

Innovation Solution

A computing device captures an input audio signal, transforms it into digital segments, identifies phoneme features, and maps these to phoneme models from a library, combining them to generate an output signal that mimics the user's speech in the voice of a specific character, maintaining the original's tone, inflection, and accent.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional filter-based audio transformation is used, then processing simplicity is maintained, but the output sounds computer-generated and fails to convincingly replicate speech patterns

Engineering Contradiction:
Improvespeech replication authenticityVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio signal is divided into phoneme segments, each representing a distinct speech sound. The system identifies and separates individual phonemes from the input audio, then maps them to corresponding phoneme models in a library, enabling precise control over speech synthesis while maintaining natural pronunciation patterns

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transforms audio processing from filter-based frequency manipulation to phoneme-based parameter mapping. By changing the fundamental approach from continuous signal filtering to discrete phoneme identification and selection, the system achieves more accurate speech replication with character voice characteristics

Inventive Principle:
Principle #35Parameter changes

2Reliability

If phoneme-based mapping with voice data is implemented, then speech pattern replication is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvespeech pattern accuracyVSAvoidaudio processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Phoneme models are pre-recorded and stored in a library before runtime processing. By preparing the phoneme models in advance and organizing them in a searchable database, the system eliminates the need for real-time phoneme synthesis, significantly reducing processing time during actual audio transformation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses pre-recorded phoneme models as templates to copy and reconstruct speech. Instead of generating phonemes algorithmically during processing, the system selects and combines existing phoneme recordings from the library, maintaining natural speech characteristics while minimizing computational overhead

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9129602B1Mimicking user speech patterns
Publication Date: 2015.09.08 AMAZON TECH INC
  • US9129602B1 patent drawing
  • US9129602B1 patent drawing
  • US9129602B1 patent drawing

AI summary

Approaches are described for generating an audio signal that mimics speech captured a computing device. An input audio signal (e.g., a speech signal) can be transformed from the time domain into another domain, to generate one or more audio signal segments, where each segment can correspond to a window of time. The device can then determine, for each audio signal segment, a feature characteristic of the audio signal, such as a phoneme. Each one of segments can be mapped, based at least in part on the respective feature characteristic, to a model audio signal. The device can then generate an output audio signal including each model audio signal as determined by the mapping, where the output audio signal is in a sequence associated with the input audio signal.