Audio Signal Phoneme Mapping for Speech Mimicry
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for processing audio signals in computing devices are limited in replicating speech in a way that mimics the original speaker, often resulting in computer-generated sounds that fail to convincingly portray speech as if it were produced by a character or entity.
Innovation Solution
A computing device captures an input audio signal, transforms it into digital segments, identifies phoneme features, and maps these to phoneme models from a library, combining them to generate an output signal that mimics the user's speech in the voice of a specific character, maintaining the original's tone, inflection, and accent.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional filter-based audio transformation is used, then processing simplicity is maintained, but the output sounds computer-generated and fails to convincingly replicate speech patterns
Solution Approach 1:
The audio signal is divided into phoneme segments, each representing a distinct speech sound. The system identifies and separates individual phonemes from the input audio, then maps them to corresponding phoneme models in a library, enabling precise control over speech synthesis while maintaining natural pronunciation patterns
Solution Approach 2:
The system transforms audio processing from filter-based frequency manipulation to phoneme-based parameter mapping. By changing the fundamental approach from continuous signal filtering to discrete phoneme identification and selection, the system achieves more accurate speech replication with character voice characteristics
2Reliability
If phoneme-based mapping with voice data is implemented, then speech pattern replication is improved, but processing time and computational resources increase
Solution Approach 1:
Phoneme models are pre-recorded and stored in a library before runtime processing. By preparing the phoneme models in advance and organizing them in a searchable database, the system eliminates the need for real-time phoneme synthesis, significantly reducing processing time during actual audio transformation
Solution Approach 2:
The system uses pre-recorded phoneme models as templates to copy and reconstruct speech. Instead of generating phonemes algorithmically during processing, the system selects and combines existing phoneme recordings from the library, maintaining natural speech characteristics while minimizing computational overhead
Data Source
AI summary
Approaches are described for generating an audio signal that mimics speech captured a computing device. An input audio signal (e.g., a speech signal) can be transformed from the time domain into another domain, to generate one or more audio signal segments, where each segment can correspond to a window of time. The device can then determine, for each audio signal segment, a feature characteristic of the audio signal, such as a phoneme. Each one of segments can be mapped, based at least in part on the respective feature characteristic, to a model audio signal. The device can then generate an output audio signal including each model audio signal as determined by the mapping, where the output audio signal is in a sequence associated with the input audio signal.


