Audio Avatar Creation from Human Voices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Small form factor electronic devices, such as smartphones and smartwatches, face challenges in implementing conventional user interfaces due to limited screen size, leading to difficulties in using voice inputs and outputs effectively, and existing voice-altering technologies primarily focus on speed and frequency changes for entertainment purposes, lacking the ability to transform one voice to mimic another for enhanced user experience or anonymity.
Innovation Solution
A voice analysis and altering module characterizes a subject voice and transforms it to mimic a target voice, maintaining the verbal message while changing the voice, allowing for the creation of audio avatars that can be used in various applications, including social networks, GPS guidance, and voice-based personal assistants, enabling users to express themselves authentically without being confined to their biological voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional user interfaces relying on text or touch input are used, then devices can maintain simplicity, but user interaction becomes difficult on devices with limited screen size
Solution Approach 1:
The patent replaces mechanical/touch-based input methods with acoustic field interaction. Users communicate with devices through voice commands, and devices respond through synthesized speech, eliminating the need for physical keyboard or touchscreen interaction and overcoming the limitations of small screen real estate.
Solution Approach 2:
The voice input/output system serves multiple functions: it enables complex commands without requiring visual display, provides hands-free operation, and creates a more natural interaction paradigm that works effectively on devices with minimal display area.
2Adaptability or versatility
If voice-altering capabilities are added for entertainment purposes, then user engagement increases, but the technology lacks practical applications for anonymity and personalization
Solution Approach 1:
The system transforms voice by modifying acoustic parameters including pitch, speed, and timbre characteristics. By adjusting these parameters, the system can alter a user's voice to sound like different people, enabling both entertainment effects and practical applications such as anonymity protection and personalized audio feedback.
Solution Approach 2:
The voice alteration system creates synthetic copies of target voices by analyzing and replicating their acoustic characteristics. This allows the system to generate realistic voice imitations that can be applied to protect user anonymity or provide personalized experiences without requiring actual recordings from all possible target voices.
3Adaptability or versatility
If voice speed and frequency are changed for entertainment, then voice variety is achieved, but the verbal message and original meaning are distorted
Solution Approach 1:
The voice processing system segments the input audio into individual phonemes or speech units. Each segment is then processed independently to apply the desired voice transformation while preserving the original phonetic content. This segmentation approach allows the system to change voice characteristics without distorting the underlying verbal message.
Solution Approach 2:
The system uses an intermediary processing layer that separates the voice characteristics from the semantic content. By treating these aspects independently and recombining them with the transformed voice, the system maintains message fidelity while achieving the desired voice variation.
Data Source
AI summary
A subject voice is characterized and altered to mimic a target voice while maintaining the verbal message of the subject voice. Thus, the words and message are the same as in the original voice, but the voice that conveys the words and message in the altered voice is different. Audio signals corresponding to the altered voice are output, for example to an application for playback to a user, or to another application or device for subsequent playback by the user or someone else. In one embodiment, the altered voice is posted to a social network. In other embodiments, the altered voice is used by other software applications or consumer electronics applications, such as GPS guidance systems, ebook readers, voice-based intelligent personal assistants, chat applications, and/or others that use voice as an input or output.


