Virtual Character Voice Generation With Real-Time Latent Parameter Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating realistic and customizable voices for virtual characters is challenging due to difficulties in replicating desired characteristics such as volume, pace, pitch, and articulation.
Innovation Solution
A system comprising a server and client devices that utilize AI models to encode and decode audio inputs, allowing users to modify voice parameters and apply templates to generate voices for virtual characters in real-time, with additional features for facial animation synchronization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If AI models are used to encode and decode audio inputs for voice generation, then voice customization and realism are improved, but system complexity increases
Solution Approach 1:
The patent introduces latent space representations as an intermediary between audio input and output. The encoder transforms audio inputs into latent space vectors, which serve as compressed intermediaries containing essential voice characteristics. This intermediary representation simplifies the overall transformation process while maintaining high fidelity in voice generation.
Solution Approach 2:
The voice generation system is segmented into distinct modular components: an encoder module that processes audio inputs, a latent space representation layer, and a decoder module that generates output audio. This segmentation allows each component to be optimized independently while working together to achieve realistic voice generation.
2Adaptability or versatility
If multiple voice parameters are modified in real-time, then voice adaptability is improved, but processing time increases
Solution Approach 1:
The system performs preliminary encoding of audio inputs into latent space representations before any parameter modifications are applied. This preliminary action organizes the audio data into a compact form that facilitates efficient real-time parameter adjustments during the decoding phase, reducing overall processing time.
Solution Approach 2:
The patent implements dynamic parameter adjustment capabilities that allow voice characteristics to be modified in real-time during the decoding process. The system can adaptively adjust multiple voice parameters on-the-fly based on input conditions, maintaining versatility while optimizing processing efficiency through the structured latent space representation.
Data Source
AI summary
The present disclosure describes techniques of generating voices for virtual characters. A plurality of source sounds may be received. The plurality of source sounds may correspond to a plurality of frames of a video. The video may comprise a virtual character. The plurality of source sounds may be converted into a plurality of representations in a latent space using a first model. Each representation among the plurality of representations may comprise a plurality of parameters. The plurality of parameters may correspond to a plurality of sound features. A plurality of sounds may be generated in real time for the virtual character in the video based at least in part on modifying at least one of the plurality of parameters of each representation.


