Virtual Character Voice Generation With Real-Time Parameter Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating realistic and customizable voices for virtual characters is challenging due to difficulties in replicating desired characteristics such as volume, pace, pitch, and articulation.
Innovation Solution
A system comprising a server and client devices that utilize AI models to encode and decode audio inputs, allowing users to modify parameters and apply templates to generate voices for virtual characters in real-time, incorporating user camera data for improved facial animation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional voice synthesis methods are used, then the system is simple, but the voice characteristics (volume, pace, pitch, articulation) are not realistic or customizable
Solution Approach 1:
The patent applies parameter changes by encoding voice characteristics into multiple adjustable parameters including volume, pace, pitch, and articulation. These parameters are modified during the voice generation process to achieve realistic and customizable voice outputs for virtual characters, directly resolving the contradiction between voice accuracy and system complexity.
Solution Approach 2:
The voice generation system is segmented into distinct functional modules: an encoder that processes input audio, a parameter modification module that adjusts voice characteristics, and a decoder that generates the final voice output. This segmentation allows complex voice synthesis to be achieved through manageable, independent components.
2Productivity
If real-time voice generation is implemented, then user interaction is enhanced, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary encoding of the input audio signal into a structured representation with extracted voice parameters before actual voice generation occurs. This preliminary processing organizes the data in advance, enabling faster real-time synthesis when the decoded voice output is needed, thus reducing overall processing time.
Data Source
AI summary
The present disclosure describes techniques of generating voices for virtual characters. A plurality of source sounds may be received. The plurality of source sounds may correspond to a plurality of frames of a video. The video may comprise a virtual character. The plurality of source sounds may be converted into a plurality of representations in a latent space using a first model. Each representation among the plurality of representations may comprise a plurality of parameters. The plurality of parameters may correspond to a plurality of sound features. A plurality of sounds may be generated in real time for the virtual character in the video based at least in part on modifying at least one of the plurality of parameters of each representation.


