Virtual Character Voice Generation With Real-Time Latent Parameter Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating realistic and customizable voices for virtual characters is challenging due to difficulties in replicating desired characteristics such as volume, pace, pitch, and articulation.

Innovation Solution

A system comprising a server and client devices that utilize AI models to encode and decode audio inputs, allowing users to modify voice parameters and apply templates to generate voices for virtual characters in real-time, with additional features for facial animation synchronization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If AI models are used to encode and decode audio inputs for voice generation, then voice customization and realism are improved, but system complexity increases

Engineering Contradiction:
Improvevoice generation realismVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent introduces latent space representations as an intermediary between audio input and output. The encoder transforms audio inputs into latent space vectors, which serve as compressed intermediaries containing essential voice characteristics. This intermediary representation simplifies the overall transformation process while maintaining high fidelity in voice generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The voice generation system is segmented into distinct modular components: an encoder module that processes audio inputs, a latent space representation layer, and a decoder module that generates output audio. This segmentation allows each component to be optimized independently while working together to achieve realistic voice generation.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If multiple voice parameters are modified in real-time, then voice adaptability is improved, but processing time increases

Engineering Contradiction:
Improvevoice parameter adaptabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary encoding of audio inputs into latent space representations before any parameter modifications are applied. This preliminary action organizes the audio data into a compact form that facilitates efficient real-time parameter adjustments during the decoding phase, reducing overall processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic parameter adjustment capabilities that allow voice characteristics to be modified in real-time during the decoding process. The system can adaptively adjust multiple voice parameters on-the-fly based on input conditions, maintaining versatility while optimizing processing efficiency through the structured latent space representation.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250371779A1Voice generation for virtual characters
Publication Date: 2025.12.04 LEMON INC(GB)
  • US20250371779A1 patent drawing
  • US20250371779A1 patent drawing
  • US20250371779A1 patent drawing

AI summary

The present disclosure describes techniques of generating voices for virtual characters. A plurality of source sounds may be received. The plurality of source sounds may correspond to a plurality of frames of a video. The video may comprise a virtual character. The plurality of source sounds may be converted into a plurality of representations in a latent space using a first model. Each representation among the plurality of representations may comprise a plurality of parameters. The plurality of parameters may correspond to a plurality of sound features. A plurality of sounds may be generated in real time for the virtual character in the video based at least in part on modifying at least one of the plurality of parameters of each representation.