Virtual Character Voice Generation With Real-Time Parameter Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating realistic and customizable voices for virtual characters is challenging due to difficulties in replicating desired characteristics such as volume, pace, pitch, and articulation.

Innovation Solution

A system comprising a server and client devices that utilize AI models to encode and decode audio inputs, allowing users to modify parameters and apply templates to generate voices for virtual characters in real-time, incorporating user camera data for improved facial animation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional voice synthesis methods are used, then the system is simple, but the voice characteristics (volume, pace, pitch, articulation) are not realistic or customizable

Engineering Contradiction:
Improvevoice characteristic accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by encoding voice characteristics into multiple adjustable parameters including volume, pace, pitch, and articulation. These parameters are modified during the voice generation process to achieve realistic and customizable voice outputs for virtual characters, directly resolving the contradiction between voice accuracy and system complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The voice generation system is segmented into distinct functional modules: an encoder that processes input audio, a parameter modification module that adjusts voice characteristics, and a decoder that generates the final voice output. This segmentation allows complex voice synthesis to be achieved through manageable, independent components.

Inventive Principle:
Principle #1Segmentation

2Productivity

If real-time voice generation is implemented, then user interaction is enhanced, but processing time and computational resources increase

Engineering Contradiction:
Improvevoice generation speedVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary encoding of the input audio signal into a structured representation with extracted voice parameters before actual voice generation occurs. This preliminary processing organizes the data in advance, enabling faster real-time synthesis when the decoded voice output is needed, thus reducing overall processing time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12412559B2Voice generation for virtual characters
Publication Date: 2025.09.09 LEMON INC(GB)
  • US12412559B2 patent drawing
  • US12412559B2 patent drawing
  • US12412559B2 patent drawing

AI summary

The present disclosure describes techniques of generating voices for virtual characters. A plurality of source sounds may be received. The plurality of source sounds may correspond to a plurality of frames of a video. The video may comprise a virtual character. The plurality of source sounds may be converted into a plurality of representations in a latent space using a first model. Each representation among the plurality of representations may comprise a plurality of parameters. The plurality of parameters may correspond to a plurality of sound features. A plurality of sounds may be generated in real time for the virtual character in the video based at least in part on modifying at least one of the plurality of parameters of each representation.