Virtual Character Motion Generation Using Audio-Text Semantic Tags

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating virtual character motions are inefficient and lack accuracy, particularly in scenarios where manual intervention is required, and they fail to effectively incorporate semantic information from both audio and text, leading to poor motion generation efficiency and accuracy.

Innovation Solution

A method and apparatus that utilize a motion library constructed by clustering sample motion clips based on semantic tags derived from text and audio, enabling precise and efficient generation of virtual character motions by retrieving and synthesizing motion data from a preset library using dual-modality information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual intervention is used for motion generation, then motion accuracy can be improved, but motion generation efficiency deteriorates

Engineering Contradiction:
Improvemotion accuracyVSAvoidmotion generation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-construction of a motion library containing motion data from multiple motion categories before actual motion generation is needed. The library is built offline by collecting, processing, and organizing motion data in advance, so that during runtime, the system can directly retrieve and synthesize motions without manual intervention, thus improving efficiency while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by transitioning from manual motion generation parameters to automated retrieval parameters. The system uses semantic tags, audio features, and text information as retrieval parameters to automatically select appropriate motion data from the pre-construction library, replacing manual control parameters with automated multi-dimensional matching parameters

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If semantic information from audio and text is incorporated, then motion accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvemotion accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex semantic information processing into distinct modules: audio feature extraction module, text semantic analysis module, semantic tag determination module, and motion retrieval module. Each module handles a specific aspect of the information processing, making the overall system more manageable and easier to implement despite the complexity of handling multi-modal semantic information

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces semantic tags as an intermediary between the raw audio/text input and the motion data retrieval. The semantic tags serve as a bridge that translates complex semantic information into a format that can be used for efficient motion category matching and retrieval, simplifying the interface between different system components

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If a preset motion library is constructed through clustering, then motion retrieval efficiency is improved, but construction time increases

Engineering Contradiction:
Improvemotion retrieval efficiencyVSAvoidconstruction time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing motion data clustering and library construction in advance before the motion generation system is put into operation. The offline construction process includes collecting motion data, extracting features, clustering into motion categories, and organizing into a searchable library structure, so that the time-consuming processing is completed beforehand rather than during runtime

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating a structured representation of motion data in the form of motion categories and semantic mappings. Instead of processing raw motion data during retrieval, the system copies and stores pre-processed motion characteristics, features, and category associations that can be quickly matched during runtime without re-processing the original complex motion data

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250278881A1Method and apparatus for generating motion of virtual character, and method and apparatus for constructing motion library of virtual character
Publication Date: 2025.09.04 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20250278881A1 patent drawing
  • US20250278881A1 patent drawing
  • US20250278881A1 patent drawing

AI summary

A method and an apparatus for generating a motion of a virtual character, and a method and an apparatus for constructing a motion library of a virtual character are provided, and belong to the field of computer technologies. The method for generating a motion of a virtual character includes: obtaining audio and text of a virtual character, the text indicating semantic information of the audio (201); determining a semantic tag of the text based on the text (202); retrieving a motion category matching the semantic tag and motion data belonging to the motion category from a preset motion library (203); and generating a motion sequence of the virtual character based on the motion data (204).