Virtual Character Audio-Video Synchronization via Dynamic Action Duration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital human technologies face challenges in synchronizing audio information and video actions, leading to a lack of expressiveness as actions are often not completed when audio finishes or vice versa.

Innovation Solution

A virtual character control method that acquires target keywords from a text, determines action start and end positions, predicts audio broadcasting duration, and matches it with a target action file to ensure the video duration matches the audio, thereby generating multimedia information that synchronizes audio and video effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Duration of action of moving object

If a virtual character performs actions with fixed duration, then the action execution is simple and straightforward, but the audio and video timing cannot be dynamically adjusted to match different speech lengths

Engineering Contradiction:
Improveaction durationVSAvoidaudio-video synchronization adaptability
Core Design Contradiction:
Duration of action of moving objectVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by making the virtual character's action duration variable rather than fixed. The system dynamically adjusts the action duration based on the predicted speech duration, allowing the virtual character to adapt its performance timing to match different speech lengths while maintaining natural expression.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the time duration parameter of virtual character actions based on predicted speech duration. By calculating the expected speech length and selecting or adjusting action files with corresponding durations, the system ensures synchronization between audio and video without requiring manual timing adjustments.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If the system selects action files without considering audio duration, then the action selection process is simple and fast, but the video actions will not be synchronized with the audio broadcasting

Engineering Contradiction:
Improveaction selection efficiencyVSAvoidaudio-video timing precision
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system performs preliminary action by predicting the speech duration before the virtual character performs the action. This advance calculation allows the system to pre-select action files with matching durations, ensuring synchronization is achieved automatically without requiring post-processing adjustments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback by comparing the predicted speech duration with the duration of available action files. Based on this comparison, the system selects the most appropriate action file that matches the speech timing, creating a closed-loop selection process that ensures synchronization while maintaining efficiency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250095257A1Virtual character control method, apparatus, device and storage medium
Publication Date: 2025.03.20 CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
  • US20250095257A1 patent drawing
  • US20250095257A1 patent drawing
  • US20250095257A1 patent drawing

AI summary

The present disclosure relates to a virtual character control method and apparatus, a device and a storage medium. In the present disclosure, one or more target keywords are acquired from a preset text, and a first target character and a second target character are determined from the preset text for each of the one or more target keywords. Further, an audio broadcasting duration from the first target character to the second target character is predicted, and a target action file is determined from one or more preset action files corresponding to the target keyword according to the audio broadcasting duration, so that a time duration of the target action video matches the audio broadcasting duration, where the target action video is obtained by driving, according to the target action file, the virtual character to perform a target action.