Multimodal Digital Human Driving for Flexible Pose Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital human driving technologies are limited by the need for single-modality data, strict pose requirements, and poor interactivity, leading to unrealistic representations and user experience issues when interacting with real persons.

Innovation Solution

A method that integrates image and audio information to extract motion features, using a character generator for flexible digital human driving, enabling seamless transitions and accurate representations through feature fusion and modality switching.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If single-modality data (image, voice, or text) is used to drive the digital human, then the system complexity is low, but the interactivity and representation effect are insufficient

Engineering Contradiction:
ImproveinteractivityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple data modalities (image, audio, text) into a unified driving system. The system integrates image information for pose estimation, audio information for voice-driven animation, and text information for lip-sync, combining these modalities to drive the digital human simultaneously, thereby enhancing interactivity and representation effect while managing system complexity through modular architecture

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal digital human driving system that can accept multiple types of input data (image, audio, text) and process them through a common framework. The system is designed to handle different modalities flexibly, allowing it to adapt to various interaction scenarios and improve versatility without requiring separate dedicated systems for each modality

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If image-based driving is used, then the pose requirements are strict, but the digital human cannot be effectively driven when the target person leaves the camera screen or the pose is too large

Engineering Contradiction:
Improvepose accuracyVSAvoiddriving reliability
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent introduces pose estimation technology as an intermediary between the captured image and the digital human driving. The pose estimation module processes the image information to extract key pose parameters, acting as a mediator that translates raw image data into actionable driving signals. This intermediary layer helps maintain driving reliability even when the target person moves out of frame or changes pose significantly, by using estimated pose data rather than requiring perfect direct observation

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If text-to-voice conversion is used to drive the digital human, then the implementation is reliable, but the generated digital human has weak interactivity

Engineering Contradiction:
Improveimplementation reliabilityVSAvoidinteractivity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent combines text-to-voice conversion with audio information processing and image-based pose estimation. Instead of relying solely on text-to-voice conversion, the system merges multiple driving inputs including captured audio and visual pose data, creating a more interactive and realistic digital human experience while maintaining the reliability of text-based input through the unified multi-modal framework

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250329094A1Digital human driving method, digital human driving device and storage medium
Publication Date: 2025.10.23 ZTE CORP
  • US20250329094A1 patent drawing
  • US20250329094A1 patent drawing
  • US20250329094A1 patent drawing

AI summary

A digital human driving method, a digital human driving device, and a storage medium are disclosed. The digital human driving method may include: acquiring image information and audio information of a target object; performing recognition and determination on the image information and the audio information to obtain a determination result; performing feature extraction processing on the image information and/or the audio information according to the determination result to obtain a first motion feature and/or a second motion feature; inputting the first motion feature and/or the second motion feature and a digital human base image into a character generator; and performing driving processing on the digital human base image through the character generator, and outputting a first digital human driving image.