Multimodal Digital Human Driving for Flexible Pose Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital human driving technologies are limited by the need for single-modality data, strict pose requirements, and poor interactivity, leading to unrealistic representations and user experience issues when interacting with real persons.
Innovation Solution
A method that integrates image and audio information to extract motion features, using a character generator for flexible digital human driving, enabling seamless transitions and accurate representations through feature fusion and modality switching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If single-modality data (image, voice, or text) is used to drive the digital human, then the system complexity is low, but the interactivity and representation effect are insufficient
Solution Approach 1:
The patent merges multiple data modalities (image, audio, text) into a unified driving system. The system integrates image information for pose estimation, audio information for voice-driven animation, and text information for lip-sync, combining these modalities to drive the digital human simultaneously, thereby enhancing interactivity and representation effect while managing system complexity through modular architecture
Solution Approach 2:
The patent creates a universal digital human driving system that can accept multiple types of input data (image, audio, text) and process them through a common framework. The system is designed to handle different modalities flexibly, allowing it to adapt to various interaction scenarios and improve versatility without requiring separate dedicated systems for each modality
2Manufacturing precision
If image-based driving is used, then the pose requirements are strict, but the digital human cannot be effectively driven when the target person leaves the camera screen or the pose is too large
Solution Approach 1:
The patent introduces pose estimation technology as an intermediary between the captured image and the digital human driving. The pose estimation module processes the image information to extract key pose parameters, acting as a mediator that translates raw image data into actionable driving signals. This intermediary layer helps maintain driving reliability even when the target person moves out of frame or changes pose significantly, by using estimated pose data rather than requiring perfect direct observation
3Reliability
If text-to-voice conversion is used to drive the digital human, then the implementation is reliable, but the generated digital human has weak interactivity
Solution Approach 1:
The patent combines text-to-voice conversion with audio information processing and image-based pose estimation. Instead of relying solely on text-to-voice conversion, the system merges multiple driving inputs including captured audio and visual pose data, creating a more interactive and realistic digital human experience while maintaining the reliability of text-based input through the unified multi-modal framework
Data Source
AI summary
A digital human driving method, a digital human driving device, and a storage medium are disclosed. The digital human driving method may include: acquiring image information and audio information of a target object; performing recognition and determination on the image information and the audio information to obtain a determination result; performing feature extraction processing on the image information and/or the audio information according to the determination result to obtain a first motion feature and/or a second motion feature; inputting the first motion feature and/or the second motion feature and a digital human base image into a character generator; and performing driving processing on the digital human base image through the character generator, and outputting a first digital human driving image.


