Virtual Action Sequence Composition Using Video Motion Clip Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for synthesizing motion sequences of virtual objects rely heavily on manual design by artists, which is costly and results in rigid motions when keyword matching fails, leading to stationary virtual objects.
Innovation Solution
A method that extracts motion information from real videos to construct a library of continuous-motion clips, using AI-based neural networks to fuse semantic and motion attribute information into representation vectors for accurate retrieval and synthesis of motion sequences, allowing for natural and flexible simulations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual design by artists is used to create motion sequences, then motion quality and naturalness can be maintained, but the cost and time consumption increase significantly
Solution Approach 1:
The patent creates a motion database by copying and storing real human motion data from video recordings. Instead of manually designing each motion, the system captures actual human movements and stores them as reusable motion clips, allowing automated retrieval and composition of natural motions while significantly reducing artist workload and cost.
Solution Approach 2:
The patent performs preliminary action by pre-recording and storing diverse human motions in a database before they are needed. Motion clips are captured, processed, and organized in advance with associated keywords and metadata, enabling rapid automated retrieval and composition when motion sequences are required, eliminating the need for real-time manual design.
2Extent of automation
If keyword matching is used to retrieve motion clips, then automated motion sequence generation is achieved, but rigid motions occur when matching fails
Solution Approach 1:
The patent implements feedback mechanisms where the system evaluates the quality and suitability of retrieved motion clips. When keyword matching retrieves clips, the system assesses whether they naturally fit the context, and can adjust retrieval parameters or select alternative clips to ensure natural motion composition, preventing rigid or inappropriate motions when initial matching fails.
Solution Approach 2:
The patent changes retrieval parameters dynamically based on context. Instead of relying solely on fixed keyword matching, the system adjusts search parameters, weights, and selection criteria based on the specific motion composition needs, allowing flexible retrieval of natural motions that fit the contextual requirements even when exact keyword matches are unavailable.
3Device complexity
If a limited motion database is used, then database management complexity is reduced, but motion diversity and flexibility decrease
Solution Approach 1:
The patent segments the motion database into organized categories based on motion types, actions, contexts, and other relevant dimensions. Each motion clip is tagged with multiple keywords and metadata, allowing the database to scale to large sizes while maintaining manageable organization and enabling flexible, diverse motion retrieval through multi-dimensional search capabilities.
Data Source
Figure 1~2A
Figure 2B~2C
Figure 3A~3B
AI summary
Disclosed are a method, device, and apparatus for compositing an action sequence of a virtual object, and a computer readable storage medium. The method comprises: obtaining description information of an action sequence of a virtual object; determining, on the basis of the description information and a continuous action clip library constructed from video materials, a set of continuous action clips similar to at least some actions in the action sequence; and compositing the action sequence of the virtual object on the basis of the set of continuous action clips, each continuous action clip in the continuous action clip library comprising a unique identifier of the continuous action clip, action information of the continuous action clip, and a representation vector corresponding to the continuous action clip.