Video Processing Apparatus for Mixed Reality Training with Dynamic Body Part Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing XR training techniques are limited in that they can only teach motions based on pre-stored motion information of a partner or specific body parts like fingers, making them not user-friendly for teaching motions of other body parts.
Innovation Solution
A video processing apparatus and system that allows a first user to acquire and transmit a real video, receive and generate motion information of a second user, and create a mixed video by aligning and combining the real video with virtual objects representing the second user's body parts, enabling more comprehensive XR training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If pre-stored motion information of a partner is used for XR training, then training can be performed, but the training is limited only to the motion information that has been stored in advance and cannot adapt to new motions
Solution Approach 1:
The system performs preliminary action by capturing and storing motion information in advance through the video processing apparatus. The apparatus records real videos of users performing various motions, extracts motion information, and stores it for future training sessions. This preliminary capture and storage of diverse motion data enables the system to adapt to different training needs without requiring real-time motion capture, thus resolving the contradiction between adaptability and time loss.
2Adaptability or versatility
If virtual objects of specific body parts like fingers are generated, then specific motion teaching can be performed, but it is impossible to teach motion of other body parts such as arms
Solution Approach 1:
The video processing apparatus is designed with universal functionality to capture, process, and generate virtual objects for any body part. The apparatus uses general-purpose video capture and motion extraction algorithms that work across different body parts (fingers, arms, legs, etc.) without requiring specialized processing for each body part. This multi-functional design enables comprehensive body part coverage while avoiding the complexity of creating separate specialized systems for each body part type.
Solution Approach 2:
The system implements dynamic virtual object generation that adapts to different body parts based on the captured motion information. Rather than using fixed, static virtual objects, the system dynamically creates and adjusts virtual representations of body parts according to the real-time or pre-captured motion data. This dynamic approach allows the system to handle various body parts flexibly, from fingers to arms to entire body movements, resolving the contradiction between versatility and complexity.
3Adaptability or versatility
If real video is transmitted to another video processing apparatus and motion information is received and processed, then comprehensive XR training can be achieved, but the system complexity increases
Solution Approach 1:
The video processing system is segmented into distinct functional modules: a first video processing apparatus for capturing real videos and extracting motion information, a second video processing apparatus for receiving and processing the data, and a virtual object generation component. This segmentation allows each module to perform its specific function efficiently, reducing overall system complexity while maintaining comprehensive training capabilities. The modular architecture enables independent optimization and easier maintenance of each component.
Data Source
AI summary
A video processing apparatus that is used by a first user and processes a video. A real video of a real space visually recognized by the first user is acquired and transmitted to another video processing apparatus used by a second user different from the first user. Motion information concerning motion of the second user is received from the other video processing apparatus. A virtual object of a body part of the second user, which can be displayed in the video, is generated based on the received motion information. A mixed video is generated by mixing the real video and the virtual object. When generating the mixed video, respective positions of the real video and the virtual object are aligned, and the mixed video is generated.


