Generative AI AR Video Synthesis for Realistic Object Movement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image and video generation systems using ML models often produce unrealistic and resource-intensive XR experiences due to limitations in generating high-quality, realistic images and videos, leading to wasted resources and reduced user engagement.

Innovation Solution

A generative ML model processes movement videos and target images to create realistic AR experiences by overlaying a user's face on an object, reducing the time and expense required for high-quality image and video production.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional ML models are used for image and video generation in XR systems, then the system can produce content, but the generated images and videos appear unrealistic and require excessive computational resources

Engineering Contradiction:
Improverealism of generated contentVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system segments the content generation process into two distinct stages: (1) generating a low-resolution skeleton structure that captures the essential movement and pose information, and (2) synthesizing high-resolution realistic images and videos by applying generative models only to the skeleton framework. This segmentation allows the computationally intensive generative modeling to be applied efficiently to simplified structural data rather than full-resolution input, thereby reducing overall computational resource consumption while maintaining high realism in the final output.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by first creating a simplified skeleton representation of the subject's movement and pose before generating the final realistic content. This preliminary skeleton generation captures the essential structural information needed for realistic rendering, allowing the subsequent high-resolution generative modeling to focus only on refining the appearance rather than creating the basic structure from scratch. This preliminary structuring significantly reduces the computational burden of the final generation step.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If high-quality realistic images and videos are generated using existing ML models, then user engagement improves, but the time and expense for production increases significantly

Engineering Contradiction:
Improvecontent generation speedVSAvoidproduction time for high-quality content
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The content generation process is segmented into efficient stages where a skeleton structure is first created to define the essential movement and pose, then high-resolution realistic images and videos are generated by applying generative models only to this simplified framework. This segmentation enables rapid production of high-quality content because the computationally intensive generative modeling operates on the compact skeleton data rather than full-resolution inputs, dramatically reducing production time while maintaining high realism and user engagement.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses the skeleton structure as a template or copy framework that captures the essential movement and pose information. The generative models then create realistic high-resolution content by copying and refining this skeletal structure, filling in the detailed appearances while maintaining the underlying movement patterns and poses. This copying approach from simplified skeleton to detailed realistic content enables fast production of high-quality media without requiring time-consuming generation from scratch.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If existing XR systems generate virtual objects, then the system can create augmented reality experiences, but the virtual objects fail to interact realistically with the real-world environment

Engineering Contradiction:
Improveinteraction realism between virtual and real objectsVSAvoidquality of virtual object rendering
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system employs dynamic skeleton structures that capture real-time movement and pose information from the real-world environment. These dynamic skeletons serve as adaptive frameworks that automatically adjust to match the motion patterns and spatial relationships of real objects. The generative models then synthesize realistic virtual objects that dynamically interact with the real-world environment by following these motion patterns, creating immersive AR experiences where virtual and real objects appear to interact naturally and realistically.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where the generated realistic images and videos are continuously refined based on the skeleton structure and real-world context. The generative models receive feedback about the skeleton's movement patterns and spatial positioning, allowing them to adjust the rendering of virtual objects to better match the real-world environment. This feedback loop ensures that virtual objects interact realistically with real objects, maintaining consistent appearance and behavior throughout the AR experience.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250329118A1Generative ai experience with movement
Publication Date: 2025.10.23 SNAP INC
  • US20250329118A1 patent drawing
  • US20250329118A1 patent drawing
  • US20250329118A1 patent drawing

AI summary

Systems and methods are provided for generating an augmented reality (AR) experience. The systems and methods receive a video depicting movement of a humanoid and a target image depicting an object. The systems and methods process, by a generative machine learning (ML) model, the video and the target image to generate a new video depicting the object performing the movement. The systems and methods generate the AR experience using the new video to overlay a face of a user on a portion of the new video.