Real-Time Occlusion Masking for Augmented Reality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Augmented Reality (AR) systems are inefficient and time-consuming in capturing image data of objects occluded by real-world objects, often requiring pre-modeling or scanning of the entire environment, which limits their flexibility and immersive interactive experiences.
Innovation Solution
An AR system that captures image data in real-time during an interactive session, using sensors to detect user movements and changes in perspective, allowing for the generation of virtual representations of occluded objects without pre-existing environmental maps, enabling flexible AR effects in various environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If pre-modeling or scanning of the entire environment is performed to capture image data of occluded objects, then the accuracy of virtual representations is improved, but the time required and system complexity increase significantly
Solution Approach 1:
The system performs preliminary capture of image data from multiple perspectives during the AR session initialization phase, storing this data for later use. This allows the actual virtual representation generation to occur quickly when needed, without requiring time-consuming pre-scanning of the entire environment.
Solution Approach 2:
The system introduces an intermediate step of capturing and storing raw image data from multiple cameras at different perspectives, which serves as a mediator between the physical environment and the final virtual representations. This intermediate data repository enables quick generation of accurate virtual objects without direct real-time scanning.
2Measurement precision
If pre-modeling or scanning of the entire environment is performed to capture image data of occluded objects, then the accuracy of virtual representations is improved, but the system complexity and resource requirements increase
Solution Approach 1:
The system divides the environment into multiple discrete perspectives captured by separate cameras positioned at different locations. Each camera captures a specific viewpoint, and these segmented views are later combined to reconstruct occluded objects. This segmentation approach simplifies the overall system architecture compared to attempting to scan and model the entire environment as a single complex unit.
Solution Approach 2:
The system uses multiple cameras that serve dual purposes: capturing the visible field of view for standard AR display and simultaneously capturing occluded regions for generating virtual representations. This multi-functionality eliminates the need for separate scanning devices or specialized equipment, reducing system complexity while maintaining accuracy.
3Measurement precision
If environmental maps are required for generating AR effects, then the quality of occluded object representation is improved, but the adaptability to various environments without pre-scanning decreases
Solution Approach 1:
The system dynamically captures image data from multiple perspectives during the AR session based on the user's current viewpoint and the detected occluded objects. Rather than relying on static pre-built environmental maps, the system adapts its data capture strategy in real-time to the specific environment and user needs, enabling operation in previously unseen locations while maintaining representation quality.
Solution Approach 2:
The system performs self-service by automatically capturing and processing image data from its own multiple camera perspectives without requiring external scanning equipment or pre-existing environmental models. The AR device itself generates the necessary data for creating accurate virtual representations of occluded objects in any environment it encounters.
Data Source
AI summary
To provide masks for augmented reality (AR) effects in real-time, the embodiments described herein capture a plurality of images of an environment that includes one or more real-world objects after initiating an AR session with an AR device located in the environment. Upon detecting a trigger to generate an AR effect targeting a first real-world object of the one or more real-world objects, a first virtual representation of the first real-world object is generated, based on the plurality of images of the environment, and a second virtual representation of a second real-world object of the one or more real-world objects is generated based on the plurality of images of the environment, wherein the second real-world object is at least partially occluded by the first real-world object.


