Spatial Document System for XR Object Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing extended reality (XR) systems face challenges in precise tracking of real-world objects and are cumbersome for users to configure XR content, leading to imprecise or unreliable tracking and time-consuming content generation.
Innovation Solution
A computing system generates and interacts with 'spatial documents' by associating text with physical objects in real-world scenes using image, sensor, and audio data, allowing users to create immersive experiences in AR, MR, and VR environments through a hierarchical document structure that can be easily authored and expanded.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional XR tracking techniques are used, then real-world objects can be tracked, but tracking precision and reliability are insufficient
Solution Approach 1:
The patent introduces physical markers as intermediary objects that facilitate precise tracking. These markers serve as mediators between the camera system and the real-world environment, providing reliable visual features for tracking algorithms to lock onto, thereby improving both tracking precision and reliability without requiring complex modifications to the underlying SLAM system.
Solution Approach 2:
The system changes the parameters of the tracking target by using specifically designed markers with high-contrast patterns and known geometric features. These markers have optimized visual parameters that make them easily detectable and trackable across varying lighting conditions and distances, thereby improving tracking precision and reliability.
2Ease of manufacture
If complex XR content configuration methods are used, then detailed XR content can be created, but user configuration becomes time-consuming and difficult
Solution Approach 1:
The system enables self-service content creation by automatically capturing spatial information, audio, and video data, then autonomously processing and assembling this data into structured XR content. Users simply need to record their narration while the system handles the complex tasks of spatial mapping, text-to-speech conversion, and content assembly, dramatically reducing configuration time and effort.
Solution Approach 2:
The system performs preliminary actions by pre-processing captured data into structured formats during the recording phase. Spatial features are identified and tagged in advance, audio is transcribed and synchronized with video frames beforehand, and all this preparation work is completed before the actual content assembly, making the final content generation rapid and efficient.
Data Source
AI summary
A computing system captures image data using a camera and captures spatial information using one or more sensors. The computing system receives voice data using a microphone. The computing system analyzes the voice data to identify a keyword. The computing system analyzes the image data and the spatial information to identify an object corresponding to the keyword. The computing system generates text based on the voice data and the keyword. The computing system stores the text in association with the object. The computing system generates and provides output comprising the text linked to the object or a derivative thereof.


