XR Spatial Scanning Using 2D Detection and Depth-Based 3D Anchoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing extended reality (XR) systems face challenges in efficiently and computationally lightweight detection of physical objects in a real-world scene for accurate anchoring of virtual objects, as conventional 3D data processing is time-consuming and intensive.
Innovation Solution
An XR system captures video frame data and determines 3D positions of physical objects using 2D positions and depth data, optionally with object identification by an object identification service, encapsulating these processes in a service for applications, and labels detected objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional 3D data processing is used to detect physical objects, then measurement precision is improved, but productivity deteriorates due to time-consuming and computationally intensive processing
Solution Approach 1:
The patent segments the 3D detection process into two distinct stages: (1) rapid 2D object detection in video frames using conventional computer vision techniques, and (2) subsequent conversion of 2D positions to 3D positions using depth data from a depth map. This segmentation allows the computationally intensive 3D processing to be minimized while maintaining detection accuracy, thereby resolving the contradiction between measurement precision and productivity.
Solution Approach 2:
The patent transitions from direct 3D object detection to a two-dimensional detection approach followed by dimensionality conversion. By detecting objects in 2D video frames and then converting these positions to 3D using depth information, the system avoids the computational burden of direct 3D processing while maintaining accurate spatial detection, thus improving processing speed without sacrificing detection precision.
2Measurement precision
If full 3D data processing is performed for object detection, then detection accuracy is improved, but use of energy worsens due to computational intensity
Solution Approach 1:
The patent divides the energy-consuming 3D processing task into a low-energy 2D detection phase and a minimal-energy conversion phase. By performing object detection in 2D video frames (which is computationally less intensive) and then converting 2D coordinates to 3D using pre-computed depth data, the system significantly reduces energy consumption while maintaining detection accuracy.
Solution Approach 2:
The patent employs dimensionality change by detecting objects in 2D space and converting to 3D space only when necessary. This approach leverages the fact that 2D image processing is less computationally demanding, thereby reducing energy consumption while still achieving accurate 3D object detection through subsequent coordinate transformation using depth map data.
Data Source
AI summary
An extended Reality (XR) system that provides services for determining 3D data of physical objects in a real-world scene. The XR system receives a request from an application to initiate a spatial scan of a real-world scene. In response, the XR system captures video frame data of the real-world scene and captures a pose of the XR system. The XR system determines a physical object in the real-world scene and determines a 2D position of the physical object, using the video frame data. The XR system determines a depth of the physical object using the 2D position and determines a 3D position of the physical object in the real-world scene using the 2D position of the physical object, the depth of the physical object, and the pose of the XR system. The XR system communicates the 3D position data to the application.


