Augmented Reality Depth Rendering Using Single Camera Feature Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality systems require multiple cameras and sensors to accurately render virtual objects with depth perspective, making them costly and impractical for use in conventional mobile devices.
Innovation Solution
The use of conventional mobile devices equipped with a single rear-facing camera, gyroscope, magnetometer, accelerometer, and location-detection systems to capture and render virtual objects within real-world images by identifying features and tracking their position relative to the environment, allowing for the placement and presentation of virtual objects with depth using feature recognition technology.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras and sensors are used to accurately render virtual objects with depth perspective, then the rendering accuracy and depth perception are improved, but the device cost and complexity increase
Solution Approach 1:
The patent segments the depth perception function across multiple software processing stages rather than requiring multiple physical sensors. The system divides the task into: (1) capturing 2D images with a single camera, (2) detecting features and their positions, (3) calculating depth information through image processing and comparison, and (4) rendering virtual objects with computed depth perspective. This segmentation allows accurate depth rendering using only one camera.
Solution Approach 2:
The patent replaces the mechanical/optical system of multiple cameras and depth sensors with a computational approach using image processing algorithms. Instead of using multiple physical devices to capture depth information, the system uses software to analyze 2D images, detect features, calculate positions, and synthesize depth perspective through processing. This substitution eliminates the need for additional hardware sensors while achieving the same functional goal.
2Manufacturing precision
If multiple cameras and sensors are used to accurately render virtual objects, then the visual quality is improved, but the device cost increases
Solution Approach 1:
The patent uses a single, inexpensive rear-facing camera that is already present in conventional mobile devices, replacing the need for expensive dedicated depth cameras or sensor arrays. The system achieves accurate virtual object placement through software processing of images from this single, low-cost camera, making the technology accessible in standard consumer devices without requiring expensive hardware upgrades.
Solution Approach 2:
The patent creates a computational model of the physical environment by detecting and copying feature positions from 2D images. The system identifies features in the captured image, determines their positions, and uses this copied spatial information to accurately place virtual objects in the augmented reality scene. This copying approach allows accurate rendering without needing multiple physical cameras to directly capture the 3D space.
3Measurement precision
If feature recognition technology is used to track position and render virtual objects, then the accuracy of virtual object placement is improved, but the processing complexity increases
Solution Approach 1:
The patent extracts only the essential information needed for accurate virtual object placement from the captured images - specifically, feature detection and position calculation. Rather than processing entire images or using complex multi-sensor data fusion, the system extracts key feature points, determines their positions relative to the camera, and uses this extracted information for rendering. This extraction approach achieves high precision while keeping processing requirements manageable for mobile devices.
Data Source
AI summary
Systems described herein apply visual computer-generated elements into real-world images with an appearance of depth by using information available via conventional mobile devices. The systems receive a reference image and reference image data collected contemporaneously with the reference image. The reference image data includes a geo-location, a direction heading, and a tilt. The systems identify one or more features within the reference image and receive a user's selection of a foreground feature from the one or more features. The systems receive a virtual object definition that includes an object type, a size, and an overlay position of the virtual object relative to the foreground feature. The virtual object is provided in the virtual layer appearing behind the foreground feature. The systems store, in a memory, the reference image data associated with the virtual object definition.


