Perspective Correction for Head-Mounted Displays Using Keyframe Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In computer-generated reality (CGR) environments, particularly in head-mounted devices (HMDs), the difference in perspective between the scene camera and the user's perspective leads to impaired distance perception, disorientation, and poor hand-eye coordination due to the offset positions of the eyes, display, and camera, resulting in incomplete image transformations and holes in the displayed image.
Innovation Solution
A method involving a processor, image sensor, and display that captures multiple images of a scene from various perspectives, obtains a depth map, transforms the current image to align with the user's perspective, and fills holes using information from other images, ensuring a more accurate and complete visual representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If the scene camera is positioned at an offset from the user's eyes in an HMD, then the device structure is simplified and manufacturing is easier, but the user experiences impaired distance perception, disorientation, and poor hand-eye coordination
Solution Approach 1:
The patent introduces an image transformation system as an intermediary between the camera and the user's eyes. The scene camera captures images from its offset position, and the processor transforms these images to correct the perspective mismatch. This intermediary transformation process resolves the contradiction by allowing the camera to remain at an offset position (simplifying manufacturing) while the user receives corrected images that match their eye perspective (improving ease of operation).
Solution Approach 2:
The patent applies parameter changes by transforming the image parameters (perspective, position, scale) to compensate for the camera-eye offset. The processor modifies the image parameters through perspective correction algorithms, changing the visual representation to match the user's actual viewing geometry. This allows the physical camera position to remain fixed and simple while the displayed image parameters are dynamically adjusted to eliminate disorientation and improve hand-eye coordination.
2Manufacturing precision
If the image is transformed to correct the perspective difference between camera and user, then the visual alignment accuracy is improved, but holes appear in the transformed image due to incomplete transformation
Solution Approach 1:
The patent applies preliminary action by capturing multiple images from different perspectives before the transformation process. The scene camera captures a plurality of images from different positions and orientations. These pre-captured images serve as source material for the transformation, allowing the system to fill holes in the transformed image by selecting appropriate source regions from the previously captured images, thus preventing information loss.
Solution Approach 2:
The patent uses copying by creating multiple copies of the scene from different perspectives (multiple captured images) and then copying relevant portions from these images to fill holes in the transformed image. The processor identifies hole regions in the transformed image and copies pixel data from corresponding regions in the source images, effectively reproducing missing visual information without losing detail.
3Loss of information
If multiple images are captured from different perspectives, then the completeness of the transformed image is improved, but the processing time and computational complexity increase
Solution Approach 1:
The patent applies partial action by selectively processing only the necessary portions of the captured images. Instead of fully transforming and processing all captured images, the system identifies specific hole regions in the transformed image and only processes the minimum necessary source regions from the captured images to fill those holes. This partial processing approach maintains image completeness while significantly reducing computational complexity and processing time compared to full image transformation.
Data Source
AI summary
In one implementation, a method of performing perspective correction is performed at a head-mounted device including one or more processors, non-transitory memory, an image sensor, and a display. The method includes capturing, using the image sensor, a plurality of images of a scene from a respective plurality of perspectives. The method includes capturing, using the image sensor, a current image of the scene from a current perspective. The method includes obtaining a depth map of the current image of the scene. The method include transforming, using the one or more processors, the current image of the scene based on the depth map, a difference between the current perspective of the image sensor and a current perspective of a user, and at least one of the plurality of images of the scene from the respective plurality of perspectives. The method includes displaying, on the display, the transformed image.


