Head-mounted display position estimation using monocular camera and motion sensor
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing head-mounted display devices face challenges in accurately grasping the position of objects in the real world using single outside scene information, leading to discomfort due to deviations between virtual and real-world object positions, and require larger size, higher costs, and more complex manufacturing processes.
Innovation Solution
A head-mounted display device equipped with an image display unit, an outside-scene acquiring unit, a position estimating unit, and an augmented-reality processing unit that uses at least two types of outside scene information over time to estimate the position of target objects in the real world, allowing for accurate alignment of virtual objects with minimal hardware requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a stereo camera with two or more lenses is used to grasp the position of the object, then the position estimation accuracy is improved, but the device size, cost, and manufacturing complexity increase
Solution Approach 1:
The patent combines multiple outside scene information acquiring means (such as a monocular camera and motion detection sensor) into a single integrated system that works together to achieve position estimation, replacing the need for a complex stereo camera system while maintaining measurement accuracy
Solution Approach 2:
The patent introduces motion detection information as an intermediary element that mediates between the simple monocular camera and the final position estimation, allowing accurate 3D position calculation without requiring multiple cameras
2Measurement precision
If a stereo camera with two or more lenses is used to grasp the position of the object, then the position estimation accuracy is improved, but the manufacturing cost increases
Solution Approach 1:
The patent merges the functions of a monocular camera and motion detection sensor to achieve position estimation accuracy previously only attainable with expensive stereo cameras, significantly reducing manufacturing costs while maintaining measurement precision
Solution Approach 2:
The patent uses a inexpensive monocular camera instead of an expensive stereo camera system, accepting that the single camera provides sufficient data when combined with motion detection information to achieve accurate position estimation
3Measurement precision
If a stereo camera with two or more lenses is used to grasp the position of the object, then the position estimation accuracy is improved, but the device size increases
Solution Approach 1:
The patent combines a compact monocular camera and motion detection sensor into a small integrated package, achieving the same position estimation accuracy as a stereo camera system but with significantly reduced device volume suitable for head-mounted display
4Device complexity
If single outside scene information acquiring means is used, then the device size and cost are reduced, but the position estimation accuracy deteriorates
Solution Approach 1:
The patent uses motion detection information as an intermediary that bridges the gap between simple single-camera observation and accurate 3D position estimation, enabling a monocular camera to achieve stereo-camera-level accuracy through temporal motion analysis
Data Source
AI summary
A head-mounted display device with which a user can visually recognize a virtual image and an outside scene includes an outside-scene acquiring unit configured to acquire outside scene information including at least a feature of the outside scene in a visual field direction of the user, a position estimating unit configured to estimate, on the basis of at least two kinds of the outside scene information acquired by the outside-scene acquiring unit over time, a position of any target object present in a real world, and an augmented-reality processing unit configured to cause the image display unit to form, on the basis of the estimated position of the target object, the virtual image representing a virtual object to be added to the target object.


