Mixed Reality Pose Estimation via Sensor Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional virtual reality systems fail to create a realistic mixed-reality environment by not rendering synthetic objects in a real-world context, leading to issues like low latency processing of real-world video data and inaccurate calculation of user pose, resulting in jittery or bouncing synthetic objects within the user's field of view.

Innovation Solution

A system and method that utilize user-worn sensors, such as video cameras, LIDAR, IMU, and GPS, to capture and process data for accurate pose estimation, combined with low latency processing techniques, to integrate synthetic objects seamlessly into the real-world environment using a see-through HMD, allowing users to interact with both real and synthetic objects in a stable and realistic manner.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional virtual reality systems are used to create a completely synthetic environment, then user interaction with synthetic objects is enabled, but the system fails to render synthetic objects in a real world context and cannot process real world video data with low latency

Engineering Contradiction:
Improverealism of mixed-reality environmentVSAvoidlatency in processing real world video data
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system merges real world video data processing with synthetic object rendering by integrating multiple sensors (video cameras, LIDAR, IMU, GPS) to capture both real environment data and user pose information simultaneously. This combination enables low latency processing of real world video data while maintaining accurate rendering of synthetic objects in the correct spatial context, resolving the contradiction between realism and latency

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary processing layer that receives data from multiple sensors and synthesizes pose estimation information. This intermediary process integrates video camera data, LIDAR measurements, IMU readings, and GPS location to calculate accurate user pose in real time, enabling low latency rendering of synthetic objects that appear stable in the user's field of view

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If conventional systems render synthetic objects without accurate pose estimation, then rendering speed is maintained, but synthetic objects appear to jitter or bounce within the user's field of vision

Engineering Contradiction:
Improverendering speedVSAvoidstability of synthetic objects in user's field of view
Core Design Contradiction:
SpeedVSStability of the object's composition

Solution Approach 1:

The system performs preliminary pose estimation by continuously tracking user head position and orientation using IMU sensors and video cameras before rendering synthetic objects. This advance calculation of user pose ensures that synthetic objects are rendered at the correct positions in the user's field of view, preventing jitter and bounce effects while maintaining high rendering speed through efficient predictive algorithms

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback loops that continuously monitor user pose changes and adjust synthetic object positions in real time. By using IMU data and video processing to detect user head movements and immediately updating the rendering accordingly, the system maintains stable positioning of synthetic objects in the user's field of view without sacrificing rendering speed

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the system uses multiple sensors for accurate pose estimation, then pose calculation accuracy is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of pose estimationVSAvoidcomplexity of sensor integration system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system achieves multi-functionality by using a unified sensor fusion architecture where video cameras, LIDAR, IMU, and GPS sensors serve multiple purposes simultaneously. The same sensors used for pose estimation also provide environment mapping, object tracking, and location services, reducing overall system complexity while maintaining high measurement precision through shared processing pipelines

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution enables a stable and realistic mixed-reality environment by accurately estimating user and device poses, reducing jitter and drift, and allowing multiple users to interact simultaneously within the same environment, enhancing user experience in applications like gaming and training.

Implementation Method 1

A mixed-reality generation system 100 includes a user-worn computer module 102 that may include one or more video cameras 108, LIDAR 112, an inertial measurement unit (IMU) 114, a global positioning system (GPS) sensor 116

Methodology Applied
Scientific EffectLIDAR: LIDAR

Implementation Method 2

A mixed-reality generation system 100 includes a user-worn computer module 102 that may include one or more video cameras 108, LIDAR 112, an inertial measurement unit (IMU) 114

Methodology Applied
Scientific EffectInertial measurement: Accelerometer

Data Source

PatentUS9892563B2System and method for generating a mixed reality environment
Publication Date: 2018.02.13 SRI INTERNATIONAL
  • US9892563B2 patent drawing
  • US9892563B2 patent drawing
  • US9892563B2 patent drawing

AI summary

A system and method for generating a mixed-reality environment is provided. The system and method provides a user-worn sub-system communicatively connected to a synthetic object computer module. The user-worn sub-system may utilize a plurality of user-worn sensors to capture and process data regarding a user's pose and location. The synthetic object computer module may generate and provide to the user-worn sub-system synthetic objects based information defining a user's real world life scene or environment indicating a user's pose and location. The synthetic objects may then be rendered on a user-worn display, thereby inserting the synthetic objects into a user's field of view. Rendering the synthetic objects on the user-worn display creates the virtual effect for the user that the synthetic objects are present in the real world.