Graphical Coordinate System Transforms for Marker-Free AR Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing augmented reality systems face challenges in displaying graphical content that adapts to the content of a video feed without requiring depth cameras or optical tags, which consume significant processing power and necessitate a prior setup of the real-world environment.

Innovation Solution

A visual approach that uses optical data to transform graphical elements across video frames, employing a graphical element coordinate system defined by a quad or bounding box, and utilizes homography transforms and six-degree-of-freedom camera poses to maintain a world-locked orientation, without specialized hardware like depth cameras or IMUs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth cameras or optical tags are used to position graphical elements, then positioning precision is improved, but processing power consumption increases

Engineering Contradiction:
Improvepositioning precisionVSAvoidprocessing power consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts and removes the need for depth cameras and optical tags from the system. Instead of using these specialized hardware components, the invention uses only standard camera feeds and processes visual data optically to achieve positioning, thereby eliminating the high processing power requirements of depth sensors while maintaining positioning precision through homography transforms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/optical sensing systems (depth cameras, IMUs) with a purely visual processing approach using standard cameras. By substituting specialized hardware with optical data processing through homography transforms and feature matching, the system achieves comparable positioning precision with significantly reduced processing power consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If depth cameras or optical tags are used to position graphical elements, then positioning precision is improved, but device complexity increases

Engineering Contradiction:
Improvepositioning precisionVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes depth cameras, IMUs, and optical tags from the hardware configuration. The system achieves positioning functionality using only a standard camera feed, thereby simplifying the device architecture while maintaining the capability to position graphical elements precisely through visual homography transforms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent makes the standard camera serve multiple functions: it is used both for capturing the video feed and for positioning graphical elements through visual processing. This eliminates the need for specialized depth-sensing hardware, reducing device complexity while maintaining positioning precision through the universal use of the standard camera's optical data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If optical tags or pre-setup markers are used in the environment, then graphical element placement accuracy is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveplacement accuracyVSAvoidsetup requirement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent enables the system to automatically establish its own coordinate system and positioning references using visual features naturally present in the video feed. Instead of requiring pre-placed optical tags or markers, the system self-generates the necessary reference frames through homography transforms applied to the camera feed itself, eliminating setup requirements while maintaining placement accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Instead of placing markers in the real world and then mapping to them, the patent inverts the approach by creating virtual reference frames directly from the camera feed through homography transforms. This reverses the traditional marker-based workflow, eliminating the need for physical tag placement while achieving the same positioning accuracy through computational geometry applied to visual data.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentEP3711025B1Graphical coordinate system transform for video frames
Publication Date: 2025.07.02 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3711025B1 patent drawingFigure 1
  • EP3711025B1 patent drawingFigure 2
  • EP3711025B1 patent drawingFigure 3

AI summary

A computing device is provided, which is configured with a processor configured to compute feature points in a new frame and a prior frame of a series of successive video frames, compute optical flow vectors between these frames, and determine a homography transform between these frames based upon the feature points and optical flow vectors. The processor is further configured to apply the homography transform to the graphical element coordinate system in the prior frame to generate an updated graphical element coordinate system in the new frame, and generate a six degree of freedom camera pose transform therebetween based on the homography transform and a camera pose of the graphical element coordinate system in the prior frame. The processor is further configured to render an updated graphical element in the new frame relative to the updated graphical element coordinate system using the six degree of freedom camera pose transform.