Scannable Code Landmarks for Precise 3D AR Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing augmented reality systems face challenges in precisely aligning virtual content with the real world, requiring accurate estimation of camera movement and object orientation to ensure seamless integration.

Innovation Solution

A system that generates a scan request based on image data containing a coded image, determines its position and orientation in a 3D Euclidean space, and accesses media content from a repository to display AR content accurately, considering device, user, and contextual attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If camera movement and object orientation are estimated to align AR content with the real world, then alignment accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvealignment accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces coded images as intermediary landmarks in the real world that encode position and orientation information. Instead of directly estimating camera movement and object orientation through complex algorithms, the system uses these coded images as mediators to provide direct reference data, thereby improving alignment accuracy while reducing the computational complexity of the tracking system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates virtual copies of real-world landmarks by placing coded images at specific positions and orientations. These coded images serve as simplified representations that capture the essential spatial information needed for AR alignment, allowing the system to work with standardized reference markers rather than complex real-world geometry, thus reducing system complexity while maintaining precision.

Inventive Principle:
Principle #26Copying

2Measurement precision

If coded images are used as landmarks to determine position and orientation, then alignment precision is improved, but implementation complexity increases

Engineering Contradiction:
Improvealignment precisionVSAvoidimplementation complexity
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The coded images serve multiple functions simultaneously: they act as visual landmarks for position determination, encode orientation information, and provide reference points for scale calculation. This multi-functionality allows a single simple component to replace multiple complex sensing and measurement systems, improving precision while simplifying implementation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent encodes spatial information (position and orientation) into the visual parameters of the coded images themselves. By changing the arrangement, color, or structure of patterns within the coded images, the system can convey multiple pieces of spatial information through a single visual marker, thereby improving alignment precision without adding physical complexity to the landmarks.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12597214B2Scannable codes as landmarks for augmented-reality content
Publication Date: 2026.04.07 SNAP INC
  • US12597214B2 patent drawing
  • US12597214B2 patent drawing
  • US12597214B2 patent drawing

AI summary

A system to perform operations that include: generating, at a client device, a scan request that comprises image data, the image data comprising a depiction of a coded image that comprises a reference to a media repository; determining a set of coordinates that indicate a position and orientation of the coded image within a three-dimensional (3D) Euclidean space responsive to the scan request; defining a reference point at the client device based on the set of coordinates that indicate the position and orientation of the coded image within the 3D Euclidean space; accessing media content from within the media repository based on the reference to the media content associated with the coded image, wherein the media content may include AR content; and causing display of a presentation of the media content at the client device based on the reference point.