3D Caption Rendering with Face Tracking for AR Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional virtual rendering systems face challenges such as environmental conditions, user actions, and unanticipated visual interruptions, leading to inconsistent presentation of virtual objects in real-world environments. Additionally, these systems are often limited in functionality for authoring AR content on mobile devices, requiring users to navigate through multiple views and windows.

Innovation Solution

The system enables the creation and rendering of virtual three-dimensional (3D) objects, such as 3D captions, within a camera feed, allowing them to appear as if they exist in real-world environments. It includes user interfaces for automatically adding 3D captions to images or videos based on context and for augmenting user-inputted 3D captions with graphical elements like emojis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional virtual rendering systems are used to present virtual objects in real-world environments, then the system can provide basic augmented reality functionality, but the presentation consistency deteriorates due to environmental conditions, user actions, and visual interruptions

Engineering Contradiction:
Improvepresentation consistencyVSAvoidenvironmental conditions and visual interruptions
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system dynamically adapts virtual object presentation by continuously tracking face positions and adjusting rendering parameters in real-time. The virtual objects move and transform according to detected face movements, maintaining consistent presentation despite user actions and environmental changes. This dynamic adaptation allows the system to respond to visual interruptions and environmental conditions while preserving the illusion of virtual objects being present in the real world.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If multiple views and windows are used for authoring AR content, then the system provides comprehensive functionality, but the ease of operation deteriorates due to complex navigation requirements

Engineering Contradiction:
Improveauthoring functionalityVSAvoidnavigation complexity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system merges multiple authoring functions and views into a single integrated interface. Users can access and manipulate all AR content creation tools, virtual object parameters, and scene configuration options within one unified workspace, eliminating the need to navigate between multiple windows and views while maintaining comprehensive authoring capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system provides automatic face detection and tracking, and intelligently positions virtual objects based on detected faces without requiring manual configuration. This self-service capability reduces the complexity of the authoring process by automatically handling tasks that would otherwise require navigating through multiple configuration screens and parameters.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250078427A13D captions with face tracking
Publication Date: 2025.03.06 SNAP INC
  • US20250078427A1 patent drawing
  • US20250078427A1 patent drawing
  • US20250078427A1 patent drawing

AI summary

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing at least one program and method for performing operations comprising: receiving, by one or more processors that implement a messaging application, a video feed from a camera of a user device; detecting, by the messaging application, a face in the video feed; in response to detecting the face in the video feed, retrieving a three-dimensional (3D) caption; modifying the video feed to include the 3D caption at a position in 3D space of the video feed proximate to the face; and displaying a modified video feed that includes the face and the 3D caption.