3D Caption Rendering with Semantic Graphical Elements for Stable AR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional virtual rendering systems face challenges such as environmental conditions, user actions, and visual interruptions, leading to inconsistent presentation of virtual objects in real-world environments. Additionally, these systems lack functionality for authoring AR content on mobile devices, requiring users to navigate through multiple views and windows.

Innovation Solution

The system enables the creation and rendering of virtual three-dimensional (3D) objects, such as 3D captions, within a camera feed, simulating their presence in real-world environments. It includes user interfaces for automatically adding 3D captions to images or videos based on context and allows users to edit and preview these captions before rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional virtual rendering systems are used to present virtual objects in real-world environments, then the system can provide augmented reality experiences, but the presentation becomes inconsistent due to environmental conditions, user actions, and visual interruptions

Engineering Contradiction:
Improveconsistency of virtual object presentationVSAvoidenvironmental conditions and visual interruptions
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system dynamically adapts the presentation of virtual objects based on detected environmental conditions, user actions, and visual interruptions. The rendering system adjusts virtual object properties in real-time to maintain consistent appearance despite changing conditions, resolving the contradiction between providing AR experiences and maintaining reliability under varying environmental factors.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If traditional virtual rendering systems are used, then augmented reality experiences can be provided, but the system lacks functionality for authoring AR content on mobile devices

Engineering Contradiction:
ImproveAR content authoring functionalityVSAvoidsystem functionality requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system integrates multiple functions into a unified platform that enables both AR content consumption and authoring on mobile devices. Users can create, edit, and share 3D captions and AR content directly on their mobile devices without requiring separate authoring tools, thereby adding versatility while managing device complexity through integrated design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If users navigate through multiple views and windows to author AR content, then comprehensive control over virtual objects is possible, but the ease of operation decreases

Engineering Contradiction:
ImproveAR content authoring processVSAvoidnumber of views and windows
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system merges multiple views and windows into a unified interface for authoring AR content. Users can access and control all AR content creation functions through a single integrated interface on mobile devices, eliminating the need to navigate through multiple separate views and windows, thereby improving ease of operation while maintaining comprehensive control capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12347045B23D captions with semantic graphical elements
Publication Date: 2025.07.01 SNAP INC
  • US12347045B2 patent drawing
  • US12347045B2 patent drawing
  • US12347045B2 patent drawing

AI summary

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing at least one program and method for performing operations comprising: receiving, by a messaging application, a video feed from a camera of a user device that depicts a face; receiving a request to add a 3D caption to the video feed; identifying a graphical element that is associated with context of the 3D caption; and displaying the 3D caption and the identified graphical element in the video feed at a position in 3D space of the video feed proximate to the face depicted in the video feed.