AR Tutorial Segmentation via Screen Capture and Gesture Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Creating tutorials for augmented reality systems is time-consuming and requires extensive coding experience, making it difficult to produce a sufficient number of AR-based tutorials for a useful library of solutions.

Innovation Solution

A system that ingests screen-based video tutorials, segments them into steps, and generates gesture- and interactor-based annotations, allowing end users to follow tutorials at their own pace through augmented reality overlays.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional coding-based methods are used to create AR tutorials, then the precision and control of tutorial creation is improved, but the time required and complexity of the process increases significantly

Engineering Contradiction:
Improvetutorial creation precisionVSAvoidtutorial creation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system enables self-service tutorial creation by automatically detecting device screens, recording interactions, generating annotations, and creating AR overlays without requiring manual coding or complex 3D model configuration. The tutorial creation process serves itself through automated screen capture and interaction analysis.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical coding process with an automated system that uses screen capture, image processing, and automatic annotation generation. Instead of manually writing code to create AR tutorials, the system automatically generates tutorial content by recording and analyzing device screen interactions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If WYSIWYG creation tools are used to simplify tutorial development, then the ease of operation is improved, but the complexity of setting up 3D models and spatial configuration remains high

Engineering Contradiction:
Improvetutorial creation easeVSAvoidspatial configuration complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system extracts the complex 3D model creation and spatial configuration steps from the tutorial development process. By directly capturing and annotating actual device screens, the patent removes the need to create separate 3D representations and configure their spatial relationships, keeping only the essential recording and annotation functions.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of creating abstract 3D models that require spatial configuration, the system uses direct copies of actual device screens as the basis for AR tutorials. This copying approach eliminates the need for complex spatial setup while maintaining visual accuracy, as the tutorials are built from real screen captures rather than reconstructed 3D models.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If manual 3D model creation is required for each tutorial, then the visual accuracy and detail are improved, but the productivity of tutorial production decreases

Engineering Contradiction:
Improvevisual accuracyVSAvoidtutorial production rate
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system creates accurate visual representations by directly copying device screens during actual usage. This approach maintains visual accuracy while dramatically improving productivity, as tutorials can be created by recording real interactions rather than manually building 3D models for each tutorial scenario.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary actions by automatically capturing screens and generating annotations during the recording phase itself, rather than requiring separate steps for model creation, texturing, and annotation. This preliminary processing enables rapid tutorial production while maintaining visual fidelity to the actual device interface.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11557065B2Automatic segmentation for screen-based tutorials using AR image anchors
Publication Date: 2023.01.17 FUJIFILM BUSINESS INNOVATION CORP
  • US11557065B2 patent drawing
  • US11557065B2 patent drawing
  • US11557065B2 patent drawing

AI summary

Example implementations described herein involve systems and methods for a mobile application device to playback and record augmented reality (AR) overlays indicating gestures to be made to a recorded device screen. A device screen is recorded by a camera of the mobile device, wherein a mask is overlaid on a user hand interacting with the device screen. Interactions made to the device screen are detected based on the mask, and AR overlays are generated corresponding to the reactions.