AI-Driven AR Mentoring for Automatic Object Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality (AR) systems require significant manual effort for environment scanning and database annotation, limiting their use in new or unknown environments and being specific to certain make and models, which hinders efficient AR-assisted mentoring and collaboration.

Innovation Solution

The implementation of AI-driven methods using deep learning algorithms for semantic feature determination and machine learning for scene understanding, enabling automatic identification of objects and their spatial relationships, and generating visual representations to guide users in tasks without the need for manual database creation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual environment scanning and database annotation are used, then AR systems can provide accurate object identification and spatial relationships, but significant manual effort and time are required

Engineering Contradiction:
Improveobject identification accuracyVSAvoidmanual annotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs automatic environment scanning and database annotation using AI algorithms without requiring manual intervention. The deep learning model automatically identifies objects, determines their spatial relationships, and annotates the scene graph, enabling the system to serve itself rather than requiring external manual annotation efforts.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical annotation processes with automated AI-based image processing and deep learning algorithms. The system uses neural networks to automatically extract semantic features and generate scene graphs, substituting human labor with intelligent automated systems.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual database annotation is performed for specific make and model, then accurate AR visualization can be achieved, but the system cannot be used in new or unknown environments

Engineering Contradiction:
ImproveAR visualization accuracyVSAvoidenvironment adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system uses a universal deep learning model that can process and annotate diverse environments and objects without requiring pre-existing annotated databases for each specific make and model. The AI model generalizes across different scenarios, enabling the system to adapt to new environments while maintaining reliable AR visualization through automatic scene graph generation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of manufacture

If pre-annotated databases are created for AR systems, then object identification can be performed, but the process becomes specific to certain make and models and requires hours of manual authoring

Engineering Contradiction:
Improvedatabase creation easeVSAvoidsystem configuration complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The system automatically generates its own annotated database through AI-based environment scanning and scene graph generation. Rather than requiring manual database creation and configuration, the system performs self-annotation by processing captured images and automatically extracting object information, spatial relationships, and semantic features to build the necessary data structures.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240096093A1Ai-driven augmented reality mentoring and collaboration
Publication Date: 2024.03.21 SRI INTERNATIONAL
  • US20240096093A1 patent drawing
  • US20240096093A1 patent drawing
  • US20240096093A1 patent drawing

AI summary

A method for AI-driven augmented reality mentoring includes determining semantic features of objects in at least one captured scene, determining 3D positional information of the objects, combining information regarding the identified objects with respective 3D positional information to determine at least one intermediate representation, completing the determined intermediate representation using machine learning to include additional objects or positional information of the objects not identifiable from the at least one captured scene, determining at least one task to be performed and determining steps to be performed using a knowledge database, generating at least one visual representation relating to the determined steps for performing the at least one task, determining a correct position for displaying the at least one visual representation, and displaying the at least one visual representation on the see-through display in the determined correct position as an augmented overlay to the view of the at least one user.