Augmented Reality Interface for 3D Task Specification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current human-machine interaction methods are inadequate for accurately and efficiently conveying complex 3D task information between humans and robots, particularly in real-world scenarios, due to limitations in menu-driven UIs, voice interfaces, and the high cost and limited availability of wearable augmented-reality devices.

Innovation Solution

A method and apparatus for human-machine interaction that utilizes a user device to receive and display object and position identification inputs in an augmented-reality manner, providing visual feedback on the 3D posture and properties of objects and positions, enabling intuitive task instruction and visualization of the robot's understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a menu-driven UI is used for human-machine interaction, then the interface is simple to implement, but it is difficult to accurately specify 3D information and complex tasks

Engineering Contradiction:
Improveinterface implementation complexityVSAvoid3D information specification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transitions from traditional 2D menu interfaces to 3D spatial interaction by overlaying virtual interface elements onto the real-world 3D environment. Users can select objects and specify positions by interacting with virtual markers and hotspots displayed in augmented reality, enabling precise 3D information input without complex menus

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces augmented reality markers and virtual hotspots as intermediary elements between the user and the robot system. These visual intermediaries are displayed through the user device's camera view, allowing users to intuitively select objects and positions by tapping on virtual elements that correspond to real-world locations, thereby accurately conveying 3D task information

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If a voice interface is used for human-machine interaction, then the interaction is hands-free, but it is not adequate to accurately specify 3D information

Engineering Contradiction:
Improvehands-free operation capabilityVSAvoid3D information specification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent combines voice recognition with augmented reality visual interfaces. Users can issue voice commands for high-level task instructions while simultaneously using visual markers and hotspots in the augmented reality view to precisely specify 3D object positions and orientations. This hybrid approach leverages the strengths of both voice (hands-free operation) and visual AR interfaces (precise 3D specification)

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent divides the interaction process into segmented components: voice commands handle abstract task intentions while visual AR markers handle concrete 3D parameter specification. This segmentation allows each interface modality to specialize in what it does best, with voice for intent and visual AR for precise spatial parameters

Inventive Principle:
Principle #1Segmentation

3Loss of information

If wearable augmented-reality devices are used for human-machine interaction, then the visualization capability is enhanced, but the cost is high and availability is low

Engineering Contradiction:
Improvevisualization information completenessVSAvoiddevice cost and availability
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The patent designs the augmented reality interface to be device-agnostic, working on standard smartphones and tablets rather than requiring specialized wearable AR devices. The same visual markers and hotspots can be displayed on any device with a camera and display, making the system universally accessible without needing expensive dedicated hardware

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent creates virtual copies of real-world objects and positions as augmented reality markers and hotspots that can be viewed on standard device screens. These visual representations provide sufficient spatial information without requiring the user to wear immersive AR equipment, achieving effective visualization through software-based virtual overlays on conventional devices

Inventive Principle:
Principle #26Copying

4Device complexity

If traditional interaction methods are used, then the implementation is straightforward, but the interaction efficiency for complex 3D tasks is low

Engineering Contradiction:
Improveimplementation straightforwardnessVSAvoidinteraction efficiency for complex tasks
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent pre-generates augmented reality markers and virtual hotspots at known 3D positions before the user needs to interact with them. The system calculates and displays these visual guides in advance, allowing users to quickly select objects and positions by simply tapping on pre-positioned virtual elements, significantly reducing the time and complexity of specifying 3D task parameters

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11164002B2Method for human-machine interaction and apparatus for the same
Publication Date: 2021.11.02 ELECTRONICS & TELECOMM RES INST
  • US11164002B2 patent drawing
  • US11164002B2 patent drawing
  • US11164002B2 patent drawing

AI summary

Disclosed herein are a method for human-machine interaction and an apparatus for the same. The method includes receiving object identification input for identifying an object related to the task to be dictated to a machine through the I/O interface of a user device that displays a 3D space; displaying an object identification visual interface, corresponding to the object identified within the space recognized by the machine, on the user device in an augmented-reality manner; receiving position identification input for identifying a position in the 3D space related to the task; displaying a position identification visual interface, corresponding to the position identified within the space recognized by the machine, on the user device in an augmented-reality manner; and receiving information related to the result of the task performed through the machine.