Augmented Reality Interface for 3D Task Specification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human-machine interaction methods are inadequate for accurately and efficiently conveying complex 3D task information between humans and robots, particularly in real-world scenarios, due to limitations in menu-driven UIs, voice interfaces, and the high cost and limited availability of wearable augmented-reality devices.
Innovation Solution
A method and apparatus for human-machine interaction that utilizes a user device to receive and display object and position identification inputs in an augmented-reality manner, providing visual feedback on the 3D posture and properties of objects and positions, enabling intuitive task instruction and visualization of the robot's understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a menu-driven UI is used for human-machine interaction, then the interface is simple to implement, but it is difficult to accurately specify 3D information and complex tasks
Solution Approach 1:
The patent transitions from traditional 2D menu interfaces to 3D spatial interaction by overlaying virtual interface elements onto the real-world 3D environment. Users can select objects and specify positions by interacting with virtual markers and hotspots displayed in augmented reality, enabling precise 3D information input without complex menus
Solution Approach 2:
The patent introduces augmented reality markers and virtual hotspots as intermediary elements between the user and the robot system. These visual intermediaries are displayed through the user device's camera view, allowing users to intuitively select objects and positions by tapping on virtual elements that correspond to real-world locations, thereby accurately conveying 3D task information
2Ease of operation
If a voice interface is used for human-machine interaction, then the interaction is hands-free, but it is not adequate to accurately specify 3D information
Solution Approach 1:
The patent combines voice recognition with augmented reality visual interfaces. Users can issue voice commands for high-level task instructions while simultaneously using visual markers and hotspots in the augmented reality view to precisely specify 3D object positions and orientations. This hybrid approach leverages the strengths of both voice (hands-free operation) and visual AR interfaces (precise 3D specification)
Solution Approach 2:
The patent divides the interaction process into segmented components: voice commands handle abstract task intentions while visual AR markers handle concrete 3D parameter specification. This segmentation allows each interface modality to specialize in what it does best, with voice for intent and visual AR for precise spatial parameters
3Loss of information
If wearable augmented-reality devices are used for human-machine interaction, then the visualization capability is enhanced, but the cost is high and availability is low
Solution Approach 1:
The patent designs the augmented reality interface to be device-agnostic, working on standard smartphones and tablets rather than requiring specialized wearable AR devices. The same visual markers and hotspots can be displayed on any device with a camera and display, making the system universally accessible without needing expensive dedicated hardware
Solution Approach 2:
The patent creates virtual copies of real-world objects and positions as augmented reality markers and hotspots that can be viewed on standard device screens. These visual representations provide sufficient spatial information without requiring the user to wear immersive AR equipment, achieving effective visualization through software-based virtual overlays on conventional devices
4Device complexity
If traditional interaction methods are used, then the implementation is straightforward, but the interaction efficiency for complex 3D tasks is low
Solution Approach 1:
The patent pre-generates augmented reality markers and virtual hotspots at known 3D positions before the user needs to interact with them. The system calculates and displays these visual guides in advance, allowing users to quickly select objects and positions by simply tapping on pre-positioned virtual elements, significantly reducing the time and complexity of specifying 3D task parameters
Data Source
AI summary
Disclosed herein are a method for human-machine interaction and an apparatus for the same. The method includes receiving object identification input for identifying an object related to the task to be dictated to a machine through the I/O interface of a user device that displays a 3D space; displaying an object identification visual interface, corresponding to the object identified within the space recognized by the machine, on the user device in an augmented-reality manner; receiving position identification input for identifying a position in the 3D space related to the task; displaying a position identification visual interface, corresponding to the position identified within the space recognized by the machine, on the user device in an augmented-reality manner; and receiving information related to the result of the task performed through the machine.


