Augmented Reality Label Generation from Conversation Transcripts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual object labeling in computer vision applications is labor-intensive and inefficient, especially in scenarios where reliable object recognition algorithms are not available, and in remotely-assisted augmented reality sessions where timely and accurate textual descriptions of annotated objects are often lacking.
Innovation Solution
A method and system that automatically generate computer vision labels by analyzing transcripts from remotely-assisted augmented reality sessions, using Natural Language Understanding algorithms to detect potential entity names mentioned during graphic annotations, and calculating confidence scores to accept or reject candidate labels, thereby providing textual descriptions for objects in 3D models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual object labeling is performed in remotely-assisted augmented reality sessions, then accurate textual descriptions of annotated objects can be obtained, but the process becomes labor-intensive and inefficient
Solution Approach 1:
The system automatically generates candidate labels by analyzing conversation transcripts and graphic annotations itself, without requiring manual intervention. The remote user's spoken words during the AR session are automatically processed to create potential labels, eliminating the need for separate manual labeling work while maintaining accuracy.
Solution Approach 2:
The system performs preliminary label generation during the AR session by capturing and analyzing the remote user's conversation and annotations in real-time. This preliminary action creates candidate labels that can be directly used or refined, avoiding the need for subsequent manual labeling steps.
2Productivity
If automated object recognition tools are used to generate labels, then labeling efficiency is improved, but reliability is reduced when no reliable object recognition algorithm is available
Solution Approach 1:
The system uses the remote user's conversation transcript as an intermediary to bridge the gap between graphic annotations and accurate labels. Instead of relying on object recognition algorithms to interpret images, the system analyzes the remote user's spoken descriptions, which serve as a reliable intermediary source of object information.
Solution Approach 2:
The system replaces the mechanical system of object recognition algorithms with a language processing approach. Instead of using computer vision algorithms to identify objects in images, the system uses natural language understanding to extract object names from the remote user's conversation, substituting one technological approach for another that is more reliable in this context.
3Productivity
If graphic annotations are drawn without accompanying textual descriptions, then annotation speed is improved, but loss of information occurs regarding object identification
Solution Approach 1:
The system captures the remote user's spoken feedback during the annotation process and automatically converts it into textual labels. The conversation transcript serves as feedback that complements the graphic annotations, ensuring that object identification information is preserved without slowing down the annotation process.
Solution Approach 2:
The system merges graphic annotations with textual information from the conversation transcript into a unified labeling system. By combining these two data sources, the system creates comprehensive labels that include both the spatial information from annotations and the descriptive information from speech, preventing information loss.
Data Source
AI summary
Receiving data recorded during a remotely-assisted augmented reality session held between a remote user and a local user, the data including: drawn graphic annotations that are associated with locations in a 3D model representing a physical scene adjacent the local user, and a transcript of a conversation between the remote and local users. Generating at least one candidate label for each location, each candidate label being textually descriptive of a physical entity that is located, in the physical scene, at a location corresponding to the respective location in the 3D model. The generation of each candidate label includes: for each graphic annotation, automatically analyzing the transcript to detect at least one potential entity name that was mentioned, by the remote and/or local user, temporally adjacent the drawing of the respective graphic annotation. Accepting or rejecting each candidate label, to define it as a true label of the respective physical entity.


