Natural Language Tagging for Visual Media
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Tagging visual data is a cumbersome process, requiring users to navigate to specific views, click on locations, and enter textual tag information, which becomes complex with detailed information such as damage type, severity, and repair actions.
Innovation Solution
The system uses natural language processing to determine tags for multi-view interactive digital media representations (MVIDMRs) by applying grammar to speech input, allowing users to specify tag locations and metadata using natural language, and automatically updating the MVIDMR with tags in corresponding positions across multiple images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users manually navigate to views and enter textual tag information, then tagging accuracy can be achieved, but the process becomes cumbersome and time-consuming
Solution Approach 1:
The patent replaces the mechanical manual interaction system (clicking, navigating, typing) with an automated natural language processing system. Users speak commands like 'tag this dent as moderate damage' and the system automatically parses the speech, identifies the location, determines the damage type and severity, and applies the tag without manual navigation or text entry.
Solution Approach 2:
The patent introduces natural language as an intermediary between the user and the tagging system. Instead of direct manual interaction with UI elements, users communicate through speech that is processed by speech recognition and natural language understanding components, which then translate the intent into structured tag data.
2Loss of information
If detailed tag information is attached to visual data, then information completeness is improved, but the complexity of the tagging process increases
Solution Approach 1:
The patent segments the tagging process into distinct components: speech recognition for converting audio to text, natural language parsing for extracting semantic meaning, object model location for identifying where to place tags, and tag generation for creating structured data. This segmentation allows each component to handle a specific aspect, reducing overall complexity while maintaining information completeness.
Solution Approach 2:
The system performs self-service by automatically generating complete tag information from natural language input. When a user says 'tag this scratch on the front bumper as light damage,' the system automatically extracts the location (front bumper), damage type (scratch), severity (light), and applies the appropriate tag without requiring the user to manually fill out multiple fields or navigate through complex forms.
3Stability of the object's composition
If tags are applied across multiple views, then consistency is improved, but the effort required increases
Solution Approach 1:
The patent implements a universal tagging approach where a single natural language command can apply tags across multiple views simultaneously. The system processes the speech input once, determines the object model location, and then automatically updates all relevant views (front, rear, left, right) with consistent tag information, eliminating the need to manually tag each view separately.
Solution Approach 2:
The system creates copies of the tag data across multiple views based on the object model. When a tag is applied to a location in the object model, the system automatically copies this tag information to all views where that location is visible, ensuring consistency across front, rear, left, and right views without requiring separate manual operations for each view.
Data Source
AI summary
A tag characterizing a portion of a multi-view interactive digital media representation (MVIDMR) may be determined by applying a grammar to natural language data. The MVIDMR may include images of an object and may be navigable in one or more dimensions. An object model location for the tag identifying a location within a three-dimensional object model may be determined by applying the grammar to the natural language data. The tag may then be applied to the MVIDMR by associating it with two or more of the images at positions determined based on the object model location.


