Natural Language Image Tagging via Gesture and NLP

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image editing applications are complex, making it difficult for users to locate and initiate specific operations, leading to frustration and reduced user experience due to the multitude of choices and complex interaction processes.

Innovation Solution

Implementing natural language image tags using a natural language processing module and gesture recognition to identify user intent and specify portions of an image, allowing users to interact with applications in a more intuitive manner by defining image portions and operations through natural language inputs and gestures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional image editing applications provide multiple operations and functionalities, then the application becomes more versatile and feature-rich, but the user interface becomes more complex and harder to navigate

Engineering Contradiction:
ImprovefunctionalityVSAvoiduser interface complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a camera interface as an intermediary layer between the user and the complex image editing operations. The camera interface provides a simplified, intuitive workspace where users can naturally interact with images through viewing and basic gestures, while the complex editing functionalities remain accessible through the application's underlying system without requiring direct user navigation through their complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conventional applications provide many operation choices, then the application becomes more versatile, but it becomes difficult for users to locate and initiate specific operations

Engineering Contradiction:
Improveoperation choicesVSAvoidoperation location and initiation
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system performs automatic image analysis and operation recommendations without requiring users to manually search through numerous operation choices. The application analyzes the captured image, identifies relevant features, and automatically presents context-appropriate editing operations, allowing the system to serve itself in identifying user needs rather than requiring users to navigate through all available options

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The application performs preliminary image analysis and operation preparation before the user actually requests an operation. By pre-processing the image and identifying potential editing opportunities in advance, the system has operations ready to be presented to the user, eliminating the need for users to search through the full operation list when they want to edit

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If conventional techniques require manual selection of image portions for operations, then users have precise control, but the process becomes inefficient and time-consuming

Engineering Contradiction:
Improveimage portion selection precisionVSAvoidoperation initiation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the manual mechanical interaction of dragging and dropping selection tools with automated image recognition technology. The system uses computer vision algorithms to automatically identify and select relevant portions of the image based on content analysis, substituting the manual selection mechanism with an intelligent automated system that maintains precision while dramatically improving efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9141335B2Natural language image tags
Publication Date: 2015.09.22 ADOBE INC
  • US9141335B2 patent drawing
  • US9141335B2 patent drawing
  • US9141335B2 patent drawing

AI summary

Natural language image tags are described. In one or more implementations, at least a portion of an image displayed by a display device is defined based on a gesture. The gesture is identified from one or more touch inputs detected using touchscreen functionality of the display device. Text received in a natural language input is located and used to tag the portion of the image using one or more items of the text received in the natural language input.