Image Editing System Using Multi-Modal Conversation for Object Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image editing applications lack the ability to efficiently process multi-modal user input and direct user conversations, limiting their effectiveness in handling complex editing tasks such as removing or replacing objects in images, especially when users provide spoken commands without semantic understanding of the image context.

Innovation Solution

The system directs a user conversation using multi-modal input, including spoken instructions and visual indicators, to identify objects and process editing queries, utilizing a vision module to ascertain pixels and a language processor to determine editing requests, and generates harmonized images by obtaining fill or replacement material from a database to create natural-looking composite images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If image editing applications use manual selection methods (e.g., painting with a paintbrush tool), then users can specify objects to edit, but the complexity and time required for editing tasks increases significantly

Engineering Contradiction:
Improveease of object selectionVSAvoidtime required for editing tasks
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical selection methods (paintbrush tools, click-to-select) with automated computer vision-based object identification. The system uses trained neural networks to automatically detect and identify objects in images based on natural language descriptions, eliminating the need for users to manually outline or click on objects while maintaining accurate selection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically analyzing the image content and identifying objects without requiring user intervention for selection. The neural network processes the image independently to locate and identify objects matching the user's natural language description, reducing both the skill level and time required for editing tasks.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If image editing applications include comprehensive voice interfaces, then users can communicate editing requests naturally, but the system lacks semantic knowledge of the image context

Engineering Contradiction:
Improveease of command inputVSAvoidaccuracy of object identification
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent merges voice interface capabilities with computer vision and semantic understanding systems. The natural language processor combines spoken commands with automated image analysis to achieve both ease of operation and accurate object identification. The system integrates multiple modalities (speech and visual processing) to understand and execute editing requests accurately.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary layer consisting of a natural language processor and semantic understanding module that bridges the gap between voice commands and image editing operations. This intermediary translates natural language descriptions into precise object identification and editing parameters, enabling accurate execution without requiring manual selection.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If image editing applications process only spoken commands, then the interface is simple, but the system cannot benefit from multi-modal input for improved accuracy

Engineering Contradiction:
Improvesimplicity of interfaceVSAvoidaccuracy of editing query processing
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements a universal interface that accepts multiple input modalities including spoken commands, natural language descriptions, and visual indicators. The system processes these different input types through a unified architecture that combines voice recognition, natural language processing, and computer vision to achieve high accuracy in editing query processing while maintaining user-friendly operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10613726B2Removing and replacing objects in images according to a directed user conversation
Publication Date: 2020.04.07 ADOBE INC
  • US10613726B2 patent drawing
  • US10613726B2 patent drawing
  • US10613726B2 patent drawing

AI summary

Systems and techniques are described herein for directing a user conversation to obtain an editing query, and removing and replacing objects in an image based on the editing query. Pixels corresponding to an object in the image indicated by the editing query are ascertained. The editing query is processed to determine whether it includes a remove request or a replace request. A search query is constructed to obtain images, such as from a database of stock images, including fill material or replacement material to fulfill the remove request or replace request, respectively. Composite images are generated from the fill material or the replacement material and the image to be edited. Composite images are harmonized to remove editing artifacts and make the images look natural. A user interface exposes images, and the user interface accepts multi-modal user input during the directed user conversation.