Scene-Based Image Editing Using AI and Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image editing systems are inflexible and inefficient, requiring users to interact with individual pixels and navigate complex menus to perform edits, which demands deep specialized knowledge and significant user interaction.

Innovation Solution

The scene-based image editing system utilizes machine learning models to process digital images, pre-processing them to identify objects, anticipate edits, and generate supplementary components, allowing users to interact with digital images as if they were real scenes, with edits automatically reflecting real-world conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional image editing systems require users to interact with individual pixels and navigate complex menus, then editing precision can be achieved, but user interaction complexity and time consumption increase significantly

Engineering Contradiction:
Improveediting precisionVSAvoiduser interaction complexity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent introduces an intermediary system (AI model and processing circuitry) that translates natural language instructions into precise image editing operations. This intermediary layer allows users to provide high-level semantic guidance while the system handles the complex pixel-level manipulations, resolving the contradiction between precision and ease of operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service editing by automatically interpreting user intent from natural language and executing the appropriate editing operations without requiring users to manually navigate menus or select individual pixels. The AI model autonomously performs the complex tasks while maintaining editing precision.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If conventional image editing systems provide detailed control over individual pixels, then editing accuracy is improved, but the system complexity and user knowledge requirements increase

Engineering Contradiction:
Improveediting accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary system (AI model and processing circuitry) that translates natural language instructions into precise image editing operations. This intermediary layer allows users to provide high-level semantic guidance while the system handles the complex pixel-level manipulations, resolving the contradiction between precision and ease of operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical interaction model (manual pixel selection and menu navigation) with an intelligent system that uses natural language processing and AI to interpret user intent and automatically execute editing operations, reducing system complexity from the user perspective while maintaining accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If conventional image editing systems require deep specialized knowledge for operation, then editing quality can be maintained, but accessibility and ease of use deteriorate

Engineering Contradiction:
Improveediting qualityVSAvoidaccessibility
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system enables self-service editing by automatically interpreting user intent from natural language and executing the appropriate editing operations without requiring users to manually navigate menus or select individual pixels. The AI model autonomously performs the complex tasks while maintaining editing precision.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent fundamentally changes the interaction parameter from technical pixel-level coordinates and menu selections to natural language semantics. This parameter transformation allows users with no specialized knowledge to communicate editing intentions effectively, maintaining quality while improving accessibility.

Inventive Principle:
Principle #35Parameter changes

4Ease of operation

If conventional image editing systems require significant user interaction to perform edits, then control over the editing process is maintained, but productivity and efficiency decrease

Engineering Contradiction:
ImprovecontrolVSAvoidediting efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system performs preliminary actions by pre-processing the image through the AI model to understand semantic content and potential edit opportunities before the user provides instructions. This allows the system to be ready to execute edits quickly once user intent is expressed, maintaining control while improving productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent enables continuous editing workflows where the AI model maintains context and can execute multiple editing operations in sequence based on ongoing user instructions, eliminating the need for repeated menu navigation and tool selection, thus maintaining control while significantly improving efficiency.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12210800B2Modifying digital images using combinations of direct interactions with the digital images and context-informing speech input
Publication Date: 2025.01.28 ADOBE INC
  • US12210800B2 patent drawing
  • US12210800B2 patent drawing
  • US12210800B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that edit digital images using combinations of speech input and gesture interactions. For instance, in some embodiments, the disclosed systems receive speech input from a client device displaying a digital image within a graphical user interface, the digital image portraying an object. Additionally, the disclosed systems detect, via the graphical user interface, one or more gesture interactions with respect to the object of the digital image. Based on the speech input, the disclosed systems determine an edit for the object of the digital image indicated by the one or more gesture interactions. Further, the disclosed systems modify the object within the digital image using the edit indicated by the one or more gesture interactions.