Language-Guided Image Editing via Cycle-Augmented GAN

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems fail to accurately modify images using language-based commands due to imbalances in training data, leading to erroneous modifications and inability to recognize a wide variety of visual modification commands.

Innovation Solution

A language-guided image-editing system utilizing a cycle-augmentation generative-adversarial neural network (CAGAN) with an editing description network and an attention algorithm to generate accurate and adaptive image modifications based on natural language requests.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional language-based image editing systems are used, then the editing process is simple and user-friendly, but the accuracy of image modifications is poor and the system cannot recognize a wide variety of visual modification commands

Engineering Contradiction:
Improveaccuracy of image modificationsVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the image editing task into multiple components: a language model processes natural language commands, a vision model analyzes image content, and a modification module applies targeted edits. This segmentation allows each component to specialize, improving overall accuracy while managing complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate representations including visual embeddings that bridge language commands and image modifications. These embeddings serve as mediators that translate natural language intent into precise visual modifications, enhancing accuracy without requiring direct complex coupling between language processing and image editing modules

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conventional language-based image editing systems are used, then the system is easy to operate, but it fails to accurately modify images and cannot utilize a wide variety of language-based commands

Engineering Contradiction:
Improvevariety of language-based commandsVSAvoidaccuracy of modifications
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system implements a universal language-processing framework that can interpret multiple types of visual modification commands (e.g., object manipulation, scene modification, style changes). The unified architecture handles diverse command types through common processing pipelines, enhancing versatility while maintaining consistent accuracy across different command categories

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent utilizes parameter adjustments in the vision and language models to adapt to different types of modification commands. By dynamically adjusting model parameters based on command type and image characteristics, the system achieves high accuracy across a wide variety of language-based commands without requiring separate specialized systems

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250190234A1Modifying digital images utilizing a language guided image editing model
Publication Date: 2025.06.12 ADOBE INC
  • US20250190234A1 patent drawing
  • US20250190234A1 patent drawing
  • US20250190234A1 patent drawing

AI summary

This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that perform language guided digital image editing utilizing a cycle-augmentation generative-adversarial neural network (CAGAN) that is augmented using a cross-modal cyclic mechanism. For example, the disclosed systems generate an editing description network that generates language embeddings which represent image transformations applied between a digital image and a modified digital image. The disclosed systems can further train a GAN to generate modified images by providing an input image and natural language embeddings generated by the editing description network (representing various modifications to the digital image from a ground truth modified image). In some instances, the disclosed systems also utilize an image request attention approach with the GAN to generate images that include adaptive edits in different spatial locations of the image.