Swapping Autoencoder Scene Layout Digital Image Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital image editing systems lack flexibility and accuracy in editing real images, often generating images with inaccurate digital content placement and limited ability to add new content without referencing existing images.

Innovation Solution

The system employs a swapping autoencoder that incorporates scene layout maps to accurately and flexibly generate modified digital images by combining structure codes and texture codes, allowing for the addition of new digital content not present in the input image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional digital image editing systems utilize neural networks based purely on unlabeled datasets, then the systems can generate digital images, but the digital content is placed in inaccurate, unrealistic, or undesirable locations

Engineering Contradiction:
Improveaccuracy of digital content placementVSAvoidflexibility in adding new digital content
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the image editing process into distinct functional modules: an encoder that extracts latent representations, a scene layout map generator that defines semantic boundaries, and a decoder that reconstructs images. This segmentation allows each module to specialize in specific tasks, with the scene layout map specifically handling semantic accuracy while the encoder-decoder pair maintains flexibility for content addition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The scene layout map serves as an intermediary between the input image and the generated output. It acts as a semantic guide that constrains where different types of digital content should be placed, mediating between the flexible latent space representations and the realistic placement requirements. This intermediary structure enables accurate content placement without limiting the ability to add new content types.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conventional digital image editing systems are limited to replication and rearrangement of existing digital content, then the systems maintain consistency with input images, but they cannot add new digital content not already found in the input images

Engineering Contradiction:
Improveability to add new digital contentVSAvoidrealism of generated images
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary action by pre-training the encoder-decoder architecture on large datasets of real images before deployment. This pre-training establishes realistic priors about how digital content should appear and be arranged, ensuring that when new content is added, it maintains realism consistent with the training data rather than requiring reference images for every new content type.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by transitioning from discrete content selection (choosing from existing image content) to continuous latent space manipulation. By representing digital content in a continuous latent space, the system can generate novel content variations and entirely new content types while maintaining realism, controlled by adjusting latent vectors rather than being constrained to replicate existing discrete content.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If conventional digital image editing systems lack semantic consideration, then the systems can process images efficiently, but they generate images with inaccurate representations of digital content

Engineering Contradiction:
Improveaccuracy of digital content representationVSAvoidcomplexity of scene layout integration
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts semantic information into a separate scene layout map structure, taking it out from the main image processing flow. This extraction allows the semantic constraints to be defined independently and applied as guidance during image generation, improving content placement accuracy without deeply entangling semantic processing with the core encoding-decoding operations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system adds another dimension by introducing the scene layout map as a separate structural layer that operates in parallel with the image data. This additional dimension provides semantic organization without interfering with the efficiency of the main image processing pipeline, allowing accurate content representation while maintaining computational tractability through clear separation of concerns.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12254545B2Generating modified digital images incorporating scene layout utilizing a swapping autoencoder
Publication Date: 2025.03.18 ADOBE INC
  • US12254545B2 patent drawing
  • US12254545B2 patent drawing
  • US12254545B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer readable media for accurately and flexibly generating modified digital images utilizing a novel swapping autoencoder that incorporates scene layout. In particular, the disclosed systems can receive a scene layout map that indicates or defines locations for displaying specific digital content within a digital image. In addition, the disclosed systems can utilize the scene layout map to guide combining portions of digital image latent code to generate a modified digital image with a particular textural appearance and a particular geometric structure defined by the scene layout map. Additionally, the disclosed systems can utilize a scene layout map that defines a portion of a digital image to modify by, for instance, adding new digital content to the digital image, and can generate a modified digital image depicting the new digital content.