Semantic Image Synthesis With Class-Specific Latent Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing generative adversarial network (GAN)-based semantic image synthesis methods lack the ability to control the synthesis process of semantic classes in a targeted and interpretable manner, limiting the capability for local edits of specific classes in generated images.

Innovation Solution

A computer-implemented method and device that utilize a generator configured to determine pixels of a synthetic image by providing a label map and latent code, with class-specific directions in the latent space to alter pixels meaningfully and interpretable, ensuring diverse and consistent changes in the appearance of specified classes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If GAN-based SIS models are used for semantic image synthesis, then high visual quality is achieved, but the ability to control synthesis process of specific semantic classes is lost

Engineering Contradiction:
Improvevisual qualityVSAvoidcontrol capability
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent segments the latent space into class-specific directions by identifying and isolating latent vectors that correspond to individual semantic classes. This segmentation allows independent control of each class while maintaining the overall image synthesis quality, resolving the contradiction between visual quality and control capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by enabling different levels of control granularity - from global image synthesis to class-specific modifications. Users can selectively adjust specific semantic classes (e.g., changing only sky color or only building structures) while keeping other classes unchanged, thus achieving both high visual quality and precise operational control.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If class-specific control is implemented in latent space, then targeted local edits are enabled, but complexity of the synthesis process increases

Engineering Contradiction:
Improvelocal edit capabilityVSAvoidsynthesis process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-identifying and storing class-specific latent directions during the training phase. This preliminary extraction of semantic class vectors allows the system to handle complex class-specific edits during inference without requiring complex real-time processing, thus enabling versatile local edits while managing synthesis process complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary layer that acts as a bridge between the user's high-level semantic instructions and the GAN's low-level image generation. This intermediary component translates simple class-specific modification requests into appropriate latent space adjustments, enabling versatile control without directly increasing the complexity of the core synthesis process.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12505554B2Device and computer-implemented method for determining pixels of a synthetic image
Publication Date: 2025.12.23 ROBERT BOSCH GMBH
  • US12505554B2 patent drawing
  • US12505554B2 patent drawing
  • US12505554B2 patent drawing

AI summary

A device and computer-implemented method for determining pixels of a synthetic image. The method comprises providing a generator that is configured to determine an output from a first input comprising a label map and a first latent code, wherein the label map comprises a mapping of at least one class to at least one of the pixels, wherein the method comprises providing the label map and a latent code, wherein the latent code comprises input data points in a latent space, providing a first direction for moving input data points in the latent space, determining the first latent code depending on at least one input data point of the latent code that is moved in the first direction, determining the synthetic image depending on an output of the generator for the first input.