Semantic Image Synthesis With Class-Specific Latent Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing generative adversarial network (GAN)-based semantic image synthesis methods lack the ability to control the synthesis process of semantic classes in a targeted and interpretable manner, limiting the capability for local edits of specific classes in generated images.
Innovation Solution
A computer-implemented method and device that utilize a generator configured to determine pixels of a synthetic image by providing a label map and latent code, with class-specific directions in the latent space to alter pixels meaningfully and interpretable, ensuring diverse and consistent changes in the appearance of specified classes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If GAN-based SIS models are used for semantic image synthesis, then high visual quality is achieved, but the ability to control synthesis process of specific semantic classes is lost
Solution Approach 1:
The patent segments the latent space into class-specific directions by identifying and isolating latent vectors that correspond to individual semantic classes. This segmentation allows independent control of each class while maintaining the overall image synthesis quality, resolving the contradiction between visual quality and control capability.
Solution Approach 2:
The patent applies local quality by enabling different levels of control granularity - from global image synthesis to class-specific modifications. Users can selectively adjust specific semantic classes (e.g., changing only sky color or only building structures) while keeping other classes unchanged, thus achieving both high visual quality and precise operational control.
2Adaptability or versatility
If class-specific control is implemented in latent space, then targeted local edits are enabled, but complexity of the synthesis process increases
Solution Approach 1:
The patent performs preliminary action by pre-identifying and storing class-specific latent directions during the training phase. This preliminary extraction of semantic class vectors allows the system to handle complex class-specific edits during inference without requiring complex real-time processing, thus enabling versatile local edits while managing synthesis process complexity.
Solution Approach 2:
The patent introduces an intermediary layer that acts as a bridge between the user's high-level semantic instructions and the GAN's low-level image generation. This intermediary component translates simple class-specific modification requests into appropriate latent space adjustments, enabling versatile control without directly increasing the complexity of the core synthesis process.
Data Source
AI summary
A device and computer-implemented method for determining pixels of a synthetic image. The method comprises providing a generator that is configured to determine an output from a first input comprising a label map and a first latent code, wherein the label map comprises a mapping of at least one class to at least one of the pixels, wherein the method comprises providing the label map and a latent code, wherein the latent code comprises input data points in a latent space, providing a first direction for moving input data points in the latent space, determining the first latent code depending on at least one input data point of the latent code that is moved in the first direction, determining the synthetic image depending on an output of the generator for the first input.


