Semantic Image Synthesis via Spatially-Adaptive Normalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing software applications require significant manual effort and expertise for users to create photorealistic images, as they involve manual cropping, alignment, and blending of image components, which can be cumbersome and often result in images with artifacts.

Innovation Solution

A system that uses semantic layouts, where users draw or create regions with associated labels, and an image synthesis network, employing a spatially-adaptive normalization layer, to generate photorealistic images seamlessly, reducing manual interaction and artifacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users manually crop, align, and blend image components using existing software, then they can create photorealistic images, but the process requires significant manual effort and expertise

Engineering Contradiction:
Improveease of creating photorealistic imagesVSAvoidtime required for manual image manipulation
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system enables self-service by automatically performing image synthesis operations. The neural network model autonomously crops, aligns, and blends image components based on semantic layouts without requiring manual user intervention for each operation, thus resolving the contradiction between ease of operation and time consumption

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical operations (cutting, pasting, blending) with an automated neural network system. The GAN-based image synthesis model substitutes the mechanical manual process with an intelligent automated system that processes images through learned transformations, eliminating the need for manual effort while maintaining image quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If users manually manipulate image components to achieve photorealistic results, then image quality can be improved, but the process becomes complicated and cumbersome

Engineering Contradiction:
Improveimage quality and visual fidelityVSAvoidcomplexity of image creation process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system extracts the complex manual manipulation steps from the user's workflow. By separating the semantic layout creation (simple user action) from the image synthesis (complex automated process), the system maintains high image quality while simplifying the user interface and reducing process complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The neural network model acts as an intermediary between the simple semantic layout input and the complex photorealistic image output. This intermediary automatically handles the complex alignment, blending, and artifact removal operations, preserving image quality while shielding users from process complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If users perform manual cropping and blending of image components, then they can control image content, but artifacts and misalignment often result

Engineering Contradiction:
Improveseamlessness of image boundariesVSAvoidsimplicity of image creation
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The generative adversarial network employs feedback mechanisms where the discriminator evaluates the synthesized images and provides feedback to the generator. This feedback loop continuously improves the alignment and blending quality, eliminating artifacts and ensuring seamless boundaries while maintaining ease of operation through automated processing

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces manual mechanical blending operations with neural network-based image synthesis. The learned transformations in the GAN model automatically handle boundary alignment and artifact removal, achieving reliable seamless results without requiring manual intervention that could introduce errors

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Extent of automation

If traditional software tools are used to blend image components, then some automation is provided, but significant manual interaction is still required

Engineering Contradiction:
Improveautomation of image synthesisVSAvoiduser effort required
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The system achieves complete automation through self-service processing. The neural network model autonomously performs all image synthesis operations including cropping, aligning, blending, and artifact removal based solely on the semantic layout input, eliminating the need for manual interaction while maintaining ease of operation through simple region drawing

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20200242771A1Semantic image synthesis for generating substantially photorealistic images using neural networks
Publication Date: 2020.07.30 NVIDIA CORP
  • US20200242771A1 patent drawing
  • US20200242771A1 patent drawing
  • US20200242771A1 patent drawing

AI summary

A user can create a basic semantic layout that includes two or more regions identified by the user, each region being associated with a semantic label indicating a type of object(s) to be rendered in that region. The semantic layout can be provided as input to an image synthesis network. The network can be a trained machine learning network, such as a generative adversarial network (GAN), that includes a conditional, spatially-adaptive normalization layer for propagating semantic information from the semantic layout to other layers of the network. The synthesis can involve both normalization and de-normalization, where each region of the layout can utilize different normalization parameter values. An image is inferred from the network, and rendered for display to the user. The user can change labels or regions in order to cause a new or updated image to be generated.