Image-Guided Model Inversion for Accurate Digital Image Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image generation systems using generative adversarial networks (GANs) lack flexibility and accuracy in synthesizing digital images, particularly failing to capture high-level semantics and generalize to domains outside specific training data, leading to inefficiencies and unrealistic outputs.

Innovation Solution

The method employs an image-guided model inversion system that utilizes a neural network image classifier to constrain features and patches of an initial image relative to a target image, using a discriminator to reduce feature and patch differences, allowing for more accurate and flexible synthesis of digital images across various domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional GANs are used to synthesize digital images, then image generation capability is provided, but flexibility and accuracy in capturing high-level semantics deteriorates

Engineering Contradiction:
Improveaccuracy in capturing high-level semanticsVSAvoidflexibility in synthesizing images across domains
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an image classifier as an intermediary component that bridges the target image and the synthesized image. The classifier constrains the feature representations at multiple layers, ensuring that high-level semantics are preserved during the synthesis process. This intermediary mechanism enables accurate semantic capture while maintaining flexibility across different domains without requiring domain-specific retraining.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent operates in the feature representation space of the image classifier rather than directly in pixel space. By constraining features at multiple layers of the classifier, the system works in a higher-dimensional semantic space that captures abstract properties of images. This dimensional transformation enables the system to generalize across domains while maintaining semantic accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If GAN projection is used to project target image into latent space, then image modification is enabled, but realism of synthesized images deteriorates

Engineering Contradiction:
Improveease of image modificationVSAvoidrealism of synthesized images
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the image classifier continuously evaluates the synthesized image and provides constraints back to the synthesis process. The classifier's feature representations guide the optimization of the synthesized image, ensuring that modifications remain within the manifold of realistic images. This feedback loop maintains realism while enabling flexible image modification.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If conventional systems retrain neural networks for each target image, then domain-specific accuracy is improved, but computational efficiency deteriorates

Engineering Contradiction:
Improvedomain-specific synthesis accuracyVSAvoidcomputational efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates a universal synthesis system that works across multiple domains without requiring retraining. The image classifier serves as a domain-agnostic constraint that can be applied to any target image. The system extracts features and applies constraints in a way that is independent of the specific domain, enabling the same synthesis pipeline to handle diverse images efficiently.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary feature extraction and constraint formulation using the pre-trained image classifier before the actual synthesis process. By preparing the feature constraints in advance based on the target image's semantic content, the system avoids the need for time-consuming retraining while maintaining domain-specific accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11842468B2Synthesizing digital images utilizing image-guided model inversion of an image classifier
Publication Date: 2023.12.12 ADOBE INC
  • US11842468B2 patent drawing
  • US11842468B2 patent drawing
  • US11842468B2 patent drawing

AI summary

This disclosure describes methods, non-transitory computer readable storage media, and systems that utilize image-guided model inversion of an image classifier with a discriminator. The disclosed systems utilize a neural network image classifier to encode features of an initial image and a target image. The disclosed system also reduces a feature distance between the features of the initial image and the features of the target image at a plurality of layers of the neural network image classifier by utilizing a feature distance regularizer. Additionally, the disclosed system reduces a patch difference between image patches of the initial image and image patches of the target image by utilizing a patch-based discriminator with a patch consistency regularizer. The disclosed system then generates a synthesized digital image based on the constrained feature set and constrained image patches of the initial image.