Diffusion Seed Translation for Precise Unpaired Image Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image-to-image translation systems struggle to modify specific attributes without altering other semantic and appearance aspects, particularly in unpaired settings, which is challenging for training perception systems in vehicles.

Innovation Solution

A method and system utilizing a denoising diffusion implicit model (DDIM) inversion, seed-to-seed generative adversarial network (sts-GAN), and spatial guidance to translate images by encoding input images into a stable diffusion latent space, preserving semantic and structural details, and applying a pre-trained stable diffusion model with target output prompts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing image-to-image translation systems are used to modify specific attributes, then attribute translation is achieved, but other semantic and appearance aspects are altered

Engineering Contradiction:
Improveattribute translation precisionVSAvoidsemantic and appearance fidelity
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent segments the image representation into latent space variables that separate semantic content from style attributes. By operating in this decomposed latent space rather than directly on pixel space, the system can modify specific attributes (e.g., weather conditions, time of day) while preserving other semantic and appearance aspects through targeted manipulation of discrete latent variables.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a latent space as an intermediary representation between the input image and the translated output. This latent space acts as a mediator that decouples different image attributes, allowing independent control and modification of specific properties without affecting others. The latent variables serve as intermediaries that bridge the source and target domains while maintaining semantic consistency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If unpaired image-to-image translation is implemented, then training data requirements are reduced, but control over specific attribute modification is lost

Engineering Contradiction:
Improveunpaired translation capabilityVSAvoidattribute control precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent changes the parameter space from pixel values to latent space variables. By representing images in terms of latent variables that capture semantic and stylistic properties, the system enables precise control over specific attributes (such as weather, time, season) even in unpaired settings. The latent parameterization allows independent adjustment of individual attributes without requiring paired training data.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamic control over latent variables that represent different image attributes. The system allows flexible, independent manipulation of latent parameters corresponding to specific attributes (e.g., changing only weather conditions while keeping other attributes fixed), providing fine-grained control in unpaired translation scenarios through dynamic adjustment of latent space coordinates.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20260057554A1System and method of image-to-image translation in diffusion seed space
Publication Date: 2026.02.26 YISSUM RESEARCH DEVELOPMENT COMPANY OF THE HEBREW UNIVERSITY OF JERUSALEM LTD
  • US20260057554A1 patent drawing
  • US20260057554A1 patent drawing
  • US20260057554A1 patent drawing

AI summary

A computer-implemented method of image-to-image translation that, when executed by data processing hardware, causes the data processing hardware to perform operations comprising applying an inversion technique to an input image to generate a source-domain seed, translating the source-domain seed to a target-domain seed using a translation module, and sampling the target-domain seed to generate a denoised code.