Diffusion Style Transfer With Representative Attention Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing style transfer techniques are limited by considering a single style image, leading to poor performance and entanglement of content and style elements, and are computationally inefficient when handling multiple style images.

Innovation Solution

A technique that generates an average embedding and representative attention map keys and values from multiple style images, using a trained machine learning model to create a stylized output image that includes content and style elements effectively, reducing computing requirements and avoiding entanglement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If multiple style images are considered to improve visual performance, then style transfer quality improves, but computing requirements increase significantly

Engineering Contradiction:
Improvestyle transfer qualityVSAvoidcomputing requirements
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the style images into discrete style tokens through clustering. Instead of processing all style images directly, the system divides them into representative clusters, selecting only key style images as representatives. This segmentation reduces the computational burden while preserving the essential style characteristics needed for high-quality style transfer.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by selecting only a subset of style images (representatives from each cluster) rather than using all available style images. This partial selection is sufficient to achieve good style transfer quality while significantly reducing computing requirements. The clustering approach ensures that the selected subset captures the diversity of the full style image collection.

Inventive Principle:
Principle #16Partial or excessive action

2Productivity

If a single style image is used to reduce computing requirements, then computing efficiency improves, but style transfer performance deteriorates

Engineering Contradiction:
Improvecomputing efficiencyVSAvoidstyle transfer performance
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent merges multiple style images into cluster representatives by combining their features through averaging. The style tokens are created by merging the visual features of multiple style images within each cluster, creating a composite representation that captures the essential style characteristics of the entire group. This merging allows the system to use fewer images while maintaining or improving style transfer performance.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates universal style tokens that can represent multiple style images simultaneously. Each style token serves as a universal representative for its cluster, capable of capturing the style characteristics of all images within that cluster. This multi-functionality allows a single style token to stand in for multiple style images, improving both computing efficiency and style transfer performance.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of manufacture

If existing style transfer techniques are used, then implementation is simple, but content and style become entangled leading to poor results

Engineering Contradiction:
Improveimplementation simplicityVSAvoidstyle transfer accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent segments the style transfer process into distinct components: content features, style features, and style tokens. By separating style extraction from style application, and introducing style tokens as intermediate representations, the system prevents entanglement between content and style. This segmentation maintains implementation simplicity while dramatically improving style transfer accuracy through the use of style banks and style tokens that decouple style representation from content processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250356540A1Style transfer using generative diffusion features
Publication Date: 2025.11.20 DISNEY ENTERPRISES INC
  • US20250356540A1 patent drawing
  • US20250356540A1 patent drawing
  • US20250356540A1 patent drawing

AI summary

The present invention sets forth techniques for performing style transfer from multiple supplied style images to a supplied content image to generate novel images that include style elements from the multiple supplied style images and content elements from the supplied content image. The techniques include guiding one or more self-attention and cross-attention layers included in a machine learning model based on the multiple supplied style images, such that content elements and style elements included in the style images are not entangled when generating the novel images. The techniques also distill a small subset of representative attention map values from multiple style images, improving performance while reducing computational costs compared to processing all attention map values from the multiple style images.