Multi-branch Harmonization Network for Composite Image Artifacts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image compositing systems face challenges in achieving accurate and realistic image harmonization, often resulting in artifacts, especially with small objects, and failing to handle extreme differences in image characteristics between the foreground and background.

Innovation Solution

The implementation of a multi-branched harmonization framework that combines a semantic feature neural network, a transformer neural network, and a convolutional neural network to extract semantic, global, and local information from composite images, along with a style normalization layer to capture and inject visual style information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional image compositing systems perform image harmonization by matching color, luminescence, saturation, temperature, or lighting conditions, then the composite image appearance is improved, but artifacts are generated especially with small objects

Engineering Contradiction:
Improveimage harmonization accuracyVSAvoidartifacts
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The system segments the image processing into three distinct neural network branches: a semantic feature network for global semantic information, a transformer network for long-range contextual dependencies, and a convolutional network for localized background context. This segmentation allows each branch to specialize in different aspects of harmonization, improving accuracy while reducing artifacts through coordinated processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system combines multiple types of neural network architectures (semantic feature extraction, transformer attention mechanisms, and convolutional processing) into a composite harmonization framework. This composite approach integrates global semantic understanding with local detail preservation, effectively resolving the contradiction between harmonization accuracy and artifact generation

Inventive Principle:
Principle #40Composite materials

2Adaptability or versatility

If conventional image compositing systems perform image harmonization to match foreground and background characteristics, then the aesthetic meshing is improved, but the system fails to handle extreme differences in image characteristics

Engineering Contradiction:
Improveaesthetic meshing capabilityVSAvoidhandling extreme differences
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system adds a global dimension to harmonization by introducing a transformer neural network that processes long-range contextual dependencies across the entire image. This complementary global perspective works alongside local convolutional processing to handle extreme differences in characteristics between foreground and background regions

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system employs a multi-branch architecture where semantic features, global context, and local context are integrated through feedback mechanisms. The transformer network provides global feedback to guide local harmonization decisions, enabling the system to adapt to extreme differences while maintaining aesthetic consistency

Inventive Principle:
Principle #23Feedback

3Reliability

If conventional image compositing systems blend or match characteristics of foreground and background portions, then the realism is improved, but the system lacks semantic context understanding

Engineering Contradiction:
ImproverealismVSAvoidsemantic context
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system performs preliminary extraction of semantic features from the input image before conducting the actual harmonization process. This preliminary semantic understanding is then used to guide the blending and matching operations, ensuring that realism is achieved while preserving semantic context throughout the process

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12223623B2Harmonizing composite images utilizing a semantic-guided transformer neural network
Publication Date: 2025.02.11 ADOBE INC
  • US12223623B2 patent drawing
  • US12223623B2 patent drawing
  • US12223623B2 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods that implement a multi-branch harmonization neural network architecture to harmonize composite images. For example, in one or more implementations, the semantic-guided transformer-based harmonization system uses a convolutional branch, a transformer branch, and a semantic branch to generate a harmonized composite image based on an input composite image and a corresponding segmentation mask. More particularly, the convolutional branch comprises a series of convolutional neural network layers followed by a style normalization layer to extract localized information from the input composite image. Further, the transformer branch comprises a series of transformer neural network layers to extract global information based on different resolutions of the input composite image. The semantic branch includes a visual neural network that generates semantic features that inform the harmonization of the composite images.