Multi-branch Harmonization Network for Composite Image Artifacts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image compositing systems face challenges in achieving accurate and realistic image harmonization, often resulting in artifacts, especially with small objects, and failing to handle extreme differences in image characteristics between the foreground and background.
Innovation Solution
The implementation of a multi-branched harmonization framework that combines a semantic feature neural network, a transformer neural network, and a convolutional neural network to extract semantic, global, and local information from composite images, along with a style normalization layer to capture and inject visual style information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional image compositing systems perform image harmonization by matching color, luminescence, saturation, temperature, or lighting conditions, then the composite image appearance is improved, but artifacts are generated especially with small objects
Solution Approach 1:
The system segments the image processing into three distinct neural network branches: a semantic feature network for global semantic information, a transformer network for long-range contextual dependencies, and a convolutional network for localized background context. This segmentation allows each branch to specialize in different aspects of harmonization, improving accuracy while reducing artifacts through coordinated processing
Solution Approach 2:
The system combines multiple types of neural network architectures (semantic feature extraction, transformer attention mechanisms, and convolutional processing) into a composite harmonization framework. This composite approach integrates global semantic understanding with local detail preservation, effectively resolving the contradiction between harmonization accuracy and artifact generation
2Adaptability or versatility
If conventional image compositing systems perform image harmonization to match foreground and background characteristics, then the aesthetic meshing is improved, but the system fails to handle extreme differences in image characteristics
Solution Approach 1:
The system adds a global dimension to harmonization by introducing a transformer neural network that processes long-range contextual dependencies across the entire image. This complementary global perspective works alongside local convolutional processing to handle extreme differences in characteristics between foreground and background regions
Solution Approach 2:
The system employs a multi-branch architecture where semantic features, global context, and local context are integrated through feedback mechanisms. The transformer network provides global feedback to guide local harmonization decisions, enabling the system to adapt to extreme differences while maintaining aesthetic consistency
3Reliability
If conventional image compositing systems blend or match characteristics of foreground and background portions, then the realism is improved, but the system lacks semantic context understanding
Solution Approach 1:
The system performs preliminary extraction of semantic features from the input image before conducting the actual harmonization process. This preliminary semantic understanding is then used to guide the blending and matching operations, ensuring that realism is achieved while preserving semantic context throughout the process
Data Source
AI summary
The present disclosure relates to systems, non-transitory computer-readable media, and methods that implement a multi-branch harmonization neural network architecture to harmonize composite images. For example, in one or more implementations, the semantic-guided transformer-based harmonization system uses a convolutional branch, a transformer branch, and a semantic branch to generate a harmonized composite image based on an input composite image and a corresponding segmentation mask. More particularly, the convolutional branch comprises a series of convolutional neural network layers followed by a style normalization layer to extract localized information from the input composite image. Further, the transformer branch comprises a series of transformer neural network layers to extract global information based on different resolutions of the input composite image. The semantic branch includes a visual neural network that generates semantic features that inform the harmonization of the composite images.


