HDR Image Generation via Deep Learning Alignment and Attention Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current high dynamic range (HDR) image generation techniques often suffer from ghosting due to hand tremor, object movement, and other jitter issues when capturing images with both high-brightness and low-brightness areas, leading to suboptimal results.

Innovation Solution

A device and method utilizing a deep learning architecture, specifically a homography network trained through a convolutional neural network (CNN), to align and merge multiple low dynamic range (LDR) images captured at different aperture values or shutter speeds, employing an attention mechanism to weight and combine image regions based on their brightness and association with objects, thereby generating a high-quality HDR image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If multiple images are taken and superimposed to form HDR image, then dynamic range is improved, but ghosting occurs due to hand tremor and object movement

Engineering Contradiction:
Improvedynamic rangeVSAvoidimage quality
Core Design Contradiction:
Illumination intensityVSReliability

Solution Approach 1:

The system performs preliminary alignment of multiple captured images before merging them to create the HDR image. By pre-aligning the images to compensate for hand tremor and object movement, the system eliminates ghosting artifacts while maintaining the dynamic range benefits of multi-image superposition.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses a deep learning architecture that processes alignment information as feedback to adjust the merging process. The attention mechanism selectively weights different image regions based on their quality and relevance, allowing the system to adapt to movement and tremor conditions while producing high-quality HDR output.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If deep learning architecture with attention mechanism is used, then alignment accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvealignment accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The deep learning architecture is segmented into functional modules: a homography network for alignment estimation, an attention mechanism for selective weighting, and a merging network for HDR synthesis. This modular segmentation allows the complex processing to be broken down into manageable stages, improving alignment accuracy while enabling more efficient implementation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes parameters dynamically during processing, including adjusting the degree of attention weighting and merging coefficients based on input image characteristics. This adaptive parameter adjustment allows the system to achieve high alignment accuracy without requiring excessive computational complexity for all possible scenarios.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12051180B2Method for generating images with high dynamic range during multiple exposures and times of capture, device employing method, and storage medium
Publication Date: 2024.07.30 CHIUN MAI COMM SYST INC
  • US12051180B2 patent drawing
  • US12051180B2 patent drawing

AI summary

A method for generating images with high dynamic range (HDR) based on multiple images captured at different aperture values, under different conditions, or at different shutter speeds is applied in a device. The method inputs the original multiple images into a predetermined model and aligns the multiple images. The method further confirms object images that need to be attended among multiple aligned images and obtains a merge weighting for each of the object images, and merges the images for a generated HDR according to the merge weighting of each image. The device utilizing the method is also disclosed.