High dynamic range image imaging network method based on pyramid structure spatial feature transformation

By using the PSFTDNet network and deformable convolution alignment techniques with shared offsets, the ghosting problem caused by image misalignment is solved, image quality is improved, the dynamic range limitation of ordinary camera sensors is overcome, and high dynamic range images that are closer to human eye observation are generated.

CN116977193BActive Publication Date: 2026-04-07NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively avoid image misalignment caused by camera movement when generating high dynamic range images, resulting in ghosting. Furthermore, the limited dynamic range of ordinary camera sensors leads to overexposure or underexposure in certain areas of the image, resulting in loss of detail.

Method used

A high dynamic range imaging network (PSFTDNet) based on pyramid structure spatial feature transformation and shared offset deformable convolution alignment is adopted. It includes a shared offset deformable convolution alignment network, a PSFT-based conditional network, and a DERDB-based fusion network. Through the processing flow of gamma correction, shared offset deformable convolution alignment, pyramid structure spatial feature transformation, and deformable convolution residual dense blocks, high-quality HDR images are generated.

Benefits of technology

It effectively reduces ghosting, improves image quality, makes full use of multi-scale image information, enhances image alignment, and generates high dynamic range images that are closer to what the human eye observes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977193B_ABST
    Figure CN116977193B_ABST
Patent Text Reader

Abstract

The high dynamic range imaging network method based on pyramid structure spatial feature transformation is based on three sub-networks: a deformable convolutional alignment network, a PSFT conditional network, and a fusion network of deformable convolutional residual dense blocks. The process is as follows: 1) The exposure of the input image is aligned using gamma correction, and then all images with and without exposure alignment are input into two convolutional layers of the deformable convolutional alignment network with shared offsets to obtain preliminary image features; 2) The exposure-aligned image features are geometrically aligned in the deformable convolutional alignment network with shared offsets, and the offset obtained in the geometric alignment process is also applied to the LDR image features without exposure alignment for geometric alignment; only gamma correction is used to align the exposure; 3) The geometrically aligned LDR image features without exposure alignment are input into the PSFT-based conditional network to obtain optimized features and obtain a high dynamic range image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of HDR (High Dynamic Range) imaging and relates to a method for a high dynamic range image imaging network (PSFTDNet) based on pyramid structure spatial feature transformation and shared offset deformable convolution alignment, especially the direction of generating a single HDR image using multiple LDR (Low Dynamic Range) images with different exposures. Background Technology

[0002] Most cameras on the market cannot capture HDR images due to the limited dynamic range of their sensors, making it difficult to reflect all the details of an image. HDR image technology was developed to overcome the limitations of LDR images (LDR stands for Low Dynamic Range, HDR for High Dynamic Range; common LDR formats include PNG and JPG. For example, an 8-bit image has RGB channels with grayscale values ​​ranging from 0 to 255, meaning it can represent 256^3 colors. While this seems like a lot, it's far from sufficient to describe the colors in the real world. Therefore, HDR is needed to describe and store images with richer color details. However, HDR is not a fixed standard; some are 12-bit, some are 16-bit, and common formats include HDR, EXR, RAW, and TIFF. Generally, LDR image channel values ​​are mapped to between 0 and 1, with the brightest value being 1, while the brightest value in HDR can exceed 1). The goal is to achieve an image effect closer to what the human eye perceives. Currently, the most effective method for generating HDR images is to synthesize a single HDR image from multiple LDR images with different exposures. In this process, a reference image needs to be selected from the input LDR image as a reference for the geometric features of the generated HDR image.

[0003] This method is highly effective for static scenes, but when dealing with misalignments (mismatches) in the input images caused by camera or object movement, the generated results often contain unnatural, ghost-like patches, commonly known as "ghosting." The key to solving the "ghosting" problem lies in addressing the misaligned portions of the image caused by motion. A common idea is to align the images before input, and many alignment methods have been proposed, but they still have limitations when handling complex scenes. Summary of the Invention

[0004] The main problems addressed by this invention are: proposing a more effective method for generating HDR images that avoids "ghosting" effects: a high dynamic range imaging network based on Pyramid Spatial Feature Transform and shared-offsets Deformable convolutional Network (PSFTDNet). Its aim is to address the problem that most cameras can only capture a relatively small dynamic range, leading to overexposure or underexposure in certain areas of the image, resulting in loss of detail. It overcomes the limitations of ordinary camera sensors to obtain images that more closely resemble what the human eye observes. This invention is primarily a method for processing multiple input LDR images to generate HDR images. Its main sub-networks are: 1. Shared-Offsets Deformable Alignment Network; 2. PSFT-Based Condition Network; 3. Merging Network.

[0005] The technical solution of this invention is: a high dynamic range imaging network (PSFTDNet) based on pyramid structure spatial feature transformation, characterized by being based on the following three sub-networks: 1) a shared-offsets deformable alignment network; 2) a PSFT-based condition network (PSFT-based condition network).

[0006] 3) Deformable convolutional residual dense block-based merging network (DERDB Based Merging) The imaging network method (denoted as a DERDB-based fusion network) is as follows: 1) Exposure alignment of the input image is achieved using gamma correction (correction parameter 2.2), that is, adjusting the exposure of images with different exposures to the same level. Then, all images with and without exposure alignment are input into two convolutional layers of a deformable convolutional alignment network with shared offsets to obtain preliminary image features; 2) Exposure-aligned image features are geometrically aligned in the deformable convolutional alignment network with shared offsets, that is, the geometric features of multiple images are aligned. The offsets obtained during the geometric alignment process are also applied to the LDR image features without exposure alignment. Exposure alignment is achieved only through gamma correction, without checking the alignment status. In this embodiment, only the alignment network performs the processing, without any alignment conditions; 3) The geometrically aligned LDR image features without exposure alignment are input into a PSFT-based conditional network to obtain optimized features; 4) All obtained features are input into a DERDB-based fusion network to obtain the final HDR image, i.e., a high dynamic range image.

[0007] In short, the process of PSFTDNet can be represented as:

[0008]

[0009] in, Indicates the alignment network. Let f(·) represent a PSFT-based conditional network, f(·) represent the mapping function represented by PFSTDNet, and θ represent the parameters in the network.

[0010] Further: The image is input into the network and features are extracted and aligned at the feature level. The network is a deformable convolutional alignment network with shared offsets. The features are refined to make full use of the supplementary information between images. The network is a conditional network based on pyramid structure spatial feature transformation, a fusion network based on deformable convolutional residual dense blocks, and a processing flow of aligning features before feature fusion and then refining features.

[0011] Furthermore, the processing steps are as follows:

[0012] 1-1) The steps for generating an HDR image from multiple LDR images with different exposures are as follows: From the input LDR images with different exposures (denoted as L1, L2, L3), gamma correction is performed to obtain corresponding exposure-aligned images, denoted as (H1, H2, H3). The gamma correction process is represented as follows:

[0013]

[0014] Where t i L represents i The exposure time, where γ is a parameter in gamma correction, is set to 2.2;

[0015] 1-2) Extract and align features from the image input using a SharedOffsets Deformable Alignment Network;

[0016] Image features that are misaligned in exposure but geometrically aligned are input into a PSFT-based condition network to obtain refined features;

[0017] 1-3) Input all the features obtained in the above steps into the DERDB-based Merging Network to obtain the final generated HDR image;

[0018] 1-4) The steps for extracting and aligning features from a deformable convolutional alignment network are as follows:

[0019] 1-4-1) In the alignment network, the input {L1,L2,L3,H1,H2,H3} is processed through two convolutional layers to obtain the corresponding preliminary features. in and They represent from L respectively i and H i Features extracted from;

[0020] 1-4-2) will A deformable convolutional alignment network with a pyramid structure is input to obtain aligned features, and the offsets learned in this process are applied to the alignment network. Alignment;

[0021] This process can be represented as:

[0022]

[0023] Where ΔP represents the learned offset, g(·) represents the mapping of several convolutional layers, [·,·] represents the connection operation, and DConv represents the deformable convolution operation. Indicates the aligned features; This represents the features of the reference exposure-aligned image; It refers to the features of the exposure-aligned image. This parameter is derived from the features of the LDR domain input image and then aligned.

[0024] 1-5) Refining features to fully utilize the complementary information between images: The steps of feature refinement using a conditional network based on pyramid structure spatial feature transformation, i.e., a PSFT-based conditional network, are as follows:

[0025] 1-5-1) Based on the pyramid structure spatial feature transformation module, i.e., the PSFT module, feature downsampling is used to obtain image features of different spatial scales;

[0026] 1-5-2) Perform Spatial Feature Transform (SFT) on the obtained image features of different scales. The SFT process is represented as follows:

[0027]

[0028] Where (α,β) represent the learned affine transformation parameters. ⊙ represents the mapping of the convolutional layer;

[0029] 1-5-3) The final refined image features are obtained by upsampling the results from each layer. This process is represented as follows:

[0030]

[0031] Among them, C l This represents the intermediate conditional network at layer l, (·) ↑2 This represents upsampling with a base of 2, where h and p represent the mapping of several convolutional layers. This represents the l-th layer features refined through the SFT process.

[0032] Alignment features of the LDR domain for image i at or after the Lth layer;

[0033] h l : A mapping function consisting of several convolutional layers at layer L;

[0034] The intermediate conditions of the i-th image in layer L;

[0035] α l ,β l : Affine transformation parameter pair at layer L;

[0036] The mapping function of the Lth convolutional layer;

[0037] The features of the i-th image after SFT spatial feature transformation at layer L;

[0038] Features in layer L+1;

[0039] It is a feature of the reference exposure aligned image. This parameter comes from the features of the Lth layer image after alignment in the LDR domain input.

[0040] The PSFT-based conditional network is the main subnetwork of PSFTDNet, hence the name PSFTDNet. "Exposure alignment" means "exposure value alignment." In this example, the alignment method is gamma correction with a correction parameter of 2.2. Exposure is aligned only through gamma correction; the alignment is not checked, and there are no alignment conditions.

[0041] Beneficial effects: Generating higher-quality HDR images from LDR images with different exposures effectively reduces ghosting. The invention also features: 1. Deformable convolution alignment with shared offsets for better image feature alignment; 2. Utilizing spatial feature transformation based on a pyramid structure to fully leverage multi-scale image information; 3. Employing Deformable Residual Dense Blocks (DERDB) to improve the final fusion effect. Attached Figure Description

[0042] Figure 1 This is a workflow view of the present invention.

[0043] Figure 2 A schematic diagram of deformable convolution alignment with shared offsets.

[0044] Figure 3 This is a schematic diagram of a conditional network based on spatial feature transformation of a pyramid structure.

[0045] Figure 4 This is a schematic diagram of a DERDB-based fusion network.

[0046] Figure 5 The results of PSFTDNet and other HDR image generation networks, AHDRNet and ADNet, are shown under the same test case. Figure 5 There are 7 columns of images from left to right. Each column corresponds to the output image after processing long exposure, medium exposure, short exposure, ground truth, PSTDNET, ADNET, and AHDRNET, respectively. Ground Truth represents the ground truth.

[0047] Figure 6 , Figure 7 , Figure 8 , Figure 9 This is the result obtained for PSFDNet and many other networks under a specific test case.

[0048] Figure 6The top row of images consists of three columns from left to right, each corresponding to an LDR image, a PSTDNET output (after tone mapping), and an LDR image slice, respectively. Figure 6 The bottom row of images has four columns from left to right, each column corresponding to AHDRNET, ADNET, PSTDNET, and the actual output image, respectively.

[0049] Figure 7 The top row of images consists of three columns from left to right, each corresponding to an LDR image, a PSTDNET output (after tone mapping), and an LDR image slice, respectively. Figure 7 The bottom row of images has four columns from left to right, each column corresponding to AHDRNET, ADNET, PSTDNET, and the actual output image, respectively.

[0050] Figure 8 The top row of images consists of three columns from left to right, each corresponding to an LDR image, a PSTDNET output (after tone mapping), and an LDR image slice, respectively. Figure 8 The bottom row of images has four columns from left to right, each column corresponding to AHDRNET, ADNET, PSTDNET, and the actual output image, respectively.

[0051] Figure 9 The top row of images consists of three columns from left to right, each corresponding to an LDR image, a PSTDNET output (after tone mapping), and an LDR image slice, respectively. Figure 9 The bottom row of images has four columns from left to right, each corresponding to AHDRNET, ADNET, PSTDNET, and the actual output image, respectively. Detailed Implementation

[0052] The key technologies involved in this invention are: deformable convolutional alignment network with shared offset, PSFT-based conditional network, and DERDB-based fusion network. These will be explained below:

[0053] 1. Shared-Offset Deformable Alignment Network

[0054] The shared offset deformable convolutional alignment network uses pyramid-structured deformable convolutional alignment modules (PCDAlignment Modules) to align images. This network is based on deformable convolutional alignment, which learns offsets from the exposure-aligned image features of the input to change the shape of the convolutional kernel (convolution filter) to perform convolution processing on the input, thereby achieving better geometric alignment.

[0055] Furthermore, since the corresponding exposure-aligned and exposure-misaligned images have the same geometric features, the offsets learned from the features of the exposure-aligned images will also be used to geometrically align the exposure-misaligned images. Moreover, learning offsets only from the features of the exposure-aligned images allows the network to focus more on the differences in geometric structure between images, avoiding the influence of different exposures on the learned offsets.

[0056] Figure 2 A schematic diagram of deformable convolution alignment with shared offsets.

[0057] 2. PSFT-based condition network

[0058] Spatial Feature Transform (SFT) refines the geometrically aligned features by concatenating the features of the reference image and the non-reference image into a conditional network to generate a priori condition containing complementary information between the two images. The output of this priori condition is then fed into an SFT convolutional layer to obtain affine transformation parameters. Finally, these parameters are used to perform an affine transformation on the non-reference image features to obtain the refined result.

[0059] The PSFT-based conditional network extends spatial feature transformation to a pyramidal spatial feature transform (PSFT). The PSFT module obtains image features of different spatial scales through downsampling, forming a pyramid structure, and then applies the SFT operation to the features of each pyramid layer. The prior conditions and refined features generated by the upper-layer conditional network will participate in the generation of prior conditions and refined features in the lower layers. Upsampling uses bilinear interpolation. In PSFTNet, the number of layers in the PSFT module is set to 3.

[0060] From SFT to PSFT, feature information at different scales is fully utilized. This coarse-to-fine process yields more accurate results and improves the ability to handle objects of different sizes and different scenes.

[0061] Figure 3 This is a schematic diagram of a conditional network based on spatial feature transformation of a pyramid structure.

[0062] 3. DERDB-Based Merging Network

[0063] The image features generated by the above network are input into the DERDB-based fusion network to obtain the final HDR image result.

[0064] The PSFTDNet fusion network, based on the Dilated Residual Dense Block (DRDB), proposes the Deformable Residual Dense Block (DERDB). The PSFTDNet fusion network is based on a series of Deformable Residual Dense Blocks (DERDBs), which use deformable convolution to improve the accuracy of generating HDR results.

[0065] Figure 4 This is a schematic diagram of a DERDB-based fusion network.

[0066] Example of use: The steps to generate an HDR image from multiple (three examples) LDR images with different exposures are as follows:

[0067] 1) The LDR images with different exposures (denoted as L1, L2, L3) are obtained by gamma correction to obtain the corresponding exposure-aligned images (denoted as H1, H2, H3). The gamma correction process can be represented as follows:

[0068]

[0069] Where t i L represents i The exposure time, γ is a parameter in gamma correction, which is set to 2.2 in this invention. Gamma correction (ideal gamma correction is achieved by reversing the output through an inverse nonlinear transformation).

[0070] 2) Shared-Offset Deformable Convolutional Alignment Networks that share offsets with image input

[0071] Features are extracted and aligned using an Alignment Network.

[0072] 3) Input the image features that are misaligned in exposure but geometrically aligned into a PSFT-based conditional network.

[0073] The refined features are obtained.

[0074] 4) Input all the features obtained in the above steps into the DERDB-based merging network. The input features are processed through a series of deformable convolutional residual dense blocks (DERDBs) to obtain the final generated HDR image.

[0075] The steps for extracting and aligning features from a deformable convolutional alignment network are as follows:

[0076] 1) In the alignment network, the input {L1,L2,L3,H1,H2,H3} is processed through two convolutional layers to obtain the corresponding preliminary features. in and They represent from L respectively i and H i Features extracted from [the data].

[0077] 2) A deformable convolutional alignment network with a pyramid structure is input to obtain aligned features, and the offsets learned in this process are applied to the alignment network. Alignment.

[0078] This process can be represented as:

[0079]

[0080] Where ΔP represents the learned offset, g(·) represents the mapping of several convolutional layers, [·,·] represents the connection operation, and DConv represents the deformable convolution operation. This indicates the aligned features.

[0081] The steps for feature refinement using a PSFT-based conditional network are as follows:

[0082] 1) The PSFT module downsamples features to obtain image features of different spatial scales.

[0083] 2) Perform Spatial Feature Transform (SFT) on the obtained image features of different scales. The SFT process can be represented as:

[0084]

[0085] Where (α,β) represent the learned affine transformation parameters. represents the mapping of the convolutional layer, and ⊙ represents the Adama product.

[0086] 3) The final refined image features are obtained by upsampling the results from each layer.

[0087] This process can be represented as:

[0088]

[0089] Among them, C l This represents the intermediate conditional network at layer l, (·) ↑2 This represents upsampling with a base of 2, where h and p represent the mapping of several convolutional layers. This represents the l-th layer feature refined through the SFT process. In this invention, the number of layers in the PSFT module is set to 3.

[0090] Sample Results

[0091] PSFTDNet was tested using two datasets: the NTIRE 2021 Multi-Frame HDRChallenge and a dataset provided by Kalantari et al. Specifically, 200 samples were randomly selected from the training set of the NTIRE dataset as part of the test set for this invention.

[0092] PSFTDNet uses two methods for testing: PSNR values ​​in the tone-mapped range (PSNR-T method) and PSNR values ​​in the HDR range (PSNR-L).

[0093] Figure 5 This section showcases the results obtained by PSFTDNet and other HDR image generation networks, AHDRNet and ADNet, under the same test case. The left side shows three images with different exposures as input, while the right side shows the results obtained using different image processing methods. Ground Truth represents the actual values.

[0094] Figure 6 , Figure 7 , Figure 8 , Figure 9 This is the result obtained for PSFDNet and many other networks under a specific test case.

[0095] As can be seen, PSFTDNet has a certain advantage over other HDR image generation networks in eliminating "ghosting".

[0096] Table 1 shows the average quantization evaluation results of PSFTDNet, AHDRNet, and ADNet on the two datasets mentioned above:

[0097]

[0098] Table 2 below shows the specific quantitative evaluation results of PSFTDNet, AHDRNet, and ADNet in different scenarios on the NTIRE dataset:

[0099] As can be seen, PSFTDNet shows a significant improvement over other networks in most cases.

Claims

1. A high dynamic range imaging network method based on pyramid structure spatial feature transformation, characterized by: Based on the following three sub-networks: 1) a deformable convolutional alignment network with shared offsets; 2) a conditional network based on the pyramid structure spatial feature transform PSFT; 3) Fusion network based on deformable convolutional residual dense blocks; The process of the imaging network method is as follows: 1) Use gamma correction to align the exposure of the input image, that is, adjust the exposure of images with different exposures to the same level, and then input all the images with aligned and unaligned exposures into two convolutional layers of a deformable convolutional alignment network with shared offsets to obtain preliminary image features. 2) Exposure-aligned image features are geometrically aligned in a deformable convolutional alignment network with shared offsets. That is, the geometric features of multiple images are aligned, and the offsets obtained in the geometric alignment process are also applied to the LDR image features that are not exposure-aligned for geometric alignment. Align the exposure only by correcting the gamma; 3) Input the geometrically aligned but misaligned LDR image features into a PSFT-based conditional network to obtain optimized features; 4) Input all the obtained features into a DERDB-based fusion network to obtain the final HDR image, i.e., a high dynamic range image; The image is input into the network, and features are extracted and aligned at the feature level. The network is a deformable convolutional alignment network with shared offsets. The feature refinement process is carried out by a conditional network based on pyramid structure spatial feature transformation, a fusion network based on deformable convolutional residual dense blocks, and a feature refinement process after feature aligning before feature fusion. The processing steps are as follows: 1-1) The steps for generating an HDR image from multiple LDR images with different exposures are as follows: The input LDR images with different exposures, denoted as L1, L2, and L3, are gamma-corrected to obtain corresponding exposure-aligned images, denoted as H1, H2, and H3. The gamma correction process is represented as follows: in express Exposure time, The parameter for gamma correction is set to 2.2; 1-2) Input the image into a deformable convolutional alignment network with shared offsets to extract and align features; input the exposed but geometrically aligned image features into a PSFT-based conditional network to obtain refined features; 1-3) Input all the features obtained in the above steps into the DERDB-based fusion network to obtain the final generated HDR image; the DERDB fusion network refers to the fusion network of deformable residual dense blocks / deformable residual dense blocks; 1-4) The steps for extracting and aligning features from a deformable convolutional alignment network are as follows: 1-4-1) In the alignment network, the input { The initial features are obtained after two convolutional layers. in and They represent from and Features extracted from; 1-4-2) will A deformable convolutional alignment network with a pyramid structure is input to obtain aligned features, and the offsets learned in this process are applied to the alignment network. Alignment; This process can be represented as: in This represents the learned offset. This represents a mapping of several convolutional layers. Indicates a connection operation. This represents a deformable convolution operation. , Indicates the aligned features; : Indicates the features of the reference exposure-aligned image; : This parameter is derived from the features of the LDR domain input image and then aligned. 1-5) Refining features to fully utilize the complementary information between images: The steps of feature refinement using a conditional network based on pyramid structure spatial feature transformation, i.e., a PSFT-based conditional network, are as follows: 1-5-1) Based on the pyramid structure spatial feature transformation module, i.e., the PSFT module, feature downsampling is used to obtain image features of different spatial scales; 1-5-2) Perform the Spatial Feature Transform (SFT) process on the obtained image features of different scales. The SFT process is represented as follows: in, This represents the learned affine transformation parameters. Represents the mapping of convolutional layers. Represents the Adama product; 1-5-3) The final refined image features are obtained by upsampling the results from each layer. This process is represented as follows: in, Indicates the first Intermediate condition network of layers, This indicates upsampling with a base of 2. and This represents a mapping of several convolutional layers. This indicates the refinement process through SFT. Layer features; Alignment features of image i in the LDR domain at or after the Lth layer; : A mapping function consisting of several convolutional layers at layer L; The intermediate conditions of the i-th image in layer L; : Affine transformation parameter pair at layer L; : The mapping function of the Lth convolutional layer; The features of the i-th image after SFT spatial feature transformation at layer L; Features in layer L+1; : This is a feature of the reference exposure aligned image. This parameter is derived from the features of the Lth layer image after alignment in the LDR domain input.