High Dynamic Range Imaging Method and Apparatus Thereof

By using multiple LDR images and a Swin Fourier convolution network, the method addresses color distortion and ghosting in HDR imaging, resulting in high-quality HDR images with preserved texture.

JP2025520989AActive Publication Date: 2025-07-04CHUNG ANG UNIV IND ACADEMIC COOP FOUND
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024529780
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-22
Filing Date
2023-12-27
Publication Date
2025-07-04
Estimated Expiration
2043-12-27

AI Technical Summary

Technical Problem

Conventional high dynamic range (HDR) imaging methods suffer from color distortion and ghost phenomena when converting low dynamic range (LDR) images into HDR images, and fail to preserve texture effectively.

Method used

A method involving the acquisition of multiple LDR images with different exposure values, application of a weighted value map estimation model to generate an input feature map, and processing through a Swin Fourier convolution network model for Fourier transform to minimize color distortion and reduce ghosting.

Benefits of technology

The method achieves HDR images with minimized color distortion and preserved texture, effectively reducing ghost phenomena.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520989000001_ABST
    Figure 2025520989000001_ABST
Patent Text Reader

Abstract

Disclosed are a high dynamic range imaging method and apparatus that minimize color distortion, preserve texture, and reduce ghosting. 【Solution means】The high dynamic range imaging method of the present invention includes: (a) obtaining a first Low Dynamic Range (LDR) image, a second LDR image, and a third LDR image each having a different exposure value; (b) applying the first, second, and third LDR images to a weighted value map estimation model to extract a weighted value map, and using the extracted weighted value map to generate an input feature map by integrating the first, second, and third LDR images; and (c) applying the input feature map to a Swin Fourier convolution network model for Fourier transform, and after convolution, connecting a reference feature map to generate a High Dynamic Range (HDR) image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a high dynamic range imaging method and an apparatus therefor.

Background Art

[0002] The dynamic range of a digital image indicates the range of measured values that can be indicated by a physical sensor in an imaging system. However, when the exposure time is insufficient or excessive, an image obtained due to the functional limitations of the physical sensor has a low dynamic range (LDR), and the subjective image quality deteriorates. To solve such problems, many studies on high dynamic range imaging methods have been conducted.

[0003] Conventional imaging systems incorporate some HDR functions. In the case of the motion perception-based method, a weighted value map, a threshold bitmap, and similar techniques are utilized to synthesize an HDR image from LDR images having multiple exposure values for motions occurring in the image.

[0004] The alignment-based method aligns LDR images using methods such as the twist around the exposure value image to expand the dynamic range.

[0005] However, the prior art has succeeded in obtaining HDR images but still has problems such as motion and color distortion occurring in LDR images.

Summary of the Invention

Problems to be Solved by the Invention

[0006] The present invention has been made in view of the above-described conventional problems, and an object of the present invention is to provide a high dynamic range imaging method and an apparatus therefor that minimize color distortion occurring when converting an LDR image into an HDR image, preserve texture, and reduce ghost phenomena.

Means for Solving the Problems

[0007] A high dynamic range imaging method according to one aspect of the present invention made to achieve the above object includes: (a) obtaining a first LDR (Low Dynamic Range) image, a second LDR image, and a third LDR image each having a different exposure value; (b) applying the first LDR image, the second LDR image, and the third LDR image to a weighted value map estimation model to extract a weighted value map, and generating an input feature map by integrating the first LDR image, the second LDR image, and the third LDR image using the extracted weighted value map; and (c) applying the input feature map to a Swin Fourier convolution network model for Fourier transform, and connecting a reference feature map after convolution to generate an HDR (High Dynamic Range) image.

[0008] The reference feature map may be a feature map of the second LDR image extracted by the weighted value map estimation model. The step (b) may include: applying the first LDR image, the second LDR image, and the third LDR image to a convolution layer of the weighted value map estimation model to extract a first feature map, a second feature map, and a third feature map respectively; applying the first feature map and the second feature map to a first attention module of the weighted value map estimation model to generate a first weighted value map, and applying the third feature map and the second feature map to a second attention module of the weighted value map estimation model to generate a second weighted value map; and reflecting the first weighted value map and the second weighted value map on the first feature map and the third feature map respectively, and then combining them with the second feature map to generate the input feature map. The Swin-Fourier convolutional network model is composed of a plurality of Swin-Fourier convolutional blocks and transpose convolutional blocks. The Swin-Fourier convolutional block includes a Swin Transformer block that divides the input feature map into patches to form a hierarchical feature map, a residual block that adds a global residual using skip connections to the feature map of the reference image, a first 3×3 convolutional layer and a second 3×3 convolutional layer connected to the backend of the Swin Transformer block. The output values of the first 3×3 convolutional layer and the second 3×3 convolutional layer are synthesized element by element to output a first sub-synthesis result value. The output values of a third 3×3 convolutional layer connected to the backend of the residual block and a spectral transformation module that Fourier-transforms the output value of the residual block are synthesized element by element to output a second sub-synthesis result value. It may include a concatenation part that concatenates the first sub-synthesis result value and the second sub-synthesis result value. The loss function of the Swin-Fourier convolutional network model can be calculated as shown in Equation 8 below.

Equation

Equation

Equation

Equation

[0009] A high dynamic range (HDR) imaging device according to one aspect of the present invention made to achieve the above object includes an image acquisition unit that acquires a first low dynamic range (LDR) image, a second LDR image, and a third LDR image having different exposure values, and applies the first LDR image, the second LDR image, and the third LDR image to a weighted value map estimation model to extract a weighted value map, and uses the extracted weighted value map to integrate the first LDR image, the second LDR image, and the third LDR image to generate an input feature map. And an HDR generation unit that applies the input feature map to a Swin Fourier convolution network model for Fourier transform, connects a reference feature map after convolution, and generates an HDR (High Dynamic Range) image.

[0010] The weighted value map estimation unit applies the first LDR image, the second LDR image, and the third LDR image to the convolutional layer of the weighted value map estimation model to extract a first feature map, a second feature map, and a third feature map respectively. The first feature map and the second feature map are applied to the first attention module of the weighted value map estimation model to generate a first weighted value map, and the third feature map and the second feature map are applied to the second attention module of the weighted value map estimation model to generate a second weighted value map. After reflecting the first weighted value map and the second weighted value map on the first feature map and the third feature map respectively, they can be combined with the second feature map to generate the input feature map.

Advantages of the Invention

[0011] According to the high dynamic range synthesis method and apparatus of the present invention, it is possible to minimize color distortion that occurs when converting an LDR image into an HDR image, and there is an effect that an HDR image with preserved texture and no ghost phenomenon can be obtained.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0013] As used in this specification, the singular expressions include plural expressions unless otherwise clearly indicated in the context. In this specification, terms such as "configured to" or "including" should not be construed as necessarily including all of the many components or many steps described in the specification, and may not include some of the components or some of the steps, or may need to be construed as further including additional components or steps. Also, terms such as "… part" and "module" described in the specification mean a unit that processes at least one function or operation, which may be embodied in hardware or software, or in a combination of hardware and software.

[0014] Hereinafter, specific examples of embodiments for carrying out the present invention will be described in detail with reference to the drawings.

[0015] FIG. 1 is a flowchart showing a high dynamic range imaging method according to an embodiment of the present invention, FIG. 2 is a diagram showing a network architecture of an HDR imaging device according to an embodiment of the present invention, FIG. 3 is a diagram showing a structure of a Swin Fourier convolution model according to an embodiment of the present invention, FIG. 4 is a diagram showing a structure of a Swin Transformer block according to an embodiment of the present invention, FIG. 5 is a diagram showing a structure of a residual block according to an embodiment of the present invention, FIG. 6 is a diagram showing a structure of a spectral conversion block according to an embodiment of the present invention, FIG. 7 is a diagram showing a weighted value map result according to an embodiment of the present invention, and FIG. 8 is a diagram showing activation results due to perceptual loss according to the prior art and an embodiment of the present invention.

[0016] In step 110, the HDR imaging device 100 acquires a plurality of LDR images having different exposure values.

[0017] For example, the HDR imaging device 100 acquires a first LDR image, a second LDR image, and a third LDR image having different exposure values. Here, it is assumed that the first LDR image, the second LDR image, and the third LDR image are consecutive images having different exposure values. For the sake of convenience, it is assumed that the exposure value of the first LDR image is the smallest and the exposure value of the third LDR image is the largest, and the description will be centered around this.

[0018] For example, the LDR image input to obtain an HDR image is represented as g i and this means images with different exposures.

[0019] Geometrically aligned images The intensity conversion version of TIFF2025520989000006.tif10151 is obtained using gamma correction as in Equation 1.

[0020]

Equation

[0021] Here, L i represents the LDR image, TIFF2025520989000008.tif11151 represents the HDR image, γ represents the gamma parameter, and in one embodiment of the present invention, it is set to 2.2. t i represents the exposure time.

[0022] TIFF2025520989000009.tif10128 is obtained by linearizing the non-linear LDR input image using the camera response function and applying gamma correction. L generated in the preprocessing process i and TIFF2025520989000010.tif10128 are defined as the input image Gi, and an HDR image is obtained as the connected LDR input image. When this is expressed by a mathematical formula, it is shown as in Formula 2.

[0023]

Equation

[0024] The overall network model of the high dynamic range imaging device according to one embodiment of the present invention is defined as in Formula 3.

[0025]

Equation

[0026] Here, TIFF2025520989000013.tif10128 represents the output (HDR image) of the overall network model, and θ represents the learning intermediate variable.

[0027] In step 115, the HDR imaging device 100 generates a weight map using the first LDR image, the second LDR image, and the third LDR image, and integrates the first LDR image, the second LDR image, and the third LDR image using the generated weight map to generate an input feature map.

[0028] Refer to FIG. 2 and explain this in more detail.

[0029] The HDR imaging device 100 applies a convolution operation to each of the first LDR image, the second LDR image, and the third LDR image to generate a first feature map, a second feature map, and a third feature map, respectively.

[0030] The HDR imaging device 100 applies the first feature map and the second feature map to an attention module to generate a first weighted value map, and applies the third feature map and the second feature map to the attention module to generate a second weighted value map. The second feature map is used as a reference feature map.

[0031] When this is expressed by a mathematical formula, it is as shown in Formula 4.

[0032]

Number

[0033] Here, G i represents an LDR image, H1 represents an extracted feature map layer, and f i represents an attention module, TIFF2025520989000015.tif11128 represents a weighted value map, and the weighted value map has values between [0, 1].

[0034] After the HDR imaging device 100 reflects the first weighted value map and the second weighted value map on the first feature map and the third feature map, respectively, it combines them with the reference feature map (the second feature map) to generate an input feature map. To minimize ghosting, the HDR imaging device 100 reflects the first weighted value map on the first feature map and the second weighted value map on the third feature map, and then combines them with the reference feature map (the second feature map) to generate an input feature map.

[0035] That is, the HDR imaging device 100 reflects the first weight value map on the first feature map by multiplying the first weight value map and the first feature map element by element, and multiplies the second weight value map and the third feature map element by element to reflect the second weight value map on the third feature map. Next, the HDR imaging device 100 combines the first feature map and the third feature map on which the weight value map is reflected with the reference feature map (second feature map) to generate an input feature map.

[0036] As a result, as shown in FIG. 7, in a region where a ghost artifact is generated due to the movement of an object different from the reference image, an undesired artifact can be minimized by lowering the weight value, which is an element-by-element multiplication of the first weight value map H1(G1) and the second weight value map H1(G2), and is expressed as Equation 5.

[0037]

Equation

[0038] Here, F 2 represents the feature map and the channel-by-channel connection, and C(·) represents an operator that performs the connection after element-by-element multiplication.

[0039] F estimated to minimize ghost 2 is used as an input feature map for learning an improved perceptual loss function.

[0040] In 120 steps, the HDR imaging device 100 applies the input feature map to a Swin-Fourier convolutional network model for Fourier transform, performs convolutional operations, and then connects the reference feature map to generate an HDR (High Dynamic Range) image.

[0041] As shown in FIG. 2, the Swin-Fourier convolutional network model is composed of a plurality of Swin-Fourier convolutional blocks (SCblock), residual blocks (RB), and transpose convolutional blocks (TCblock).

[0042] The input feature map is reconstructed into an HDR image through a U-shaped network model (i.e., a Swin-Fourier convolutional network model) composed of a Swin-Fourier convolutional block and a transposed convolutional block.

[0043] The Swin-Fourier convolutional block is composed of a Swin-Transformer block based on a transformer structure that divides the feature map into patches to form a hierarchical feature map, and a residual block that adds a global residual using skip connections to the feature map of the reference image.

[0044] As shown in Figure 4, the Swin-Transformer block is composed of a layer normalization block, a window-based multi-head self-attention module, a multi-layer perceptron module, and a structure in which 1×1 convolution is performed after the convolutional layer.

[0045] The Swin-Fourier convolutional block allows a feature map that calls the U-shaped encoder module, so it is downsampled by 2×2 stride convolution to further reduce the resolution, and the transformer is integrated into the U-shaped network block to make better use of hierarchical information. Also, the residual block is as shown in Figure 5.

[0046] Furthermore, the Swin-Fourier convolutional network model has a first 3×3 convolutional layer and a second 3×3 convolutional layer connected to the rear end of the Swin-Transformer block, synthesizes the output values of the first 3×3 convolutional layer and the second 3×3 convolutional layer element by element to output a first sub-synthesis result value, and synthesizes the output values of the third 3×3 convolutional layer connected to the rear end of the residual block and the output value of the spectral conversion module that performs Fourier transform on the output value of the residual block element by element to output a second sub-synthesis result value, and further includes a concatenation part that concatenates the first sub-synthesis result value and the second sub-synthesis result value.

[0047] The transposed convolutional block used in the decoder module is composed of a transposed convolutional block and a Swin-Fourier convolutional block. To hierarchically learn global information, it upsamples with a 2×2 transposed convolution to increase the resolution.

[0048] Also, according to an embodiment of the present invention, the Swin-Fourier convolutional network model adds the feature map of the reference image as a global residual using skip connections to assist learning.

[0049] The spectral transformation block combined with the Swin-Fourier convolutional block performs convolution after FFT (Fast-Fourier Transform), and its structure is as shown in FIG. 6. Since the spectral transformation block uses FFT, it is easy to process the periodic patterns represented in the image, and it can be used together with the Swin-Fourier convolutional block to reinforce not only local information but also global context. FFT converts the image signal into a periodic frequency signal, thereby enabling the learning of the periodic patterns of the image.

[0050] According to an embodiment of the present invention, the Swin-Fourier convolutional network model adds the second feature map for the reference image (i.e., the second LDR image) as a global residual using global skip connections for residual learning.

[0051] The Swin-Fourier convolutional network model according to an embodiment of the present invention is learned using a logarithmic perceptual loss function.

[0052] Expressing this in a mathematical formula is as shown in Formula 6.

[0053]

Equation

[0054] Here, x represents the target HDR image, and y represents the generated HDR image.

Equation

[0055]

Number

[0056] Here, μ represents the compression intermediate variable.

[0057] Therefore, the total loss function is shown as in Equation 8.

[0058]

Number

[0059] Here, l1 represents the MAE loss, and VGG19() represents a Swin-Fourier convolutional neural network model as a convolutional neural network.

[0060] As shown in FIG. 8, by applying the log perceptual loss function according to an embodiment of the present invention, compared with the prior art, the activation is enhanced so that the region having a low value has a high value, and the region having a high value maintains a high value.

[0061] FIG. 9 is a diagram schematically showing the internal configuration of an HDR imaging device according to an embodiment of the present invention.

[0062] Referring to FIG. 9, an HDR imaging device 100 according to an embodiment of the present invention includes an image acquisition unit 910, a weight value map estimation unit 920, an HDR generation unit 930, a memory 940, and a processor 950.

[0063] The image acquisition unit 910 is means for acquiring a first LDR (Low Dynamic Range) image, a second LDR image, and a third LDR image each having a different exposure value.

[0064] The weight value map estimation unit 920 has a weight value map estimation model, and after applying the first, second, and third LDR images acquired by the image acquisition unit 910 to the weight value map estimation model to extract a weight value map, it is a means for generating an input feature map by integrating the first, second, and third LDR images using the extracted weight value map.

[0065] For example, the weight value map estimation unit 920 applies the first, second, and third LDR images to the convolutional layer of the weight value map estimation model to extract a first feature map, a second feature map, and a third feature map respectively. Next, the weight value map estimation unit 920 applies the first feature map and the second feature map to the first attention module of the weight value map estimation model to generate a first weight value map, and applies the third feature map and the second feature map to the second attention module of the weight value map estimation model to generate a second weight value map. The weight value map estimation unit 920 reflects the first weight value map and the second weight value map on the first feature map and the third feature map respectively, and then combines them with the second feature map to generate an input feature map.

[0066] The HDR imaging device 930 is a means for applying an input feature map to a Swin-Fourier convolutional network model for Fourier transform, performing convolutional operations, and then connecting a reference feature map to generate an HDR (High Dynamic Range) image.

[0067] The Swin-Fourier convolutional network model is a U-shaped network model composed of a plurality of Swin-Fourier convolutional blocks and transpose convolutional blocks.

[0068] As described above, the Swin-Fourier convolution block includes a Swin Transformer block that divides the input feature map into patches to construct a hierarchical feature map, a residual block that adds a global residual to the feature map of the reference image using skip connections, and a first 3×3 convolution layer and a second 3×3 convolution layer connected to the rear end of the Swin Transformer block. The output values of the first 3×3 convolution layer and the second 3×3 convolution layer are combined element-wise to output a first sub-combined result value. The output values of a third 3×3 convolution layer connected to the rear end of the residual block and a spectral conversion module that Fourier-transforms the output value of the residual block are combined element-wise to output a second sub-combined result value. It is configured to include a concatenation unit that concatenates the first sub-combined result value and the second sub-combined result value.

[0069] With such a Swin-Fourier convolution network model, there is an advantage that HDR synthesis can be easily performed by dividing the global region of the input feature map into patches, performing Fourier transform, and learning using periodic patterns.

[0070] Also, as described above, by learning the U-shaped network model using logarithmic perceptual loss, activation is enhanced so that regions with low values have high values and regions with high values maintain high values.

[0071] The memory 940 stores various instruction words for performing the HDR imaging method according to an embodiment of the present invention.

[0072] The processor 950 is a means for controlling internal components of the HDR imaging device 100 according to an embodiment of the present invention (for example, the image acquisition unit 910, the weighted value map estimation unit 920, the HDR generation unit 930, the memory 940, etc.).

[0073] The apparatus and method according to embodiments of the present invention are embodied in the form of program instructions executed by various computer means and recorded on a computer-readable recording medium. The computer-readable recording medium includes program instructions, data files, data structures, etc., alone or in combination. The program instructions recorded on the computer-readable recording medium are either specially designed and configured for the present invention or are known and usable by those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices such as ROMs, RAMs, and flash memories that are specially configured to store and execute program instructions. Examples of program instructions include not only machine language codes such as those created by compilers but also high-level language codes that can be executed by a computer using an interpreter or the like.

[0074] The above-described hardware device is configured to operate with at least one software module in order to perform the operations of the present invention, and vice versa.

[0075] As described above, the present invention has been described centering on its embodiments. Those skilled in the art will understand that the present invention can be embodied in a modified form without departing from the essential characteristics of the present invention. Therefore, the disclosed embodiments should be considered from an illustrative rather than a limiting perspective. The scope of the present invention is shown in the claims rather than the above description, and all differences within the scope equivalent thereto should be construed as being included in the present invention.

Description of Reference Numerals

[0076] 100 HDR imaging device 910 Image acquisition unit 920 Weight value map estimation unit 930 HDR generation unit 940 Memory 950 Processor

Claims

1. (a) Obtaining a first Low Dynamic Range (LDR) image, a second LDR image, and a third LDR image each having different exposure values; (b) Applying the first LDR image, the second LDR image, and the third LDR image to a weighted value map estimation model to extract a weighted value map, and using the extracted weighted value map to generate an input feature map by integrating the first LDR image, the second LDR image, and the third LDR image; (c) Applying the input feature map to a Swin Fourier convolutional network model for Fourier transform, performing convolutional operation, and then connecting a reference feature map to generate a High Dynamic Range (HDR) image. A high dynamic range imaging method characterized by comprising the above steps.

2. The high dynamic range imaging method according to claim 1, wherein the reference feature map is a feature map of the second LDR image extracted by the weighted value map estimation model.

3. The step (b) Applying the first LDR image, the second LDR image, and the third LDR image to a convolutional layer of the weighted value map estimation model to extract a first feature map, a second feature map, and a third feature map respectively; Applying the first feature map and the second feature map to a first attention module of the weighted value map estimation model to generate a first weighted value map, and applying the third feature map and the second feature map to a second attention module of the weighted value map estimation model to generate a second weighted value map; After reflecting the first weighted value map and the second weighted value map on the first feature map and the third feature map respectively, combining them with the second feature map to generate the input feature map. The high dynamic range imaging method according to claim 1, characterized by comprising the above steps.

4. The Swin Fourier convolutional network model is composed of a plurality of Swin Fourier convolutional blocks and transpose convolutional blocks. The Swin Fourier convolutional block A Swin Transformer block that divides the input feature map into patches to form a hierarchical feature map A residual block that adds a global residual using skip connections to the feature map of a reference image, and a first 3×3 convolutional layer and a second 3×3 convolutional layer connected to the rear end of the Swin Transformer block, synthesizing the output values of the first 3×3 convolutional layer and the second 3×3 convolutional layer element by element to output a first sub-synthesis result value, and synthesizing the output values of a third 3×3 convolutional layer connected to the rear end of the residual block and the output value of the spectral transformation module that performs Fourier transform on the output value of the residual block element by element to output a second sub-synthesis result value, and a concatenation unit that concatenates the first sub-synthesis result value and the second sub-synthesis result value. The high dynamic range imaging method according to claim 3 is characterized by comprising the above.

5. The loss function of the Swin Fourier convolutional network model is calculated as shown in the following mathematical formula 8. The high dynamic range imaging method according to claim 1 is characterized by this. 【Number 8】 Here, 【Number 6】 where x represents the target HDR image and y represents the generated HDR image. 【Number 61】 represents a Gaussian kernel that generalizes the Euclidean distance to a manifold using geodesic distance, and T(·) is a tone mapping operator. 【Number 7】 Calculated as such, μ represents the compression medium variable, and l 1 represents the MAE loss, and VGG19() represents the convolutional neural network.

6. A computer-readable recording medium recording program code for executing the method according to any one of claims 1 to 5.

7. An image acquisition unit that acquires a first LDR (Low Dynamic Range) image, a second LDR image, and a third LDR image each having a different exposure value; a weighted value map estimation unit that applies the first LDR image, the second LDR image, and the third LDR image to a weighted value map estimation model to extract a weighted value map, and generates an input feature map by integrating the first LDR image, the second LDR image, and the third LDR image using the extracted weighted value map; an HDR generation unit that applies the input feature map to a Swin Fourier convolutional network model to perform Fourier transform, performs convolutional operation, and then connects a reference feature map to generate an HDR (High Dynamic Range) image. The high dynamic range imaging apparatus is characterized by comprising the above.

8. The high dynamic range imaging device according to claim 7, wherein the reference feature map is a feature map of the second LDR image extracted by the weighted value map estimation model.

9. The weighted value map estimation unit applies the first LDR image, the second LDR image, and the third LDR image to the convolutional layer of the weighted value map estimation model to extract a first feature map, a second feature map, and a third feature map respectively, applies the first feature map and the second feature map to the first attention module of the weighted value map estimation model to generate a first weighted value map, and applies the third feature map and the second feature map to the second attention module of the weighted value map estimation model to generate a second weighted value map, After reflecting the first weighted value map and the second weighted value map on the first feature map and the third feature map respectively, they are combined with the second feature map to generate the input feature map. The high dynamic range imaging device according to claim 7 is characterized in that.

10. The Swin Fourier convolutional network model is composed of a plurality of Swin Fourier convolutional blocks and transpose convolutional blocks, The Swin Fourier convolutional block includes a Swin Transformer block that divides the input feature map into patches to form a hierarchical feature map, a residual block that adds a global residual using skip connections to the feature map of the reference image, has a first 3×3 convolutional layer and a second 3×3 convolutional layer connected to the rear end of the Swin Transformer block, synthesizes the output values of the first 3×3 convolutional layer and the second 3×3 convolutional layer element by element to output a first sub-synthesis result value, and synthesizes the output values of the third 3×3 convolutional layer connected to the rear end of the residual block and the output value of the spectral conversion module that Fourier-transforms the output value of the residual block element by element to output a second sub-synthesis result value, and a concatenation part that concatenates the first sub-synthesis result value and the second sub-synthesis result value. The high dynamic range imaging device according to claim 9 is characterized by comprising.

Citation Information

Patent Citations

  • High-dynamic-range ghosting-eliminating imaging system and method based on attention module

    CN113160178A

  • U-shaped image segmentation network based on convolution enhanced cross self-attention deformer

    CN115908805A

  • Sealing gasket

    KR1020200140490A

  • Generation of high dynamic range visual media

    US20190096046A1

  • Image processing apparatus and method

    US20200134787A1