An ultra-high dynamic range imaging system and method based on a diffusion model image prior
By employing pre-alignment and feature fusion techniques based on a diffusion model, the challenges of image alignment and fusion in high dynamic range scenes are solved, achieving natural tone mapping effects and high-quality low dynamic range image generation.
Patent Information
- Application Number
- CN202411958196.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing technologies struggle to generate high-quality fusion results when dealing with high dynamic range scenes with significant exposure differences, and the tone mapping process is prone to introducing artifacts and contrast loss.
The algorithm employs a pre-alignment module based on a diffusion model, a variational autoencoder, a decomposition-fusion control module, and a fidelity control module. It aligns images using an optical flow estimation algorithm, performs feature fusion by combining cross-attention and multi-scale cross-attention mechanisms, and utilizes prior knowledge from the diffusion model for tone mapping to ensure the fidelity of the fusion result.
Generates natural and realistic tone mapping effects in extremely high dynamic range scenes, improves the robustness of alignment and blending, reduces artifacts and contrast loss, and produces high-quality low dynamic range images.
Smart Images

Figure CN119743679B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of high dynamic range imaging, and in particular to an ultra-high dynamic range imaging system and method based on a diffusion model image prior. BACKGROUND
[0002] Capturing high dynamic range (HDR) scenes is one of the key challenges in camera system design. At present, most cameras use exposure fusion technology to expand the dynamic range by fusing images taken at different exposure levels.
[0003] High dynamic range (HDR) imaging can be divided into two methods according to the domain where fusion occurs. One is the HDR fusion method, which first fuses multiple images at different exposures to generate an HDR image, and then compresses it into a low dynamic range (Low-Dynamic Range, LDR) image through tone mapping for visualization or display; the other is the multi-exposure fusion method, which directly fuses multiple images at different exposures to output a low dynamic range image, saving the step of tone mapping.
[0004] In the field of HDR fusion, Kalantari et al. [1] proposed an alignment and fusion-based HDR fusion method, laying an important foundation for this field. Subsequently, researchers have continuously improved the alignment process by developing more advanced modules to address the artifacts caused by motion between different exposures. Steven Tel et al. [2] further proposed an HDR fusion method without explicit alignment, significantly simplifying the alignment step. Recently, Kong et al. [3] proposed a novel and efficient HDR fusion technique based on optical flow. In addition, some research [4,5] attempts to implement HDR fusion algorithms based on diffusion model frameworks, but has not fully utilized the powerful prior knowledge contained in pre-trained diffusion models. Although the above methods have made some progress in the field of HDR fusion, they are usually only applicable to scenes with small exposure differences (usually 3-4 stops). When applied to ultra-high dynamic range scenes (i.e., scenes with large exposure differences), these methods often struggle to generate high-quality fusion results due to alignment errors, inconsistent lighting conditions, or artifacts in the tone mapping process.
[0005] In the field of multi-exposure fusion, existing research [6,7,8] mainly focuses on static scenes, directly fusing multiple images to generate low dynamic range images through supervised or self-supervised methods. This method often struggles to generate high-quality fusion results when dealing with moving scenes or complex scenes with large exposure differences.
[0006] In summary, to handle dynamic scenes, most HDR fusion algorithms in the prior art usually first perform an alignment operation on the input frames. When there is a large brightness difference between the input frames, alignment becomes exceptionally difficult, which often leads to the appearance of the "ghosting" problem.
[0007] In addition, most HDR algorithms assume that underexposed images are simply darker versions of normally exposed images. However, with changes in exposure levels, the appearance of objects can change significantly, which often leads to unnatural fusion results, affecting image quality.
[0008] And the result of prior art fusion is usually an HDR image, but since ordinary low dynamic range display devices cannot directly present HDR images, these images need to be further compressed through tone mapping. When the dynamic range is too high, tone mapping can introduce additional problems, such as the loss of natural contrast and details, thereby affecting the final output effect.
[0009] Prior art references:
[0010] [1] Kalantari N K, Ramamoorthi R. Deep high dynamic range imaging of dynamic scenes [J]. ACM Trans. Graph., 2017, 36(4): 144: 1-144: 12.
[0011] [2] Tel S, Wu Z, Zhang Y, et al. Alignment-free HDR Deghosting with Semantics Consistent Transformer [C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023: 12836-12845.
[0012] [3] Kong L, Li B, Xiong Y, et al. SAFNet: Selective Alignment Fusion Network for Efficient HDR Imaging [C] / / European Conference on Computer Vision. Springer, Cham, 2025: 256-273.
[0013] [4] Yan Q, Hu T, Sun Y, et al. Towards high-quality hdr deghosting with conditional diffusion models [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2023.
[0014] [5] Hu T, Yan Q, Qi Y, et al. Generating content for hdr deghosting from frequency view [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024:25732-25741.
[0015] [6] Liang P, Jiang J, Liu X, et al. Fusion from decomposition: A self-supervised decomposition approach for image fusion [C] / / European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022:719-735.
[0016] [7] Jiang T, Wang C, Li X, et al. Meflut: Unsupervised 1d lookup tables for multi-exposure image fusion [C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023:10542-10551.
[0017] [8] Wu G, Fu H, Liu J, et al. Hybrid-supervised dual-search: Leveraging automatic learning for loss-free multi-exposure image fusion [C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(6): 5985-5993. SUMMARY
[0018] The purpose of the present application is to overcome the defects of the prior art and provide an ultra-high dynamic range imaging system and method based on a diffusion model image prior, which can improve the alignment and fusion robustness in the case of large exposure level difference and generate natural and realistic tone mapping effects in the ultra-high dynamic range scene.
[0019] The purpose of the present application can be achieved by the following technical solutions: an ultra-high dynamic range imaging system based on a diffusion model image prior, comprising a pre-alignment module, a variational autoencoder, a decomposition and fusion control module, a fidelity control module and a variational auto-decoder, the pre-alignment module is used to align the short exposure image to the long exposure image to generate a pre-alignment result;
[0020] The variational autoencoder is used to encode the long exposure image to obtain an encoding result;
[0021] The decomposition and fusion control module controls the fusion process of the diffusion model based on the pre-alignment result and the encoding result, and performs information fusion and tone mapping on the input exposure image to obtain a fusion result;
[0022] The fidelity control module is based on the long exposure image and the pre-alignment result, and is used to control the fidelity of the fusion result;
[0023] The variational auto-decoder is used to decode the fusion result to obtain a reconstructed image.
[0024] Further, the pre-alignment module uses an optical flow estimation algorithm to align the short exposure image to the long exposure image and mask the occluded area to generate a pre-alignment result.
[0025] Further, the decomposition and fusion control module includes a short exposure feature extraction unit, a long exposure feature extraction unit and a fusion unit, the short exposure feature extraction unit is used to extract image features of the short exposure image;
[0026] The long exposure feature extraction unit is used to extract image features of the long exposure image from the encoding result;
[0027] The fusion unit fuses and tone maps the image features of the short-exposure image and the long-exposure image based on a diffusion model and using prior knowledge.
[0028] Further, the short-exposure feature extraction unit includes a plurality of cascaded short-exposure feature extractors.
[0029] The long-exposure feature extraction unit includes a plurality of cascaded long-exposure feature extractors.
[0030] Further, the fusion unit employs a cross-attention mechanism and a multi-scale cross-attention mechanism to fuse the image features of the short-exposure image and the long-exposure image.
[0031] An ultra-high dynamic range imaging method based on diffusion model image prior, comprising the following steps:
[0032] S1, using an optical flow estimation algorithm to transform and align the short-exposure image to the long-exposure image, and masking the occluded area to obtain a pre-alignment result;
[0033] S2, decompose the information of the short-exposure image and the long-exposure image, and fuse and tone map the pre-alignment result by controlling the fusion process of the diffusion model to obtain a fusion result;
[0034] S3, according to the long-exposure image and the pre-alignment result, controlling the fidelity of the fusion result, and then decoding to obtain a reconstructed image.
[0035] Further, the step S1 includes the following steps:
[0036] S11, using an optical flow estimation algorithm to process the long-exposure image and the short-exposure image respectively to obtain corresponding first optical flow and second optical flow;
[0037] S12, performing consistency check on the first optical flow and the second optical flow to obtain an occlusion mask;
[0038] S13, aligning the short-exposure image and the first optical flow, and masking the occluded area using the occlusion mask to generate a pre-alignment result.
[0039] Further, the step S2 includes the following steps:
[0040] S21, encoding the long-exposure image to obtain an encoded result;
[0041] S22, performing short-exposure feature extraction and long-exposure feature extraction on the short-exposure image and the encoded result respectively;
[0042] S23, the short exposure feature and the long exposure feature are fused by using a diffusion model, fusion and tone mapping of the pre-alignment result are completed, and a fusion result is obtained.
[0043] Further, the step S22 specifically includes the following steps:
[0044] For the short exposure image, the brightness component of the short exposure image is removed, the structure and color information are extracted, the same network structure as ControlNet is adopted, and the short exposure feature of the short exposure image is extracted through several convolution layers.
[0045] For the long exposure image, the long exposure feature is extracted from the hidden variable of the encoding result.
[0046] Further, the step S23 specifically includes the following steps:
[0047] Compared with the prior art, the present application has the following advantages:
[0048] The present application provides an ultra-high dynamic range imaging system based on a diffusion model image prior, which includes a pre-alignment module, a variational autoencoder, a decomposition fusion control module, a fidelity control module and a variational auto-decoder, the short exposure image is aligned to the long exposure image by using the pre-alignment module, and a pre-alignment result is generated; the long exposure image is encoded by using the variational autoencoder, and an encoding result is obtained; the decomposition fusion control module is used to control the fusion process of the diffusion model based on the pre-alignment result and the encoding result, information fusion and tone mapping are performed on the input exposure image, and a fusion result is obtained; and the fidelity control module is used to control the fidelity of the fusion result based on the long exposure image and the pre-alignment result. Therefore, the alignment and fusion problems caused by large exposure level difference can be solved, the tone mapping can be adaptively learned through the image prior of the diffusion model, the ultra-high dynamic range tone mapping problem can be solved, and the fidelity of the reconstructed image can be effectively maintained.
[0049] The present application uses an optical flow estimation algorithm to transform and align the short exposure image to the long exposure image, and to mask the occluded area, so as to obtain a pre-alignment result, thereby having better robustness in dealing with potential alignment errors and illumination changes.
[0050] The present application designs a decomposition fusion control scheme, removes the brightness component of the short exposure image, extracts the structure and color information thereof, combines the cross attention mechanism and the multi-scale cross attention mechanism to enhance the feature fusion with the long exposure image, and uses the prior information of the diffusion model to improve the robustness of the fusion.
[0051] The present application designs an additional fidelity control scheme, inputs the long-exposure image and the pre-alignment result into the fidelity control module to ensure the fidelity of the fusion result, and controls the fusion result to be highly consistent with the input texture. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 A method flowchart of the present application is shown in the figure.
[0053] Figure 2 A pre-alignment process schematic diagram in the embodiment is shown in the figure.
[0054] Figure 3 A decomposition and fusion control process schematic diagram in the embodiment is shown in the figure.
[0055] Figure 4 An image fusion and tone mapping process schematic diagram in the embodiment is shown in the figure.
[0056] Figure 5 And Figure 6 An imaging effect comparison schematic diagram in the embodiment is shown in the figure. DETAILED DESCRIPTION
[0057] The present application will be described in detail below in combination with the figures and specific embodiments.
[0058] EMBODIMENT
[0059] A super-high dynamic range imaging system based on a diffusion model image prior includes a pre-alignment module, a variational autoencoder, a decomposition and fusion control module, a fidelity control module, and a variational auto-decoder. The pre-alignment module is used to align a short-exposure image to a long-exposure image to generate a pre-alignment result. Specifically, an optical flow estimation algorithm can be used to align the short-exposure image to the long-exposure image, and the occluded area is masked to generate the pre-alignment result.
[0060] The variational autoencoder is used to encode the long-exposure image to obtain an encoding result.
[0061] The decomposition and fusion control module controls the fusion process of the diffusion model based on the pre-alignment result and the encoding result, performs information fusion and tone mapping on the input exposure image to obtain a fusion result.
[0062] Specifically, the decomposition and fusion control module includes a short-exposure feature extraction unit, a long-exposure feature extraction unit, and a fusion unit. The short-exposure feature extraction unit is used to extract image features of the short-exposure image, and a plurality of cascaded short-exposure feature extractors can be used.
[0063] The long-exposure feature extraction unit is used to extract image features of the long-exposure image from the encoding result, and a plurality of cascaded long-exposure feature extractors can be used.
[0064] The fusion unit is based on a diffusion model and employs a cross-attention mechanism and a multi-scale cross-attention mechanism to fuse and tone map the image features of short-exposure and long-exposure images using prior knowledge.
[0065] The fidelity control module, based on long-exposure images and pre-alignment results, is used to control the fidelity of the fusion result;
[0066] Variational autodecoders are used to decode the fusion results to obtain the reconstructed image.
[0067] Based on the above system, a method for ultra-high dynamic range imaging based on diffusion model image priors is implemented, such as... Figure 1 As shown, it includes the following steps:
[0068] S1. The optical flow estimation algorithm is used to transform and align the short exposure image onto the long exposure image, and the occluded area is masked to obtain the pre-alignment result;
[0069] S2. Information decomposition is performed on short-exposure and long-exposure images, and the pre-alignment results are fused and tone mapped by controlling the fusion process of the diffusion model to obtain the fusion result;
[0070] S3. Based on the long exposure image and the pre-alignment result, the fidelity of the fusion result is controlled, and then the reconstructed image is obtained through decoding.
[0071] This embodiment applies the above scheme, constructing a pre-alignment module, a diffusion model network, a decomposition and fusion control module, and a fidelity control module to achieve the following: Figure 2 The pre-alignment process shown (using optical flow estimation to transform and align the short-exposure image onto the long-exposure image, and masking occluded areas to obtain the pre-alignment result) is as follows: Figure 3 The decomposition and fusion control process shown (by removing the luminance component of the short-exposure image, extracting its structural and color information, and combining cross-attention and multi-scale cross-attention mechanisms to enhance feature fusion with the long-exposure image), is as follows: Figure 4 The image fusion and tone mapping process shown (utilizing the prior knowledge of the diffusion model, the results of pre-alignment of long exposure and short exposure images are fused and tone mapped, and an additional fidelity control module is used to control the fidelity of the reconstruction results).
[0072] Specifically, first, the pre-alignment module is used to align the short exposure image to the long exposure image to generate a pre-alignment result. Then, the long exposure image is encoded by the variational autoencoder to obtain an encoding result. Subsequently, the encoding result and the pre-alignment image are input into the decomposition and fusion control module, which is responsible for controlling the diffusion model network to fuse and tone map different exposure images to finally generate a fusion result. Finally, the long exposure image and the pre-alignment result are input into the fidelity control module to ensure the fidelity of the fusion result, and finally the reconstructed image is obtained.
[0073] The pre-alignment module is used to align two different exposure images to generate a pre-alignment result, providing a basis for subsequent information fusion. As shown in Figure 2 , the short exposure image is aligned to the long exposure image using the optical flow estimation algorithm RAFT, and the occluded area is masked to generate a pre-alignment result.
[0074] The decomposition and fusion control module is used to decompose the input different exposure images and control the fusion process of the diffusion model, thereby effectively realizing information fusion and tone mapping. As shown in Figure 4 , the module extracts the structure and color information of the short exposure image by removing its brightness component, and uses the same network structure as ControlNet to extract the information of the short exposure image through several convolution layers. Then, the cross-attention mechanism is used to fuse the features of the long exposure image. In addition, the module also uses a multi-scale cross-attention mechanism to enhance the feature fusion with the long exposure image and preserve more details when generating guide information, ensuring the accuracy of the fusion result.
[0075] The diffusion model network is used to fuse and tone map different exposure images. As shown in Figure 3 , the module uses the pre-trained diffusion model Stable Diffusion to fuse and tone map the pre-alignment result of the long exposure image and the short exposure image based on its prior knowledge.
[0076] The fidelity control module is used to ensure the fidelity of the fusion result. As shown in Figure 3 , the module controls the fidelity of the reconstructed image to ensure that the fusion result is highly consistent with the input texture. The fidelity control module uses the same network structure as the decomposition and fusion control module to further improve the authenticity of the fusion effect.
[0077] To verify the effectiveness of the present scheme, the present scheme and existing HDR fusion methods and multi-exposure fusion methods are used in this embodiment to process different exposure images, and the imaging results are compared, as shown in Figure 5 and Figure 6As shown, it can be seen that the scheme has strong robustness to potential alignment errors or illumination changes, and can generate natural tone mapping effects in very high dynamic range scenes. Even in complex scenes with exposure differences up to 9 stops (-6EV to +3EV), it can still generate beautiful and high-quality imaging results.
[0078] In summary, based on the image prior of the diffusion model, the scheme proposes a fusion and tone mapping imaging scheme suitable for ultra-high dynamic range dynamic scenes:
[0079] 1. To solve the alignment problem of large exposure level difference, the short exposure image is transformed and aligned to the long exposure image, and the occluded area is masked.
[0080] 2. To solve the fusion problem of large exposure level difference, a new decomposition-fusion control branch is proposed. By removing the brightness component of the short exposure image, its structure and color information are extracted, and cross-attention mechanism and multi-scale cross-attention mechanism are used to enhance feature fusion with the long exposure image. At the same time, the prior information of the diffusion model is used to improve the robustness of fusion.
[0081] 3. To solve the problem of ultra-high dynamic range tone mapping, the image prior of the diffusion model is used to adaptively learn the tone mapping.
[0082] 4. To ensure that the generated output is highly consistent with the real scene, an additional fidelity control branch is proposed, which can effectively maintain the fidelity of the result image.
[0083] Compared with traditional HDR fusion technology, the fusion effect of the scheme is better, and it follows the multi-exposure fusion process. Exposure fusion directly generates LDR output, avoiding cascading errors.
[0084] The scheme has stronger robustness in dealing with potential alignment errors and illumination changes. Short exposure images are used as soft guidance rather than hard constraints, so it has high robustness to alignment errors and illumination changes.
[0085] The image tone result output by the scheme is better, and the image prior of the generation model can ensure the natural appearance of the output image, thereby reducing the generation of potential artifacts, and can generate natural and realistic tone mapping effects in very high dynamic range scenes.
Claims
1. A high dynamic range imaging system based on diffusion model image prior, characterized in that, It includes a pre-alignment module, a variational autoencoder, a decomposition and fusion control module, a fidelity control module, and a variational autodecoder. The pre-alignment module is used to align a short-exposure image to a long-exposure image and generate a pre-alignment result. The variational autoencoder is used to encode long exposure images to obtain encoding results; The decomposition and fusion control module, based on the pre-alignment and encoding results, controls the fusion process of the diffusion model to perform information fusion and tone mapping on the input exposure image, thereby obtaining the fusion result. The decomposition and fusion control module includes a short-exposure feature extraction unit, a long-exposure feature extraction unit, and a fusion unit. The short-exposure feature extraction unit is used to extract image features of short-exposure images. The long exposure feature extraction unit is used to extract image features of long exposure images from the encoding results; The fusion unit is based on a diffusion model and uses prior knowledge to fuse and tone map the image features of short-exposure and long-exposure images; The fidelity control module, based on long-exposure images and pre-alignment results, is used to control the fidelity of the fusion result. The variational autodecoder is used to decode the fusion result to obtain the reconstructed image.
2. The ultra-high dynamic range imaging system based on diffusion model image prior as described in claim 1, characterized in that, The pre-alignment module employs an optical flow estimation algorithm to align short-exposure images to long-exposure images and mask occluded areas, generating pre-alignment results.
3. The ultra-high dynamic range imaging system based on diffusion model image prior as described in claim 1, characterized in that, The short-explosion feature extraction unit includes multiple cascaded short-explosion feature extractors; The long-explosion feature extraction unit includes multiple cascaded long-explosion feature extractors.
4. The ultra-high dynamic range imaging system based on diffusion model image prior as described in claim 1, characterized in that, The fusion unit employs a cross-attention mechanism and a multi-scale cross-attention mechanism to fuse image features from short-exposure and long-exposure images.
5. A method for ultra-high dynamic range imaging based on diffusion model image priors, characterized in that, Includes the following steps: S1. The optical flow estimation algorithm is used to transform and align the short exposure image onto the long exposure image, and the occluded area is masked to obtain the pre-alignment result; S2. Information decomposition is performed on short-exposure and long-exposure images, and the pre-alignment results are fused and tone mapped by controlling the fusion process of the diffusion model to obtain the fusion result; S3. Based on the long exposure image and the pre-alignment result, the fidelity of the fusion result is controlled, and then the reconstructed image is obtained through decoding; Step S2 includes the following steps: S21. Encode the long exposure image to obtain the encoding result; S22. Extract short-exposure features and long-exposure features from short-exposure images and encoding results, respectively. S23. Use the diffusion model to fuse the short-explosion features and long-explosion features, complete the fusion of the pre-aligned results and tone mapping, and obtain the fusion result.
6. The ultra-high dynamic range imaging method based on diffusion model image prior as described in claim 5, characterized in that, Step S1 includes the following steps: S11. The long exposure image and the short exposure image are processed by the optical flow estimation algorithm respectively to obtain the corresponding first optical flow and second optical flow; S12. Perform a consistency check on the first optical flow and the second optical flow to obtain the occlusion mask; S13. Align the short exposure image and the first optical flow, and use an occlusion mask to cover the occluded area to generate a pre-alignment result.
7. The ultra-high dynamic range imaging method based on diffusion model image prior as described in claim 5, characterized in that, The specific process of step S22 is as follows: For short-exposure images, the luminance component of the short-exposure image is removed, and the structural and color information is extracted. The same network structure as ControlNet is used to extract the short-exposure features of the short-exposure image through several layers of convolution. For long-exposure images, long-exposure features are extracted from the latent variables of the encoding results.
8. The ultra-high dynamic range imaging method based on diffusion model image prior as described in claim 5, characterized in that, Specifically, step S23 involves a fusion process that combines cross-attention mechanism and multi-scale cross-attention mechanism.
Citation Information
Patent Citations
Image multi-frame fusion method and device, electronic equipment and storage medium
CN115578273A
High dynamic range imaging method based on multi-scale progressive reconstruction network
CN118799230A