Unsupervised high dynamic range imaging method based on diffusion model
By training the network under unsupervised conditions using brightness and structure constraints from multi-exposure low dynamic range images, and iteratively completing the motion region using a diffusion model, the problem of information loss in dynamic scenes is solved, and high-quality high dynamic range image reconstruction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-29
AI Technical Summary
Existing high dynamic range imaging methods struggle to effectively recover areas of missing information caused by object movement, occlusion changes, or camera shake in dynamic scenes. Furthermore, deep learning-based reconstruction methods are prone to producing blurry results and are highly dependent on labeled data.
By training a high dynamic range reconstruction network under unsupervised conditions using the brightness and structural consistency constraints of multi-exposure low dynamic range images, preliminary reconstruction results are generated. Based on exposure information and reconstruction differences, motion regions are detected, and a diffusion model is introduced for iterative completion to generate high dynamic range images with complete structure and reasonable brightness.
It achieves high-quality, high-fidelity high dynamic range image reconstruction in dynamic scenes, significantly restores motion region information, and improves image brightness and structural integrity.
Smart Images

Figure CN122115233A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of unsupervised high dynamic range imaging technology, and more particularly to an unsupervised high dynamic range imaging method based on a diffusion model. Background Technology
[0002] High dynamic range (HDR) imaging technology reconstructs a wide brightness range of a real scene by fusing multiple low dynamic range images with different exposures. This addresses the overexposure, underexposure, and detail loss issues inherent in single-exposure imaging in high-contrast scenes and has been widely applied in fields such as photography, autonomous driving, and security monitoring. In static scenes, images with different exposures typically exhibit good complementarity, and high-quality reconstruction results can be obtained through alignment and fusion. However, in dynamic scenes, due to object movement, occlusion changes, or camera shake, spatial misalignment and information loss often coexist between images with different exposures. Especially in areas where strong motion overlaps with overexposure or underexposure, some scene content may be unobservable in all input images, forming irreversible information loss regions that no longer satisfy the traditional assumption of multi-exposure complementarity. Existing HDR imaging methods often rely on strategies such as image alignment, motion compensation, or pixel suppression, but these methods essentially assume that at least one image contains valid information. When all exposed images lack true content, only conservative suppression or simple interpolation can be performed, making it difficult to recover the true structure. While deep learning-based reconstruction methods improve overall performance, they often employ regression modeling, which can lead to blurred results in areas with missing information and are heavily reliant on labeled data. Diffusion models, through a progressive denoising generation process, learn the prior distribution of the image. They can proactively generate structurally sound and semantically consistent content based on context, even when input information is insufficient or missing, offering a new approach to addressing the problem of unobservable information in dynamic scenes. However, current technologies still lack a comprehensive solution for effectively integrating diffusion models into high dynamic range (HDR) imaging processes. How to combine imaging constraints with the generative capabilities of diffusion models to effectively fill in missing information regions remains a key challenge in HDR image reconstruction technology.
[0003] Therefore, it is necessary to improve one or more of the problems existing in the above-mentioned related technical solutions.
[0004] It should be noted that this section is intended to provide background or context for the technical solutions of this disclosure as set forth in the claims. The description herein does not constitute an admission that it is prior art simply because it is included in this section. Summary of the Invention
[0005] The purpose of this disclosure is to provide an unsupervised high dynamic range imaging method based on a diffusion model, thereby overcoming at least to some extent one or more problems caused by the limitations and defects of related technologies.
[0006] According to embodiments of this disclosure, an unsupervised high dynamic range imaging method based on a diffusion model is provided, comprising: Step S1: Under unsupervised conditions, the high dynamic range reconstruction network is trained by utilizing the brightness consistency and structural consistency constraints between multi-exposure low dynamic range images to obtain preliminary high dynamic range reconstruction results. Step S2: Enhance the details of the static area based on the exposure information of the reference frame, and generate a motion region mask based on the reconstruction difference detection of motion or unreliable information regions. Step S3: Using the enhanced high dynamic range image, motion region mask, and text information as conditions, a pre-trained diffusion model is introduced. While maintaining the integrity of the known static regions, the regions marked by the motion region mask are iteratively completed, and finally, a high dynamic range image with complete structure and reasonable brightness is obtained by fusion.
[0007] Furthermore, step S1 specifically includes: The input multi-exposure low dynamic range image sequence is mapped to the linear domain through gamma correction to obtain the corresponding high dynamic range image representation; The intermediate exposure image is selected as the reference frame, and the remaining exposure images are aligned with it using optical flow. A pseudo-labeled high dynamic range image is then generated through a weighted fusion method. The fusion weights are determined based on the exposure confidence function of the reference frame pixel brightness values. A high dynamic range reconstruction network is constructed. Each low dynamic range image is stitched together with its corresponding high dynamic range representation and then input into the network to obtain a preliminary high dynamic range image. The network is trained under joint constraints using content loss and structural loss. Content loss is supervised based on the pseudo-labeled high dynamic range image within a reliably aligned region, while structural loss is supervised based on the reference frame within a reasonably exposed region, in order to optimize the network parameters and generate an optimized preliminary high dynamic range image.
[0008] Furthermore, the multi-exposure low dynamic range image sequence is as follows:
[0009] The expression for a high dynamic range image is:
[0010] in, For gamma parameters, Exposure time; The initial high dynamic range image is as follows:
[0011] in, For input features, Represents the learnable parameters of the network; Content loss is:
[0012] in, For tone mapping functions, For element-wise multiplication, This is a pseudo-label image. It is a binary mask; The structural loss is:
[0013] in, The triangular fusion coefficient; The total loss function is:
[0014] in, This is a pseudo-label image.
[0015] Furthermore, step S2 specifically includes: The reference frame is converted into a grayscale image, an initial exposure mask is generated based on the pixel brightness and a preset threshold, and a continuous mask is obtained through morphological processing. By using an exposure mask, the preliminary high dynamic range image and the pseudo-label high dynamic range image are fused together, so as to use the pseudo-label information to supplement and enhance the overexposed or underexposed details in the static area, and obtain the enhanced high dynamic range image. The pixel-level difference between the tone-mapped pseudo-labeled high dynamic range image and the enhanced high dynamic range image is calculated, and pixels with a difference value greater than a preset threshold are marked as motion or unreliable reconstruction regions, generating a motion region mask.
[0016] Furthermore, the initial exposure mask is:
[0017] in, Indicates pixel position, Represents a grayscale image. Indicates the preset threshold;
[0018] The enhanced high dynamic range image is as follows:
[0019] The motion region mask is:
[0020] in, It is a fixed threshold.
[0021] Furthermore, step S3 specifically includes: The enhanced high dynamic range image is mapped to the standard dynamic range domain through inverse gamma transform to obtain the standard dynamic range conditional image, which is then input into the pre-trained diffusion model along with the motion region mask and text condition. In the reverse denoising process of the diffusion model, based on the Markov chain property, the known and unknown regions marked by the motion region mask are updated differentially and iteratively; the pixel values of the known regions are directly predicted from the initial input, while the pixel values of the unknown regions are gradually generated by sampling through the noise prediction network of the diffusion model. After iteration, the output of the diffusion model is mapped back to the high dynamic range domain to obtain the content of the motion region. By combining motion region masks, the motion region content is fused with the enhanced high dynamic range image to generate the final high dynamic range image.
[0022] Furthermore, in the t-th iteration, the update formulas for the known and unknown regions are:
[0023]
[0024]
[0025] in, Given the region, For initial noise-free image estimation, For the first The signal retention coefficient of the step, It is the first Gaussian noise. For the first Noisy variables of the step, For the first The noise figure of the step. The diffusion coefficient is... The predicted noise output by the noise prediction network. For text conditions, It is the second Gaussian noise. For the first After one iteration -1 update result; The content of the movement area is as follows:
[0026] in, For decoders; The final high dynamic range image is as follows:
[0027] in, This refers to the content of the motion area.
[0028] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: In the embodiments of this disclosure, the unsupervised high dynamic range (HMR) imaging method based on the diffusion model described above achieves the following: First, under unsupervised conditions, the HMR reconstruction network is guided to learn by the brightness consistency and structural constraints among multiple exposure low dynamic range (LMR) images, resulting in preliminary HMR reconstruction results with reasonable brightness and reliable structure. Subsequently, static and moving regions are distinguished based on exposure information and reconstruction differences. Detail enhancement is performed on static regions, and moving or unreliable information regions are accurately marked. Finally, a diffusion model is introduced to perform high-quality content completion on unobservable regions caused by motion, occlusion, or exposure issues, while preserving known static regions. This achieves structural integrity reconstruction and visual quality improvement of HMR images in dynamic scenes. Second, the HMR image obtained by this method not only has reasonable brightness and complete structure, but also significantly recovers information from moving regions. By fully leveraging the prior capabilities of the generative model, high-quality, high-fidelity HMR image reconstruction in dynamic scenes is achieved, demonstrating broad application value and promotion potential. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0030] Figure 1 A flowchart illustrating the steps of an unsupervised high dynamic range imaging method based on a diffusion model in an exemplary embodiment of this disclosure is shown. Figure 2 A flowchart illustrating an unsupervised high dynamic range imaging method based on a diffusion model in an exemplary embodiment of this disclosure is shown. Figure 3 This illustration shows multiple frames of low dynamic range images with different exposures in an exemplary embodiment of this disclosure; Figure 4 The final high dynamic range image in an exemplary embodiment of this disclosure is shown. Detailed Implementation
[0031] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0032] Furthermore, the accompanying drawings are merely illustrative diagrams of embodiments of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0033] This example implementation provides an unsupervised high dynamic range imaging method based on a diffusion model. (Reference) Figure 1 As shown, this unsupervised high dynamic range imaging method based on a diffusion model may include: Step S1: Under unsupervised conditions, the high dynamic range reconstruction network is trained by utilizing the brightness consistency and structural consistency constraints between multi-exposure low dynamic range images to obtain preliminary high dynamic range reconstruction results. Step S2: Enhance the details of the static area based on the exposure information of the reference frame, and generate a motion region mask based on the reconstruction difference detection of motion or unreliable information regions. Step S3: Using the enhanced high dynamic range image, motion region mask, and text information as conditions, a pre-trained diffusion model is introduced. While maintaining the integrity of the known static regions, the regions marked by the motion region mask are iteratively completed, and finally, a high dynamic range image with complete structure and reasonable brightness is obtained by fusion.
[0034] The unsupervised high dynamic range (HMR) imaging method based on the diffusion model described above achieves two main objectives. First, under unsupervised conditions, the HMR reconstruction network is guided to learn by considering the brightness consistency and structural constraints among multiple exposure low dynamic range (LMR) images, resulting in preliminary HMR reconstructions with reasonable brightness and reliable structure. Subsequently, static and moving regions are distinguished based on exposure information and reconstruction differences. Detail enhancement is applied to static regions, while moving or unreliable regions are precisely labeled. Finally, a diffusion model is introduced to complete unobservable regions caused by motion, occlusion, or exposure issues with high-quality content, ensuring the preservation of known static regions. This achieves structural integrity reconstruction and visual quality improvement of HMR images in dynamic scenes. Second, the final HMR image obtained through this method not only has reasonable brightness and complete structure but also significantly recovers information from moving regions. By fully leveraging the prior capabilities of the generative model, high-quality, high-fidelity HMR image reconstruction in dynamic scenes is achieved, demonstrating broad application value and potential for widespread adoption.
[0035] Below, we will refer to Figures 1 to 4 The steps of the above-described unsupervised high dynamic range imaging method based on the diffusion model in this example embodiment will be described in more detail.
[0036] like Figure 2 The diagram shows a flowchart of an unsupervised high dynamic range imaging method based on a diffusion model.
[0037] In step S1, under unsupervised conditions, the high dynamic range reconstruction network is trained using the brightness consistency and structural consistency constraints between multi-exposure low dynamic range images to obtain preliminary high dynamic range reconstruction results.
[0038] Specifically, preliminary reconstruction of high dynamic range images based on unsupervised learning. In this application, step S1 is used to construct high dynamic range pseudo-labels by aligning and weighting low dynamic range images under unsupervised conditions, and to jointly constrain the high dynamic range reconstruction network from the two levels of content consistency and structural consistency, so as to obtain preliminary high dynamic range reconstruction results with reasonable brightness distribution and stable structural information, providing basic input for subsequent region segmentation and generative completion.
[0039] The specific process is as follows: Sub-step 1: Preprocessing of low dynamic range images Given a sequence of multi-exposure low dynamic range images of the same scene Mapping it to the linear domain using gamma correction yields the corresponding high dynamic range image representation:
[0040] in, This represents a high dynamic range image under the corresponding exposure conditions. For gamma parameters, The exposure time is used. Through the above mapping operation, the original low dynamic range image is uniformly transformed into the linear high dynamic range domain to serve as the input for the subsequent high dynamic range reconstruction network.
[0041] Sub-step 2: Generation of multi-exposure pseudo-high dynamic range labels based on optical flow alignment Select intermediate exposure image Using the reference frame, the image is calculated using the optical flow method. and Compared to Pixel-level alignment transformation, and obtain the pixel-level alignment transformation through reverse mapping. Aligned images Based on this, a pseudo-labeled high dynamic range image is generated through weighted fusion, and its expression is:
[0042] in, The fusion weights for the corresponding images are defined as follows:
[0043] The aforementioned weighting coefficients are used to characterize the reliability and contribution of images with different exposures at the current pixel location. Among them, Based on normal exposure image An exposure confidence function, constructed from pixel brightness values, is used to measure the reliability of information about a pixel within its corresponding brightness range. This exposure confidence function employs a piecewise continuous triangular function form, allowing the weights to adjust smoothly with changes in pixel brightness, thereby avoiding abrupt changes or artifacts in the fusion result.
[0044] Sub-step 3: Unsupervised high dynamic range network training based on reliable region constraints During training, the multi-exposure high dynamic range representation obtained in sub-step 1 is utilized. Each low dynamic range image Its corresponding high dynamic range representation By concatenating along the channel dimension, a 6-channel input feature is constructed:
[0045] Then, the high dynamic range reconstruction network Accept the stitched multi-exposure input and generate a preliminary high dynamic range image:
[0046] in, For the initially generated high dynamic range image, This represents the learnable parameters of the network.
[0047] To obtain preliminary high dynamic range images To ensure the rationality of its brightness distribution and structural information, this application introduces content loss and structural loss to jointly constrain the network output during the training process.
[0048] First, introduce content loss. This is used to guide the network to learn reliable high dynamic range information from non-reference frames. Due to the pseudo-label images... During the generation process, optical flow alignment errors or large-scale motion may affect the reconstruction, and inaccurate information may still be present in local regions. Directly applying supervision to the entire image can easily introduce artifacts into the reconstruction result. Therefore, this application only performs supervision in regions with reliable alignment, using binary masks. The definition for masking out poorly aligned pixels is as follows:
[0049] in, , is the tone mapping function. The compression factor is 1. This is an element-wise multiplication operation.
[0050] Secondly, structural loss is introduced. Used to constrain the network to maintain the reference frame The integrity of the structural information in the middle. This loss is achieved through the fusion coefficient. Emphasis is placed on properly exposed areas in the reference frame, while mitigating the negative impact of underexposed or overexposed areas, thereby ensuring that the network output is structurally consistent with the reference frame.
[0051] in, The triangular fusion coefficients are used to highlight the reasonably bright areas of the reference frame, allowing the network to give these areas higher weights in its structure.
[0052] Combining the two types of losses mentioned above, the training objective of the network is defined as:
[0053] Through the above design, the network can accurately restore details in areas of reasonable brightness while maintaining overall structural consistency, thereby generating preliminary high dynamic range images with stable brightness and reliable structure. .
[0054] In step S2, static areas are enhanced with detail based on the exposure information of the reference frame, and motion or unreliable information areas are detected based on reconstruction differences to generate motion area masks.
[0055] Specifically, static region enhancement and motion region marking In this application, step S2 is used to perform regional adaptive processing on the preliminary high dynamic range image generated in step S1. By utilizing the exposure information of the reference frame to enhance the details of the static region, and combining motion region detection to identify motion or unreliable reconstructed regions, the motion region is detected while enhancing the details of the static region, providing reliable conditional input for the diffusion generation model in step S3.
[0056] The specific process is as follows: Sub-step 1: Static area detail enhancement: Reference frame Convert to grayscale And based on the threshold Generate initial exposure mask :
[0057] in Indicates the pixel location. To eliminate isolated noise and smooth region boundaries, ... Morphological processing (erosion and dilation) is performed to obtain a continuous mask. This mask is then used to process the initial high dynamic range image. High dynamic range images with pseudo-labels The results of the fusion are obtained as static region enhancement:
[0058] in This is an element-wise multiplication operation. This process can use pseudo-labels to supplement overexposure or underexposure information in static areas while preserving details from the initial reconstruction.
[0059] Sub-step 2: Motion region detection: In dynamic scenes, due to object movement or camera shake, low dynamic range (LR) images with different exposures may exhibit spatial misalignment or information loss at the same location, leading to unreliable reconstruction results for some regions in the initial high dynamic range (HDR) image. This paper utilizes the enhanced HDR image obtained from static region detail enhancement. High dynamic range images with pseudo-labels The difference generates a motion region mask. :
[0060] in For a fixed threshold, This is the tone mapping function. Pixels with a value of 1 in the mask indicate that the region is a moving region or an unreliable reconstruction region. After generating this moving region mask, it can be used as a conditional input for the diffusion model in step S3 to perform high-quality content completion on the marked region, achieving accurate restoration of the moving region.
[0061] In step S3, the enhanced high dynamic range image, motion region mask, and text information are used as conditions to introduce a pre-trained diffusion model. While maintaining the integrity of the known static region, the region marked by the motion region mask is iteratively completed, and finally, a high dynamic range image with complete structure and reasonable brightness is obtained by fusion.
[0062] Specifically, motion region completion based on diffusion model In this application, step S3 is used to generate and complete high-quality content for the motion region marked in step S2. By using a pre-trained diffusion model to infer the information missing regions caused by motion or underexposure while maintaining the integrity of the known regions, a final high-quality high dynamic range image is generated.
[0063] The specific process is as follows: Sub-step 1: Input conditions for the diffusion model: Before entering the diffusion model, the enhanced high dynamic range image generated in step S2 is first processed. A standard dynamic range conditional image is obtained by mapping the image to the standard dynamic range domain using inverse gamma transform. Subsequently, this standard domain image and motion region mask are... And text-conditional input diffusion model, in which the integrity of the known region is determined directly from... predict This was achieved by utilizing the characteristics of a Markov process with additive Gaussian noise. For the unknown region, the [process name] is used in the [process name]. Obtained in step iteration Sampling is performed. The encoder and decoder of the diffusion model are denoted as follows: and The noise scheduling parameters are Optional text conditions are .
[0064] Sub-step 2: Iterative completion of the motion region: In the In the iterative steps, the update formulas for the known and unknown regions are:
[0065]
[0066]
[0067] in It is Gaussian noise. The diffusion coefficient is... This is element-wise multiplication. The iteration starts from... arrive The motion area content is generated gradually.
[0068] Sub-step 3: High Dynamic Range Image Fusion and Output: go through After one iteration, the reconstruction result can be obtained. And map it back to the high dynamic range domain:
[0069] Finally, combine the motion region mask. Combine the generated motion region content with the enhanced high dynamic range image The images are fused together to obtain the final high dynamic range image.
[0070] Through the above steps, missing information in the moving regions is fully restored with high quality, while ensuring that the details and structure of the static regions are not destroyed, thereby generating a final high dynamic range image with reasonable brightness and complete structure. .
[0071] In a specific embodiment, such as Figure 3 The image shown is a low dynamic range image with multiple frames at different exposures; Figure 4 This is a schematic diagram of the high dynamic range reconstruction results obtained using the method described in this application. It can be seen that this application maintains good brightness consistency and detail in static regions, while achieving reasonable information completion in moving regions.
[0072] In summary, this application proposes a high dynamic range (HMR) image imaging method based on unsupervised learning and a diffusion model, aiming to address the problem of image information loss caused by object movement, overlapping of occluded regions and overexposed or underexposed regions in dynamic scenes. Existing techniques often suffer from detail loss, blurring, or structural inconsistencies in such scenarios. To overcome these shortcomings, this application designs a unified HMR image imaging framework. First, it utilizes the inherent brightness and structural relationships of multi-exposure low dynamic range (LMR) images to generate preliminary HMR images under unlabeled conditions, enhancing details in static regions while identifying moving or unreliable regions to form a reliable motion region mask. Based on this mask, this application further introduces a diffusion model to complete the content of moving regions, generating high-quality information with reasonable structure and consistent semantics even if these regions are unobservable in all input images. Through this method, the final HMR image not only has reasonable brightness and complete structure, but also significantly recovers information from moving regions. Compared with traditional methods, this application fully leverages the prior capabilities of the generative model to achieve high-quality, high-fidelity HMR image reconstruction in dynamic scenes, possessing broad application value and promotion potential.
[0073] It should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise" in the above description indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of this disclosure and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this disclosure.
[0074] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0075] In the embodiments of this disclosure, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.
[0076] In embodiments of this disclosure, unless otherwise expressly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature being directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature being directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0077] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0078] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. An unsupervised high dynamic range imaging method based on a diffusion model, characterized in that, include: Step S1: Under unsupervised conditions, the high dynamic range reconstruction network is trained by utilizing the brightness consistency and structural consistency constraints between multi-exposure low dynamic range images to obtain preliminary high dynamic range reconstruction results. Step S2: Enhance the details of the static area based on the exposure information of the reference frame, and generate a motion region mask based on the reconstruction difference detection of motion or unreliable information regions. Step S3: Using the enhanced high dynamic range image, motion region mask, and text information as conditions, a pre-trained diffusion model is introduced. While maintaining the integrity of the known static regions, the regions marked by the motion region mask are iteratively completed, and finally, a high dynamic range image with complete structure and reasonable brightness is obtained by fusion.
2. The unsupervised high dynamic range imaging method based on a diffusion model according to claim 1, characterized in that, Step S1 specifically includes: The input multi-exposure low dynamic range image sequence is mapped to the linear domain through gamma correction to obtain the corresponding high dynamic range image representation; The intermediate exposure image is selected as the reference frame, and the remaining exposure images are aligned with it using optical flow. A pseudo-labeled high dynamic range image is then generated through a weighted fusion method. The fusion weights are determined based on the exposure confidence function of the reference frame pixel brightness values. A high dynamic range reconstruction network is constructed. Each low dynamic range image is stitched together with its corresponding high dynamic range representation and then input into the network to obtain a preliminary high dynamic range image. The network is trained under joint constraints using content loss and structural loss. Content loss is supervised based on the pseudo-labeled high dynamic range image within a reliably aligned region, while structural loss is supervised based on the reference frame within a reasonably exposed region, in order to optimize the network parameters and generate an optimized preliminary high dynamic range image.
3. The unsupervised high dynamic range imaging method based on the diffusion model according to claim 2, characterized in that, The multi-exposure low dynamic range image sequence is as follows: The expression for a high dynamic range image is: in, For gamma parameters, Exposure time; The initial high dynamic range image is as follows: in, For input features, Represents the learnable parameters of the network; Content loss is: in, For tone mapping functions, For element-wise multiplication, This is a pseudo-label image. It is a binary mask; The structural loss is: in, The triangular fusion coefficient; The total loss function is: in, This is a pseudo-label image.
4. The unsupervised high dynamic range imaging method based on the diffusion model according to claim 3, characterized in that, Step S2 specifically includes: The reference frame is converted into a grayscale image, an initial exposure mask is generated based on the pixel brightness and a preset threshold, and a continuous mask is obtained through morphological processing. By using an exposure mask, the preliminary high dynamic range image and the pseudo-label high dynamic range image are fused together, so as to use the pseudo-label information to supplement and enhance the overexposed or underexposed details in the static area, and obtain the enhanced high dynamic range image. The pixel-level difference between the tone-mapped pseudo-labeled high dynamic range image and the enhanced high dynamic range image is calculated, and pixels with a difference value greater than a preset threshold are marked as motion or unreliable reconstruction regions, generating a motion region mask.
5. The unsupervised high dynamic range imaging method based on the diffusion model according to claim 4, characterized in that, The initial exposure mask is: in, Indicates pixel position, Represents a grayscale image. Indicates the preset threshold; The enhanced high dynamic range image is as follows: The motion region mask is: in, It is a fixed threshold.
6. The unsupervised high dynamic range imaging method based on the diffusion model according to claim 5, characterized in that, Step S3 specifically includes: The enhanced high dynamic range image is mapped to the standard dynamic range domain through inverse gamma transform to obtain the standard dynamic range conditional image, which is then input into the pre-trained diffusion model along with the motion region mask and text condition. In the reverse denoising process of the diffusion model, based on the Markov chain property, the known and unknown regions marked by the motion region mask are updated differentially and iteratively; the pixel values of the known regions are directly predicted from the initial input, while the pixel values of the unknown regions are gradually generated by sampling through the noise prediction network of the diffusion model. After iteration, the output of the diffusion model is mapped back to the high dynamic range domain to obtain the content of the motion region. By combining motion region masks, the motion region content is fused with the enhanced high dynamic range image to generate the final high dynamic range image.
7. The unsupervised high dynamic range imaging method based on the diffusion model according to claim 6, characterized in that, In the t-th iteration, the update formulas for the known and unknown regions are: in, Given the region, For initial noise-free image estimation, For the first The signal retention coefficient of the step, It is the first Gaussian noise. For the first Noisy variables of the step, For the first The noise figure of the step. The diffusion coefficient is... The predicted noise output by the noise prediction network. For text conditions, It is the second Gaussian noise. For the first After one iteration -1 update result; The content of the movement area is as follows: in, For decoders; The final high dynamic range image is as follows: in, This refers to the content of the motion area.