Perceptual artifact detection and elimination method for high dynamic range reconstruction
By constructing HDR-AD dataset and HDR-ADet detector, the artifact problem in HDR images is explicitly solved, image quality is improved, and new image evaluation indicators are provided, solving the shortcomings of artifact detection and evaluation in the prior art.
Patent Information
- Application Number
- CN202510176766.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The prior art is difficult to explicitly solve the artifact problem in HDR image reconstruction, and there is a lack of evaluation indicators that can reflect the image quality of the artifact.
By detecting and positioning artifact areas in HDR images, an HDR-AD data set and a VisionTransformers-based artifact detector HDR-ADet are constructed, and the artifact mask is output to improve the HDR model and improve image quality.
Successfully detecting and eliminating artifact areas in HDR images to improve HDR imaging quality, providing a new image evaluation index - artifact score, reflecting human perception of image quality.
Smart Images

Figure CN120031852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of HDR reconstruction, and in particular to the task of reconstructing HDR images from multi-exposure LDR images. Background Art
[0002] Due to the limited sensitivity of camera sensors, images are usually captured with a low dynamic range. These low dynamic range (LDR) images lose details in very bright or very dark areas. In contrast, high dynamic range (HDR) images capture a wider spectrum of light, retaining rich details and vivid colors. This capability makes HDR imaging valuable in fields such as medical imaging, remote sensing, and photography, driving the growing demand for high-quality HDR images. Some hardware devices can directly capture HDR images, but their high prices prevent them from being widely used. A more feasible alternative is to merge LDR images of different exposures to create HDR images, which can extend the dynamic range and recover more details. However, camera or object motion often causes misalignment between LDR images, resulting in ghosting artifacts in the merged HDR image. In addition, saturation in extremely bright or dark areas introduces distortion artifacts, affecting image quality and reducing visual appeal.
[0003] Among existing methods, those based on convolutional neural networks (CNNs) focus on fusing LDR inputs after optical flow alignment; end-to-end models introduce various attention mechanisms to suppress artifact effects in the fusion stage.
[0004] Recently, Transformers have been shown to be effective in this task because they can capture global context information. A specific approach to using Transformers for HDR reconstruction is to first use a spatial attention mechanism to extract coarse features of LDR images with different exposures, and then use multiple context-aware transformer blocks (CTBs) to perform HDR reconstruction.
[0005] All of these existing methods attempt to solve the artifact problem in HDR imaging by improving the quality of alignment and fusion. The specific approach is to extract information from high-exposure and low-exposure images to fill in the missing parts of the intermediate-exposure image to implicitly solve the artifact problem. However, there is currently no approach to explicitly solve this problem by detecting and locating artifact areas in reconstructed HDR images. At the same time, because these artifact areas are usually small, they will be ignored in the calculation of the average value during the calculation of quantitative indicators, but they have a great impact on image quality. There is currently no HDR image quality evaluation indicator for artifacts.
[0006] In general, the disadvantages of the prior art are:
[0007] 1. There is a lack of HDR reconstruction methods that can explicitly solve the artifact problem, and most of them solve it implicitly by extracting information.
[0008] 2. There is a lack of evaluation indicators that can reflect the impact of artifacts on image quality and visual experience. Summary of the invention
[0009] In view of the problems of the prior art, the present invention aims to explicitly solve the artifact problem in the multi-exposure HDR reconstruction task by detecting and locating the artifact area in the reconstructed HDR image. Because the artifact area has a great influence on human visual perception, the artifact mask obtained by detection can be introduced into the HDR reconstruction process as human feedback, thereby improving the visual effect of the reconstructed image. This artifact mask can also be used as an image evaluation indicator to reflect the visual quality of the HDR image through the size of the artifact area.
[0010] The present invention treats HDR artifacts as unique and detectable entities, explicitly addressing them rather than relying solely on alignment and fusion methods. The goal of the present invention is to identify artifact areas in HDR outputs and feed this information back to the HDR model for improvement. To this end, the present invention first collects a diverse set of LDR images in various scenarios and reconstructs HDR outputs using multiple HDR models. Then, the present invention analyzes the artifacts present and creates detailed pixel-by-pixel annotations, ultimately forming the HDR-AD (HDR Artifacts Dataset) dataset. With HDR-AD, the present invention introduces an artifact detector HDR-ADet (HDR Artifacts Detector) based on Vision Transformers (ViTs), which takes an HDR image as input and outputs a mask that highlights the artifact area. Finally, the HDR model is improved by feeding back the detected artifacts during the training process to improve quality.
[0011] Technical Solution
[0012] A method for perceptual artifact detection and elimination for high dynamic range reconstruction includes the following steps:
[0013] Step 1: Create an HDR-AD dataset
[0014] In order to achieve the goal of detecting artifact areas in HDR images, a high-quality HDR-AD dataset is first constructed: LDR images acquired under different exposures, HDR images inferred using the existing HDR reconstruction model, and masks obtained by manual pixel-by-pixel annotation. The HDR images and masks are used to train the detection model.
[0015] Step 2: Construct and train the HDR-ADet model
[0016] We build an artifact detector HDR-ADet based on ViT, and use the HDR-AD dataset to train this detector. HDR-ADet, which has been trained with a large amount of data, can accurately locate artifact areas in images.
[0017] After training, an HDR image is input and HDR-ADet will output a binary mask representing the artifact area.
[0018] Step 3: Use the artifact mask obtained by HDR-ADet to fine-tune the HDR reconstruction model to improve the quality of the HDR image reconstructed by the model.
[0019] Because artifacts are highly correlated with human visual perception, the artifact mask obtained by HDR-ADet can be considered as human feedback on the quality of HDR image reconstruction. By integrating HDR-ADet into the loss function calculation process of the HDR reconstruction framework, it is used to improve the quality of HDR images reconstructed by the model.
[0020] The artifact mask output by HDR-ADet can also be used as an evaluation indicator for HDR images: the existing HDR image evaluation indicators cannot reflect the impact of artifacts on image quality. The present invention proposes a novel image evaluation indicator. The artifact score (AS) obtained by the artifact mask can accurately reflect human perception of image quality.
[0021] Beneficial Effects
[0022] Through the above technical solution, the present invention can successfully detect artifact areas in HDR images obtained by multiple HDR reconstruction models, and use these areas as feedback to fine-tune the HDR reconstruction model, thereby improving HDR imaging quality, not only achieving breakthroughs in quantitative indicators, but also being closer to human visual perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1A schematic diagram of the overall processing flow of the method of the present invention;
[0024] Figure 2 Detailed flow chart of the method of the present invention;
[0025] Figure 3 Statistical information of a data set obtained by an embodiment of the present invention;
[0026] Figure 4 Detailed analysis of the qualitative and quantitative analysis of the HDR reconstruction model and artifact ratio in the embodiments of the present invention;
[0027] Figure 5 Specific network architecture diagram of the artifact detector HDR-ADet of the present invention;
[0028] Figure 6 Visual comparison results of an embodiment of the present invention;
[0029] Figure 7 Schematic diagram of visualization results of image evaluation indicators according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The technical solution provided by this application will be further described below in conjunction with specific embodiments and their accompanying drawings. The advantages and features of the present application will become more clear from the following description.
[0031] A method for detecting and eliminating perceptual artifacts for high dynamic range reconstruction includes the following steps: Figure 1 , Figure 2 )
[0032] Step 1: Create an HDR-AD dataset
[0033] Unlike existing HDR reconstruction models that usually solve artifacts indirectly through fusion and alignment, our work takes a novel approach to deal with the problem by directly detecting and solving artifacts. Current LDR-HDR datasets are limited in scale, lack diversity, and do not specifically target artifacts.
[0034] To address this gap, we first construct a comprehensive HDR artifact dataset based on real-world scenes, which includes the following steps:
[0035] Step 1.1 collect a multi-exposure LDR image set;
[0036] The present invention first uses mobile phones and digital cameras to capture a variety of image content to support HDR across multiple platforms. All images are initially captured in RAW format at the native resolution of each device and then converted to TIFF format for further processing. In order to focus on difficult scenes in the HDR image generation process, the present invention archives the collected LDR image set to ensure that each final set of LDR images contains different exposures and object movement.
[0037] Step 1.2: reconstructing an HDR image using an existing HDR reconstruction model;
[0038] Next, the present invention applies mature and high-performance HDR reconstruction models (such as AHDR, HDR-Transformer and SCTNet) to generate HDR images. The sources of the generated images include LDR images collected using the above process, as well as existing Kal's and Tel's images.
[0039] Step 1.3: Manually label artifact masks.
[0040] The online platform labelbox is used for manual labeling. During the labeling process, images with too high or too low reconstruction quality will be excluded. If multiple annotators cannot find the artifact area in the same HDR image, the reconstruction quality is considered to be too high; if an HDR image is almost entirely an artifact area and no successfully reconstructed area can be found, the reconstruction quality is considered to be too low. Images with significant disagreements between annotators will also be ignored to prevent model bias.
[0041] Specifically, the present invention collected 1,213 LDR image sets, each of which contained low-, medium-, and high-exposure images, and obtained 1,765 HDR images with pixel-by-pixel annotations through a rigorous screening process.
[0042] In addition to the LDR image set collected in step 1.1, the present invention also processes Kal's and Tel's datasets to increase the diversity of the dataset and promote the robust generalization of the detector.
[0043] The statistical information of the data set of the present invention is as follows Figure 3 As shown, specifically:
[0044] like Figure 3 As shown in (a), the dataset covers both indoor and outdoor scenes to cover various lighting conditions.
[0045] like Figure 3 (b) shows the quantitative analysis of brightness and contrast of the dataset of the present invention and Kal's and Tel's datasets.
[0046] like Figure 3 (c) shows examples at different time periods—morning, sunrise / sunset, and night—to capture a wide range of lighting conditions.
[0047] like Figure 3As shown in (d), the LDR image set also includes a variety of motion types that can produce noticeable artifacts, creating challenging cases. The present invention introduces motion of people or objects in the scene to simulate misalignment during bracketing (i.e., foreground motion). In addition, the present invention deliberately moves the camera to produce camera shake (i.e., background motion), and in some image sets, these two motion types are combined (i.e., panoramic motion). This diversity of motion types is critical to capturing a wide range of artifact variations that significantly affect HDR quality.
[0048] Step 2: Construct HDR-ADet network and train it
[0049] The HDR-ADet network includes: a feature backbone extraction network, a feature fusion bottleneck, and a detection head. The backbone network and head architecture perform detection tasks and are combined with a fusion bottleneck to better integrate features across multiple scales. Figure 5 .
[0050] Step 2.1 Construct feature backbone extraction network
[0051] The feature extraction backbone network adopts a hybrid window attention strategy, which combines window attention (Window Self-Attention, WSA) with multiple cross-window blocks (Global Self-Attention, GSA). Specifically, given an input HDR image I, the image is processed by patch embedding and Transformer blocks. The output is the final feature map F with a patch size of 16. This process is formally expressed as follows:
[0052] F=V(I),
[0053] Among them, V(·) is the feature extraction backbone network based on the window ViT.
[0054] Step 2.2 Construct feature fusion bottleneck
[0055] Since artifacts in HDR images vary in shape and size, this paper proposes a multi-scale feature fusion network to extract features of different scales from the output of the feature extraction backbone network. Specifically, the feature map (F) from the backbone network generates multi-scale features through a series of convolution, pooling, and deconvolution layers. The feature map of each scale represents the information of the image at different resolutions, thereby being able to capture a variety of features from large-scale coarse information to local details. The convolution layer is responsible for extracting local features, the pooling layer helps reduce the spatial dimension and enhance the focus on the overall information of the image, and the deconvolution layer restores the detail information through upsampling operations.
[0056] p i =S i(F),
[0057] The value range of i is 1, 2, 3, 4, p i Indicates features of different scales, which are 1 / 32, 1 / 16, 1 / 8, and 1 / 4 feature maps, respectively. i (·) is a series of convolution, pooling, and deconvolution layers.
[0058] However, since artifacts in HDR images often have complex context inconsistencies and sometimes involve subtle details in the image, relying solely on local features may not be enough to capture global information. Therefore, the present invention also introduces a global information extraction mechanism. This mechanism extracts global context information from the feature map, specifically through a global pooling or global feature extraction module. Global information p 4 Capturing the macroscopic structure and background information within the entire image is particularly important for artifact detection, because artifacts may be affected by the global scene and show different forms. The extraction process is as follows:
[0059] f g =Maxpooling(p 4 ),
[0060] Next, the present invention will extract the global information f g Upsample to each scale feature p i The upsampling operation converts the global information into a spatial dimension that matches the feature map of each scale, so that the global information can be effectively combined with the local features. In this way, the global information provides complementary contextual information for each scale feature, thereby helping the network better understand the overall structure of the image and enhance the detection accuracy of artifacts:
[0061] f i =U(f g )+p i ,
[0062] where U(·) represents the upsampling operation.
[0063] Through the design of the above-mentioned multi-scale feature fusion network, the present invention can not only effectively process the differences in artifacts at different scales, but also improve the network's perception of artifacts by introducing global information, making artifact detection in HDR images more accurate and reliable.
[0064] Step 2.3 Construct the detection head
[0065] In order to strike a balance between performance and computational cost, the present invention uses an MLP decoder as the detection head. First, the multi-scale features are upsampled to 1 / 4 of the input HDR image size and concatenated. Then, two MLP layers are used to fuse the features and generate the final prediction. The MLP decoder MLP(·) can be formulated as follows:
[0066]
[0067] in is the predicted artifact mask.
[0068] Through this design, the MLP decoder can not only efficiently fuse features from different scales and levels, but also generate accurate artifact predictions at a low computational cost. Compared with the traditional convolutional neural network decoder, the MLP decoder has the advantages of simple structure and low computational resource requirements, which makes it very suitable for fast processing in practical applications while still maintaining high detection accuracy.
[0069] Step 2.4 Train the HDR-ADet network
[0070] The artifact detector HDR-ADet is trained in a supervised manner using the HDR-AD dataset, freezing the backbone network and only updating the parameters of the feature fusion bottleneck and the detection head. The number of training rounds is set to 200.
[0071] Step 3: Integrate HDR-ADet into the HDR reconstruction framework to improve the visual quality of the reconstructed HDR image.
[0072] Human visual perception is very sensitive to visual artifacts, so the artifact masks output by HDR-ADet can provide valuable guidance for visual enhancement models (i.e., HDR models) to align them with human perceptual responses. Intuitively, integrating HDR-ADet into the HDR reconstruction framework can improve the visual quality of reconstructed HDR images and better reflect human perception of light and color, such as Figure 4 As shown, (a) is the HDR image reconstructed by different models from the same LDR image, and (b) is the average artifact ratio of HDR images of various models on multiple datasets.
[0073] This paper proposes a simple, effective and unified method to optimize the loss function calculation by integrating HDR-ADet into the HDR reconstruction framework: the predicted HDR image Artifact mask It is used to calculate the L1 loss penalty for the corresponding area during the loss calculation process. The corresponding loss function formula is as follows:
[0074]
[0075] Among them, H is the true value, L r is the original loss of the HDR reconstruction framework, is a hyperparameter, L h It is the feedback fine-tuning loss. By increasing the loss penalty in this way, the HDR model can pay more attention to the areas with poor quality, that is, the artifact areas, during the fine-tuning process, thereby improving the image quality in a targeted manner.
[0076] Compared with the prior art, the present invention has the following innovations:
[0077] (1) HDR Artifact Dataset (HDR-AD): We created the first dataset for HDR artifact detection, including a set of multi-exposure LDR images, HDR images with artifacts, and pixel-by-pixel artifact masks. Based on this dataset, we comprehensively analyze how different scenes and models lead to the formation of artifacts, providing a valuable benchmark for HDR model evaluation.
[0078] (2) HDR Artifact Detector (HDR-ADet): Our novel detector HDR-ADet can accurately identify artifact areas in HDR images. By incorporating human perceptual feedback, HDR-ADet allows fine-tuning of the HDR model to improve visual quality. This is the first study to explicitly and explicitly address the HDR artifact problem by detecting artifacts.
[0079] (3) Robust performance and evaluation: Extensive experiments confirm the robustness of HDR-ADet in various scenarios and HDR reconstruction models, demonstrating its effectiveness in enhancing HDR reconstruction through intra-domain and cross-domain fine-tuning. In addition, user studies verify the reliability of AS as an evaluation metric.
[0080] Example 1
[0081] An experimental platform was built to simulate the method of the present invention, as follows:
[0082] Datasets and Models
[0083] First, for HDR artifact detection, we conduct extensive experiments on the created HDR-AD dataset to verify the performance of the proposed HDR-ADet in different scenes / image contents and HDR reconstruction methods. Second, to evaluate the effectiveness of the proposed human feedback mechanism, we integrate HDR-ADet into various HDR reconstruction models (i.e., AHDR, HDR-Trans, and SCTNet) and fine-tune it on widely used HDR reconstruction datasets (i.e., Kal's and Tel's).
[0084] Comparison of implementation effects
[0085] Tables 1 and 2 show the improvement effect of using HDR-ADet on the HDR reconstruction model on different data sets. Among the evaluation indicators used, PSNR is used to evaluate the image quality by calculating the ratio between the maximum possible signal strength and noise (error) between the original image and the reconstructed image, with a value range of [0, ∞]; SSIM is a perception-based image quality evaluation indicator that aims to measure the similarity of image structure information, brightness, and contrast, with a value range of [0, 1]; HDR-VDP-2 is a quality evaluation indicator for high dynamic range images (HDR) that aims to predict the human eye's perception of image quality differences, with a value range of [0, 100].
[0086] For each reconstruction model, Tables 1 and 2 show the results of training on HDR reconstruction datasets (i.e., Kal's and Tel's) using HDR reconstruction models (including AHDR, HDR-Trans, and SCTNet), including: the results of the original algorithm without any operation, the results of fine-tuning using the original algorithm settings, the fine-tuning effects of HDR-ADet* trained using only part of the dataset of the corresponding model, and HDR-ADet trained using the complete dataset.
[0087] Experimental results show that on the two datasets, three different HDR reconstruction models achieved improvements in indicators after fine-tuning using HDR-ADet.
[0088] Table 1 Quantitative results of HDR-ADet as a human feedback library fine-tuning HDR reconstruction model
[0089]
[0090] Table 2 Quantitative results of HDR-ADet as a human feedback cross-library fine-tuning HDR reconstruction model
[0091]
[0092]
[0093] Figure 6 The visual qualitative results before and after fine-tuning using HDR-ADet are shown, and it can be seen that these fine-tuned models retain more details and show better visual effects.
[0094] Introducing a non-reference HDR image evaluation metric, Artifacts Score (AS), based on the artifact mask generated by HDR-ADet. Unlike current HDR-specific evaluation metrics, AS highlights the impact of artifacts on HDR image quality. Since the perception of image quality is not linearly related to its size, the present invention applies a nonlinear mapping function to the artifact area ratio AR, which is defined as follows:
[0095] AS=min(log(1+α·AR),1),
[0096] α is used as a scaling factor to control the influence of the artifact area ratio. In the implementation of the present invention, α is set to 10 and AS is truncated at a maximum value of 1. If the score exceeds 1, the HDR image reconstruction quality is considered to be very poor.
[0097] This embodiment calculates the AS of all HDR images in HDR-AD-O, such as Figure 7 As shown in Figure 2, there is a clear degradation in image quality as AS increases. The results are highly consistent with human visual perception, providing a clearer understanding of image quality based on artifacts.
[0098] A user study was also conducted to examine the correlation between AS and human perception. The Spearman correlation coefficient was calculated and the participants were divided into expert and non-expert groups to improve the accuracy of image quality assessment. The quantitative results in Table 3 show that AS is highly consistent with human visual perception.
[0099] Table 3 User survey results of HDR-ADet as image evaluation index
[0100] Tester Same scene Different scenarios total expert 0.7109 0.6093 0.6601 Non-expert 0.6701 0.4980 0.5840
Claims
1. A method for detecting and eliminating perceptual artifacts for high dynamic range reconstruction, characterized in that: The following steps are involved: Step 1: Create an HDR-AD dataset In order to achieve the goal of detecting artifact areas in HDR images, a high-quality HDR-AD dataset is first constructed: LDR images are acquired at different exposures, HDR images are inferred using the existing HDR reconstruction model, and masks are obtained by manual pixel-by-pixel annotation; Step 2: Construct and train the HDR-ADet model Build an artifact detector HDR-ADet based on ViT and use the HDR-AD dataset to train this detector to locate artifact areas in the image; Step 3: Use the artifact mask obtained by HDR-ADet to fine-tune the HDR reconstruction model to improve the quality of the HDR image reconstructed by the model. Because artifacts are highly correlated with human visual perception, the artifact mask obtained by HDR-ADet can be considered as human feedback on the quality of HDR image reconstruction. By integrating HDR-ADet into the loss function calculation process of the HDR reconstruction framework, it is used to improve the quality of HDR images reconstructed by the model.
2. A method for detecting and eliminating perceptual artifacts for high dynamic range reconstruction as claimed in claim 1, characterized in that: Step 1 includes the following steps: Step 1.1 collect a multi-exposure LDR image set; A variety of image content was captured using mobile phones and digital cameras to support HDR across multiple platforms; all images were initially captured in RAW format at the native resolution of each device and then converted to TIFF format for further processing; in order to focus on difficult scenes in the HDR image generation process, the collected LDR image set was archived and organized to ensure that each final set of LDR images contained different exposures and object movement; Step 1.2: reconstructing an HDR image using an existing HDR reconstruction model; Step 1.3: Manually annotate artifact masks. The online platform labelbox is used for manual labeling. During the labeling process, images with too high or too low reconstruction quality will be excluded. If multiple annotators cannot find the artifact area in the same HDR image, its reconstruction quality is considered to be too high; if an HDR image is almost entirely artifact area and the successfully reconstructed area cannot be found, its reconstruction quality is considered to be too low; and images with significant differences between annotators will also be ignored to prevent model bias.
3. A method for detecting and eliminating perceptual artifacts for high dynamic range reconstruction as claimed in claim 1, characterized in that: The HDR-ADet network in step 2 includes: a feature backbone extraction network, a feature fusion bottleneck, and a detection head; the backbone network and the head architecture perform the detection task and are combined with a fusion bottleneck to better integrate features across multiple scales; The specific steps are as follows: Step 2.1 Construct feature backbone extraction network The feature extraction backbone network adopts a hybrid window attention strategy, which combines window attention with multiple cross-window blocks. Specifically, given an input HDR image I, the image is processed by patch embedding and Transformer blocks. The output is the final feature map F with a patch size of 16. This process is formally expressed as follows: F=V(I), Among them, V(·) is the feature extraction backbone network based on the window ViT; Step 2.2 Construct feature fusion bottleneck Since artifacts in HDR images vary in shape and size, a multi-scale feature fusion network is proposed to extract features of different scales from the output of the feature extraction backbone network; Specifically, the feature map (F) from the backbone network generates multi-scale features through a series of convolution, pooling and deconvolution layers. The feature map of each scale represents the information of the image at different resolutions, so as to capture various features from large-scale coarse information to local details. The convolution layer is responsible for extracting local features, the pooling layer helps reduce the spatial dimension and enhance the focus on the overall information of the image, and the deconvolution layer restores the detail information through upsampling operations. p i =S i (F), The value range of i is 1, 2, 3, 4, p i Indicates features of different scales, which are 1 / 32, 1 / 16, 1 / 8, and 1 / 4 feature maps, respectively. i (·) is a series of convolution, pooling, and deconvolution layers; A global information extraction mechanism is introduced to extract global context information from the feature map, which is done through a global pooling or global feature extraction module; the global information p4 captures the macroscopic structure and background information within the entire image, and the global information f g The extraction process is as follows: f g =Maxpooling(p4), Next, we will extract the global information f g Upsample to each scale feature p i The upsampling operation converts the global information into a spatial dimension that matches the feature map of each scale, so that the global information can be effectively combined with the local features. In this way, the global information provides supplementary contextual information for each scale feature, thereby helping the network to better understand the overall structure of the image and enhance the detection accuracy of artifacts: f i =U(f g )+p i , Where U(·) represents the upsampling operation; Step 2.3 Construct the detection head First, the multi-scale features are upsampled to 1 / 4 of the input HDR image size and concatenated; Then, two MLP layers are used to fuse the features and generate the final prediction; The MLP decoder MLP(·) can be formulated as follows: in is the predicted artifact mask. Step 2.4 Train the HDR-ADet network The artifact detector HDR-ADet is supervisedly trained using the dataset HDR-AD, freezing the backbone network and only updating the parameters of the feature fusion bottleneck and detection head.
4. A method for detecting and eliminating perceptual artifacts for high dynamic range reconstruction as claimed in claim 1, characterized in that: In step 3, HDR-ADet is integrated into the HDR reconstruction framework to optimize the loss function calculation: the predicted HDR image Artifact mask It is used to calculate the L1 loss penalty for the corresponding area during the loss calculation process. The corresponding loss function formula is as follows: Among them, H is the true value, L r is the original loss of the HDR reconstruction framework, is a hyperparameter, L h It is the feedback fine-tuning loss. By increasing the loss penalty in this way, the HDR model can pay more attention to the areas with poor quality, that is, the artifact areas, during the fine-tuning process, thereby improving the image quality in a targeted manner.
5. A method for detecting and eliminating perceptual artifacts for high dynamic range reconstruction as claimed in claim 1, characterized in that: In step 3, the HDR reconstruction framework is AHDR, HDR-Trans or SCTNet.
6. A method for detecting and eliminating perceptual artifacts for high dynamic range reconstruction as claimed in claim 1, characterized in that: The artifact score AS, an evaluation indicator for no-reference HDR images, is defined as follows: AS=min(log(1+α·AR),1), Here α is used as a scaling factor to control the impact of the artifact area ratio.
Citation Information
Patent Citations
Image artifact detection and automatic removal method based on U-net structure
CN111583152A
High dynamic range image artifact removing method based on attention mechanism
CN114998138A
Ghosting detection method and electronic equipment
CN119206262A
Constant Bracket High Dynamic Range (cHDR) Operations
US20150350513A1
Efficient inverse tone mapping network for standard dynamic range (SDR) to high dynamic range (HDR) conversion on HDR display
US20230059233A1