Solar flare removal method based on dynamic frequency domain guidance and contrast learning module
Through the dynamic frequency domain guidance and contrast learning module method, the problems of artifact separation and local content damage of light source in flare removal are solved, and efficient flare removal and image detail recovery are achieved.
Patent Information
- Application Number
- CN202510503639.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
Existing methods are difficult to effectively remove large-scale artifacts caused by flares, and often damage local content information near the light source when removing flares.
Using a method based on dynamic frequency domain guidance and contrast learning module, the flare artifacts and content information are separated by the global frequency domain dynamic guidance module, and the local details guidance module is aligned with the local details guidance module, and local details damage is suppressed in combination with the contrast learning strategy.
Effectively remove flare artifacts, while retaining local details near the light source, improving image quality and improving the performance of downstream computer vision tasks.
Smart Images

Figure CN120374461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of camera lenses, and in particular to a flare removal method based on dynamic frequency domain guidance and contrast learning module. Background Art
[0002] Strong light sources often cause flares in night-time photography, seriously degrading the visual quality of images and affecting the performance of downstream tasks. Although there have been preliminary advances in this field, existing methods still struggle to effectively remove large-scale artifacts caused by flares and ignore structural damage in local regions of the light source. Due to the damage of flare artifacts to the local content structure, unnatural details will appear near the light source in the restored image.
[0003] Lens flare is an optical phenomenon that typically occurs when a strong light source enters a camera lens. Due to light being scattered or reflected in the optical system, it forms radial bright regions, light spots, blurred spots or stripes in the image. This phenomenon is very common in photography and computer vision, especially in night-time environments where the presence of multiple artificial light sources exacerbates the impact of flares. Flares not only reduce the contrast of the image but also suppress details around the light source, resulting in a significant decrease in visual quality and affecting downstream computer vision tasks (such as object detection, semantic segmentation, and optical flow estimation, etc.).
[0004] Flares are mainly divided into two types: scattering flares and reflection flares. Scattering flares are caused by dust, scratches or abrasions on the lens surface and usually appear as radial stripes that always surround the light source and do not change with the movement of the camera or the light source. Reflection flares, on the other hand, are produced by multiple reflections inside the lens system and appear as circular or polygonal patterns related to the shape of the light source and move in the opposite direction to the light source as the camera moves. Although lens anti-reflection coatings can partially alleviate this, in simplified lens systems such as smartphones, a contaminated lens surface will exacerbate the flare problem. Designing a general flare removal method is extremely challenging because of the diverse shapes, colors and sizes of flares, and the complex interaction between them and the scene content increases the difficulty of distinguishing real targets from artifacts. Eliminating flares without compromising the overall quality of the image remains an urgent problem in this field.
[0005] Traditional flare removal methods are divided into two stages: detection and removal. This method first assumes the shape and position of the flare to hypothesize the flare and uses sample blocks to repair the flare area. However, in reality, the shapes and types of flares are diverse, and traditional methods cannot effectively remove them. Recently, some deep learning-based methods have been proposed.
[0006] Flare Removal
[0007] Physics-based Flare Removal:
[0008] The most common optical solution to avoid lens flare is to coat the lens surface with an anti-reflection coating. It uses the principle of destructive interference to weaken the reflection of light in the lens system and greatly enhance the transmission. However, it cannot completely reduce reflection, especially failing in the case where the light source is very bright. Another common method is to reduce flare by improving the camera lens material. Boynton et al. proposed a liquid-filled camera lens to mitigate flare artifacts caused by light reflection. MacLeod et al. employed a neutral density filter to minimize reflected flare artifacts. Although these specific physical methods can eliminate some lens glare artifacts, they generally cannot solve unforeseen glare.
[0009] Deep learning-based flare removal:
[0010] Most early methods were two-stage: detection and repair. These methods detect flares based on strong assumptions about the illuminance, shape, and location of flares, and then use sample patches to repair the regions. For example, Chabert et al. binarized the image using a series of thresholds, calculated the contour features of the binarized image to obtain a series of potential flare candidate regions, and then reconstructed these candidate regions. Vittoria et al. detected flare spots by overexposing the features near the flare spots and created a flare spot mask to remove the flares. These handcrafted feature-based methods are only applicable to limited types of flares, prone to treating local bright regions as flares, and difficult to distinguish different types of flares.
[0011] Recently, data-driven learning methods have been proposed. Wu et al. proposed a method for synthesizing paired training data by directly adding flare images to scene images to synthesize flare-corrupted images for training neural networks, but it did not generalize well to real-world data. Qiao et al. proposed a network trained on unpaired flare data, consisting of a light source detection module, flare detection and removal, and a generation module. Dai et al. created the first nighttime flare removal benchmark dataset, the Flare7K dataset, providing a valuable benchmark for studying this challenging nighttime flare removal task. Since the artificial spectrum has a different diffraction pattern from the solar spectrum, 7k++ uses the newly real-captured Flare-R dataset to enhance the synthetic Flare 7K dataset. To further improve the image restoration quality, FFformer uses the fast Fourier transform to extract global frequency features, enhancing the model's perception ability, and uses the global features to further enhance the removal of nighttime flares. MFDNet proposed a lightweight multi-frequency delay network based on the Laplace pyramid, decomposing the flare-damaged image into low-frequency and high-frequency bands, effectively separating the illumination and content information in the image. Although these methods can initially remove flares, they fail to effectively separate the content information from the artifact regions and have limited recovery ability for large-scale artifact regions. In addition, when dealing with the region near the light source, these methods often cause problems of local content damage to the light source.
[0012] Application of the Frequency Domain in Low-Level Vision Tasks
[0013] Frequency domain analysis has been widely applied in low-level vision tasks such as image restoration and low-light enhancement. By transforming the image from the spatial domain to the frequency domain, the details, noise, and structural information of the image can be processed more effectively. DFFormer constructs a global filter using the Fourier transform, significantly reducing the computational amount and demonstrating good performance in low-level vision tasks. He et al. achieved effective low-light video enhancement by optimizing a 4D lookup table through wavelet frequency domain construction of prior information. A proposed a wavelet-based diffusion model to learn the distribution of clean images in the frequency domain after wavelet transform, greatly improving the inference efficiency. B utilized the frequency domain prior information in the image to enhance the restoration quality of zero-shot low-light image enhancement. D proposed a fast Fourier transform network based on frequency domain discrimination, which introduced a gating mechanism based on the Joint Photographic Experts Group (JPEG) compression algorithm to discriminatively determine which low-frequency and high-frequency information of the features should be retained, thus restoring a clear image.
[0014] To further enhance the quality of the restored image, Harmonizing proposed a plug-and-play Adaptive Focusing Module (AFM), which can adaptively mask the clean background area and help the model focus on the areas severely affected by flares. FFformer uses the Fast Fourier Transform to extract global frequency features, enhancing the model's receptive field and further enhancing the elimination of night-time flares using global features. However, these methods fail to effectively separate the content information from the artifact regions and have limited recovery ability for large-scale artifact regions. In addition, when processing the regions near the light source, these methods often cause problems of damage to the local content of the light source.
[0015] Therefore, a flare removal method based on dynamic frequency domain guidance and contrast learning module is designed to provide a technical solution to the above technical problems. Summary of the Invention
[0016] Based on this, it is necessary to provide a flare removal method based on dynamic frequency domain guidance and contrast learning module to solve the technical problems raised in the above background technology.
[0017] To solve the above technical problems, the present invention adopts the following technical solutions:
[0018] A flare removal method based on dynamic frequency domain guidance and contrast learning module, 1. A flare removal method based on dynamic frequency domain guidance and contrast learning module, characterized in that the steps are as follows:
[0019] S1: The global frequency domain dynamic guidance module separates the flare artifacts and content information by dynamically optimizing the global frequency domain weights of the features, guiding the network to retain the content information while removing the artifacts;
[0020] S2: The local detail guidance module guides the local features of the light source to align with the reference image to promote the refined recovery of the image.
[0021] As a preferred embodiment of the flare removal method based on dynamic frequency domain guidance and contrast learning module provided by the present invention, in step S1, the global frequency domain dynamic guidance module separates the flare artifacts and content information by dynamically optimizing the global frequency domain weights of the features, guiding the network to retain the content information while removing the artifacts, and the steps are as follows:
[0022] Replace the window attention in Uformer with the dynamic frequency domain guidance module;
[0023] Let the flare-damaged image features pass through N encoders;
[0024] Initialize the learnable parameters as the initial weights, apply the weights to each channel of the spectrogram, and guide the model to perceive the flare artifact regions;
[0025] Convert the features to the spatial domain through inverse Fourier transform, and establish residual connections to enhance local detail feature identification.
[0026] As a preferred embodiment of the flare removal method based on the dynamic frequency domain guidance and contrast learning module provided by the present invention, each encoder includes a GDFG module and a downsampling layer;
[0027] The GDFG module is used to convert the features from the spatial domain to the frequency domain through Fourier transform, and dynamically optimize the weights of the feature frequency domain through the information between different channels of the input features and learnable parameters, so as to perceive the characteristics of flare artifacts;
[0028] The downsampling layer is used to reshape the flattened features into a 2D spatial feature map, and then downsample the map, doubling the number of channels using a 4×4 convolution with a stride of 2.
[0029] As a preferred embodiment of the flare removal method based on the dynamic frequency domain guidance and contrast learning module provided by the present invention, initialize the learnable parameters as the initial weights, and the steps are as follows:
[0030] Take the input features as the weight coefficients, dynamically optimize the coefficient of the weights according to the characteristics of the input data, and perform targeted enhancement on the target features. The input X passes through MPL, and the expression is as follows:
[0031] M(X) = W2starRelu(W1LN(X));
[0032] Among them, both W1 and W2 represent learnable weight matrices.
[0033] As a preferred embodiment of the flare removal method based on the dynamic frequency domain guidance and contrast learning module provided by the present invention, in step S2, the local detail guidance module is used to guide the local features of the light source to align with the reference image to promote the refined restoration of the image, and the steps are as follows:
[0034] Construct a local detail guidance module according to the framework of contrast learning;
[0035] For each positive sample, find the target feature block that best matches it.
[0036] As a preferred embodiment of the flare removal method based on the dynamic frequency domain guidance and contrast learning module provided by the present invention, construct a local detail guidance module according to the framework of contrast learning, and the steps are as follows:
[0037] Given the image I' output by the restoration network and the reference image I, and convert them into the form of the two-dimensional spatial domain;
[0038] Randomly extract local feature blocks and extract smaller feature subsets for contrast learning.
[0039] As a preferred embodiment of the flare removal method based on the dynamic frequency domain guidance and contrast learning module provided by the present invention, for each positive sample, find the target feature block that best matches it, and the steps are as follows:
[0040] A positive sample pair is formed between each reference feature block and the corresponding target feature block, and its similarity
[0041] metric is:
[0042]
[0043] where z1 and z2 both represent the reference feature block and the corresponding target feature block;
[0044] Use other patches of the input image X as negative samples;
[0045] Calculate the normalized similarity between the positive sample and all samples through the softmax function, and the expression
[0046] is as follows:
[0047]
[0048] where p i represents the probability of the i-th category, T is the temperature parameter used to control the smoothness of the softmax function, and s i represents the positive sample similarity;
[0049] Constrain the matching distribution of the positive sample through the cross-entropy loss, and the expression is as follows:
[0050]
[0051] where N represents the total number of samples.
[0052] It can be undoubtedly seen that through the above technical solutions of the present application, the technical problems to be solved by the present application can surely be solved.
[0053] At the same time, through the above technical solutions, the present invention has at least the following beneficial effects:
[0054] 1. The flare removal method based on the dynamic frequency domain guidance and contrast learning module provided by the present invention proposes a global dynamic frequency domain guidance module. By dynamically optimizing the global frequency domain features, the flare information is decoupled from the content information, and while eliminating artifacts, the interference of artifacts on the surrounding content information is suppressed.
[0055] 2. The present invention introduces a contrastive learning strategy and designs a local detail guidance module to guide the local area of the light source to align with the reference image, so as to effectively suppress the local detail damage caused by flare removal.
[0056] 3. The method of the present invention explores from both global and local perspectives, and through a large number of experiments, it is proved that it can effectively remove flare artifacts and repair them. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0058] Figure 1 It is a schematic overview diagram of the DFDNet of the present invention;
[0059] Figure 2 It is a schematic diagram of the LDGM of the present invention;
[0060] Figure 3 It is a schematic diagram of the restoration results of the present invention on real and synthetic night flare-damaged datasets;
[0061] Figure 4 It is a schematic diagram of the comparison of the restoration results between the method of the present invention and the state-of-the-art method on real-world night flare-damaged datasets and consumer electronics test datasets;
[0062] Figure 5 It is a schematic diagram of the visualization of the ability of the GDFG module of the present invention to perceive flares;
[0063] Figure 6 It is a schematic diagram of the dimension experiment of the filter in the present invention;
[0064] Figure 7 It is a comparison diagram of the restoration results with and without the GDFG module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following will further describe the present invention in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0066] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings.
[0067] It should be noted that, without conflict, the embodiments in the present invention and the features and technical solutions in the embodiments may be combined with each other.
[0068] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0069] Referring to Figures 1-7 , a flare removal method based on a dynamic frequency domain guidance and contrast learning module, at least includes the following steps:
[0070] S1: The global frequency domain dynamic guidance module separates flare artifacts and content information by dynamically optimizing the global frequency domain weights of features, guiding the network to retain content information while removing artifacts.
[0071] S2: The local detail guidance module guides the local features of the light source to align with the reference image to promote the refined recovery of the image.
[0072] S1 at least includes the following steps:
[0073] Replace the window attention in Uformer with the dynamic frequency domain guidance module;
[0074] To alleviate the situation of large-area artifacts and local damage near the light source, exploration was carried out from two levels. First, from the perspective of the global frequency domain, an attempt was made to decouple flare artifacts and the background area, guiding the recovery network to suppress the negative impact of artifacts on surrounding background pixels while removing flares. Then, considering local information, the local light source was aligned with the reference image to further promote the positively oriented refined recovery of the image.
[0075] As Figure 1 The overall framework is shown in the figure. Among them, DFDNet consists of multiple global dynamic frequency domain guidance (GDFG) modules. A global frequency domain dynamic guidance module was designed and embedded in Uformer to dynamically optimize the global frequency domain weights of features using learnable parameters, guiding the network to analyze the frequency domain characteristics of artifacts. Then, a contrast learning strategy was introduced, and a local guidance module was designed to suppress local detail damage caused by flare removal.
[0076] The network of the present invention is based on Uformer and is an overall U-shaped structure. The flare removal method based on Uformer has been verified to be effective by Flare7k++. It gradually extracts global features through multiple downsamplings in the spatial domain of features. Due to the use of the local window self-attention mechanism, some global information will inevitably be ignored, and it is impossible to well analyze the characteristics of flares and normal texture images in the spatial domain, which is not conducive to removing large-area flare artifacts.
[0077] Let the flare-damaged image features pass through N encoders;
[0078] As Figure 1 shown, referring to the U-shaped structure of Uformer, given the flare-damaged image I, first apply a 3×3 convolutional layer of LeakyReLU to extract the low-level feature X_0 ∈ RC×H×W, where C is the number of channels, and H, W are the height and width. Next, this feature passes through N encoders (in the present invention, N is set to 4). Each encoder sub-stage includes a GDFG module and a downsampling layer. The GDFG module transforms the feature from the spatial domain to the frequency domain through Fourier transform, and dynamically optimizes the weights in the frequency domain of the feature by the information between different channels of the input feature and learnable parameters, gradually perceiving the characteristics of flare artifacts.
[0079] In the downsampling layer, first reshape the flattened feature into a 2D spatial feature map, and then downsample the map, doubling the number of channels using a 4×4 convolution with a stride of 2. The encoder stage can be expressed as:
[0080] E n = Encoder(E n-1 ) = Down(GDFG(E n-1 ))
[0081] where GDFG is the global dynamic frequency domain guidance module, and Down is the downsampling layer.
[0082] At the end of the encoder, GDFG is still used as the bottleneck stage of the network to separate the flare features from the latent space. In the decoding stage, a combination of N layers of GDFG and upsampling layers is used to reconstruct the image. First, upsample using a 2×2 transposed convolution with a stride of 2, which reduces the number of feature channels by half and doubles the size of the feature map. After that, the feature input to the GDFG block is a combination of the upsampled feature and the corresponding feature from the encoder through a skip connection. After N decoder stages, the flattened feature is reshaped into a 2D feature map, and a 3×3 convolutional layer is applied to obtain the restored image and the predicted flare image. To suppress the influence of flares on the local light source, a contrastive learning strategy is introduced, and LDGM is designed to further promote the positively oriented refined restoration of the image by maximizing the mutual information between the local light source of the restored image and the reference image.
[0083] Initialize the learnable parameters as the initial weights, and apply the weights to each channel of the spectrogram to guide the model to perceive the flare artifact area;
[0084] To separate the characteristics of flare artifacts from the perspective of the global frequency domain, the GDFG module is designed, as Figure 1As shown, it consists of a 2D Fourier transform, an MLP, a Normal layer, and residual connections. In the early global filter method, learnable global filter weights are used to optimize and adjust the Fourier spectrogram. However, considering the deep fusion of large-area flare artifact features and background information, it is difficult to separate the artifact characteristics with only a single global filter. Therefore, a global dynamic weight with N dimensions is proposed and applied to each channel of the features. The coefficients of these weights are determined by the input features, so that the model can perceive the target features more quickly.
[0085] Specifically, first define the multi-channel global filter D, and the expression is as follows:
[0086] G = F -1 (W ⊙ F(X))
[0087] Among them, F -1 represents the inverse Fourier transform, F represents the Fourier transform, G represents the filter, and W represents the calculated dynamic multi-channel weights.
[0088] Next, initialize the weight W. It is a series of learnable parameters. Use the input features as the weight coefficients, and dynamically optimize the coefficients of the weights according to the characteristics of the input data, so as to enhance the target features in a targeted manner. The input X passes through the MPL and is defined as:
[0089] M(X) = W2starRelu(W1LN(X))
[0090] Among them, both W1 and W2 represent learnable weight matrices. Then, apply the weights to each channel of the spectrogram to guide the model to perceive the flare artifact area. After execution, transform the spectrogram to the spatial domain through the inverse Fourier transform, and establish a residual connection to enhance the local detail features.
[0091] S2 includes at least the following steps:
[0092] Construct a local detail guidance module according to the contrast learning framework;
[0093] Previous methods will cause damage to the local light source when eliminating flares, such as Figure 1 shown in the second row. This is because the local detail information of the light source is lost while eliminating the flares. Based on this, a local detail guidance module is designed using the contrast learning framework, as Figure 2 shown. A patch-based method is used instead of operating on the entire image, constraining the local consistency of the target image and reference image features, focusing on local detail alignment and accelerating the model convergence speed.
[0094] For each positive sample, find the target feature block that best matches it;
[0095] Specifically, given the restored network output image I’ and the reference image I, first convert them into the form of two-dimensional spatial domain, randomly extract local feature blocks, and extract smaller feature subsets for contrastive learning. The goal of using contrastive learning is to match the corresponding input-output patches at a specific location. For each positive sample, the goal is to find the target feature block that best matches it. A positive sample pair is formed between each reference feature block and the corresponding target feature block, and its
[0096] similarity metric is:
[0097]
[0098] where z1 and z2 both represent the reference feature block and the corresponding target feature block;
[0099] Other patches of the input image X can be used as negative samples. For example, in the figure: the light source center should be more closely associated with other patches of the same input (such as the halo around the light source, the dark night background)
[0100] than with the local input light source. A negative sample set is formed by other feature blocks in the batch, and the similarity of the negative samples is calculated. Then, the normalized similarity of the positive sample to all samples is calculated through the softmax function:
[0101]
[0102] where p i represents the probability of the i-th category, T is the temperature parameter used to control the smoothness of the softmax function, and s i represents the positive sample similarity.
[0103] The matching distribution of the positive sample is constrained by the cross-entropy loss, and the expression is as follows:
[0104]
[0105] where N represents the total number of samples;
[0106] The loss function
[0107] During the training process, first add the flare image F to the background image IB to obtain the flare-corrupted image I. The flare removal network can be defined as Φ, which takes the flare-corrupted image I as the input. Then, the estimated flare-free image and the flare image can be expressed as:
[0108]
[0109] Supervise the flare and background images using the L1 loss and the perceptual loss Lvgg. The background image and the flare loss can be written as:
[0110]
[0111] L F = L1(I′, I0)
[0112] where L B represents the background image reconstruction loss, and L F represents the flare reconstruction loss;
[0113] The reconstruction loss L rec is defined as:
[0114]
[0115] where Clip represents clipping, represents the addition operation. Then, the addition is clipped to the range [0, 1].
[0116] The presence of flares often interferes with the global frequency distribution of the image, and traditional pixel-level losses (such as L1 or L2) are difficult to effectively capture this global frequency domain information. Therefore, by introducing a loss in the frequency domain and through the joint constraint of the frequency domain and the spatial domain, the model's ability to recover the global structure and detailed texture is improved, and the recovery deviation is reduced. Specifically, for the recovered image L rec and the reference image L rec , the fast Fourier transform is performed respectively to obtain their representations in the frequency domain, including amplitude and phase. The L1 loss is calculated for the amplitude and phase of both, respectively, and is defined as:
[0117] L FFT = L1(A rec , A ref ) + L1(P rec , P ref )
[0118] where A rec represents the reconstructed image amplitude, A ref represents the reference image amplitude, P rec represents the reconstructed image phase, and P ref represents the reference image phase;
[0119] Generally speaking, the final loss function aims to minimize the weighted sum of all these losses:
[0120] L = L B + L F + L rec + L LDGM + L FFT .
[0121] Dataset
[0122] A paired set of flare-corrupted images and flare-free images was generated using the Flare7k++ synthesis pipeline as the training set. The background images were sampled from 24K Flickr images.
[0123] Flare images and their corresponding light sources were sampled from the Flare 7K and Flare-R datasets with a 50% probability.
[0124] For fair comparison, the data augmentation strategy also followed Flare7k++. To verify the robustness of the method of the present invention, experiments were conducted on four test sets: paired test sets: Flare7K++ real test dataset, Flare7K++ synthetic test dataset, unpaired test sets: consumer electronics test dataset, Flare-corrupted images, as shown in Tables 1 - 3.
[0125] Table 1: Quantitative comparison of real and synthetic night flare corruption data
[0126]
[0127]
[0128] Table 2: Quantitative comparison of real-world night flare corruption datasets
[0129] Method Published NIQE MUSIQ PI Flare7k NeurIPS'22 2.725 64.157 1.856 BracketFlare CVPR'23 2.984 63.889 1.963 IR-SDE ICML'23 3.005 64.689 2.017 Flare7k++ TPAMI'24 2.867 64.419 1.926 Kotp ICASSP'24 2.750 64.222 1.87 Ours 2.714 64.702 1.849
[0130] Table 3: Quantitative comparison of consumer electronics test datasets
[0131]
[0132]
[0133] Comparative experiment
[0134] Qualitative analysis
[0135] First, Figure 3 shows a visual comparison of the Flare7k++ real-world and synthetic test sets. The results show that the method of the present invention is most effective in eliminating large-scale artifacts and restoring local details of the light source. As shown in the first row of Figure 1 , the existing method causes damage to the local part of the light source after removing the flare, while the method of the present invention can still restore the content information of the local part of the light source. As shown in Figure 3As shown in the second line of , although the existing method eliminates the glare effect, there are still large-scale flare artifacts. The method of the present invention has produced satisfactory results. Next, a visual comparison on the unpaired test set Flare-corrupted images is shown, as Figure 4 shown. For such real-world night-time flare-corrupted images, the method of the present invention can also recover very well. To verify the robustness of the method of the present invention, tests were also carried out on the consumer electronics test dataset, as Figure 4 shown. This dataset contains flare-corrupted images captured by 10 types of consumer electronics products. The results show that the method of the present invention still has good generalization performance in the face of daytime flares and strong flares generated by mobile phone lenses.
[0136] Ablation experiments
[0137] Night-time flares are usually accompanied by large-scale artifacts, and it is difficult for existing methods to eliminate this phenomenon. The present invention designs a GDFG module to dynamically separate flare artifacts, Figure 3 and the results of show that the method of the present invention effectively alleviates this problem. To verify the role of this module in the network of the present invention, the present invention conducts ablation experiments, as shown in Table 4. The results show that GDFG can effectively improve the performance of the model in removing flares. If the designed frequency-domain loss is added at the end of the network, through the joint constraint of the spatial domain and the frequency domain, the perceptual quality of the image will be further improved.
[0138] The ability of the GDFG module to perceive flares is also visually demonstrated, as Figure 5 shown. Method 1 with and without the GFDG module, as well as the flare region annotation, are respectively shown. The results show that GDFG can effectively perceive the flare and artifact regions.
[0139] In addition, experiments were carried out on the dimensions of the filters in the module, and the results are as Figure 6 shown. When N = 2 or 3, there are slight differences in the pixel distributions between the restored image and the reference image. When N = 4, the pixel value distributions are the closest. Therefore, in the experimental settings, N = 4.
[0140] Table 4: Ablation study of different module combinations in the network of the present invention
[0141]
[0142] To alleviate the problem of losing content information due to flare removal, LDGM was designed based on a contrastive learning strategy to maximize the mutual information between the local light source of the restored image and the reference image. Ablation experiments of this module were conducted, as shown in Table 4. From the results in the table, it can be seen that when this module was introduced, the value of S-PSNR increased significantly, improving by 0.607 compared to the baseline method, which proves that this module is effective in restoring these stripe flare regions. At the same time, there was also an improvement in the G-PSNR metric, indicating that this module also has a positive effect on the content restoration near the light source. A comparison chart of the restoration results with and without this module is shown, as Figure 7 shown.
[0143] From the visualization results, it can be clearly seen that some detail information in the light source area of the figure was lost due to flare removal. When LDGM was added, the method of the present invention could effectively restore this detail information.
[0144] In addition, experiments were conducted on the model convergence speed, as Figure 2 shown. When LDGM was introduced, the model of the present invention had reached the best effect at 150,000 iteration times, which was nearly 40% faster than without this module.
[0145] Conclusion
[0146] The Dynamic Frequency-guided Flare Network (DFDNet) was proposed to address the challenge of flare artifacts in night photography. DFDNet learns flare features from the frequency domain perspective and integrates a contrastive learning strategy to adjust local features, thus effectively removing large-scale artifacts and significantly reducing the damage to the structure caused by strong light sources in the local area. The proposed Global Dynamic Frequency-guided (GDFG) module dynamically optimizes global features using frequency domain information, realizes the precise decoupling of flare information and content information, and minimizes the interference to the surrounding content. In addition, a Local Detail-guided Module (LDGM) based on the contrastive learning strategy was designed to make the local features consistent with the reference image, reduce the loss of local details, and ensure the fine restoration of the image. Experimental results show that DFDNet performs excellently in flare removal and image quality restoration, making outstanding contributions to the flare removal task.
[0147] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art in the relevant technical field can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A flare removal method based on a dynamic frequency domain guidance and contrast learning module, characterized in that The steps are as follows: S1: The global frequency-domain dynamic guidance module separates flare artifacts and content information by dynamically optimizing the global frequency-domain weights of features, guiding the network to remove artifacts while retaining content information; S2: The local detail guidance module guides the local features of the light source to align with the reference image to promote the refined recovery of the image.
2. The flare removal method based on a dynamic frequency domain guidance and contrast learning module according to claim 1, wherein In step S1, the global frequency-domain dynamic guidance module separates flare artifacts and content information by dynamically optimizing the global frequency-domain weights of features, guiding the network to remove artifacts while retaining content information. The steps are as follows: Replace the window attention in Uformer with the dynamic frequency-domain guidance module; Let the flare-damaged image features pass through N encoders; Initialize the learnable parameters as the initial weights, apply the weights to each channel of the spectrogram, and guide the model to perceive the flare artifact area; Convert the features to the spatial domain through inverse Fourier transform and establish a residual connection to enhance the local detail feature identification.
3. A flare removal method based on a dynamic frequency domain guidance and contrast learning module according to claim 2, characterized in that, Each encoder includes a GDFG module and a downsampling layer; The GDFG module is used to convert the features from the spatial domain to the frequency domain through Fourier transform, and dynamically optimize the weights of the feature frequency domain through the information between different channels of the input features and the learnable parameters to perceive the flare artifact characteristics; The downsampling layer is used to reshape the flattened features into a 2D spatial feature map, then downsample the map, and double the number of channels using a 4×4 convolution with a stride of 2.
4. A flare removal method based on a dynamic frequency domain guidance and contrast learning module according to claim 1, characterized in that Initialize the learnable parameters as the initial weights. The steps are as follows: Take the input features as the weight coefficients, dynamically optimize the coefficient of the weights according to the characteristics of the input data, and perform targeted enhancement on the target features. The input X passes through MPL, and the expression is as follows: M(X) = W2starRelu(W1LN(X)); Among them, both W1 and W2 represent learnable weight matrices.
5. A flare removal method based on a dynamic frequency domain guidance and contrast learning module according to claim 1, characterized in that, In step S2, the local detail guidance module guides the local features of the light source to align with the reference image to promote the refined recovery of the image. The steps are as follows: Construct the local detail guidance module according to the framework of contrastive learning; For each positive sample, find the target feature block that best matches it.
6. A flare removal method based on a dynamic frequency domain guidance and contrast learning module according to claim 5, characterized in that, Construct the local detail guidance module according to the framework of contrastive learning. The steps are as follows: Given the image I' output by the recovery network and the reference image I, and convert them into the form of a two-dimensional spatial domain; Randomly extract local feature blocks and extract smaller feature subsets for contrastive learning.
7. A flare removal method based on a dynamic frequency domain guidance and contrast learning module according to claim 5, characterized in that, For each positive sample, find the target feature block that best matches it. The steps are as follows: A positive sample pair is formed between each reference feature block and the corresponding target feature block, and its similarity measure is: Among them, both z1 and z2 represent the reference feature block and the corresponding target feature block; Use other patches of the input image X as negative samples; Calculate the normalized similarity between the positive sample and all samples through the softmax function. The expression is as follows: Among them, p i represents the probability of the i-th category, T is the temperature parameter used to control the smoothness of the softmax function, and s i represents the positive sample similarity; Constrain the matching distribution of the positive sample through the cross-entropy loss. The expression is as follows: Among them, N represents the total number of samples.