Image glare removing method based on codebook prior and double-domain alignment
By adopting deep learning methods based on codebook prior and dual domain alignment in the image glare removal method, the problems of insufficient prior utilization, difficulty in artifact separation and insufficient detail recovery in traditional methods are solved, and high-quality image glare removal and detail recovery are achieved.
Patent Information
- Application Number
- CN202411930121.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional image glare removal methods lack high-quality prior information, making it difficult to obtain robust results in complex scenarios, and model design is limited by fixed features and simple assumptions, poor generalization, nonlinear mixing of glare and real image content leads to difficulty in separation, and artifact residue or loss of details are common.
Deep learning methods based on codebook priors and dual domain alignment are adopted, high-quality prior representation is extracted through vector quantization technology, and the features of glare images and clear images are aligned in the spatial domain and frequency domain through the dual domain alignment mechanism. The feature weight is dynamically adjusted in combination with the glare perception attention mechanism to achieve accurate separation and removal of artifacts and image content.
It significantly improves the accuracy and detail recovery ability of image glare removal, improves image quality and subsequent processing reliability, and solves the problems of insufficient prior utilization, difficulty in artifact separation and insufficient detail recovery in traditional methods.
Smart Images

Figure CN119941564A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer vision and image processing, and in particular relates to an image glare removal method based on codebook prior and dual-domain alignment. Background Art
[0002] Lens flare is an optical artifact phenomenon caused by multiple reflections or scattering of a strong light source in the lens. It manifests as bright spots, halos, color overflow and reduced contrast, which seriously affects image quality and subsequent processing.
[0003] Although traditional hardware methods (such as anti-reflective coatings or special lens designs) can reduce glare to a certain extent, they are costly and have limited applicability. Glare removal methods based on deep learning show potential, but still face the following problems: First, there is a lack of effective use of high-quality clear image priors, and it is difficult to obtain robust results in complex scenes by directly fitting end-to-end models; second, the model design is limited by fixed features or simple assumptions, and the generalization is poor; third, the nonlinear mixing of glare and real image content makes separation difficult, and artifacts or detail loss are common. Summary of the invention
[0004] The purpose of the present invention is to provide a deep learning method for single image glare removal, aiming to solve the problems of lack of high-quality priors, insufficient separation of artifacts and image content, and insufficient restoration of details in the area affected by glare in traditional methods. Specifically, the present invention adopts the following technical solutions:
[0005] 1. Clear image codebook prior extraction:
[0006] In the image glare removal task, the present invention maps the potential features of the clear image to a discrete space through vector quantization technology, generates a high-quality prior representation, and optimizes the extraction process through a residual quantization strategy to efficiently capture semantic consistency and local details. The prior representation provides structured guidance for subsequent artifact separation and removal, greatly improving the accuracy of feature alignment and glare removal.
[0007] 2. Dual-domain alignment mechanism:
[0008] Based on high-quality priors, this paper proposes a dual-domain alignment mechanism, which captures local features in the spatial domain and global features in the frequency domain through convolutional networks and Fourier transform modeling, effectively separating glare artifacts from image content. This mechanism aligns glare images with clear prior features at multiple scales, ensuring the consistency of the overall image structure while restoring real textures and details.
[0009] 3. Glare perception attention mechanism:
[0010] Glare artifacts often cause different degrees of serious interference to image brightness and details in different areas. In this invention, a glare mask is generated to identify the area affected by glare, and a glare perception attention module is introduced to dynamically adjust the feature weights and focus on repairing key areas. The mask is generated by the difference between the glare image and the initial restoration result, thereby quantifying the distribution of artifacts, allowing the network to concentrate resources on processing the affected area, complementing the global alignment mechanism to achieve accurate restoration of brightness and details. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 :The overall framework of the method, including prior extraction, dual-domain alignment, and attention enhancement module.
[0012] Figure 2 :Schematic diagram of the block-based residual quantization strategy workflow
[0013] Figure 3 : Detailed structural diagram of the dual-domain interaction module.
[0014] Figure 4 : Schematic diagram of the workflow of the glare perception attention mechanism. DETAILED DESCRIPTION
[0015] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0016] The present invention is implemented including:
[0017] The overall process of the method proposed by the present invention is as follows Figure 1 As shown, the specific construction steps of the image glare removal method in the embodiment of the present invention are as follows:
[0018] Step 1: Extraction of clear image codebook prior based on deep learning and vector quantization: Figure 1 As shown in (a), we first focus on extracting high-quality clear image priors based on deep learning methods on clear image datasets to form discrete semantic feature representations that are instructive for the subsequent glare removal process. The process is divided into the following steps:
[0019] Step 1-1, build a clear image dataset: collect high-quality clear image dataset D containing a variety of scenes (including indoor, outdoor, low light, high dynamic range) c The dataset is standardized, including resizing, pixel value normalization, and data augmentation (flipping, cropping, and color perturbation) to improve the generalization ability of the model.
[0020] Step 1-2, input image and encoder to extract features: Input clear image Through the encoder E c Extract feature representation z:
[0021]
[0022] Among them, the encoder consists of multi-layer convolutional units and self-attention blocks, which can extract the global features and local details of the image.
[0023] Steps 1-3, quantized features generate discrete priors: The continuous latent features are converted into discrete priors using a block-based residual quantization strategy workflow. Convert to discretized features like Figure 2 ,The workflow of the block-based residual quantization strategy is as follows:
[0024] First, Divide into local feature blocks P;
[0025] Then for each feature block P i Perform iterative quantization:
[0026]
[0027]
[0028] in, is the residual of the d-th quantization, c k is a vector in the codebook C.
[0029] Steps 1-4, decoder reconstructs a clear image: discretization features Through the decoder D c Decoded into reconstructed image I rec :
[0030]
[0031] Optimization objective combined with pixel reconstruction loss I rec , Perceptual loss L perc and quantization loss L vq , the optimization model maximizes the representation ability of the prior features:
[0032]
[0033] The total loss function is:
[0034]
[0035] Through multiple rounds of iterations, the parameters of the encoder and codebook are dynamically updated to generate a high-quality codebook that can efficiently represent clear image features.
[0036] Step 2: Latent space projection based on dual-domain alignment mechanism: Figure 1 As shown in (b), the present invention uses latent space projection and dual-domain alignment mechanism to accurately align the features of the glare image with the prior features of the clear image in the spatial domain and frequency domain to achieve the separation of glare artifacts from the real image content. To achieve this goal, the process is divided into the following steps:
[0037] Step 2-1, encoder initialization: Initialize two encoder models with the training results of step 1, the clear image encoder E c Used to extract clear image features Glare Image Encoder E f Used to extract latent features of glare images In the glare image encoder E f The space-frequency domain interaction module is additionally introduced in E. c With E f Share the same codebook and decoder, while freezing E c , codebook C and decoder D c Parameters.
[0038] Step 2-2, feature extraction and projection: For paired input images, the clear image I c Through the clear image encoder E c Extracting high-quality prior features Glare Image I f Extracting latent features via glare image encoder The characteristics of the glare image Projection to discrete prior features The latent space where the projection result is located is optimized by using the feature alignment module to make it consistent with The alignment process is guided by a loss function to reduce the difference in feature distribution between the glare image and the clear prior.
[0039] Step 2-3, space-frequency domain interactive workflow: Based on the optimization of spatial domain features, the present invention further utilizes Fourier transform to extract the global characteristics of illumination artifacts. By combining the advantages of spatial domain and frequency domain, complementarity is achieved in local detail modeling and global feature extraction, thereby ensuring that the global distribution characteristics of artifacts are consistent with the clarity prior and filtering out glare artifacts.
[0040] like Figure 2 As shown in Figure 1, the spatial-frequency domain interaction module consists of two parts: spatial feature optimization and channel feature optimization. Spatial feature optimization realizes the spatial information interaction between spatial domain features and frequency domain features. For the input feature X, the output feature X′ evolves according to the following formula:
[0041]
[0042] in and Represents discrete Fourier transform and inverse discrete Fourier transform, Conv 3×3 Represents 3*3 convolution, DWConv 3×3 It represents a depth-wise separable convolution with a convolution kernel size of 3*3. Channel feature optimization realizes the interaction of different channels of spatial domain features and frequency domain features, following the following rules:
[0043]
[0044] Among them, Conv 1×1 Represents channel convolution.
[0045] Step 2-4, alignment loss optimization: The alignment process is constrained in the spatial domain by the following loss function to achieve layer-by-layer alignment of features:
[0046]
[0047] Among them, s∈{1, 2, 3} represents the three different scales of feature extraction by the encoder. and Represents the features extracted from the clear image and the glare image at the corresponding scale. The extracted feature information introduces the feature constraints of Fourier transform in the frequency domain to further optimize the global feature alignment effect:
[0048]
[0049] The loss functions of spatial alignment and frequency domain alignment are combined to guide the feature alignment process:
[0050]
[0051] Step 2-5, after training, the features of the glare image The quantified After the decoder D c The restored clear image is denoted as I rec .
[0052] Step 3: Detail enhancement based on glare perception attention: Figure 1 As shown in (c), step 3, based on the preliminary restoration result and feature flow of step 2, explicitly calculates the difference area between the input image and the restoration result, identifies the main impact range of the glare artifact, gradually generates a glare distribution attention map, and uses a dynamic feature enhancement mechanism to refine and repair the key areas, thereby achieving detail restoration and quality improvement of the overall image.
[0053] Step 3-1, difference area identification and preliminary mask generation: Input glare image I fand the preliminary recovery result I generated in step 2 rec , subtract the two pixel by pixel and calculate the difference:
[0054] Δ(x,y)=|I f (x,y)-I rec (x,y)|
[0055] The difference value Δ(x, y) reveals the degree of interference of the glare artifact on the original image, especially in the highlight or edge area, where the difference value is larger, and the main distribution range of the artifact can be preliminarily located.
[0056] The difference map Δ(x, y) is thresholded to filter out subtle errors and retain only significant artifact areas:
[0057]
[0058] Where, T is a preset threshold used to determine which areas are significantly affected by artifacts. The resulting preliminary glare mask M(x, y) is a binary image that marks the main artifact distribution areas.
[0059] Step 3-2, distributed attention generation and dynamic enhancement: Taking the preliminary mask M(x, y) as the anchor point, use the convolutional network to further optimize the mask, extract context information, and generate a more accurate distributed attention map:
[0060] M refined =ReLU(Conv 3×3 (M))
[0061] Optimized attention map M refined It not only contains the intensity distribution of the artifact area, but also has a more accurate boundary description, which enables fine-grained attention to the artifact area.
[0062] Step 3-3, feature reuse and glare attention enhancement: reuse the encoder E in step 2 f and D f Feature flow, where the encoder feature F enc Capturing the original details of the glare image, the decoder features F enc Reflects the initial recovery of local details and structural information. Follow:
[0063] F mix =F enc +F dec ,
[0064] F′ mix =F mix ⊙M refined +F mix
[0065] Dynamically apply distributed attention M_{refined} to the fusion features, where F mix represents the fusion feature, and ⊙ represents the multiplication of corresponding positions. The dynamic enhancement process gives higher weights to the artifact areas, strengthening the focus on these areas while maintaining global consistency.
[0066] The output of the final model is added to I in the form of residuals rec The final result is:
[0067] I res =Refine(F enc , F dec ),
[0068] I final =I rec +I res
[0069] Among them, Refine represents the optimized network of step 3.
[0070] Step 3-4, glare focus loss optimization: based on distribution attention M refined The high-weighted area introduces glare focus loss:
[0071]
[0072] The focal loss works together with the pixel-level and perceptual-level losses to optimize the training process in step 3:
[0073] L total =L rec +L perc +L focus
[0074] Step 4: Final result output: Combine steps 1-3 to generate the final high-quality restored image I final , the restored image has higher detail fidelity and natural texture.
[0075] The present invention provides an image glare removal method based on codebook prior and dual-domain alignment. The above description is only an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for removing image glare based on codebook prior and dual-domain alignment, characterized in that: The following steps are involved: Step 1-1, using the vector quantization method to extract the high-quality codebook prior of the clear image, generating potential features through the encoder and quantizing the features using the quantization codebook; Step 1-2, projecting the image affected by glare into the discrete latent space, and aligning its features by combining the spatial-frequency dual-domain interaction module; Steps 1-3, generate a glare mask based on the glare perception attention mechanism and perform regional enhancement on the feature stream of the restoration network; In steps 1-4, the restoration process is optimized by combining the loss functions to finally generate a high-quality glare-free image.
2. The method according to claim 1, characterized in that The vector quantization method adopts a residual quantization strategy, which specifically includes: Step 2-1, divide the latent features generated by the encoder into multiple local blocks; Step 2-2, gradually optimizing the residual of each local block through iterative quantization until the final quantized feature is generated; Step 2-3, use the decoder to restore the quantized features to a clear image.
3. The method according to claim 1, characterized in that The space-frequency dual-domain interaction module includes: Step 3-1, spatial interaction branch, captures the global dependency of the image through Fourier transform and extracts spatial features in combination with deep convolution operation; Step 3-2, frequency domain interaction branch, captures the inter-channel dependency through local convolution operation of frequency domain features, and restores the frequency domain information using inverse Fourier transform; Step 3-3, the fusion module combines spatial features and frequency domain features to generate aligned image feature representations.
4. The method according to claim 1, characterized in that The glare perception attention mechanism includes: Step 4-1, generating a glare mask by taking the difference between the input image and the preliminary restored image; Step 4-2, using the glare mask to perform regional weighting on the network features to enhance the detail recovery capability of the area affected by glare; In step 4-3, the attention result is fed back to the feature map as a residual to further improve the restoration quality.
5. The method according to claim 1, characterized in that The combined loss includes: Step 5-1, alignment loss, including alignment errors in the spatial domain and the frequency domain, is used to guide the feature matching between the image affected by glare and the clear image; Step 5-2, glare focus loss, enhances the recovery ability of glare area through a weighted strategy based on glare mask.
Citation Information
Cited By
Night traffic low light enhancement and overexposure suppression method based on codebook guidance
CN122415403A
A codebook-guided night traffic low-light enhancement and overexposure suppression method
CN122415403B