Night image halo removing method

By constructing prior knowledge and deep unfolding networks, combined with the UBFormer model and LWTB network architecture, the problem of incomplete removal of light spots and stripes artifacts in night images is solved, and better image texture preservation and detail recovery are achieved without increasing costs.

CN120634898AActive Publication Date: 2025-09-12SOUTHWEST UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510795385.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-15
Publication Date
2025-09-12
Estimated Expiration
2045-06-15

AI Technical Summary

Technical Problem

Existing technologies cannot fully utilize prior information in night image processing, resulting in incomplete removal of light spots and streaks, affecting visual quality and increasing application costs.

Method used

Prior knowledge is constructed, including pixel-level spatial consistency terms, feature-level detail enhancement terms, and feature-level noise suppression terms. Multiple iterative operations are performed through initialization of the network and deep expansion of the network. The UBFormer model and LWTB network architecture are combined with the proximal network to learn implicit prior information to extract the halo-removed image.

Benefits of technology

Without increasing costs, it effectively removes light spots and streaks from nighttime images, preserves image texture and details, and improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634898A_ABST
    Figure CN120634898A_ABST
Patent Text Reader

Abstract

The invention provides a nighttime image halo removing method, which comprises the following steps of: constructing prior knowledge, constructing an initialization network and a deep expansion network according to the prior knowledge, and extracting a variable initial value set from a nighttime halo polluted image by adopting the initialization network; inputting the nighttime halo pollution image and the variable initial value set into a deep expansion network for multiple iterative operations, and extracting a final halo-removed image from the nighttime halo pollution image; wherein the deep expansion network comprises a plurality of near-end networks, and each near-end network participates in one iterative operation. Under the condition that priori knowledge is fully utilized, a deep expansion network with a plurality of near-end networks is adopted to carry out iteration on a mapping graph, a halo-free feature graph, a constraint variable and a halo-removed image, so that a final halo-removed image is extracted, image textures are better reserved, image details are better recovered, and the image quality is improved. On the premise that the application cost is not increased, light spots and stripe artifacts in the image are well removed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method for removing halo from nighttime images. Background Art

[0002] When photographing at night, strong light scattering or reflections within the lens can create flare and streak artifacts, reducing image contrast and detail clarity, impacting visual quality and the performance of image processing algorithms. For example, a stereo camera in a nighttime driving scenario might mistakenly identify halos as obstacles. A drone tracking an aerial target might even track the wrong target.

[0003] The first method to reduce halo is to optimize hardware design, such as using special lenses or applying anti-reflective coatings on the lenses. However, these hardware improvement measures cannot fully remove halo, and increase the application cost. In addition, unpredictable light spots will appear when the lens is contaminated. The second method to reduce halo is to design image processing algorithms to remove halo in images. Recently, learning-based methods have achieved certain results in the field of image restoration. For example, several image restoration models are selected as baseline models; based on the Swin Transformer, fast Fourier transform is introduced to extract global frequency features; cascade neural networks are combined with fine-tuning strategies, and triples and contrastive learning are constructed to optimize the model. Existing methods do not fully utilize prior information when processing images, resulting in incomplete removal of light spots and stripe artifacts in the image, affecting visual quality.

[0004] Therefore, a method is needed that can make full use of prior information to better preserve image texture and restore image details, and better remove light spots and stripe artifacts in images without increasing application costs. Summary of the Invention

[0005] In order to overcome the problems existing in the related art, the purpose of the present invention is to provide a method for removing halo from night images, which can make full use of prior information to better preserve image texture and restore image details, and better remove light spots and streak artifacts in the image without increasing the application cost.

[0006] A method for removing halo from a nighttime image, comprising: Constructing prior knowledge, the prior knowledge including a pixel-level spatial consistency term, a feature-level detail enhancement term, and a feature-level noise suppression term; An initialization network and a depth expansion network are constructed based on prior knowledge, and a set of initial variable values ​​is extracted from the nighttime halo-contaminated image using the initialization network; wherein the set of initial variable values ​​includes an initial halo mask image, an initial halo-removed image, an initial mapping image, an initial halo-free feature map, and initial constraint variables; The night halo pollution image and the variable initial value set are input into a deep expansion network for multiple iterative operations, and a final halo-removed image is extracted from the night halo pollution image; wherein the deep expansion network includes multiple proximal networks, and each proximal network participates in one iterative operation.

[0007] In a preferred technical solution of the present invention, the constructing of prior knowledge includes: The halo layer decomposition model is constructed according to the following formula: ; Among them, I is the night halo pollution image, B is the halo removal image, and F is the halo mask image; Construct a first optimization function corresponding to the pixel-level spatial consistency term: ; Where f represents the Fibonacci norm, is a Fibonacci norm term, which is used to ensure the fidelity of image data. E1 is a first optimization function.

[0008] In a preferred technical solution of the present invention, the constructing of prior knowledge includes: The convolution dictionary is used to extract the feature map according to the following formula: ; ; ; Among them, ZI is the feature map of the night halo pollution image, ZB is the feature map without halo, ZF is the feature map of halo mask; D is the convolution dictionary, is the convolution operator; Construct a second optimization function including the feature-level detail enhancement term: ; Among them, M is the mapping graph, * is element-by-element multiplication, is the first penalty coefficient, is the second penalty coefficient, is the first regularization coefficient, is the first prior term, and E2 is the second optimization function.

[0009] In a preferred technical solution of the present invention, the constructing of prior knowledge includes: The feature-level noise prior is constructed according to the following formula: ; in, is the second regularization coefficient, It is the norm of the element-by-element multiplication of the halo-free feature map and the halo mask feature map; Construct the third optimization function of the feature-level noise suppression term: ; in, is the convolution of the convolution dictionary and the halo mask image, and the norm of the element-wise product of the halo-free feature map, and E3 is the third optimization function.

[0010] In a preferred technical solution of the present invention, before extracting the variable initial value set from the night halo pollution image using the initialization network, the method further includes: Calculate the Charbonnier consistency loss between the halo-removed image and the true value to obtain the first loss function; Calculate the Charbonnier consistency loss between the halo mask image and the training set to obtain the second loss function; Add the first loss function and the second loss function to obtain the third loss function; Calculate the Vgg consistency loss between the halo removed image and the true value to obtain the fourth loss function; Calculate the Vgg consistency loss between the halo mask image and the training set to obtain the fifth loss function; Add the fourth loss function and the fifth loss function to obtain the sixth loss function; Performing a weighted summation on the third loss function and the sixth loss function to obtain a final loss function; The final loss function is used to train the initialized network and the deeply expanded network end-to-end simultaneously.

[0011] In a preferred technical solution of the present invention, the step of inputting the nighttime halo pollution image and the variable initial value set into a deep expansion network for multiple iterative operations includes: At the t-th iteration, the gradient of the t-th map is calculated, and the t-th map and the gradient of the t-th map are input into the t-th proximal network of the depth expansion network to calculate the t+1-th map; Convolve the convolution dictionary with the halo mask image to obtain the tth halo mask feature; calculate the gradient of the tth halo-free feature map, input the tth halo-free feature map, the gradient of the tth halo-free feature map, and the tth halo mask feature into the soft threshold operator to obtain the t+1th halo-free feature map; Calculate the gradient of the t-th constraint variable, input the t-th constraint variable and the gradient of the t-th constraint variable into the t-th proximal network of the depth expansion network, and calculate the t+1-th constraint variable; The gradient of the t-th halo-removed image is calculated, and the t+1-th halo-removed image is calculated according to the gradient of the t-th halo image and the t-th halo-removed image.

[0012] In a preferred technical solution of the present invention, the step of inputting the t-th mapping image and the gradient of the t-th mapping image into the t-th proximal network of the depth expansion network to calculate the t+1-th mapping image includes: The t+1th mapping graph is calculated according to the following formula: ; Among them, M t+1 is the t+1th mapping graph, M t is the tth mapping graph, is the gradient of the t-th mapping, Update parameters for the mapping graph, * is element-by-element multiplication, prox t1 is the t1th proximal network of the map of the deeply unfolded network.

[0013] In a preferred technical solution of the present invention, the step of inputting the t-th halo-free feature map, the gradient of the t-th halo-free feature map, and the t-th halo mask feature into a soft threshold operator to obtain the t+1-th halo-free feature map comprises: The t+1th halo-free feature map is calculated according to the following formula: ; Among them, ZB t+1 is the t+1th halo-free feature map, soft is the soft threshold operator, ZB t is the t-th halo-free feature map, Update parameters for the halo-free feature map, is the gradient of the t-th halo-free feature map, is the second regularization coefficient, is the first penalty coefficient, is the second penalty coefficient, D t is the t-th convolution dictionary, F is the halo mask image, is the convolution operator and * is the element-wise multiplication.

[0014] In a preferred technical solution of the present invention, the calculation of the t+1th constraint variable includes: The t+1th constraint variable is calculated according to the following formula: ; Among them, N t+1is the t+1th constraint variable, prox t2 is the t2th proximal network of the constraint variable of the deep expansion network, N t is the t-th constraint variable, is the update parameter of the constraint variable, is the gradient of the t-th constraint variable, and * is element-wise multiplication.

[0015] In a preferred technical solution of the present invention, the step of calculating the t+1th halo-removed image includes: The t+1th halo-removed image is calculated according to the following formula: ; Among them, B t+1 For the t+1th halo removed image, B t For the t-th halo removed image, Update parameters for halo removal image, Remove the gradient of the image for the t-th halo, * is element-wise multiplication.

[0016] The beneficial effects of the present invention are: The method for removing halo from nighttime images provided by the present invention involves constructing prior knowledge, including a pixel-level spatial consistency term, a feature-level detail enhancement term, and a feature-level noise suppression term. The pixel-level spatial consistency term, based on the principle of image additivity, decomposes the nighttime halo-contaminated image into a halo-removed image and a halo mask image, thereby simplifying computations. The pixel-level spatial consistency term is constrained by a Fibonacci norm term to ensure data fidelity. The feature-level detail enhancement term uses a convolutional dictionary to extract richer features from the nighttime halo-contaminated image, the halo-removed image, and the halo mask image, thereby enhancing the learning capabilities of the initialization network and the deep expansion network. The feature-level detail enhancement term uses the Fibonacci norm to constrain the relationship between the halo-removed image, the halo-free feature map, and the convolutional dictionary D, thereby adaptively suppressing streak artifacts in the nighttime halo-contaminated image and enhancing details in halo areas. The feature-level noise suppression term uses a y-norm to constrain the halo-free feature map and the halo mask feature map, thereby suppressing the halo noise content of the halo-free feature map by increasing the halo mask feature map. A nighttime halo-contaminated image and prior knowledge are input into an initialization network, which extracts a set of initial variable values ​​from the nighttime halo-contaminated image. This set of variable values ​​includes an initial halo mask image, an initial halo-removed image, an initial mapping graph, an initial halo-free feature map, and initial constraint variables. The nighttime halo-contaminated image and the initial variable value set are then input into a deep unfolding network for multiple iterations to extract the final halo-removed image from the nighttime halo-contaminated image. Each iteration of the deep unfolding network utilizes a proximal network to learn implicit priors. This proximal network combines the advantages of CNNs in local feature modeling with the advantages of Transformers in global information modeling. Leveraging prior knowledge, the present invention utilizes a deep unfolding network with multiple proximal networks to iteratively process the mapping graph, halo-free feature map, constraint variables, and halo-removed image to extract the final halo-removed image. This method better preserves image texture and restores image detail, effectively removing speckle and streak artifacts from the image without increasing application cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flow chart of a method for removing halo from nighttime images according to the present invention; Figure 2 This is a flow chart of the present invention, which inputs a nighttime halo pollution image and a variable initial value set into a deep expansion network for multiple iterative operations; Figure 3 It is a schematic diagram of the network architecture of the UBFormer model of the present invention; Figure 4 Schematic diagram of the network architecture of LWTB of the present invention. DETAILED DESCRIPTION

[0018] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present invention more thorough and complete and to fully convey the scope of the present invention to those skilled in the art.

[0019] Example 1 like Figure 1 As shown, this embodiment provides a method for removing halo from nighttime images, comprising: S1: Constructing prior knowledge, wherein the prior knowledge includes a pixel-level spatial consistency term, a feature-level detail enhancement term, and a feature-level noise suppression term.

[0020] S2: constructing an initialization network and a deep expansion network based on prior knowledge, and using the initialization network to extract a set of initial variable values ​​from the nighttime halo pollution image; wherein, the set of initial variable values ​​includes an initial halo mask image, an initial halo removal image, an initial mapping image, an initial halo-free feature map, and initial constraint variables.

[0021] S3: Inputting the night halo pollution image and the variable initial value set into the deep expansion network for multiple iterative operations, and extracting the final halo-removed image from the night halo pollution image; wherein, the deep expansion network includes multiple proximal networks, and each proximal network participates in one iterative operation.

[0022] The deep expansion network of the present invention includes multiple modules, each of which has a proximal network, and each module performs an iterative operation. Both the initialization network and the proximal network use the UBFormer model, which uses multi-scale features to learn the implicit prior of the image. The UBFormer model includes a residual module and a large-window Transformer module. Night halo-contaminated images refer to images with halo and streak artifacts in night scenes. Removing halo and streak artifacts from night halo-contaminated images can improve image resolution.

[0023] The constructing of prior knowledge includes: The halo layer decomposition model is constructed according to the following formula: ; (1) Among them, I is the night halo pollution image, B is the halo removed image, and F is the halo mask image.

[0024] The pixel-level spatial consistency term uses a halo layer decomposition model. The halo layer decomposition model is based on the principle of image additiveness and decomposes the nighttime halo-contaminated image into a halo-removed image and a halo mask image. That is, the nighttime halo-contaminated image is composed of a halo-removed image and an additive halo mask image. The halo mask image is a noise image superimposed on the halo-removed image. Given a nighttime halo-contaminated image I and a halo mask image F estimated through a network, MAP estimation based on Bayes' theorem can transform the image halo removal problem into an energy function minimization problem. The following formula is used to construct the first optimization function corresponding to the pixel-level spatial consistency term: ; (2) Where f represents the Fibonacci norm, is a Fibonacci norm term, which is used to ensure the fidelity of image data. E1 is a first optimization function, and the Fibonacci norm term is used to ensure the fidelity of data.

[0025] The constructing of prior knowledge includes: The convolution dictionary is used to extract the feature map according to the following formula: ; (3) ; (4) ; (5) Among them, ZI is the feature map of the night halo pollution image, ZB is the feature map without halo, ZF is the feature map of halo mask; D is the convolution dictionary, is the convolution operator; Formula (3)-Formula (5) can be rewritten as: ; (6) ; (7) ; (8) Where D is the convolution dictionary, is the convolution operator, c represents the image channel, n represents the feature channel, and total represents the total number of feature channels. When c=1, it represents the red channel of the image, also known as the R channel; when c=2, it represents the green channel of the image, also known as the G channel; when c=3, it represents the blue channel of the image, also known as the B channel.

[0026] Combine the pixel-level spatial consistency term and the feature-level detail enhancement term, and introduce the first penalty coefficient and the second penalty coefficient , construct a second optimization function including the feature-level detail enhancement term: ; (9) Among them, M is the mapping graph, * is element-by-element multiplication, is the first penalty coefficient, is the second penalty coefficient, is the first regularization coefficient, is the first prior term, and E2 is the second optimization function.

[0027] The convolution dictionary is used to ensure that richer features ZI, ZB, and ZF are extracted from I, B, and F to enhance the learning ability of the model. In addition, the convolution dictionary can reduce information loss in each iterative update by extracting the image into a multi-channel feature map.

[0028] The difference between the convolution of ZB with D and B constrains the relationship between the feature ZB and the image to be restored, B. This maps the learned implicit prior ZB back to the halo-removed image B, thereby enhancing detail information in the halo-removed image. The mapping map M balances the removal of streaking artifacts with the enhancement of detail. When streaking artifacts are detected, M is increased, resulting in a smaller value for ZB, thereby suppressing them. When halo regions are detected, M is decreased, resulting in a larger value for ZB, thereby enhancing detail in the halo region.

[0029] By leveraging the Frobenius norm constraint on the relationship between the halo removal image B, the halo-free feature ZB, and the convolution dictionary D, ZB can learn this prior constraint, i.e., adaptively suppress streak artifacts and enhance halo area details.

[0030] The constructing of prior knowledge includes: The feature-level noise prior is constructed according to the following formula: ; (10) in, is the second regularization coefficient, It is the norm of the element-by-element multiplication of the halo-free feature map and the halo mask feature map; Construct the third optimization function of the feature-level noise suppression term: ; (11) in, is the convolution of the convolution dictionary and the halo mask image, and the norm of the element-wise product of the halo-free feature map, and E3 is the third optimization function.

[0031] By using the feature-level noise suppression term, the halo noise content of the halo-free feature map can be suppressed when the halo mask feature map has a large value. Introducing the constraint variable N and the third penalty coefficient , so that N=ZB, the optimization problem is converted to: ; (12) Among them, I is the night halo pollution image, B is the halo removal image, F is the halo mask image, D is the convolution dictionary, is the convolution operator, ZB is the halo-free feature map, M is the mapping map, * is the element-by-element multiplication, is the first penalty coefficient, is the second penalty coefficient, is the third penalty coefficient. is the first regularization coefficient, is the second regularization coefficient, is the third regularization coefficient, is the first prior term, The proximal operator of the mapping graph M is estimated by the corresponding proximal network. The proximal operator of the mapping graph M is used to represent the first prior term, and the proximal gradient algorithm is used to solve the first prior term. The proximal operator of the constraint variable N is estimated by the corresponding proximal network. The proximal operator of the constraint variable N is used to represent the second prior term, and the proximal gradient algorithm is used to solve the second prior term.

[0032] Constructing an initialization network and a deep expansion network based on prior knowledge involves inputting prior knowledge into the initialization network and the deep expansion network. The different modules of the initialization network are interconnected during the process of extracting the initial value set of variables from nighttime halo-contaminated images. As the deep expansion network iteratively operates on the prior knowledge, the different modules within the deep expansion network are iteratively linked. This operation enables the initialization network and the deep expansion network to form a data processing pathway, thereby guiding the initialization network and the deep expansion network to learn the prior knowledge. The initialization network and the deep expansion network constructed based on prior knowledge can effectively remove halo mask images from nighttime halo-contaminated images, that is, remove halo and streak artifacts from nighttime halo-contaminated images, resulting in the final halo-removed image.

[0033] The method for removing halo from nighttime images provided in this embodiment involves constructing prior knowledge, including a pixel-level spatial consistency term, a feature-level detail enhancement term, and a feature-level noise suppression term. The pixel-level spatial consistency term, based on the principle of image additivity, decomposes the nighttime halo-contaminated image into a halo-removed image and a halo mask image, thereby simplifying computations. The pixel-level spatial consistency term is constrained by a Fibonacci norm term to ensure data fidelity. The feature-level detail enhancement term uses a convolutional dictionary to extract richer features from the nighttime halo-contaminated image, the halo-removed image, and the halo mask image, thereby enhancing the learning capabilities of the initialization network and the deep expansion network. The feature-level detail enhancement term uses the Fibonacci norm to constrain the relationship between the halo-removed image, the halo-free feature map, and the convolutional dictionary D, thereby adaptively suppressing streaking artifacts in the nighttime halo-contaminated image and enhancing details in halo areas. The feature-level noise suppression term uses a y-norm to constrain the halo-free feature map and the halo mask feature map, thereby suppressing the halo noise content of the halo-free feature map by increasing the halo mask feature map. A nighttime halo-contaminated image and prior knowledge are input into an initialization network, which extracts a set of initial variable values ​​from the nighttime halo-contaminated image. This set of variable values ​​includes an initial halo mask image, an initial halo-removed image, an initial mapping graph, an initial halo-free feature map, and initial constraint variables. The nighttime halo-contaminated image and the initial variable value set are then input into a deep unfolding network for multiple iterations to extract the final halo-removed image from the nighttime halo-contaminated image. Each iteration of the deep unfolding network utilizes a proximal network to learn implicit priors. This proximal network combines the advantages of CNNs in local feature modeling with the advantages of Transformers in global information modeling. Leveraging prior knowledge, the present invention utilizes a deep unfolding network with multiple proximal networks to iteratively process the mapping graph, halo-free feature map, constraint variables, and halo-removed image to extract the final halo-removed image. This method better preserves image texture and restores image detail, effectively removing speckle and streak artifacts from the image without increasing application cost.

[0034] Example 2 like Figure 1 As shown, this embodiment provides a method for removing halo from nighttime images. This embodiment is based on Example 1 and describes the differences from Example 1. The method includes: S1: Constructing prior knowledge, wherein the prior knowledge includes a pixel-level spatial consistency term, a feature-level detail enhancement term, and a feature-level noise suppression term.

[0035] S2: constructing an initialization network and a deep expansion network based on prior knowledge, and using the initialization network to extract a set of initial variable values ​​from the night halo pollution image; wherein, the set of initial variable values ​​includes an initial halo mask image, an initial halo removal image, an initial mapping image, an initial halo-free feature map and initial constraint variables.

[0036] S3: Inputting the night halo pollution image and the variable initial value set into the deep expansion network for multiple iterative operations, and extracting the final halo-removed image from the night halo pollution image; wherein, the deep expansion network includes multiple proximal networks, and each proximal network participates in one iterative operation.

[0037] like Figure 2 As shown, the night halo pollution image and the variable initial value set are input into the deep expansion network for multiple iterative operations, including: S31: In the t-th iterative operation, the gradient of the t-th mapping map is calculated, the t-th mapping map and the gradient of the t-th mapping map are input into the t-th proximal network of the deep expansion network, and the t+1-th mapping map is calculated.

[0038] S32: Convolve the convolution dictionary with the halo mask image to obtain the tth halo mask feature; calculate the gradient of the tth halo-free feature map, input the tth halo-free feature map, the gradient of the tth halo-free feature map and the tth halo mask feature into the soft threshold operator to obtain the t+1th halo-free feature map.

[0039] S33: Calculate the gradient of the t-th constraint variable, input the t-th constraint variable and the gradient of the t-th constraint variable into the t-th proximal network of the deep expansion network, and calculate the t+1-th constraint variable.

[0040] S34: Calculate the gradient of the t-th halo-removed image, and calculate the t+1-th halo-removed image according to the gradients of the t-th halo image and the t-th halo-removed image.

[0041] Formulas (13)-(16) are used to iteratively update the mapping image, halo-free feature map, constraint variables, and halo-removed image: ; Among them, M t+1 is the t+1th mapping graph, M t is the tth mapping graph, is the gradient of the t-th map, Update parameters for the mapping graph, * is element-by-element multiplication, prox t1 is the t1th proximal network of the map of the deep unfolded network, t1=t. ZB t+1is the t+1th halo-free feature map, soft is the soft threshold operator, ZB t is the t-th halo-free feature map, Update parameters for the halo-free feature map, is the gradient of the t-th halo-free feature map, is the second regularization coefficient, is the first penalty coefficient, is the second penalty coefficient, D t is the t-th convolution dictionary, F is the halo mask image, is the convolution operator, and * is the element-by-element multiplication. N t+1 is the t+1th constraint variable, prox t2 is the t2th proximal network of the constraint variable of the deep expansion network, t2=t, N t is the t-th constraint variable, is the update parameter of the constraint variable, is the gradient of the t-th constraint variable, and * is element-wise multiplication. t+1 For the t+1th halo removed image, B t For the t-th halo removed image, Update parameters for halo removal image, Remove the gradient of the image for the t-th halo, * is element-wise multiplication.

[0042] t1 and t2 are numerically equal to t. In the process of the deep expansion network performing the tth iteration of the prior knowledge, there are two proximal networks, one of which is used to calculate M t+1 , another proximal network is used to calculate N t+1 For example, during the fifth iteration of the deep unfolding network on the prior knowledge, the fifth proximal network of the deep unfolding network's mapping graph is used to calculate M 6 , the fifth proximal network of the constraint variable of the deep expansion network is used to calculate N 6 , at this time, t1 and t2 are both equal to 5, t1 represents the 5th proximal network of the mapping graph of the deep unfolding network, and t2 represents the 5th proximal network of the constraint variable of the deep unfolding network.

[0043] The calculation formula of the soft operator is as follows: (17); Where soft(input, r) is the soft operator, sign is the sign function, input is the independent variable of the soft operator, max is the maximum value function, and r is the adjustment parameter of the soft operator. This example uses r = 1 as an example. When iuput > 1, the soft operator grows linearly with a slope of 1; when -1 ≤ input ≤ 1, soft(input) = 0; when input < -1, the soft operator grows linearly with a slope of 1.

[0044] Calculate using formula (18)-formula (21) 、 、 and : ; in, is the third penalty coefficient, I is the night halo pollution image, B is the halo removed image, and F is the halo mask image.

[0045] In formulas (13)-(16) and (18)-(21), t is the number of iterations, 0≤t≤num, num+1 is the total number of iterations, num≥2. When t=0, M 0 is the initial mapping, ZB 0 is the initial halo-free feature map, N 0 is the initial constraint variable, B 0 is the initial halo-removed image. Substitute the initial mapping image, initial halo-free feature map, initial constraint variables, initial halo-removed image and initial convolution dictionary into formula (18)-formula (21) to calculate 、 、 and , and then 、 、 and Substitute into formula (13)-formula (16) and calculate 、 、 and , repeat the above process until num+1 iterations are completed to obtain the final halo-removed image. The mapping graph, constraint variables, and halo-removed image are all optimized using the proximal gradient descent algorithm, while the halo-free feature map is optimized using the proximal gradient descent algorithm and the iterative shrinkage threshold algorithm.

[0046] The UBformer model is a four-layer U-shaped network that introduces multi-scale features to learn implicit prior information. The UBformer model includes a downsampling module, an upsampling module, a residual module, and a large window transformer module (LWTB).

[0047] like Figure 3 As shown in the figure, the overall framework of the UBFormer model is based on the encoder-decoder structure design. Given an input feature map , where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of feature channels. In the backbone network, this paper uses 1×1 convolution to extract features to reduce the amount of computation when the number of feature channels C is large.

[0048] For the downsampling and upsampling modules, this paper uses pixel reorganization and pixel shuffling operations. For the residual module, this paper replaces the 3×3 convolution module of U-Net with a 3×3 depthwise separable convolution (DWConv) and a 1×1 convolution (Conv) to reduce the amount of computation and parameters. It also replaces the ReLU activation function with the LeakyReLU activation function to preserve information corresponding to negative values.

[0049] At the minimum-scale feature layer, the present invention designs LWTB, which uses the Transformer model to improve the learning ability of the network and uses a larger window to improve the global receptive field of the model. Due to the high degree of parallelism of the GPU, the present invention finds that larger windows do not bring excessive delays and computational costs compared to small windows. While balancing speed and accuracy, this embodiment sets the window size to 32×32. Although the window Transformer uses multiple heads to increase the diversity of features, the features are still relatively similar. Therefore, the present invention adds a spatial differentiation feature to each feature of the group after the self-attention mechanism to increase the spatial differences between features, and introduces a gating mechanism to enhance the nonlinear expression ability of the network.

[0050] At the minimum feature scale layer, the present invention adopts LWTB, Figure 4 This is a schematic diagram of the network architecture of LWTB. The window attention adopted by LWTB is different from that of Uformer. The present invention does not use relative position encoding because relative position encoding will occupy more video memory resources in the case of larger windows and has limited performance gains.

[0051] SABlock is based on the window-based Transformer mechanism. The present invention designs DNet (Differentiated Network) to solve the low-rank problem of the self-attention mechanism, and designs a Conv module to enhance the nonlinear expression ability of the model.

[0052] The initial feature map X0 is transformed into the first feature map X1 through the window transformation. The first feature map X1 is transformed into the second feature map X2 through the SABlock module. The second feature map X2 is transformed into the third feature map X3 through the FFN module. The second and third feature maps are calculated using the following formula: ;(twenty two) ;(twenty three) In the self-attention block, the present invention first uses a Conv module to encode the query matrix Q, key matrix K, value matrix V, and attention score matrix A. The encoding network for these four feature matrices does not share weights. The Conv module consists of a 3×3 depthwise convolution, a ReLU activation function, and a 1×1 convolution to encode spatial and channel information. Because the self-attention mechanism uses multiple feature sharing heads, the present invention introduces a DNet to enhance the diversity of features after self-attention and improve the model's expressiveness. The DNet consists of a 3×3 depthwise convolution, a ReLU activation function, and a 3×3 depthwise convolution. This increases feature diversity while reducing parameter redundancy and computational complexity while maintaining a lightweight architecture.

[0053] The fourth characteristic map is calculated using the following formula: ;(twenty four) Among them, X4 is the fourth feature map, softmax is the normalized exponential function, Q is the query matrix, K T is the transposed matrix of the key matrix, d k is the dimension of the key matrix, V is the value matrix, and A is the attention score matrix.

[0054] This embodiment utilizes a proximal network (UBFormer) to learn implicit priors. UBFormer comprises a residual module and a large-window Transformer module. The residual module employs depthwise separable convolution and pointwise convolution to efficiently capture information over a large receptive field, while preventing gradient vanishing through skip connections. The large-window Transformer module comprises a self-attention block and a feedforward neural network. Because the self-attention mechanism utilizes multiple feature-sharing heads, this invention introduces a DNet to enhance the diversity of features after self-attention, thereby improving the model's expressiveness.

[0055] Example 3 This embodiment provides a method for removing halo from nighttime images. This embodiment is based on Example 1 and describes the differences from Example 1. The method includes: S1: Constructing prior knowledge, wherein the prior knowledge includes a pixel-level spatial consistency term, a feature-level detail enhancement term, and a feature-level noise suppression term.

[0056] S2: constructing an initialization network and a deep expansion network based on prior knowledge, and using the initialization network to extract a set of initial variable values ​​from the night halo pollution image; wherein, the set of initial variable values ​​includes an initial halo mask image, an initial halo removal image, an initial mapping image, an initial halo-free feature map and initial constraint variables.

[0057] S3: Inputting the night halo pollution image and the variable initial value set into the deep expansion network for multiple iterative operations, and extracting the final halo-removed image from the night halo pollution image; wherein, the deep expansion network includes multiple proximal networks, and each proximal network participates in one iterative operation.

[0058] Before extracting the variable initial value set from the night halo pollution image using the initialization network, the method further includes: S21': Calculate the Charbonnier consistency loss between the halo removed image and the true value to obtain the first loss function.

[0059] S22': Calculate the Charbonnier consistency loss between the halo mask image and the training set to obtain the second loss function.

[0060] S23': Add the first loss function and the second loss function to obtain a third loss function.

[0061] S24′: Calculate the Vgg consistency loss between the halo removed image and the true value to obtain a fourth loss function.

[0062] S25': Calculate the Vgg consistency loss between the halo mask image and the training set to obtain the fifth loss function.

[0063] S26′: Add the fourth loss function and the fifth loss function to obtain a sixth loss function.

[0064] S27′: performing weighted summation on the third loss function and the sixth loss function to obtain a final loss function.

[0065] S28′: Using the final loss function, perform end-to-end training on the initialized network and the deep expanded network simultaneously.

[0066] The first loss function is calculated according to the following formula: ; (25) in, is the first loss function, B is the halo removed image, gt is the true value, The square of the two-norm of the difference between the halo-removed image and the true value, Tune parameters for the loss function.

[0067] The second loss function is calculated according to the following formula: ; (26) in, is the second loss function, F is the halo mask image, flare is the training set, It is the square of the two-norm of the difference between the halo mask image and the training set.

[0068] The third loss function is calculated according to the following formula: ; (27) in, is the third loss function, is the first loss function, is the second loss function, B is the halo removed image, gt is the true value, F is the halo mask image, and flare is the training set.

[0069] The fourth loss function is calculated according to the following formula: ; (28) in, is the fourth loss function, VggNet(B) is the output of the Vgg network when the input is B, VggNet(gt) is the output of the Vgg network when the input is gt, It is the norm of the difference between the output of the Vgg network when the input is B and the output of the Vgg network when the input is gt.

[0070] The fifth loss function is calculated according to the following formula: ; (29) in, is the fifth loss function, VggNet(F) is the output of the Vgg network when the input is F, VggNet(flare) is the output of the Vgg network when the input is flare, It is the norm of the difference between the output of the Vgg network when the input is F and the output of the Vgg network when the input is flare.

[0071] The sixth loss function is calculated according to the following formula: ; (30) in, is the sixth loss function, is the fourth loss function, is the fifth loss function.

[0072] The final loss function is calculated according to the following formula: ; (31) Among them, l final is the final loss function, w1 is the first adjustment weight, w2 is the second adjustment weight, and this embodiment takes w1 and w2 as 1 as an example.

[0073] The final loss function is used to simultaneously train the initialization network and the deep expansion network end-to-end to optimize the network parameters of the initialization network and the deep expansion network. The night halo removal model of this embodiment is trained on the Flare7K++ dataset, which includes the synthetic dataset Flare7K and the real dataset Flare-R. Flare7K provides 5,000 synthetic halo images, and Flare-R has 962 real halo images. This embodiment uses the synthesis process of Flare7K++ to generate halo images and corresponding halo-removed images. The background image is 23,949 images randomly selected from the 24K Flickr dataset, and the halo image and light source are randomly selected from Flare7K and Flare-R, respectively, accounting for 50% each. The enhanced halo image in the Flare7K++ dataset is obtained by gamma correction and image enhancement of the synthetic halo image in Flare7K and the real halo image in Flare-R. Image enhancement includes rotation, translation, shearing, scaling, Gaussian blur and random flipping. For each background image, the RGB values ​​are multiplied by a random factor and Gaussian noise is added. The result is normalized to [0, 1] to serve as the true value of the corresponding halo-free image. This embodiment subtracts the light source from the enhanced halo image to generate a halo-removed image. The enhanced halo image is combined with the background image to generate a night halo-contaminated image, which is used as input data for initializing the network. During testing, this embodiment uses the Flare7K test set, which includes 100 pairs of real images, some of which are night halo-contaminated images and others are halo-removed images.

[0074] This embodiment uses an end-to-end training method to train the initialization network and the deep expansion network at the same time, uses the Adam optimizer for multiple iterations of training, and the learning rate is set to 2e -4. This embodiment compares the halo mask image learned by the initialization network with the halo mask image in the Flare7K++ synthesis process, and uses the second loss function and the fifth loss function to train the initialization network and adjust the parameters of the initialization network. This embodiment uses the initial halo mask image extracted by the initialization network as a priori image for the deep expansion network. This embodiment uses the deep expansion network to iterate the halo removal image, mapping map, constraint variables and halo-free feature map for multiple times, and uses the first loss function and the fourth loss function to train the deep expansion network and adjust the parameters of the deep expansion network. This embodiment uses evaluation indicators such as PSNR, SSIM, and LPIPS to evaluate the image quality of the final halo removal image, and determines whether to continue or stop training the initialization network and the deep expansion network based on the evaluation results and the total number of iterations num.

[0075] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0076] The above description is only a preferred embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for removing halo from nighttime images, characterized in that: include: Constructing prior knowledge, the prior knowledge including a pixel-level spatial consistency term, a feature-level detail enhancement term, and a feature-level noise suppression term; An initialization network and a depth expansion network are constructed based on prior knowledge, and a set of initial variable values ​​is extracted from the nighttime halo-contaminated image using the initialization network; wherein the set of initial variable values ​​includes an initial halo mask image, an initial halo-removed image, an initial mapping image, an initial halo-free feature map, and initial constraint variables; The night halo pollution image and the variable initial value set are input into a deep expansion network for multiple iterative operations, and a final halo-removed image is extracted from the night halo pollution image; wherein the deep expansion network includes multiple proximal networks, and each proximal network participates in one iterative operation.

2. The method for removing halo from nighttime images according to claim 1, characterized in that: The constructing of prior knowledge includes: The halo layer decomposition model is constructed according to the following formula: ; Among them, I is the night halo pollution image, B is the halo removal image, and F is the halo mask image; Construct a first optimization function corresponding to the pixel-level spatial consistency term: ; Where f represents the Fibonacci norm, is a Fibonacci norm term, which is used to ensure the fidelity of image data. E1 is a first optimization function.

3. The method for removing halo from nighttime images according to claim 2, wherein: The constructing of prior knowledge includes: The convolution dictionary is used to extract the feature map according to the following formula: ; ; ; Among them, ZI is the feature map of the night halo pollution image, ZB is the feature map without halo, ZF is the feature map of halo mask; D is the convolution dictionary, is the convolution operator; Construct a second optimization function including the feature-level detail enhancement term: ; Among them, M is the mapping graph, * is element-by-element multiplication, is the first penalty coefficient, is the second penalty coefficient, is the first regularization coefficient, is the first prior term, and E2 is the second optimization function.

4. The method for removing halo from nighttime images according to claim 3, wherein: The constructing of prior knowledge includes: The feature-level noise prior is constructed according to the following formula: ; in, is the second regularization coefficient, It is the norm of the element-by-element multiplication of the halo-free feature map and the halo mask feature map; Construct the third optimization function of the feature-level noise suppression term: ; in, is the convolution of the convolution dictionary and the halo mask image, and the norm of the element-wise product of the halo-free feature map, and E3 is the third optimization function.

5. The method for removing halo from nighttime images according to claim 4, characterized in that: Before extracting the variable initial value set from the night halo pollution image using the initialization network, the method further includes: Calculate the Charbonnier consistency loss between the halo-removed image and the true value to obtain the first loss function; Calculate the Charbonnier consistency loss between the halo mask image and the training set to obtain the second loss function; Add the first loss function and the second loss function to obtain the third loss function; Calculate the Vgg consistency loss between the halo removed image and the true value to obtain the fourth loss function; Calculate the Vgg consistency loss between the halo mask image and the training set to obtain the fifth loss function; Add the fourth loss function and the fifth loss function to obtain the sixth loss function; Performing a weighted summation on the third loss function and the sixth loss function to obtain a final loss function; The final loss function is used to train the initialized network and the deeply expanded network end-to-end simultaneously.

6. The method for removing halo from nighttime images according to claim 1, wherein: The step of inputting the nighttime halo pollution image and the variable initial value set into a deep expansion network for multiple iterative operations includes: At the t-th iteration, the gradient of the t-th map is calculated, and the t-th map and the gradient of the t-th map are input into the t-th proximal network of the depth expansion network to calculate the t+1-th map; Convolve the convolution dictionary with the halo mask image to obtain the tth halo mask feature; calculate the gradient of the tth halo-free feature map, input the tth halo-free feature map, the gradient of the tth halo-free feature map, and the tth halo mask feature into the soft threshold operator to obtain the t+1th halo-free feature map; Calculate the gradient of the t-th constraint variable, input the t-th constraint variable and the gradient of the t-th constraint variable into the t-th proximal network of the depth expansion network, and calculate the t+1-th constraint variable; The gradient of the t-th halo-removed image is calculated, and the t+1-th halo-removed image is calculated according to the gradient of the t-th halo image and the t-th halo-removed image.

7. The method for removing halo from nighttime images according to claim 6, wherein: The step of inputting the t-th mapping image and the gradient of the t-th mapping image into the t-th proximal network of the depth expansion network to calculate the t+1-th mapping image comprises: The t+1th mapping graph is calculated according to the following formula: ; Among them, M t+1 is the t+1th mapping graph, M t is the tth mapping graph, is the gradient of the t-th map, Update parameters for the mapping graph, * is element-by-element multiplication, prox t1 is the t1th proximal network of the map of the deeply unfolded network.

8. The method for removing halo from nighttime images according to claim 6, wherein: The step of inputting the t-th halo-free feature map, the gradient of the t-th halo-free feature map, and the t-th halo mask feature into a soft threshold operator to obtain the t+1-th halo-free feature map comprises: The t+1th halo-free feature map is calculated according to the following formula: ; Among them, ZB t+1 is the t+1th halo-free feature map, soft is the soft threshold operator, ZB t is the t-th halo-free feature map, Update parameters for the halo-free feature map, is the gradient of the t-th halo-free feature map, is the second regularization coefficient, is the first penalty coefficient, is the second penalty coefficient, D t is the t-th convolution dictionary, F is the halo mask image, is the convolution operator and * is the element-wise multiplication.

9. The method for removing halo from nighttime images according to claim 6, wherein: The calculation of the t+1th constraint variable includes: The t+1th constraint variable is calculated according to the following formula: ; Among them, N t+1 is the t+1th constraint variable, prox t2 is the t2th proximal network of the constraint variable of the deep expansion network, N t is the t-th constraint variable, is the update parameter of the constraint variable, is the gradient of the t-th constraint variable, and * is element-wise multiplication.

10. The method for removing halo from nighttime images according to claim 6, wherein: The calculating of the t+1th halo-removed image includes: The t+1th halo-removed image is calculated according to the following formula: ; Among them, B t+1 For the t+1th halo removed image, B t For the t-th halo removed image, Update parameters for halo removal image, Remove the gradient of the image for the t-th halo, * is element-wise multiplication.

Citation Information

Patent Citations

  • Single image defogging method based on edge and hue fidelity joint constraint optimization

    CN114936973A

  • Image denoising method based on convolutional dictionary learning and deep expansion

    CN118537569A

  • Underwater halo image sharpening method based on radial gradient iterative network

    CN119180760A

  • Method for detecting infrared ship target based on improved yolov7

    US20250078541A1