Optimization Method for Nighttime Light Images Based on Spatial Information Fusion and Multi-Attention Network
Through the night light image optimization method based on spatial information fusion multi-attention network, the generation adversarial network and attention mechanism are used to remove night light interference and generate pseudo-day images, the problems of night image blur and light interference are solved, and the image quality and accuracy of detail reconstruction are improved.
Patent Information
- Application Number
- CN202210306335.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-27
AI Technical Summary
The prior art is difficult to effectively remove night light interference and noise information, resulting in blurred night images and interfering with the details of areas with strong light, and lack of automatic enhancement methods.
The night light image optimization method based on spatial information fusion multi-attention network is adopted. By generating an adversarial network, the spatial attention and self-attention mechanism are used to identify feature areas with high brightness, remove light interference, and generate pseudo-daytime images.
Effectively remove night light interference, restore original scene information, and generate more refined pseudo-daytime images, improving the effectiveness of detail reconstruction and the accuracy of scene recognition.
Smart Images

Figure CN114926349B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to the technical fields of an enhancement and optimization method for images interfered by night lights and a method for generating pseudo-daytime, specifically an optimization method for night light images based on a multi-attention network that fuses spatial information. Background Art
[0002] From the perspective of requirements: With the development of social economy, human nocturnal activities have become increasingly rich. Many activities and researches of people are difficult to carry out due to reasons such as line of sight and night photography effects. Cameras are also difficult to capture key information hidden in the darkness due to light reasons. Moreover, randomly captured videos and images are often interfered by strong light, which also causes the loss of feature information. The launch of wars, the march and sneak attacks of the military mostly occur at night, and most lawbreakers also choose to carry out illegal activities under the cover of night. In the field of public security, by restoring night images to images similar to daytime, it is possible to achieve reconnaissance and monitoring without difference between night and day, resist the malicious interference of suspects with strong light on cameras, and at the same time help restore dim dead corners into clearly visible images, restore a clear scene for case handlers, quickly find important information such as the colors of people's clothes and vehicles, and provide assistance for case detection.
[0003] From the perspective of the current technical situation: Modern camera devices collect light information reflected from the surface of an object through an image sensor to obtain digital image information. Therefore, the imaging quality of camera devices is affected by the light source distribution around the target. The prior knowledge obtained from observational statistics in traditional image processing lacks universality for night images.
[0004] With the development of deep learning technology, generative models are a more effective method for restoring missing information based on learned distributions. However, due to the weak generalization ability of traditional generative adversarial models for actual data, such models cannot fully restore image information. Summary of the Invention
[0005] To solve the above technical problems, in the case where it is impossible to increase the brightness of the surrounding environment and impossible to change the imaging principle of the sensor, in order to offset the light interference and noise information at night, the present invention proposes an optimization method for night light images based on a multi-attention network that fuses spatial information, enhances and optimizes the captured night light videos and images, and finally obtains pseudo-daytime images.
[0006] This method first trains and tests on a data set to obtain an accurate image pseudo-daytime generation model, and then uses the model to process night light videos and images and outputs pseudo-daytime images.
[0007] This method uses the spatial attention (SPA) mechanism and the self-attention mechanism to identify and focus on the feature regions with high brightness, enhancing the image information.
[0008] Based on the daytime image information, the self-attention non-local neural network is used to guide the generation of the global detail information of the image.
[0009] A generative adversarial model is constructed by the spatial attention recursive network and the self-attention residual network. In the generator part, the spatial context information is used to extract features and generate an attention map, so as to remove the interference of night lights and enhance the corresponding dark areas. And gradient penalty is adopted to enhance the robustness of the model, and the discriminator uses the self-attention mechanism to guide the output of the generator.
[0010] The present invention solves the problem that there is a lack of a method for automatically enhancing night light images in the prior art. Since the light images taken at night are often affected by light, the images in the dark areas are blurred, and in the areas with strong light, the image details are interfered by the strong light.
[0011] This method can be widely applied to various night activity fields such as night security reconnaissance, natural night biological activity observation, and night UAV observation.
[0012] The specific technical solution of the present invention is: first, construct an image pseudo-daytime generation model based on the multi-attention network for spatial information fusion; then train the image pseudo-daytime generation model; finally, use the trained image pseudo-daytime generation model to process the captured night light videos and images, and output pseudo-daytime images.
[0013] Construct a network model: adopt the overall framework of the generative adversarial network. This generative adversarial network is mainly divided into a generator network and
[0014] Input the blurred night light image interfered by light into the generator network, and use the spatial attention recursive network model to process the image to remove the strong light interference and reconstruct all the features of the blurred dark light.
[0015] The discriminator network uses the optimized image reconstructed by the spatial attention recursive network model to make judgment and prediction, then uses the self-attention mechanism to capture the probability distribution of strong light, and focuses the attention on the meaningful regions.
[0016] Construct a loss function: by constructing a hybrid loss function, use the feature map of strong light information as the supervision signal, and optimize the training process by means of gradient penalty;
[0017] According to the idea of the attention map generated by the generator spatial model guided by the self-attention map of the discriminator network, construct a multi-attention loss cost function, so that the generator pays attention to more detailed local blurred information, and as much as possible makes the generated network reconstruct samples similar to the real image data distribution.
[0018] Network model training:
[0019] First, collect night-time light image data and daytime clear image data at the same location. Use the daytime clear images as training labels to train the generator network and the discriminator network; first binarize the data containing night-time blurred data and clean daytime data to obtain a label map with only strong light information and a pseudo-blurred layer;
[0020] Input the data into the initialized spatial information fusion multi-attention network (generator network) to output an initially optimized pseudo-clear image;
[0021] Perform complex geometric constraints on the global image structure through the self-attention layer mechanism of the discriminator network to output the attention feature distribution;
[0022] Subsequently, under the guidance of the loss function, the spatial information fusion multi-attention network detects and removes strong light interference in the next iteration.
[0023] Finally, train to obtain a spatial information fusion multi-attention network (generator network) that can fully enhance and optimize night-time light image information.
[0024] Spatial attention uses the SPA Net model; it uses horizontal / vertical neighborhood information in space to model missing information. It is a two-round four-direction (up, down, left, right) recurrent neural network composed of ReLU and a recurrent neural network.
[0025] The discriminator network is composed of a residual network and a self-attention network:
[0026] First, extract features through a three-layer convolutional network with a normalization layer and a LeakyReLU layer;
[0027] Finally, output the result by stacking two layers of self-attention networks and one layer of convolutional layer. Different from other methods, the present invention uses the correlation coefficient after softmax conversion in the last self-attention network as the correlation judgment between strong light interference and the dark mask information, and extracts the feature layer output in the self-attention layer. Finally, the discriminator outputs a global feature map, a self-attention feature map, and an interference feature correlation matrix D i,j .
[0028] The adversarial and attention hybrid loss function is used to optimize the training process to enable the generator network to better reconstruct the information of real photographed objects. The generator network hybrid loss function (i.e., the adversarial and attention hybrid loss function) is based on the adversarial loss function L GAN1Based on (G, D), it is designed by combining the L1 loss function and the attention loss function. The discriminator network loss function applies gradient penalty to the discriminator network, constructs a loss function using the correlation matrix of light factors extracted from the discriminator network, and combines it with the attention feature map output by the self-attention layer.
[0029] The beneficial effects of the present invention are as follows: Based on the generative adversarial network, the present invention proposes an optimization method for nighttime light images based on a multi-attention network with spatial information fusion, which restores the original scene information by removing light interference and dark mask occlusion, and generates more refined pseudo-daytime images. The present invention solves the problem of weak generalization ability of the existing technology for actual data, improves the ability of the model to remove interference from nighttime lights and generate pseudo-daytime, further improves the effectiveness of detail reconstruction and the accuracy of scene recognition, and obtains a stable model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic diagram of the process of designing the image pseudo-daytime generation model of the present invention;
[0031] Figure 2 It is a schematic diagram of the generator network of the present invention;
[0032] Figure 3 It is a schematic diagram of the discriminator network of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention and the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0034] The present invention is based on the generative adversarial network (GAN). The generator uses the ideas of the spatial attention network (SPA Net) and the residual network, and the discriminator uses the idea of the self-attention mechanism to construct a generative adversarial network that fuses spatial attention and self-attention to remove strong light interference and dark background. And an image pseudo-daytime generation model obtained by training the generative adversarial network is used as a nighttime light image optimization model for nighttime light videos and images.
[0035] The generative adversarial network is trained and tested on the public dataset and the collected dataset, and directly tested with actual real environment data to obtain better strong light removal effects and pseudo-daytime generation effects, greatly restoring the original daytime scene information.
[0036] Specifically, the generative adversarial network that fuses spatial attention and self-attention of the present invention is composed of a generator network and a discriminator network. As Figure 1 shown, the generator network removes interference from the input image, restores scene information, and generates a pseudo-daytime image, and the discriminator network discriminates and guides the image.
[0037] The training set inputs 1000 512*512*3 images containing strong light interference and natural night images. First, a binary image of the interference and the dark mask is calculated, and the night image and the day image are combined and input into the network for training at the same time.
[0038] The total number of epochs is set to 300 times. First, it is trained through the generator network to initially generate a pseudo-daytime image. As Figure 2 shown, in this step, the input image is first passed through three residual modules composed of convolutional neural networks to extract eigenvalues; subsequently, it passes through four spatial attention modules to remove strong light interference (this spatial attention module is composed of three spatial attention residual blocks and a spatial attention network); finally, it passes through two residual networks and a convolutional block to output a clear pseudo-daytime image and an attention map.
[0039] The discriminator network includes four convolutional blocks and two self-attention blocks. As Figure 3 shown, the first half is composed of three convolutional blocks to form a residual structure, and features are extracted through three convolutional blocks; finally, through two self-attention layers and a convolutional layer, a self-attention map and a judgment result are output. The use of self-attention layers makes the judgment of the discriminator network on the generator network more rigorous, and at the same time helps to distinguish and judge interference data and surrounding dim scene information.
[0040] In order to make the generated pictures as similar as possible to the training pictures, according to the network characteristics of the generator network and the discriminator network, an adversarial and attention hybrid loss function is constructed on the basis of the generative adversarial network loss function, so that the generator network can better reconstruct real scene information.
[0041] The generator loss function of the hybrid loss function is
[0042] Based on the adversarial loss function L GAN1 (G,D), it is designed by combining the L1 loss function and the attention loss function. This adversarial and attention hybrid loss function can be expressed as:
[0043]
[0044] The adversarial loss function is expressed as:
[0045]
[0046] Here, x is the input image containing strong light interference and dark masks, E is the mathematical expectation, and D(x, G(x)) represents the probability that the discriminator judges the reconstructed clear image as the reference image. G(x) is the prediction result of the generator to remove interference and noise. After several rounds of iterative training, the distributions of the real data x and the fake data G(x) generated by the generator are similar, denoted as p data (x).
[0047] The L1 loss function described above is used to measure the accuracy of each reconstructed pixel. This function is expressed as;
[0048]
[0049] R is the real daytime scene image, and λ c is the contribution of the weight of each channel to the loss. C, H, and W represent the number of channels, image height, and width respectively.
[0050] The attention loss function is expressed as:
[0051]
[0052] Matrix G Att is the two-dimensional attention feature generated by the spatial attention model, and matrix M is the binary image label of the strong light interference and dark mask areas, which is obtained through difference calculation.
[0053] To better guide the training of the generator, gradient penalty is applied to the model discriminator network. A loss function is constructed using the correlation matrix of light factors extracted from the discriminator network, and combined with the attention feature map output by the self-attention layer. The final discriminator network loss function can be expressed as:
[0054]
[0055] Among them, λ is the gradient penalty coefficient, represents interpolation on the real daytime image and the pseudo-daytime clear image data generated by the generator, and ε ∼ U[0, 1]. N is the amount of data in the training. D i,j represents the correlation between the model and region i when judging region j. D SA represents the final output of the self-correlation layer feature map of the discriminator.
Claims
1. A method for optimizing night light images based on a multi-attention network with spatial information fusion, characterized by the steps including: First, construct an image pseudo-daytime generation model based on a multi-attention network with spatial information fusion; Then, train the image pseudo-daytime generation model; Finally, use the trained image pseudo-daytime generation model to process the captured night light videos and images, and output pseudo-daytime images; The steps for designing the image pseudo-daytime generation model include: 1) First, construct an image pseudo-daytime generation model based on a multi-attention network with spatial information fusion; 2) Then, train the image pseudo-daytime generation model; In step 1), adopt the overall framework of the generative adversarial network to construct a generative adversarial network that fuses spatial attention and self-attention; this generative adversarial network includes a generator network and a discriminator network; Input the blurred night lights interfered by light into the generator network; In the generator network, use the spatial attention recurrent network to process the image to remove strong light interference and reconstruct all the features of the blurred and dim light; In the discriminator network, perform judgment and prediction on the optimized image reconstructed by the spatial attention recurrent network, use the self-attention mechanism to capture the probability distribution of strong light, and focus the attention on the meaningful areas; In step 2): First, construct a generator network loss function and a discriminator network loss function: The generator network loss function is to construct a hybrid loss function by combining the L1 loss function and the attention loss function on the basis of the adversarial loss function; The discriminator network loss function is to apply gradient penalty to the discriminator network; construct a loss function using the correlation matrix of light factors extracted in the discriminator network, and combine the attention feature map output by the self-attention layer to finally obtain the discriminator network loss function; Then, perform model training, and the steps include: 2.1) Collect night light images and clear daytime images at the same location, use the clear daytime images as training labels, and train the generator network and the discriminator network; First, perform binarization processing on the night blurred data and clean daytime data to obtain a label map with only strong light information and a pseudo-blurred layer; then input it into the initialized multi-attention network with spatial information fusion, that is, the generator network, to output an initial optimized pseudo-clear image; then perform complex geometric constraints on the global image structure through the self-attention layer mechanism of the discriminator network to output the attention feature distribution; 2.2) Guide the generator network to detect and remove strong light interference in the next iteration through the hybrid loss function; finally, train to obtain the multi-attention network with spatial information fusion that finally fully enhances and optimizes the night light image information, that is, the generator network.
2. The method for optimizing night light images based on a multi-attention network with spatial information fusion according to claim 1, characterized in that The generator network uses three feature compression and excitation residual networks to initially extract features, constructs four spatial attention modules to gradually identify and remove light interference and dark masks in four stages, and two residual blocks to reconstruct a clean background.
3. The method for optimizing night light images based on a multi-attention network with spatial information fusion according to claim 1, characterized by the space Note that the recursive network uses horizontal / vertical neighborhood information in the image space to model the missing information. It is a two-round four-direction recursive neural network composed of ReLU and a recursive neural network. The information of each pixel is further obtained by aggregating the information in four directions of each pixel position in the image. Finally, the model outputs the reconstructed image and the two-dimensional interference attention matrix M(x). M(x) = Wtotal × F(x) + x x is the input feature map, and F(x) is the pixel feature value extracted in four rounds of direction extraction. W total is the total weight value of the four rounds of direction; The four directions refer to the four directions of up, down, left, and right of the pixel position.
4. The method for optimizing night light images based on a multi-attention network integrating spatial information according to claim 1, characterized in that the discriminator network is composed of a residual convolutional network and a self-attention network; first, features are extracted by a three-layer convolutional network of a normalization layer and a LeakyReLU layer; then, two layers of self-attention networks are stacked to output a self-attention feature map and a correlation matrix D i,j , and the self-attention feature map is the final output; Correlation matrix D i,j It is calculated by softmax and extracted from the self-attention layer, and is used to judge the correlation of light factors.
5. The method for optimizing night light images based on a multi-attention network with spatial information fusion according to claim 1,[[]] wherein The night light image is a natural color band with a size of 512*512*3.
6. The method for optimizing night light images based on a multi-attention network with spatial information fusion according to claim 1,[[]] wherein The construction method of the hybrid loss function of the generator network is as follows: Based on the adversarial loss function \(L\) GAN1 (G, D), the hybrid loss function designed by combining the L1 loss function and the attention loss function is expressed as: wherein: a. The adversarial loss function is expressed as: x is an image with strong light interference and a dark mask, E is the mathematical expectation, and D(x, G(x)) represents the probability that the discriminator network judges the reconstructed clear image as the reference image; G(x) is the prediction result of the generator network for removing interference and noise; after multiple rounds of iterative training, the distributions of the real data x and the fake data G(x) generated by the generator network are similar, denoted as p data (x); b. The L1 loss function is used to measure the accuracy of each reconstructed pixel, and it is expressed as: R is the real daytime scene image, and λ c is the contribution of the weight of each channel to the loss, where C, H, and W represent the number of channels, the image height, and the width, respectively; c. The attention loss function is expressed as: Matrix G Att is a two-dimensional attention feature generated by a spatial attention model, and matrix M is a binary image label of strong light interference and dark mask regions.
Citation Information
Patent Citations
Fine-grained cross-media retrieval method based on generative adversarial network
CN112800249A
Image synthesis method based on dynamic self-attention generative adversarial network
CN113379655A