A Deep Learning-Based Method for Enhancing Details of Aerial Images

By building a cyclic convolutional neural network with memory state, combining a multi-camera fusion model of event cameras and traditional cameras, the problems of low brightness, noise and blur of long-distance aerial images are solved, and visual perception enhancement images with high dynamic range and high texture details are generated, improving the quality and perceptibility of aerial images.

CN119624810BActive Publication Date: 2025-07-08POWERCHINA BEIJING ENG CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411714364.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-07-08
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problems of low brightness, large proportion of noise signals and blurred images in long-distance aerial images, resulting in poor image quality and unable to meet the needs of computer vision tasks.

Method used

Using aerial image detail enhancement method based on deep learning, by constructing a cyclic convolutional neural network with memory state, combining a multi-camera fusion model of event cameras and traditional cameras, image denoising, reconstruction and fusion is performed to generate visual perception enhancement images with high dynamic range information and high texture details.

Benefits of technology

Improve the overall brightness of aerial images, reduce noise signals, enhance image details, improve image perceptibility and quality, and support subsequent visual object analysis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624810B_ABST
    Figure CN119624810B_ABST
Patent Text Reader

Abstract

The present invention provides a method for enhancing the details of aerial images based on deep learning, including: constructing a recurrent convolutional neural network with a memory state; the recurrent convolutional neural network with a memory state includes a denoising network, an event reconstruction image network, and a fusion network; in a low-light scene, through a visual camera and a normal camera, in each acquisition period, a low-quality video frame sequence VL<supgt;*< / supgt; and a low-quality event stream FL<supgt;*< / supgt> of a target area are respectively acquired and input into the trained recurrent convolutional neural network with a memory state, and a reconstructed first high-quality video frame sequence VHR1<supgt;*< / supgt> is output, which is the finally obtained video frame sequence with enhanced image details. The model designed by the present invention simultaneously utilizes the high dynamic range information possessed by the event stream data and the high texture detail information possessed by the image frame sequence data, and can fuse and construct a visually enhanced image with both high dynamic range and high texture detail information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an aerial image detail enhancement method based on deep learning. Background Art

[0002] With the rapid development of modern industry and the wide application of computer technology, the structure of image acquisition devices is becoming more and more precise, the market inventory is increasing continuously, and the use is becoming more and more convenient. At the same time, along with the development of storage devices and Internet technology, the number of digital images has shown an explosive growth worldwide, and the quality of digital images is also getting higher and higher, which is mainly reflected in the fact that the resolution of images is getting larger, the colors are getting more delicate, and the pictures are getting clearer. Therefore, the huge amount of image data that can be easily obtained is sufficient to support various computer vision tasks, which also lays a solid foundation for the development of the computer vision field. However, due to various environmental factors that may occur in the shooting scene, such as long distance, low light, etc., and the inevitable signal errors and noises in the process of capturing images by digital image acquisition devices, the obtained digital images sometimes may be distorted or even unrecognizable. For this reason, in recent years, image enhancement has attracted wide attention in the field of computer vision.

[0003] The existing image enhancement technologies mainly focus on tasks such as reducing the noise signals contained in images, correcting the local blurring existing in images, and increasing the brightness contrast of low-light images. The specific implementation methods can be mainly divided into traditional methods and machine learning methods. For traditional methods, specifically, traditional image enhancement methods such as gray-scale transformation, spatial domain filtering, frequency domain filtering, histogram enhancement, etc., are relatively simple to implement and the effects are relatively fixed, and it is difficult to meet the current requirements for image enhancement. For machine learning methods, specifically including various methods based on deep learning, which are the focus of current academic research.

[0004] Currently, in the field of deep learning, traditional cameras are generally used to shoot scene images and then used in deep learning networks. This method has the following problems:

[0005] When the distance between a traditional camera and a visual target in a scene to be photographed is large, exceeding the concept of "close-up shooting" that people usually consider, the visual target is considered a long-distance visual target at this time. For example, aerial photography is a relatively typical shooting scenario for long-distance visual targets. When a traditional camera is far from the scene to be photographed and the visual target therein, the visual target has a small range in the scene, and the amount of light reflected by the target entering the camera is small, so it is easy to cause the overall brightness of the target and the image to be low. At the same time, the reduction in the amount of light entering the camera will cause the noise signal of the image to occupy a larger proportion of all the information signals of the image, that is, it will result in a lower signal-to-noise ratio. In addition, since the image range included in the long-distance aerial photography scenario is large, if the camera shakes slightly at this time, it will cause the entire image to change significantly in a short time, so it is easier to produce blurring in the case of a longer exposure time.

[0006] Therefore, how to process the aerial photography images captured by a traditional camera and effectively improve the quality of aerial photography images is the key problem to be solved at present. Summary of the Invention

[0007] Aiming at the defects existing in the prior art, the present invention provides a method for enhancing the details of aerial photography images based on deep learning, which can effectively solve the above problems.

[0008] The technical solution adopted by the present invention is as follows:

[0009] The present invention provides a method for enhancing the details of aerial photography images based on deep learning, including the following steps:

[0010] Step S1, obtaining a training sample set; each training sample in the training sample set simultaneously includes a high-quality event stream, a low-quality event stream, a low-quality video frame sequence, and a high-quality video frame sequence that are truly collected for the same area corresponding to the same time interval;

[0011] Wherein: the high-quality event stream and the low-quality event stream are collected by a visual camera; the high-quality video frame sequence and the low-quality video frame sequence are collected by an ordinary camera;

[0012] Step S2, for each training sample, performing data preprocessing to obtain a preprocessed training sample; in the preprocessed training sample, it includes a preprocessed high-quality event stream, a low-quality event stream, a low-quality video frame sequence, and a high-quality video frame sequence, and the high-quality event stream, the low-quality event stream, the low-quality video frame sequence, and the high-quality video frame sequence are respectively represented as: high-quality event stream FH, low-quality event stream FL, low-quality video frame sequence VL, and high-quality video frame sequence VH;

[0013] Step S3, construct a recurrent convolutional neural network with a memory state; the recurrent convolutional neural network with a memory state includes a denoising network, an event reconstruction image network, and a fusion network;

[0014] Step S4, use the preprocessed training samples to train the recurrent convolutional neural network with a memory state to obtain the trained recurrent convolutional neural network with a memory state. The specific training method is as follows:

[0015] Step S4.1, input the low-quality event stream FL into the denoising network, and perform denoising processing by the denoising network to generate an event stream FL1;

[0016] Step S4.2, input the event stream FL1 generated in Step S4.1 into the event reconstruction image network, and the event reconstruction image network outputs a reconstructed video frame sequence VR;

[0017] Step S4.3, input the reconstructed video frame sequence VR obtained in Step S4.2 and the low-quality video frame sequence VL into the fusion network, and the fusion network outputs a reconstructed first high-quality video frame sequence VHR1;

[0018] Step S4.4, input the high-quality event stream FH into the event reconstruction image network, and the event reconstruction image network outputs a reconstructed second high-quality video frame sequence VHR2;

[0019] Step S4.5, calculate the first loss function value between the reconstructed video frame sequence VR and the reconstructed second high-quality video frame sequence VHR2;

[0020] Calculate the second loss function value between the reconstructed video frame sequence VR and the high-quality video frame sequence VH;

[0021] Calculate the third loss function value between the reconstructed first high-quality video frame sequence VHR1 and the high-quality video frame sequence VH;

[0022] Integrate the first loss function value, the second loss function value, and the third loss function value to obtain a comprehensive loss function value;

[0023] Step S4.6, update the network parameters of the recurrent convolutional neural network with a memory state according to the comprehensive loss function value;

[0024] Return to Step S4.1, use the next training sample to continue training the recurrent convolutional neural network with a memory state until the recurrent convolutional neural network with a memory state reaches the training termination condition;

[0025] Step S5, in a low-light scenario, through a vision camera and a normal camera, in each acquisition cycle, respectively acquire a low-quality video frame sequence VL * and a low-quality event stream FL * , input them into the trained recurrent convolutional neural network with a memory state, and output a reconstructed first high-quality video frame sequence VHR1 * , which is the finally obtained video frame sequence with enhanced image details.

[0026] Preferably, in step S2, the preprocessing method for the low-quality video frame sequence and the high-quality video frame sequence is:

[0027] Convert each frame image in the actually acquired low-quality video frame sequence and high-quality video frame sequence into a grayscale image.

[0028] Preferably, convert it into a grayscale image through formula (1):

[0029] I gray = 0.2989×I rgb [r]+0.5870×I rgb [g]+0.1140×I rgb [b] (1)

[0030] where: I gray is the grayscale image, and I rgb [r], I rgb [g] and I rgb [b] are the RGB three-channel image components of each frame image respectively.

[0031] Preferably, in step S2, the preprocessing method for the high-quality event stream and the low-quality event stream is: convert the event stream into the form of a voxel grid.

[0032] Preferably, converting the event stream into the form of a voxel grid is specifically:

[0033] ① Divide the entire event stream into a number of event stream units with a time interval length of ΔT according to the event timestamps;

[0034] ② For each event stream unit, its start time is represented as t0, then the end time is: t0 + ΔT. Suppose there are B timestamps in the time interval length [t0, t0 + ΔT], which are respectively represented as: t0, t1, …, t B-1 ; Therefore, any timestamp is represented as t i , i = 0, 1, …, B - 1;

[0035] ③ For each event stream unit, generate a container for all events corresponding to each of its timestamps t i .

[0036] ④ Using formula (2), the timestamp t i is normalized to obtain the normalized timestamp

[0037]

[0038] ⑤ Convert all events in each event stream unit into a voxel grid E;

[0039] The voxel grid E has dimensions B×H×W, where B is the number of timestamps and also the number of containers; H represents the number of pixels in the length direction, and W represents the number of pixels in the width direction; then for each pixel, its two-dimensional coordinates are (x l , y m ), where l = 0, 1, …, H - 1, m = 0, 1, …, W - 1; its three-dimensional coordinates with time attribute are: (x l , y m , t i ). Using formula (3), obtain its voxel pixel weighted value E(x l , y m , t i ):

[0040]

[0041] Where:

[0042] ω is the pixel value at the three-dimensional coordinates (x l , y m , t i );

[0043] p jki means that the event camera senses the change in the brightness signal of the incident light asynchronously at each pixel position and generates an event when the absolute value of the change in the brightness signal exceeds the threshold CT. It is represented by the quadruple e jki = (x j , y k , p jki , t i ), where (x j , y k ) represents the horizontal and vertical coordinates of the pixel at which the event is triggered, t i represents the timestamp of the event; p jki represents the polarity of the event, p jki ∈{-1, 1}, that is: p jki is -1, representing dimming; p jki is 1, representing brightening; if no event is generated at the corresponding pixel position, then p jkiIs a null value;

[0044] α is a mixing weight coefficient, the magnitude of which is determined by the proportion of events generated in the voxel grid E, and its value ranges from 0 to 1, i.e., α ∈ [0, 1];

[0045] Therefore, after determining the voxel pixel weighted value E(x l , y m , t i ), the voxel grid E is determined.

[0046] Preferably, when calculating the loss function value between each video frame sequence, first regularize the voxel grid E of each video frame sequence, and then calculate the loss using the Euclidean distance for the regularized voxel grid using the Euclidean distance to calculate the loss.

[0047] Preferably, the method for regularizing the voxel grid E is as follows:

[0048] ① For each voxel pixel weighted value E(x l , y m , t i ) in the voxel grid E, use formula (4) to obtain the intermediate variable Mask(x l , y m , t i ):

[0049]

[0050] ② Use formula (5) to sum all the intermediate variables Mask(x l , y m , t i ) obtained in the voxel grid E, and count the number of non-zero voxel pixel weighted values E(x l , y m , t i ), denoted as N;

[0051]

[0052] ③ Use formula (6) to obtain the mean μ:

[0053]

[0054] ④ Use formula (7) to obtain the variance σ:

[0055]

[0056] ⑤ Use formula (8) to obtain the regularization value of the voxel pixel weighted value E(x l , y m , t i )

[0057]

[0058] By regularizing each voxel pixel weight value E(x l , y m , t i ), that is, regularizing the voxel grid E.

[0059] The method for enhancing details of aerial images based on deep learning provided by the present invention has the following advantages:

[0060] 1. The model designed by the present invention simultaneously utilizes the high dynamic range information possessed by event stream data and the high texture detail information possessed by image frame sequence data, and can fuse and construct a visually perceived enhanced image that simultaneously possesses high dynamic range and high texture detail information;

[0061] 2. In the process of converting event stream data into a voxel grid, the contribution of pixel values is simultaneously considered, and the method of the present invention simultaneously utilizes the information of the voxel grid and pixel coordinates through a mixed weight method.

[0062] 3. The method of the present invention has the ability to improve the image quality and enhance the image details to a certain extent, and has application value in the field of aerial photography. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic diagram of the overall model structure of a method for enhancing details of aerial images based on deep learning provided by the present invention;

[0064] Figure 2 It is a schematic diagram of the network structure used in the present invention;

[0065] Figure 3 It is a schematic diagram of the fusion of the reconstructed image and the low-quality image provided by the present invention;

[0066] Figure 4 It is the image enhancement effect of the real data low-light scene provided by the present invention;

[0067] Figure 5 It is the fusion effect of the event reconstruction image and the low-light scene image provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0069] An event camera is a new type of visual sensor, also known as a neuromorphic visual sensor. Inspired by neuroscience, it imitates the working principle of the eye, asynchronously captures the brightness changes of each pixel, and organizes the data in an "event stream" structure. This is fundamentally different from the way traditional cameras expose at a fixed time to capture fixed frame rate light intensity images. Due to the various advantages of event cameras, event cameras have broad development prospects in many application fields such as autonomous driving, robots, and wearable devices. At present, in many scientific research fields of computer vision, event cameras have also produced many related research and application results.

[0070] Event cameras also have their shortcomings. Limited by their imaging principles, the event data collected when the scene changes little is relatively sparse, making it difficult to parse the target. In addition, the current event camera imaging noise is large, the price is expensive, and there is little public data. Research in the field of event cameras still faces many problems. By combining event cameras with traditional cameras and using the obtained multi-camera fusion model to capture high temporal resolution, high dynamic range event streams and low frame rate, high texture detail intensity image sequences in the same scene, the advantages of the two can be complemented to generate high temporal resolution, high dynamic range, and high texture detail intensity images, which can then be used to support subsequent visual target parsing tasks.

[0071] Therefore, the present invention uses a multi-camera fusion model of event cameras and traditional cameras and the heterogeneous visual stream data generated by them to improve the various problems caused by the long distance between the camera and the scene to be photographed, such as improving the overall brightness of the image, reducing the blurred part in the image, and reducing the noise signal in the image signal, thereby improving the perceptibility of long-distance visual targets and providing better support for subsequent visual target parsing tasks.

[0072] The present invention is particularly suitable for the field of aerial remote sensing image data processing technology. It is aimed at the practical application of aerial image processing and takes improving the texture details of aerial images as the main purpose. It proposes an image texture enhancement method optimized for aerial scenes based on deep learning. The method includes two core steps: first, low-quality event denoising and image reconstruction; second, the fusion of the reconstructed image and the actual collected low-quality image. The method aims to utilize the high dynamic range information of the event stream data and the high texture detail information of the actual collected image frame sequence data to fuse and construct a perception-enhanced image with both high dynamic range and high texture detail information. The method of the present invention has a certain ability to improve image quality and enhance image perception, and has application value in the field of aerial photography.

[0073] See also Figure 1 The present invention provides a method for enhancing aerial image details based on deep learning, comprising the following steps:

[0074] Step S1, obtain a training sample set; each training sample in the training sample set simultaneously includes a high-quality event stream, a low-quality event stream, a low-quality video frame sequence, and a high-quality video frame sequence corresponding to the same time interval obtained by true aerial photography of the same region;

[0075] Among them: the high-quality event stream and the low-quality event stream are collected by a vision camera; the high-quality video frame sequence and the low-quality video frame sequence are collected by an ordinary camera;

[0076] Step S2, for each training sample, perform data preprocessing to adapt to the input format of the network model, and obtain a preprocessed training sample;

[0077] In the preprocessed training sample, it includes a preprocessed high-quality event stream, a low-quality event stream, a low-quality video frame sequence, and a high-quality video frame sequence. The high-quality event stream, the low-quality event stream, the low-quality video frame sequence, and the high-quality video frame sequence are respectively represented as: high-quality event stream FH, low-quality event stream FL, low-quality video frame sequence VL, and high-quality video frame sequence VH;

[0078] In step S2, the preprocessing method for the low-quality video frame sequence and the high-quality video frame sequence is:

[0079] Convert each frame image in the actually collected low-quality video frame sequence and high-quality video frame sequence into a grayscale image. Specifically: through formula (1), convert it into a grayscale image:

[0080] I gray = 0.2989×I rgb [r]+0.5870×I rgb [g]+0.1140×I rgb [b] (1)

[0081] Among them: I gray is the grayscale image, I rgb [r], I rgb [g] and I rgb [b] are respectively the RGB three-channel image components of each frame image.

[0082] In step S2, the preprocessing method for the high-quality event stream and the low-quality event stream is: convert the event stream into the form of a voxel grid (Voxel Grid); converting the event stream into the form of a voxel grid is specifically:

[0083] A voxel is analogous to a pixel and is the smallest unit in three-dimensional space. A voxel grid can be analogous to an image plane composed of pixels and is a three-dimensional cube composed of voxels. The event voxel grid contains three dimensions, two of which are the horizontal and vertical coordinates of the pixel respectively, and the third dimension is set as the container variable (Bin), indicating the degree of temporal subdivision of the voxel grid.

[0084] ① Divide the entire event stream into several event stream units with a time interval length of ΔT according to the event timestamps;

[0085] ② For each event stream unit, its start time is represented as t0, then the end time is: t0 + ΔT. Suppose there are B timestamps in the time interval length [t0, t0 + ΔT], which are respectively represented as: t0, t1, …, t )-1 ; Therefore, any timestamp is represented as t i , i = 0, 1, …, B - 1;

[0086] ③ For each event stream unit, a container is generated for all events corresponding to each timestamp t i ;

[0087] ④ Using formula (2), normalize the timestamp t i to obtain the normalized timestamp

[0088]

[0089] ⑤ Convert all events in each event stream unit into a voxel grid E;

[0090] The size of the voxel grid E is B × H × W, where B is the number of timestamps and also the number of containers; H represents the number of pixels in the length direction, and W represents the number of pixels in the width direction; then for each pixel, its two-dimensional coordinates are (x l , y m ), where l = 0, 1, …, H - 1, m = 0, 1, …, W - 1; its three-dimensional coordinates with time attributes are: (x l , y m , t i ). Using formula (3), obtain its voxel pixel weighted value E(x l , y m , t i ):

[0091]

[0092] Among them:

[0093] ω is the three-dimensional coordinate (x l, y m , t i ) pixel value at;

[0094] p jki It means that the event camera asynchronously senses the change of the brightness signal of the incident light at each pixel position, and generates an event when the absolute value of the change of the brightness signal exceeds the threshold CT (Contrast Threshold, CT), and uses the quadruple e jki = (x j , y k , p hki , t i ) to represent, where, (x j , y k ) represents the horizontal and vertical coordinates of the pixel at which the event is triggered, t i represents the timestamp of the event; p jki represents the polarity of the event, p jki ∈ {-1, 1}, that is: p jki is -1, representing dimming; p jki is 1, representing brightening; if no event is generated at the corresponding pixel position, then p jki is a null value;

[0095] α is the mixing weight coefficient, and its size is determined by the proportion of events generated in the voxel grid E, and its value ranges from 0 to 1, that is: α ∈ [0, 1];

[0096] Therefore, after determining the voxel pixel weighted value E(x l , y m , t i ), the voxel grid E is determined.

[0097] In the present invention, each event within ΔT assigns its polarity value to two adjacent containers according to the chronological distance relationship, and has no effect on the containers outside the adjacent containers. The voxel grid generated in this way conforms to the fixed size characteristics and, to a certain extent, retains the chronological information of the event stream data, which enables the subsequent training process to make the most of the useful information of the event stream data. At the same time, the influence of the pixel value ω is considered, and the proportion of the pixel value in the grid E can be weighted and mixed with the voxel value through user-defined control, making the event contain richer information and the manifestation method more flexible than a single method.

[0098] Step S3, construct a recurrent convolutional neural network with a memory state; the recurrent convolutional neural network with a memory state includes a denoising network, an event reconstruction image network, and a fusion network;

[0099] The recurrent convolutional neural network with memory state is used to map the input heterogeneous visual stream data into a perceptually enhanced image through function calculation by designing the network structure and output it.

[0100] As a specific embodiment, the denoising network, the event reconstruction image network, and the fusion network have the following structural characteristics:

[0101] (1) Denoising network

[0102] The denoising network is used to remove noise from the input low-quality event stream and is implemented using an improved U-Net network structure. Through experimental tests, the best denoising effect is achieved when the number of downsampling units and upsampling units in both the downsampling module and the upsampling module is set to 4; the U-Net network is as Figure 2 shown. The left side is the downsampling module, which contains multiple downsampling units. Each downsampling unit sequentially includes two convolutional layers with a convolution kernel size of 3×3, a rectified linear unit (ReLU), and a pooling layer with a size of 2×2 and a stride of 2. The right side of the network is the upsampling module, which includes several upsampling units. Each upsampling unit sequentially includes an upsampling layer, a transposed convolution layer, and a cropping operation, and after cropping, the feature map is concatenated with the corresponding downsampled feature map.

[0103] Considering the temporal characteristics of event data, the convolutional layers in the downsampling module are modified to long short-term memory (LSTM) layers, so that it can remember the features of previous data to better complete the denoising task;

[0104] Adjust the output of the denoising grid structure, set the number of input channels in the denoising network structure to 3, and ensure that it is consistent with the output format dimension of E(x l ,y m ,t i ), that is: E(x l ,y m ,t i ) contains 3 output data. Therefore, the number of input channels in the denoising network structure is set to 3, and the number of output channels is 3 to meet the format of the event voxel grid data.

[0105] (2) Event reconstruction image network

[0106] The high-quality event stream FH and the generated event stream FL1 share the same set of event reconstruction image networks and share weights.

[0107] (3) Fusion network

[0108] The specific method is as follows: The reconstructed video frame sequence VR and the low-quality video frame sequence VL obtained from the low-quality event stream are concatenated by channel and then input into the fusion network, and finally a high-quality video frame image with both high dynamic range and high texture detail features is synthesized.

[0109] The fusion network structure refers to the U-Net network structure, and the number of downsampling units and upsampling units in the downsampling module and the upsampling module are both set to 3. In addition, in the fusion network structure, the number of input channels is set to 2, and the number of output channels is set to 1 to meet the requirement that the input is two concatenated grayscale images and the output is a fused grayscale image. The complete structure of the reconstructed image and the low-quality image fusion is as Figure 3 shown.

[0110] Step S4: Use the preprocessed training samples to train the recurrent convolutional neural network with memory state to obtain the trained recurrent convolutional neural network with memory state. The specific training method is as follows:

[0111] Step S4.1: Input the low-quality event stream FL into the denoising network, and perform denoising processing by the denoising network to generate the event stream FL1;

[0112] Step S4.2: Input the event stream FL1 generated in step S4.1 into the event reconstruction image network, and the event reconstruction image network outputs the reconstructed video frame sequence VR;

[0113] Step S4.3: Input the reconstructed video frame sequence VR obtained in step S4.2 and the low-quality video frame sequence VL into the fusion network, and the fusion network outputs the reconstructed first high-quality video frame sequence VHR1;

[0114] Step S4.4: Input the high-quality event stream FH into the event reconstruction image network, and the event reconstruction image network outputs the reconstructed second high-quality video frame sequence VHR2;

[0115] Step S4.5: Calculate the first loss function value between the reconstructed video frame sequence VR and the reconstructed second high-quality video frame sequence VHR2;

[0116] Calculate the second loss function value between the reconstructed video frame sequence VR and the high-quality video frame sequence VH;

[0117] Calculate the third loss function value between the reconstructed first high-quality video frame sequence VHR1 and the high-quality video frame sequence VH;

[0118] Integrate the first loss function value, the second loss function value and the third loss function value to obtain the comprehensive loss function value;

[0119] Step S4.6, update the network parameters of the recurrent convolutional neural network with a memory state according to the value of the comprehensive loss function;

[0120] Return to step S4.1, and use the next training sample to continue training the recurrent convolutional neural network with a memory state until the recurrent convolutional neural network with a memory state reaches the training termination condition;

[0121] For example, by setting the loss function, calculate the gradient and backpropagate it to each weight, and update the network weights by the gradient descent method to achieve iterative optimization of the entire network.

[0122] The specific method is as follows:

[0123] Measure the voxel grid generated by the event stream according to the calculation formula of the mean absolute error loss between images. First, regularize the voxel grid, and then calculate the error of the event loss between two voxel grids using the Euclidean distance (L2Loss). By setting the loss function, calculate the gradient and backpropagate it to each weight, and update the weights by the gradient descent method to achieve iterative optimization of the entire network.

[0124] The specific practice is as follows:

[0125] The event loss measures the gap between event data. Specifically, the present invention measures the voxel grid generated by the event stream according to the calculation formula of the mean absolute error loss between images. Between the voxel grids of different data, the voxel value distributions are quite different and there is no maximum value limit. Before measuring the loss between voxel grids, it is necessary to regularize the voxel grid.

[0126] In the present invention, if the voxel grid is used as a unit, when calculating each loss function, it is necessary to regularize the voxel grid E of each video frame sequence, and then Calculate the loss using the Euclidean distance.

[0127] The method for regularizing the voxel grid E is as follows:

[0128] ① For each voxel pixel weight value E(x l ,y m ,t i ) in the voxel grid E, use formula (4) to obtain the intermediate variable Mask(x l ,y m ,t i ):

[0129]

[0130] ② Using formula (5), sum all the intermediate variable Masks(x l , y m , t i ) obtained in the voxel grid E, and count the number of non-zero voxel pixel weighted values E(x l , y m , t i ), which is denoted as N;

[0131]

[0132] ③ Using formula (6), obtain the mean value μ:

[0133]

[0134] ④ Using formula (7), obtain the variance σ:

[0135]

[0136] ⑤ Using formula (8), obtain the regularization value of the voxel pixel weighted value E(x l , y m , t i )

[0137]

[0138] By regularizing each voxel pixel weighted value E(x l , y m , t i ), it means regularizing the voxel grid E.

[0139] Step S5, in a low-light scene, through a visual camera and an ordinary camera, in each acquisition period, respectively acquire the low-quality video frame sequence VL * and the low-quality event stream FL * , and input them into the trained recurrent convolutional neural network with a memory state, and output the reconstructed first high-quality video frame sequence VHR1 * , which is the finally obtained video frame sequence with enhanced image details.

[0140] The aerial scene visual perception enhancement method proposed by the present invention has been tested and verified on real data. It can be Figure 4 seen that the method proposed by the present invention has a good effect on image enhancement of the low-light heterogeneous visual stream data of the aerial scene taken on site, and is clearer and brighter than the aerial low-light images.

[0141] Figure 5It shows the fusion effect of the present invention on the event-reconstructed image and the low-light scene image. Comparing the event-reconstructed image with the fused image, the fused image has more texture details than the event-reconstructed image. That is, the method of the present invention obtains texture detail information from low-quality images and combines it with the dynamic range information of the event-reconstructed image, and finally generates a high-quality image with both high dynamic range and high texture details.

[0142] The present invention relates to the field of aerial remote sensing image data enhancement in image processing. Aiming at the actual application of aerial image processing and with the main purpose of improving the texture details of aerial images, an image texture enhancement method optimized for aerial scenes based on deep learning is proposed.

[0143] The present invention aims at aerial remote sensing images with poor quality, increases the proportion of useful information in the images, emphasizes certain overall features or local features in the images, clarifies the originally unclear images, emphasizes certain specific image features while suppressing other image features, thereby improving the quality of the images, enriching the detailed information possessed by the images, and enhancing the perception effect of the images.

[0144] The present invention utilizes the advantages of both video frame sequence data and event stream data, complements their advantages, generates high-quality images with both high dynamic range and high texture detail features, so as to enhance the visual detail effect of aerial scenes.

[0145] The principle of the present invention lies in:

[0146] From the perspective of image enhancement, for low-quality video frame sequences captured when the scene illumination is low, with the help of the high dynamic range information introduced by the paired low-quality event stream data, the image perception effect can be enhanced, including brightness improvement, noise removal, etc.; from the perspective of event stream reconstructed images, for low-quality event stream data captured when the scene illumination is low, the reconstruction effect of the event stream reconstructed images can be improved with the help of the high texture detail information introduced by the paired low-quality video frame sequence data. Utilizing the advantages of both video frame sequence data and event stream data, high-quality images with both high dynamic range and high texture detail features are generated, thus realizing the innovation of enhancing the visual detail effect of aerial scenes.

[0147] The advantages of the present invention compared with the prior art are as follows:

[0148] 1. The model designed by the present invention simultaneously utilizes the high dynamic range information possessed by event stream data and the high texture detail information possessed by image frame sequence data, and can fuse and construct a visual perception enhanced image with both high dynamic range and high texture detail information;

[0149] 2. During the process of converting event stream data into a voxel grid, the contribution of pixel values is considered simultaneously. The method of the present invention makes use of the information of both the voxel grid and pixel coordinates by means of a mixing weight.

[0150] 3. The method of the present invention has a certain ability to improve image quality and enhance image details, and has application value in the field of aerial photography.

[0151] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for enhancing details of aerial images based on deep learning, characterized in that, Including the following steps: Step S1, obtaining a training sample set; each training sample in the training sample set simultaneously includes a high-quality event stream, a low-quality event stream, a low-quality video frame sequence, and a high-quality video frame sequence that are truly collected from the same region corresponding to the same time interval; Wherein: the high-quality event stream and the low-quality event stream are collected by a vision camera; the high-quality video frame sequence and the low-quality video frame sequence are collected by an ordinary camera; Step S2, for each training sample, performing data preprocessing to obtain a preprocessed training sample; the preprocessed training sample includes a preprocessed high-quality event stream, a low-quality event stream, a low-quality video frame sequence, and a high-quality video frame sequence, and the high-quality event stream, the low-quality event stream, the low-quality video frame sequence, and the high-quality video frame sequence are respectively represented as: high-quality event stream FH, low-quality event stream FL, low-quality video frame sequence VL, and high-quality video frame sequence VH; Step S3, constructing a recurrent convolutional neural network with a memory state; the recurrent convolutional neural network with a memory state includes a denoising network, an event reconstruction image network, and a fusion network; Step S4, using the preprocessed training sample to train the recurrent convolutional neural network with a memory state to obtain the trained recurrent convolutional neural network with a memory state. The specific training method is as follows: Step S4.1, inputting the low-quality event stream FL into the denoising network, and performing denoising processing by the denoising network to generate an event stream FL1; Step S4.2, inputting the event stream FL1 generated in Step S4.1 into the event reconstruction image network, and the event reconstruction image network outputs a reconstructed video frame sequence VR; Step S4.3, inputting the reconstructed video frame sequence VR obtained in Step S4.2 and the low-quality video frame sequence VL into the fusion network, and the fusion network outputs a reconstructed first high-quality video frame sequence VHR1; Step S4.4, inputting the high-quality event stream FH into the event reconstruction image network, and the event reconstruction image network outputs a reconstructed second high-quality video frame sequence VHR2; Step S4.5, calculating a first loss function value between the reconstructed video frame sequence VR and the reconstructed second high-quality video frame sequence VHR2; Calculating a second loss function value between the reconstructed video frame sequence VR and the high-quality video frame sequence VH; Calculating a third loss function value between the reconstructed first high-quality video frame sequence VHR1 and the high-quality video frame sequence VH; Combining the first loss function value, the second loss function value, and the third loss function value to obtain a combined loss function value; Step S4.6, updating the network parameters of the recurrent convolutional neural network with a memory state according to the combined loss function value; Return to Step S4.1, use the next training sample to continue training the recurrent convolutional neural network with a memory state until the recurrent convolutional neural network with a memory state reaches the training termination condition; Step S5, in a low-light scenario, through a vision camera and a normal camera, in each acquisition cycle, respectively acquire a low-quality video frame sequence VL of the target area * and a low-quality event stream FL * , input them into the trained recurrent convolutional neural network with a memory state, and output a reconstructed first high-quality video frame sequence VHR1 * , which is the finally obtained video frame sequence with enhanced image details.

2. The method for enhancing details of aerial images based on deep learning according to claim 1, wherein In step S2, the preprocessing methods for the low-quality video frame sequence and the high-quality video frame sequence are as follows: Convert each frame image in the actually collected low-quality video frame sequence and high-quality video frame sequence into a grayscale image.

3. The method for enhancing the details of aerial images based on deep learning according to claim 2, characterized in that, Convert to a grayscale image through formula (1): I gray = 0.2989 × I rgb [r] + 0.5870 × I rgb [g] + 0.1140 × I rgb [b] (1) where: l gray is a grayscale image, and I rgb [r], I rgb [g] and I rgb [b] are the RGB three-channel image components of each frame image, respectively.

4. A method for enhancing details of aerial images based on deep learning according to claim 1, characterized in that, In step S2, the preprocessing method for the high-quality event stream and the low-quality event stream is to convert the event stream into the form of a voxel grid.

5. A method for enhancing details of aerial images based on deep learning according to claim 4, characterized in that, Convert the event stream into the form of a voxel grid, specifically: ① Divide the entire event stream into a number of event stream units with a time interval length of ΔT according to the event timestamps. ② For each event stream unit, its start time is represented as t0, and the end time is: t0 + ΔT. Suppose there are B timestamps in the time interval length [t0, t0 + ΔT], which are respectively represented as: t0, t1, …, t B-1 ; Therefore, any timestamp is represented as t i , where i = 0, 1, …, B - 1; ③ For each event stream unit, a container is generated for all events corresponding to each of its timestamps t i ; ④ Using formula (2), timestamp t i is normalized to obtain the normalized timestamp ⑤ Convert all events in each event stream unit into a voxel grid E. The size of the voxel grid E is B×H×W, where B is the number of timestamps and also the number of containers; H represents the number of pixels in the length direction, and W represents the number of pixels in the width direction; then for each pixel, its two-dimensional coordinates are (x l , y m ), where l = 0, 1, …, H−1, m = 0, 1, …, W−1; its three-dimensional coordinates with time attribute are: (x l , y m , t i ). Using formula (3), its voxel pixel weighted value E(x l , y m , t i ) is obtained: Where: ω is the pixel value at the three-dimensional coordinates (x l , y m , t i ); p jki It means that the event camera asynchronously senses the change in the brightness signal of the incident light at each pixel position, and generates an event when the absolute value of the change in the brightness signal exceeds the threshold CT. The event is represented by a quadruple e jki =(x j , y k , p jki , t i ), where (x j , y k ) represents the horizontal and vertical coordinates of the pixel at which the event is triggered, t i represents the timestamp of the event; p jki represents the polarity of the event. p jki ∈{-1, 1}, that is: p jki being -1 represents dimming; p jki being 1 represents brightening; if no event is generated at the corresponding pixel position, then p jki is a null value; α is the mixing weight coefficient, and its value is determined by the proportion of events generated in the voxel grid E, with a value range of 0 to 1, that is, α ∈ [0, 1]; Therefore, after determining the voxel pixel weighting value E(x l , y m , t i ), the voxel grid E is determined.

6. The method for enhancing details of aerial images based on deep learning according to claim 5, characterized in that When calculating the loss function values between each video frame sequence, first regularize the voxel grid E of each video frame sequence, and then perform calculations on the regularized voxel grid Use the Euclidean distance to calculate the loss.

7. A method for enhancing details of aerial images based on deep learning according to claim 6, characterized in that The method for regularizing the voxel grid E is: ①For each voxel pixel weighting value E(x l , y m , t i ) in the voxel grid E, using formula (4), the intermediate variable Mask(x l , y m , t i ) is obtained: ② Using formula (5), sum all the intermediate variables Mask(x k , y m , t i ) obtained in the voxel grid E, and count the number of non-zero weighted voxel pixel values E(x k , y m , t i ), denoted as N; ③ Use formula (6) to obtain the mean value μ: ④ Use formula (7) to obtain the variance σ: ⑤ Using formula (8), the regularization value of the voxel pixel weight value W(x l , y m , t i ) is obtained By regularizing each voxel pixel weight value E(x l , y m , t i ), that is, regularizing the voxel grid E.

Citation Information

Patent Citations

  • Self-supervised video deblurring and image frame insertion method based on event camera

    CN114494050A

  • Event camera video reconstruction method based on deep learning

    CN115484410A