A no-reference enhancement method for drones to collect images at night
By designing a reference-free enhancement method, using enhancement curve estimation network, enhancement calculation formula and perceptual restoration network, the problems of poor visualization and noise block effect of drones at night are solved, and high-quality image enhancement effect is achieved.
Patent Information
- Application Number
- CN202310355905.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-04-06
AI Technical Summary
When drones acquire images at night, the image visibility is significantly poorer, and there is noise and block effects, so the existing simple image enhancement methods are difficult to effectively improve.
A reference-free enhancement method is designed, and by building a data set for training, the enhanced network is designed, including an enhancement curve estimation network, an enhancement calculation formula and a perceptual restoration network, and the enhancement of low-illumination images is learned using the reference-free loss function to suppress highlights, improve block effects and noise.
It effectively improves the visibility and quality of images collected at night by the drone, and the generated images have better perception capabilities and suppresses noise and block effects.
Smart Images

Figure CN116385298B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and computer vision, and in particular to a no-reference enhancement method for image acquisition at night by an unmanned aerial vehicle (UAV). Background Art
[0002] In recent years, the progress and development of science and technology have significantly promoted the improvement of society and people's lives. With the miniaturization of the volume of image acquisition devices and the high efficiency of acquisition capabilities, images and image systems are closely related to people's daily lives and production development. Image-based processing systems have a wide range of uses in life scenarios. The image system can facilitate users to record and observe in real time intuitively, with high convenience. However, limited by the environment and usage status of the acquisition device, the obtained images lack ideal observability, which is reflected in phenomena such as motion blur and poor illuminance of the images. Especially when the lighting conditions are poor, the situation of poor illuminance is very common. At this time, the images taken often have multiple dark areas, which bring difficulties to human eye reading or machine vision processing. Therefore, designing an enhancement method for low-illuminance images has important theoretical and application significance.
[0003] Low-illuminance image enhancement aims to improve the visibility of the content of images with insufficient lighting conditions. In low-illuminance images, the visibility and distinguishability of objects, scenes, and textures are poor, and it is difficult to be recognized by the human eye or used for advanced computer vision tasks. For enhancing low-illuminance images, there are usually two typical scenarios: one is to enhance the entire image, which is usually seen when the overall lighting conditions are very poor or the exposure parameters (shutter time or white balance parameters) of the camera are set incorrectly; the other is to enhance some dark areas in the image, which is more common in naturally captured images.
[0004] In the real application scenario of the UAV scenario, the visibility of images collected at night is significantly poor. At the same time, since the image device system carried by the UAV is relatively lightweight, the quality of the images it collects is often not high. In particular, images collected in the night scene will also show image degradation phenomena such as noise and blocking effects. For these manifestations of night image degradation, simple image enhancement methods are difficult to effectively improve image perception, and may even enhance the degree of some image degradation. In addition, since no true-value images with normal illuminance can be collected in the night UAV scenario, it is more difficult to design a low-illuminance enhancement method for this scenario. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a reference-free enhancement method for images collected by drones at night. This method predicts enhancement coefficients for each pixel of the input image, and a specially designed enhancement operation formula can effectively suppress highlights. The designed network can further improve block effects and noise, thereby generating high-quality results. The designed network uses a reference-free loss function to learn the enhancement of low-light images. These specially designed loss functions indirectly constrain key factors such as the brightness, color, illumination smoothness, noise, and semantic consistency of the image enhancement results, thereby obtaining results with better perceptual ability.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A reference-free enhancement method for images collected by drones at night, comprising the following steps:
[0007] Step S1, construct a training dataset, and preprocess each image collected by the drone at night to obtain a training dataset;
[0008] Step S2, design an enhancement network for images collected by drones at night, which is composed of an enhancement curve estimation network, an enhancement operation formula, and a perceptual restoration network;
[0009] Step S3, design a reference-free loss function for training the network designed in step S2;
[0010] Step S4, use the training dataset to train the enhancement network for images collected by drones at night;
[0011] Step S5, input the night image collected by the drone to be tested into the designed network, and use the trained network to predict and generate a final result with better perception.
[0012] In a preferred embodiment, in step S1, construct a training dataset, and preprocess each image collected by the drone at night to obtain a training dataset; including the following steps:
[0013] Step S11, scale each image in the dataset to an image of the same size with dimensions H×W;
[0014] Step S12, perform normalization processing on the training images; given an image I train , calculate the normalized image using the following formula:
[0015]
[0016] where I train is an image with 8-bit color depth and size H×W, and I bit_max is an image with size H×W and all pixel values being 255.
[0017] In a preferred embodiment, in step S2, for the enhancement network of the UAV to collect images at night, the enhancement network consists of an enhancement curve estimation network, an enhancement operation, and a perceptual restoration network; the following steps are included:
[0018] Step S21: Design an enhancement curve estimation network, and use the designed network to generate parameters for image enhancement; the enhancement curve estimation network accepts the normalized UAV night-time collected image X with a size of H×W as input. The specific structure of the enhancement curve estimation network consists of 3 convolutional blocks, and each convolutional block is composed of a convolutional layer and an activation layer in sequence; the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, and the activation layer uses the ReLU activation function; the enhancement curve estimation network outputs the parameters A for image enhancement with a size of H×W, that is, each pixel of the input image has its corresponding image enhancement parameter;
[0019] Step S22: Design an enhancement operation formula to enhance the night-time image collected by the UAV, where the parameters of the formula depend on the result output by the enhancement curve estimation network in step S21; for the normalized UAV night-time collected image X with a size of H×W and the image enhancement parameter A obtained in step S21, the enhancement operation formula is as follows:
[0020]
[0021] where represents the preliminary enhancement result, α is a scaling factor that controls the curve shape and should be set to a constant greater than 0, and g(X; A) is a non-linear function with parameter A; the specific calculation method of g(X; A) is:
[0022]
[0023] where X is the normalized UAV night-time collected image with a size of H×W, where the cos(·) function calculates the cosine value for each element of the input tensor, and θ is a preset curve inflection point value, and the effective value range is [0, 2];
[0024] Step S23: Design a perceptual restoration network, which includes an encoder and a decoder; the input of the perceptual restoration network is the preliminary enhancement result and the output is the final result The designed network can further improve the quality of the night-time image enhancement collected by the UAV.
[0025] In a preferred embodiment, step S23 is specifically implemented as follows:
[0026] Step S231: Design an encoder structure, and the encoder consists of convolutional block C 1, C 2 , pooling layer P 1 , convolutional block C 3 and pooling layer P 2 stacked; where the convolutional block C 1 , C 2 , C 3 is composed of a convolutional layer and an activation layer in sequence; the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, and the activation layer is a LeakyReLU activation function with a negative slope set to 0.2; the pooling layer P 1 , P 2 are both max-pooling operations with a pooling size of 2×2;
[0027] When the input is a normalized drone night acquisition image X with a size of H×W, the formula for its feedforward operation is expressed as:
[0028] F 1 = C 1 (X)
[0029] F 2 = C 2 (F 1 )
[0030] F 3 = P 1 (F 2 )
[0031] F 4 = C 3 (F 3 )
[0032] F 5 = P 2 (F 4 )
[0033] where F 1 ,..., F 5 are the feature outputs of the operations of each intermediate layer of the encoder;
[0034] Step S232, design the decoder structure. The decoder is composed of upsampling blocks U 1 , convolutional blocks C 4 , C 5 , upsampling blocks U 2 , convolutional blocks C 6 , C 7 , C 8 , C 9 stacked in sequence. The output of the decoder is the final result where the upsampling block U 1 , U 2has the same structure, that is, it consists of 1 transposed convolutional layer, the nearest neighbor linear interpolation operation I(·), and the feature concatenation operation Cat(·); the transposed convolutional layer is a transposed convolution with a kernel size of 2×2 and a stride of 2; the convolutional block C 6 , C 7 , C 8 , C 9 is composed of a convolutional layer and an activation layer in sequence; the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, and the activation layer is the LeakyReLU activation function, and the negative slope of this activation function is set to 0.2; among them, the feature output of the convolutional block C 5 is denoted as F 8 ; the upsampling block accepts two feature inputs of different sizes and finally outputs the concatenated features with a larger size; specifically, the upsampling block U 1 accepts F 5 , F 3 as inputs, F 3 obtains F 5 ′ through the transposed convolutional layer, F 3 obtains F 3 ′ through the nearest neighbor linear interpolation operation I(F 3 ), F 3 ′ and F 5 ′ have the same size, and the feature concatenation Cat(F′ 5 , F′ 3 ) concatenates the two features along the channel dimension; similarly, the upsampling block U 2 accepts F 8 , X as inputs, F 8 obtains F 8 ′ through the transposed convolutional layer, X obtains X′ through the nearest neighbor linear interpolation operation I(X), X′ and F 8 ′ have the same size, and the feature concatenation Cat(F′ 8 , X′) concatenates the two features along the channel dimension.
[0035] In a preferred embodiment, in step S3, a reference-free loss function is designed for training the network designed in step S2; it includes the following steps:
[0036] Step S31, design the overall optimization objective of the entire network. The optimization objective is to minimize the total loss function
[0037]
[0038] where, represents the brightness control loss function, and λ bri represents the weight of the brightness control loss function; represents the color loss function, and λ colRepresents the weight of the color loss function; Represents the image enhancement parameter smoothing loss function, λ tv Represents the weight of the image enhancement parameter smoothing loss function; Represents the semantic-driven deblocking loss function, λ blo Represents the weight of the semantic-driven deblocking loss function; Represents the semantic self-regularized denoising loss function, λ noi Represents the weight of the semantic self-regularized denoising loss function;
[0039] Step S32: Design the brightness control loss function; The calculation formula is as follows:
[0040]
[0041] Among them, the preliminarily enhanced image is divided into N = 16×16, that is, 256 regions of the same size, and the average brightness value within the region is denoted as 0.6 is a preset target brightness constant value, || || 1 is the absolute value operation;
[0042] Step S33: Design the color loss function; The calculation formula is as follows:
[0043]
[0044] In the above formula respectively represent the preliminary enhancement results in the red, green, and blue channels of the RGB color space;
[0045] Step S34: Design the illumination smoothing loss function; The calculation formula is as follows:
[0046]
[0047] Among them, A is the color channel c corresponding to the image enhancement parameter A output by the enhancement curve estimation network, and c is taken from the RGB color space; || || 2 is the operation of obtaining the Euclidean norm; represents the operation of obtaining the first-order gradient, that is and are respectively the first-order differences of A c in the vertical and horizontal directions, || || 1 is the absolute value operation;
[0048] Step S35: Design the semantic-driven deblocking loss function The deblocking loss function can suppress the blocking effect during the image enhancement process.
[0049] Step S36: Design a semantic self-regularized denoising loss function The denoising loss function can effectively suppress the noise pixels during the image enhancement process.
[0050] In a preferred embodiment, in step S35, design a semantic-driven deblocking loss function The deblocking loss function can suppress the blocking effect during the image enhancement process; it includes the following steps:
[0051] Step S351: For the input normalized nighttime drone-acquired image X, use the VGG-16 model pre-trained on the classification task of the ImageNet dataset to extract features from X, and the result of summing the output features of the Conv4-2 layer by channel is denoted as f X ;
[0052] Step S352: For the preliminarily enhanced image use the VGG-16 model pre-trained on the classification task of the ImageNet dataset to extract features, and the result of summing the output features of the Conv4-2 layer by channel is denoted as
[0053] Step S353: For the output result of the enhancement network designed for nighttime drone-acquired images in step S2, that is, the final result use the VGG-16 model pre-trained on the classification task of the ImageNet dataset to extract features, and the result of summing the output features of the Conv4-2 layer by channel is denoted as
[0054] Step S354: Design a semantic-driven deblocking loss function which is represented by the formula:
[0055]
[0056] where * represents element-wise multiplication of matrices, e is the base of the natural logarithm, μ is a parameter that can control the semantic intensity perception, and the default setting is 20. || || 1 is the absolute value operation.
[0057] In a preferred embodiment, in step S36, design a semantic self-regularized denoising loss function The denoising loss function Can effectively suppress noise pixels during image enhancement; including the following steps:
[0058] Step S361: Use the subgraph sampling method in Neighbor2Neighbor to construct the subgraph sampler G N =(g 1 , g 2 ); The specific method is as follows: For any input image T with size H T ×W T ; Divide it into tuples, where represents the floor operation, and each tuple contains 2×2 = 4 pixels; In all tuple partitions, the pixels in its j-th tuple are respectively denoted as The subgraph sampler G N =(g 1 , g 2 ) When sampling in the j-th tuple, for g 1 and g 2 , they respectively randomly correspond to "g 1 uses row sampling, g 2 uses column sampling" or "g 1 uses column sampling and g 2 uses row sampling"; where row sampling means selecting 1 element from the set , and column sampling means selecting 1 element from the set .
[0059] Step S362: For the input normalized night-time drone-acquired image X, use the VGG-16 model pre-trained on the ImageNet dataset for classification tasks to extract features from X, and the result of summing the output features of its Conv4-2 layer by channel is denoted as f X ;
[0060] Step S363: The semantic self-regularized denoising loss function is designed as follows:
[0061]
[0062] where * represents the element-wise multiplication of matrices. is the subgraph difference loss, is the regularization term of the semantic self-regularized denoising loss function; The calculation method of
[0063]
[0064] where Φ is the perceptual restoration network described in step S23, is the preliminarily enhanced image g 1associated with g 2 defined by the sub - graph sampler G; || || 2 is an operation to obtain the Euclidean norm;
[0065] The calculation method of is as follows:
[0066]
[0067] where Φ is the perceptual restoration network described in step S23, is the preliminarily enhanced image g 1 associated with g 2 defined by the sub - graph sampler G; || || 2 is an operation to obtain the Euclidean norm.
[0068] In a preferred embodiment, in the step S4, an enhancement network for training the images collected by the drone at night using a training data set; includes the following steps:
[0069] Step S41: Select a random training image X from the data set constructed in step S1;
[0070] Step S42: Training image encoding and enhancement. Input the image X, obtain the image enhancement parameter A through the enhancement curve estimation network, and obtain the preliminarily enhanced image through the enhancement operation formula Finally, obtain the final result through the perceptual restoration network Calculate the loss of the total loss function in step S31
[0071] Step S43: Use the backpropagation method to calculate the gradients of the parameters in the enhancement curve estimation network and the perceptual restoration network, and update the parameters using the Adam optimization method;
[0072] Step S44: The above steps are one iteration of the training process. The entire training process requires 100 iterations, and in each iteration process, multiple image pairs are randomly sampled as a batch for training.
[0073] In a preferred embodiment, in the step S5, input the night - time image collected by the drone to be measured into the designed network, and use the trained network to predict and generate a final result with better perception.
[0074] Compared with the prior art, the present invention has the following beneficial effects: Aiming at the problems that the existing method for enhancing night images collected by drones lacks highlight retention and image quality improvement, the present invention aims to improve the visibility of night images collected by drones and at the same time improve the image quality. The present invention proposes a reference-free enhancement method for images collected by drones at night, which enhances the brightness by designing a network and suppresses noise and blocking effects. The enhancement curve estimation network in the network is used to generate parameters for pixel-level image enhancement; the enhancement operation formula is used to enhance the night images collected by drones to generate a preliminary result; the perceptual restoration network is used to denoise and suppress blocking effects, and finally can output high-quality images with normal illumination. Description of the Drawings
[0075] Figure 1 It is a flowchart of the method according to a preferred embodiment of the present invention.
[0076] Figure 2 It is the enhancement process of the reference-free enhancement method for images collected by drones at night according to a preferred embodiment of the present invention.
[0077] Figure 3 It is the enhancement curve estimation network according to a preferred embodiment of the present invention.
[0078] Figure 4 It is the perceptual restoration network according to a preferred embodiment of the present invention. Detailed Description of the Invention
[0079] The present invention will be further described below with reference to the drawings and embodiments.
[0080] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0081] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0082] The present invention provides a reference-free enhancement method for images collected by drones at night, as Figures 1-4 shown, including the following steps:
[0083] Step S1, constructing a training dataset, preprocessing each image collected by the drone at night to obtain a training dataset;
[0084] Step S2: Design an enhancement network for night-time image acquisition by drones. This network consists of an enhancement curve estimation network, an enhancement operation formula, and a perceptual restoration network;
[0085] Step S3: Design a reference-free loss function for training the network designed in Step S2;
[0086] Step S4: Use the training dataset to train the enhancement network for night-time image acquisition by drones;
[0087] Step S5: Input the night-time images collected by the drone to be tested into the designed network, and use the trained network to predict and generate the final result with better perception.
[0088] Furthermore, Step S1 includes the following steps:
[0089] Step S11: Scale each image in the dataset to an image of the same size with dimensions H×W.
[0090] Step S12: Normalize the training images. Given an image I train , the formula for calculating the normalized image is as follows:
[0091]
[0092] where I train is an image of size H×W with 8-bit color depth, and I bit_max is an image of size H×W with all pixel values being 255.
[0093] Furthermore, as Figure 2 shown, Step S2 includes the following steps:
[0094] Step S21: As Figure 3 shown, design an enhancement curve estimation network and use the designed network to generate parameters for image enhancement. This network takes the normalized night-time image X collected by the drone with dimensions H×W as input. The specific structure of the enhancement curve estimation network consists of 3 convolutional blocks, and each convolutional block is composed of a convolutional layer and an activation layer in sequence. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. The activation layer uses the ReLU activation function. The enhancement curve estimation network outputs parameters A for image enhancement with dimensions H×W, that is, each pixel of the input image has its corresponding image enhancement parameter.
[0095] Step S22: Design an enhanced operation formula to enhance the night images collected by the UAV. The parameters of the formula depend on the results output by the enhancement curve estimation network in step S21. For the normalized UAV night-time collected image X of size H×W, and the image enhancement parameter A obtained in step S21, the enhanced operation formula is as follows:
[0096]
[0097] where represents the preliminary enhancement result, α is a scaling factor that controls the curve shape and should be set as a constant greater than 0, and g(X; A) is a non-linear function with parameter A. The specific calculation method of g(X; A) is:
[0098]
[0099] where X is the normalized UAV night-time collected image of size H×W. The cos(·) function calculates the cosine value for each element of the input tensor, and θ is the preset curve inflection point value, with a valid value range of [0, 2].
[0100] Step S23: Design a perceptual restoration network, which includes an encoder and a decoder. The input of the perceptual restoration network is the preliminary enhancement result and the output is the final result The designed network can further improve the quality of the enhancement of the night images collected by the UAV.
[0101] Furthermore, as Figure 4 shown, step S23 includes the following steps:
[0102] Step S231: Design the encoder structure. The encoder is composed of convolutional blocks C 1 , C 2 , pooling layer P 1 , convolutional block C 3 and pooling layer P 2 stacked together. Among them, convolutional blocks C 1 , C 2 , C 3 are composed of a convolutional layer and an activation layer in sequence. The convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1. The activation layer is a LeakyReLU activation function, and the negative slope of this activation function is set to 0.2. Pooling layers P 1 , P 2 are both max-pooling operations with a pooling size of 2×2.
[0103] When the input is the normalized UAV night-time collected image X of size H×W, the formula for its feed-forward operation is expressed as:
[0104] F1 = C 1 (X)
[0105] F 2 = C 2 (F 1 )
[0106] F 3 = P 1 (F 2 )
[0107] F 4 = C 3 (F 3 )
[0108] F 5 = P 2 (F 4 )
[0109] where F 1 ,..., F 5 are the characteristic outputs of each layer operation in the encoder middle.
[0110] Step S232, design the decoder structure. The decoder is composed of upsampling blocks U 1 , convolutional blocks C 4 , C 5 , upsampling blocks U 2 , convolutional blocks C 6 , C 7 , C 8 , C 9 stacked in sequence. The output of the decoder is the final result where the upsampling blocks U 1 , U 2 have the same structure, that is, they are composed of 1 transposed convolutional layer, nearest neighbor linear interpolation operation I(·) and feature concatenation Cat(·) operation. The transposed convolutional layer is a transposed convolution with a convolutional kernel size of 2×2 and a stride of 2. The convolutional blocks C 6 , C 7 , C 8 , C 9 are composed of convolutional layers and activation layers in sequence. The convolutional layer is a convolution with a convolutional kernel size of 3×3, a stride of 1, and a padding of 1. The activation layer is the LeakyReLU activation function, and the negative slope of this activation function is set to 0.2. Among them, the characteristic output of the convolutional block C 5 is denoted as F 8 . The upsampling block accepts two feature inputs of different sizes and finally outputs a concatenated feature with a larger size. Specifically, the upsampling block U 1 accepts F 5 , F 3 as inputs, and F 3 obtains F5 ′,F 3 After the nearest neighbor linear interpolation operation I(F 3 ) to get F 3 ′,F 3 ′ and F 5 ′ has the same size, feature splicing Cat(F′ 5 , F′ 3 ) concatenates the two features according to the channel dimension. Similarly, the upsampling block U 2 Accept F 8 ,X as input,F 8 After the transposed convolution layer, we get F 8 ′, X is obtained by the nearest neighbor linear interpolation operation I(X), X′ and F 8 ′ has the same size, feature splicing Cat(F′ 8 , X′) concatenates the two features according to the channel dimension.
[0111] Further, step S3 includes the following steps:
[0112] Step S31: Design the overall optimization goal of the entire network. The optimization goal is to minimize the total loss function
[0113]
[0114] in, represents the brightness control loss function, λ bri Represents the weight of the brightness control loss function; represents the color loss function, λ col Represents the weight of the color loss function; represents the image enhancement parameter smoothing loss function, λ tv Represents the weight of the smoothing loss function of the image enhancement parameters; represents the semantically driven deblocking loss function, λ blo represents the weight of the semantically driven deblocking loss function; represents the denoising loss function of semantic self-regularization, λ noi Represents the weight of the denoising loss function of semantic self-regularization.
[0115] Step S32: design a brightness control loss function. The calculation formula is as follows:
[0116]
[0117] Among them, the image after preliminary enhancement Divided into N = 16 × 16, that is, 256 regions of the same size, the average brightness value in the region is recorded as 0.6 is a preset target brightness constant value, || || 1 is an absolute value operation.
[0118] Step S33: Design a color loss function. The calculation formula is as follows:
[0119]
[0120] In the above formula respectively represent the preliminary enhancement result The mean values of the red, green, and blue channels in the RGB color space.
[0121] Step S34: Design a lighting smoothing loss function. The calculation formula is as follows:
[0122]
[0123] Among them, A is the color channel c corresponding to the image enhancement parameter A output by the enhancement curve estimation network, and c is taken from the RGB color space. || || 2 is an operation to obtain the Euclidean norm. represents an operation to obtain the first-order gradient, that is and are respectively the first-order differences of A c in the vertical and horizontal directions, || || 1 is an absolute value operation.
[0124] Step S35: Design a semantic-driven deblocking loss function This loss function can suppress the blocking effect during the image enhancement process.
[0125] Step S36: Design a semantic self-regularized denoising loss function This loss function can effectively suppress the noise pixels during the image enhancement process.
[0126] Furthermore, step S35 includes the following steps:
[0127] Step S351: For the input normalized drone night acquisition image X, use the VGG-16 model pre-trained on the classification task of the ImageNet dataset to extract features from X, and the result of summing the output features of its Conv4-2 layer by channel is denoted as f X .
[0128] Step S352: For the preliminarily enhanced image Use the VGG-16 model pre-trained on the classification task of the ImageNet dataset for Extract features, and denote the result of summing the output features of the Conv4-2 layer by channel as
[0129] Step S353: For the output result of the enhancement network designed for nighttime image acquisition by the drone in Step S2, that is, the final result Use the VGG-16 model pre-trained on the classification task of the ImageNet dataset for Extract features, and denote the result of summing the output features of the Conv4-2 layer by channel as
[0130] Step S354: Design a semantic-driven deblocking loss function It is represented by the formula:
[0131]
[0132] In the formula, * represents element-wise multiplication of matrices, e is the base of the natural logarithm, μ is a parameter that can control the perception of semantic intensity, and the default setting is 20. || || 1 is the absolute value operation.
[0133] Furthermore, Step S36 includes the following steps:
[0134] Step S361: Use the subgraph sampling method in Neighbor2Neighbor to construct a subgraph sampler G N =(g 1 , g 2 ). The specific method is as follows: For any input image T with a size of H T ×W T . Divide it into tuples, where represents the floor operation, and each tuple contains 2×2 = 4 pixels. Among all tuple partitions, the pixels in the j-th tuple are respectively denoted as The subgraph sampler G N =(g 1 , g 2 ) When sampling in the j-th tuple, for g 1 and g 2 respectively randomly correspond to "g 1 uses row sampling, g 2 uses column sampling" or "g 1 uses column sampling and g 2 uses row sampling". Among them, row sampling means selecting 1 element from the set , and column sampling means selecting 1 element from the set .
[0135] Step S362: For the input normalized night-time drone acquisition image X, use the VGG-16 model pre-trained for classification tasks on the ImageNet dataset to extract features from X, and denote the result of summing the output features of its Conv4-2 layer by channel as f X .
[0136] Step S363: The semantic self-regularized denoising loss function is designed as follows:
[0137]
[0138] where * represents element-wise multiplication of matrices. is the subgraph difference loss, is the regularization term of the semantic self-regularized denoising loss function. The calculation method of
[0139]
[0140] is as follows: where Φ is the perceptual restoration network described in Step S23, is the preliminarily enhanced image g 1 and g 2 are defined by the subgraph sampler G. || || 2 is the operation to obtain the Euclidean norm.
[0141] The calculation method of
[0142]
[0143] is as follows: where Φ is the perceptual restoration network described in Step S23, is the preliminarily enhanced image g 1 and g 2 are defined by the subgraph sampler G. || || 2 is the operation to obtain the Euclidean norm.
[0144] Furthermore, Step S4 includes the following steps:
[0145] Step S41: In the dataset constructed through Step S1, select a random training image X.
[0146] Step S42: Training image encoding and enhancement. Input the image X, obtain the image enhancement parameter A through the enhancement curve estimation network, and obtain the preliminarily enhanced image through the enhancement operation formula Finally, obtain the final result through the perceptual restoration network Calculate the loss of the total loss function in Step S31
[0147] Step S43: Calculate the gradients of the parameters in the enhanced curve estimation network and the perception restoration network using the backpropagation method, and update the parameters using the Adam optimization method.
[0148] Step S44: The above steps are one iteration of the training process. The entire training process requires 100 iterations, and in each iteration process, multiple image pairs are randomly sampled as a batch for training.
[0149] Further, step S5 includes the following steps:
[0150] Step S5: Input the night image collected by the UAV to be measured into the designed network, and use the trained network to predict and generate the final result with better perception.
[0151] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. A reference - free enhancement method for drones to collect images at night, characterized in that, it includes the following steps: Step S1: Construct a training dataset, pre - process each image collected by the drone at night to obtain a training dataset; Step S2: Design an enhancement network for images collected by the drone at night. This enhancement network consists of an enhancement curve estimation network, an enhancement operation formula, and a perceptual restoration network; Step S3: Design a reference - free loss function for training the network designed in Step S2; Step S4: Use the training dataset to train the enhancement network for images collected by the drone at night; Step S5: Input the night - time images collected by the drone to be measured into the designed network, and use the trained network to predict and generate a final result with better perception; In the said Step S2, for the enhancement network for images collected by the drone at night, this enhancement network consists of an enhancement curve estimation network, an enhancement operation, and a perceptual restoration network; it includes the following steps: Step S21: Design an enhancement curve estimation network, and use the designed network to generate parameters for image enhancement; This enhancement curve estimation network receives the normalized drone - collected night - time image X with a size of H×W as input. The specific structure of the enhancement curve estimation network consists of 3 convolutional blocks, and each convolutional block is composed of a convolutional layer and an activation layer in sequence; the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, and the activation layer uses the ReLU activation function; the enhancement curve estimation network outputs the parameters A for image enhancement with a size of H×W, that is, each pixel of the input image has its corresponding image enhancement parameter; Step S22: Design an enhancement operation formula to enhance the night - time images collected by the drone, where the parameters of the formula depend on the results output by the enhancement curve estimation network in Step S21; for the normalized drone - collected night - time image X with a size of H×W and the image enhancement parameter A obtained from Step S21, the enhancement operation formula is as follows: Among them represents the preliminary enhancement result, α is a scaling factor that controls the curve shape and should be set as a constant greater than 0, and g(X; A) is a non-linear function with parameter A; the specific calculation method of g(X; A) is as follows: where X is the normalized drone - collected night - time image with a size of H×W, and the cos(H) function calculates the cosine value for each element of the input tensor, θ is a preset curve inflection point value, and the effective value range is [0, 2]; Step S23: Design a perception restoration network, which includes an encoder and a decoder; the input of the perception restoration network is the preliminary enhancement result The output is the final result The designed network can further improve the quality of night image enhancement collected by drones.
2. The reference - free enhancement method for drones to collect images at night according to claim 1, characterized in that, in the said Step S1, when constructing a training dataset and pre - processing each image collected by the drone at night to obtain a training dataset, it includes the following steps: Step S11: Scale each image in the dataset to an image with the same size of H×W; Step S12: Normalize the training images; Given an image I train , calculate the formula for the normalized image as follows: where I train is an image of size H×W with an 8-bit color depth, and I bit_max is an image of size H×W with all pixel values being 255.
3. The reference - free enhancement method for drones to collect images at night according to claim 1, characterized in that, the specific implementation of the said Step S23 is as follows: Step S231: Design the encoder structure. The encoder is composed of convolutional blocks C 1 , C 2 , pooling layers P 1 , convolutional blocks C 3 and pooling layers P 2 stacked together; among them, convolutional block C 1 , C 2 , C 3 is composed of a convolutional layer and an activation layer in sequence; the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, and the activation layer is a LeakyReLU activation function, and the negative slope of this activation function is set to 0.2; pooling layer P 1 , P 2 are both max-pooling operations with a pooling size of 2×2; When inputting the normalized drone - collected night - time image X with a size of H×W, the formula for its feed - forward operation is expressed as: F 1 = C 1 (X) F 2 = C 2 (F 1 ) F 3 = P 1 (F 2 ) F 4 = C 3 (F 3 ) F 5 = P 2 (F 4 ) Among which F 1 ,…,F 5 are the characteristic outputs of the operations of each intermediate layer of the encoder; Step S232: Design the decoder structure. The decoder is composed of upsampling blocks U 1 , convolutional blocks C 4 , C 5 , upsampling blocks U 2 , convolutional blocks C 6 , C 7 , C 8 , C 9 stacked in sequence. The output of the decoder is the final result Among them, the upsampling blocks U 1 , U 2 have the same structure, that is, they are composed of 1 transposed convolutional layer, nearest neighbor linear interpolation operation I(·), and feature concatenation Cat(·) operation; the transposed convolutional layer is a transposed convolution with a kernel size of 2×2 and a stride of 2; the convolutional blocks C 6 , C 7 , C 8 , C 9 are composed of a convolutional layer and an activation layer in sequence; the convolutional layer is a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, and the activation layer is the LeakyReLU activation function, and the negative slope of this activation function is set to 0.2; among them, the feature output of the convolutional block C 5 is denoted as F 8 ; the upsampling block receives two feature inputs of different sizes and finally outputs the concatenated feature with a larger size; specifically, the upsampling block U 1 receives F 5 , F 3 as inputs, F 3 obtains F 5 ' through the transposed convolutional layer, F 3 obtains F 3 ' through the nearest neighbor linear interpolation operation I(F 3 ), F 3 ' and F 5 ' have the same size, and the feature concatenation Cat(F′ 5 , F′ 3 ) concatenates the two features along the channel dimension; similarly, the upsampling block U 2 receives F 8 , X as inputs, F 8 obtains F 8 ' through the transposed convolutional layer, X obtains X' through the nearest neighbor linear interpolation operation I(X), X' and F 8 ' have the same size, and the feature concatenation Cat(F′ 8 , X') concatenates the two features along the channel dimension.
4. The reference - free enhancement method for drones to collect images at night according to claim 1, characterized in that, In step S3, a reference - free loss function is designed for training the network designed in step S2, including the following steps: Step S31: Design the overall optimization objective of the entire network; the optimization objective is to minimize the total loss function Among them, represents the brightness control loss function, and λ bri represents the weight of the brightness control loss function; represents the color loss function, and λ col represents the weight of the color loss function; represents the image enhancement parameter smoothing loss function, and λ tv represents the weight of the image enhancement parameter smoothing loss function; represents the semantic-driven deblocking loss function, and λ blo represents the weight of the semantic-driven deblocking loss function; represents the semantic self-regularized denoising loss function, and λ noi represents the weight of the semantic self-regularized denoising loss function; Step S32: Design a brightness control loss function; The calculation formula is as follows: Among them, the preliminarily enhanced image is divided into N = 16×16, that is, 256 regions of the same size, and the average brightness value within the region is denoted as 0.6 is a preset target brightness constant value, |||| 1 is an absolute value operation; Step S33: Design a color loss function; The calculation formula is as follows: In the above formula respectively represent the preliminary enhancement results the means of the red, green, and blue channels in the RGB color space; Step S34: Design a lighting smoothing loss function; The calculation formula is as follows: Wherein, A is the color channel c corresponding to the image enhancement parameter A output by the enhanced curve estimation network, and c is taken from the RGB color space; |||| 2 is an operation to obtain the Euclidean norm; represents an operation to obtain the first-order gradient, that is and are respectively the first-order differences of A c in the vertical and horizontal directions, |||| 1 is an absolute value operation; Step S35: Design a semantic-driven deblocking loss function This deblocking loss function can suppress blocking artifacts during the image enhancement process; Step S36: Design a semantic self-regularized denoising loss function This denoising loss function can effectively suppress noise pixels during the image enhancement process.
5. According to a reference - free enhancement method for unmanned aerial vehicle (UAV) night - time image acquisition as claimed in claim 4, characterized in that In the step S35, design a semantic-driven deblocking loss function This deblocking loss function can suppress the blocking effect during the image enhancement process; it includes the following steps: Step S351: For the input normalized nighttime acquisition image X of the drone, use the VGG-16 model pre-trained for the classification task on the ImageNet dataset to extract features from X. The result of summing the output features of the Conv4-2 layer by channel is denoted as f X ; Step S352: For the preliminarily enhanced image use the VGG-16 model pre-trained on the classification task of the ImageNet dataset to extract features, and the result of summing the output features of the Conv4-2 layer by channel is denoted as Step S353: For the output result of the enhancement network designed for night-time image acquisition by the drone in step S2, that is, the final result Use the VGG-16 model pre-trained on the classification task of the ImageNet dataset for Feature extraction, and the result of summing the output features of the Conv4-2 layer by channel is denoted as Step S354: Design a semantic-driven deblocking loss function It is represented by the formula as follows: where * represents the element-by-element multiplication of matrices, e is the base of the natural logarithm, μ is a parameter that can control the perception of semantic intensity, and the default setting is 20; |||| 1 is the absolute value operation.
6. According to a reference - free enhancement method for unmanned aerial vehicle (UAV) night - time image acquisition as claimed in claim 4, characterized in that In the step S36, a denoising loss function with semantic self-regularization is designed. This denoising loss function can effectively suppress noise pixels during the image enhancement process and includes the following steps: Step S361: Use the subgraph sampling method in Neighbor2Neighbor to construct the subgraph sampler G N =(g 1 ,g 2 ); The specific method is as follows: For any input image T, its size is H T ×W T ; Divide it into tuples, where represents the floor operation, and each tuple contains 2×2 = 4 pixels; Among all tuple partitions, the pixels in its j-th tuple are respectively denoted as When the subgraph sampler G N =(g 1 ,g 2 ) samples in the j-th tuple, for g 1 and g 2 , they respectively randomly correspond to "g 1 uses row sampling, g 2 uses column sampling" or "g 1 uses column sampling and g 2 uses row sampling"; where row sampling means selecting 1 element from the set , and column sampling means selecting 1 element from the set ; Step S362: For the input normalized nighttime acquisition image X of the drone, use the VGG-16 model pre-trained for the classification task on the ImageNet dataset to extract features from X, and denote the result of summing the output features of the Conv4-2 layer by channel as f X ; Step S363, Semantic Self-Regularized Denoising Loss Function The design is as follows: where * represents the element-wise multiplication of matrices; is the subgraph difference loss, is the regularization term of the semantic self-regularized denoising loss function; The calculation method of where Φ is the perception restoration network described in step S23, is the preliminarily enhanced image g 1 and g 2 is defined by the subgraph sampler G; || || 2 is the operation of obtaining the Euclidean norm; The calculation method is as follows: where Φ is the perception restoration network described in step S23, is the initially enhanced image g 1 and g 2 is defined by the subgraph sampler G; || || 2 is the operation of obtaining the Euclidean norm.
7. According to a reference - free enhancement method for unmanned aerial vehicle (UAV) night - time image acquisition as claimed in claim 1, characterized in that In step S4, a training dataset is used to train the enhancement network for UAV night - time image acquisition, including the following steps: Step S41: In the dataset constructed through step S1, a random training image X is selected; Step S42: Training Image Encoding and Enhancement; For the input image X, the image enhancement parameter A is obtained through the enhancement curve estimation network, and the preliminarily enhanced image is obtained through the enhancement operation formula Finally, the final result is obtained through the perceptual restoration network Calculate the loss of the total loss function in step S31 Step S43: The back - propagation method is used to calculate the gradients of the parameters in the enhancement curve estimation network and the perceptual restoration network, and the Adam optimization method is used to update the parameters; Step S44: The above steps are one iteration of the training process. The entire training process requires 100 iterations. And in each iteration process, multiple image pairs are randomly sampled as a batch for training.
Citation Information
Patent Citations
No-reference low-illumination image enhancement method based on local scene perception
CN115205160A
Low-illumination video enhancement method and system and storage medium
CN115619674A