Plume spectral image contrast enhancement method based on improved U-net network
By improving the U-net network structure, using dense feature fusion module and deep multi-layer residual module, the shortcomings of image contrast enhancement in the prior art in complex goals and harsh environments are solved, and higher restoration accuracy and image quality are achieved.
Patent Information
- Application Number
- CN202510030414.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-03
AI Technical Summary
Existing deep learning-based image contrast enhancement methods do not recover image quality and detail when facing complex goals (such as plume shapes) and harsh environments (such as strong radiation backgrounds).
Improve the U-net network structure, and by setting up dense feature fusion modules and deep multi-layer residual modules, enhance the feature transmission of the encoder to the decoder, improve the model optimization effect, and apply dense feature fusion modules in the encoding and decoding stages to preserve spatial information.
Improve image contrast enhancement recovery accuracy, especially in the case of complex targets and harsh environments, significantly improving image quality and detail recovery effects.
Smart Images

Figure CN120088174A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a plume spectral image contrast enhancement method based on an improved U-net network. Background Art
[0002] Due to the influence of the environment, light, etc., the clarity and contrast of the captured photos are relatively low, and the key points in the image cannot be highlighted. Image contrast enhancement is to enhance the contrast of the image by certain means, making the targets therein more obvious, which is beneficial to subsequent recognition and other processing. Traditional image contrast enhancement algorithms are based on gray-scale transformation to achieve contrast enhancement. Typical algorithms include using linear transformation to increase the gray-scale range, using non-linear transformation to stretch or compress high or low gray-scale regions, using histogram equalization to average the gray-scale distribution, using Laplace transformation to highlight edge details, and the Retix algorithm based on optical physical characteristics, etc. These methods have good effects on specific scenes and specific target images, but have poor robustness to complex backgrounds and targets with complex shapes such as plumes.
[0003] Due to the good adaptability and intelligence of deep learning algorithms, many progresses have been made in the field of image processing in recent years. The Low-Light Network (LL-Net) introduced deep learning into the low-light image enhancement task, constructed a stacked sparse denoising autoencoder, and artificially synthesized low-light data to simulate the low-light environment, enhancing and denoising low-light noisy images. Compared with traditional algorithms, the image quality has been significantly improved, but the model structure is simple, and the advantages of deep learning are not fully utilized, and there is still great room for improvement.
[0004] The Multi-Branch Low-Light Enhancement Network (MBLLEN) proposed a multi-branch network, which respectively corresponds to various functional requirements such as brightness enhancement, contrast enhancement, denoising and artifact removal in image enhancement, and fuses features at different levels to achieve the effect of improving the image quality in multiple aspects; ALIE proposed an architecture combining an attention mechanism-guided enhancement method and a multi-branch network, guiding region-adaptive low-light enhancement and denoising by generating attention maps and noise maps. Wang et al. combined the Retinex theory and neural networks, estimated the illumination component through a convolutional neural network, adjusted the exposure degree to obtain the desired normally exposed image, and added a smoothing loss to improve the contrast and a three-channel color loss to improve the vividness; EnlightenGAN first used GAN technology in the low-light field, enabling training and learning even with unpaired data, and using local and global discriminators to handle local and global illumination conditions.
[0005] However, the existing technologies mainly have the following problems: Although the deep learning-based methods can improve the enhancement effect to a certain extent through diversified means such as changing the network structure, learning different types of features, and optimizing the loss function, there is still much room for improvement in the restoration of image quality and details in the face of complex targets (plume shapes) and harsh environments (strong radiation backgrounds). Summary of the Invention
[0006] In order to solve the problem that the existing deep learning-based methods lack targeted design for the degradation model, resulting in low restoration accuracy, and there is a certain degree of loss of spatial information in the basic U-Net structure. The purpose of the present invention is to propose a plume spectral image contrast enhancement system and method based on an improved U-net network. By improving the U-net network structure, the problems that the existing deep learning-based methods lack targeted design for the degradation model, resulting in low restoration accuracy, and there is a certain degree of loss of spatial information in the basic U-Net structure are solved.
[0007] The present invention provides a plume spectral image contrast enhancement method based on an improved U-net network, including the following steps:
[0008] Step 1: Input the initial image into the encoder to obtain initial encoded features;
[0009] Step 2: Input the initial encoded features into the encoding optimization module to obtain encoded optimization features;
[0010] Step 3: Input the encoded optimization features into the decoder to obtain an enhanced image.
[0011] Further, the specific process of the step 1: Input the initial image into the encoder to obtain initial encoded features; is as follows:
[0012] Step 201: Input the initial image into the first encoding module l 1 to obtain first optimized features;
[0013] Step 202: Input the first optimized features into the second encoding module l 2 to obtain second optimized features;
[0014] Step 203: Input the second optimized features into the third encoding module l 3 to obtain third optimized features;
[0015] Step 204: Input the third optimized features into the fourth encoding module l 4 to obtain fourth fusion features.
[0016] Further, the first encoding module l 1 includes a single-layer convolution module and a multi-layer residual module arranged in sequence; the second encoding module l2 It includes a convolutional downsampling module, a dense feature fusion module, and a multi-layer residual module arranged in sequence; the third encoding module l 3 It includes a convolutional downsampling module, a dense feature fusion module, and a multi-layer residual module arranged in sequence; the fourth encoding module l 4 It includes a convolutional downsampling module and a dense feature fusion module arranged in sequence.
[0017] Furthermore, the single-layer convolutional module is mainly used to extract local features of the image, and the calculation formula is:
[0018] ConvBlock(x) = ReLU(BatchNorm(Conv2D(x))) (9)
[0019] Where, ConvBlock is defined as the single-layer convolutional module, x represents the input image, Conv2D represents the convolutional operation, BatchNorm represents batch normalization, and the activation function is set as the ReLU function.
[0020] Furthermore, the calculation formula of the multi-layer residual module is:
[0021] i k = ConvBlock(ConvBlock(x)) + x k∈(1,..., n) (10)
[0022] Where, k represents the ordinal number of the encoding module, and i k represents the output result of the k-th level encoder l k of.
[0023] Furthermore, the calculation formula of the dense feature fusion module is:
[0024]
[0025] Where, i n is the feature directly output by the n-th level encoding layer, is the dense fusion enhanced feature output by all previous encoding layers, represents the n-th level dense fusion enhanced feature.
[0026] Furthermore, the convolutional downsampling module can effectively reduce the size of the feature map and extract more abstract high-level features. The calculation formula is as follows:
[0027]
[0028] Where, the input dimension of this convolutional operation is C k , and the output dimension is C k+1 .
[0029] Further, the encoding optimization module utilizes the VGG-16 neural network structure and uses global residuals to fuse four modules with their respective input features to obtain encoding optimization features.
[0030] Further, the calculation formula for the encoding optimization features:
[0031]
[0032] Where, represents the k-th level of densely fused enhanced features, vgg represents the VGG-16 neural network, and its output result performs a residual calculation with its input to obtain the k-th level of encoding optimization feature i k .
[0033] Further, the specific process of step 3, inputting the encoding optimization features into the decoder to obtain the enhanced image is as follows:
[0034] Step 501: Input the fourth encoding optimization feature into the first decoding module j 1 to obtain the first decoding feature;
[0035] Step 502: Input the first decoding feature and the third encoding optimization feature into the second decoding module j 2 to obtain the second decoding feature;
[0036] Step 503: Input the second decoding feature and the second encoding optimization feature into the third decoding module j 3 to obtain the third decoding feature;
[0037] Step 504: Input the third decoding feature and the first encoding optimization feature into the fourth decoding module j 4 to obtain the contrast enhancement output feature.
[0038] Further, the first decoding module j 1 includes a feature extraction module and an atrous convolution module arranged in sequence; the second decoding module j 2 includes an enhanced feature fusion module, a dense feature fusion module, and a convolutional upsampling module arranged in sequence; the third decoding module j 3 includes an enhanced feature fusion module, a dense feature fusion module, and a convolutional upsampling module arranged in sequence; the fourth decoding module j 4 includes an enhanced feature fusion module, a dense feature fusion module, and a single-layer convolutional module arranged in sequence.
[0039] Further, the feature extraction module decodes effective feature information from the optimized features output by the encoder using convolution. Dilated convolution enlarges the receptive field of the convolution kernel, enabling the network to capture context information over a larger range. The calculation formula is as follows:
[0040]
[0041] where ConD represents dilated convolution and α represents the dilation rate. j n-1 ′ represents the result after feature extraction and dilated convolution in the first-layer decoder.
[0042] Further, each level of the decoding module is entirely composed of a 3-layer fully convolutional residual structure. The decoding calculation formula is as follows:
[0043]
[0044] where represents the decoding result of the (n - 1)-th level, θ n-1 represents the optimizable parameters in the decoder, j n ′ is the result of the encoded optimized features after feature extraction and dilated convolution.
[0045] Further, in the second decoding module j 2 and the third decoding module j 3 the upsampling module is used to enlarge the size of the feature map to match the size of the encoded optimized result at the corresponding level. The calculation formula is as follows:
[0046] j n-1 ′ = ↑ 2 (j n-1 ) (15)
[0047] where ↑ 2 represents the upsampling operator with a scale factor of 2, j n-1 ′ and i n-1 have the same size and dimension.
[0048] j n-1 represents the output result of the decoder module at the (n - 1)-th level.
[0049] Further, the calculation formula of the enhanced feature fusion module is as follows:
[0050]
[0051] where ↑ 2 represents the upsampling operator with a scale factor of 2, represents the decoding module composed of a 3-layer fully convolutional residual structure, θ nGenerally refers to the optimization parameters in the trainable structure; n represents the feature level, and CMR(i n-1 ) represents the encoding correction of the i n-1 -level features.
[0052] Furthermore, the calculation formula defined by the dense feature fusion module is:
[0053]
[0054] where j n is the output of the nth-level enhanced feature of the decoder, is the enhanced output after feature fusion, represents the dense connection of the enhanced feature outputs of all previous levels before this level of decoder.
[0055] Furthermore, the calculation formula of the loss function is:
[0056]
[0057] where W * is the perceptual loss, which is achieved through the l 2 loss of the intermediate feature layer of the pre-trained VGG-16, y is the true reference data in the GT dataset, is the enhanced data generated by the model.
[0058] The advantages of the present invention are as follows: The present invention provides this method for enhancing the contrast of plume spectral images based on the improved U-net network. By setting up a dense feature fusion module, it is applied to both the encoding and decoding stages. A deep multi-layer residual module is designed to address the problem of incomplete fusion of different-scale features between the same-scale encoding-decoding levels, replacing the skip connections in the original U-Net structure, strengthening the feature screening transmitted from the encoder to the decoder, improving the model optimization effect, and reducing the network training time and error. Since the present invention is highly targeted at the degradation model and has high restoration accuracy, it has good performance in enhancing the contrast of plume spectral images.
[0059] The encoder uses the form of residual groups to construct the refinement units for each level of reconstruction, and the decoder contains a total of 5 residual groups; the encoding optimization module adopts the VGG-16 neural network structure, which has a relatively deep network structure and a relatively large number of convolutional layers and pooling layers, enabling the network to learn more high-level abstract features than the output of the encoder to achieve the optimization effect on the initial encoding features; at the end of the decoder, a single convolutional layer is used to reconstruct the estimated enhanced image from the final features, and an error feedback mechanism of progressive resampling is used, which can better extract the spatial high-frequency information of the image from the previous-level scale space. By gradually performing reconstruction residual fusion, the spatial information lost during the resampling process can be effectively retained.
[0060] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0061] Figure 1 It is a flowchart of a plume spectral image contrast enhancement method based on an improved U-net network.
[0062] Figure 2 It is an optimized diagram of the decoder grading.
[0063] Figure 3 It is a network structure diagram of the nth-level MSDF module.
[0064] Figure 4 It is the original image before enhancement.
[0065] Figure 5 It is an image diagram of the plume spectral image contrast enhancement method based on an improved U-net network after enhancement. Detailed Embodiments
[0066] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the specific embodiments, structural features, and their effects of the present invention are described in detail below with reference to the accompanying drawings and embodiments.
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0068] All features disclosed in this specification, or all steps in the disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
[0069] Any feature disclosed in this specification (including any additional claims, abstract, and drawings), unless specifically recited, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically recited, each feature is only an example of a series of equivalent or similar features.
[0070] Embodiment 1
[0071] This embodiment provides a plume spectral image contrast enhancement method based on an improved U-net network. By improving the U-net network structure, it solves the problems that the existing deep learning-based methods lack targeted design for the degradation model, resulting in low restoration accuracy, and there is a certain degree of loss of spatial information in the basic U-Net structure.
[0072] First, introduce the plume spectral image contrast enhancement system based on the improved U-net network, which includes an encoder for initial feature extraction, optimization, fusion, and dense connection of the input image, and finally obtaining the initial encoded features; an encoding optimization module for learning higher-level abstract features based on the initial encoded features and outputting the optimized encoded features to the decoder; and a decoder for decoding, upsampling, dense connection, and feature fusion of the obtained optimized encoded features to obtain the final contrast enhancement output.
[0073] The encoder uses the form of residual groups to construct the refinement units for reconstruction at each level. The decoder contains a total of 5 residual groups; the encoding optimization module adopts the VGG-16 neural network structure, which has a relatively deep network structure and a relatively large number of convolutional layers and pooling layers, enabling the network to learn more high-level abstract features than the output of the encoder to achieve the optimization effect of the initial encoded features; at the end of the decoder, a single convolutional layer is used to reconstruct the estimated enhanced image from the final features, and an error feedback mechanism of progressive resampling is used, which can better extract the spatial high-frequency information of the image from the previous scale space. By gradually performing reconstruction residual fusion, the spatial information lost during the resampling process can be effectively retained.
[0074] There are a total of 4 different-scale encoding modules with a downsampling rate of 2 in the encoder stage, denoted as the first encoding module l 1 , the second encoding module l 2 , the third encoding module l 3 and the fourth encoding module l 4 . The first encoding module l 1 performs initial feature extraction on the input image using a single-layer convolution. The number of convolution kernels used is 16, the stride is 1, and zero-padding is performed on the edges. Then, a multi-layer residual module is used to optimize the initial features, obtaining 16 optimized features and inputting them into the l2 module.
[0075] The second encoding module l 2 first performs a spatial scale transformation on the output of the first encoding module l 1 using a single-layer convolution. The number of convolution kernels used is 32, the stride is 2, and no zero-padding is performed on the edges. Then, a dense feature fusion module is used to fuse the initial features. Finally, the same multi-layer residual module as in the first encoding module l 1 is used to deeply optimize the fused features and input them into the third encoding module l 3 . The third encoding module l 3 and the fourth encoding module l 4 use the same operations as the second encoding module l 2 and perform dense connection on all the features of the previous layers, finally obtaining the initial encoded features f ori .
[0076] The encoding optimization module uses VGG-16 to optimize the initial encoding features, and its spatial scale is the same as that of the fourth encoding module l 4 The output features are consistent, the number of feature channels is 256, and global residuals are used to fuse the four modules with their respective input features. For the first encoding module l 1 , the second encoding module l 2 , the third encoding module l 3 The encoded outputs are connected to the corresponding decoder-level outputs in the same way.
[0077] The decoder uses a structure that is basically symmetric to the encoder, including 4 decoding modules with different scales and an upsampling rate of 2, denoted as the first decoding module j 1 , the second decoding module j 2 , the third decoding module j 3 and the fourth decoding module j 4 . The first decoding module j 1 uses dilated convolution to perform initial decoding upsampling on the output of the encoding optimization module. For the second decoding module j 2 First, the feature enhancement and fusion module is used to fuse the output of the third encoding module l 3 module, then the dense feature fusion module is used to connect the decoder output, and finally, a single-layer convolution is used to optimize the fused features and input them into the third decoding module j 3 . The third decoding module j 3 and the fourth decoding module j 4 use the same connection method as the second decoding module j 2 to densely connect the outputs of all previous decoding stages, and finally, a single-layer convolution is used for feature fusion, with the number of convolution kernels being 16, to obtain the final contrast-enhanced output. All convolution kernels in this network are 3×3 in size, the default stride is 1, and the number of feature channels at each level of the network is input-output: [1, 16, 32, 64, 128, 256, 128, 64, 32, 16, 1].
[0078] As Figure 1 shown, the plume spectral image contrast enhancement method based on the improved U-net network includes the following steps:
[0079] Step 1: Input the initial image into the encoder to obtain the initial encoding features;
[0080] Step 2: Input the initial encoding features into the encoding optimization module to obtain the encoding optimization features;
[0081] Step 3: Input the encoding optimization features into the decoder to obtain the enhanced image.
[0082] Further, the specific process of Step 1, inputting the initial image into the encoder to obtain the initial encoded features, is as follows:
[0083] Step 201, input the initial image into the first encoding module l 1 to obtain the first optimized feature;
[0084] Step 202, input the first optimized feature into the second encoding module l 2 to obtain the second optimized feature;
[0085] Step 203, input the second optimized feature into the third encoding module l 3 to obtain the third optimized feature;
[0086] Step 204, input the third optimized feature into the fourth encoding module l 4 to obtain the fourth fused feature.
[0087] The encoder includes the first encoding module l 1 , the second encoding module l 2 , the third encoding module l 3 and the fourth encoding module l 4 . The first encoding module l 1 is used to perform initial feature extraction on the input image using single-layer convolution, and use a multi-layer residual module to optimize the initial features, obtain the first optimized feature and input it into the second encoding module l 2 ;
[0088] The second encoding module l 2 is used to perform spatial scale transformation on the first optimized feature using single-layer convolution, and use a dense feature fusion module to fuse the first optimized feature, obtain the second fused feature, and finally use a multi-layer residual module to perform deep optimization on the second fused feature, obtain the second optimized feature and input it into the third encoding module l 3 ;
[0089] The third encoding module l 3 is used to perform spatial scale transformation on the second optimized feature using single-layer convolution, and use a dense feature fusion module to fuse the second optimized feature, obtain the third fused feature, and finally use a multi-layer residual module to perform deep optimization on the third fused feature, obtain the third optimized feature and input it into the fourth encoding module l 4 ;
[0090] The fourth encoding module l 4 is used to perform spatial scale transformation on the third optimized feature using single-layer convolution, and use a dense feature fusion module to fuse the third optimized feature, obtain the fourth fused feature;
[0091] Densely connect all the features of the previous layer to finally obtain the initial encoded features.
[0092] The first encoding module l 1 , the second encoding module l 2 , the third encoding module l 3 and the fourth encoding module l 4 have different encoding scales.
[0093] Furthermore, the first encoding module l 1 includes a single-layer convolution module and a multi-layer residual module arranged in sequence; the second encoding module l 2 includes a convolutional downsampling module, a dense feature fusion module, and a multi-layer residual module arranged in sequence; the third encoding module l 3 includes a convolutional downsampling module, a dense feature fusion module, and a multi-layer residual module arranged in sequence; the fourth encoding module l 4 includes a convolutional downsampling module and a dense feature fusion module arranged in sequence.
[0094] Furthermore, the single-layer convolution module is mainly used to extract local features of the image, and the calculation formula is:
[0095] ConvBlock(x) = ReLU(BatchNorm(Conv2D(x))) (9)
[0096] where ConvBlock is defined as the single-layer convolution module, x represents the input image, Conv2D represents the convolution operation, BatchNorm represents batch normalization, and the activation function is set to the ReLU function.
[0097] Furthermore, the calculation formula of the multi-layer residual module is:
[0098] i k = ConvBlock(ConvBlock(x)) + x k∈(1,..., n) (10)
[0099] where k represents the encoding module ordinal number, and i k represents the output result of the k-th level encoder l k .
[0100] Furthermore, the calculation formula of the dense feature fusion module is:
[0101]
[0102] where i n is the feature directly output by the n-th level encoding layer, is the densely fused and enhanced feature output by all the previous encoding layers, Represents the nth-level dense fusion enhanced feature.
[0103] Furthermore, the convolutional downsampling module can effectively reduce the size of the feature map and extract more abstract high-level features. The calculation formula is as follows:
[0104]
[0105] Among them, The input dimension of this convolutional operation is C k , and the output dimension is C k+1 .
[0106] Furthermore, the encoding optimization module utilizes the VGG-16 neural network structure and uses global residuals to fuse the four modules with their respective input features to obtain encoded optimization features; the spatial scale of the encoding optimization module is the same as the spatial size of the output features of the four encoding modules. For the first encoding module l 1 , the second encoding module l 2 , the third encoding module l 3 and the fourth encoding module l 4 's encoded outputs, all use global residuals to fuse the four modules with their respective input features and connect them to the corresponding decoder-level outputs.
[0107] Specifically, input the first optimized feature into the encoding optimization module to obtain the first encoded optimization feature;
[0108] Input the second optimized feature into the encoding optimization module to obtain the second encoded optimization feature;
[0109] Input the third optimized feature into the encoding optimization module to obtain the third encoded optimization feature;
[0110] Input the fourth fusion feature into the encoding optimization module to obtain the fourth encoded optimization feature.
[0111] Furthermore, the calculation formula for the encoded optimization feature:
[0112]
[0113] Among them, Represents the kth-level dense fusion enhanced feature, vgg represents the VGG-16 neural network, and its output result performs a residual calculation with its input to obtain the kth-level encoded optimization feature i k .
[0114] Furthermore, the specific process of step 3, inputting the encoded optimization feature into the decoder to obtain the enhanced image is:
[0115] Step 501: Input the fourth encoded optimization feature into the first decoding module j 1 Obtain the first decoding feature;
[0116] Step 502: Input the first decoding feature and the third encoded optimization feature into the second decoding module j 2 Obtain the second decoding feature;
[0117] Step 503: Input the second decoding feature and the second encoded optimization feature into the third decoding module j 3 Obtain the third decoding feature;
[0118] Step 504: Input the third decoding feature and the first encoded optimization feature into the fourth decoding module j 4 Obtain the contrast-enhanced output feature.
[0119] The decoder includes the first decoding module j 1 , the second decoding module j 2 , the third decoding module j 3 and the fourth decoding module j 4 , the first decoding module j 1 is used to perform initial decoding upsampling on the output of the encoded optimization module using dilated convolution, and the second decoding module j 2 uses the enhanced feature fusion module to fuse the output of the third encoding module l 3 , then uses the dense feature fusion module to connect the decoder output, and finally uses a single-layer convolution to optimize the fused features and input them into the third decoding module j 3 ; j 3 、j 4 modules use the same connection method as the j 2 module to densely connect the outputs of all previous decoding stages, and finally use a single-layer convolution for feature fusion to output the contrast-enhanced image.
[0120] The encoder uses the form of residual groups to construct the refinement units for each level of reconstruction.
[0121] The structure of the decoder is basically symmetric to that of the encoder.
[0122] Furthermore, the first decoding module j 1 includes a feature extraction module and a dilated convolution module arranged in sequence; the second decoding module j 2 includes an enhanced feature fusion module, a dense feature fusion module, and a convolutional upsampling module arranged in sequence; the third decoding module j 3 includes an enhanced feature fusion module, a dense feature fusion module, and a convolutional upsampling module arranged in sequence; the fourth decoding module j 4It includes an enhanced feature fusion module, a dense feature fusion module, and a single-layer convolution module that are set in sequence.
[0123] Furthermore, the feature extraction module decodes effective feature information from the optimized features output by the encoder using convolution. Dilated convolution enlarges the receptive field of the convolution kernel, enabling the network to capture context information over a larger range. The calculation formula is:
[0124]
[0125] where ConD represents dilated convolution, α represents the dilation rate, and j n-1 ′ represents the result after feature extraction and dilated convolution in the first layer of the decoder.
[0126] Furthermore, each level of the decoding module is entirely composed of a 3-layer fully convolutional residual structure. The decoding calculation formula is as follows:
[0127]
[0128] where represents the decoding result of the (n - 1)th level, and θ n-1 represents the optimizable parameter in the decoder, and j n ′ is the result of the encoded optimized features after feature extraction and dilated convolution.
[0129] Furthermore, in the second decoding module j 2 and the third decoding module j 3 , the upsampling module is used to enlarge the size of the feature map to match the size of the encoded optimized result at the corresponding level. The calculation formula is as follows:
[0130] j n-1 ′ = ↑ 2 (j n-1 ) (15)
[0131] where ↑ 2 represents the upsampling operator with a scale factor of 2, and j n-1 ′ and i n-1 have the same size and dimension.
[0132] j n-1 represents the output result of the decoder module at the (n - 1)th level.
[0133] Furthermore, the calculation formula of the enhanced feature fusion module is:
[0134]
[0135] where ↑ 2 represents the upsampling operator with a scale factor of 2, Denotes a decoding module composed of a 3-layer fully convolutional residual structure, θ n Generally refers to the optimization parameters in the trainable structure; n represents the feature level, and CMR(i n-1 ) represents the encoding correction of the i n-1 -level features.
[0136] Furthermore, the calculation formula defined by the dense feature fusion module is:
[0137]
[0138] Among them, j n is the output of the nth-level enhanced feature of the decoder, is the enhanced output after feature fusion, represents the dense connection of the enhanced feature outputs of all previous levels before this level of decoder.
[0139] Furthermore, the calculation formula of the loss function is:
[0140]
[0141] Among them, W * is the perceptual loss, which is realized through the l 2 loss of the intermediate feature layer of the pre-trained VGG-16. y is the true reference data in the GT dataset, is the enhanced data generated by the model.
[0142] Adjacent-level progressive enhancement
[0143] The idea of the Boosting Algorithm is to refine the enhanced image based on the previously estimated image. In the image enhancement task, the main purpose is to eliminate the effects of the degradation matrix A and the introduced environmental noise n on the original high-contrast image. Traditional algorithms will cause the loss of some detailed information of the original high-contrast image during the elimination process of A and N. From the perspective of energy conservation, the optimized image of each time is combined with the original image to be processed recursively as the next input, obtaining the joint representation of the updated adjacent two iterative results, and updating the current iterative result to participate in the next recursive calculation. Each image calculation is combined with the optimized result of the previous adjacent step, gradually narrowing the target search space until the previously set convergence threshold condition is met, and finally obtaining the optimized output. This algorithm model can be expressed as Equation (5).
[0144]
[0145] Among them, represents the estimated image at the nth iteration, and g(·) is the image optimization method, Indicates the joint input using the input image I(x) and the image enhanced at the nth time. A low-contrast image can be regarded as the result of the linear superposition of the degraded part A on the original image. The degraded part is defined as R(J(x))=(1 - T)A / J(x), and for the same low-contrast image, the degradation degree ratio between its iterative images at each level is (1 - T).
[0146] When performing the R operation on images of the same scene, better enhancement results can be obtained in each iteration, but the improvement in image contrast is relatively small. That is, if J 1 (x) and J 2 (x) are images of the same scene, the following conditions are satisfied.
[0147]
[0148] In the boosting algorithm, expanding Equation (6) into each iterative process can be expressed as Equation (7).
[0149]
[0150] Based on the hypothesis conditions proposed by Equations (6) and (7), a deep learning-based feature enhancement fusion module is constructed based on the boosting algorithm model to learn the optimized output pattern between levels in end-to-end training.
[0151] In the U-Net network for image contrast enhancement, the decoder mainly restores the original high-contrast image from the encoded data. In order to perform step-by-step optimization on j L , the Boosting strategy and the encoding optimization module are embedded in the decoder network. The original corresponding scale connection mode between the encoder and decoder in U-Net and the structure of the sampled optimized Boosting module are as Figure 2 shown.
[0152] As Figure 2 (a) is the feature fusion model between the same-scale encoder-decoder modules in the original U-Net network. By merging the i n -level features in the encoder and the j n -level features in the decoder and sending them into the decoding module for optimization, the decoding module is a 3-layer fully convolutional residual network, which is used to optimize and reconstruct the features at this level to obtain the j n+1 -level decoded feature output. As Figure 2 (b) is the Boosting structure, which satisfies the mathematical model represented by Equation (5) by adding and subtracting residual connections between the j n -level features and the j n+1 -level features and changing the feature merging method in the original U-Net to the addition of spatially corresponding signals. On the basis of Boosting, further optimize the i output by the encodern Level features, such as Figure 2 (c), through the CMR module, using the spatial attention mechanism to encode and correct the i n level features, and using the corrected map as the level reconstruction output in Equation (5). Based on the designed decoder hierarchical optimization mode, namely the feature enhancement fusion module, it can be expressed in the form shown in Equation (2).
[0153]
[0154] Among them, ↑ 2 represents the upsampling operator with a scale factor of 2, represents the decoding module composed of 3 layers of fully convolutional residual structures, θ n generally refers to the optimization parameters in the trainable structure; n represents the feature level, CMR(i n-1 ) represents encoding and correcting the i n-1 level features.
[0155] Based on the above enhancement module design, the designed encoder uses the form of residual groups to construct the refinement units for each level of reconstruction, and there are 5 residual groups in the decoder. At the end of the decoder, a single convolutional layer is used to reconstruct the estimated enhanced image from the final features.
[0156] Dense feature fusion module
[0157] The performance of the U-Net architecture is limited in many aspects. For example, there is a loss of spatial information during the downsampling process of the encoder, and there is a lack of sufficient correlation representation between features at non-adjacent levels. For features at non-adjacent levels, all features can be first resampled to the same scale, downsampling for the encoder and upsampling for the decoder, and then fused with a dense layer, including 1 dense connection layer and 1 convolutional layer, to construct a dense network block (DenseNet). However, due to the different spatial scales and spatial sizes of features at different levels, there is still a partial loss of spatial information during the resampling process, and simply using the resampling cascade method still has certain limitations for improving the performance of feature fusion.
[0158] Regarding the problem of spatial information loss caused by spatial scale transformation in the encoding and decoding model, the present invention combines the super-resolution reconstruction related model to optimize the design of the feature information protection mechanism in this process. In the super-resolution reconstruction method, the back-projection technique generates a high-resolution image by minimizing the reconstruction error between the estimated high-resolution result and multiple observed low-resolution inputs Construct a multi-scale spatial information fusion model to ensure the effective reconstruction of spatial detail information.
[0159] In the back-projection algorithm, for a single-frame low-resolution image input, it satisfies the computational model shown in Equation (8).
[0160]
[0161] Among them, represents the high-resolution output estimated at the t-th iteration, L ob represents the low-resolution image observed using the downsampling operator f, and h represents the back-projection operator.
[0162] Based on the back-projection algorithm in (8), the present invention proposes a Cross-Scale Dense Feature Fusion (MSDF) module to effectively retain the information lost during the resampling process and utilize the features of non-adjacent layers. The MSDF module is applied to both the U-Net encoder and decoder simultaneously. Through an error feedback mechanism, the feature enhancement performance at the current level is improved. As Figure 3 shown, 1 MSDF module is used in both the encoder and decoder. One is before the multi-layer residual block in the encoder, and the other is after the enhanced feature fusion module in the decoder. Moreover, the outputs of the MSDF modules in the encoder / decoder are directly connected to the MSDF modules at all subsequent levels to form a dense connection structure to achieve dense feature fusion. Taking the n-th level of the decoder as an example, the network structure of the MSDF module is as Figure 3 shown.
[0163] The definition of the MSDF at the n-th level of the decoder is shown in Equation (3).
[0164]
[0165] Among them, j n is the enhanced feature output at the n-th level of the decoder, is the enhanced output after feature fusion, represents the dense connection of the enhanced feature outputs at all previous levels before the current level of the decoder. First, for the output of the feature enhancement fusion module, a multi-layer convolutional downsampling module (with a convolutional kernel size of 3×3 and a stride of 2) is used to gradually downsample the features at the j n level to obtain a feature space with the same spatial scale as j 0 , and based on Equation (8), the output features of j 0 are connected using a residual connection. Then, a convolutional upsampling module (a dilated convolution with a size of 3×3 and a stride of 1) is used to gradually upsample the features at the j 0 spatial scale to the j n spatial scale, and finally, the enhanced feature output at the n-th level For the same method is used for dense connection to obtain the outputs at each level.
[0166] For the encoder, the dense fusion feature at its n-th level can be expressed as Equation (1).
[0167]
[0168] where i n is the feature directly output by the n-th level encoding layer, is the dense fusion enhanced feature output by all previous encoding layers, represents the dense fusion enhanced feature at the n-th level. This process also follows the feature scale space transformation rule as Figure 3 shown.
[0169] Compared with the method of directly downsampling and fusing high-scale space features, this method uses a step-by-step resampling error feedback mechanism, which can better extract the spatial high-frequency information of the image from the previous scale space. By gradually performing reconstruction residual fusion, it can effectively retain the spatial information lost during the resampling process.
[0170] Loss function
[0171] To reduce the overfitting of the model to the reference data and considering the certain difference between subjective image quality improvement and reference-based loss, perceptual loss and mean squared error loss (MSELoss) are jointly used for network training, and the overall loss function is expressed in the form shown in Equation (4).
[0172]
[0173] where W * is the perceptual loss, which is achieved through the l 2 loss of the intermediate feature layer of the pre-trained VGG-16, y is the true reference data in the GT dataset, is the enhanced data generated by the model.
[0174] In summary, the plume spectral image contrast enhancement method based on the improved U-net network provided in this embodiment, by setting a dense feature fusion module, which is applied to both the encoding and decoding stages, and aiming at the problem of incomplete fusion of different scale features between the same scale encoding-decoding levels, designs a deep multi-layer residual module to replace the skip connection in the original U-Net structure, strengthens the feature screening transmitted from the encoder to the decoder, improves the model optimization effect, and at the same time reduces the network training time and error. Since the present invention has strong pertinence to the degradation model and high restoration accuracy, it has good performance in plume spectral image contrast enhancement.
[0175] Embodiment 2
[0176] Using the image contrast enhancement method based on the improved U-net network shown in Embodiment 1, forFigure 4 Process the image shown as follows.
[0177] Image acquisition
[0178] Select an infrared image with low-contrast characteristics as the input sample. The image has a size of 640x512 pixels and a depth of 1 (grayscale image).
[0179] Network configuration
[0180] Use the improved U-net network structure described above. The network includes an encoder, an encoding optimization module (based on the VGG-16 structure), and a decoder. The network is implemented through the PyTorch framework and deployed in a computing environment with an NVIDIA RTX 3080 GPU.
[0181] Preprocessing
[0182] The input image is first converted into a tensor format acceptable to the model and normalized so that the pixel value range is mapped from [0, 255] to [-1, 1].
[0183] Model inference
[0184] Input the preprocessed image into the improved U-net network, and the network outputs an enhanced image tensor.
[0185] Post-processing: Denormalize the tensor output by the network, convert it back to the pixel value range of [0, 255], and convert it into a visual image format.
[0186] Through the above steps, we obtain the enhanced image as Figure 5 shown. Compared with the original image, the enhanced image visually shows higher contrast and clearer details. Through the comparison of objective evaluation metrics (such as the structural similarity index SSIM and the peak signal-to-noise ratio PSNR), the quality of the enhanced image is significantly better than that of the original image.
[0187] Through the above image contrast enhancement method based on the improved U-net network, the effectiveness of the network in contrast enhancement can be demonstrated. The network can not only improve the visual quality of the image but also retain important detail information, which has important application value in the field of image analysis and processing.
[0188] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A plume spectral image contrast enhancement method based on an improved U-net network, characterized in that: The steps include: Step 1, input the initial image into the encoder to obtain the initial encoding features; Step 2: Input the initial coding features into the coding optimization module to obtain coding optimization features; Step 3: Input the encoded optimized features into the decoder to obtain an enhanced image.
2. The plume spectral image contrast enhancement method based on the improved U-net network according to claim 1, characterized in that: The step 1, inputting the initial image into the encoder to obtain the initial encoding feature; The specific process is: Step 201: input the initial image into the first encoding module l1 to obtain the first optimized feature; Step 202: input the first optimized feature into the second encoding module l2 to obtain the second optimized feature; Step 203: input the second optimized feature into the third encoding module l3 to obtain the third optimized feature; Step 204: input the third optimized feature into the fourth encoding module l4 to obtain the fourth fusion feature.
3. The plume spectral image contrast enhancement method based on the improved U-net network according to claim 2, characterized in that: The first encoding module l1 includes a single-layer convolution module and a multi-layer residual module arranged in sequence; the second encoding module l2 includes a convolution downsampling module, a dense feature fusion module and a multi-layer residual module arranged in sequence; the third encoding module l3 includes a convolution downsampling module, a dense feature fusion module and a multi-layer residual module arranged in sequence; the fourth encoding module l4 includes a convolution downsampling module and a dense feature fusion module arranged in sequence.
4. The plume spectral image contrast enhancement method based on the improved U-net network according to claim 3, characterized in that: The calculation formula of the dense feature fusion module is: Among them, i n is the feature directly output by the n-th level encoding layer, is the dense fusion enhancement feature output by all previous encoding layers, Represents the n-th level dense fusion enhanced features.
5. The plume spectral image contrast enhancement method based on an improved U-net network according to claim 1, characterized in that: The coding optimization module utilizes the VGG-16 neural network structure and uses the global residual to fuse the four modules with the respective input features to obtain coding optimization features.
6. The plume spectral image contrast enhancement method based on an improved U-net network according to claim 1, characterized in that: The specific process of step 3, inputting the coding optimization feature into the decoder to obtain the enhanced image is: Step 501: input the fourth encoding optimization feature into the first decoding module j1 to obtain the first decoding feature; Step 502: Input the first decoding feature and the third encoding optimization feature into the second decoding module j2 to obtain a second decoding feature; Step 503: input the second decoding feature and the second encoding optimization feature into the third decoding module j3 to obtain a third decoding feature; Step 504: Input the third decoding feature and the first encoding optimization feature into the fourth decoding module j4 to obtain a contrast enhancement output feature.
7. The plume spectral image contrast enhancement method based on the improved U-net network according to claim 6, characterized in that: The first decoding module j1 includes a feature extraction module and a hole convolution module arranged in sequence; the second decoding module j2 includes an enhanced feature fusion module, a dense feature fusion module and a convolution upsampling module arranged in sequence; the third decoding module j3 includes an enhanced feature fusion module, a dense feature fusion module and a convolution upsampling module arranged in sequence; the fourth decoding module j4 includes an enhanced feature fusion module, a dense feature fusion module and a single-layer convolution module arranged in sequence.
8. The plume spectral image contrast enhancement method based on the improved U-net network according to claim 7, characterized in that: The calculation formula of the enhanced feature fusion module is: Among them, ↑2 represents the upsampling operator with a scaling factor of 2. represents the decoding module consisting of a 3-layer fully convolutional residual structure, θ n Generally refers to the optimization parameters in the trainable structure; n represents the feature level, CMR(i n-1 ) indicates i n-1 Level features are coded and corrected.
9. The plume spectral image contrast enhancement method based on an improved U-net network according to claim 7, characterized in that: The calculation formula defined by the dense feature fusion module is: Among them, j n is the nth level enhanced feature output of the decoder, is the enhanced output after feature fusion, Represents the dense connection of the enhanced feature outputs of all previous levels of the decoder.
10. The plume spectral image contrast enhancement method based on an improved U-net network according to claim 1, characterized in that: The calculation formula of the loss function is: Among them, W * is the perceptual loss, which is implemented by the l2 loss of the pre-trained VGG-16 intermediate feature layer. y is the real reference data in the GT dataset. Augmented data generated for the model.