Image enhancement network based on multi-scale feature fusion
By introducing multi-scale feature fusion and feature guidance technology into the image enhancement network, the problem of poor image enhancement effect in low-light environments is solved, and image color and details are enhanced simultaneously, the calculation load is reduced, and the original color information is retained.
Patent Information
- Application Number
- CN202411099552.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-12
AI Technical Summary
The existing image enhancement methods are not effective in low-light environments, especially in complex image processing, where the enhancement effect is reduced and it relies too much on manual parameter adjustment, and ignores the impact of the color signal characteristics of the reflective component on the enhanced image.
An image enhancement network based on multi-scale feature fusion is proposed, including feature guidance enhancement module, feature learning module, decomposition sub-network and color adjustment module. Through technical means such as cross feature attention and gated feature feedforward block, the calculation amount of the feature extraction process is reduced and the feature information is effectively fused and converged.
The network can simultaneously enhance the color and details of the image, improve the image quality in low-light environments, reduce the computing load, and effectively retain the original color information of the image.
Smart Images

Figure CN119067900B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an image enhancement network based on multi-scale feature fusion. Background Art
[0002] The development of artificial intelligence has placed increasingly stringent requirements on the quality of collected images. Although current imaging technology is developing rapidly, obtaining high-quality images is still affected by external factors, especially at night and in low-light environments. In most cases, people try to obtain enhanced images by delaying the exposure time or using artificial light to assist. However, this may cause blurry and distorted images, and if overexposed, the image details may be lost. In order to address the influence of external factors, after decades of development, many enhancement methods for low-light environments have been developed.
[0003] The enhancement effect of traditional enhancement methods drops sharply when facing complex images (such as high noise, low contrast, etc.) and they rely too much on manual parameter adjustment. Recently, people have studied the enhancement method of deep learning based on Retinex theory. By decomposing the input image into an illumination component containing the image brightness distribution and a reflection component containing rich color information of the image, and enhancing and denoising them separately. However, most of these methods simply process the reflection component and do not consider the influence of color signals on the enhanced image, which will cause unnatural colors and artifacts in the image. Therefore, it is necessary to strengthen the extraction of color details to retain the original color information of the image. Existing deep learning enhancement methods usually adopt three architectures, namely encoder-decoder structure, high-resolution image feature processing, and multi-scale cross-resolution structure. However, in the encoder-decoder structure, the previous U-Net model ignores the use of small resolution features and does not effectively collect the feature information of jump connections, resulting in the loss of some spatial information in the whole process, and it is difficult to recover it by subsequent measures. Based on the above analysis, the shortcomings of the enhancement method based on Retinex and codec architecture can be simply summarized as follows: ① The degree of enhancement of potential detail information is low. ② The enhancement method based on the codec structure lacks the effective use of the underlying resolution and jump connection features. ③ The influence of the color signal characteristics of the reflection component on the enhanced image is ignored. Based on the above problems, the present invention proposes an image enhancement network based on multi-scale feature fusion. Summary of the invention
[0004] The purpose of the present invention is to provide an image enhancement network based on multi-scale feature fusion to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the image enhancement network based on multi-scale feature fusion includes:
[0006] The feature-guided enhancement module, which consists of a cross-feature attention (CFA) and a gated feature feed-forward block (GFFB), is used to reduce the computational complexity of the feature extraction process based on the number of channels;
[0007] The feature learning module consists of a feature enhancement block FEB and a feature convergence block FCB, with the feature L guided between each layer. t As input, it first passes through the feature enhancement block for adaptive enhancement to obtain the preliminary enhancement result. It is then sent to the feature convergence block to Perform fusion convergence to obtain a clear and convergent enhanced feature map The whole process shares parameters, and the feature learning module maps the feature output of each layer All converge to the same state and participate in the subsequent upsampling and fusion operations of each layer of features;
[0008] The decomposition sub-network consists of three Conv+LeakyReLU, one Conv+Sigmoid and CA, which is used to denoise the image before decomposition to obtain the illumination component and the reflection component;
[0009] The color adjustment module includes a global feature extraction block and a feature cross-learning block. The global extraction module GFB is embedded in the bottom layer of the U-Net model to retain the features, and the feature cross-learning block is used to fuse the information of each upsampling and jump connection to extract color features.
[0010] Preferably: the CFA of the feature-guided enhancement module can simultaneously accept three feature inputs, namely, the initial and optimized feature maps of the bottom layer With the initial feature map of the previous layer
[0011] Then, the features from the bottom layer are upsampled and processed with the features from the previous layer by 3×3Conv+LeakyReLU+1×1Conv, and then the dimension is transformed to generate
[0012]
[0013] The process can be expressed as:
[0014]
[0015] Among them, R represents the dimension transformation Reshape, up represents the upsampling operation, and C 1×1 represents a 1×1 convolution kernel, C 3×3 Represents a 3×3 convolution kernel. Then, matrix multiplication is used to calculate the cross-channel covariance matrix between Q and K to obtain a size of P 256×256The obtained attention map P expresses the changes of the underlying resolution features before and after enhancement at the same pixel position, and is used to guide V to perform feature enhancement. The attention map P is matrix multiplied with V in the same way to obtain a size of The feature enhancement map of , the process of obtaining it can be expressed as:
[0016]
[0017] Among them, Softmax represents the Softmax function, represents matrix multiplication, and c represents the learning scaling parameter that can control the size. Then, the feature enhancement map is transformed into The enhanced features are fused through 3×3Conv+LeakyReLU+1×1Conv and residual connection, and then the feature information flow is constrained by GFFB, which consists of three 1×1Convs, two 3×3Convs and GELU activation functions, so that the high-resolution feature layer above can focus on restoring its own structural information and retaining useful features.
[0018] Preferably, the FEB of the feature learning module consists of three parts: an initial layer 3×3Conv+LeakyReLU, an intermediate layer 3×3Conv+GroupNormal+LeakyReLU and a multi-scale extraction module MSEM, and an output layer 3×3Conv+Sigmoid. MSEM connects 3×3 convolution kernels in a multi-branch parallel manner. The calculation process of MSEM can be expressed as:
[0019]
[0020] in, Represents MSEM, C 3×3 Represents a 3×3 convolution kernel.
[0021] Preferably: the FEB has the characteristic L t As input, output initial enhanced features And as the input of the feature convergence module in the next stage, the overall process of FEB can be expressed as:
[0022]
[0023] in, represents FEB, θ represents the shared weight, L t Represents the input features of a certain layer, Represents MSEM, C 3×3 represents a 3×3 convolution kernel, represents the LeakyReLU activation function, represents GroupNormal, σ represents Sigmoid function.
[0024] Preferably: the FCB includes an initial layer of 3×3Conv+GroupNormal+LeakyReLU, an intermediate layer feature fusion module FFM, an output layer of 3×3Conv+Sigmoid, and introduces a correction image S t And add it to the input features of FCB, the overall calculation process of FCB can be expressed as:
[0025]
[0026] Among them, C 3×3 represents a 3×3 convolution kernel, represents the LeakyReLU activation function, represents GroupNormal, σ represents Sigmoid function, and the entire FLM calculation process can be summarized as:
[0027]
[0028] in, represents FCM, θ represents the shared weight, On behalf of FLM,
[0029] Represents FEM.
[0030] Preferably: the calculation process of GFEB can be expressed as:
[0031]
[0032] Among them, DC stands for DConv, Avg stands for average pooling layer, Max stands for maximum pooling layer, c stands for Concat, σ stands for Sigmoid function, C 1×1 Represents a 1×1 convolution kernel.
[0033] Preferably: the calculation process of the FCLM-LM can be expressed as:
[0034]
[0035] Among them, FC stands for FCLM, c stands for Concat, C 3×3 represents a 3×3 convolution kernel, C 1×1 Represents a 1×1 convolution kernel.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] 1) The network proposed in the present invention can enhance both color and details, and it includes a multi-scale inter-layer guidance sub-network and a decomposition sub-network;
[0038] 2) The FGEM of the present invention uses lower resolution features to guide the enhancement of features in the previous layer. In addition, the FLM aims to enhance and converge guided features from different scales. The FLM includes FEB and FCB, and designs the FFM, which consists of multiple fusion units that are cascaded in a new residual structure, making full use of features at different scales while reducing the computational load of the network.
[0039] 3) A CAM based on a U-Net model of the present invention, the CAM includes FCLM and GFEB, which are respectively used to extract skip connection information and high-dimensional spatial features from the decoder and retain low-dimensional spatial features to capture rich color detail features in the reflectance component. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the overall framework diagram of the network of the present invention;
[0041] Figure 2 It is the network structure diagram of FGEM of the present invention;
[0042] Figure 3 It is the network structure diagram of FLM of the present invention;
[0043] Figure 4 is a network structure diagram of the FCLM of the present invention;
[0044] Figure 5 A subjective visual comparison of the network of the present invention with other cutting-edge methods on three images in the LOL dataset;
[0045] Figure 6 These are the objective evaluation results of the five different algorithms in the example on the LOL dataset. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the implementation regulations described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] Example
[0048] See also Figure 1 , the figure is a network framework diagram of the present invention, an image enhancement network based on multi-scale feature fusion, including:
[0049] The feature-guided enhancement module, which consists of a cross-feature attention (CFA) and a gated feature feed-forward block (GFFB), is used to reduce the computational complexity of the feature extraction process based on the number of channels;
[0050] The feature learning module consists of a feature enhancement block FEB and a feature convergence block FCB, with the feature L guided between each layer. t As input, it first passes through the feature enhancement block for adaptive enhancement to obtain the preliminary enhancement result. It is then sent to the feature convergence block to Perform fusion convergence to obtain a clear and convergent enhanced feature map The whole process shares parameters, and the feature learning module maps the feature output of each layer All converge to the same state and participate in the subsequent upsampling and fusion operations of each layer of features;
[0051] The decomposition sub-network consists of three Conv+LeakyReLU, one Conv+Sigmoid and CA, which is used to denoise the image before decomposition to obtain the illumination component and the reflection component;
[0052] The color adjustment module includes a global feature extraction block and a feature cross-learning block. The global extraction module GFB is embedded in the bottom layer of the U-Net model to retain the features, and the feature cross-learning block is used to fuse the information of each upsampling and jump connection to extract color features.
[0053] Furthermore, in the multi-scale inter-layer guided subnetwork (MSIG), the original image first undergoes a layer of 3×3Conv for preliminary feature extraction, and the obtained feature map is continuously downsampled 4 times to obtain feature maps of different channels and sizes. In order to better extract the underlying resolution features, a spatial channel attention (SCA) with both spatial and channel dimensions is designed on the underlying branch. Among them, the channel attention focuses on the input features in terms of meaning, while the spatial attention focuses on the information part. If the channel attention and the spatial attention are made independent of each other, it is not conducive to the sharing of features between the spatial and channel information. In addition, there is also spatial information in the channel information, and the two are not independent of each other. Therefore, SCA obtains representative attention by fusing the information between the channel dimension and the spatial dimension, so that the network can share the features in the channel and spatial dimensions. Therefore, a stacked structure of three consecutive SCAs is adopted and connected in a residual manner. The features of different stages are fused and the feature map is mapped to the [0,1] interval through a 3×3Conv+Sigmoid operation, which effectively retains the illumination changes of the feature map. The features of feature maps of different resolution sizes at the same position are highly correlated, and the underlying features contain rich feature information. Therefore, the underlying resolution features can be used to guide the feature map of the previous layer resolution, so as to accurately and efficiently enhance the feature map of the previous layer resolution. At the same time, since the underlying resolution features have obvious brightness and texture differences before and after enhancement, the underlying resolution features before and after enhancement are further used to guide the enhancement of the feature map of the previous layer, so that the network can adaptively notice the difference between the underlying resolution features before and after enhancement. Therefore, after the underlying resolution features are enhanced, the optimized feature map and the initial underlying resolution features are upsampled and sent to the feature guided enhancement module (FGEM) based on Transformer at the same time as the previous layer features. The underlying resolution features are used to guide the previous layer features to enhance the details. After the initial enhancement of FGEM, the result passes through the feature learning module (FLM), whose structure is mainly composed of 3×3Conv+GroupNormal+LeakyReLU. FLM can not only effectively enhance the input features, but also converge the features output at each stage to the same state. The whole process shares parameters to reduce the computational burden of the network. The converged features are weighted fused with the initial features of the same resolution size. The upsampling operation is also used and passed through the residual block. The residual block structure is as follows: Figure 1As shown in the figure, it consists of a 1×1Conv, two 3×3Conv+LeakyReLU and SE, which are connected in a residual way. SE can assign different weights to each channel in the multi-channel input to retain more features to prevent the loss of detail information during upsampling. At the same time, the output result of the residual block is spliced with the feature map of the previous layer in the channel dimension, and also enters the FLM module for feature enhancement convergence operation. The subsequent output feature map is similar to the previous process. Rich detail features are obtained through the guidance of each layer of features. The enhanced result of each layer is also fused with the features of the previous layer through upsampling operation to obtain the final feature with rich detail information.
[0054] In the decomposition subnetwork, the original image is decomposed into illumination component and reflection component by connecting three layers of Conv+LeakyReLU and CA. CA can assign different weights to pixels in the feature map according to coordinates, assigning smaller weights to pixel coordinates with more noise and larger weights to pixel coordinates with less noise, so as to achieve the effect of noise suppression in the feature map. The obtained reflection component enters the color adjustment module constructed in a shallow U-Net manner, and a global feature extraction block (GFEB) is introduced in the bottom layer of the module to prevent the loss of underlying detail information. At the same time, a feature cross-learning module (FCLM) is designed in the subsequent upsampling process. This module receives the input of each layer of jump connection features and the output of upsampling at the same time, so as to better fuse the features of each layer and prevent noise and artifacts. Finally, the optimized features of the two subnetworks are spliced in the channel and passed through Conv+LeakyReLU to obtain the final enhanced image.
[0055] From the above, it can be seen that the downsampling operation of the present invention adopts stride convolution instead of maximum pooling, and the upsampling operation is implemented by transposed convolution, which can further improve the performance of the network. MSIG uses progressive enhancement to ensure that feature maps of various resolutions are enhanced to a consistent degree. At the same time, it adopts a feature-guided method for each layer to maintain a low amount of calculation, thereby achieving a balance between image enhancement effect and running speed.
[0056] Furthermore, in terms of image processing, Transformer has efficient parallel computing capabilities and can capture the global information of long-distance context features. It can capture the global changes of the underlying resolution features and connect with each pixel in the upper-level resolution features, thereby guiding the upper-level resolution features to perform effective detail enhancement. However, when extracting features, the original Transformer needs to calculate the amount of computation based on the height and width of the features, which greatly increases the computational complexity of the network. Therefore, in order to better apply the underlying resolution features and reduce the amount of computation, the present invention designs FGEM based on the lightweight Transformer, extracts features in the channel dimension, and reduces its computational complexity from O(b / 2) with a height h and a width w. 2 w 2 ) is reduced to O(c according to the number of channels c 2 ), FGEM can link the bottom-level resolution features with the upper-level features to realize the network’s use of the bottom-level features. The structure of its module is as follows: Figure 2 As shown, the feature size of the output after each operation, the number of channels c, height h, and width w are all shown in the figure.
[0057] FGEM consists of two parts: Cross Feature Attention (CFA) and Gate Feature Feedforward Block (GFFB). To achieve the connection between the bottom-level resolution features and the upper-level features, CFA can simultaneously accept the input of three features, namely the initial and optimized feature maps of the bottom layer. With the initial feature map of the previous layer Then, the features from the bottom layer are upsampled and processed with the features from the previous layer by 3×3Conv+LeakyReLU+1×1Conv, and then the dimension is transformed to generate
[0058] The process can be expressed as:
[0059]
[0060] Among them, R represents the dimension transformation Reshape, up represents the upsampling operation, and C 1×1 represents a 1×1 convolution kernel, C 3×3 Represents a 3×3 convolution kernel.
[0061] Furthermore, matrix multiplication is used to calculate the cross-channel covariance matrix between Q and K to obtain a matrix of size P 256×256Compared with the original Transformer, the lightweight Transformer can effectively reduce the computational complexity when using the cross attention map P. The obtained attention map P expresses the changes of the underlying resolution features before and after enhancement at the same pixel position, and is used to guide V to perform feature enhancement. Further, the attention map P is matrix multiplied with V in the same way to obtain a size of The feature enhancement map. The process of obtaining it can be expressed as:
[0062]
[0063] Among them, Softmax represents the Softmax function, · represents matrix multiplication, and c represents a learning scaling parameter that can control the size. Then, the feature enhancement map is transformed into The enhanced features are fused through 3×3Conv+LeakyReLU+1×1Conv and residual connection, and then the feature information flow is constrained by GFFB, which consists of three 1×1Convs, two 3×3Convs and GELU activation functions, so that the high-resolution feature layer above can focus on restoring its own structural information and retaining useful features.
[0064] The feature learning module (FLM) proposed in the present invention is as follows Figure 3 As shown in Figure 1, it is placed in other layers except the bottom resolution, and the full convolution strategy is used to adaptively enhance the features of each layer and converge them to the same state. FLM consists of two parts: feature enhancement block (FEB) and feature convergence block (FCB). It uses the feature L guided between each layer. t As input, it first passes through the feature enhancement block for adaptive enhancement to obtain the preliminary enhancement result. It is then sent to the feature convergence block to Perform fusion convergence to obtain a clear and convergent enhanced feature map And the whole process shares parameters, FLM maps the feature output of each layer They all converge to the same state and participate in the subsequent upsampling and fusion operations of each layer of features to reduce the computational complexity of the network.
[0065] In this embodiment, FEB consists of three parts, an initial layer 3×3Conv+LeakyReLU, an intermediate layer 3×3Conv+GroupNormal+LeakyReLU and a multi-scale extraction module (MSEM), and an output layer 3×3Conv+Sigmoid. MSEM connects 3×3 convolution kernels in a multi-branch parallel manner to expand the network's receptive field and extract rich features. Specifically, although a large convolution kernel can obtain rich feature information, it will also increase the network calculation burden. Therefore, two 3×3 convolution kernels and three 3×3 convolution kernels are connected in series to replace 5×5 convolution kernels and 7×7 convolution kernels, which reduces the network parameters while realizing the effect of the large convolution kernel. The calculation process of MSEM can be expressed as:
[0066]
[0067] in, Represents MSEM, C 3×3 Represents a 3×3 convolution kernel.
[0068] Subsequently, two 3×3Conv+GroupNormal+LeakyReLU are densely connected to extract features. Compared with BatchNormal, GroupNormal is more suitable for normalizing multi-channel feature maps. It can group the channels of the feature map and normalize each group of channels to reduce the dependency between channels and accelerate network convergence. Finally, the 3×3Conv+Sigmoid function is used to ensure that the output is in the range of [0,1]. FEB is based on the feature L t As input, output initial enhanced features And it serves as the input of the feature convergence module in the next stage. The overall process of FEB can be expressed as:
[0069]
[0070] in, represents FEB, θ represents the shared weight, L t Represents the input features of a certain layer, Represents MSEM, C 3×3 represents a 3×3 convolution kernel, represents the LeakyReLU activation function, represents GroupNormal, σ represents Sigmoid function.
[0071] Furthermore, the modules used by FCB are similar to those of FEB, with an initial layer of 3×3Conv+GroupNormal+LeakyReLU, an intermediate layer feature fusion module FFM, and an output layer of 3×3Conv+Sigmoid, and the whole process shares weights. In order to enable the network to learn image features without increasing computational complexity, the present invention introduces the correction image S t And add it to the input features of FCB to show the output feature map With the initial input feature L t The difference between them can indirectly realize the convergence behavior of the output feature map of each layer and ensure that they are all in the same state.
[0072] Furthermore, first, the output of FEB With the initial input feature L t Divide element by element to get the corrected feature S t , and is corrected by a layer of 3×3Conv+GroupNormal+LeakyReLU to obtain S1. Subsequently, S1 is sent to FFM, where a cascade method is used to connect the fusion unit. The fusion unit consists of 3×3Conv+GroupNormal+LeakyReLU and is connected with a new residual structure. Its structure is as follows Figure 3 As shown in the figure, specifically, considering that the traditional full residual connection method can realize cross-information transmission and improve the reusability of network information, if it is connected with a stacked structure, it will increase the model calculation burden. Therefore, taking the first fusion unit as an example, the features before and after activation of 3×3Conv+GroupNormal are respectively connected with the input features through residual connections. Similarly, the output of the previous fusion unit is used as the subsequent input. Each fusion unit can use the features after activation of the previous residual block to guide the fusion of its own features before and after activation to achieve the effect of a multi-cascade residual structure. The calculation process of FFM can be expressed as:
[0073]
[0074] in, Represents FFM, C 3×3 represents a 3×3 convolution kernel, represents the LeakyReLU activation function, Represents GroupNormal.
[0075] Finally, the output result is controlled within the range of [0,1] by 3×3Conv+Sigmoid and compared with the correction feature S t Element-wise subtraction obtains a clear and convergent feature map It will participate in the subsequent upsampling fusion process and add the initial resolution features of the same layer to the resolution features of the previous layer as the input of the previous layer FLM. Therefore, the overall calculation process of FCB can be expressed as:
[0076]
[0077] Among them, C 3×3 represents a 3×3 convolution kernel, represents the LeakyReLU activation function, represents GroupNormal, σ represents Sigmoid function.
[0078] Finally, the entire FLM calculation process can be summarized as:
[0079]
[0080] in, represents FCM, θ represents the shared weight, On behalf of FLM, Represents FEM.
[0081] Enhancement methods based on Retinex theory, such as Retinex-Net, introduce a large amount of random noise when decomposing images, resulting in serious noise and color deviation in the enhanced images. Therefore, it is very important to denoise the images before decomposition. In order to obtain clear illumination components and reflection components, the decomposition subnetwork of the present invention is mainly composed of three Conv+LeakyReLU, one Conv+Sigmoid and CA. The convolution size is set to 3×3. The original image is first subjected to Conv+LeakyReLU for preliminary feature extraction, and then CA is introduced to suppress the noise in the image and give different weights. Finally, a clear decomposition result is obtained through the Conv+Sigmoid function.
[0082] Retinex theory believes that the inherent properties of objects on the reflection component contain the color information of the image. Most of the previous enhancement methods based on Retinex did not consider the loss of color detail features during transmission when processing the reflection component. Based on this, a color adjustment module was proposed based on the U-Net model. The global extraction module GFB was embedded in the bottom layer of the U-Net model to retain the features and prevent the loss of details of the bottom layer features during upsampling and downsampling. The feature cross-learning module was designed to fuse the information of each upsampling and jump connection to fully extract the color features. CAM uses the reflection component after decomposition of the original image as input, and finally obtains a clear and low-noise reflection component and uses it to guide the final image fusion.
[0083] The U-Net structure and convolution operation are limited in their ability to capture long-distance information. Therefore, the present invention proposes a global extraction block GFEB that can fully retain the underlying global information. Its structure is as follows: Figure 4 As shown in the figure, firstly, a dual-branch structure is used to extract global information based on maximum pooling and average pooling. Then, a DConv constructed by three convolution kernels of different sizes is connected in series to perform deep feature extraction. 1×1Conv with different numbers of channels is applied to obtain the features between each channel and multiply them element by element with the input features to obtain a feature map with rich local global information. After that, the output of each branch is connected to the input features through 1×1Conv, and the final output of GFEB is obtained through 1×1Conv+LeakyReLu. The calculation process of GFEB can be expressed as:
[0084]
[0085] Among them, DC stands for DConv, Avg stands for average pooling layer, Max stands for maximum pooling layer, c stands for Concat, σ stands for Sigmoid function, C 1×1 Represents a 1×1 convolution kernel.
[0086] The previous method of directly fusing the skip connection information with the upsampling features in the U-Net model is too simple and easily leads to the loss of effective information. The accompanying enhancement results are mostly unsatisfactory, which is undesirable. Therefore, in order to make full use of the skip connection information and the upsampling features, a feature cross learning module FCLM is proposed to fully fuse the feature information of the two. The structure of FCLM is as follows Figure 4 As shown in the figure, FCLM uses four branches to process the input features respectively and fuse the output features of each branch. First, by introducing CBAM and SE, local global features are extracted from the input features in the spatial channel dimension. Then, the input features of 1×1Conv+Sigmoid are multiplied element by element and the weights with different dimensional information are obtained through LeakyReLU. The input features are subjected to 1×1 and 3×3 convolutions in the other two branches respectively, and are added element by element with the outputs of CBAM and SE. The obtained features are also subjected to activation functions and connected with the input features on the channel to participate in subsequent feature fusion. Finally, the weights are added to the concatenated features to get the final output.
[0087] The FCLM-LM connection module designed by the present invention can fully fuse the output features of each level of FCLM and reduce the loss of features. FCLM-LM takes each layer of skip connection and upsampling features as input, first fuses them by element-by-element addition, and then processes them in sequence by three series-connected FCLMs. The features output by each FCLM are spliced and fused by skip connection, fully retaining the feature information, and finally pass through 3×3Conv and adopt residual structure to add with initial fusion features to achieve deep fusion of long-distance feature information. The calculation process of FCLM-LM can be expressed as:
[0088]
[0089] Among them, FC stands for FCLM, c stands for Concat, C 3×3 represents a 3×3 convolution kernel, C 1×1 Represents a 1×1 convolution kernel.
[0090] To verify the effectiveness of the network of the present invention, it is compared with 13 currently advanced classical image enhancement methods, including Retinex-Net (2018-BMVC), KinD (2019-ACMMM), EnlightenGAN, RRDNet (2020-ICME), Zero-DCE (2020-CVPR), Zero-DCE++, RUAS (2021-CVPR), SCI (2022-CVPR), URetinex-Net (2022-CVPR), UNIENet (2022-ECCV), PSENet (2023-WACV), PairLIE (2023-CVPR), QuadPrior (2024-CVPR), and experiments are carried out on the LOL public dataset. In terms of quantification, seven image evaluation indicators are selected for objective evaluation, including PSNR, SSIM, MS-SSIM, LPIPS, NIQE, BRISQUE, and PIQE.
[0091] The subjective visual comparison between the network of the present invention and other methods is shown in Figure 2. Figure 5As shown in the figure, it can be intuitively felt that the classic Retinex-Net method failed to effectively process the reflection component, resulting in unnatural artifacts and a lot of noise in the enhanced image. The KinD, RRDNet, ZeroDCE, ZeroDCE++ and SCI methods have insufficient exposure effects, blur and a lot of noise in the detail area, the EnlightenGAN and PSENet methods have color distortion and poor noise suppression, the images restored by the RUAS and UNIENet methods are overexposed, resulting in excessive blurring of the detail area, and although the PairLIE method brightens the image overall, there are still noise and blur problems that interfere with the visual quality. Compared with other methods, the URetinex-Net and QuadPrior networks of the present invention show better visual effects in terms of image exposure and color detail information retention.
[0092] Five evaluation indicators were used to objectively evaluate the images of different methods, as shown in Table 1 ( Figure 6 ), red bold represents the best result, black bold represents the second best result. Compared with the URetinex-Net method, the network of the present invention ranks second in SSIM and NIQE indicators. However, it is worth noting that the network of the present invention reaches 23.2966, 0.8723 and 0.2542 respectively in other indicators, which further shows that the enhanced image of the network of the present invention on the LOL dataset has better subjective visual effects in terms of color information and details.
[0093] The above content is a further detailed description of the present invention in combination with specific implementation methods. It cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, some simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the scope of protection determined by the claims submitted for the present invention.
Claims
1. Image enhancement network based on multi-scale feature fusion, characterized by: include: The feature-guided enhancement module, which consists of a cross-feature attention (CFA) and a gated feature feed-forward block (GFFB), is used to reduce the computational complexity of the feature extraction process based on the number of channels; The feature learning module consists of a feature enhancement block FEB and a feature convergence block FCB, with the feature L guided between each layer. t As input, it first passes through the feature enhancement block for adaptive enhancement to obtain the preliminary enhancement result. It is then sent to the feature convergence block to Perform fusion convergence to obtain a clear and convergent enhanced feature map The whole process shares parameters, and the feature learning module maps the feature output of each layer All converge to the same state and participate in the subsequent upsampling and fusion operations of each layer of features; The decomposition sub-network consists of three Conv+LeakyReLU, one Conv+Sigmoid and CA, which is used to denoise the image before decomposition to obtain the illumination component and the reflection component; The color adjustment module includes a global feature extraction block and a feature cross-learning module. The global feature extraction block GFEB is embedded in the bottom layer of the U-Net model to retain the features. The feature cross-learning module is used to fuse the information of each upsampling and jump connection to extract color features.
2. The image enhancement network based on multi-scale feature fusion according to claim 1, characterized in that: The CFA of the feature-guided enhancement module can simultaneously accept three feature inputs, namely, the initial and optimized feature maps of the bottom layer With the initial feature map of the previous layer Then, the features from the bottom layer are upsampled and processed with the features from the previous layer by 3×3Conv+LeakyReLU+1×1Conv, and then the dimension is transformed to generate The process can be expressed as: Among them, R represents the dimension transformation Reshape, up represents the upsampling operation, and C 1×1 represents a 1×1 convolution kernel, C 3×3 Represents a 3×3 convolution kernel. Then, matrix multiplication is used to calculate the cross-channel covariance matrix between Q and K to obtain a size of P 256 ×256 The obtained attention map P expresses the changes of the underlying resolution features before and after enhancement at the same pixel position, and is used to guide V to perform feature enhancement. The attention map P is matrix multiplied with V in the same way to obtain a size of The feature enhancement map of , the process of obtaining it can be expressed as: Among them, Softmax represents the Softmax function, represents matrix multiplication, and c represents the learning scaling parameter that can control the size. Then, the feature enhancement map is transformed into The enhanced features are fused through 3×3Conv+LeakyReLU+1×1Conv and residual connection, and then the feature information flow is constrained by GFFB, which consists of three 1×1Convs, two 3×3Convs and GELU activation functions, so that the high-resolution feature layer above can focus on restoring its own structural information and retaining useful features.
3. The image enhancement network based on multi-scale feature fusion according to claim 1, characterized in that: The FEB of the feature learning module consists of three parts: the initial layer 3×3Conv+LeakyReLU, the intermediate layer 3×3Conv+GroupNormal+LeakyReLU and the multi-scale extraction module MSEM, and the output layer 3×3Conv+Sigmoid. MSEM connects the 3×3 convolution kernel in a multi-branch parallel manner. The calculation process of MSEM can be expressed as: in, Represents MSEM, C 3×3 Represents a 3×3 convolution kernel.
4. The image enhancement network based on multi-scale feature fusion according to claim 3, characterized in that: The FEB is characterized by L t As input, output initial enhanced features And as the input of the next stage feature convergence module, the overall process of FEB can be expressed as: in, represents FEB, θ represents the shared weight, L t Represents the input features of a certain layer, Represents MSEM, C 3×3 represents a 3×3 convolution kernel, represents the LeakyReLU activation function, represents GroupNormal, σ represents Sigmoid function.
5. The image enhancement network based on multi-scale feature fusion according to claim 4, characterized in that: The FCB includes an initial layer of 3×3Conv+GroupNormal+LeakyReLU, an intermediate layer feature fusion module FFM, an output layer of 3×3Conv+Sigmoid, and introduces a correction image S t And add it to the input features of FCB, the overall calculation process of FCB can be expressed as: Among them, C 3×3 represents a 3×3 convolution kernel, represents the LeakyReLU activation function, represents GroupNormal, σ represents Sigmoid function, and the calculation process of the entire feature learning module FLM can be summarized as: in, represents FCM, θ represents the shared weight, On behalf of FLM, Represents FEM.
6. The image enhancement network based on multi-scale feature fusion according to claim 1, characterized in that: The calculation process of the global feature extraction block GFEB can be expressed as: Among them, DC stands for DConv, Avg stands for average pooling layer, Max stands for maximum pooling layer, c stands for Concat, σ stands for Sigmoid function, and C1×1 stands for 1×1 convolution kernel.
7. The image enhancement network based on multi-scale feature fusion according to claim 6, characterized in that: The connection module FCLM-LM based on the feature cross-learning module FCLM can fully fuse the output features of each level of FCLM and reduce the loss of features. FCLM-LM takes the jump connection and up-sampling features of each layer as input, first fuses them by element-by-element addition, and then processes them in sequence by three series-connected FCLMs. The features output by each FCLM are spliced and fused by jump connection to fully retain the feature information. Finally, after 3×3Conv and the residual structure is used to add the initial fusion features, the deep fusion of long-distance feature information is realized. The calculation process of FCLM-LM can be expressed as: Among them, FC stands for FCLM, c stands for Concat, C 3×3 represents a 3×3 convolution kernel, C 1×1 Represents a 1×1 convolution kernel.
Citation Information
Patent Citations
Image enhancement method based on improved multi-scale fusion generative adversarial network
CN115223004A
Image processing method based on multi-scale frequency feature fusion Transform model
CN117876293A