Noise estimation image enhancement method based on Transformer network

By constructing a feature extraction decomposition subnet and a noise estimation optimization subnet based on Transformer network, the problems of information loss and feature redundancy usage caused by pooling operations in image enhancement are solved, and efficient image feature extraction and denoising processing are achieved, and image quality is improved.

CN115965555BActive Publication Date: 2025-05-23XIDIAN UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202310042018.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-28
Publication Date
2025-05-23
Estimated Expiration
2043-01-28

AI Technical Summary

Technical Problem

The prior art has problems such as pooling operations in image enhancement that lead to the loss of original image information, which cannot be compensated by fusion of multi-scale features, and the use of U-net as the basic block of the network leads to the problem of feature redundancy use and high time complexity.

Method used

A feature extraction and decomposition subnet is built based on the Transformer network, combining the convolution operator to extract local information and the global modeling ability of Transformer to avoid feature redundancy use. Design the noise estimation optimization subnet to estimate the noise variance by fitting the convolutional layer to perform denoising processing to improve the computational efficiency.

Benefits of technology

Effectively extracting image features improves the model's expression ability and computing efficiency, reduces noise and color distortion, and improves the contrast and clarity of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965555B_ABST
    Figure CN115965555B_ABST
Patent Text Reader

Abstract

The present invention proposes a noise estimation image enhancement method based on a Transformer network. The present invention constructs a feature extraction decomposition subnetwork, uses a convolution operator to extract local information of the image, and combines the global modeling capability of the Transformer to give full play to the expression ability of the model, better extract image features, and solve the problem that the prior art only relies on parameter redundancy of convolutional layer stacking and easily causes gradient disappearance. The present invention can accurately retain image details and restore a clearer image. The noise estimation optimization subnetwork designed by the present invention performs a fitting estimate on the noise, and uses a residual structure for denoising optimization, which simply and effectively removes the noise in the enhanced image, so that the present invention can quickly and effectively perform denoising, and the enhanced image becomes smoother, improving the visual effect and image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and further relates to a noise estimation image enhancement method based on a Transformer network in the field of image enhancement technology. The present invention can be used for low-light illumination images in the fields of remote sensing, night image video, biomedicine, etc., to improve the contrast and clarity of the image, reduce the noise and color distortion generated by amplification through image enhancement, and thus enhance the image quality. Background Art

[0002] Low-light image enhancement is one of the basic tasks in image processing, and it has a wide range of applications in different fields, including visual surveillance, autonomous driving, and mobile photography. Images, as an important carrier, have penetrated into people's daily lives, especially smartphone photography has become ubiquitous and prominent. In such applications, low-light images containing a lot of noise are often obtained due to inevitable environmental or technical limitations, such as night image photography and limited exposure time. Therefore, enhancing low-light images and videos has important significance and practical application value. Unsupervised algorithms based on deep learning are mainly divided into two categories: supervised learning and unsupervised learning. Supervised learning image enhancement methods rely on large-scale training datasets to learn the mapping relationship between reference images and real images through deep learning modeling, so as to enhance the real low-light images to make them close to the reference image level. This type of method usually has a good enhancement effect on a certain type of data, but it is difficult to achieve good results for data with characteristics different from the training dataset, and the generalization ability is weak. In addition, the denoising method in image enhancement is mainly based on the Gaussian theory of uniform image noise distribution, and does not consider the situation where the noise level changes due to uneven illumination. The biggest advantage of unsupervised learning image enhancement methods is that they get rid of the difficulty of obtaining large-scale high-quality training data sets, saving a lot of manpower and material resources. Unsupervised image algorithms based on deep learning can improve the quality of low-light images while balancing computational efficiency. However, existing unsupervised algorithms tend to pursue algorithm flexibility and speed, and it is difficult for enhanced images to achieve good quality performance.

[0003] Tianjin University proposed a method for enhancing low-light images in its patent application "Low-light image enhancement method based on attention guidance and multi-scale feature fusion" (application date: March 18, 2022, application number: 202210267898.3, application publication number: CN 114596233A). The implementation steps of the method include: 1. Constructing a data set, which includes a large number of test sets synthesized in simulated low-light environments and low-light images in BayerRaw format taken in real environments. 2. Preprocessing the image data, including reducing the dimension, removing the black level, and amplifying the image coefficients. 3. Constructing a feature extraction network, using five stages to form a feature extraction network for learning feature information at different scales, and each stage consists of several residual blocks with dense connections introduced. 4. Constructing a feature fusion network including a multi-resolution fusion block and a chain residual pooling layer, and finally obtaining a feature map that fully contains the detail information of shallow features and the semantic information of deep features, realizing the fusion between feature maps at each level. 5. Design an interpretable attention loss function to guide network training. 6. Post-processing, output enhanced high-resolution images. Although this method proposes an interpretable attention guidance mechanism and does not bring additional neural network reasoning burden. However, the method still has shortcomings. The feature fusion network uses multi-resolution fusion blocks and chained residual pooling layers, which can effectively integrate shallow and deep feature information. The pooling operation is prone to lose the original image information, which cannot be compensated by fusing multi-scale features, resulting in a decline in the performance of image enhancement indicators.

[0004] Dalian Maritime University proposed an unsupervised low-light image enhancement method in its patent document "Unsupervised low-light image enhancement method and device based on illumination information guidance" (application date: June 8, 2022, application number CN202210646447.0, application publication number: CN115115540A). The implementation steps of the method include: collecting training data sets and test data sets; constructing a low-light enhancement model based on illumination information guidance, and inputting the images in the training data set into the low-light enhancement model for training; the low-light enhancement model based on illumination information guidance is a generative adversarial network including a generator and a discriminator; the generator includes an illumination estimation module and an enhancement network, and the illumination estimation module generates a single-channel illumination information map to guide the learning of the enhancement network; the enhancement network uses U-net as the basic block of the network, and adds illumination information map guidance at different scale layers of the U-net network; the image in the test data set is input into the trained low-light enhancement model to obtain an enhanced normal light image. This method uses a generative adversarial network for image enhancement, getting rid of the limitation of the supervised algorithm that requires paired image data, and adds illumination information to the features to guide the low-light enhancement process to avoid structural loss and color deviation. However, the method still has the disadvantage that it uses U-net as the basic block of the network and extracts features multiple times at different scales, resulting in redundant use of features and reducing the efficiency of the model.

[0005] Xi'an University of Technology proposed a very low illumination image enhancement method in its patent document "Extremely Low Illumination Image Enhancement Method Based on Retinex Model" (application date: April 29, 2022, application number: CN 202210475548.6, application publication number: CN 114723638 A). This method uses a bilateral filter to estimate the light brightness value of each pixel in the color channel map of the low-illuminance image, and determines the illumination component of the color channel map based on the light brightness value, and constructs a quaternion-based low-rank matrix constraint to denoise the reflection component; calculates the first reflection component corresponding to the color channel map based on the color channel map and the corresponding illumination component; removes the noise in the first reflection component based on the first denoising constraint model to obtain the second reflection component; and generates an enhanced low-illuminance image based on the illumination component and the second reflection component. This method can effectively retain the intrinsic connection between the three RGB color channels by constructing a denoising constraint model, increase the denoising effect, and thus solve the problem of severe color distortion after low-illuminance image enhancement. However, the method still has the disadvantage that the denoising constraint model uses a more complex quaternion matrix form, which has a high time complexity and thus causes the entire model to process images more slowly. Summary of the invention

[0006] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a noise estimation image enhancement method based on a Transformer network to solve the problems that the pooling operation easily loses the original image information and cannot be compensated by fusing multi-scale features, and that U-net is used as the basic block of the network to extract features multiple times at different scales, resulting in redundant use of features, and the denoising constraint model uses a complex quaternion matrix with high time complexity.

[0007] The idea of ​​realizing the purpose of the present invention is: the present invention constructs a fully resolved feature extraction decomposition subnetwork, which helps to retain the spatial details of the image and is used to solve the problem of lost image information caused by pooling or downsampling operations. The network uses convolution operators to extract local information of the image, and combined with the global modeling ability of Transformer, it can fully exert the expression ability of the model, avoid the problem that the current U-net module is prone to redundant use of features, and can more effectively extract image features. The noise estimation optimization subnetwork designed by the present invention estimates the noise level of the enhanced image by fitting the estimated noise variance of the convolution layer to achieve the purpose of denoising, which is simple and efficient, and improves the shortcomings of the prior art that the actual noise distribution of the image is not considered, the denoising is blind and a complex four-element matrix denoising model is used, and the time complexity is high. The feature extraction decomposition subnetwork constructed by the present invention based on Transformer uses the characteristics of Transformer and CNN to capture global and local information, give full play to the expression ability of the model, and improve the calculation efficiency of the model. The input image is transformed by the feature extraction decomposition subnetwork to obtain an illuminated image and a reflected image, wherein the reflected image represents the essential content of the image such as color and details, and a large amount of noise generated by the decomposition transformation operation is also included. Therefore, by constructing a noise estimation optimization subnetwork, the noise variance of the reflected image is approximately estimated, and finally the residual network is introduced to learn denoising to obtain the enhanced image.

[0008] The specific steps of the method of the present invention include the following:

[0009] Step 1: Generate training set:

[0010] Select at least 500 noisy low-light images, crop each image into 128*128 image blocks, perform data augmentation on each image block, and then convert all augmented image blocks into 3*128*128 tensor forms, and form all image blocks in tensor form into a training set;

[0011] Step 2: Construct a Transformer-based feature extraction decomposition subnetwork:

[0012] Build a feature extraction decomposition subnetwork consisting of the first convolution layer, the first Transformer layer, the second Transformer layer, the third Transformer layer, the second convolution layer, and the divider in series; the activation functions of the first and second convolution layers are implemented using the ReLU function and the Sigmoid function respectively, the sizes of the convolution kernels are set to 3*3 and 1*1 respectively, and the step sizes are all set to 1; the number of channels of the first to third Transformer layers is set to 16;

[0013] Step 3: Construct the noise estimation optimization subnetwork:

[0014] Step 3.1, build a noise estimation module consisting of the first convolutional layer and the second convolutional layer in series. The activation functions of the first and second convolutional layers both use the ReLU function, the convolution kernel sizes are set to 3*3 and 1*1 respectively, and the step sizes are set to 1;

[0015] Step 3.2: Build an optimization module consisting of the first convolutional layer, the second convolutional layer, and the channel attention layer in series; the activation functions of the first and second convolutional layers use the ReLU function, the convolution kernel size is set to 3*3, and the step size is set to 1; the number of channels in the channel attention layer is set to 16;

[0016] Step 3.3, build a fusion module consisting of a multiplier, a convolution layer, and an adder in series. The activation function of the convolution layer uses the ReLU function, the convolution kernel size is set to 1*1, and the step size is set to 1;

[0017] Step 3.4, the noise estimation module and the optimization module are connected in parallel and then connected in series with the fusion module to form a noise estimation optimization subnetwork;

[0018] Step 4, the feature extraction decomposition sub-network and the noise estimation optimization sub-network are connected in series to form an enhanced network;

[0019] Step 5: Train the enhanced network:

[0020] The training set is input into the enhanced network, and the back propagation algorithm is used to perform gradient descent, and the parameters of the enhanced network are iteratively updated until the loss function of the enhanced network converges, thereby obtaining a trained enhanced network.

[0021] Step 6: Perform image enhancement on low-light images:

[0022] In the same way as step 1, the low-light image to be enhanced is cropped and augmented, and then input into the trained enhancement network for image enhancement, and the image with normal light is output.

[0023] Compared with the prior art, the present invention has the following advantages:

[0024] First, since the Transformer-based feature extraction decomposition sub-network constructed by the present invention utilizes the convolution operator to extract local information of the image, and combines the global modeling capability of the Transformer, it fully exerts the expression ability of the model and better extracts image features, thereby improving the problem of redundant model parameters, low efficiency, and easy gradient disappearance caused by relying solely on stacked convolutional layers to expand the receptive field in the prior art. The present invention adopts a full-resolution feature extraction network, which can accurately retain image details and restore a clearer image.

[0025] Second, due to the noise estimation optimization subnetwork constructed by the present invention, the noise level of the enhanced image is estimated by fitting the estimated noise variance through the convolutional layer to achieve the denoising purpose, which is simple and efficient, and overcomes the shortcomings of the prior art that the actual noise distribution of the image is not considered, the denoising is blind and a complex denoising model is used. The present invention can perform denoising quickly and effectively, and the enhanced image becomes smoother, with better visual effects and image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flow chart of the present invention;

[0027] Figure 2 It is a structural schematic diagram of the network model of the present invention. DETAILED DESCRIPTION

[0028] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0029] Reference Figure 1 , further describing the implementation steps of the embodiment of the present invention.

[0030] Step 1: Generate a training set.

[0031] An embodiment of the present invention selects 500 noisy low-light images from a low-light dataset LOLdataset, crops each image into an image block of 128*128, performs data augmentation on each image block, and then converts all the augmented image blocks into a tensor form of 3*128*128, and forms all the image blocks in tensor form into a training set.

[0032] The operation of performing data augmentation on each image block is as follows.

[0033] One of the four image transformations, namely, clockwise rotation of 90 degrees, counterclockwise rotation of 90 degrees, rotation of 180 degrees, and image center symmetric transformation, is randomly selected for data augmentation. In the embodiment of the present invention, a total of 500 image blocks are generated through data augmentation, and the training set is composed of 500 image blocks without data augmentation, for a total of 1000 image blocks.

[0034] Reference Figure 2 , the enhanced network structure constructed by the present invention is further described.

[0035] Step 2: Construct a Transformer-based feature extraction decomposition sub-network.

[0036] A feature extraction decomposition subnetwork is built, which consists of the first convolutional layer, the first Transformer layer, the second Transformer layer, the third Transformer layer, the second convolutional layer, and a divider connected in series in sequence; the activation functions of the first and second convolutional layers are implemented using ReLU function and Sigmoid function respectively, the sizes of the convolution kernels are set to 3*3 and 1*1 respectively, and the step sizes are all set to 1; the number of channels of the first to third Transformer layers is set to 16.

[0037] The Transformer structure adopted in the embodiment of the present invention is composed of a first Layer norm layer, a multi-head self-attention layer, a first adder, a second Layer norm layer, an MLP feedforward layer, and a second adder connected in series, wherein the number of channels of the Layer norm layer is set to 16.

[0038] Step 3: Construct the noise estimation optimization subnetwork.

[0039] Step 3.1, build a noise estimation module consisting of the first convolutional layer and the second convolutional layer in series. The activation functions of the first and second convolutional layers both use the ReLU function, the convolution kernel sizes are set to 3*3 and 1*1 respectively, and the step sizes are set to 1.

[0040] In step 3.2, build an optimization module consisting of the first convolutional layer, the second convolutional layer, and the channel attention layer in series; the activation functions of the first and second convolutional layers both use the ReLU function, the convolution kernel size is set to 3*3, and the step size is set to 1; the number of channels of the channel attention layer is set to 16.

[0041] Step 3.3, build a fusion module consisting of a multiplier, a convolution layer, and an adder in series. The activation function of the convolution layer uses the ReLU function, the convolution kernel size is set to 1*1, and the step size is set to 1.

[0042] Step 4: Connect the feature extraction decomposition subnetwork and the noise estimation optimization subnetwork in series to form an enhanced network.

[0043] Step 5: Train the enhanced network.

[0044] The training set is input into the enhanced network, and the back-propagation algorithm is used for gradient descent. The parameters of the enhanced network are iteratively updated until the loss function of the enhanced network converges to obtain a trained enhanced network.

[0045] The specific training steps of the enhanced network are as follows:

[0046] In the first step, the training set is input into the enhancement network, the first convolutional layer in the feature extraction decomposition subnetwork is used to transform the number of channels, the first to third Transformer layers are used to extract features, and the second convolutional layer is used to reduce the dimension to obtain the illumination image; the input training set image is divided by the illumination image through a divider to obtain the reflection image, and the feature extraction decomposition subnetwork outputs the reflection image.

[0047] In the second step, the reflected image is used as the input of the noise estimation optimization subnetwork. The first convolution layer in the noise estimation module in the subnetwork extracts features from the input to obtain a feature map. After filtering the feature map in a small range, the feature map is subtracted from the filtered image, and the square mean of the difference is calculated to approximate the variance of the noise. Finally, the variance after noise estimation is corrected through learning by the second convolution layer, and the noise variance feature map is output.

[0048] In the third step, the reflection map is input into the first and second convolutional layers in the optimization module to complete further feature extraction, and then the features are corrected through the channel attention layer to retain valuable features and obtain approximate ideal image features without degradation.

[0049] In the fourth step, the feature map obtained in the third step and the noise variance feature map obtained in the second step are multiplied by the multiplier in the fusion module to fit the noise features, and the noise map is input into the convolution layer in the fusion module to obtain the noise map. The training set image and the noise map input in the first step are added using the adder to obtain the output image.

[0050] The loss function formula is as follows:

[0051]

[0052] Where L represents the loss function, ∑ represents the summation operation, M represents the total number of pixels in the low-light image block x, N(i) represents the window area centered on the i-th pixel, the window size is 5*5, j represents the j-th pixel in the N(i) window, and w i,j represents the weight of the j-th pixel for the i-th pixel, representing the 1-norm operation, x i represents the i-th pixel of the low-light image, x j represents the jth pixel in the window area centered on the i-th pixel in the low-light image, s represents the enhanced image, 2 Represents a 2-norm operation.

[0053] The w i,j It is obtained by the following formula:

[0054]

[0055] Among them, exp(·) represents the exponential operation with natural number e as the base, s i,c represents the i-th pixel in the color channel of the enhanced image in the YUV color space, s j,c represents the jth pixel in the window area centered on the i-th pixel of the color channel of the enhanced image in the YUV color space, σ represents the standard deviation of the Gaussian kernel, and its value is 0.1.

[0056] Step 6: Perform image enhancement on the low-light image.

[0057] In the same way as step 1, the low-light image to be enhanced is cropped and augmented, and then input into the trained enhancement network for image enhancement, and the image with normal light is output.

[0058] The effect of the present invention is further described below in conjunction with simulation experiments.

[0059] 1. Simulation experiment conditions:

[0060] The hardware platform of the simulation experiment of the present invention includes: the processor is Intel(R) Core(TM) CPUi9-10900X@3.70GHz, the memory is 32GB, and the graphics card is NVIDIA RTX2080Ti.

[0061] The software platform of the simulation experiment of the present invention is as follows: a code running including the Torch-1.11.0+cu100 environment library is built in the Python 3.9.0 virtual environment of Anaconda.

[0062] The data used in the simulation experiment of the present invention is the low-light dataset LOLdataset. 500 images are randomly selected from the dataset and cropped to obtain image blocks of 128*128 size, and data augmentation is performed on each cropped image block. All the augmented image blocks are converted into 3*128*128 tensor forms, and all the image blocks in tensor form are used as a training set; and 15 low-light images and one-to-one corresponding normal lighting reference images are randomly selected from the remaining dataset as image pairs to form a test set.

[0063] 2. Simulation content and results analysis:

[0064] The simulation experiment of the present invention uses the present invention and an existing technology (SCI image enhancement method) to perform image enhancement on the input test set low-light illumination images respectively.

[0065] In the simulation experiment, an existing technology used is:

[0066] The prior art SCI image enhancement method refers to: the image enhancement algorithm proposed by Long Ma et al. in “Toward Fast, Flexible, and Robust Low-Light Image Enhancement” (Published as a conference paper at CVPR2022), referred to as the SCI image enhancement method.

[0067] In order to verify the simulation experiment effect of the present invention, the image quality evaluation index average peak signal to noise ratio (PSNR) is used to evaluate the enhanced result images of the two methods of the simulation experiment. The average peak signal to noise ratio of 15 test images is calculated using the following formula, and the calculation results are plotted in Table 1:

[0068]

[0069] in, It means that the image is summed according to width m and length n respectively, and I(i,j) and K(i,j) represent the pixel values ​​of the enhanced image and the reference image under normal illumination respectively.

[0070] Table 1 Comparison of the image quality index PSNR between the present invention and the prior art

[0071]

[0072] It can be seen from Table 1 that the average peak signal-to-noise ratio PSNR of the present invention is 17.04 dB, which is higher than the prior art method SCI, proving that the present invention can obtain a better quality enhanced image.

[0073] The above simulation experiments show that: by constructing a feature extraction decomposition subnetwork based on Transformer, the present invention can efficiently extract image features, improve the parameter redundancy and gradient vanishing problem that only relies on the stacking of convolutional layers, and improve the expression ability and parameter utilization efficiency of the network model. In addition, the designed noise estimation optimization subnetwork performs fitting estimation on the noise and uses the residual structure for denoising optimization, which simply and effectively removes the noise in the enhanced image and improves the overall performance of the network model.

Claims

1. A noise estimation image enhancement method based on Transformer network, It is characterized in that Construct a Transformer-based feature extraction decomposition network and a noise estimation optimization network; the steps of the image enhancement method include the following: Step 1: Generate training set: Select at least 500 noisy low-light images, crop each image into 128*128 image blocks, perform data augmentation on each image block, and then convert all augmented image blocks into 3*128*128 tensor forms, and form all image blocks in tensor form into a training set; Step 2: Construct a Transformer-based feature extraction decomposition subnetwork: Build a feature extraction decomposition subnetwork consisting of the first convolution layer, the first Transformer layer, the second Transformer layer, the third Transformer layer, the second convolution layer, and the divider in series; the activation functions of the first and second convolution layers are implemented using the ReLU function and the Sigmoid function respectively, the sizes of the convolution kernels are set to 3*3 and 1*1 respectively, and the step sizes are all set to 1; the number of channels of the first to third Transformer layers is set to 16; Step 3: Construct the noise estimation optimization subnetwork: Step 3.1, build a noise estimation module consisting of the first convolutional layer and the second convolutional layer in series. The activation functions of the first and second convolutional layers both use the ReLU function, the convolution kernel sizes are set to 3*3 and 1*1 respectively, and the step sizes are set to 1; Step 3.2: Build an optimization module consisting of the first convolutional layer, the second convolutional layer, and the channel attention layer in series; the activation functions of the first and second convolutional layers use the ReLU function, the convolution kernel size is set to 3*3, and the step size is set to 1; the number of channels in the channel attention layer is set to 16; Step 3.3, build a fusion module consisting of a multiplier, a convolution layer, and an adder in series. The activation function of the convolution layer uses the ReLU function, the convolution kernel size is set to 1*1, and the step size is set to 1; Step 3.4, the noise estimation module and the optimization module are connected in parallel and then connected in series with the fusion module to form a noise estimation optimization subnetwork; Step 4, the feature extraction decomposition sub-network and the noise estimation optimization sub-network are connected in series to form an enhanced network; Step 5: Train the enhanced network: The training set is input into the enhanced network, and the back propagation algorithm is used to perform gradient descent, and the parameters of the enhanced network are iteratively updated until the loss function of the enhanced network converges, thereby obtaining a trained enhanced network. Step 6: Perform image enhancement on low-light images: In the same way as step 1, the low-light image to be enhanced is cropped and augmented, and then input into the trained enhancement network for image enhancement, and the image with normal light is output.

2. The noise estimation image enhancement method based on Transformer network according to claim 1, It is characterized in that The data augmentation for each image block in step 1 refers to randomly selecting one of four image transformations: clockwise rotation of 90 degrees, counterclockwise rotation of 90 degrees, rotation of 180 degrees, and central symmetric transformation of the image to perform data augmentation.

Citation Information

Patent Citations

  • Low-illumination image enhancement method based on attention guidance and multi-scale feature fusion

    CN114596233A

  • Extremely low illumination image enhancement method based on Retinex model

    CN114723638A

  • Retinex Model-Based Image Enhancement Method for Extremely Low Light

    CN114723638B

  • Unsupervised low-light image enhancement method and device based on illumination information guidance

    CN115115540A

  • Unsupervised low-light image enhancement method and device based on illumination information guidance

    CN115115540B