An image enhancement method based on curve learning and Transformer

Through the image enhancement method based on curve learning and Transformer, the features of low-illumination images are extracted and pixel mapping relationships are learned, which solves the problem of low-illumination images and the contrast reduction, and a significant improvement in image quality is achieved.

CN114881873BActive Publication Date: 2025-05-06NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210411643.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-05-06
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of brightness and contrast reduction in low-illumination images, especially in uneven lighting or inclement weather conditions, image quality is difficult to guarantee.

Method used

Using the image enhancement method based on curve learning and Transformer, by constructing a Vision Transformer network model, local features and global relationships of low-illumination images are extracted, pixel mapping relationships are learned, and the network is trained using L1-Loss and second-order gradient regularity to output pixel transformation curves for image enhancement.

Benefits of technology

Effective enhancement of low-illumination images is achieved, the brightness and contrast of the images are improved, noise is removed, and image quality is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881873B_ABST
    Figure CN114881873B_ABST
Patent Text Reader

Abstract

The present invention discloses an image enhancement method based on curve learning and Transformer. The present invention extracts the features of low-light images in a learning-based manner, and automatically generates a pixel transformation curve to achieve image brightness enhancement and local contrast stretching. The curve transformation is abstracted into the corresponding transformation of pixels. Taking an 8-bit digital image as an example, its pixel range is an integer of 0 to 255, and the corresponding pixel can be expressed as a 256-dimensional vector. In order to improve the naturalness of the image, a one-dimensional Laplace regularization term for the vector is added to the loss function to ensure the smoothness and monotonicity of the transformation curve, so that the pixel size relationship after the transformation remains unchanged. In addition, in order to reduce the noise in the enhanced image, the BM3D denoising algorithm is introduced. According to the slope of the transformation curve that can reflect noise recognition to a certain extent, different regions are subjected to different degrees of denoising to better remove the noise in the enhanced image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to an image enhancement method based on curve learning and Transformer. Background Art

[0002] In today's highly digital age, digital images have become an important information carrier, playing an important role in both daily life and production environments. However, the quality of images cannot be guaranteed during the acquisition process. For example, insufficient lighting, uneven lighting, or bad weather will seriously reduce the brightness and contrast of the image. During the transmission and storage of images, the image will also experience a decrease in accuracy and information loss, which seriously affects the restoration of information and the performance of image processing algorithms. Image enhancement is to restore and amplify the image information according to certain requirements, improve the contrast of the image, and remove noise from the image.

[0003] Image enhancement algorithms mainly include four categories, which are divided into methods based on histogram equalization, methods based on Retinex theory, methods based on physical models, and methods based on deep learning. Among them, the algorithm based on histogram equalization is a relatively classic method. It uses the statistical characteristics of the histogram to change the distribution of pixels to achieve image contrast stretching and dynamic range expansion. The method based on Retinex theory is based on the theoretical model of Retinex, and estimates different factors that affect the color of the object to restore a higher quality image. The method based on the physical model is based on specific physical observations. For example, the inverted image of a low-light image has similar characteristics to a foggy image, and then the defogging algorithm is used to achieve image enhancement. Since 2017, with the rapid development of deep learning, more and more image enhancement methods based on deep learning have begun to appear, and have become a hot topic of current research.

[0004] Appropriate curve adjustment can effectively enhance the brightness and local contrast of the image. Curve color adjustment is similar to gamma transformation, but more flexible, and different curves can be constructed according to different needs. Curve transformation is a corresponding mapping relationship between corresponding pixel values, without changing the order of pixel size, which meets the requirements of low-light image enhancement. Curve transformation can well control the relative size of pixels in the enhanced image, but it is very difficult to manually construct a curve transformation suitable for different images. Summary of the invention

[0005] In order to solve the technical problems mentioned in the above background technology, the present invention proposes an image enhancement method based on curve learning and Transformer.

[0006] In order to achieve the above technical objectives, the technical solution of the present invention is:

[0007] An image enhancement method based on curve learning and Transformer, comprising the following steps:

[0008] Step 1, concretize the low illumination image enhancement into corresponding pixel mapping from the low illumination image to the normal illumination image;

[0009] Step 2: Build a Vision Transformer network model to extract local features and global relationships in low-light images;

[0010] Step 3, obtaining all coordinates of the pixel values ​​in the low-light image that are equal to the specified value, then taking out all pixels at these coordinates in the corresponding normal-light image, and taking the value with the highest frequency among these pixels as the reference pixel value corresponding to the specified pixel in the low-light image after enhancement;

[0011] Step 4, use L1-Loss as the main loss function, and add the second-order gradient regularization of the output vector, that is, the pixel correspondence; train the deep network model on the LOLdataset dataset until convergence; Step 5, input the low-light image to be enhanced into the trained network model, output the pixel transformation curve, and use the curve to transform the input low-light image to obtain the enhanced image;

[0012] Step 6, combining the BM3D denoising algorithm, taking the gradient and pixel value of the pixel transformation curve output by the model in step 5 as a reference, perform denoising processing to different degrees on each area of ​​the enhanced image in step 5.

[0013] Preferably, in step 1, the mapping relationship between the low-light image and the normal-light image is simplified to a pixel mapping relationship and represented by a one-dimensional vector containing 256 elements, as shown in the following formula:

[0014] Output(x,y)=V[Input(x,y)],

[0015] V=[m0,m1,m2,…,m 255 ] T

[0016] Among them, Output(x, y) represents the pixel value of the enhanced image at the (x, y) coordinate, Input(x, y) represents the pixel value of the input image at (x, y); V is a one-dimensional vector, whose elements

[0017] m0,m1,m2,…,m 255 In turn, they represent the pixel values ​​of pixels 0, 1, 2, ..., 255 in the low-light image corresponding to the pixel values ​​in the normal-light image.

[0018] Preferably, the specific steps in step 2 are as follows:

[0019] (21) Divide the low-light image into different image blocks, flatten each image block into a one-dimensional vector, add a position code to each vector, and concatenate a trainable pixel mapping vector;

[0020] (22) Input into the Vision Transformer main framework, and use the multi-head attention mechanism to learn the features of the low-light image and the relationship between different image blocks; finally, extract the pixel mapping vector as the output result of the network.

[0021] Preferably, the specific steps in step 3 are as follows:

[0022] (31) Establish a 256*256 matrix A, where the row coordinates represent the pixel values ​​in the low-light image, the column coordinates represent the corresponding pixels in the normal-light image, and the element value represents the number of pixel values ​​corresponding to the row coordinates, as shown in the following formula:

[0023]

[0024] Where P(x=i) represents the coordinates of the element with pixel value i in the low-light image, V(·) represents the pixel value at coordinate “·” in the normal-light image, and C(v=j) represents the number of pixels with value equal to j;

[0025] (32) Establish a one-dimensional vector M of 256 elements, take the column coordinates with the largest value in each row of matrix A as the element value of M, and use M as the reference value for model training.

[0026] Preferably, the specific steps in step 4 are as follows:

[0027] (41) Use represents the output of the deep network model, that is, the mapping relationship of pixels, M represents the reference value of model training, calculates the L1-Loss of the two vectors as the main loss function term, and adds The Laplace regularization of is as follows:

[0028]

[0029]

[0030] in and M is a one-dimensional vector, Represents the pixel transformation curve output by the proposed network model;

[0031] M i Respectively and the i-th element of M, for The previous and next elements of ;

[0032] (42) The deep network is trained on the public dataset LOLDataset and the loss function is back-propagated to make the network gradually converge.

[0033] Preferably, in step 5, the specific steps are as follows:

[0034] (51) Input the low-light image to be enhanced into the trained network model, and the model outputs the pixel transformation curve

[0035] (52) Using the curve output by the model The low-light image is transformed into a curve as shown in the following formula:

[0036]

[0037] Where L(x, y) represents the pixel value of the input low-light image L at the coordinate (x, y), Represents the enhanced image The pixel value at (x, y).

[0038] Preferably, in step 6, the specific steps are as follows:

[0039] (61) Enhanced Image Use the BM3D denoising algorithm to globally denoise the image and obtain the denoised image

[0040] (62) A weight coefficient W is determined according to the slope of the transformation curve, where W is positively correlated with the slope of the curve and negatively correlated with the size of the pixel value;

[0041] (63) The weight coefficient W is used to enhance the image and denoised image The fusion process is expressed as:

[0042]

[0043] Where I is the final output result.

[0044] The beneficial effects brought by adopting the above technical solution are:

[0045] The present invention discloses an image enhancement method based on curve learning and Transformer, which abstracts the enhancement problem of low-light images into a pixel mapping problem of curve transformation. The Transformer network architecture is introduced to establish the internal connection of the image from the global and local perspectives, and learn the mapping relationship between the pixels in the low-light image and the pixels in the enhanced image. This method can be used in the fields of mobile phone photography, monitoring, security, etc. It is fast and effective, the method is ingenious, the conception is novel, and it has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is the overall process framework of the method adopted by the present invention;

[0047] Figure 2 It is a schematic diagram of the curve transformation adopted by the present invention;

[0048] Figure 3 This is a demonstration of the low-light image enhancement effect of the present invention. DETAILED DESCRIPTION

[0049] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0050] The present invention discloses an image enhancement method based on curve learning and Transformer. The overall framework is as follows: Figure 1 As shown, the specific implementation steps are as follows:

[0051] Step (A) uses curve transformation to enhance the low-light image, and further materializes the curve transformation into a pixel mapping relationship. The schematic diagram of the curve transformation is shown in Figure 2 As shown;

[0052] Step (B), using the Vision Transformer network framework to extract local features and global relationships in the original image, and automatically generate pixel mapping relationships of the enhanced image;

[0053] Step (C) trains the deep network in a supervised learning manner, counts the pixels at corresponding coordinates of the low-light image and the normal-light image, finds all coordinates in the low-light image whose pixel values ​​are equal to the specified values, then takes out all pixels at the corresponding coordinates in the normal-light image, and uses the value with the highest frequency among these pixels as the reference pixel value corresponding to the specified pixel in the low-light image after enhancement.

[0054] Step (D), using L1-Loss as the main loss function to train the network, and adding the second-order gradient regularization of the output vector, i.e., the pixel correspondence, to ensure the smoothness of the transformation curve;

[0055] Step (E), inputting the low-light image to be enhanced into the trained network model, outputting a pixel transformation curve, and using the curve to transform the input low-light image to obtain an enhanced image;

[0056] Step (F), combining the BM3D denoising algorithm, taking the gradient and pixel value of the transformation curve as reference, performs denoising processing on the enhanced image to different degrees.

[0057] In the aforementioned low-light image enhancement method based on curve learning and Transformer, step (A) regards the enhancement process of the low-light image as a curve transformation process, and further abstracts it into pixel mapping, which includes the following steps:

[0058] (A1) The mapping relationship between low-light images and normal-light images is simplified to a pixel mapping relationship, thereby simplifying the learning process. Taking an 8-bit image as an example, the mapping relationship can be represented by a one-dimensional vector of 256 elements, where the value of the element represents the corresponding pixel value after enhancement, as shown in the following formula:

[0059] Output(x,y)=V[Input(x,y)],

[0060] V=[m0,m1,m1,…,m 255 ] T

[0061] Among them, Output(x, y) represents the pixel value of the normal illumination image at the (x, y) coordinate, Input(x, y) represents the pixel value of the input image at (x, y), and V is the mapping relationship of the pixels.

[0062] The aforementioned low-light image enhancement method based on curve learning and Transformer, step (B), uses the Vision Transformer framework to learn and extract local and global information and relationships of low-light images, and automatically generates pixel mapping relationships of normal-light images. The steps are as follows:

[0063] (B1) The low-light image is divided into different blocks and flattened into one-dimensional vectors. Each vector is added with a position code and concatenated with a trainable pixel mapping vector.

[0064] (B2) Input into the Vision Transformer main framework, use the multi-head attention mechanism to learn the relationship between the features of the input image and different image blocks. Finally, extract the pixel mapping vector as the output result of the network.

[0065] In the aforementioned low-light image enhancement method based on curve learning and Transformer, step (C) is to train the deep network in a supervised learning manner, count the pixels of the corresponding coordinates of the low-light image and the normal-light image, and take the pixel in the original image that corresponds to the largest proportion of the pixels in the normal-light image as the pixel corresponding to the pixel. The steps are as follows:

[0066] (C1) Create a 256*256 matrix A, where the row coordinates represent the pixel values ​​in the low-light image, the column coordinates represent the corresponding pixels in the normal-light image, and the element value represents the number of pixel values ​​corresponding to the original pixel values ​​of the row coordinates and the column coordinates. As shown in the following formula:

[0067]

[0068] Where P(x=i) represents the coordinates of the element with pixel value i in the low-light image, V(·) represents the pixel value at coordinate “·” in the normal-light image, and C(v=j) represents the number of pixels with value equal to j.

[0069] (C2) Establish a one-dimensional vector M with 256 elements, and take the column coordinate with the largest value in each row of matrix A as the element value of M. M can be used as a reference value for training.

[0070] In the aforementioned low-light image enhancement method based on curve learning and Transformer, step (D) uses L1-Loss as the main loss function to train the network, and adds a second-order gradient regularization to the output vector, that is, the pixel correspondence, to ensure the smoothness of the transformation curve. The steps are as follows:

[0071] (D1) represents the output of the network, that is, the mapping relationship of pixels, M represents the reference value, calculates the L1-Loss of the two vectors as the main loss function term, and imposes The Laplace regularization of is as follows:

[0072]

[0073]

[0074] (D2) Train the network on the public dataset LOLdataset and back-propagate the loss function so that the network gradually converges.

[0075] In the aforementioned low-light image enhancement method based on curve learning and Transformer, step (E) is to input the low-light image to be enhanced into the trained network model, and the model will output a pixel transformation curve. The input low-light image is transformed using the curve to obtain an enhanced image. The steps are as follows:

[0076] (E1) Input the low-light image to be enhanced into the trained network model in claim 5, and the model outputs a pixel transformation curve

[0077] (E2) Using the curve output by the model The low-light image is transformed into a curve as shown in the following formula:

[0078]

[0079] Where L(x, y) represents the pixel value of the input low-light image L at the coordinate (x, y). Represents the enhanced image The pixel value at (x, y). Representation curve The value of the “·”th element of .

[0080] In the aforementioned low-light image enhancement method based on curve learning and Transformer, step (F) combines the BM3D denoising algorithm and uses the gradient and pixel value of the transformation curve as a reference to perform denoising on the enhanced image to varying degrees. The steps are as follows:

[0081] (F1) for enhanced image Use the BM3D denoising algorithm to perform global denoising on the image and obtain the image

[0082] (F2) determining a weight coefficient W according to the slope of the transformation curve, where W is positively correlated with the slope of the curve and negatively correlated with the size of the pixel value;

[0083] (F3) Use the weight coefficient W to enhance the image and denoised image Fusion, while retaining the bright details, removes the noise of the method. The process can be expressed as:

[0084]

[0085] I is the final output result. The enhanced result is as follows: Figure 3 shown.

[0086] The embodiments are only for illustrating the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. An image enhancement method based on curve learning and Transformer, characterized in that: The following steps are involved: Step 1, concretize the low illumination image enhancement into corresponding pixel mapping from the low illumination image to the normal illumination image; Step 2: Build a Vision Transformer network model to extract local features and global relationships in low-light images; Step 3, obtaining all coordinates of the pixel values ​​in the low-light image that are equal to the specified value, then taking out all pixels at these coordinates in the corresponding normal-light image, and taking the value with the highest frequency among these pixels as the reference pixel value corresponding to the specified pixel in the low-light image after enhancement; Step 4, use L1-Loss as the main loss function and add the output vector, i.e., the second-order gradient regularization of the pixel correspondence; train the deep network model on the LOL dataset until convergence; Step 5: input the low-light image to be enhanced into the trained network model, output a pixel transformation curve, and use the curve to transform the input low-light image to obtain an enhanced image; Step 6, combining the BM3D denoising algorithm, taking the gradient and pixel value of the pixel transformation curve output by the model in step 5 as a reference, performing denoising processing to different degrees on each area of ​​the enhanced image in step 5; The specific steps in step 2 are as follows: (21) Divide the low-light image into different image blocks, flatten each image block into a one-dimensional vector, add a position code to each vector, and concatenate a trainable pixel mapping vector; (22) Input into the Vision Transformer main framework, and use the multi-head attention mechanism to learn the relationship between the features of the low-light image and different image blocks; finally, extract the pixel mapping vector as the output result of the network; The specific steps in step 4 are as follows: (41) Use represents the output of the deep network model, that is, the mapping relationship of pixels, M represents the reference value of model training, calculates the L1-Loss of the two vectors as the main loss function term, and adds The Laplace regularization of is as follows: in, and M is a one-dimensional vector, Represents the pixel transformation curve output by the proposed network model; Respectively and the i-th element of M, for The previous and next elements of ; (42) The deep network is trained on the public dataset LOLDataset and the loss function is back-propagated to make the network gradually converge.

2. The image enhancement method based on curve learning and Transformer according to claim 1, characterized in that: In step 1, the mapping relationship between the low-light image and the normal-light image is simplified to a pixel mapping relationship and represented by a one-dimensional vector containing 256 elements, as shown in the following formula: Output(x,y)=V[Input(x,y)], V=[m0,m1,m2,…,m 255 ] T Among them, Output(x,y) represents the pixel value of the enhanced image at the (x,y) coordinate, Input(x,y) represents the pixel value of the input image at (x,y); V is a one-dimensional vector, whose elements m0,m1,m2,…,m 255 They represent the pixel values ​​of 0, 1, 2, ..., 255 in the low-light image corresponding to the pixel values ​​in the normal-light image.

3. The image enhancement method based on curve learning and Transformer according to claim 1, characterized in that: The specific steps in step 3 are as follows: (31) Establish a 256*256 matrix A, where the row coordinates represent the pixel values ​​in the low-light image, the column coordinates represent the corresponding pixels in the normal-light image, and the element value represents the number of pixel values ​​corresponding to the row coordinates, as shown in the following formula: Where P(x=i) represents the coordinates of the element with pixel value i in the low-light image, V(·) represents the pixel value at coordinate "·" in the normal-light image, and C(v=j) represents the number of pixels with value equal to j; (32) Establish a one-dimensional vector M of 256 elements, take the column coordinates with the largest value in each row of matrix A as the element value of M, and use M as the reference value for model training.

4. The image enhancement method based on curve learning and Transformer according to claim 1, characterized in that: In step 5, the specific steps are as follows: (51) Input the low-light image to be enhanced into the trained network model, and the model outputs the pixel transformation curve (52) Using the curve output by the model The low-light image is transformed into a curve as shown in the following formula: Where L(x,y) represents the pixel value of the input low-light image L at the coordinate (x,y), Represents the enhanced image The pixel value at position (x,y).

5. The image enhancement method based on curve learning and Transformer according to claim 1, characterized in that: In step 6, the specific steps are as follows: (61) Enhanced Image Use the BM3D denoising algorithm to globally denoise the image and obtain the denoised image (62) A weight coefficient W is determined according to the slope of the transformation curve, where W is positively correlated with the slope of the curve and negatively correlated with the size of the pixel value; (63) The weight coefficient W is used to enhance the image and denoised image The fusion process is expressed as: Where I is the final output result.

Citation Information

Patent Citations

  • Model structure, model training method, image enhancement method and equipment

    CN112529150A

  • Transform-based twin network image denoising method and system, medium and equipment

    CN114359109A