Low-light Image Enhancement Method Based on Reinforcement Learning and Aesthetic Evaluation

The low-light image enhancement method uses reinforcement learning with an expanded action space and aesthetic quality scoring to address the limitations of existing methods, achieving improved flexibility and user-centric image enhancements in challenging lighting conditions.

JP2025518390AActive Publication Date: 2025-06-12NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024572237
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-10
Filing Date
2023-02-07
Publication Date
2025-06-12
Estimated Expiration
2043-02-07

AI Technical Summary

Technical Problem

Existing low-light image enhancement methods focus primarily on improving brightness and contrast, but struggle with scenarios involving uneven illumination and backlight, and often rely on objective evaluation metrics that neglect user subjective experience.

Method used

A low-light image enhancement method based on reinforcement learning that expands the action space to include both brightness enhancement and reduction, and incorporates aesthetic quality scoring as part of the loss function to simulate user subjective evaluation.

Benefits of technology

The method achieves higher flexibility and effectiveness in real-world low-light scenarios by allowing for a wider range of enhancement operations and better aligning image enhancements with user aesthetic preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025518390000001_ABST
    Figure 2025518390000001_ABST
Patent Text Reader

Abstract

The present invention discloses a low-light image enhancement method based on reinforcement learning and aesthetic evaluation. First, abnormal luminance images in different lighting scenarios are generated, and a training dataset for the reinforcement learning system is constructed based on the images. Next, the training dataset, policy network, and value network in the reinforcement learning system are initialized, and the policy network and value network are updated based on a reward value without reference and an aesthetic evaluation reward value. After the training is completed, the enhanced image result is output. By expanding the range of the action space defined in reinforcement learning, the enhancement operation obtained from the input low-light image in the present invention has a larger dynamic range, has higher flexibility for real scenarios, and can better meet the low-light image enhancement requirements in real scenarios. Additionally, by introducing the score of aesthetic quality evaluation as part of the loss function, the enhanced image can have a better visual effect and the user's subjective evaluation score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image enhancement technology, and mainly relates to a low-light image enhancement method based on reinforcement learning and aesthetic evaluation.

Background Art

[0002] Pictures taken under poor lighting conditions have a low dynamic range of the image due to insufficient incident light on the digital camera sensor, and are severely interfered by noise, making it difficult to obtain high-quality images. However, low-light image enhancement plays a very important role in the field of computer vision. Pictures taken under low illumination often have many adverse effects. For example, the captured image is blurred and the image body is uncertain, the face is blurred and the recognition is inaccurate, and the details are blurred and the meaning of the image representation is incorrect. In this way, it not only affects the experience of humans using imaging devices, but also reduces the quality of photos, and sometimes leads to the transmission of incorrect information. Low-light image enhancement is advantageous for subsequent higher-level operations, such as target detection, face recognition, image classification, etc., by making the brightness of the captured image brighter, the contrast higher, and the structural information clearer, and has strong practical significance.

[0003] In recent years, methods based on deep learning generally learn how to improve and enhance low-light images using high-quality normal light images as guidance. LL-Net (Kin Gwn Lore, Adedotun Akintayo, and Soumik Sarkar. 2017. LLNet: A deep autoencoder approach to natural low-light image enhancement. Pattern Recognition 61(2017), 650-662.) proposed a stacked autoencoder that simultaneously performs noise removal and enhancement using synthetic low-light / normal light image pairs. However, due to the difference from real images, it is inevitable that the distribution of synthetic data deviates from real-world images, leading to a significant performance degradation when transferring to real situations. Subsequently, Wei et al. (Chen Wei, Wenjing Wang, Wenhan Yang, and Jiaying Liu. 2018. Deep retinex decomposition for low-light enhancement. arXiv preprint arXiv:1808.04560(2018).) collected a real dataset with low-light / normal light image pairs and proposed a retina network that decomposes images into illumination and reflectance in a data-driven manner based on this.After that, many other teacher-aided low-light image enhancement neural networks have been proposed (Wenhan Yang, Shiqi Wang, Yuming Fang, Yue Wang, and Jiaying Liu. 2020. From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement. In IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR)., Yonghua Zhang, Jiawan Zhang, and Xiaojie Guo. 2019. Kindling the darkness: A practical low-light image enhancer. In Proceedings of the 27th ACM International Conference on Multimedia. 1632-1640.). Recent methods focus on teacherless low-light image enhancement, and models can be trained directly using low-light images without any paired training data. The recent Zero-DCE (Chunle Guo, Chongyi Li, Jichang Guo, Chen Change Loy, Junhui Hou, Sam Kwong, and Runmin Cong. 2020. Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 1780-1789.) trains a deep low-light image enhancement model using a reference-free loss. However, conventional deep learning methods often only focus on enhancing low-light images with insufficient brightness, but in low-light images in backlight conditions and scenarios with uneven illumination, there may also be phenomena of normal brightness or overexposure.

[0004] On the one hand, for the image enhancement task, the most important evaluation criterion is the subjective evaluation of the user. However, in the conventional method, in the training stage of the model, objective evaluation indicators (loss functions) with or without reference are often adopted to guide the training of the model. Here, the loss function with reference mainly includes L 1 loss, L 2 loss and SSIM loss, and the loss function without reference mainly uses spatial consistency loss, exposure control loss, color constancy loss, and illumination smoothness loss. The loss functions with or without reference described above focus on the difference between low-light images and normal-brightness images and the characteristics of the images themselves, ignoring the subjective evaluation of the user.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Object of the Invention: In view of the problems existing in the above background art, the present invention provides a low-light image enhancement method based on reinforcement learning and aesthetic evaluation. First, considering the low-light image imaging method and the complexity of the scenario, a wider operating space range is defined, which includes not only operations to improve the pixel brightness of the image but also operations to reduce the pixel brightness of the image. The enhancement operation can be performed multiple times, and by learning a single random enhancement policy, it has higher flexibility for real-world scenarios. Next, when calculating the loss function, a more flexible loss without reference is used, and at the same time, a new aesthetic quality scoring that can approximate the subjective evaluation index of the user is introduced as part of the loss function.

Means for Solving the Problems

[0006] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is as follows.

[0007] A low-light image enhancement method based on reinforcement learning and aesthetic evaluation, comprising the following steps.

[0008] Step S1: Generate abnormal luminance images in different lighting scenarios, and construct a training dataset for the reinforcement learning system based on the images. Step S2: Initialize the training dataset, policy network, and value network in the reinforcement learning system. Step S3: Update the policy network and value network based on the reward value without reference and the aesthetic evaluation reward value. Step S4: When the training of all samples is completed and all training iteration times are completed, the training of the model is completed. Step S5: Output the image result after enhancing the low-light image.

[0009] Furthermore, the specific method for initializing the policy network and value network in step S2 includes the following.

[0010] JPEG2025518390000002.jpg45170

[0011] Furthermore, the specific steps for updating the policy network and value network in step S3 include the following.

[0012] Step S3.1: Train the training dataset based on the historical stage images to obtain the following environmental reward values. JPEG2025518390000003.jpg24170 Step S3.2: Train the training dataset based on the historical stage images to obtain the output value of the value network. Step S3.3: Update the value network based on the environmental reward value and the output value of the value network. JPEG2025518390000004.jpg98170 Here, θ p represents the policy network parameters.

[0013] JPEG2025518390000005.jpg65170

[0014] Furthermore, the environmental reward value in step S3.1 considers the following influencing factors: (1) Spatial consistency loss JPEG2025518390000006.jpg27170 Here, K represents the size of the local area, and Ω(i) represents the four adjacent areas centered on area i. Y represents the pixel average tone value of the local area in the enhanced image, and I represents the pixel average tone value of the local area in the input image. (2) Exposure control loss JPEG2025518390000007.jpg23170 Here, E represents the tone level in the RGB color space of the image pixels, M represents a plurality of non-overlapping local areas, and Y represents the pixel average tone value of one local area in the enhanced image. JPEG2025518390000008.jpg75170(4) Luminance smoothing loss JPEG2025518390000009.jpg56170(5) Aesthetic quality loss To score the aesthetic quality of the enhanced image, an additional deep learning network model for image aesthetic scoring is introduced to perform aesthetic scoring on the image, and further calculate the aesthetic quality loss. Two independent aesthetic scoring models are trained using the color, luminance attributes, and quality attributes of the image respectively. JPEG2025518390000010.jpg82170 The goal of image enhancement is to maximize the reward value r. The smaller the spatial consistency loss, exposure control loss, color constancy loss, and luminance smoothing loss, the better the image quality. The larger the aesthetic quality loss, the better the image quality. Therefore, JPEG2025518390000011.jpg23170 The environmental reward value at time t is expressed as follows under the condition of introducing influencing factors: JPEG2025518390000012.jpg14170

[0015] Furthermore, in the spatial consistency loss, the size K of the local area is set to 4×4.

[0016] Furthermore, in the exposure control loss, E is set to 0.6, and M represents a non-overlapping local region of size 16×16.

Advantages of the Invention

[0017] The beneficial effects are as follows.

[0018] (1) By expanding the range of the action space defined by reinforcement learning, the reinforcement operations obtained from the input low-light images have a larger dynamic range and higher flexibility for real scenarios. Considering the uneven illumination and backlight in low-light scenarios, not only is the operation of enhancing the image brightness set, but the action space also further includes the behavior of darkening the image brightness. Such a definition method can better meet the low-light image enhancement requirements in real scenarios.

[0019] (2) By introducing the score of aesthetic quality evaluation as part of the loss function, the enhanced image can have a better vision effect and the user's subjective evaluation score. In conventional deep learning-based low-light enhancement methods, most methods rely on paired training datasets and adopt reference loss functions, and some methods use the information of the image itself to design a reference-free loss function to guide the training of the enhancement network. However, all the above losses are objective evaluation indicators. The present invention can better guide the low-light image enhancement network to generate high-quality images that satisfy users by introducing the aesthetic evaluation score as an indicator to simulate the user's subjective evaluation.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Embodiments for Carrying Out the Invention

[0021] Hereinafter, the present invention will be further described with reference to the drawings. Obviously, the described embodiments are some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained on the premise that those skilled in the art do not pay creative labor belong to the protection scope of the present invention.

[0022] The present invention provides a low-light image enhancement method based on reinforcement learning and aesthetic evaluation. The specific principle is as shown in FIG. 1 and includes the following steps.

[0023] Step S1: Generate abnormal luminance images in different lighting scenarios and construct a training dataset for the reinforcement learning system based on the images.

[0024] Step S2: Initialize the training dataset, policy network, and value network in the reinforcement learning system.

[0025] JPEG2025518390000013.jpg59170

[0026] Step S3: Update the policy network and value network based on the reward value without reference and the aesthetic evaluation reward value. Specifically, Step S3.1: Train the training dataset based on the historical stage images to obtain the following environmental reward values, JPEG2025518390000014.jpg95170

[0027] (2) Exposure control loss JPEG2025518390000015.jpg60170

[0028] (3) Color constancy loss JPEG2025518390000016.jpg52170

[0029] (4) Luminance smoothing loss JPEG2025518390000017.jpg56170

[0030] (5) Aesthetic quality loss Currently, aesthetic image analysis has been attracting more and more attention in the field of computer vision. It is related to the high perception of vision aesthetics. The machine learning model used for image aesthetic quality evaluation has broad application prospects, such as in image search, photo management, image editing and photography. For humans, aesthetic quality evaluation is always related to the color and brightness of the image, the quality of the image, composition and depth, and the semantic content. It is difficult to regard aesthetic quality evaluation as an isolated task. In order to score the aesthetic quality of the enhanced image, the present invention additionally introduces an image aesthetic scoring deep learning network model to perform aesthetic scoring on the image, and further calculates the aesthetic quality loss. JPEG2025518390000018.jpg80170

[0031] The goal of image enhancement is to make the reward value r as large as possible. The smaller the spatial consistency loss, exposure control loss, color constancy loss, and luminance smoothing loss are, the better the image quality is, and the larger the aesthetic quality loss is, the better the image quality is. JPEG2025518390000019.jpg48170

[0032] Step S3.2: Train the training dataset based on the historical stage image to obtain the value network output value.

[0033] Step S3.3: Update the value network based on the environmental reward value and the value network output value. JPEG2025518390000020.jpg28170

[0034] Step S3.4: Update the policy network based on the environmental reward value and the value network output value. JPEG2025518390000021.jpg82170

[0035] In each step of the enhancement operation, first, the low-light image in this state is input into the policy network. The policy network outputs, for each pixel in the image, a reinforcement policy based on the currently input image. The input image executes the enhancement operation according to the policy determined by the policy network. This enhancement operation needs to be repeated multiple times according to a predefined plan.

[0036] The action space A is extremely important for the performance of the network. If the range is too small, the enhancement of the low-light image is limited. If the range is too large, it will lead to a very large search space, making the network training extremely difficult. In this embodiment, based on experience, the range A ∈ [-0.5, 0.5] is set, and the step distance is set to 0.05. This setting acts on the predefined output representation, specifically as follows: JPEG2025518390000022.jpg57170

[0037] Similar to the adjustment of the luminance curve of the image used in photo editing software, the predefined output representation here is a quadratic curve expressed as follows: JPEG2025518390000023.jpg44170

[0038] With the above settings, the following can be ensured.

[0039] a. Each pixel is within the normalization range of [0, 1].

[0040] b. Reduce the cost of finding an appropriate reinforcement policy. For different selections of the number of reinforcement iterations, our reinforcement curve can effectively cover the pixel value space in this action space setting.

[0041] Step S4: When all sample training is completed and all training iterations are completed, the training of the model is completed.

[0042] Step S5: Output the image result after enhancing the low-light image.

[0043] As described above, it is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make some improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be regarded as within the protection scope of the present invention.

Claims

1. A low-light image enhancement method based on reinforcement learning and aesthetic evaluation, comprising: generating abnormal luminance images in different lighting scenarios, and constructing a training dataset of a reinforcement learning system based on the images in step S1; initializing a training dataset, a policy network, and a value network in the reinforcement learning system in step S2; updating the policy network and the value network based on a reward value without reference and an aesthetic evaluation reward value in step S3; when the training of all samples is completed and all training iteration times are completed, completing the training of the model in step S4; and outputting an image result after enhancing the low-light image in step S5. A low-light image enhancement method based on reinforcement learning and aesthetic evaluation is characterized by the above.

2. The specific method for initializing the policy network and the value network in step S2 is: The low-light image enhancement method based on reinforcement learning and aesthetic evaluation according to claim 1, characterized by the above.

3. The specific steps for updating the policy network and the value network in step S3 are: training the training dataset based on the historical stage image to obtain an environmental reward value as follows in step S3.1, training the training dataset based on the historical stage image to obtain a value network output value in step S3.2; updating the value network based on the environmental reward value and the value network output value in step S3.3; updating the policy network based on the environmental reward value and the predicted value in step S3.

4. The low-light image enhancement method based on reinforcement learning and aesthetic evaluation according to claim 2, characterized by the above.

4. In step S3.4, The low-light image enhancement method based on reinforcement learning and aesthetic evaluation according to claim 3, characterized by the above.

5. The environmental reward value in step S3.1 considers the following influencing factors: (1) Spatial consistency loss Y represents the pixel average tone value of a local area in the enhanced image, and I represents the pixel average tone value of a local area in the input image. (2) Exposure control loss (3) Color constancy loss (4) Luminance smoothing loss (5) Aesthetic quality loss To score the aesthetic quality of the enhanced image, an additional image aesthetic scoring deep learning network model is introduced to perform aesthetic scoring on the image, and further, the aesthetic quality loss is calculated. Two independent aesthetic scoring models are trained using the color and luminance attributes and the quality attributes of the image respectively. The goal of image enhancement is to maximize the reward value r. The smaller the spatial consistency loss, exposure control loss, color constancy loss, and luminance smoothing loss, the better the image quality. The larger the aesthetic quality loss, the better the image quality. Therefore, A low-light image enhancement method based on reinforcement learning and aesthetic evaluation according to claim 3, characterized in that.

6. In the spatial consistency loss, the size K of the local area is set to 4×4. A low-light image enhancement method based on reinforcement learning and aesthetic evaluation according to claim 5, characterized in that.

7. In the exposure control loss, E is set to 0.6, and M represents a non-overlapping local area of size 16×16. A low-light image enhancement method based on reinforcement learning and aesthetic evaluation according to claim 5, characterized in that.

Citation Information

Patent Citations

  • Underwater image enhancement method based on imaging model and reinforcement learning

    CN114037622A