Multi-scale and attention mechanism-based low-light image anti-overexposure enhancement method

Through the low-light image enhancement method based on multi-scale and attention mechanism, the problems of insufficient image brightness and overexposed under low-light conditions are solved, and the image enhancement effect with high delicateness and strong sense of hierarchy is achieved, and the generalization ability and processing speed of the model are improved.

CN120107134AInactive Publication Date: 2025-06-06CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510269568.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Images captured under low light conditions often have problems such as insufficient brightness, significant noise, and color distortion. Traditional image enhancement algorithms are prone to local overexposure, which is difficult to meet the dual needs of dynamic range compression and detail retention.

Method used

The low-light image anti-overexposure enhancement method based on multi-scale and attention mechanisms is adopted to obtain information at different frequency levels through Laplace pyramid decomposition, and combine the multi-scale fusion network and SE/CA attention mechanism of the U-Net architecture to perform cross-frequency fusion and detail enhancement.

Benefits of technology

Effectively capture different levels of information of the image, improve the delicateness and sense of hierarchy of the enhancement effect, improve the generalization ability of the method, reduce the complexity of the model, significantly improve the processing speed, and show stronger robustness and adaptability in different low-light scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107134A_ABST
    Figure CN120107134A_ABST
Patent Text Reader

Abstract

The invention provides a low-light image anti-overexposure enhancement method based on multiple scales and an attention mechanism. The problem of overexposure after image enhancement in the existing method is solved. According to the method, the Laplacian pyramid of the image is utilized, the deep neural network is used for extracting multi-scale information of the low-light image, and detail features of the image are effectively mined and reserved. And adjusting and fusing the multi-scale information by adopting a targeted processing strategy, and recovering to obtain a fine illumination image capable of accurately representing external illumination distribution. And a targeted attention mechanism is added during processing of each sub-network for adjustment so as to realize image enhancement. According to the method, the Laplacian pyramid information of the image is fully utilized, and the effect of enhancing the dark light image is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image enhancement technology, and in particular to a low-light image overexposure prevention enhancement method based on multi-scale and attention mechanism. Background Art

[0002] In real-world scenarios such as smartphone photography, security monitoring, and autonomous driving, images captured under low-light conditions often suffer from insufficient brightness, significant noise, and color distortion, which seriously restrict subsequent visual analysis tasks. Traditional image enhancement algorithms (such as histogram equalization and gamma correction) can easily lead to local overexposure by globally adjusting pixel intensity distribution, resulting in loss of detail in highlight areas, and are unable to meet the dual requirements of dynamic range compression and detail retention in real-world scenarios.

[0003] Traditional image enhancement methods can be divided into two types: histogram equalization-based methods and Retinex-based methods. The former improves the brightness of the image by adjusting the grayscale range of the image, but the effect on detail enhancement is poor. The latter is based on the Retinex visual theory and believes that the low-light image is the product of the object's reflection image and the external illumination image. The low-light image is processed by Gaussian blurring to obtain the illumination image, and then the illumination effect is removed by mathematical operations to finally obtain the reflection image, thereby achieving image enhancement. However, the Retinex-based method has limited effect in enhancing color images and is prone to color distortion.

[0004] Although deep learning-based methods (such as LLNet and KinD) can model nonlinear mappings end-to-end, they have the following problems: single-scale feature extraction is difficult to take into account both global illumination distribution and local detail recovery; existing networks lack the ability to perceive dynamic highlights, which can easily lead to pixel saturation due to over-enhancement; some models improve performance through cascade modules, but sacrifice real-time performance due to high computational complexity. Attention mechanisms (such as SENet and CBAM) improve feature expression efficiency through channel / spatial weight allocation and have been initially applied in the field of image enhancement. Channel attention: used to suppress noise-related feature channels (such as DRBN); spatial attention: locate dark areas in low-light enhancement (such as Zero-DCE). However, there are also limitations: existing attention modules focus more on dark area enhancement and lack predictive suppression of potential overexposed areas. Multi-scale architectures (such as U-Net and pyramid networks) achieve detail retention by fusing features of different resolutions, while dynamic attention mechanisms can adaptively adjust the intensity of regional enhancement. Recent studies (such as SNR-Aware networks) have shown: multi-scale feature complementarity: shallow features retain texture details, and deep features model global illumination; attention-guided enhancement: suppress overexposed areas through learnable weight distribution; joint optimization potential: the combination of the two can simultaneously achieve light restoration, noise suppression and overexposure prevention.

[0005] In summary, both traditional methods and deep learning methods have certain limitations. Traditional methods are simple to implement and have a fast processing speed. They do not rely on training data sets, but their generalization ability is weak and it is difficult to achieve effective enhancement in low-light images in multiple scenes. Although deep learning methods can improve the generalization ability of the model with the help of deep networks and large-scale data sets, thereby obtaining more accurate image enhancement effects, they often rely too much on domain knowledge and have high model complexity, which may lead to problems such as loss of image texture details or color distortion. Therefore, in order to overcome these challenges, a low-light image enhancement method based on multi-scale processing and attention mechanism is proposed. Multi-scale processing can effectively capture different levels of information in the image and improve the delicacy and layering of the enhancement effect, while the attention mechanism can guide the model to focus on the key detail areas in the image, thereby ensuring the clarity and realism of the image enhancement while improving the generalization ability of the method, reducing the complexity of the model, and significantly improving the processing speed. This new method can not only enhance image details, but also improve performance in different low-light scenes, showing stronger robustness and adaptability. Summary of the invention

[0006] In view of the shortcomings and improvement needs of existing deep learning-based methods, the present invention provides a low-light image overexposure prevention and enhancement method based on multi-scale and attention mechanism. The specific scheme includes the following steps:

[0007] Step 1: Perform Laplacian pyramid decomposition on the low-light image to obtain low, medium and high frequency level information of the low-light image;

[0008] Step 2: Use the network to extract low-frequency global color information and medium- and high-frequency image coarse and detailed information of the low-light image;

[0009] Step 3: Process the image details in a targeted manner by using the attention mechanism to process information at different frequency levels;

[0010] Step 4: Cross-frequency fusion of information at different frequency levels to obtain a reconstructed image;

[0011] Furthermore, the specific process of the above step 1 is: using the low-light image as the bottom image, a Gaussian image pyramid is established from bottom to top. The specific implementation process is shown in formulas (1), (2), (3), and (4):

[0012]

[0013]

[0014] In formula (1), G i represents the i-th layer image of the pyramid, there are n layers in total, Down() represents the downsampling operation, represents the convolution operation, g k×krepresents a Gaussian convolution kernel of size k×k; Formula (2) is a Gaussian kernel function, where σ represents a scale parameter; the above process is iterated multiple times to obtain image pyramids of different heights; in Formula (3), Up() represents an upsampling operation; in Formula (4), L i-1 Represents the Laplacian pyramid image of the i-1th layer, which contains the detailed information of this layer. Repeat the above process until the Laplacian pyramid images L of all layers are calculated. i (i=0,1,2,…n-2).

[0015] Furthermore, the above step 2 is specifically as follows: design a multi-scale fusion network based on the U-Net architecture, which consists of three independent sub-networks, each of which is responsible for extracting the image information of each layer of the Laplacian pyramid of the low-light image. The main structure of the network adopts the encoder-decoder framework, and the Swish and Mish activation functions are introduced in the encoder-decoder framework through multi-level downsampling and upsampling operations to realize feature extraction and reconstruction, respectively extracting the color distribution and brightness information of the low-light image, object contours and coarse-grained texture, fine texture and noise features.

[0016] Furthermore, the above step three is specifically as follows: the color distribution and brightness information obtained in step two is passed through the SE attention mechanism, the dependencies between channels are learned through the fully connected layer, channel weights are generated, the learned channel weights are multiplied by the original feature map, and the channel features are recalibrated; the object contour and coarse-grained texture obtained in step two are passed through the CA attention mechanism, horizontal and vertical attention weights are generated through convolution and activation functions, the generated attention weights are multiplied by the original feature map, and the edge and texture areas are enhanced; the fine texture and noise features obtained in step two are passed through global average pooling and a fully connected layer to generate channel attention weights to enhance the features of important channels, spatial attention weights are generated through convolution operations, and the areas with rich details are focused on. The outputs of channel attention and spatial attention are added to obtain the final high-frequency features.

[0017] Furthermore, the above step 4 is specifically as follows: the image information of each layer of the Laplacian pyramid finally processed in step 3 is integrated by upsampling, and the contribution of each sub-network to the final result is calculated by learning, and the network capacity is allocated in the form of weights. The low-frequency level is processed by the sub-network, and the output of the sub-network is expanded to twice the original using strided transposed convolution to generate an enlarged image I; then the intermediate frequency level is added to the image I in turn, and this refinement sampling process will continue until the final output image is generated.

[0018] Beneficial Effects

[0019] a: The present invention improves the U-Net network framework and the encoder-decoder network structure, and simultaneously captures global and local information when processing images. The jump connection of U-Net directly integrates the low-level features of the encoder with the high-level features of the decoder, avoiding the loss of information in the deep network. The Swish and Mish activation functions are introduced to enhance the nonlinear characteristics of the network. The mixed activation functions are tried between different modules of U-Net to increase the expressiveness and flexibility of the model.

[0020] b: It comprehensively considers the unique features of each level of the image, uses multiple attention mechanisms to fully explore the unique image details of different scales, and dynamically focuses on important areas.

[0021] c: Design a variety of loss functions to adjust their constraints and improve the visual quality of the enhanced image. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A flow chart of the low-light image overexposure prevention and enhancement method based on multi-scale and attention mechanism provided by the present invention;

[0023] Figure 2 A network model flow chart of the low-light image enhancement model provided by the present invention;

[0024] Figure 3 A schematic diagram of a network model of a low-light image enhancement model provided by the present invention;

[0025] Figure 4 The extraction network framework diagram of the low-light image enhancement model provided by the present invention; Specific implementation plan

[0026] The technical solutions in the embodiments of the present invention are fully described below in conjunction with the accompanying drawings of the examples of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0027] The present invention provides a low-light image anti-overexposure enhancement method based on multi-scale and attention mechanism, combined with Figure 1 The method is specifically implemented by the following specific steps:

[0028] Step 1: construct a multi-scale representation of the low-light image to obtain multi-scale information of the low-light image. Image pyramid is an important means of image preprocessing and is widely used. In this step, the Laplacian pyramid is selected as a multi-scale transformation tool to perform multi-scale representation of the low-light image.

[0029] Specifically, the low-light image is used as the bottom image of the pyramid to establish the Laplacian pyramid. The implementation process is shown in formulas (1), (2), (3), and (4):

[0030]

[0031]

[0032] In formula (1), G i Represents the i-th layer image of the pyramid, with a total of n layers, Down() downsampling operation, represents the convolution operation, g k×k represents a Gaussian convolution kernel of size k×k; Formula (2) is a Gaussian kernel function, where σ represents a scale parameter; the above process is iterated multiple times to obtain image pyramids of different heights; in Formula (3), Up() represents an upsampling operation; in Formula (4), L i-1 Represents the Laplacian pyramid image of the i-1th layer, which contains the detailed information of this layer. Repeat the above process until the Laplacian pyramid images L of all layers are calculated. i (i=0,1,2,…n-2).

[0033] Step 2: Use the network to extract the low-frequency global color information and the mid- and high-frequency image coarse and detailed information of the low-light image.

[0034] Specifically, the Laplacian pyramid images of each layer obtained in step 1 are input into the network model. The network model here is composed of multiple sub-networks, and each layer of the Laplacian pyramid image will be processed by an independent sub-network. Each sub-network will output feature maps of its own level, and these feature maps contain information of different frequency bands. In specific implementation, Figure 3 As shown in the figure, the structures of each sub-network are basically the same. They are all network models based on the U-Net network framework using the encoder-decoder main structure. The skip connections of U-Net directly fuse the low-level features of the encoder with the high-level features of the decoder, avoiding the loss of information in the deep network. The encoder-decoder structure can gradually extract high-level semantic features by stacking convolutional layers and pooling layers, and use the Swish activation function in the Encoder layer to retain more useful information and alleviate the gradient disappearance problem. The Mish activation function is used in the Decoder layer to enhance the model's ability to capture details and improve resolution.

[0035] Step 3: Process the image details in a targeted manner by using the attention mechanism to process information at different frequency levels.

[0036] Specifically, the low-frequency feature map obtained in step 2 Encode the global color distribution and light intensity, perform global average pooling on each channel, compress the spatial information, and obtain the channel description vector

[0037]

[0038] The nonlinear relationship between channels is learned through the fully connected layer to generate channel weights:

[0039] s=σ(W 2 ·δ(W 1 ·z)) (6)

[0040] in is a learnable parameter, r is the compression ratio set to 16, δ is the ReLU activation function, and σ is the Sigmoid function.

[0041] Multiply the channel weight s by the original feature channel by channel to enhance the key color channels (such as the dominant color system) and suppress redundant brightness information:

[0042]

[0043] The intermediate frequency feature map obtained in step 2 The feature map is averaged in the horizontal (X) and vertical (Y) directions to generate position encodings in two directions:

[0044]

[0045] Concatenate horizontal and vertical encodings to obtain position-sensitive features Generate attention map through 1×1 convolution + batch normalization + nonlinear activation (Swish): A = Swish (BN (Conv 1×1 (Z))).

[0046] Split into horizontal attention along the spatial dimension and vertical attention Multiply the attention map by the original features element-wise to strengthen the contour response:

[0047]

[0048] The high-frequency feature map obtained in step 2 Global average pooling generates channel description The weights generated by the fully connected layer are: s channel =σ(W channel ·δ(W reduce ·z channel )) Output channel weighted features The mean and standard deviation are calculated along the channel dimension, and then concatenated and subjected to 7×7 convolution to generate a spatial weight map: s space =σ(Conv 7×7 ([μ(F high );σ(F high )])); Output spatial weighted features Add the channel and spatial attention results to preserve details while suppressing noise:

[0049]

[0050] Concatenate the optimized low, medium, and high frequency features according to the weights:

[0051]

[0052] Where W low , W mid , W high It is a learnable 1×1 convolution kernel used to dynamically adjust the contribution of each frequency band. The final output is sent to the subsequent reconstruction module to generate an enhanced image.

[0053] Step 4: Perform cross-frequency fusion on information of different frequency levels to obtain a reconstructed image.

[0054] Specifically, the enhanced multi-frequency feature map output from step 3 is input into the reconstruction network, and the low-frequency feature F low Use strided transposed convolution (deconvolution) to perform 2x upsampling:

[0055] I up =Deconv 3×3 (F low ,stride=2,pad=1) (13)

[0056] Generate initial reconstructed image I base ; The upsampling result is fused with the original low-frequency features through residual connection:

[0057]

[0058] The intermediate frequency feature F mid and Weighted addition after alignment:

[0059]

[0060] Perform 2x upsampling again to generate an intermediate image The high frequency feature F high and To inject details:

[0061]

[0062] Preserve original high-frequency details through skip connections:

[0063] I out =I final +SkipConn(F high ) (17)

[0064] The above contents are all embodiments of the present invention. In the network model training stage, low-light and normal-light image pairs in the low-light image dataset are respectively input into the network, the former is used as the network input, and the latter is used as the label, and the loss function is used to measure the loss between the enhanced image and the label. The network is gradient-transferred based on the loss, the network parameters are updated, and the network is iteratively trained multiple times.

[0065] In the network model training stage, the image enhancement method provided by the present invention inputs low-light and normal-light image pairs in the low-light image dataset into the network respectively, with the former serving as the network input and the latter serving as labels. A loss function is used to measure the loss between the enhanced image and the label, and the network is gradient-transferred based on the loss to update the network parameters, and the network is trained iteratively for multiple times.

[0066] A composite loss function is used in the network training phase. Specifically, it includes two loss sub-items: the similarity loss L between each frequency level feature and the target feature. Fre ; Reconstruction loss L that measures the difference between the reconstructed image and the normal image Rec .

[0067]

[0068] Among them, G i is the low-frequency image of the i-th layer in the low-light image, is the low-frequency information of the target image. i is the coarse and detailed information of the intermediate frequency image, is the corresponding intermediate frequency detail information in the target image. i is the fine detail information of the high-frequency image, is the high frequency part of the target image. j (·) is the j-th layer feature extracted by the pre-trained network, G is the reconstructed image, is the target image. G(i,j) and are the pixel values ​​of the reconstructed image and the target image at position i, j respectively. By minimizing this total loss function, the model can effectively enhance low-light images while avoiding overexposure and retaining more detail information and naturalness. The final loss function L is composed of two weighted loss function sub-items, which effectively supervises the training process of the network model.

[0069] As described above, the present invention can be better implemented

[0070] The above content is only a detailed description of the calculation model and processing flow of the present invention, and is not a limitation on the implementation of the present invention. Any equivalent structure and equivalent process replacement made using the contents of the present invention specification and drawings is still within the protection scope of the present invention.

Claims

1. A low-light image overexposure prevention and enhancement method based on multi-scale and attention mechanism, characterized in that: The steps include: Step 1: Perform Laplacian pyramid decomposition on the low-light image to obtain low, medium and high frequency level images of the low-light image; Step 2: Use the network to extract low-frequency global color information and medium- and high-frequency image coarse and detailed information of the low-light image; Step 3: Process the image details by using the attention mechanism to process the information at different frequency levels; Step 4: Cross-frequency fusion of information at different frequency levels to obtain a reconstructed image; Step 5. Construct two loss functions to ensure that the detail information of each frequency level (low frequency, medium frequency, high frequency) is effectively processed and the overall quality of the reconstructed image is optimized. The first loss function is based on the difference between the features of each frequency level and the target features, and uses the L1 norm to measure the reconstruction error of low-frequency color information, medium-frequency coarse detail information, and high-frequency fine detail information. By minimizing this loss function, the model better preserves and enhances the multi-scale details of the image. The second loss function combines perceptual loss and pixel-level loss to ensure that the reconstructed image is visually natural and rich in details. The perceptual loss extracts features and calculates feature differences through a pre-trained convolutional neural network, while the pixel-level loss measures the pixel differences between the reconstructed image and the target image. By jointly optimizing these two loss functions, the model can effectively prevent overexposure while enhancing low-light images and retain the details and naturalness of the image.

2. The low-light image anti-overexposure enhancement method based on multi-scale and attention mechanism as claimed in claim 1, characterized in that: The process of step 2 is as follows: A multi-scale fusion network based on an improved U-Net architecture is designed. The network contains three independent sub-networks, each of which is specifically used to extract information from low-light images at different Laplacian pyramid levels. The core structure of the network adopts an encoder-decoder framework, and realizes feature extraction and reconstruction through multi-level downsampling and upsampling operations, improves the nonlinear activation function in the network, solves the gradient vanishing and dead zone problems, and improves the model's expressiveness and training efficiency.

3. The low-light image anti-overexposure enhancement method based on multi-scale and attention mechanism as claimed in claim 1, characterized in that: The process of step three is as follows: The SE attention mechanism is used in the low-frequency information processing of low-light images. By explicitly modeling the channel relationship, it effectively improves the global consistency of the low-frequency illumination distribution and avoids overexposure. The CA attention mechanism is used in the mid-frequency information processing. By capturing position-aware information, it accurately enhances the edges and texture areas of the mid-frequency region. The DANet attention mechanism is used in the high-frequency information processing. Through parallel channel-spatial attention, it adaptively enhances high-frequency details.

4. The low-light image overexposure prevention and enhancement method based on multi-scale and attention mechanism as claimed in claim 1, characterized in that: The loss function in step 5 has two sub-items, namely: Among them, G i is the low-frequency image of the i-th layer in the low-light image, is the low-frequency information of the target image. i is the coarse and detailed information of the intermediate frequency image, is the corresponding intermediate frequency detail information in the target image. i is the fine detail information of the high-frequency image, is the high frequency part of the target image. j (·) is the j-th layer feature extracted by the pre-trained network, G is the reconstructed image, is the target image. G(i, j) and are the pixel values ​​of the reconstructed image and the target image at positions i and j respectively. By minimizing this total loss function, the model can effectively enhance low-light images while avoiding overexposure and retaining more detail information and naturalness.

Citation Information

Patent Citations

  • Low-light image enhancement method and device fusing high and low frequency feature information

    CN116152120A

  • Laplace three-layer cyclic high-definition image enhancement method

    CN117611465A

  • Gray and color image fusion method based on Laplacian pyramid and self-attention mechanism

    CN118071613A

  • Lightweight low-illumination image enhancement method based on convolutional neural network

    CN118351041A

Cited By

  • Low-light image enhancement method based on multilevel feature fusion

    CN121190325A