Improved EnlighttenGAN-based water surface floating object image enhancement method in low-light environment

Through the improved EnlightenGAN architecture, combined with feature extraction and detail preservation subnets, the contradiction between brightness and details in image enhancement of floating objects on the water surface in low-light environments is resolved, efficient image enhancement effects are achieved, and the efficiency and accuracy of water area monitoring are improved.

CN120707406APending Publication Date: 2025-09-26HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410346083.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing GAN-based image enhancement methods for floating objects on the water surface in low-light environments find it difficult to effectively retain image details while improving image brightness and suppress noise amplification, affecting the efficiency and accuracy of water area monitoring.

Method used

An improved EnlightenGAN architecture is adopted, combining the feature extraction subnet and the detail preservation subnet. Through the feature smoothing block, the explicit enhancement center module and the spatial attention module, the image brightness is enhanced and the details of floating objects on the water surface are preserved.

Benefits of technology

While improving the image brightness, it effectively retains the edge and texture information of floating objects on the water surface, improving the image quality and detection accuracy in low-light environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707406A_ABST
    Figure CN120707406A_ABST
Patent Text Reader

Abstract

The invention discloses a method for enhancing a water surface floating object image in a low-light environment based on an improved EnlighttenGAN, and aims to provide a method for providing a water surface floating object image with better quality in the low-light environment. A network improved by the method comprises two parallel networks, namely a feature extraction subnet and a detail preservation subnet. Wherein the feature extraction sub-network adopts an encoder-decoder architecture, refers to the design of an EnlithenGAN network, and is improved by using a proposed explicit enhancement center module, so that the enhanced picture can highlight the features of a to-be-detected floating object while the overall quality is good. The detail preservation subnet does not adopt any downsampling operation, and can maintain the resolution of an original image and preserve image details in the image enhancement process. The method provided by the invention can effectively enhance the quality of the water surface floating object image in the low-light scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular relates to a method for enhancing images of floating objects on a water surface in a low-light environment. Background Art

[0002] With the rapid development of artificial intelligence and image processing technology, low-light image enhancement has become a hot topic of research, especially in areas such as water monitoring and environmental protection. Effective identification of floating objects on the water surface in low-light environments is particularly critical. However, images of floating objects captured at night or in low-light conditions often suffer from low brightness, high noise, and loss of detail due to insufficient light, which seriously affects subsequent image recognition and analysis. While traditional image enhancement techniques such as histogram equalization and Retinex theory have some effectiveness, they often struggle to balance the contradiction between brightness enhancement and detail preservation when processing images from these special scenes, resulting in poor enhancement results.

[0003] In response to the above problems, image enhancement methods based on deep learning have received widespread attention in recent years. In particular, the powerful capabilities of generative adversarial networks (GANs) in unsupervised learning scenarios have provided a new solution for low-light image enhancement. However, existing GAN-based methods still face challenges when dealing with image enhancement of floating objects on the water surface in low-light environments. How to significantly improve the image brightness while effectively retaining the details of floating objects in the image and suppressing noise amplification has become a key technical problem that needs to be solved urgently. Therefore, developing an image enhancement method that can effectively improve the quality of low-light images of floating objects on the water surface while ensuring that the details of floating objects are clearly discernible is of great significance to improving the efficiency and accuracy of water area monitoring. Summary of the Invention

[0004] Purpose of the invention: The purpose of the present invention is to address the deficiencies in the prior art and provide an image enhancement method for floating objects on the water surface in low-light environments based on an improved EnlightenGAN. Combining relevant image processing technologies, this invention provides an image enhancement method that can effectively improve the overall brightness and contrast of the image while retaining the detailed information of floating objects.

[0005] Technical solution: The present invention provides a method for enhancing images of floating objects on the water surface in a low-light environment based on an improved EnlightenGAN, comprising the following steps:

[0006] Step S1: The input low-light image first undergoes basic preprocessing operations, including denoising and data augmentation operations, to improve the stability and robustness of image processing;

[0007] Step S2: The image after basic preprocessing is then input into the feature smoothing block. This module uses a small-sized convolution kernel for convolution operation, combined with batch normalization and nonlinear activation function to effectively smooth the image and reduce noise while retaining important edge and texture information.

[0008] In step S3, the image processed by the feature smoothing block is fed into a feature extraction subnetwork based on the improved EnlightenGAN architecture. In this step, the Explicit Enhancement Center (EEC) module is used to further enhance key image features. The encoder-decoder structure extracts and reconstructs global and local features of the image to effectively improve image brightness and detail.

[0009] Step S4: At the same time, the image processed by the feature smoothing block is also fed into the detail preservation subnet in parallel. This subnet does not use any pooling operation and emphasizes preserving the original image information at its original resolution during the image enhancement process, especially the edges and textures of floating objects on the water surface, so as to maintain the naturalness and realism of the image while improving the brightness.

[0010] Step S5: The outputs of the feature extraction subnet and the detail preservation subnet are fed into a fusion layer, which concatenates the 32-channel feature maps of the two subnets to form a 64-channel feature map. This layer then passes through a 3×3 convolutional layer and a tanh activation function to output the enhanced image.

[0011] In step S6, the output image of the fusion layer is then post-processed, including color correction, contrast adjustment, and sharpening, to further improve the visual quality of the image. Furthermore, specific optimization processes, such as dynamic range adjustment, can be performed as needed to ensure that the enhanced image maintains good visual quality on different display devices.

[0012] The step S3 is to extract the features of the image of floating objects on the water surface in a low-light environment, and the step S3 is further as follows:

[0013] Step S3.1: In the feature extraction subnet, the input image first passes through the encoder. The encoder consists of multiple convolutional layers, each followed by batch normalization and ReLU activation functions to extract deep features of the image.

[0014] Step S3.2: After preliminary feature extraction, the image features are fed into the Explicit Enhancement Center (EEC) module. The EEC module explicitly enhances the key features of the image by combining a lightweight multi-layer perceptron (MLP) and a learnable enhancement center (LEC).

[0015] Step S3.3: In the explicit enhancement center module, the image features are first passed through a lightweight multi-layer perceptron, which is designed to capture global long-range dependencies and strengthen the feature representation of important areas in the image. is the output of the depthwise convolution module, GN(·) represents the group normalization operation, and DConv(·) is a depthwise convolution with a kernel size of 1×1. The process of this step is shown in the following formula:

[0016]

[0017] For channel MLP based modules, from depthwise convolution based modules The output features are first group normalized, and then channel MLP is implemented on these features. Finally, channel scaling, path random dropout and The residual connection of this step is as follows:

[0018]

[0019] Among them, CMLP represents the channel MLP operation;

[0020] Step S3.4, then, the features are passed through the learnable enhancement center, which aggregates the local features of the image and further refines and enhances the local details of the image content, especially for key objects such as floating objects on the water surface. in First, the encoding is performed through a combination of a set of convolutional layers. The encoded features are then processed by a CBR block, which consists of a 3×3 convolutional layer followed by a batch normalization layer and a ReLU activation function. The encoded features are then processed by a CBR unit. This unit consists of a 3×3 convolutional layer, a BN layer for batch normalization, and a ReLU function for activation. Through the above steps, the encoded features are is input into the codebook. Then, a set of scaling factors S is used to make and b k Mapped to the corresponding position information. The information of the entire image relative to the kth codeword can be calculated as follows:

[0021]

[0022] in, is the i-th pixel, b k is the kth learnable visual codeword, s k is the kth scaling factor, It describes the position information of each pixel relative to the codeword, and K is the total number of enhancement centers. Then use the Φ function to integrate all e k, where the Φ function includes a batch normalization layer, a ReLU activation function, and a mean layer. Based on this, the complete information of the entire image relative to the kth codeword can be calculated:

[0023]

[0024] After obtaining the output of the codebook, it is further input into a fully connected layer and a 1×1 convolutional layer to enhance the features of key salient categories. Afterwards, the input features X from the feature smoothing block are in Multiply the channel by channel with the scale factor coefficient δ(·). The above process can be expressed as:

[0025]

[0026] Among them, Conv 1×1 represents 1×1 convolution, δ(·) represents the sigmoid activation function, Represents channel-by-channel multiplication. Finally, the feature X output from the feature smoothing block in And the local key feature z is added channel by channel, and the formula is as follows:

[0027]

[0028] in Represents a channel-wise addition operation.

[0029] The step S4 is to preserve the details of the image of floating objects on the water surface in a low-light environment, and the step S4 is further as follows:

[0030] Step S4.1, use two convolutional layers of size 3×3 and stride 1 to extract features from the input image;

[0031] Step S4.2: Use 15 residual blocks to operate on the convolution feature map to maintain the color and texture information of the floating object image;

[0032] Step S4.3: Use the spatial attention module to further process the feature map, combining the average pooling and maximum pooling methods to extract different information and more effectively fit the illumination distribution of the normal brightness image.

[0033] Beneficial effects: Compared with the prior art, the advantages of the present invention are:

[0034] This paper is based on the existing EnlightenGAN network structure, comprehensively considers the characteristics of floating objects on the water surface, combines a dual-path parallel subnetwork structure, uses a feature extraction subnetwork to extract different levels of image features, and a detail preservation subnetwork to maintain image resolution and retain image details, which can effectively enhance the image of floating objects on the water surface in low-light environments.

[0035] Compared with traditional enhancement methods, the present invention overcomes the problem of limited receptive field of CNN convolutional network. At the same time, compared with the comprehensive and non-discriminatory learning of transformer network, the present invention can enhance the feature representation of local key areas with less computational overhead, which is crucial for maintaining the consistency of details of floating objects on the water surface under extreme lighting conditions. It has important practical significance in low-light image enhancement and detection of floating objects on the water surface at night. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is the overall flow chart of the present invention.

[0037] Figure 2 Comparison of image enhancement results in partial exposure scenarios.

[0038] Figure 3 Comparison of image enhancement results in extremely dark scenes.

[0039] Figure 4 Comparison of image enhancement results of the algorithm in overall low-brightness scenes. DETAILED DESCRIPTION

[0040] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0041] like Figure 1 As shown, a method for enhancing images of floating objects on the water surface in a low-light environment based on an improved EnlightenGAN in this embodiment includes the following steps:

[0042] Step S1: Float-night, a dataset of floating objects at night, was constructed using real-life photography and online downloading. The dataset was augmented through rotation, horizontal mirroring, and image distortion. The resulting dataset contains 3,000 images, covering five categories: bottles, aquatic plants, accumulated garbage, humans, and boats. Of these, 7,678 objects are small, accounting for 5,988 (78%).

[0043] Step S2: Use Python program to build the improved model of the present invention, which mainly includes two parallel networks: feature extraction subnet and detail preservation subnet;

[0044] Step S3: Initialize the training parameters of the model, set the image size to 320×320, the batch size to 32, the learning rate to 0.0001, and the number of iterations to 200;

[0045] Step S4: start model training;

[0046] Step S5: Update the training parameters according to the iteration results of each round and the parameter update strategy;

[0047] Step S6: determine whether the model has converged;

[0048] Step S7: If it has not converged yet, jump to step S2;

[0049] Step S8: If the model has converged, test the test set and verify the validation set.

[0050] To verify the effectiveness of this method, we conducted experiments using the Float-night dataset. We also compared the proposed method with the Equalized, Zero-DCE, RUAS, and EnlightenGAN methods. The evaluation metrics for the different enhancement methods are shown in Table 1. Figure 2 、 Figure 3 and Figure 4 The results of low-light image enhancement for floating objects on the water surface are shown in the following figures: a partially overexposed scene, an extremely dark scene, and an overall low-brightness scene. The above charts demonstrate that the low-light floating object enhancement algorithm based on the improved EnlightenGAN in this embodiment can enhance the quality of floating object images while maintaining a short training time, with significantly better enhancement efficiency than other methods.

[0051] Table 1 Comparison of different enhancement algorithms

[0052]

Claims

1. A method for enhancing images of floating objects on the water surface in low-light environments based on an improved EnlightenGAN, characterized by: The following steps are involved: Step S1: The input low-light image first undergoes basic preprocessing operations, including denoising and data augmentation operations, to improve the stability and robustness of image processing; Step S2: The image after basic preprocessing is then input into the feature smoothing block. This module uses a small-sized convolution kernel for convolution operation, combined with batch normalization and nonlinear activation function to effectively smooth the image and reduce noise while retaining important edge and texture information. In step S3, the image processed by the feature smoothing block is fed into a feature extraction subnetwork based on the improved EnlightenGAN architecture. In this step, the Explicit Enhancement Center (EEC) module is used to further enhance key image features. The encoder-decoder structure extracts and reconstructs global and local features of the image to effectively improve image brightness and detail. Step S4: At the same time, the image processed by the feature smoothing block is also fed into the detail preservation subnet in parallel. This subnet does not use any pooling operation and emphasizes preserving the original image information at its original resolution during the image enhancement process, especially the edges and textures of floating objects on the water surface, so as to maintain the naturalness and realism of the image while improving the brightness. Step S5: The outputs of the feature extraction subnet and the detail preservation subnet are fed into the fusion layer, which splices the two subnets 32 The feature map of the channel is formed into a feature map containing 64 channels, which is then processed by a 3×3 convolution layer and a tanh activation function to output the enhanced image; In step S6, the output image of the fusion layer is then post-processed, including color correction, contrast adjustment, and sharpening, to further improve the visual quality of the image. Furthermore, specific optimization processes, such as dynamic range adjustment, can be performed as needed to ensure that the enhanced image maintains good visual quality on different display devices.

2. The method for enhancing images of floating objects on water surfaces in low-light environments based on the improved EnlightenGAN according to claim 1 is characterized by: Step S3.1: In the feature extraction subnet, the input image first passes through the encoder. The encoder consists of multiple convolutional layers, each followed by batch normalization and ReLU activation functions to extract deep features of the image. Step S3.2: After preliminary feature extraction, the image features are fed into the Explicit Enhancement Center (EEC) module. The EEC module explicitly enhances the key features of the image by combining a lightweight multi-layer perceptron (MLP) and a learnable enhancement center (LEC). Step S3.3: In the explicit enhancement center module, the image features are first passed through a lightweight multi-layer perceptron, which is designed to capture global long-range dependencies and strengthen the feature representation of important areas in the image. is the output of the depthwise convolution module, GN(·) represents the group normalization operation, and DConv(·) is a depthwise convolution with a kernel size of 1×1. The process of this step is shown in the following formula: For channel MLP based modules, from depthwise convolution based modules The output features are first group normalized, and then channel MLP is implemented on these features. Finally, channel scaling, path random dropout and The residual connection of this step is as follows: Among them, CMLP represents the channel MLP operation; Step S3.4, then, the features pass through the learnable enhancement center, which aggregates the local features of the image and further refines and enhances the local details of the image content, especially for key objects such as floating objects on the water surface. The feature Xin processed by the feature smoothing block is first encoded by a combination of a set of convolutional layers. Then, the encoded features are processed by the CBR block, which includes a 3×3 convolutional layer, followed by a batch normalization layer and a ReLU activation function. The encoded features are then processed by the CBR unit. This unit consists of a 3×3 convolutional layer, and also includes a BN layer for batch normalization and a ReLU function for activation. Through the above steps, the encoded features is input into the codebook. Then, a set of scaling factors S is used to make and b k Mapped to the corresponding position information. The information of the entire image relative to the kth codeword can be calculated as follows: in, is the i-th pixel, b k is the kth learnable visual codeword, s k is the kth scaling factor, It describes the position information of each pixel relative to the codeword, and K is the total number of enhancement centers. Then use the Φ function to integrate all e k , where the Φ function includes a batch normalization layer, a ReLU activation function, and a mean layer. Based on this, the complete information of the entire image relative to the kth codeword can be calculated: After obtaining the output of the codebook, it is further input into a fully connected layer and a 1×1 convolutional layer to enhance the features of key salient categories. Afterwards, the input features X from the feature smoothing block are in Multiply the channel by channel with the scale factor coefficient δ(·). The above process can be expressed as: Among them, Conv 1×1 represents 1×1 convolution, δ(·) represents the sigmoid activation function, Represents channel-by-channel multiplication. Finally, the feature X output from the feature smoothing block in And the local key feature z is added channel by channel, and the formula is as follows: in Represents a channel-wise addition operation.

3. The method for enhancing images of floating objects on water surfaces in low-light environments based on the improved EnlightenGAN according to claim 1 is characterized by: Step S4.1, use two convolutional layers of size 3×3 and stride 1 to extract features from the input image; Step S4.2: Use 15 residual blocks to operate on the convolution feature map to maintain the color and texture information of the floating object image; Step S4.3: Use the spatial attention module to further process the feature map, combining the average pooling and maximum pooling methods to extract different information and more effectively fit the illumination distribution of the normal brightness image.