Method for identifying floating objects on water surface in low-light environment

Through the improved Retinex-Net model and Gold-YoLo feature fusion module, combined with the MPDIoU loss function, the Yolov7 model is optimized, and the problem of low accuracy and insufficient robustness of water surface floating objects in low-light environments is solved, and high-precision and stable floating objects detection are achieved.

CN120088643APending Publication Date: 2025-06-03JIANGSU OCEAN UNIV +1

Patent Information

Application Number
CN202510123393.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has low accuracy in water surface floating objects in low light environments, insufficient robustness, and insufficient detection of small floating objects.

Method used

The improved Retinex-Net model is used for low-light image enhancement, combined with the CBAM attention mechanism and multi-scale feature interaction mechanism, the Yolov7 model is optimized through the Gold-YoLo feature fusion module and MPDIoU loss function, improving the capabilities of feature extraction and object detection.

Benefits of technology

It significantly improves the detection accuracy and robustness of floating objects on the water surface in low-light environments, reduces the leakage detection rate, improves the recognition accuracy of small floating objects, and enhances the stability of the model in complex water surface environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088643A_ABST
    Figure CN120088643A_ABST
Patent Text Reader

Abstract

According to the water surface floating object identification method in the low-light environment provided by the invention, through an improved Yov7 target detection model, a Gold-YoLo feature fusion module, a CBAM-Enhanced Retinx-Net image enhancement network and an MPDIOU loss function, the detection precision of floating objects in the low-light environment is remarkably improved. The CBAM attention mechanism and the adaptive enhancement strategy effectively enhance the brightness, contrast and detail information of the image, and solve the problem of excessive enhancement of the traditional Retinex algorithm under low illumination. The Gold-YoLo module improves the feature extraction capability of the floating object through multi-scale feature interaction. The MPDIOU loss function optimizes the bounding box regression precision and improves the stability in a complex environment. Experimental results show that the method can effectively reduce the omission ratio and improve the small floater identification precision, and is suitable for real-time floater detection in the fields of water quality monitoring, ocean protection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and image processing, and particularly relates to a method for identifying floating objects on the water surface in a low-light environment. Background Art

[0002] The detection of floating objects on the water surface has important application values in many fields such as environmental monitoring, pollution control, marine protection, and water quality monitoring. With the increasing severity of the global water pollution problem, the real-time and accurate detection of floating objects on the water surface has become a difficult problem to be solved urgently. Timely and accurately identifying and locating floating objects is crucial for marine pollution control, especially for cleaning up marine plastic waste. Traditional methods for identifying floating objects on the water surface mostly rely on image processing-based techniques, such as taking water surface images with a camera and using image processing algorithms to analyze the floating objects on the water surface. However, the identification of floating objects on the water surface faces significant challenges in low-light environments (such as at night, on rainy and cloudy days). Due to the reflection of water surface waves and the interference of background noise, traditional image recognition methods often cannot effectively distinguish floating objects from the background, resulting in a significant decrease in recognition accuracy. To solve this problem, many advanced technologies have gradually entered this field, including machine learning methods such as deep learning and convolutional neural network (CNN). With the rapid development of computer vision technology, deep learning-based methods for detecting floating objects on the water surface have been widely used. These methods extract and analyze features of water surface images by training deep neural networks, thereby realizing the detection and classification of floating objects. However, although deep learning methods perform well under normal lighting conditions, their performance often drops sharply in low-light environments, and they have poor robustness to water surface reflection and wave interference, unable to meet the requirements of practical applications.

[0003] The patent with the patent number CN202110789818.6 proposes a multi-camera real-time water surface floating object detection method based on the SSD (Single Shot MultiBox Detector) network. By video recording, camera shooting, and network data collection, using data denoising and enhancement algorithms, and combining transfer learning for training, the real-time performance and accuracy of water surface floating object detection are improved. This method can reduce the interference caused by lighting, weather, and dynamic background to real-time detection, and makes up for the defects of single-camera detection through the collaborative work of multiple cameras. The patent with the patent number CN202111494709.8 proposes a deep learning-based method for detecting floating objects on the water surface, aiming to solve the problem of identifying floating objects in low-light environments. This method selects surveillance videos containing floating objects and uses deep learning inference to detect and classify images. In particular, the model trained under low-light, blurred, and low-contrast backgrounds can effectively improve the robustness of the system, and the application of transfer learning reduces the need for a large amount of sample data.

[0004] Currently, the detection technology of floating objects on the water surface mainly relies on two types of methods: traditional image processing-based technology and deep learning-based technology. Traditional methods for detecting floating objects on the water surface usually include techniques such as edge detection, background modeling, and color space conversion. These methods can provide good results under good lighting conditions. However, in low-light environments, especially at night or in rainy and cloudy weather, the contrast and clarity of the image decrease, making it difficult to distinguish the background from the target. At this time, the recognition accuracy of traditional image processing methods drops significantly and cannot cope with complex environmental changes.

[0005] The deep learning-based method for detecting floating objects on the water surface extracts features and recognizes targets by training a convolutional neural network (CNN). After being trained on a large-scale dataset, this type of method can achieve good detection results under normal lighting conditions. However, the performance of most current deep learning models in low-light environments is not satisfactory, mainly manifested in the following aspects:

[0006] Performance degradation in low light: Due to insufficient image brightness, floating objects in low-light conditions are prone to blend with the background, making it difficult for the network to accurately extract target features and reducing the recognition accuracy.

[0007] Interference from waves and reflections: Interference factors such as wave reflections and light spots on the water surface seriously affect the stability and accuracy of target detection.

[0008] Lack of real-time performance: The processing speed of existing technologies is relatively slow and difficult to meet the real-time requirements. Especially when multiple cameras need to be processed for multitasking, real-time performance becomes an important technical bottleneck.

[0009] The existing methods for detecting floating objects on the water surface have obvious deficiencies in feature extraction ability in low-light environments. When traditional convolutional neural networks (CNNs) are trained on low-light images, it is difficult to fully capture the effective features in the images. Especially at night or under complex weather conditions, the features of floating objects are easily lost or interfered by background noise. Existing methods generally lack targeted optimization for complex lighting and dynamic backgrounds (such as waves and reflections) in the water surface environment. This results in unstable performance of existing detection models in practical applications, especially in environments such as the ocean or lakes where lighting changes frequently and background interference is large. Most existing low-light image enhancement methods are limited to simple image brightness adjustment and contrast enhancement and lack an adaptive enhancement mechanism suitable for complex water surface environments. The current technology fails to fully utilize the attention mechanism and multi-scale feature fusion in deep learning methods, resulting in low target localization accuracy and inability to meet the need for accurate recognition of floating objects.

[0010] These deficiencies limit the popularization and application of surface floating object detection technology in practical scenarios. Therefore, there is an urgent need to introduce more advanced technologies, especially optimize image enhancement, feature extraction, and multi-task processing capabilities under low-light conditions, to address the bottlenecks of existing technologies. Summary of the Invention

[0011] In view of the deficiencies of existing technologies, such as low accuracy in identifying surface floating objects in low-light environments, insufficient robustness to complex light changes and wave reflections on the water surface, and high missed detection rates for small floating objects, the present invention proposes a method for identifying surface floating objects in low-light environments. The method includes the following steps:

[0012] S1: Low-light image acquisition: Use a drone or other image acquisition device to obtain the original image of the water surface under low-light conditions, and use the original image as the input for subsequent processing;

[0013] S2: Low-light image enhancement: Input the original image obtained in step S1 into the improved Retinex-Net model for brightness and contrast enhancement; wherein, a CBAM attention mechanism with multiple layers of convolution and pyramid-shaped kernel scales is embedded in the Retinex-Net model, and through a three-layer convolution structure of the channel attention and spatial attention modules, an adaptive enhanced output image with multi-scale features is obtained;

[0014] S3: Multi-scale feature extraction: Input the adaptively enhanced image output in step S2 into the Backbone network of the Yolov7 model;

[0015] Among them, the Backbone network extracts features from the image and outputs multi-scale feature maps B2, B3, B4, and B5, which correspond to feature information of different scales respectively;

[0016] Among them, the multi-scale feature maps B2, B3, B4, and B5 are sequentially input into the Gold-YoLo feature fusion module, and a shallow feature extraction branch, a middle-layer feature fusion branch, and a deep-layer feature optimization branch are used to... Combined multi-scale interaction mechanism and fuse and strengthen the features through an adaptive weighting mechanism to obtain Multi-scale feature maps after feature enhancement;

[0017] S4: Anchor box optimization: Before training the Yolov7 model, use the Kmeans++ algorithm to perform clustering analysis on the labeled boxes of surface floating objects in the training set to obtain the optimal anchor box size; apply the optimal anchor box size to the multi-scale feature maps in step S3 to make the improved Yolov7 model more suitable for detecting floating objects of different sizes;

[0018] S5: Replacement of the MPDIoU loss function: During the training of the Yolov7 model, replace the original CIoU loss function with the MPDIoU loss function, and measure the similarity between the predicted bounding box and the ground truth annotation box based on the minimum point distance; by minimizing the MPDIoU loss, reduce the positioning error in complex lighting and wave environments.

[0019] S6: Model training: Use the enhanced images obtained in step S2 as the model input, combine the optimal anchor boxes obtained in step S4 and the MPDIoU loss function in step S5 to perform end-to-end training on the Yolov7 model integrated with the Gold-YoLo feature fusion module.

[0020] S7: Detection of floating objects on the water surface: Input the low-light water surface image to be detected into the CBAM attention mechanism in step S2 for enhancement processing, and then input it into the trained improved Yolov7 model.

[0021] S8: Result output: Use the detection results in step S7 to label the floating object targets on the image, and output their position information, confidence level, and target category information to achieve accurate identification and positioning of floating objects on the water surface in low-light environments.

[0022] As a preferred technical solution of the present invention, the CBAM attention mechanism based on multi-layer convolution introduced in the Retinex-Net model in step S2 combines channel attention and spatial attention to perform multi-scale feature mining and adaptive enhancement on low-light images; specifically, it includes the following steps:

[0023] S2-1: Decomposition of low-light images by Retinex-Net. Consider the input low-light image I as the pixel-by-pixel product of the reflection component R and the illumination component L, and its calculation formula is:

[0024] I(x) = R(x) ⊙ L(x)

[0025] where I(x), R(x), and L(x) respectively represent the image intensity, reflection component, and illumination component at pixel x, and ⊙ represents element-wise multiplication.

[0026] S2-2: Embedding of the multi-layer convolution CBAM attention mechanism. During the feature extraction process of Retinex-Net, for the obtained intermediate feature map jointly model channel attention and spatial attention through multi-layer convolution CBAM; among them, the spatial attention part uses pyramid-shaped multi-core convolution, and its calculation formula is:

[0027] M s (F) = σ(Conv 3×3 (F) + Conv 5×5(F)+Conv 7×7 (F))

[0028] Among them, Conv k×k (·) represents a convolution operation using a k×k convolution kernel; σ is an activation function; C×H×W represents the size of the feature map; the obtained spatial attention weight and channel attention weight are combined and applied to the intermediate feature F to complete the adaptive weighting of the feature in the spatial dimension and channel dimension;

[0029] S2-3: Enhance the image output by capturing and weighting the low-light features through the above-mentioned multi-layer convolutional CBAM.

[0030] As a preferred technical solution of the present invention, the structure of the Gold-YoLo feature fusion module described in step S3 is specifically as follows:

[0031] The multi-scale feature maps B2, B3, B4, and B5 extracted by multiple ELAN structures in the backbone network Backbone are respectively input into the Gold-YoLo feature fusion module, and the B2, B3, B4, and B5 respectively correspond to feature outputs with different resolutions and channel depths;

[0032] The Gold-YoLo feature fusion module is internally provided with multiple branches for multi-scale feature interaction and adaptive weight allocation of the input features; its core includes: a shallow feature adaptation branch: for the levels with higher resolutions of B2 and B3 features, it can suppress redundant features while maintaining texture and edge information; a deep feature adaptation branch: for the levels with lower resolutions and larger receptive fields of B4 and B5 features, it introduces more global semantic information and reduces information loss through cross-level channel fusion;

[0033] A cross-scale information injection mechanism is introduced in each branch to enable two-way interaction between low-level and high-level features, retaining both high-resolution details and using high-semantic information for feature correction;

[0034] Finally, the fused multi-scale feature map is subjected to adaptive weighting processing, and the enhanced B3, B4, and B5 features are output.

[0035] As a preferred technical solution of the present invention, the calculation method of the MPDIoU loss function described in step S5 is:

[0036] Given two bounding boxes A and B, which respectively represent the predicted box and the true annotation box of the model, in the form of convex shapes, the calculation formulas are respectively:

[0037]

[0038] Among them, and respectively represent the upper left and lower right coordinates of the bounding box A, and are the upper left and lower right coordinates of the bounding box B;

[0039] By calculating the intersection and union areas of the bounding boxes and combining the normalization processing of the distance, the MPDIoU value is obtained, and the calculation formula is:

[0040]

[0041] where A∩B represents the intersection area of the bounding boxes A and B, and A∪B represents their union area, and are the distances between the endpoints of the bounding box, and w and h are the width and height of the image, which are used for normalization calculation.

[0042] As a preferred technical solution of the present invention, the replacement of the MPDIoU loss function in step S5 specifically includes the following steps:

[0043] S5-1: Bounding box prediction and annotation, set the predicted bounding box B prd =[x prd ,y prd ,w prd ,h prd T and the true annotated bounding box B gt =[x gt ,y gt ,w gt ,h gt T , where x and y respectively represent the center coordinates of the bounding box, and w and h represent the width and height of the bounding box; B gt is the set of true annotated bounding boxes, and B prd is the bounding box obtained by model prediction;

[0044] S5-2: Loss function optimization, during the training process, the loss function optimizes the model parameters Θ by minimizing the following objective function, and its calculation formula is:

[0045]

[0046] where B gt represents the set of all true annotated bounding boxes, is the loss function used to calculate the regression error of each bounding box;

[0047] S5-3: The loss function adopts a calculation method based on MPDIoU, and the specific calculation formula is:

[0048] ​​

[0049] Among them, MPDIoU is a bounding box similarity metric based on the minimum point distance.

[0050] As a preferred technical solution of the present invention, the Yolov7 model described in steps S3 - S6 obtains the final network architecture by introducing the CBAM attention mechanism and the residual network structure.

[0051] Compared with the related prior art, the beneficial effects of the present invention are:

[0052] Improve the detection accuracy of floating objects in low - light environments: By improving the traditional Yolov7 network and introducing the Gold - YoLo feature fusion module and the multi - scale feature interaction mechanism, the present invention effectively improves the ability to extract the features of floating objects on the water surface. Under low - light conditions, the model can more accurately capture the key information of floating objects, thus greatly reducing the missed detection rate, improving the recognition accuracy of small floating objects, and significantly enhancing the overall detection accuracy.

[0053] Optimize the image enhancement effect and solve the problems of over - enhancement and detail loss: The existing Retinex algorithm is prone to over - enhancement and detail loss when processing low - light images. To address this defect, the CBAM - Enhanced Retinex - Net proposed by the present invention effectively enhances the brightness and contrast in low - light images by introducing the CBAM attention mechanism and the adaptive enhancement strategy, while avoiding the loss of details, thereby improving the quality of the images and providing a clearer and more reliable input for subsequent floating object detection.

[0054] Improve the detection accuracy of small floating objects: The Kmeans++ algorithm is used to optimize the anchor boxes, and the anchor box sizes more suitable for small floating objects are selected, solving the problem of insufficient accuracy in detecting small floating objects by traditional methods. Through this optimization, the present invention significantly improves the recognition ability of floating objects of different sizes. Especially when detecting small floating objects on the water surface, the model performs more precisely.

[0055] Enhance the stability and positioning accuracy of the model: By introducing the MPDIoU loss function, the accuracy of bounding box regression is further improved. This loss function is based on the minimum point distance metric, which can effectively solve the positioning error in complex wave backgrounds and enhance the stability of the model in complex water surface environments. This improvement ensures that the model can still maintain high accuracy and robustness under changing environmental conditions and adapt to the floating object recognition tasks in different water surface environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flowchart of a method for recognizing floating objects on the water surface in a low - light environment according to the present invention;

[0057] Figure 2 It is the optimized Yolov7 model structure diagram of the embodiment provided by the present invention;

[0058] Figure 3 It is the neck integrated GoldYoLo structure diagram of the embodiment provided by the present invention;

[0059] Figure 4 It is the schematic diagram of the improved CBAM attention mechanism of the embodiment provided by the present invention;

[0060] Figure 5 It is the schematic diagram of the Retinex-Net network of the embodiment provided by the present invention;

[0061] Figure 6 It is the comparison diagram of the improved Yolov7 results of the embodiment provided by the present invention;

[0062] Figure 7 It is the result diagram of CBAM-Enhanced Retinex-Net of the embodiment provided by the present invention;

[0063] Figure 8 It is the comparison diagram of recognition before and after enhancement of the embodiment provided by the present invention;

[0064] Figure 9 It is the comparison diagram of recognition effects of the embodiment provided by the present invention. Specific implementation manners

[0065] The following further illustrates the present invention in conjunction with the accompanying drawings and embodiments. However, the present invention can be implemented in many different ways and should not be construed as limited to the illustrated embodiments; on the contrary, these embodiments provide implementation manners that meet the applicable legal requirements for those skilled in the art.

[0066] Embodiment 1: As shown in Figure 1 The following, this embodiment provides a specific implementation process of a method for identifying floating objects on the water surface in a low-light environment. The steps are as follows:

[0067] S1: Low-light image acquisition: Use a drone or other image acquisition device to obtain the original image of the water surface under low-light conditions, and use the original image as the input for subsequent processing;

[0068] S2: Low-light image enhancement: Input the original image obtained in step S1 into the improved Retinex-Net model for brightness and contrast enhancement; wherein, a CBAM attention mechanism with multiple layers of convolution and pyramid-shaped kernel scales is embedded in the Retinex-Net model, and through the three-layer convolution structure of the channel attention and spatial attention modules, an adaptive enhanced output image with multi-scale features is obtained;

[0069] S3: Multi-scale Feature Extraction: Input the adaptively enhanced image output in step S2 into the Backbone network of the Yolov7 model;

[0070] Among them, the Backbone network extracts features from the image and outputs multi-scale feature maps B2, B3, B4, and B5, corresponding to feature information of different scales respectively;

[0071] Among them, input the multi-scale feature maps B2, B3, B4, and B5 into the Gold-YoLo feature fusion module in sequence, adopt a multi-scale interaction mechanism that combines a shallow feature extraction branch, a middle-layer feature fusion branch, and a deep-layer feature optimization branch, and fuse and strengthen the features through an adaptive weighting mechanism to obtain a multi-scale feature map with enhanced features;

[0072] S4: Anchor Box Optimization: Before training the Yolov7 model, use the Kmeans++ algorithm to perform clustering analysis on the annotation boxes of water surface floating objects in the training set to obtain the optimal anchor box size; Apply the optimal anchor box size to the multi-scale feature maps in step S3 to make the improved Yolov7 model more suitable for detecting floating objects of different sizes;

[0073] S5: MPDIoU Loss Function Replacement: When training the Yolov7 model, replace the original CIoU loss function with the MPDIoU loss function, and measure the similarity between the predicted bounding box and the true annotation box based on the minimum point distance; By minimizing the MPDIoU loss, reduce the positioning error in complex lighting and wave environments;

[0074] S6: Model Training: Use the enhanced image obtained in step S2 as the model input, combine the optimal anchor box obtained in step S4 and the MPDIoU loss function in step S5 to perform end-to-end training on the Yolov7 model integrated with the Gold-YoLo feature fusion module;

[0075] S7: Water Surface Floating Object Detection: Input the low-light water surface image to be detected into the CBAM attention mechanism in step S2 for enhancement processing, and then input it into the trained improved Yolov7 model;

[0076] S8: Result Output: Use the detection result in step S7 to label the floating object target on the image, and output its position information, confidence level, and target category information to achieve accurate identification and positioning of water surface floating objects in low-light environments.

[0077] Such as Figure 2As shown, the core objective of this method is to improve the recognition performance of UAV images of floating objects on water under low-light conditions. This is mainly achieved by improving the Retinex-Net low-light image enhancement model through the CBAM attention mechanism, and optimizing the Yolov7 model with the Gold-YoLo feature fusion module, MPDIoU loss function optimization, and Kmeans++ anchor box selection. The entire object detection framework uses a variety of advanced technical means to work together, from low-light image processing to feature extraction, and then to object detection, and finally to recognition and output. All modules are organically combined to ensure higher accuracy and robustness in low-light environments.

[0078] Among them, the YOLOv7 model: As a classic algorithm in object detection tasks, YOLOv7 undertakes the main object recognition task in this system. This network has been modified and optimized, incorporating multiple enhancement modules, especially in the design of feature fusion and loss functions, aiming to improve the object detection accuracy under low-light conditions.

[0079] The core structure of the Gold-YoLo feature fusion module includes three branches: a shallow feature extraction branch, a middle feature fusion branch, and a deep feature optimization branch. Each branch dynamically adjusts the feature weights through an adaptive weighting mechanism to ensure that effective feature information can be extracted under different lighting conditions.

[0080] MPDIOU Loss function: To address the performance issues of the original CIOU loss function of YOLOv7 in low-light environments, an improved MPDIOU loss function is introduced to optimize the bounding box regression accuracy, thereby improving the detection accuracy.

[0081] Kmeans++ anchor box optimization: The Kmeans++ algorithm is used for anchor box optimization to select the most suitable anchor box size for detecting low-light floating objects, thereby enhancing the model's detection ability for floating objects.

[0082] CBAM-Enhanced Retinex-Net: Integrate the CBAM (Convolutional Block Attention Module) attention mechanism into Retinex-Net. By enhancing the brightness and contrast of the image, the image fill light effect under low-light conditions is improved, making the floating objects in the image more obvious and the detection more accurate.

[0083] Such as Figure 3As shown in the figure, the outputs of different sizes obtained by performing "ELAN" four times on the backbone network Backbone are respectively used as B2 (160*160*256), B3 (80*80*512), B4 (40*40*1024), and B5 (20*20*1024) and input into GoldYoLo. The GoldYoLo module can perform multi-scale feature interaction and adaptive weight allocation on the input features, and finally output B3 (80*80*512), B4 (40*40*1024), and B5 (20*20*1024) with enhanced features. Through GoldYoLo, the feature extraction ability of the model for floating objects on the water surface can be significantly improved.

[0084] The calculation method of the MPDIoU loss function is as follows: Given two bounding boxes A and B, which respectively represent the predicted box and the true annotation box of the model, in the form of convex shapes, the calculation formulas are respectively:

[0085]

[0086] Among them, and respectively represent the upper left and lower right coordinates of the bounding box A, and are the upper left and lower right coordinates of the bounding box B;

[0087] By calculating the intersection and union areas of the bounding boxes and combining the normalization processing of the distance, the MPDIoU value is obtained, and the calculation formula is:

[0088]

[0089] Among them, A∩B represents the intersection area of the bounding boxes A and B, and A∪B represents their union area, and are the distances between the endpoints of the bounding boxes, and w and h are the width and height of the image, which are used for normalization calculation.

[0090] As Figure 4 shown, the traditional CBAM attention mechanism only uses single-layer convolution in channel attention and spatial attention. Such an attention mechanism cannot effectively extract and remember the features of low-light images. Therefore, the present invention proposes to change the single-layer convolution of the channel attention part and spatial attention to three-layer convolution, and the kernel_sizes (field of view kernels) of the convolutional layers increase sequentially from 3, 5, 7 to form a pyramid-shaped convolutional layer, so that the Retinex-Net network can obtain higher and broader image features when using the CBAM attention mechanism.

[0091] As Figure 5As shown, the final network structure of the present invention is obtained by introducing the CBAM attention mechanism and the residual network based on the traditional Retinex-Net network.

[0092] Example 2: This example proposes a method for detecting floating objects on water surface under low light based on deep learning. By introducing an improved Yolov7 model, a Gold-YoLo feature fusion module, a CBAM-Enhanced Retinex-Net image enhancement network, and an MPDIoU loss function, the detection accuracy and robustness of floating objects on the water surface under low light are significantly improved. Through experimental verification, the method of the present invention is superior to the traditional method under low light and complex water surface conditions, and the specific effects are as follows:

[0093] Figure 6 The detection effects of small floating objects before and after optimization are shown. After optimizing the anchor boxes through the Kmeans++ algorithm, the model can better adapt to the size of small floating objects, significantly improving the accuracy of small target detection. The optimized detection boxes more closely surround the target, reducing the cases of missed detection and false detection. Especially in the case of low light and complex background, the optimized model can more accurately identify floating objects. From left to right are the original Yolov7, the model optimized by GoldYoLo, and the model combined with the MPDIoU loss function. It can be seen that in the original Yolov7 model, the missed detection rate is relatively high under low light conditions, especially the recognition of small floating objects is poor. In contrast, the GoldYoLo module enhances the feature extraction ability of the model for floating objects, especially the recognition of complex water surface backgrounds, and the floating objects are more accurately labeled. After adding the MPDIoU loss function, the regression accuracy of the bounding box is further improved, and the model can more stably detect floating objects in a complex wave environment and accurately label their positions.

[0094] Figure 7 and Figure 8 The image enhancement effect of the present invention under low light conditions is shown. The low light image input is on the far left, and the image enhanced by CBAM-Enhanced Retinex-Net is on the far right. By introducing the CBAM attention mechanism and the adaptive enhancement strategy, the enhanced image has been significantly improved in terms of brightness, contrast, and details. The boundaries of floating objects are clearer, and background noise is effectively suppressed. Especially under low light and complex lighting conditions, the enhanced image is significantly better than the output of the traditional Retinex-Net, making the subsequent target detection more accurate.

[0095] Figure 9Shows the comparison of positioning accuracy after introducing the MPDIoU loss function. The traditional Yolov7 model has a missed detection rate as high as 37.8% under low-light conditions, and the recognition accuracy for small floating objects is less than 50%. Therefore, in the present invention, the neck network of the traditional Yolov7 is replaced with a Gold-YoLo feature fusion module, which improves the feature extraction ability of Yolov7 by introducing a multi-scale feature interaction mechanism. In terms of image enhancement, the existing Retinex algorithm has problems of over-enhancement and detail loss when dealing with complex water surface illumination. The CBAM-Enhanced Retinex-Net network proposed in the present invention introduces a CBAM attention mechanism and an adaptive enhancement strategy. This study also optimizes the model training strategy to improve the enhancement effect of low-light images. The present invention uses the Kmeans++ algorithm to optimize the anchor boxes, improving the recognition accuracy for small floating objects. Secondly, the MPDIOU loss function is also introduced to improve the positioning accuracy of the model, enhancing the stability of the model in complex wave environments. The above improvements make the present invention have excellent performance, providing reliable technical support for the real-time monitoring and accurate recognition of water surface floating objects.

[0096] The experimental results show that after adopting the method of the present invention, the recognition accuracy in low-light environments has been greatly improved, the missed detection rate has dropped from 37.8% to a lower level, and at the same time, the recognition accuracy for small floating objects has increased from 50% to over 75%. Compared with the traditional Yolov7 model, the improved model shows more stable and accurate detection capabilities in the case of complex water surface backgrounds, low light, and wave interference, and can effectively cope with various challenges, providing reliable technical support for the real-time monitoring and accurate recognition of water surface floating objects.

[0097] Through these experimental results, it is proved that the present invention has significant advantages in the detection of water surface floating objects in low-light environments, especially in terms of target detection accuracy, model stability, and adaptability to complex backgrounds, showing more excellent performance compared with traditional methods.

[0098] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for identifying floating objects on the water surface in a low-light environment, characterized in that: The method comprises the following steps: S1: Low-light image acquisition: Use a drone or other image acquisition equipment to obtain the original image of the water surface under low-light conditions, and use the original image as input for subsequent processing; S2: Low-light image enhancement: The original image obtained in step S1 is input into the improved Retinex-Net model for brightness and contrast enhancement; wherein the Retinex-Net model is embedded with a multi-layer convolution and pyramid-type kernel-scale CBAM attention mechanism, and a three-layer convolution structure of channel attention and spatial attention modules is used to obtain an adaptively enhanced output image with multi-scale features; S3: Multi-scale feature extraction: The adaptive enhanced image output in step S2 is input into the Backbone network of the Yolov7 model; The Backbone network extracts features from the image and outputs multi-scale feature maps B2, B3, B4, and B5, which correspond to feature information of different scales respectively; Among them, the multi-scale feature maps B2, B3, B4, and B5 are sequentially input into the Gold-YoLo feature fusion module, and a multi-scale interaction mechanism combining shallow feature extraction branches, middle-level feature fusion branches, and deep feature optimization branches is adopted. The features are fused and enhanced through an adaptive weighting mechanism to obtain a multi-scale feature map after feature enhancement. S4: Anchor frame optimization: Before training the Yolov7 model, the Kmeans++ algorithm is used to perform cluster analysis on the water surface floating object annotation frames in the training set to obtain the optimal anchor frame size; the optimal anchor frame size is applied to the multi-scale feature map in step S3, so that the improved Yolov7 model is more suitable for detecting floating objects of different sizes; S5: Replacement of MPDIoU loss function: When training the Yolov7 model, the original CIoU loss function is replaced by the MPDIoU loss function, and the similarity between the predicted bounding box and the true annotation box is measured based on the minimum point distance; by minimizing the MPDIoU loss, the positioning error in complex lighting and wave environments is reduced; S6: Model training: Using the enhanced image obtained in step S2 as the model input, combined with the optimal anchor box obtained in step S4 and the MPDIoU loss function in step S5, the Yolov7 model with integrated Gold-YoLo feature fusion module is trained end-to-end; S7: Detection of floating objects on the water surface: The low-light water surface image to be detected is input into the CBAM attention mechanism of step S2 for enhancement processing, and then input into the trained improved Yolov7 model; S8: Result output: Use the detection result of step S7 to mark the floating object target on the image, output its position information, confidence level and target category information, and realize accurate recognition and positioning of floating objects on the water surface in low light environment.

2. The method for identifying floating objects on the water surface in a low-light environment according to claim 1, characterized in that: The multi-layer convolution-based CBAM attention mechanism introduced in the Retinex-Net model described in step S2 combines channel attention with spatial attention to perform multi-scale feature mining and adaptive enhancement on low-light images; specifically, the following steps are included: S2-1: Retinex-Net’s low-light image decomposition, the input low-light image I is regarded as the pixel-by-pixel product of the reflection component R and the illumination component L, and its calculation formula is: I(x)=R(x)⊙L(x) Among them, I(x), R(x), L(x) represent the image intensity, reflection component and illumination component at pixel x respectively, and ⊙ represents element-wise multiplication; S2-2: Multi-layer convolution CBAM attention mechanism embedding, Retinex-Net feature extraction process, for the obtained intermediate feature map The channel attention and spatial attention are jointly modeled through multi-layer convolution CBAM. Among them, the spatial attention part adopts pyramid multi-kernel convolution, and its calculation formula is: M s (F)=σ(Conv 3×3 (F)+Conv 5×5 (F)+Conv 7×7 (F)) Among them, Conv k×k (·) represents the convolution operation with a k×k convolution kernel; σ is the activation function; C×H×W represents the size of the feature map; the obtained spatial attention weight is combined with the channel attention weight and applied to the intermediate feature F to complete the adaptive weighting of the features in the spatial dimension and channel dimension; S2-3: Enhanced image output, capturing and weighting low-light features through the above multi-layer convolutional CBAM.

3. The method for identifying floating objects on the water surface in a low-light environment according to claim 1, characterized in that: The structure of the Gold-YoLo feature fusion module in step S3 is as follows: The multi-scale feature maps B2, B3, B4, and B5 extracted by multiple ELAN structures in the backbone network are respectively input into the Gold-YoLo feature fusion module. B2, B3, B4, and B5 correspond to feature outputs of different resolutions and channel depths respectively. The Gold-YoLo feature fusion module has multiple branches set inside for multi-scale feature interaction and adaptive weight allocation of input features; Its core includes: shallow feature adaptation branch: for layers with higher feature resolutions such as B2 and B3, it can suppress redundant features while maintaining texture and edge information; deep feature adaptation branch: for layers with lower feature resolutions and larger receptive fields such as B4 and B5, it introduces more global semantic information and reduces information loss through cross-level channel fusion; A cross-scale information injection mechanism is introduced in each branch to enable bidirectional interaction between low-level and high-level features, which not only retains high-resolution details but also uses high-semantic information for feature correction; Finally, the fused multi-scale feature maps are adaptively weighted to output feature-enhanced B3, B4, and B5.

4. The method for identifying floating objects on the water surface in a low-light environment according to claim 1, characterized in that: The MPDIoU loss function in step S5 is calculated as follows: Given two bounding boxes A and B, representing the model's predicted box and the true annotation box respectively, in the form of convex shapes, the calculation formulas are: in, and Respectively represent the coordinates of the upper left corner and lower right corner of the bounding box A, and are the coordinates of the upper left and lower right corners of the bounding box B; By calculating the intersection and union areas of the bounding boxes and combining the normalized distance, the MPDIoU value is obtained. The calculation formula is: Among them, A∩B represents the intersection area of ​​bounding boxes A and B, and A∪B represents their union area. and is the distance between the endpoints of the bounding box, w and h are the width and height of the image, which are used for normalization calculation.

5. The method for identifying floating objects on the water surface in a low-light environment according to claim 1, characterized in that: The MPDIoU loss function replacement described in step S5 specifically includes the following steps: S5-1: Bounding box prediction and annotation, setting the predicted bounding box B prd =[x prd ,y prd ,w prd ,h prd ] T And the ground-truth bounding box B gt =[x gt ,y gt ,w gt ,h gt ] T , where x and y represent the center coordinates of the bounding box, w and h represent the width and height of the bounding box respectively; B gt is the set of bounding boxes of true annotations, B prd is the bounding box predicted by the model; S5-2: Loss function optimization. During the training process, the loss function optimizes the model parameter Θ by minimizing the following objective function, which is calculated as follows: Among them, B gt represents the set of all ground-truth annotated bounding boxes, is the loss function used to calculate the regression error for each bounding box; S5-3: The loss function adopts the calculation method based on MPDIoU. The specific calculation formula is: Among them, MPDIoU is a bounding box similarity measure based on the minimum point distance.

6. The method for identifying floating objects on the water surface in a low-light environment according to claim 1, characterized in that: The Yolov7 model described in steps S3-S6 obtains the final network architecture by introducing the CBAM attention mechanism and residual network structure.

Citation Information

Patent Citations

  • Water surface floating object multi-camera real-time detection method based on SSD network

    CN113469097A

  • Water surface floating object detection method based on deep learning

    CN114170549A

Cited By

  • Intelligent identification method and system for line diagram of power distribution network

    CN120496118A

  • A method and system for intelligent identification of distribution network circuit diagram

    CN120496118B