Water surface floating object detection method based on improved YOLOv9

By introducing the BiFormer dynamic sparse attention mechanism, BoTNet hybrid architecture and CA attention mechanism into the YOLOv9 algorithm, the problems of insufficient backbone network modeling and computational redundancy in surface target detection are solved, and the accuracy and robustness of small target detection are improved.

CN120726438APending Publication Date: 2025-09-30FUJIAN ZHONGRUI HANDING DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510889866.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing surface target detection algorithms in complex water environments have problems such as insufficient global feature modeling capabilities of the backbone network, computational redundancy of traditional convolutional modules, and weak suppression of complex background interference by the attention mechanism, resulting in insufficient accuracy and robustness in small target detection.

Method used

A feature extraction module based on the BiFormer dynamic sparse attention mechanism is introduced, combined with the BoTNet hybrid architecture and the CA attention mechanism. Through double-layer routing attention, multi-head self-attention and spatial-channel dual-branch interaction, feature extraction is optimized and noise interference is suppressed, thereby enhancing the feature response of small target areas.

Benefits of technology

It significantly improves the accuracy of small surface target detection and adaptability to complex environments, improves the model's ability to grasp the characteristics of targets of different scales, reduces the amount of calculation and maintains high efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726438A_ABST
    Figure CN120726438A_ABST
Patent Text Reader

Abstract

The invention relates to a water surface floating object detection method based on improved YOLOv9, and belongs to the technical field of computer vision and target detection. The method comprises the steps that a feature extraction module based on a BiFormer dynamic sparse attention mechanism is introduced, and through double-layer routing attention and pyramid structure design, the retention capacity of shallow small target features is improved while the calculated amount is reduced; according to the method, a hybrid architecture fused with BoTNet is constructed, the BoTNet is embedded in a backbone network, and by means of collaborative optimization of a multi-head self-attention mechanism and depth separable convolution, the model can more accurately grasp features of targets of different scales; an anti-interference strategy guided by a CA attention mechanism is introduced, noise interference caused by water surface wave reflection is suppressed through space-channel double-branch cross-dimension interaction and a dynamic activation mechanism sensitive to coordinates, and the feature response intensity of a small target area is enhanced. According to the invention, the detection performance and complex environment adaptability of the model to small targets on the water surface can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and target detection, and particularly relates to a method for detecting floating objects on a water surface based on an improved YOLOv9. Background Art

[0002] Due to the complex physical characteristics of water scenes, such images are generally subject to challenges such as non-uniform lighting conditions, dynamic noise interference, variable climate conditions, and complex background textures, which makes the task of detecting targets on the water surface unique[1]. Target detection in water scenes has always been a challenging task in the field of computer vision, and it is difficult to simultaneously ensure both detection efficiency and recognition accuracy. Currently, research methods in this field are mainly divided into two categories: traditional methods[2] and deep learning-based algorithms[3].

[0003] Traditional methods rely heavily on feature construction and classifiers to achieve target classification. However, in complex aquatic environments, these methods have significant drawbacks. Affected by environmental factors such as surface reflections and wave interference, the robustness of features decreases significantly, leading to severe degradation of classifier performance and making it difficult to meet the real-time and accuracy requirements of practical engineering applications.

[0004] Traditional target classification methods typically rely on manually designed feature engineering, combined with classifiers such as support vector machines (SVMs) and random forests to achieve target recognition. However, in complex aquatic scenarios, these methods exhibit significant limitations: environmental interference such as bright noise caused by dynamic reflections on the water surface, target morphological distortion caused by waves, and color shift caused by illumination changes can severely weaken the stability and generalization capabilities of features, leading to problems such as feature mismatching and blurred decision boundaries in the classifier.

[0005] In response to the difficulty of multi-scale target detection, researchers have proposed a variety of algorithms based on deep learning for improvement. Zhang et al. [4] constructed an improved FasterR-CNN framework that fuses shallow and deep features to enhance the robustness of the model to complex water environments. Liu Wei et al. [5] constructed a multi-feature extraction network based on MaskR-CNN to achieve simultaneous improvement in average accuracy and recall rate in the floating object recognition task. Li Guojin et al. [6] optimized the positioning accuracy of YOLOv3 for surface targets by improving the K-means clustering and category activation mapping mechanism. Wang Lin et al. [7] used deep separable convolution and adaptive feature fusion modules to achieve high-precision real-time detection of marine ships under the YOLOv4 framework.

[0006] Existing methods generally overlook the problem of detecting small objects in water scenes. Small objects have distinct characteristics: low resolution leading to loss of detail, insufficient pixel coverage leading to feature blurring, high positioning error rates, and a lack of semantic information. In complex water environments, these inherent limitations are susceptible to the cumulative effects of dynamic interference, leading to frequent missed detections in real-world scenarios.

[0007] Current deep learning methods still face significant technical bottlenecks in solving low-contrast target detection, multi-scale deformation adaptation, and small target feature enhancement. Specifically, the low-contrast features between the target and the background lead to feature extraction failure, the high variability of shape scale exacerbates the difficulty of model generalization, and the weak feature expression ability of small targets directly affects detection accuracy.

[0008] In the field of surface target detection, researchers have conducted innovative research targeting specific needs. The YOLO series of models has been continuously iteratively optimized. YOLOv8[8] significantly improves multi-scale target detection capabilities while maintaining high efficiency through lightweight architecture and improved training strategies. YOLOv8 adopts a modular network architecture consisting of four parts: input module, backbone network, feature fusion neck, and detection head. The input module is responsible for image preprocessing and improves the generalization ability of the model through dynamic size scaling and data enhancement technology. The backbone network uses the C2f module to replace the traditional C3 structure, optimizes feature propagation efficiency through the gradient diversion mechanism, and enhances the expression of deep semantic features while maintaining lightweight characteristics.

[0009] The feature fusion process constructs a dual-stream FPN-PAN pyramid network, enabling multi-scale information exchange through a cross-level feature fusion path. Specifically, the top-down feature pyramid network (FPN) is responsible for conveying high-level semantic information, while the bottom-up path aggregation network (PAN) enhances low-level detailed features, forming a bidirectional feature flow mechanism. This design effectively solves the problem of small object information loss caused by traditional single-path feature fusion.

[0010] YOLOv9[9] introduces an adaptive convolution module and a lightweight attention mechanism to improve the performance of the model without increasing the amount of computation. YOLOv9 is an iterative version of the YOLO series algorithm proposed in February 2024. While maintaining real-time detection capabilities, the algorithm uses direct regression through reversible gradient path design, feature reuse mechanism and lightweight architecture optimization, significantly improving the detection accuracy and generalization ability of the model in complex scenarios.

[0011] YOLOv9 addresses the information bottleneck problem in deep neural networks by proposing a programmable gradient information mechanism. By establishing an explicit mapping between the objective function and the input data, it ensures the integrity of gradient information during backpropagation. This mechanism effectively mitigates the conflict between information decay and reversible function constraints in deep networks by dynamically adjusting feature propagation paths. Furthermore, a generalized efficient layer aggregation network is designed, achieving an optimal balance between model lightweighting and feature utilization through a gradient path planning strategy.

[0012] As a new-generation detection framework, YOLOv9 has improved model efficiency through reversible gradient path design and feature reuse mechanism, but its native network structure still faces challenges in the water surface floating object detection scenario: First, the backbone network lacks the ability to model the global features of the dynamic water surface environment, making it difficult to capture the associated features of distant floating targets; second, the traditional convolution module has computational redundancy in the shallow feature extraction stage, which limits the model's inference speed on edge devices; finally, the existing attention mechanism has weak suppression ability for complex background interference, especially in waters with strong reflectivity, which can easily misactivate irrelevant areas and affect the ability to focus on small target features. Summary of the Invention

[0013] The purpose of the present invention is to provide a surface floating object detection method based on improved YOLOv9 to address the problems existing in the surface target detection algorithm in the detection of surface floating objects, such as insufficient global feature modeling capabilities of the backbone network, computational redundancy of traditional convolution modules, and weak suppression of complex background interference by the attention mechanism. Through lightweight feature extraction optimization, global context modeling enhancement and dynamic attention weight calibration, the model's detection performance for small surface targets and adaptability to complex environments are significantly improved.

[0014] To achieve the above objectives, the technical solution of the present invention is: a method for detecting floating objects on a water surface based on an improved YOLOv9, comprising:

[0015] A feature extraction module based on the BiFormer dynamic sparse attention mechanism is introduced. Through dual-layer routing attention and pyramid structure design, it reduces the amount of computation while improving the ability to retain shallow small object features.

[0016] A hybrid architecture integrating BoTNet is constructed, embedding BoTNet in the neck network. By leveraging the collaborative optimization of the multi-head self-attention mechanism and depthwise separable convolution, the model can more accurately grasp the features of objects at different scales.

[0017] An anti-interference strategy guided by the CA attention mechanism is introduced. A CA attention mechanism module is introduced at the end of the neck network. Through the spatial-channel dual-branch cross-dimensional interaction and the coordinate-sensitive dynamic activation mechanism, the noise interference caused by the reflection of water waves is suppressed and the feature response intensity of small target areas is enhanced.

[0018] Furthermore, the feature extraction module based on the BiFormer dynamic sparse attention mechanism is introduced, that is, the BiFormer dynamic sparse attention mechanism is introduced at the end of the neck network of YOLOv9. The BiFormer dynamic sparse attention mechanism is based on the deep optimization of the Transformer architecture.

[0019] Furthermore, the BiFormer dynamic sparse attention mechanism proposes a two-layer routing attention mechanism BRA, which realizes the intelligent allocation of computing resources through region-level affinity graph construction and dynamic pruning strategy; the BiFormer dynamic sparse attention mechanism consists of two steps: routing step and token-level attention step; the BiFormer dynamic sparse attention mechanism adopts a four-stage pyramid architecture to realize hierarchical feature extraction.

[0020] Furthermore, BoTNet is a backbone architecture that combines the self-attention mechanism and convolutional neural network. BoTNet is constructed by replacing the last three spatial 3×3 convolutions with MHSA in the last three bottleneck blocks of ResNet.

[0021] Furthermore, BoTNet is embedded after the RepNCSPELAN4 module of the neck network of YOLOv9.

[0022] Furthermore, the CA attention mechanism module achieves accurate encoding of spatial position information through a decomposition-based global average pooling strategy. The decomposition-based global average pooling strategy generates direction-aware position encoding vectors along the horizontal / vertical directions, and fuses the position encoding vectors in these two directions through the coordinate information interaction layer to establish an explicit mapping relationship between spatial position and channel features, thereby generating position-sensitive attention weights in the channel dimension.

[0023] Furthermore, the CA attention mechanism module achieves accurate encoding of spatial position information through a decomposition-based global average pooling strategy. The specific implementation formula is as follows:

[0024] Given x c , use the two spaces of the pool kernel (h,i) or (j,w) to average the output feature map in the horizontal and vertical directions respectively and The formula is as follows:

[0025]

[0026] H and W are the height and width of the input feature map x; c represents the c-th channel in the input feature map x;

[0027] Next, the feature maps of the global receptive field obtained in the horizontal and vertical directions are concatenated, and then a 1×1 convolution is performed. Finally, a nonlinear operation is performed using the Sigmoid activation function, as shown in the following formula:

[0028] f=δ(F1([z h ,z w ]))

[0029] Among them, [z k , z w ] means z h and z w Perform splicing operation, F1(·) represents 1×1 convolution operation, δ(·) represents Sigmoid activation function, that is, normalize and nonlinearize the feature map after convolution, and finally obtain the feature map f∈R C / r×(H+W) It is an intermediate feature containing horizontal and vertical spatial information, and r represents the step size of the pooling;

[0030] Then the feature map f is divided into two parts f h ∈R C / r×H and f w ∈R C / r×W , then adjust the number of channels to f h Adjust the number of channels to be consistent with the input feature map x, as shown in the following formula:

[0031] g h =σ(F h (f h ))

[0032] g w =σ(F w (f w ))

[0033] Among them, F h (·), F w (·) is the use of two 1×1 convolution F h and F w f h and f w The channel adjustment function for feature conversion, σ(·) is the activation function, which is used to adjust the f h and f w Perform linear transformation;

[0034] Then, the final output g is calculated by multiplying the attention weight matrix h and g w , the attention weight matrix calculation method is as follows:

[0035]

[0036] Among them, i and j are the indices in the height and width directions of the traversal feature map, which are used to indicate the position of the pixel in the feature map, specifically described as the input feature x c (i, j) and channel adjusted features and Perform point multiplication to obtain the output feature y c (i,j).

[0037] The present invention also provides a surface floating object detection system based on improved YOLOv9, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the steps of any of the above-described methods can be implemented.

[0038] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the steps of any of the above methods can be implemented.

[0039] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any of the above methods.

[0040] Compared with the existing technology, the present invention has the following beneficial effects: Compared with the existing surface floating object target detection algorithm and the YOLOv9 native algorithm, the contribution of the algorithm of the present invention is: (1) Introducing a feature extraction module based on the BiFormer dynamic sparse attention mechanism. Through the dual-layer routing attention and pyramid structure design, the module greatly improves the ability to retain shallow small target features while reducing the amount of calculation, improving the calculation efficiency without losing key information; (2) Carefully constructing a hybrid architecture integrating BoTNet. The BoTNet module is embedded in the neck network, and with the help of the coordinated optimization of the multi-head self-attention mechanism and the depth-separable convolution, the model can grasp the features of targets of different scales more accurately; (3) Introducing the anti-interference strategy guided by the CA attention mechanism, introducing the CA (coordinate attention mechanism) module at the end of the neck network, through the spatial-channel dual-branch cross-dimensional interaction and the coordinate-sensitive dynamic activation mechanism, effectively suppressing the noise interference caused by the reflection of water waves, and significantly enhancing the feature response intensity of the small target area. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a diagram of the network model architecture of the method of the present invention.

[0042] Figure 2 This is the BiFormer structure diagram.

[0043] Figure 3 It is a multi-head self-attention layer (MHSA) structure.

[0044] Figure 4 This is the module diagram of the CA attention mechanism.

[0045] Figure 5 This is the result graph of the YOLOv9 algorithm.

[0046] Figure 6 This is the result graph of the YOLOv9-BBC algorithm.

[0047] Figure 7 This is the mAP curve of the YOLOv9-BBC algorithm training process on the validation set.

[0048] Figure 8 This is the test result diagram of the YOLOv9 algorithm and the YOLOv9-BBC algorithm. DETAILED DESCRIPTION

[0049] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0050] The present invention provides a method for detecting floating objects on a water surface based on an improved YOLOv9, comprising:

[0051] A feature extraction module based on the BiFormer dynamic sparse attention mechanism is introduced. Through dual-layer routing attention and pyramid structure design, it reduces the amount of computation while improving the ability to retain shallow small object features.

[0052] A hybrid architecture integrating BoTNet is constructed, embedding BoTNet in the backbone network. By leveraging the collaborative optimization of the multi-head self-attention mechanism and depthwise separable convolution, the model can more accurately grasp the features of objects at different scales.

[0053] An anti-interference strategy guided by the CA attention mechanism is introduced. A CA attention mechanism module is introduced at the front end of the detection head. Through the spatial-channel dual-branch cross-dimensional interaction and the coordinate-sensitive dynamic activation mechanism, the noise interference caused by wave reflections on the water surface is suppressed and the feature response intensity of small target areas is enhanced.

[0054] The following is a specific implementation process of the present invention.

[0055] like Figure 1 As shown in FIG, a network model architecture of a method for detecting floating objects on a water surface based on an improved YOLOv9 is provided. The method of the present invention mainly includes:

[0056] 1. Feature extraction module based on BiFormer dynamic sparse attention mechanism

[0057] In the field of target detection, the application of attention mechanism is extremely extensive. In traditional lightweight network design, although the method based on depthwise separable convolution significantly reduces the computational cost by decoupling the spatial and channel dimensions, it still faces the problem of insufficient feature expression ability in the detection scenario of floating objects on the water surface. Although such methods can improve the efficiency of model inference, when dealing with small targets disturbed by waves, their channel compression strategy and simplified convolution operation can easily lead to the loss of shallow detail features. In addition, although improved methods such as dynamic convolution enhance feature diversity through dynamic parameter generation, their additional computational overhead limits the real-time performance of the model on edge devices, making it difficult to meet the stringent efficiency requirements of drone inspection scenarios.

[0058] In view of this, and to address the limitations of YOLOv9 in small target detection, the present invention introduces the BiFormer dynamic sparse attention mechanism

[10] at the end of the neck network. This mechanism is based on the deep optimization of the Transformer architecture and implements intelligent allocation of computing resources through a two-layer routing strategy. By introducing BiFormer, the network can more keenly capture the information of small targets on the water surface, further significantly improve the network's ability in small target detection, greatly enhance the accuracy and efficiency of detection, and better adapt to complex water surface target detection scenarios.

[0059] The core innovation of BiFormer lies in the proposal of a bi-level routing attention mechanism (BRA), which realizes the intelligent allocation of computing resources through the construction of region-level affinity graphs and dynamic pruning strategies.

[0060] BiFormer uses a four-stage pyramid architecture to achieve hierarchical feature extraction. Its core building block is BRA (BiFormer Residual Attention). The overall structure is as follows Figure 2 As shown, input an image, X∈R H*W*C , first divide it into S*S different regions, where each region contains H*W / S 2 eigenvectors. That is, X becomes Then we can obtain as follows:

[0061] Q=X r W q ,K=X r W k ,V=X r W v

[0062] Among them, W q , W k , W v ∈R C*CThey are the projection weights of query, key, and value respectively.

[0063] The BiFormer architecture enhances the model's ability to abstract and represent input data through hierarchical downsampling of spatial dimensions and progressive expansion of channel dimensions.

[0064] In summary, BiFormer achieves efficient modeling of multi-scale target features by integrating the dynamic sparse attention mechanism with the hierarchical feature pyramid. The introduction of BiFormer contributes to the present invention as follows:

[0065] (1) BiFormer uses two-layer routing to achieve more flexible computing power allocation, save computing resources, realize model lightweighting, and effectively alleviate the computing power constraint problem of surface equipment. (2) The core idea of ​​BiFormer combines the advantages of Transformer and pyramid network. Local details are retained through overlapping block embedding, scale abstraction is achieved through the spatial downsampling module, and cross-level semantic associations are captured by the dynamic routing module. This design enables the model to achieve breakthrough progress in small target detection tasks. (3) Spatial attention weights are generated by global average pooling and convolution kernels, combined with dynamic recalibration of channel dimensions, to achieve deep fusion of multimodal features and improve the detection ability of small targets.

[0066] 2. BoTNet Network

[0067] In the detection of floating objects on the water surface, traditional global context modeling methods have difficulty in effectively modeling the spatial correlation of fragmented floating objects when capturing them at a distance. Although such methods enhance local feature extraction capabilities by expanding the receptive field, their fixed sampling patterns cannot adaptively adjust spatial dependencies when faced with non-rigidly deformed targets caused by wave disturbances in dynamic waters. For example, in a scenario where floating plastic bottles diffuse with the water flow, the distribution of fragmented targets is continuous, but traditional global modeling methods lack a joint perception mechanism for target deformation and motion trajectory, resulting in insufficient modeling of contextual correlations for distant targets. In addition, although the Transformer-based method achieves long-distance dependency modeling through a self-attention mechanism, its high computational complexity and sensitivity to the spatial resolution of small targets make the densely distributed floating debris features easily overwhelmed by background noise, further exacerbating the problem of missed detection.

[0068] To overcome these challenges, BoTNet

[11] was introduced, which is a backbone architecture that combines the self-attention mechanism with a convolutional neural network. The design of BoTNet is simple. By replacing the last three spatial 3×3 convolutions in the last three bottleneck blocks of ResNet with MHSA (Multi-Head Self-Attention) layers, it significantly improves the baseline performance and reduces the number of parameters while maintaining low latency.

[0069] MHSA layer such as Figure 3 To significantly improve the YOLOv9 algorithm's ability to model the long-range distribution of floating objects on the water surface, this paper introduces the BoTNet architecture, embedding the Bottleneck Transformer (BoTNet) structure after the RepNCSPELAN4 module of the Neck network. This module builds a hybrid Transformer and CNN architecture. While preserving local details, it leverages the Multi-Head Self-Attention (MHSA) mechanism to accurately capture the global spatial correlation of distant objects. For example, it can keenly perceive the continuous distribution of plastic debris spreading with water flow.

[0070] The introduction of BoTNet improves the recall rate of small target detection in dense water scenarios while maintaining inference speed, significantly outperforming the traditional FPN-PAN architecture and opening up a new technical path for high-precision, real-time detection in complex dynamic environments. Unlike traditional FPN-PAN, which relies on fixed convolution kernels to model local features and struggles to capture the relevance of distant targets, the enhanced BoTNet architecture achieves spatial continuity modeling of fragmented targets across regions by synergizing Transformer global perception with CNN local extraction, while leveraging sparse computing to ensure real-time performance.

[0071] 3. CA Attention Mechanism

[0072] CA (CoordinateAttention)

[12] is a direction-aware lightweight attention mechanism that improves the model's expressiveness by fusing feature information from the channel and spatial dimensions. This mechanism breaks through the limitations of traditional channel attention and achieves accurate encoding of position information through a decomposition-based global pooling strategy while maintaining low computational complexity. Figure 4 shown.

[0073] The CA module achieves precise encoding of spatial position information through a factorized global average pooling strategy. This factorized pooling design breaks the homogenous compression of spatial information imposed by traditional global pooling and can capture the distribution patterns of targets in both the horizontal and vertical dimensions. The two directional encoding vectors are fused through a coordinate information interaction layer, establishing an explicit mapping between spatial position and channel features, thereby generating position-sensitive attention weights in the channel dimension.

[0074] The processed output of the feature information of the cth channel and height h in the input feature map x is calculated as follows:

[0075]

[0076] Similarly, for the cth channel with width w in the feature map, its output is calculated as follows:

[0077]

[0078] Given x c , use the two spaces of the pool kernel (h, i) or (j, w) to average the output feature map in the horizontal and vertical directions respectively and

[0079] Next, the feature maps of the global receptive field obtained in the horizontal and vertical directions are concatenated, and then a 1×1 convolution is performed. Finally, a nonlinear operation is performed using the Sigmoid activation function, as shown in the following formula:

[0080] f=δ(F1([z h ,z w ]))

[0081] Among them, [z k , z w ] means z h and z w Perform splicing operation, F1(·) represents 1×1 convolution operation, δ(·) represents Sigmoid activation function, that is, normalize and nonlinearize the feature map after convolution, and finally obtain the feature map f∈R C / r×(H+W) It is an intermediate feature containing horizontal and vertical spatial information, and r represents the step size of the pooling; this method enhances the network's ability to represent input data.

[0082] Then the feature map f is divided into two parts f h ∈R C / r×H and f w ∈R C / r×W , then adjust the number of channels to f h Adjust the number of channels to be consistent with the input feature map x, as shown in the following formula:

[0083] g h =σ(F h (f h ))

[0084] g w =σ(F w (f w ))

[0085] Among them, F h (·), F w (·) is the use of two 1×1 convolution F h and F w f h and f w The channel adjustment function for feature conversion, σ(·) is the activation function, which is used to adjust the f h and f w Perform linear transformation;

[0086] Then, the final output g is calculated by multiplying the attention weight matrix h and g w , the attention weight matrix calculation method is as follows:

[0087]

[0088] Among them, i and j are the indices in the height and width directions of the traversal feature map, which are used to indicate the position of the pixel in the feature map, specifically described as the input feature x c (i, j) and channel adjusted features and Perform point multiplication to obtain the output feature y c (i, j). This step enables the network to distinguish important feature information more accurately.

[0089] This paper adds an attention module to the neck network of YOLOv9. By integrating spatial position encoding and channel feature calibration mechanisms, it achieves precise focus on target features in complex water scenes. The contributions of the CA attention mechanism to this paper are:

[0090] (1) CA is a spatial-channel collaborative anti-interference attention mechanism that achieves position-sensitive feature enhancement through decomposition-based global pooling and coordinate information interaction. This mechanism generates direction-aware position encoding vectors along the horizontal / vertical direction, breaking through the traditional attention's homogeneous compression of spatial information and establishing an explicit mapping relationship between spatial position and channel features. Position-sensitive channel attention weights are generated through matrix outer product operations, enabling the model to accurately locate small target areas under strong reflective interference and suppress pseudo-activation responses from background noise. The attention diffusion network constructed with lightweight convolution kernels maintains computational efficiency while enhancing adaptability to target morphological distortion caused by wave disturbances, effectively solving the detection accuracy problem in complex water scenes.

[0091] (2) The CA mechanism improves the multi-scale target detection capability through a cross-dimensional feature fusion strategy. It combines position encoding information with the feature pyramid network (FPN) and dynamically allocates attention weights on feature maps at different levels: shallow features focus on spatial position information to enhance small target positioning, while deep features focus on semantic information to enhance complex target recognition. The coordinate information embedding mechanism of CA provides a new technical path for multi-scale feature fusion, especially showing significant advantages in the detection of fuzzy targets such as semi-submerged floating objects, laying the foundation for high-precision detection in complex dynamic environments.

[0092] experiment

[0093] To evaluate the effectiveness of the present invention in detecting floating objects on water surfaces, we conducted experiments on the FloW-img dataset (containing 2000 images) and the Floating Objects-2400 dataset (containing 2400 images), and compared it with 14 object detection algorithms. We compared YOLOv5, ATD-CNN, DES-YOLOv5, YOLOv6, YOLOv7, UltraWS, YOLOv8, MFG-YOLOv8, YOLOv8-DBS, SOE-YOLO, Global-YOLO, YOLOv9, YOLOv9-Slim, and YOLOv9-B.

[0094] 1. Parameter selection

[0095] Table 1 Experimental environment configuration table

[0096]

[0097]

[0098] The environment configuration used in this experiment is shown in Table 1. Experimental hyperparameters were set as follows: input image resolution was 640×640, batch size was set to 8, and the total number of iterations was set to 500. The Average Precision (AP) metric comprehensively evaluates the model's performance across confidence thresholds by calculating the integrated area of ​​the precision-recall curve (PR curve), as shown below:

[0099] AP = ∫0 1 P(R)dR

[0100] mAP refers to the average detection accuracy obtained by taking the average AP value of all categories. Its calculation process is as follows:

[0101]

[0102] Where N represents the number of all categories and APi represents the target detection accuracy of category i.

[0103] 2. Quantitative analysis

[0104] To evaluate the advantages of the proposed model over other object detection algorithms, this section presents a comprehensive comparative analysis, showing the results of comparing YOLOv9-BBC with 14 object detection algorithms on the FloW-Img and Floating Objects-2400 datasets. Compared to these eight object detection algorithms, the proposed YOLOv9-BBC algorithm achieved the highest mAP@0.5 value and the second highest mAP@0.5-95 value on the FloW-Img data set, and the highest mAP@0.5 and mAP@0.5-95 values ​​on the Floating Objects-2400 dataset.

[0105] Table 2 Comparison of training results between YOLOv9-BBC and various mainstream algorithms

[0106]

[0107]

[0108] As shown in Table 2, on the FloW-Img dataset, the YOLOv9-BBC algorithm demonstrates outstanding technical advantages, achieving mAP@0.5 of 81.75% and mAP@0.5:0.95 of 37.85%. With the exception of mAP@0.5:0.95, which is slightly lower than YOLOv9's 38.49%, all other metrics significantly outperform the competing algorithms. YOLOv9-BBC achieves a 1.58% improvement in mAP@0.5 over YOLOv9, thanks to a triple optimization strategy: a feature extraction module based on the BiFormer dynamic sparse attention mechanism, improvements to the network structure incorporating BoTNet, and the introduction of the CA attention mechanism.

[0109] In a comparative experiment on the Floating Objects-2400 dataset, the YOLOv9-BBC algorithm performed exceptionally well, achieving 82.01% mAP@0.5 and 38.99% mAP@0.5:0.95. Compared to Global-YOLO's 79.49% and 36.13%, mAP@0.5 improved by 2.52% and mAP@0.5:0.95 improved by 2.86%. Compared to YOLOv8-DBS's 79.85% and 39.30%, mAP@0.5 improved by 2.16%. Compared to YOLOv9's 80.53% and 37.06%, mAP@0.5 improved by 1.48% and mAP@0.5:0.95 improved by 1.93%.

[0110] In a comparative experiment on the Floating Objects-2400 dataset, the YOLOv9-BBC algorithm performed exceptionally well, achieving 82.01% mAP@0.5 and 38.99% mAP@0.5:0.95. Compared to Global-YOLO's 79.49% and 36.13%, mAP@0.5 improved by 2.52% and mAP@0.5:0.95 improved by 2.86%. Compared to YOLOv8-DBS's 79.85% and 39.30%, mAP@0.5 improved by 2.16%. Compared to YOLOv9's 80.53% and 37.06%, mAP@0.5 improved by 1.48% and mAP@0.5:0.95 improved by 1.93%.

[0111] In comparative experiments on the Floating Objects-2400 dataset, the YOLOv9-BBC algorithm performed exceptionally well, achieving a mAP@0.5 score of 82.01% and a mAP@0.5:0.95 score of 38.99%. Compared to Global-YOLO's 79.49% and 36.13%, the mAP@0.5 score increased by 2.52% and the mAP@0.5:0.95 score increased by 2.86%. Compared to YOLOv8-DBS's 79.85%, the mAP@0.5 score increased by 2.16%. Compared to YOLOv9's 80.53% and 37.06%, the mAP@0.5 score increased by 1.48% and the mAP@0.5:0.95 score increased by 1.93%. The algorithm also significantly outperformed algorithms like SOE-YOLO.

[0112] The experimental results reveal the root causes of the differences: Glod-YOLO lags behind in detection indicators due to its insufficient adaptability to complex interference in water areas and shortcomings in feature modeling; although YOLOv8-DBS optimizes detection capabilities through multiple modules, the limitations of its feature extraction architecture make it difficult to surpass YOLOv9-BBC in mAP@0.5:0.95 when facing wave-deformed targets; although YOLOv9 has basic innovative designs, there is room for improvement in the feature fusion efficiency in complex water scenes. YOLOv9-BBC, with its more optimized architectural design, has significant advantages in target detection in dynamic water scenes.

[0113] YOLOv9-BBC achieves a performance leap through innovative module improvements: BiFormer-based feature extraction: Dynamically screen key feature areas with the help of adaptive sparse attention, while reducing the amount of computation and strengthening the focus on edge features of small targets; BoTNet network structure integration: Construct a hybrid architecture of Transformer and CNN, capture the global spatial correlation of distant targets through the multi-head self-attention mechanism, and combine deep separable convolution to enhance the perception of local details; CA attention mechanism is introduced: Through the coordinated calibration of spatial position and channel features, the interference of wave reflections is suppressed and the response strength of small target areas is enhanced.

[0114] Experimental data demonstrates that the synergistic effect of these three improvements enables YOLOv9-BBC to maintain stable detection accuracy in complex water scenes. This multi-dimensional architectural innovation significantly improves the robustness of detection in complex water scenes while ensuring lightweight design, fully demonstrating its strong adaptability to dynamically deformed objects.

[0115] 3. Qualitative analysis

[0116] Figure 5 、 Figure 6 The training curves of the YOLOv9 algorithm and the YOLOv9-BBC algorithm are generated in the figure. The following conclusions can be drawn from these curves: (1) In the bounding box loss, classification loss and distribution focus loss curves, the loss of the YOLOv9-BBC algorithm decreases more smoothly and the final convergence value is lower, indicating that it optimizes the classification task more thoroughly, has better error control of target positioning and classification during generalization, and has a lower risk of overfitting; (2) The precision and recall curves show that the precision and recall of the YOLOv9-BBC algorithm remain at a higher level in the later stage of training, and the curves are smoother, indicating that it has a stronger ability to discriminate targets and fewer problems of missed detection and false detection. (3) In the mAP indicator curve, the YOLOv9-BBC algorithm finally achieves a higher mAP value, reflecting that its comprehensive accuracy in target detection is better than that of the YOLOv9 algorithm.

[0117] The YOLOv9-BBC algorithm outperforms the YOLOv9 algorithm in terms of loss function convergence and detection performance indicators, fully demonstrating the effectiveness of its improved module in improving model training stability and detection performance. Ultimately, it is concluded that the YOLOv9-BBC algorithm performs better than the YOLOv9 algorithm.

[0118] In order to intuitively analyze the training results of the model, the training process on the FloW-Img dataset is visualized to further clarify the performance advantages of the YOLOv9-BBC algorithm. This section uses the curve chart of the change of mAP@0.5 value on the validation set during the training process of the YOLOv9-BBC algorithm as the training round (epoch) to carry out a visual analysis. Figure 7 As shown in the figure, the horizontal axis represents the training round, and the vertical axis represents the mAP@0.5 value of the model on the validation set.

[0119] In the early stages of training, YOLOv9-BBC quickly gained momentum, with its mAP@0.5 curve rising rapidly, far exceeding YOLOv8, YOLOv8-DBS, and YOLOv9, demonstrating its efficient extraction of data features and parameter optimization. Near the end of training, YOLOv9-BBC achieved the highest mAP@0.5 value, leading in detection accuracy and accurately identifying and locating targets, highlighting the significant effectiveness of the improvements. In the second half of training, the algorithm's curve remained stable, with minimal fluctuations and a high level of accuracy. It demonstrated strong interference resistance, was less susceptible to overfitting and underfitting, and exhibited excellent data adaptability. Furthermore, compared to other models, YOLOv9-BBC demonstrated superior convergence speed, final detection accuracy, and stability during training. This not only strongly demonstrates its superior performance in object detection tasks, but also fully validates the effectiveness of the algorithm's innovative module combination.

[0120] like Figure 8As shown in the figure, by comparing the detection visualization results of the YOLOv9 algorithm and the YOLOv9-BBC algorithm, the advantages and disadvantages can be analyzed from the following dimensions: (1) The YOLOv9-BBC algorithm has a more complete detection of floating objects on the water surface. In some pictures, the YOLOv9 algorithm has the phenomenon of missing small targets or targets in the edge area, while the YOLOv9-BBC algorithm can stably identify more targets and cover a wider target area. In particular, it has better detection effects on dense or hidden floating objects, effectively reducing the problem of missed detection. (2) The detection confidence annotation of the YOLOv9-BBC algorithm is more scientific. The confidence of some detection frames of the YOLOv9 algorithm is low, and even the same target is repeatedly labeled or the confidence is abnormal; the confidence of the YOLOv9-BBC algorithm is mostly concentrated in a reasonable range, and the annotation logic is clear, indicating that its positioning and classification judgment of the target is more accurate, and the detection results are more reliable. (3) YOLOv9-BBC has stronger anti-interference ability. In areas with complex backgrounds such as water surface reflections and plant shadows, the YOLOv9 algorithm is prone to misjudging objects as floating objects; however, the YOLOv9-BBC algorithm can more accurately distinguish real targets from environmental interference, reduce false detections, and highlight its superior detection robustness in complex water scenes.

[0121] Overall, YOLOv9-BBC outperforms YOLOv9 in target detection completeness, confidence accuracy, and anti-interference performance in complex scenarios, fully demonstrating the significant effectiveness of its improved module in improving detection performance. Ultimately, it is concluded that the performance of the YOLOv9-BBC algorithm is better than that of the YOLOv9 algorithm.

[0122] 4. Ablation Analysis

[0123] In order to verify the effectiveness of the proposed BiFormer-based dynamic sparse attention mechanism feature extraction module, the network structure improvement integrated with BoTNet, and the introduction of the CA attention mechanism, this section conducted ablation experiments on the FloW-Img and Water Surface Floating Objects-2400 datasets respectively.

[0124] The experiment is based on the YOLOv9 model. The YOLOv9-BBC algorithm verifies the effect of the improved model using the original YOLOv9 model as a control. The total number of training epochs is 500. "√" indicates that the improved module is used in the YOLOv9 model, and "-" indicates that the improved module is not used in the YOLOv9 model. mAP is used as the evaluation indicator. The ablation experiments are shown in Tables 3 and 4, which list the ablation results of several modules of the algorithm, with the best results indicated in bold.

[0125] Table 3 Ablation test results of YOLOv9-BBC and each module on the FloW-Img dataset

[0126]

[0127] Table 4 Ablation test results of YOLOv9-BBC and each module on the water surface floating objects-2400 dataset

[0128]

[0129]

[0130] By comparing the baseline model YOLOv9 and the YOLOv9-BBC algorithm with the FloW-Img dataset, we can draw the following conclusions:

[0131] Introducing each module individually improves model performance, demonstrating the unique role of each module. When the BiFormer module is introduced alone, the mAP@0.5 on the FloW-Img dataset increases from 80.17% to 81.02%. Although the mAP@0.5:0.95 ratio decreases from 38.49% to 37.65%, its dynamic screening of key feature regions enhances the ability to focus on small target edge features during lightweight feature extraction. When the BoTNet module is incorporated alone, the mAP@0.5 ratio increases to 81.18%, while the mAP@0.5:0.95 ratio decreases to 37.60%. This module leverages a hybrid Transformer and CNN architecture to accurately capture the spatial correlation of distant targets and enhance global feature modeling. When the CA module is added alone, the mAP@0.5 ratio reaches 80.99%, while the mAP@0.5:0.95 ratio is 37.63%, demonstrating its fundamental effectiveness in spatial-channel feature co-calibration and interference suppression.

[0132] The combination of multiple modules triggers a synergistic optimization effect. When BiFormer is combined with BoTNet, mAP@0.5 reaches 81.20% and mAP@0.5:0.95 reaches 37.62%, achieving a two-way optimization of local feature preservation and global correlation perception. When BiFormer is combined with CA, mAP@0.5 increases to 81.28% and mAP@0.5:0.95 increases to 37.66%. By combining dynamic sparse attention with anti-interference mechanisms, it suppresses noise while highlighting small target features. When BoTNet is combined with CA, mAP@0.5 reaches 81.26% and mAP@0.5:0.95 reaches 37.62%, reflecting the synergistic promotion of detection accuracy by global modeling and anti-interference mechanisms.

[0133] When BiFormer, BoTNet, and CA modules are combined to form YOLOv9-BBC, the model achieves optimal performance, with mAP@0.5 reaching 81.75% and mAP@0.5:0.95 improving to 37.85%, both significantly outperforming other configurations. This demonstrates that the deep collaboration of the three modules—lightweight feature extraction, global context modeling, and dynamic anti-interference attention—not only improves mAP@0.5 but also optimizes the demanding mAP@0.5:0.95 ratio, comprehensively enhancing the model's detection capabilities for targets in complex aquatic scenes. This fully demonstrates the necessity of improving each module of YOLOv9-BBC and the significant effectiveness of combining multiple modules in improving algorithm robustness and detection accuracy.

[0134] To validate the generalization performance of the proposed method on a general floating object dataset, we conducted multiple ablation experiments on the Floating Objects-2400 dataset, using the same experimental setup as the ablation experiments on the FloW-Img dataset. The experimental results demonstrate similar performance improvements to those achieved on the FloW-Img dataset, validating the necessity, effectiveness, and generalization of each module in the YOLOv9-BBC algorithm.

[0135] The present invention also provides a surface floating object detection system based on improved YOLOv9, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the steps of any of the above-described methods can be implemented.

[0136] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the steps of any of the above methods can be implemented.

[0137] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any of the above methods.

[0138] References:

[0139] [1] Li Chang. Research on surface target detection and recognition algorithm[D]. Huazhong University of Science and Technology, 2016.

[0140] [2] Fang Jing, Feng Shunshan, Feng Yuan. Image detection method of ship targets by surface vehicles[J]. Journal of Beijing Institute of Technology, 2017, 37(12): 1235-1240.

[0141] [3] Wei Jianrong. Research on application of background texture model in sea surface ship target detection[J]. Ship Science and Technology, 2017, 39(20): 159-161.

[0142] [4] Zhang L, Zhang Y, Zhang Z, et al. Real-Time Water Surface ObjectDetection Based on Improved Faster R-CNN[J]. Sensors, 2019, 19(16): 3523-3535.

[0143] [5] Liu Wei, Wang Yuannan, Jiang Shan, et al. Research on surface floating object recognition method based on Mask R-CNN[J]. People's Yangtze River, 2021, 52(11): 226-233.

[0144] [6] Li Guojin, Yao Dongyi, Ai Jiaoyan, et al. Surface floating object detection method based on improved YOLOv3 algorithm [J]. Journal of Guangxi University (Natural Science Edition), 2021, 46(06): 1569-1578.

[0145] [7] Wang Lin, Wang Yuting. Lightweight ship target detection based on enhanced feature fusion[J]. Computer Systems and Applications, 2023, 32(02): 288-294.

[0146] [8]Varghese R,Sambath M.YOLOv8:A Novel Object Detection Algorithmwith Enhanced Performance and Robustness[C] / / Proceedings of 2024International Conference on Advances in Data Engineering and IntelligentComputing Systems,2024:1-6.

[0147] [9]Wang CY, Yeh IH, Mark Liao H Y.YOLOv9: Learning What You Want toLearn Using Programmable Gradient Information[C] / / Proceedings of the European Conference on Computer Vision, 2025:1-21.

[0148]

[10] Zhu L, Wang

[0149]

[11] Srinivas A,Lin TY,Parmar N,et al.Bottleneck Transformers forVisual Recognition[C] / / Proceedings of the 2021IEEE / CVF Conference on ComputerVision and Pattern Recognition,2021:16514-16524.

[0150]

[12] Hou Q, Zhou D, Feng J. Coordinate Attention for Efficient MobileNetwork Design[C] / / Proceedings of the 2021IEEE / CVF Conference on ComputerVision and Pattern Recognition, 2021:13708-13717.

[0151] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A method for detecting floating objects on a water surface based on improved YOLOv9, characterized in that: include: A feature extraction module based on the BiFormer dynamic sparse attention mechanism is introduced. Through dual-layer routing attention and pyramid structure design, it reduces the amount of computation while improving the ability to retain shallow small object features. A hybrid architecture integrating BoTNet is constructed, embedding BoTNet in the neck network. By leveraging the collaborative optimization of the multi-head self-attention mechanism and depthwise separable convolution, the model can more accurately grasp the features of objects at different scales. An anti-interference strategy guided by the CA attention mechanism is introduced. A CA attention mechanism module is introduced at the end of the neck network. Through the spatial-channel dual-branch cross-dimensional interaction and the coordinate-sensitive dynamic activation mechanism, the noise interference caused by the reflection of water waves is suppressed and the feature response intensity of small target areas is enhanced.

2. The method for detecting floating objects on the water surface based on improved YOLOv9 according to claim 1, wherein: The feature extraction module based on the BiFormer dynamic sparse attention mechanism is introduced, that is, the BiFormer dynamic sparse attention mechanism is introduced at the end of the neck network of YOLOv9. The BiFormer dynamic sparse attention mechanism is based on the deep optimization of the Transformer architecture.

3. The method for detecting floating objects on a surface of water based on improved YOLOv9 according to claim 1 or 2, wherein: BiFormer dynamic sparse attention mechanism proposes a two-layer routing attention mechanism BRA, which realizes the intelligent allocation of computing resources through region-level affinity graph construction and dynamic pruning strategy; BiFormer dynamic sparse attention mechanism consists of two steps: routing step and token-level attention step; BiFormer dynamic sparse attention mechanism adopts a four-stage pyramid architecture to realize hierarchical feature extraction.

4. The method for detecting floating objects on the water surface based on improved YOLOv9 according to claim 1, wherein: BoTNet is a backbone architecture that combines the self-attention mechanism and convolutional neural network. BoTNet is constructed by replacing the last three spatial 3×3 convolutions with MHSA in the last three bottleneck blocks of ResNet.

5. The method for detecting floating objects on a water surface based on improved YOLOv9 according to claim 1 or 4, wherein: BoTNet is embedded after the RepNCSPELAN4 module of the neck network of YOLOv9.

6. The method for detecting floating objects on the water surface based on improved YOLOv9 according to claim 1, wherein: The CA attention mechanism module achieves accurate encoding of spatial position information through a decomposition-based global average pooling strategy. The decomposition-based global average pooling strategy generates direction-aware position encoding vectors along the horizontal / vertical directions, and fuses the position encoding vectors in these two directions through the coordinate information interaction layer to establish an explicit mapping relationship between spatial position and channel features, thereby generating position-sensitive attention weights in the channel dimension.

7. The method for detecting floating objects on the water surface based on improved YOLOv9 according to claim 6, characterized in that: The CA attention mechanism module achieves accurate encoding of spatial position information through a decomposition-based global average pooling strategy. The specific implementation formula is as follows: Given x c , use the two spaces of the pool kernel (h, i) or (j, w) to average the output feature map in the horizontal and vertical directions respectively and The formula is as follows: H and W are the height and width of the input feature map x; c represents the c-th channel in the input feature map x; Next, the feature maps of the global receptive field obtained in the horizontal and vertical directions are concatenated, and then a 1×1 convolution is performed. Finally, a nonlinear operation is performed using the Sigmoid activation function, as shown in the following formula: f=δ(F1([z h ,z w ])) Among them, [z k , z w ] means z h and z w Perform splicing operation, F1(·) represents 1×1 convolution operation, δ(·) represents Sigmoid activation function, that is, normalize and nonlinearize the feature map after convolution, and finally obtain the feature map f∈R C / r×(H+W) It is an intermediate feature containing horizontal and vertical spatial information, and r represents the step size of the pooling; Then the feature map f is divided into two parts f h ∈R C / r×H and f w ∈R C / r×W , then adjust the number of channels to f h Adjust the number of channels to be consistent with the input feature map x, as shown in the following formula: g h =σ(F h (f h )) g w =σ(F w (f w )) Among them, F h (·), F w (·) is the use of two 1×1 convolution F h and F w f h and f w The channel adjustment function for feature conversion, σ(·) is the activation function, which is used to adjust the f h and f w Perform linear transformation; Then, the final output g is calculated by multiplying the attention weight matrix h and g w , the attention weight matrix calculation method is as follows: Among them, i and j are the indices in the height and width directions of the traversal feature map, which are used to indicate the position of the pixel in the feature map, specifically described as the input feature x c (i, j) and channel adjusted features and Perform point multiplication to obtain the output feature y c (i,j).

8. A water surface floating object detection system based on improved YOLOv9, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 7 can be implemented.

9. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 7 can be implemented.

10. An electronic device comprising a processor and a memory, wherein: The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.