A method for quickly detecting water hyacinth targets based on improved YOLO11
By improving the YOLO11 model and introducing the EfficientViT, Lite-BiFPN, and SEAMHead modules, the problem of insufficient detection accuracy of water hyacinth targets in aquatic environments was solved, and efficient and accurate water hyacinth target identification and real-time detection were achieved.
Patent Information
- Application Number
- CN202510927810.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing YOLO models lack sufficient accuracy in detecting water hyacinth targets in aquatic environments, especially in complex water backgrounds where they struggle to effectively identify small and large targets. Furthermore, they are not adaptable to interference factors such as water surface reflection and wave textures.
By optimizing the YOLO11 model, introducing the EfficientViT module for feature extraction, and combining Lite-BiFPN multi-scale feature fusion and SEAMHead attention mechanism, the model's ability to detect water hyacinth targets is enhanced.
It significantly improves the detection accuracy and speed of water hyacinth targets, enhances the model's adaptability in complex water environments, improves the recognition accuracy of small targets, and reduces the number of model parameters and computational complexity, making it suitable for resource-constrained embedded platforms.
Smart Images

Figure CN120635395B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ecosystem monitoring technology, and in particular to a rapid detection method for water hyacinth targets based on an improved YOLO11. Background Technology
[0002] Water hyacinth (Eichhornia crassipes) is a perennial floating aquatic plant native to the Amazon basin in South America, and is now widely distributed in warm and humid regions worldwide. Due to its rapid growth and strong reproductive capacity, water hyacinth can quickly form a dense vegetation cover after invading water bodies, severely hindering water flow, reducing oxygen levels, and consequently causing significant damage to aquatic organisms and ecosystems. Furthermore, large-scale accumulation of water hyacinth can also affect shipping, fisheries, and water conservancy projects, posing significant challenges to the ecological environment and socio-economic development. Therefore, timely and accurate monitoring of water hyacinth distribution is of great importance for aquatic ecological environment protection and water resource management.
[0003] With the rapid development of deep learning and computer vision technologies, the automated processing capabilities of remote sensing data are constantly improving, especially in target detection. Deep learning-based detection methods, due to their powerful feature learning capabilities and high-precision recognition, have become an important research direction in vegetation monitoring. Among numerous target detection methods, the YOLO (YouOnly Look Once) series of models, with their efficient real-time detection capabilities and good target recognition performance, have shown broad application potential in aquatic vegetation monitoring. YOLO employs a single-stage convolutional neural network (CNN), enabling it to process the entire image in real time and efficiently complete target detection tasks.
[0004] Although the YOLO model has been widely used for detecting invasive terrestrial plants, its application in aquatic environments, especially for the detection of floating plants such as water hyacinth, still faces many challenges. Water hyacinth grows in complex aquatic environments, and its spectral differences from other aquatic vegetation and water surface reflections are relatively small, resulting in room for improvement in target recognition accuracy. Therefore, how to improve the detection accuracy of the YOLO model in aquatic environments, especially for the accurate identification of water hyacinth targets, is the key issue of this patent.
[0005] In existing technologies, research has utilized the YOLO algorithm for the detection of invasive alien plants. For example, Chinese patent CN119399688A provides a method for detecting invasive alien plants based on an improved YOLOv9 algorithm. This method improves upon YOLOv9 by introducing the DynamicConv module, TripletAttention mechanism, and MPDIoU loss function to enhance the detection capability of invasive alien plants, especially those against complex backgrounds. Specific steps include: acquiring images of invasive alien plants, annotating the images, constructing a dataset, training the improved YOLOv9 model, and testing and evaluating the detection results.
[0006] Although the existing technology uses the YOLO algorithm to detect invasive alien plants, the method still has the following shortcomings:
[0007] a) Lack of optimization specifically for aquatic plants such as water hyacinth: Existing technologies mainly target terrestrial invasive plants such as ragweed and lantana. For the specific detection needs of floating aquatic plants like water hyacinth, there is a lack of targeted optimization, especially under interference from water surface reflection and background, where detection results may be unsatisfactory. Floating plant communities exhibit significant size variations, and existing methods are ineffective at detecting both small and large targets, leading to missed detections of multi-scale targets.
[0008] b) Insufficient adaptability to complex aquatic backgrounds: Traditional detection models are inadequate in feature extraction under interference from water surface reflections, wave textures, etc. Although the YOLOv9 model has a certain background suppression capability, it may not achieve ideal accuracy when dealing with interference factors in the aquatic environment, such as water surface fluctuations and reflected light. Therefore, existing technologies have not fully considered the complexity of water hyacinth in aquatic environments.
[0009] c) Insufficient accuracy in identifying small targets: Aquatic plants such as water hyacinth are usually widely and densely distributed. Water hyacinth is often mixed with duckweed and algae, and partial occlusion leads to poor sensitivity. Detecting plant targets in a small area often faces the problem of insufficient recognition accuracy. Summary of the Invention
[0010] To overcome or alleviate one or more of the above-mentioned technical problems, this invention proposes a fast detection method for water hyacinth targets based on an improved YOLO11 model. The WH-YOLO11 model is constructed by optimizing the structure and parameters of the YOLO model, combining efficient convolutional feature extraction, a lightweight local attention mechanism, an optimized feature fusion strategy, and an enhanced spatial-channel attention module. This significantly improves the detection accuracy and speed of water hyacinth targets in complex water environments, enhances the model's adaptability to complex water conditions, and improves the detection accuracy of water hyacinth targets.
[0011] This invention provides the following technical solution:
[0012] This invention provides a rapid detection method for water hyacinth targets based on an improved YOLO11, which includes the following steps:
[0013] S1. Data Preparation:
[0014] Image data is collected, covering different weather, lighting and aquatic environments. The collected images are screened, sorted and classified, and data annotation is performed to clarify the category and location information of targets in the images, thus constructing a dataset. The dataset is then subjected to format conversion, size adjustment and enhancement processing, and divided into training set, validation set and test set.
[0015] S2. Model Optimization and Training:
[0016] The backbone network of the original YOLO11 model was optimized using EfficientViT; a Lite-BiFPN multi-scale feature fusion structure was introduced into the Neck part of the original YOLO11 model to improve the information flow and fusion effect of features at different scales, thereby enhancing the model's ability to detect small targets; the convolutional modules of the original YOLO11 model were replaced with SEAMHead convolutional modules that introduce channel and spatial attention mechanisms to improve the extraction and processing capabilities of key target features, thus constructing the WH-YOLO11 model; the WH-YOLO11 model was trained using preprocessed data, and hyperparameters were adjusted to improve detection performance;
[0017] S3. Testing and Evaluation:
[0018] The trained WH-YOLO11 model was used to detect targets in test image data, automatically identifying water hyacinth targets in the images. The performance was evaluated using precision, recall, F1 score, and mAP metrics to quantify the accuracy and reliability of the detection.
[0019] According to some implementation methods, step S2 model optimization and training includes the following steps:
[0020] S21. Backbone network optimization: The backbone network of the original YOLO11 model is optimized using EfficientViT. By combining local convolution and self-attention mechanism, the ability to extract image features is improved.
[0021] S22, Neck structure optimization: The Lite-BiFPN multi-scale feature fusion structure is introduced into the Neck part of the YOLO11 model to improve the fusion effect of multi-scale features and enhance the detection capability of small targets;
[0022] S23. Convolutional module replacement: Replace the original convolutional module of the YOLO11 model with the C3K2 convolutional module and the SEAMHead convolutional module to improve the model's ability to extract key features of the target.
[0023] S24. Model Training and Tuning: Train the WH-YOLO11 model using the preprocessed dataset from step S1, and improve the model's detection performance and accuracy by adjusting hyperparameters.
[0024] According to some implementation methods, the process of optimizing the backbone network structure of the original YOLO11 model using EfficientViT in step S21 is as follows:
[0025] Assuming the input feature is F, the EfficientViT module first obtains the local feature representation F through local convolution. local Then, a lightweight self-attention mechanism is used to obtain the global feature representation F. global Finally, the local and global features are fused through weighted fusion to obtain the enhanced feature F. enhanced The specific calculation formula is as follows:
[0026] F enhanced =β·F local +(1-β)·F global (1)
[0027] Where β is the fusion weight for dynamic learning, and F is the weight for obtaining global features. g The lobal model, based on a self-attention mechanism, has the following core computational process:
[0028]
[0029] Where Q, K, and V represent the query matrix, key matrix, and value matrix of the input features, respectively, and d k is the dimension of the key matrix, used for scaling and normalization. Equation 7 achieves the fusion of contextual information by calculating the attention weights between each position and all global positions, and outputs Z as a global context-aware feature representation.
[0030] According to some implementation methods, the process of introducing a Lite-BiFPN multi-scale feature fusion structure into the Neck part of the original YOLO11 model as described in step S22 is as follows:
[0031] Suppose that the multi-scale features are represented as F1, F2, ..., F at a certain layer. n Then, bidirectional feature fusion is performed using BiFPN to finally obtain the fused feature F. fused It can be expressed by the following formula:
[0032]
[0033] Among them, w i This represents the weight at each scale.
[0034] According to some implementation methods, the process of replacing the convolutional module of the original YOLO11 model with the SEAMHead convolutional module that introduces channel and spatial attention mechanisms to construct WH-YOLO11 in step S23 is as follows:
[0035] In the SEAMHead module, assuming the input feature is F, the weighted feature calculated through channel and spatial attention mechanisms is Fi. att It can be expressed by the following formula:
[0036] Channel attention:
[0037]
[0038] Where σ is the Sigmoid function, and C(F) represents the channel weights output by the channel attention network. These are the characteristics after channel weighting;
[0039] Spatial attention:
[0040]
[0041] Where S(F) is the spatial weight of the spatial attention network output. These are spatially weighted features;
[0042] Finally, by weighted fusion of channel and spatial factors, the final attention features are obtained, calculated as follows:
[0043]
[0044] Here, α is a fusion factor that adjusts the channel and spatial attention weights. It is a fusion weight dynamically generated by a lightweight multilayer perceptron and has a value range of (0, 1).
[0045] According to some implementation methods, a C3K2 module is used before the SEAMHead module to enhance feature extraction capabilities. The C3K2 module combines multiple convolutional kernels with residual connections, and optimizes the feature extraction process through effective residual learning, thereby improving the network's feature refinement capabilities and detection accuracy. The calculation formula is as follows:
[0046] F out =σ(W2·σ(W1·F) in +B1)+B2) (6)
[0047] F inThe input feature map; W1 and W2 are the convolution kernel weights, B1 and B2 are bias terms, and σ is the activation function; residual connections achieve cross-layer propagation through +B2; F out This is for outputting feature maps.
[0048] According to some implementation methods, step S1 data preparation includes the following steps:
[0049] S11. Image Acquisition: Acquire relevant image data containing the target object, covering image samples under different weather conditions, lighting conditions, and different aquatic environments;
[0050] S12. Constructing a dataset: Filter and organize the collected image data, and establish a dataset according to certain standards;
[0051] S13. Data annotation: Use professional annotation tools to annotate the images in the dataset, clarify the location and category information of the targets in each image, and generate annotated image data that meets the training requirements;
[0052] S14. Image preprocessing: Perform preprocessing operations on the labeled image data, including format conversion, size adjustment, and data augmentation, to improve data quality and the effectiveness of model training. The processed data is used as the input to the model.
[0053] According to some implementation methods, step S3 detection and evaluation includes the following steps:
[0054] S31. Object Detection: The trained WH-YOLO11 model is applied to the test image data to perform object detection tasks, automatically identifying target objects in the images. The evaluation metrics are set as follows:
[0055]
[0056] Precision is the overall accuracy, TP is the number of true positives, and FP is the number of false positives; Recall is the overall recall, and FN is the number of false negatives; mAP@0.5 is the mean of the average precision across all classes with an IoU threshold of 0.5; N cls The total number of categories participating in the assessment, AP c Let c be the average precision of class c, where c is from 1 to N. cls _n_ natural numbers;
[0057] S32. Precision Evaluation: The performance of the target detection results is evaluated using precision, recall, F1 score and mean precision as evaluation indicators, which quantifies the accuracy and reliability of the detection.
[0058] S33. Water hyacinth distribution detection: Using the WH-YOLO11 model, which has been verified by accuracy evaluation, water hyacinth distribution detection is performed on image data in actual application scenarios. The location distribution information of water hyacinth in the image is output to achieve effective monitoring of water hyacinth invasion in water areas.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] (1) Improvement of the feature extraction module based on EfficientViT
[0061] The EfficientViT module is introduced into the YOLO11 backbone network structure, fully combining the advantages of Convolutional Neural Networks (CNNs) in local feature extraction with the capabilities of Visual Transformers (ViTs) in global modeling. This module utilizes the synergistic effect of lightweight local convolutions and self-attention mechanisms to effectively enhance the model's ability to extract features of water hyacinth targets in complex backgrounds while maintaining inference efficiency, significantly improving detection accuracy and robustness.
[0062] (2) Employing a multi-scale feature fusion structure based on Lite-BiFPN
[0063] To address the challenges posed by the large variations in target size and uneven distribution density of water hyacinth, this invention introduces a Lite-BiFPN module into the YOLO11 Neck structure. This module constructs a bidirectional feature pathway (top-down + bottom-up) to enable information flow and dynamic fusion of multi-scale features. Through a learnable weighting mechanism, this module improves feature fusion efficiency, effectively enhancing the model's performance in small target detection and complex background processing.
[0064] (3) Detection head design integrating SEAMHead attention mechanism
[0065] The detection head incorporates a structure that integrates the C3K2 and SEAMHead modules. The C3K2 module employs a 3×3 convolutional and residual connection structure to enhance the expressive power of deep semantic features; while the SEAMHead module introduces spatial and channel attention mechanisms to further strengthen the focus on key feature regions of water hyacinths and effectively suppress redundant background interference such as water reflection, thereby comprehensively improving the accuracy and robustness of target detection.
[0066] (4) Overall lightweight design and efficient inference optimization
[0067] By integrating lightweight modules such as EfficientViT, Lite-BiFPN, and SEAMHead, this invention significantly improves detection performance while significantly reducing the number of model parameters and computational complexity, resulting in faster computation and more efficient detection. This lightweight design ensures stable operation of the model on resource-constrained embedded platforms or mobile devices, meeting the application requirements for real-time and high-precision target detection in aquatic environments. Attached Figure Description
[0068] Figure 1 This is a flowchart illustrating the rapid detection method for water hyacinth targets based on the improved YOLO11 provided in an embodiment of the present invention.
[0069] Figure 2 The diagram shows the WH-YOLO11 structure constructed in step S2 of this embodiment of the invention.
[0070] Figure 3 The recall, precision, mAP50, and mAP50 95 curves provided for embodiments of the present invention.
[0071] Figure 4 The figure shows the experimental results provided in the embodiments of the present invention. Detailed Implementation
[0072] For a detailed introduction to the existing YOLO11 model, please see Ultralytics, YOLOv11, GitHub Repository: https: / / github.com / ultralytics / ultralytics.
[0073] This invention constructs a WH-YOLO11 model for detecting water hyacinth as a target. It improves the original YOLO11 model by combining three enhancements: EfficientViT feature extraction, Lite-BiFPN multi-scale fusion, and SEAMHead attention decoding. The technical features of each module are described below:
[0074] I. EfficientViT Design
[0075] EfficientViT is a highly efficient and lightweight visual transformer architecture that combines the local modeling capabilities of convolutional neural networks (CNNs) with the global perception capabilities of transformers (ViTs). It can significantly improve the modeling effect of key features in images while maintaining low computational complexity.
[0076] In this invention, the application of the EfficientViT module is specifically reflected in the following aspects:
[0077] 1. Local and Global Feature Fusion: Fine-grained features are extracted through local convolution, and a self-attention mechanism is used to model long-distance dependencies in the image, achieving effective fusion of local and global features. This enhances the model's ability to extract water hyacinth targets in complex backgrounds (such as water ripples, reflections, etc.).
[0078] 2. Efficient attention computation: EfficientViT adopts a lightweight attention computation method, which significantly reduces computational overhead compared to the traditional ViT structure, making it suitable for deployment in scenarios requiring real-time detection.
[0079] 3. Dynamic Feature Modeling: During the feature extraction process, EfficientViT introduces dynamic convolution kernels and adaptive attention mechanisms to dynamically adjust the feature extraction strategy based on the content of the input image, thereby improving the model's adaptability and detection accuracy in different aquatic environments.
[0080] Assuming the input feature is F, the EfficientViT module first obtains the local feature representation F through local convolution. local Then, a lightweight self-attention mechanism is used to obtain the global feature representation F. global Finally, the local and global features are weighted and fused to obtain the enhanced feature F. enhanced The specific calculation formula is as follows:
[0081] F enhanced =β·F local +(1-β)·F global (1)
[0082] Here, β is the fusion weight of dynamic learning, which can adaptively adjust the contribution ratio of local and global features according to different inputs.
[0083] By introducing the EfficientViT module, this invention not only improves the accuracy of water hyacinth detection, but also ensures the inference speed and computational efficiency of the model, meeting the dual requirements of real-time performance and high precision in practical applications.
[0084] II. Lite-BiFPN Design
[0085] Lite-BiFPN (Lightweight Bidirectional Feature Pyramid Network) is mainly used for multi-scale feature fusion. It can effectively improve the detection effect at different scales for water hyacinth targets of different sizes.
[0086] Multi-scale feature fusion: Water hyacinth targets have significant size differences in different images. Lite-BiFPN processes features at different scales by building a multi-level feature pyramid and enhances the perception ability of high-level features to low-level features through bidirectional information transmission, while maintaining computational efficiency.
[0087] Depthwise Separable Convolution (DSConv): Lite-BiFPN incorporates depthwise separable convolution (DSConv) to reduce computational complexity. DSConv significantly reduces computation by decomposing the convolution operation into two independent operations (depthwise convolution and pointwise convolution) while maintaining feature representation capabilities.
[0088] Suppose that the multi-scale features are represented as F1, F2, ..., F at a certain layer. n Then, bidirectional feature fusion is performed using BiFPN to finally obtain the fused feature F. fused It can be expressed by the following formula:
[0089]
[0090] Among them, w i This represents the weight at each scale. The weights are learned to ensure more efficient feature transfer between different scales.
[0091] III. Optimize the detection head
[0092] The SEAMHead attention mechanism is used to enhance the model's focus on water hyacinth targets, reduce background interference, and thus improve target recognition accuracy.
[0093] Channel attention mechanism: By adaptively adjusting the weights of channels, the model can automatically enhance its attention to important channels (features related to the water hyacinth target) and suppress interference from irrelevant channels.
[0094] Spatial attention mechanism: In the spatial dimension, attention to the image region is dynamically adjusted through the attention mechanism, automatically enhancing the response of the target area of water hyacinth and reducing the influence of the background area.
[0095] In the SEAMHead module, assuming the input feature is F, the weighted feature calculated through channel and spatial attention mechanisms is Fi. att It can be expressed by the following formula:
[0096] Channel attention:
[0097]
[0098] Where σ is the Sigmoid function, and C(F) represents the channel weights output by the channel attention network. These are the characteristics after channel weighting.
[0099] Spatial attention:
[0100]
[0101] Where S(F) is the spatial weight of the spatial attention network output. These are spatially weighted features.
[0102] Finally, through weighted fusion of channels and space, the final attention characteristics are obtained:
[0103]
[0104] Here, α is the fusion factor that adjusts the channel and spatial attention weights. It is a fusion weight dynamically generated by a lightweight multilayer perceptron (Conv1×1→GELU→Conv1×1→Sigmoid), and its value ranges from (0,1).
[0105] Before the SEAMHead module, the C3K2 module is used to further enhance feature extraction capabilities. The design principle of the C3k2 module is as follows:
[0106] The C3k2 module is a commonly used feature extraction module in the YOLO series, and its structure has been optimized in this invention. This module uses a combination of multiple 3×3 convolutional kernels (kernel size = 3) and residual connections.
[0107] F out =σ(W2·σ(W1·F) in +B1)+B2) (6)
[0108] in:
[0109] F in Input feature map; W1 and W2 are convolutional kernel weights, B1 and B2 are bias terms, σ is the activation function (such as SiLU or ReLU); residual connections achieve cross-layer propagation through +B2; F out This is for outputting feature maps.
[0110] The residual term F in the structure in This indicates a direct connection across layers, used to fuse lower-level information and improve training convergence speed and network depth.
[0111] The present invention will now be described in detail with reference to embodiments and accompanying drawings. However, it should be understood that the embodiments and drawings are for illustrative purposes only and do not constitute any limitation on the scope of protection of the present invention. All reasonable modifications and combinations included within the inventive spirit of the present invention fall within the scope of protection of the present invention.
[0112] Example 1
[0113] According to such Figure 1 The process shown in this embodiment, and the specific implementation steps are as follows:
[0114] S1. Data Acquisition and Preprocessing
[0115] 1. Data Acquisition: First, images of water bodies containing water hyacinths were acquired using high-resolution camera equipment (such as a high-definition camera mounted on a drone). To ensure dataset diversity, the acquired images covered various environmental conditions, including different lighting conditions, different water surface states (e.g., calm water, undulating water), and different water hyacinth coverage (e.g., sparse distribution, dense coverage). A total of 3000 images were collected for each scenario to ensure sufficient training data for the deep learning model. The constructed dataset provides fundamental data support for subsequent model training.
[0116] 2. Image Preprocessing: The acquired raw images are orthorectified and stitched together to ensure consistent ground feature scale and reduce distortion caused by shooting angle. They are then cropped to a standard 640×640 pixel image to ensure consistent image size for the input model. This step effectively reduces the impact of image size on the model training process and improves the efficiency of subsequent image processing.
[0117] 3. Data Labeling: Professional labeling tools (such as LabelImg) were used to accurately label the acquired images. Labeling included the species name, location coordinates, and morphological features of the water hyacinth target. To further improve the model's ability to detect water hyacinth targets at different densities and growth stages, image samples were classified according to the water hyacinth's growth morphology: flowering stage, sparse distribution, and dense coverage. The labeled data was divided into training, validation, and test sets in a 7:2:1 ratio to ensure balanced data distribution and the model's generalization ability.
[0118] S2, YOLO11 model optimization: This process is the specific technical implementation path of the WH-YOLO11 model optimization and training process. It involves structurally improving the original YOLO11 model and obtaining a customized model suitable for water hyacinth detection through optimized training.
[0119] The specific procedure for this step is as follows:
[0120] 1. The input image passes through a preliminary convolutional layer for local feature extraction.
[0121] 2. The EfficientViT module further extracts local and global features and dynamically fuses the two types of features to enhance the feature representation capability.
[0122] 3. The Lite-BiFPN module enables further adaptive spatial feature fusion optimization using multi-scale features.
[0123] 4. After further refining and optimizing the feature representation using the C3k2 module, the SEAMHead module is used to introduce channel and spatial attention mechanisms to enhance the feature response of the target region. Finally, the detection prediction convolutional layer generates the location coordinates, category, and confidence information of the water hyacinth target, thus completing the target detection task.
[0124] This embodiment performs joint optimization on the three key steps S21, S22, and S23, which belongs to the structural-level end-to-end optimization design, and is briefly summarized as follows:
[0125] S21 (Backbone Network): Replaces YOLO11 Backbone, introduces EfficientViT, and improves global awareness capabilities;
[0126] S22 (Neck structure): Introduces Lite-BiFPN and sub-modules (such as CARAFE, ASFF, BiFPN_Attn) to enhance feature fusion;
[0127] S23 (Detection Head): Replaces the traditional detection head with C3K2 and SEAMHead to enhance feature extraction and attention mechanisms.
[0128] like Figure 2 The improvement and construction process of the WH-YOLO11 model architecture is as follows:
[0129] S21 and EfficientViT are used in the backbone network to enhance the network's global feature extraction capabilities. Unlike traditional convolutional neural networks, EfficientViT achieves adaptive perception of different regions through a self-attention mechanism, essentially modeling the global correlation between input feature maps. Its core formula is as follows:
[0130]
[0131] in:
[0132] -Q, K, and V represent the query matrix, key matrix, and value matrix of the input features, respectively.
[0133] -d k It represents the dimension of the key matrix, used for scaling and normalization.
[0134] The introduction of EfficientViT enhances the ability to capture global information about water hyacinth targets.
[0135] In this embodiment, the combined use of formulas (1) to (6) is as follows: first, local and global features are integrated, and then the expression is refined to improve the feature representation capability.
[0136] The combined use of formula (7) and formula (2) is as follows: first, global feature capture is performed, and then multi-scale fusion is performed to enhance the detection capability of targets at different scales.
[0137] S22, Lite-BiFPN introduction:
[0138] The Lite-BiFPN multi-scale feature fusion structure is introduced into the Neck part of the original YOLO11 model to improve the information flow and fusion effect of features at different scales, thereby enhancing the model's ability to detect small targets. Compared with traditional BiFPN, Lite-BiFPN reduces computation by using depthwise separable convolution (DSConv), significantly improving inference efficiency while maintaining feature fusion capabilities. Lite-BiFPN employs bidirectional feature flow (top-down and bottom-up), optimizing the feature fusion process through a weighted mechanism, maintaining efficient fusion capabilities while reducing computation.
[0139] The feature fusion formula for Lite-BiFPN is:
[0140]
[0141] in:
[0142] -P i This represents the fused feature map.
[0143] -F j This represents the input multi-scale feature map (feature maps from different levels).
[0144] -w i,j These are learnable weight parameters, referring to the learnable weights of Lite-BiFPN at the i-th layer and from the j-th scale. They are normalized by Softmax and used to adjust the fusion ratio of different input feature maps. They are usually learned through the training process.
[0145] -n is the number of feature maps involved in the fusion.
[0146] S23. Optimized detection head:
[0147] The detection head incorporates the C3K2 and SEAMHead modules. The C3K2 module combines multiple convolutional layers with residual structures to enhance the refinement of feature extraction and improve detection accuracy. The SEAMHead module introduces channel and spatial attention mechanisms, using adaptive learning to optimize the response of key target regions, further improving target recognition accuracy. Its basic formula is:
[0148] F out =σ(W2·σ(W1·F) in +B1)+B2) (6)
[0149] in:
[0150] -W represents the kernel weights.
[0151] -F in It is the input feature map.
[0152] -B is the bias term.
[0153] -σ is the activation function.
[0154] -F out It outputs the feature map.
[0155] The convolutional module replacement is intended to further improve the model's performance in extracting and processing key target features.
[0156] S24, Model Training and Tuning
[0157] 1. Training dataset: The improved YOLO11 network is trained using a labeled dataset, with standard training and validation sets, and the cross-entropy loss function is used to optimize the model parameters.
[0158] 2. Training Process: Stochastic Gradient Descent (SGD) is used as the optimizer, with appropriate learning rate, momentum, and weight decay settings to optimize the network. Data augmentation (such as image rotation, translation, and cropping) is employed to improve the model's generalization ability.
[0159] 3. Parameter Tuning and Optimization: The trained model is tuned using a validation set. Hyperparameters such as the learning rate and batch size are adjusted based on the model's precision and recall on the validation set. Finally, the best-performing model is selected for further testing.
[0160] S3, Testing and Performance Evaluation
[0161] S31 Object Detection: The trained WH-YOLO11 model is applied to test image data to perform object detection tasks, automatically identifying target objects in the images. Evaluation metrics are set as follows:
[0162]
[0163] Precision is the overall accuracy, TP is the number of true positives, and FP is the number of false positives; Recall is the overall recall, and FN is the number of false negatives; mAP@0.5 is the mean of the average precision across all classes with an IoU threshold of 0.5; Ncls The total number of categories participating in the assessment, AP c Let c be the average precision of class c, where c is from 1 to N. cls The natural number.
[0164] S32 was used to evaluate accuracy and compare experimental results.
[0165] like Figure 3 To compare the recall, precision, mAP50, and mAP50 95 curves in the experiment, Table 1 below is an adjusted precision comparison table, highlighting the advantages of the improved model in water hyacinth target detection. The WH-YOLO11 model achieves better performance on specific tasks, especially in complex environments and in the context of multi-scale feature fusion.
[0166] Table 1
[0167]
[0168] As shown in Table 1, the improved YOLO11 (WH-YOLO11) model significantly outperforms the original YOLO11 model in water hyacinth target detection. Specifically:
[0169] The Precision (P) improved by 8.3%, from 80.4% to 88.7%, indicating that the improved model provides higher accuracy and fewer false alarms when identifying water hyacinths.
[0170] The Recall(R) improved by 3.3%, from 82.0% to 85.3%, indicating that the improved model performed better in capturing all water hyacinth targets and the problem of missed detection was alleviated.
[0171] The F1 score improved by 5.7%, from 81.2% to 86.9%, resulting in a significant improvement in overall performance and indicating that the model has achieved a better balance between precision and recall.
[0172] The mAP@0.5 improved by 5.6%, from 85.5% to 91.1%, indicating that the model's detection accuracy was significantly improved in multiple scenarios, especially in complex backgrounds.
[0173] The number of params was reduced from 4.0M to 3.1M, significantly reducing the number of model parameters. This means the improved model is more efficient and can perform inference faster in practical applications.
[0174] The GFLOPs have been reduced from 8.5G to 5.8G, which reduces computational complexity and thus improves inference speed and efficiency, making it more suitable for more complex scenarios and real-time requirements.
[0175] S33 was used to detect the distribution of water hyacinth, and the monitoring results were obtained:
[0176] Using the WH-YOLO11 model, which has been validated for accuracy evaluation, water hyacinth distribution detection is performed on image data from real-world application scenarios. The model outputs the location and distribution information of water hyacinth in the image, achieving effective monitoring results for water hyacinth intrusion in water bodies. Figure 4 .
[0177] In summary, the WH-YOLO11 model in this embodiment significantly improves accuracy and efficiency compared to the original YOLO11 model in the water hyacinth target detection task. In particular, it demonstrates stronger adaptability and higher detection performance in important indicators such as precision, recall, and mAP.
[0178] The above embodiments are merely preferred embodiments of the present invention, and the scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. An improved YOLO11-based method for quickly detecting water hyacinth targets, characterized in that: Includes the following steps: S1. Data Preparation: Image data is collected, covering different weather, lighting and aquatic environments. The collected images are screened, sorted and classified, and data annotation is performed to clarify the category and location information of targets in the images, thus constructing a dataset. The dataset is then subjected to format conversion, size adjustment and enhancement processing, and divided into training set, validation set and test set. S2. Model Optimization and Training: The backbone network of the original YOLO11 model is optimized using EfficientViT, specifically as follows: Assuming the input feature is F, the EfficientViT module first obtains a local feature representation through local convolution , and then obtains a global feature representation through a lightweight self-attention mechanism ; finally, the local feature and the global feature are fused through weighting to obtain an enhanced feature , and the specific calculation formula is as follows: (1) wherein, is a dynamically learned fusion weight that can adaptively adjust the contribution ratio of local and global features according to different inputs; The Lite-BiFPN multi-scale feature fusion structure is introduced into the Neck part of the original YOLO11 model to improve the information flow and fusion effect of features at different scales, thereby enhancing the model's ability to detect small targets. The convolutional modules of the original YOLO11 model are replaced with the SEAMHead convolutional module, which introduces channel and spatial attention mechanisms, to improve the extraction and processing capabilities of key target features, thus constructing the WH-YOLO11 model, as follows: In the SEAMHead module, assuming the input feature is F, the weighted feature calculated through the channel and spatial attention mechanism is , which is expressed by the following formula: Channel attention: (3) wherein, is a Sigmoid function, C(F) represents the channel weights output by the channel attention network, is the channel-weighted feature; Spatial attention: (4) where S(F) is the spatial weight output by the spatial attention network, is the spatially weighted feature; Finally, through weighted fusion of channels and space, the final attention characteristics are obtained: (5) wherein, is a fusion factor for adjusting the channel and spatial attention weights, which is a dynamically generated fusion weight by a lightweight multi-layer perception Conv1x1→GELU→Conv1x1→Sigmoid, with a value range (0, 1); Before the SEAMHead module, the C3K2 module is used to further enhance feature extraction capabilities; the C3K2 module employs a combination of multiple 3×3 convolutional kernels and residual connections. (6) in: : input feature map; and are convolution kernel weights, and are bias terms, is an activation function; residual connection is implemented by + to pass information across layers; is an output feature map; the residual term in the structure represents a direct connection across layers, which is used to integrate low-level information and improve training convergence speed and network depth; The WH-YOLO11 model was trained using preprocessed data, and hyperparameters were adjusted to improve detection performance. S3. Testing and Evaluation: The trained WH-YOLO11 model was used to detect targets in test image data, automatically identifying water hyacinth targets in the images. The performance was evaluated using precision, recall, F1 score, and mAP metrics to quantify the accuracy and reliability of the detection.
2. The method according to claim 1, wherein the improved YOLO11-based quick detection method for water hyacinth targets is characterized in that: Step S2 model optimization and training includes the following steps: S21. Backbone network optimization: The backbone network of the original YOLO11 model is optimized using EfficientViT. By combining local convolution and self-attention mechanism, the ability to extract image features is improved. S22, Neck structure optimization: The Lite-BiFPN multi-scale feature fusion structure is introduced into the Neck part of the YOLO11 model to improve the fusion effect of multi-scale features and enhance the detection capability of small targets; S23. Convolutional module replacement: Replace the original convolutional module of the YOLO11 model with the C3K2 convolutional module and the SEAMHead convolutional module to improve the model's ability to extract key features of the target. S24. Model Training and Tuning: Train the WH-YOLO11 model using the preprocessed dataset from step S1, and improve the model's detection performance and accuracy by adjusting hyperparameters.
3. The method according to claim 2, wherein the improved YOLO11-based quick detection method for water hyacinth targets is characterized by: Step S21, which optimizes the backbone network structure of the original YOLO11 model using EfficientViT, includes: To obtain global features are modeled based on self-attention mechanism, the core computation process of which is shown as follows: (7) wherein Q, K, V represent query matrix, key matrix, value matrix of input features respectively, is the dimension of the key matrix, used for scaling normalization, the above formula realizes the fusion of context information by calculating the attention weight between each position and all positions globally, and the output Z is the global context-aware feature representation.
4. The rapid detection method for water hyacinth targets based on the improved YOLO11 according to claim 3, characterized in that: Step S22 describes the process of introducing the Lite-BiFPN multi-scale feature fusion structure into the Neck part of the original YOLO11 model as follows: Assuming that the multi-scale feature at a certain layer is represented as , bidirectional feature fusion is performed through the BiFPN, and finally the fused feature is obtained, which is represented by the following formula: (2) wherein, denotes the weight at each scale.
5. The method according to claim 1, wherein the improved YOLO11-based quick detection method for water hyacinth targets is characterized by: Step S1, data preparation, includes the following steps: S11. Image Acquisition: Acquire relevant image data containing the target object, covering image samples under different weather conditions, lighting conditions, and different aquatic environments; S12. Constructing a dataset: Filter and organize the collected image data, and establish a dataset according to certain standards; S13. Data annotation: Use professional annotation tools to annotate the images in the dataset, clarify the location and category information of the targets in each image, and generate annotated image data that meets the training requirements; S14. Image preprocessing: Perform preprocessing operations on the labeled image data, including format conversion, size adjustment, and data augmentation, to improve data quality and the effectiveness of model training. The processed data is used as the input to the model.
6. The rapid detection method for water hyacinth targets based on the improved YOLO11 according to any one of claims 1 to 5, characterized in that: Step S3, testing and evaluation, includes the following steps: S31. Object Detection: The trained WH-YOLO11 model is applied to the test image data to perform object detection tasks, automatically identifying target objects in the images. The evaluation metrics are set as follows: (9) (10) (11) Precision is the accuracy, TP is the number of true positives, and FP is the number of false positives; Recall is the recall, and FN is the number of false negatives; mAP@0.5 is the mean of the average precision across all classes with an IoU threshold of 0.
5. The total number of categories participating in the evaluation. Let c be the average precision of class c, where c is 1 to 1. _n_ natural numbers; S32. Precision Evaluation: The performance of the target detection results is evaluated using precision, recall, F1 score and mean precision as evaluation indicators, which quantifies the accuracy and reliability of the detection. S33. Water hyacinth distribution detection: Using the WH-YOLO11 model, which has been verified by accuracy evaluation, water hyacinth distribution detection is performed on image data in actual application scenarios. The location distribution information of water hyacinth in the image is output to achieve effective monitoring of water hyacinth invasion in water areas.
Citation Information
Patent Citations
Method for detecting alien invasive plants based on improved YOLOv9 algorithm
CN119399688A
Infrared ship target detection method based on improved yolov7
CN117152691A
Target detection method based on improved YOLOv5s network model
CN117372684A