Aquatic floating plant detection method and device, electronic equipment and storage medium

By improving the RT-DETR model to the RT-DETR-LSK model, the problem of large parameters is solved, and efficient aquatic floating plants detection on memory-limited equipment is realized, which is suitable for cleaning of floating objects in waters.

CN120431456APending Publication Date: 2025-08-05GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510315498.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Due to the large amount of parameters, the existing RT-DETR algorithm is difficult to be applied to the detection of aquatic floating plants on embedded devices with small memory, resulting in an error or crash in the device.

Method used

The LSKblock module was introduced to improve the RT-DETR model, and the RT-DETR-LSK model can be replaced by deep separation convolution and hollow convolution, reducing the amount of parameters, and dynamically fusion branch results to capture the characteristics of different scales, and combining local and global information to build the RT-DETR-LSK model.

Benefits of technology

The number of parameters is reduced, making the model easy to deploy in embedded devices with smaller memory, maintaining high accuracy, and is suitable for detection and cleaning of aquatic floating plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431456A_ABST
    Figure CN120431456A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of visual target detection, and provides an aquatic floating plant detection method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining an image of an aquatic floating plant; constructing a data set according to the image; an LSKblock module is introduced to improve the RT-DETR model, and an RT-DETR-LSK model is obtained; the data set is input into an RT-DETR-LSK model for training and evaluation, and an aquatic floating plant detection model is obtained; and inputting the data set into an aquatic floating plant detection model for testing to obtain an aquatic floating plant detection map. According to the scheme provided by the invention, the RT-DETR model is enabled to maintain relatively high accuracy while performing multi-scale feature extraction to reduce the parameter quantity, so that the RT-DETR model can be easily deployed in embedded equipment with a relatively small memory, and the RT-DETR algorithm can be well applied to detection of aquatic floating plants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual target detection, and in particular relates to a method, device, electronic equipment and storage medium for detecting aquatic floating plants. Background Art

[0002] Currently, the removal of aquatic floating plants from rivers and lakes utilizes various detection technologies based on deep learning and computer vision. Equipment then removes aquatic floating plants based on the detection results, thereby achieving river and lake management and protection. YOLO technologies are widely used due to their fast detection speed. However, YOLO has poor detection capabilities for unknown plants and in complex scenarios. Furthermore, its processing is complex, requiring region proposals followed by classification and bounding box regression, resulting in low real-time performance.

[0003] Therefore, in order to address such situations, the existing technology uses the RT-DETR algorithm instead of the YOLO algorithm for detection. The RT-DETR algorithm has higher accuracy and speed than YOLO, and has stronger generalization ability. However, the existing RT-DETR algorithm has the problem of a large number of parameters. The large number of parameters will cause the device to report errors or crash directly when the algorithm is subsequently deployed on an embedded device with smaller memory for detection, which will make the algorithm unable to be well applied to the detection of aquatic floating plants. Summary of the Invention

[0004] The present invention provides a method to solve the problem that the RT-DETR algorithm in the prior art is difficult to apply to aquatic floating plant detection due to the large number of parameters.

[0005] In order to solve the above technical problems, in a first aspect, the present invention provides a method for detecting aquatic floating plants, the method comprising:

[0006] Acquire images of aquatic floating plants;

[0007] constructing a data set based on the images;

[0008] The LSKblock module is introduced to improve the RT-DETR model to obtain the RT-DETR-LSK model; the LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer, and an average pooling layer;

[0009] Inputting the data set into the RT-DETR-LSK model for training and evaluation to obtain an aquatic floating plant detection model;

[0010] The data set is input into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection map.

[0011] Optionally, the method of introducing the LSKblock module to improve the RT-DETR model includes:

[0012] Construct an LSK-attention module based on the LSKblock module, wherein the LSK-attention module includes a two-dimensional convolutional layer, an activation function, and the LSKBlock module;

[0013] Constructing a Block module based on the LSK-attention module, wherein the Block module includes a normalization layer, a regularization layer, a multi-layer perceptron, and the LSK-attention module;

[0014] The backbone network of the RT-DETR-LSK model is constructed based on the Block module, and the backbone network includes the OPE module and the Block module.

[0015] Optionally, the step of constructing a data set based on the image includes:

[0016] The image is annotated using Labeling software to obtain annotation information and save it in YOLO format to obtain the dataset, wherein the annotation includes the coordinate position information of the object in the image and the name of the target;

[0017] The data set is randomly divided according to a proportion to obtain a training set, a validation set and a test set, and the test set is used to input into the aquatic floating plant detection model for testing.

[0018] Optionally, the training and evaluation steps of the RT-DETR-LSK model include:

[0019] Setting training hyperparameters and performing data augmentation on the training set;

[0020] The data-enhanced training set is input into the RT-DETR-LSK model for training;

[0021] Evaluate the trained RT-DETR-LSK model using the validation set to obtain an evaluation result;

[0022] Determine whether the evaluation result meets the evaluation criteria. If not, adjust the hyperparameters according to the evaluation result and repeat the above training process. If yes, select the optimal model parameters.

[0023] The optimal model parameters are output as the aquatic floating plant detection model.

[0024] Optionally, the evaluation indicators are: using average precision mAP50 and mAP50-95, parameter quantity Params, detection speed FPS and parameter operation quantity GFLOPs as evaluation indicators to evaluate the trained RT-DETR-LSK model.

[0025] In a second aspect, the present invention provides an aquatic floating plant detection device, comprising:

[0026] an image acquisition unit, for acquiring images of aquatic floating plants;

[0027] A data set construction unit, configured to construct a data set based on the image;

[0028] A model improvement unit is used to improve the RT-DETR model by introducing the LSKblock module to obtain the RT-DETR-LSK model; the LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer, and an average pooling layer;

[0029] A model training unit, configured to input the data set into the RT-DETR-LSK model for training and evaluation to obtain an aquatic floating plant detection model;

[0030] The image testing unit is used to input the data set into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection image.

[0031] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned aquatic floating plant detection method when executing the program.

[0032] In a fourth aspect, the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the steps of the above-mentioned aquatic floating plant detection method.

[0033] In the technical solution of the present invention, the traditional RT-DETR model is improved by introducing the LSKbolck module. Since the improved RT-DETR-LSK model decomposes the standard convolution into a depth-separable convolution, replaces the large kernel with the hole convolution, and dynamically fuses the results of the two branches, the parameter doubling caused by the parallel branches is reduced; the problem of large parameter quantity is solved; at the same time, it can capture features of different scales, and adaptively adjust the receptive field of the network by selectively using convolution kernels of different sizes to better capture targets of different sizes, combine local spatial information with global information, emphasize the capture of local features, and not rely solely on global information to achieve effective capture of detail features. The RT-DETR-LSK model of the present invention reduces the parameter quantity while performing multi-scale feature extraction to maintain accuracy, making it easy to deploy in embedded devices with smaller memory, and then deployed on a fully automatic cleaning robot for floating objects in water areas, so that it can be well applied to the detection and cleaning of aquatic floating plants. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0035] Figure 1 This is a flow chart of an embodiment of a method for detecting aquatic floating plants according to the present invention;

[0036] Figure 2 for Figure 1 Flowchart of a specific implementation of step S130 in the process

[0037] Figure 3 This is a flow chart of another embodiment of the aquatic floating plant detection method of the present invention;

[0038] Figure 4 It is a structural diagram of the traditional RT-DETR model in the prior art;

[0039] Figure 5 Schematic diagram of the structure of the RT-DETR-LSK model of the present invention;

[0040] Figure 6 for Figure 5 Schematic diagram of the structure of OPE module and Block module;

[0041] Figure 7 for Figure 6Schematic diagram of the structure of the LSK-attention module and the MLP module in the Block module;

[0042] Figure 8 for Figure 7 Schematic diagram of the structure of the LSKblock module in the LSK-attention module

[0043] Figure 9 1. This is a diagram showing the detection training results of the RT-DETR-LSK model in one embodiment of the present invention;

[0044] Figure 10 This is a detection effect diagram of an aquatic floating plant detection model in one embodiment of the present invention;

[0045] Figure 11 This is a schematic structural diagram of an aquatic floating plant detection device according to an embodiment of the present invention;

[0046] Figure 12 FIG. 1 is a schematic structural diagram of an electronic device according to an embodiment of the present invention.

[0047] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this invention belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the accompanying drawings are intended to cover non-exclusive inclusions. The terms "first" and "second" in the specification and claims of the present invention and the accompanying drawings are used to distinguish different objects, not to describe a specific order.

[0049] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0050] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0051] The English names and abbreviations used in the present invention are explained as follows:

[0052] RT-DETR: Real-Time Detection Transformer, a real-time object detector based on Vision Transformer;

[0053] Backbone: backbone network, one of the components of the RT-DETR model;

[0054] Head: Head network, which is one of the components of the RT-DETR model;

[0055] RTDETRDecoder: detection head, one of the components of the RT-DETR model;

[0056] Conv: Convolutional Layer module, which is a standard convolution module;

[0057] BasicBlock: The basic residual block in the RseNet network, used to extract image features;

[0058] OPE: Overlap Patch Embed module, an image processing module in deep learning, is used to split the input image into multiple overlapping patches and convert each patch into a corresponding feature vector representation;

[0059] UpSample: Feature upsampling module, used to upsample feature maps;

[0060] AIFI: Attention-based Intrascale Feature Interaction module, a component of the RT-DETR model;

[0061] RepC3: A structure built on RepConv, consisting of multiple RepConv modules;

[0062] Concat: Concatenate Module, a feature concatenation module used to concatenate feature maps;

[0063] CCFM: Cross-Scale Feature Fusion Module, a cross-scale feature fusion module, is a feature fusion technology used in target detection models.

[0064] Example 1

[0065] The RT-DETR model mainly consists of a backbone network (Backbone) and a head network (Head). The head network (Head) consists of a hybrid encoder (Hybrid Encoder) and a decoder (RTDETRDecoder). When in use, the backbone network is the feature extraction stage, responsible for extracting image features; the hybrid encoder is the feature fusion stage, processing multi-scale features by decoupling intra-scale interactions and cross-scale fusion; the decoder is the prediction and optimization stage, responsible for generating prediction results and optimizing them through the auxiliary prediction head.

[0066] like Figure 4 As shown in the figure, the backbone network in the traditional RT-DETR model is mainly composed of the Conv module and the BasicBlock module, where the BasicBlock module includes two convolutional layers, a batch normalization layer and a residual connection; since the backbone network of the traditional RT-DETR model adopts a deeper network structure, the deep convolutional network includes multiple convolutional layers, a normalization layer and an activation layer, etc., it relies on the stacking of a large number of standard 3×3 convolutional kernels. The parameters of each layer will increase exponentially with the increase of network depth. A large number of convolutional kernel parameters will greatly increase the number of parameters of the backbone network, and it only relies on shallow convolution information, resulting in a limited understanding of the global context. It only relies on the local receptive field of the convolutional layer to extract features, and may ignore the effective fusion of local details.

[0067] like Figures 1 to 12 As shown, this embodiment provides a method for detecting aquatic floating plants, comprising the following steps:

[0068] S110: Acquire an image of aquatic floating plants;

[0069] S120: constructing a dataset based on the image;

[0070] S130: Introduce the LSKblock module to improve the RT-DETR model to obtain the RT-DETR-LSK model, which specifically includes the following steps:

[0071] S131: Construct LSK-attention module based on LSKblock module;

[0072] S132: Build a Block module based on the LSK-attention module;

[0073] S133: Construct the backbone network of the RT-DETR-LSK model based on the Block module;

[0074] Specifically, the RT-DETR-LSK model of the present invention improves the traditional RT-DETR model by introducing the LSKblock module, and reconstructs a backbone network based on the LSKblock module. The newly constructed backbone network is defined as LSKNet, which serves as the backbone network of the RT-DETR-LSK model.

[0075] Among them, the LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer and an average pooling layer;

[0076] The LSK-attention module consists of two 1*1 two-dimensional convolutions, a GELU activation function, and an LSKBlock module;

[0077] The Block module includes a batch normalization layer, an LSK-attention module, an MLP multi-layer perceptron, and a regularization layer;

[0078] LSKNet includes 4 OPE modules and 3 Block modules;

[0079] Through the above steps, the traditional RT-DETR model is transformed to obtain the RT-DETR-LSK model. The structure of the RT-DETR-LSK model is as follows: Figure 5 shown.

[0080] S140: Input the dataset into the RT-DETR-LSK model for training and evaluation to obtain an aquatic floating plant detection model;

[0081] S150: Inputting the data set into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection map.

[0082] In the technical solution of the present invention, the traditional RT-DETR model is improved by introducing the LSKbolck module. Since the improved RT-DETR-LSK model decomposes the standard convolution into a depth-separable convolution, replaces the large kernel with the hole convolution, and dynamically fuses the results of the two branches, the parameter doubling caused by the parallel branches is reduced; the problem of large parameter quantity is solved; at the same time, it can capture features of different scales, and adaptively adjust the receptive field of the network by selectively using convolution kernels of different sizes to better capture targets of different sizes, combine local spatial information with global information, emphasize the capture of local features, and not rely solely on global information to achieve effective capture of detail features. The RT-DETR-LSK model of the present invention reduces the parameter quantity while performing multi-scale feature extraction to maintain accuracy, making it easy to deploy in embedded devices with smaller memory, and then deployed on a fully automatic cleaning robot for floating objects in water areas, so that it can be well applied to the detection and cleaning of aquatic floating plants.

[0083] Example 2

[0084] like Figures 1 to 12 As shown, this embodiment provides a method for detecting aquatic floating plants. In this embodiment, it is used for detecting water hyacinth, and specifically includes the following steps:

[0085] S210: Acquire an image of water hyacinth;

[0086] Images of water hyacinths are captured and intercepted using an industrial camera, or obtained through an online image database, data platform, etc. It is understood that this embodiment uses water hyacinth detection as an example only; in other embodiments, other aquatic floating plants or other visual detection targets suitable for the method of the present invention may be used, and this is not limited here.

[0087] S220: Constructing a dataset based on the image, specifically including the following steps:

[0088] S221: Label the image using Labeling software, obtain the labeling information and save it in YOLO format to obtain a data set. The labeling includes the coordinate position information of the water hyacinth in the image and the name of the target.

[0089] Specifically, the steps for labeling through Labeling are as follows: (1) open the virtual Python environment; (2) import the dataset image into the Labeling software and select the YOLO labeling mode; (3) frame each target in each image, and each rectangular frame contains the coordinates and length and width information of the target; (4) add a target name to each target, which includes Water Hyacinth and other in this embodiment. It can be understood that the labeled target name depends on the detection object. Other refers to other plant objects except the target (water hyacinth), which is not limited here; (5) store the labeled information content in the form of .txt, and finally save the dataset in YOLO format and exit.

[0090] S222: Divide the data set randomly according to the proportion to obtain a training set, a validation set, and a test set.

[0091] Specifically, the images in the data set are randomly divided into a training set, a validation set, and a test set in a ratio of 7:2:1. It can be understood that the division ratio of the data set is only preferred in this embodiment and can be adjusted to other ratios according to actual conditions, which is not limited here.

[0092] S230: Introduce the LSKblock module to improve the RT-DETR model to obtain the RT-DETR-LSK model, which specifically includes the following steps:

[0093] S231: Construct LSK-attention module based on LSKblock module;

[0094] S232: Build a Block module based on the LSK-attention module;

[0095] S233: Construct the backbone network of the RT-DETR-LSK model based on the Block module;

[0096] Specifically, such as Figures 5 to 8 As shown, the backbone network of the improved RT-DETR-LSK model in this embodiment includes:

[0097] 4 OPE modules, the OPE module includes a convolution layer and a normalization layer, which are used to convert the input image into patches and embed them; and

[0098] 9 integrated network basic blocks Block modules, Block modules include batch normalization layer, LSK-attention module, MLP multi-layer perceptron and regularization layer; It should be noted that, Figure XThe backbone network in the example includes three Block modules. In actual operation, the three Block modules can be regarded as having three stages, each of which is repeated multiple times, with a default value of three times. Therefore, in this embodiment, it can be understood as having nine Block modules.

[0099] Among them, the LSK-attention module contains two 1*1 two-dimensional convolutions, a GELU activation function and an LSKblock module;

[0100] Among them, the LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, one maximum pooling layer and one average pooling layer.

[0101] The head network (Head) of the RT-DETR-LSK model includes: five 1*1 two-dimensional convolutions, two UpSample modules, an AIFI feature fusion module, four RepC3 modules, four Concat modules, two 3*2 convolutions and an RTDETRDecoder decoder.

[0102] After completing the above steps, the traditional RT-DETR model is improved to obtain the RT-DETR-LSK model.

[0103] S240: Input the dataset into the RT-DETR-LSK model for training and evaluation to obtain an aquatic floating plant detection model, which specifically includes the following steps:

[0104] S241: Set training hyperparameters and perform data augmentation on the training set;

[0105] Specifically, we set training hyperparameters, including the learning rate (used to control the step size of each parameter update), batch size (the number of training samples used for each parameter update), and training rounds (the number of times the dataset is fully trained). We then converted the input training set images into RGB three-channel format, and used Mosaic data augmentation and Cutout data augmentation to increase the complexity of the dataset. Finally, we cropped the expanded dataset images to a uniform size of 640*640.

[0106] Furthermore, in this embodiment, the initial learning rate Ir0 is set to 0.01, the final learning rate Irf is set to 0.1, the batch size is set to 64, and the number of epochs is set to 300. It should be noted that in some other embodiments, the above parameters can be set to other values. This is only a preferred embodiment of this embodiment and is not limited here.

[0107] S242: Input the data-enhanced training set into the RT-DETR-LSK model for training;

[0108] Specifically, in this embodiment, the following steps are included:

[0109] S2421: Input the processed image, i.e., the image cropped to 640*640 size, into the first OPE module in the backbone network of the RT-DETR-LSK model. The input image is divided into 7×7 patches with a stride of 4. Each patch is embedded in a 32-dimensional feature space, and the output feature map is 160*160*32 in size.

[0110] S2422: The feature map of size 160*160*32 is input into the three Block modules of the first stage of the backbone network. It first passes through the normalization layer to standardize the distribution of the feature map, then enters the LSK-attention module for feature extraction, and then enters a feature scaling layer to adjust the feature strength. Finally, the scaled feature is residually connected with the original feature map;

[0111] S2423: The feature map after the residual connection is normalized again through the normalization layer, and a nonlinear transformation is performed through the multi-layer perceptron to enhance the expressiveness of the features. The output of the multi-layer perceptron is then input to another feature scaling layer, and the scaled features are residually connected with the normalized feature map. Finally, a feature map with a size of 160*160*32 is output. The multi-layer perceptron includes two 1*1 convolutions and one GELU activation function. The GELU activation function is specifically: GELU(x)=x*Φ(x); wherein Φ(x) is the cumulative distribution function (CDF) of the standard normal distribution, defined as: Where erf(x) is the error function, defined as:

[0112] It should be noted that, in some other embodiments, the activation function may also adopt other activation functions, such as the ReLU activation function, etc. The GELU activation function is only preferred in this embodiment and is not limited here;

[0113] S2424: The feature map of size 160*160*32 is input into the second OPE module and split into 3×3 patches with a stride of 2. Each patch is embedded in a 64-dimensional feature space. The output 64-dimensional feature map passes through the three Block modules of the second stage, and finally outputs a feature map of size 80*80*64.

[0114] S2425: The feature map of size 80*80*64 is input into the third OPE module and split into 3×3 patches with a stride of 2. Each patch is embedded into a 128-dimensional feature space. The output 128-dimensional feature map passes through the three Block modules of the third stage, and finally outputs a feature map of size 40*40*128.

[0115] S2426: The feature map of size 40*40*128 is input into the fourth OPE module and split into 3×3 patches with a stride of 2. Each patch is embedded into a 256-dimensional feature space. The output 256-dimensional feature map passes through the three Block modules of the fourth stage, and finally outputs a feature map of size 20*20*256.

[0116] S2427: The output of step S2426 is used as the input of the AIFI module. After intra-scale interaction, F6 is obtained. The feature maps output by steps S2424 and S2425 (hereinafter referred to as F4 and F5) and F6 are used as the input of the CCFM module for multi-scale feature fusion. The image feature F6 undergoes 1*1 convolution, normalization, and activation function calculation, and is used together with F4 as the input of the fusion block. After fusion, the first output of the fusion block (the first fusion result) is obtained. The fusion result is transmitted to the detection head (RTDETRDecoder). It is supplemented that after entering the detection head, the bounding box coordinates and category probability distribution are generated for each target. The mechanism of the fusion block is as follows: features of different scales are used as the input of the fusion block and convolved with a 1*1 convolution kernel. The new feature map is fused by element-wise addition and used as the output of the fusion block.

[0117] It should be noted that the above modules are arranged in order from top to bottom and from left to right according to the RT-DETR-LSK model network architecture.

[0118] S243: Evaluate the trained RT-DETR-LSK model using the validation set to obtain evaluation results;

[0119] Specifically, the validation set images are input into the trained RT-DETR-LSK model, and the average precision mAP50 and mAP50-95, parameter amount Params, detection speed FPS and parameter operation amount GFLOPs are used as evaluation indicators to evaluate the trained RT-DETR-LSK. In this embodiment, four experiments are designed for analysis and comparison. Each group of experiments uses the same data set, training hyperparameters and pre-training weights, with a training cycle of 300, a learning rate of 0.01, an IOU (Intersection Over Union) threshold of 0.5, and an optimizer of Adam. The specific calculation of each evaluation indicator is as follows:

[0120] Accuracy calculation formula:

[0121]

[0122] Precision represents the proportion of positive examples among all instances detected as positive by the model. It is used to measure the accuracy of the model, that is, how many of the positive examples predicted by the model are truly positive examples. TP (True Positives) refers to true positives, which are positive examples correctly detected by the model; FP (False Positives) refers to false positives, which are positive examples incorrectly detected by the model.

[0123] Recall calculation formula:

[0124]

[0125] Among them, the recall rate (Recall) represents the proportion of instances that are actually positive examples that are correctly detected by the model. It is used to measure the sensitivity of the model, that is, the model's ability to detect actual targets; FN (False Negatives): False negatives are positive examples that the model fails to detect, that is, they are actually targets but the model fails to identify them.

[0126] The formula for calculating average accuracy is:

[0127]

[0128] Among them, P(R) is the Precision-Recall curve, which means that for each category, the precision-recall curve is drawn under different IOU thresholds; the average precision AP (Average Precision) means that for each category, the area under the PR curve is calculated; the evaluation model needs to comprehensively consider P and R, and the PR curve (P is the vertical axis and R is the horizontal axis) is selected to represent the average accuracy AP.

[0129] Among them, mAP50 represents the average precision calculated when the IOU threshold is 0.5; mAP50-95 represents the average precision calculated under multiple IOU thresholds (from 0.5 to 0.95, with an interval of 0.05); N is the number of categories; mAP is the average of the average precision AP of multiple categories. The larger the value, the higher the overall accuracy of the model.

[0130] Among them, IOU (Intersection over Union) is a commonly used evaluation metric in the field of target detection, which is used to measure the degree of overlap between the model detection results and the actual target position. Its expression is as follows:

[0131]

[0132] Among them, A represents the real object box in the dataset; B represents the predicted box after model detection.

[0133] Parameter calculation formula:

[0134] Parameters = C0×(k w ×k h ×C i +1)

[0135] Among them, the number of parameters (Parameters) represents the total number of parameters in the model that need to be trained and optimized.

[0136] Detection speed FPS calculation formula:

[0137]

[0138] Among them, the detection speed FPS (Frames Per Second) indicates the number of image frames that the model can process per second, which is used to measure the real-time performance of the model.

[0139] Parameter operation calculation formula:

[0140] FLOPs = ∑(Layer FLOPs);

[0141]

[0142] The parameter GFLOPs (Giga Floating Point Operations per Second) represents the number of floating-point operations performed per second, in units of one billion floating-point operations.

[0143] S244: Determine whether the evaluation results meet the evaluation criteria. If not, adjust the hyperparameters based on the evaluation results and repeat the above training process steps S242 to S243. If yes, select the optimal model parameters and proceed to step S245.

[0144] S245: Outputting the optimal model parameters as the aquatic floating plant detection model.

[0145] Specifically, such as Figure 9As shown in the figure, when the curve converges and stabilizes, the model parameters meet the evaluation criteria. These parameters are selected as the optimal model parameters and output as the aquatic floating plant detection model. As can be seen from the figure, the aquatic floating plant detection model achieves an mAP50 of 0.74, an mAP50-95 of 56.5, 12.56 parameters, 37.5 GFLOPs, and a detection speed of 116 FPS. The train / giou_loss function reflects the model's gradual improvement in positioning accuracy during training, the train / cls_loss function reflects the model's gradual improvement in classification performance, and the train / l1_loss function reflects the model's gradual improvement in regression performance. Lower losses indicate better training results.

[0146] In this embodiment, the deep learning framework is PyTorch, and the implementation hardware conditions and parameters are: GPU model is NVIDIA RTX A6000 (48GB); CUDA version is CUDA11.1; operating system is Ubuntu20.04.3LTS; Python version is Python3.8; deep learning framework is Pytorch 1.10.1.

[0147] S250: Input the test set into the aquatic floating plant detection model for testing to obtain a water hyacinth detection image.

[0148] By using the aquatic floating plant detection model, the water hyacinth image can be detected, and the detection effect is as follows: Figure 10 As shown, it can be seen from the detection results that the detection accuracy and the fitting degree of the drawn boxes are relatively high; it can be understood that the obtained aquatic floating plant detection model can be deployed in embedded devices with smaller memory to be applied to the detection of aquatic floating plants.

[0149] Furthermore, the aquatic floating plant detection model of the present invention was compared with the RT-DETR-r18, RT-DETR-r34, and RT-DETR-r50 models. The test results are shown in the following table. It can be seen that the aquatic floating plant detection model of the present invention can maintain a high accuracy while reducing the number of parameters.

[0150] Model mAP50 mAP50-95 parameters / M GFLOPs FPS RT-DETR-r18 74.2 55.2 19.87 56.9 145 RT-DETR-r34 73.4 53.9 31.10 88.8 112 RT-DETR-r50 70.4 53.4 41.95 129.5 124 RT-DETR-LSK 74.7 56.5 12.56 37.5 116

[0151] Example 3

[0152] like Figure 11 FIG. 1 is a schematic diagram showing the structure of an embodiment of the aquatic floating plant detection device of the present invention. Figure 1The present embodiment provides an aquatic floating plant detection device, which can be applied to various devices for aquatic floating plant detection. Figure 1 Corresponding to the method embodiment shown, the aquatic floating plant detection device includes: an image acquisition unit 310, a data set construction unit 320, a model improvement unit 330, a model training unit 340, and an image testing unit 350. Each module is connected through a data transmission interface to achieve data circulation and sharing.

[0153] An image acquisition unit 310 is used to acquire images of aquatic floating plants;

[0154] A data set construction unit 320 is used to construct a data set based on the image;

[0155] The model improvement unit 330 is used to introduce the LSKblock module to improve the RT-DETR model to obtain the RT-DETR-LSK model; the LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer, and an average pooling layer;

[0156] The model training unit 340 is used to input the data set into the RT-DETR-LSK model for training and evaluation to obtain an aquatic floating plant detection model;

[0157] The image testing unit 350 is used to input the data set into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection image.

[0158] The aquatic floating plant detection device of the embodiment of the present invention and the above-mentioned aquatic floating plant detection method can refer to each other, and its beneficial effects are equivalent to the beneficial effects of the above-mentioned aquatic floating plant detection method, which will not be repeated here.

[0159] Example 4

[0160] The embodiment of the present invention further provides an electronic device for executing the above-mentioned aquatic floating plant detection method of embodiment 1 and / or embodiment 2, such as Figure 12 As shown, the electronic device includes a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 can call the logic instructions in the memory 430 to execute the aquatic floating plant detection method, which includes:

[0161] Acquire images of aquatic floating plants;

[0162] Build a dataset based on images;

[0163] The LSKblock module is introduced to improve the RT-DETR model to obtain the RT-DETR-LSK model. The LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer, and an average pooling layer.

[0164] The dataset is input into the RT-DETR-LSK model for training and evaluation to obtain the aquatic floating plant detection model;

[0165] The dataset is input into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection map.

[0166] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0167] The beneficial effects of the electronic device of the embodiment of the present invention are equivalent to the beneficial effects of the above-mentioned aquatic floating plant detection method, and will not be repeated here.

[0168] Example 5

[0169] An embodiment of the present invention further provides a computer-readable storage medium, the computer-readable storage medium including a stored program, wherein when the program is executed, the device containing the computer-readable storage medium is controlled to execute the above-mentioned aquatic floating plant detection method, the method comprising:

[0170] Acquire images of aquatic floating plants;

[0171] Build a dataset based on images;

[0172] The LSKblock module is introduced to improve the RT-DETR model to obtain the RT-DETR-LSK model. The LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer, and an average pooling layer.

[0173] The dataset is input into the RT-DETR-LSK model for training and evaluation to obtain the aquatic floating plant detection model;

[0174] The dataset is input into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection map.

[0175] The beneficial effects of the computer-readable storage medium of the present invention are equivalent to the beneficial effects of the above-mentioned aquatic floating plant detection method, and will not be described in detail here.

[0176] The invention is operational with numerous general purpose or special purpose computer system environments or configurations.

[0177] For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, etc.

[0178] The invention may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer.

[0179] Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network.

[0180] In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.

[0181] Specifically, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0182] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps and they may be performed in other orders.

[0183] Moreover, at least part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0184] Obviously, the embodiments described above are only some of the embodiments of the present invention, rather than all of them. The accompanying drawings provide preferred embodiments of the present invention, but do not limit the scope of the present invention. The present invention can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to facilitate a more thorough and comprehensive understanding of the disclosure of the present invention.

[0185] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art may still modify the technical solutions described in the aforementioned specific embodiments or replace some of the technical features therein with equivalents. Any equivalent structure made using the contents of the present invention's description and drawings, directly or indirectly applied to other related technical fields, shall also fall within the scope of protection of the present invention.

Claims

1. A method for detecting aquatic floating plants, characterized in that: include: Acquire images of aquatic floating plants; constructing a data set based on the images; The LSKblock module is introduced to improve the RT-DETR model to obtain the RT-DETR-LSK model. The LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer, and an average pooling layer. Inputting the data set into the RT-DETR-LSK model for training and evaluation to obtain an aquatic floating plant detection model; The data set is input into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection map.

2. The aquatic floating plant detection method according to claim 1, wherein: The method of introducing the LSKblock module to improve the RT-DETR model includes: Construct an LSK-attention module based on the LSKblock module, wherein the LSK-attention module includes a two-dimensional convolutional layer, an activation function, and the LSKBlock module; Constructing a Block module based on the LSK-attention module, wherein the Block module includes a normalization layer, a regularization layer, a multi-layer perceptron, and the LSK-attention module; The backbone network of the RT-DETR-LSK model is constructed based on the Block module, and the backbone network includes the OPE module and the Block module.

3. The aquatic floating plant detection method according to claim 1 or 2, characterized in that: The step of constructing a data set according to the image comprises: The image is annotated using Labeling software to obtain annotation information and save it in YOLO format to obtain the dataset, wherein the annotation includes the coordinate position information of the object in the image and the name of the target; The data set is randomly divided according to a proportion to obtain a training set, a validation set and a test set, and the test set is used to input into the aquatic floating plant detection model for testing.

4. The aquatic floating plant detection method according to claim 3, wherein: The training and evaluation steps of the RT-DETR-LSK model include: Setting training hyperparameters and performing data augmentation on the training set; The data-enhanced training set is input into the RT-DETR-LSK model for training; Evaluate the trained RT-DETR-LSK model using the validation set to obtain an evaluation result; Determine whether the evaluation result meets the evaluation criteria. If not, adjust the hyperparameters according to the evaluation result and repeat the above training process. If yes, select the optimal model parameters. The optimal model parameters are output as the aquatic floating plant detection model.

5. The aquatic floating plant detection method according to claim 4, wherein: The evaluation indicators are: using the average precision mAP50 and mAP50-95, parameter quantity Params, detection speed FPS and parameter operation quantity GFLOPs as evaluation indicators to evaluate the trained RT-DETR-LSK model.

6. An aquatic floating plant detection device, characterized in that: include: an image acquisition unit, for acquiring images of aquatic floating plants; A data set construction unit, configured to construct a data set based on the image; A model improvement unit is used to improve the RT-DETR model by introducing the LSKblock module to obtain the RT-DETR-LSK model; the LSKblock module includes a 5*5 two-dimensional convolution, three 7*7 two-dimensional convolutions, three 1*1 two-dimensional convolutions, a maximum pooling layer, and an average pooling layer; A model training unit, configured to input the data set into the RT-DETR-LSK model for training and evaluation to obtain an aquatic floating plant detection model; The image testing unit is used to input the data set into the aquatic floating plant detection model for testing to obtain an aquatic floating plant detection image.

7. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the aquatic floating plant detection method according to any one of claims 1 to 5 when executing the program.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the steps of the aquatic floating plant detection method according to any one of claims 1 to 5.