Insect target detection method and device based on improved YOLOv9 model
By improving the YOLOv9 model, the lightweight architecture and feature fusion module are adopted to solve the problems of large amount of computation and dependence on high-performance devices in insect target detection tasks, and efficient and accurate insect target detection is achieved, suitable for resource-constrained scenarios.
Patent Information
- Application Number
- CN202510009944.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, insect target detection tasks have problems such as large calculation volume, high parameters, long time-consuming, and strong dependence on high-performance computing devices. Especially in scenarios where resource limitations are not allowed to fully utilize the potential of the detection model.
By improving the YOLOv9 model, lightweight Starnet is used to replace the original backbone, the LFRepBlock module is designed to replace part of the RepNCSPELAN4 module, and GSConv is used to replace part of the convolutional layer, reducing redundant calculations and ensuring the transmission and fusion of feature information.
It effectively reduces the parameter quantity and calculation cost of the model, improves detection efficiency and accuracy, reduces missed detection and missed pickup. It is suitable for resource-constrained scenarios and has significant real-time detection advantages.
Smart Images

Figure CN120070943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and particularly to an insect object detection method, device, equipment and medium based on an improved YOLOv9 model. Background Art
[0002] The biodiversity of soil animals is a key factor in maintaining soil health and ecosystem functions. These tiny organisms play an irreplaceable role in decomposing organic matter, nutrient cycling, improving soil structure and promoting plant growth. Scientists have confirmed the huge biodiversity and important ecosystem functions of soil animals. However, the monitoring, investigation, classification, analysis and statistics of soil animals still pose a huge challenge to humans.
[0003] AI technology can provide great assistance for this professional, time-consuming and laborious work. By means of certain learning and training and establishing a suitable algorithm model, as long as the computer scans animal pictures, soil animal classification can be carried out efficiently, thus making it possible for ordinary non-professionals to monitor, investigate, classify and analyze soil animals at any location. With the development of computer vision technology, object detection models are constantly improving in accuracy and complexity. The YOLO (You Only Look Once) series of models are widely used in various object detection tasks due to their end-to-end structure and real-time processing ability. As Figure 1 shown in YOLOv9, although the detection efficiency and accuracy have been improved, in resource-constrained scenarios such as drones or edge devices, the high consumption of computing resources limits its practical application. Especially in environments with high real-time requirements or limited hardware resources, its potential cannot be fully exerted.
[0004] The insect object detection task, as a complex object detection problem, has challenges such as small detection targets and complex backgrounds, and has problems of large computational volume, high number of parameters, long time consumption, and strong dependence on high-performance computing devices.
[0005] In view of this, there is an urgent need to provide a lightweight real-time positioning and detection scheme for insect objects that can be used in resource-constrained scenarios such as drones or edge devices. Summary of the Invention
[0006] To overcome the problems existing in the related art, the present disclosure provides an insect object detection method, device, equipment and medium based on an improved YOLOv9 model to solve the problems that the insect object detection task, as a complex object detection problem, has challenges such as small detection targets and complex backgrounds, and has problems of large computational volume, high number of parameters, long time consumption, and strong dependence on high-performance computing devices.
[0007] One or more embodiments of this specification provide an insect target detection method based on an improved YOLOv9 model, including the following steps:
[0008] Dataset construction: Collect magnified insect sample images through a high-resolution microscope and set category labels to obtain a sample dataset;
[0009] Detection model construction: Taking the YOLOv9 model as the basic model, replace the backbone network of the YOLOv9 model with a multi-layer StarNet backbone network to map the input features to an implicit high-dimensional non-linear feature space; replace the 13th, 19th, 28th, and 34th RepNCSPELAN4 modules of the original YOLOv9 model with LFRepBlock modules, and replace the Conv modules of the 26th and 27th layers with GSConv modules to obtain an improved YOLOv9-SFF model. Among them, the LFRepBlock module is a module that realizes multiple feature fusions and integrates lightweight components to reduce the redundant computational amount between consecutive convolutional layers, ensure the effective transmission and fusion of feature information, and ensure the extraction ability of important features;
[0010] Model training: Train the YOLOv9-SFF model through the constructed sample dataset until the model converges based on a preset convergence condition to obtain a model for insect target detection and classification.
[0011] One or more embodiments of this specification provide an insect target detection device based on an improved YOLOv9 model, including:
[0012] A dataset construction module for collecting magnified insect sample images through a high-resolution microscope and setting category labels to obtain a sample dataset;
[0013] A detection model construction module for taking the YOLOv9 model as the basic model, replacing the backbone network of the YOLOv9 model with a multi-layer StarNet backbone network to map the input features to an implicit high-dimensional non-linear feature space; replacing the 13th, 19th, 28th, and 34th RepNCSPELAN4 modules of the original YOLOv9 model with LFRepBlock modules and replacing the Conv modules of the 26th and 27th layers with GSConv modules to obtain an improved YOLOv9-SFF model. Among them, the LFRepBlock module is a module that realizes multiple feature fusions to reduce the redundant computational amount between consecutive convolutional layers, ensure the effective transmission and fusion of feature information, and ensure the extraction ability of important features.
[0014] One or more embodiments of this specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the insect target detection method based on the improved YOLOv9 model as described above.
[0015] One or more embodiments of this specification provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the insect target detection method based on the improved YOLOv9 model as described above.
[0016] An insect target detection method, device, equipment, and medium based on the improved YOLOv9 model provided by this disclosure have the advantage that a detection model with a lightweight architecture is proposed in the design, that is, the lightweight Starnet is used to replace the backbone of the original YOLOv9 to reduce the computational cost, and GSConv is used to replace some convolutional layers in the PGI of YOLOv9 to further improve the computational efficiency of the model. Some of the RepNCSPELAN4 modules for feature extraction and fusion are replaced with the LFRepBlock module, which can achieve multiple feature fusions, can achieve high precision in complex backgrounds and small target tasks, reduce the occurrence of missed detections and misclassifications, and after replacing some RepNCSPELAN4 modules with LFRepBlock, the number of parameters and GFLOPs of the model are greatly reduced, effectively solving the problems of long detection time and strong dependence on computing devices in the prior art, and having significant real-time detection advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is the structure diagram of the existing YOLOv9 network model provided by one or more embodiments of this specification;
[0019] Figure 2 It is the flowchart of an insect target detection method based on the improved YOLOv9 model provided by one or more embodiments of this specification;
[0020] Figure 3 It is the structure diagram of the improved YOLOv9-SFF network model provided by one or more embodiments of this specification;
[0021] Figure 4 The structural diagram of the StarNet backbone network provided for one or more embodiments of this specification;
[0022] Figure 5 The structural diagram of the GSConv module provided for one or more embodiments of this specification;
[0023] Figure 6 The structural diagram of the LFRepBlock module provided for one or more embodiments of this specification;
[0024] Figure 7 The structural diagram of the RepNCSPELAN4 module in the existing YOLOv9 model provided for one or more embodiments of this specification;
[0025] Figure 8 The comparison diagram of the detection results of insect images in most scenarios using the existing YOLOv9 and the YOLOv9-SFF of this embodiment provided for one or more embodiments of this specification;
[0026] Figure 9 The comparison diagram of the detection results of insect images in some scenarios with complex backgrounds, dense detection targets or small targets using the existing YOLOv9 and the YOLOv9-SFF of this embodiment provided for one or more embodiments of this specification;
[0027] Figure 10 The block diagram of an insect target detection device based on an improved YOLOv9 model provided for one or more embodiments of this specification; and
[0028] Figure 11 The structural schematic diagram of a computer device provided for one or more embodiments of this specification. Detailed implementation manners
[0029] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] There is an extremely rich variety of insect species on land, in water, and in the air. For example, there is an extremely rich and diverse group of animals inhabiting the soil. In terms of categories, soil animals cover eight phyla of animals, namely protozoa, platyhelminthes, rotifera, nematoda, annelida, mollusca, tardigrada, and arthropoda. Most of the invertebrates in all terrestrial ecosystems live in the soil or spend some periods of their life cycles in the soil. From the perspective of species records, although 25,000 species of nematodes, 8,000 species of springtails, 3,800 species of earthworms, 11,000 species of millipedes, 45,000 species of mites, etc. have been discovered and reported, however, except for earthworms, it is estimated that the other reported species only account for about 10% of the total number of potential species globally. Due to the hidden living environment of soil animals, they cannot be directly discovered, and they are also tiny (except for large soil animals, the body width of small and medium-sized soil animals is less than 2 mm). Even if specimens are collected, it is difficult to directly observe them with the naked eye or ordinary equipment. Recording soil animals often requires good professional knowledge, professional reference books, and a clear microscope.
[0031] Although AI technology can provide great assistance for insect detection work, for detection models, detecting insects in the soil, for example, faces challenges such as small detection targets and complex backgrounds. Moreover, insect detection algorithms have problems such as large computational volume, high number of parameters, long time consumption, and strong dependence on high-performance computing devices. The present invention effectively reduces the computational cost by making lightweight improvements to Yolov9 and using it for insect detection: First, use Starnet to replace the traditional backbone; Second, design and adopt the LFRepBlock module to replace some of the RepNCSPELAN4 modules to achieve the lightweight of the network; Replace the standard convolution in PGI with the GSConv module. In this way, the improved YOLOv9-SFF model for insect classification and detection is obtained; Then, compared with the original YOLOV9, the number of parameters and GFLOPs of the improved YOLOv9-SFF model are reduced by 30.1% and 29.8% respectively. In the comparative experiment, when only some of the RepNCSPELAN4 modules are replaced with LFRepBlock, the number of parameters and GFLOPs of the model are reduced by 30.2% and 23.8% respectively (further reduced to 47.8% and 52.4% after re-parameterization). This technical solution effectively solves the problems of long time consumption and strong dependence on computing devices in the prior art for insect detection, and has significant real-time detection advantages.
[0032] The following makes a detailed description of the present invention in combination with the specific implementation manners and the accompanying drawings of the specification.
[0033] According to an embodiment of the present invention, there is provided an insect target detection method based on an improved YOLOv9 model, as Figure 2As shown in the figure, it is a flowchart of an insect target detection method based on an improved YOLOv9 model provided by this embodiment. According to the insect target detection method based on the improved YOLOv9 model of the embodiment of the present invention, the following steps are included:
[0034] Step S1, dataset construction: Collect magnified insect sample images through a high-resolution microscope and set class labels to obtain a sample dataset.
[0035] Step S2, detection model construction: Using the YOLOv9 model as the basic model, replace the backbone network of the YOLOv9 model with a multi-layer StarNet backbone network to map the input features to an implicit high-dimensional non-linear feature space; replace the 13th, 19th, 28th, and 34th RepNCSPELAN4 modules of the original YOLOv9 model with LFRepBlock modules, and replace the Conv modules of the 26th and 27th layers with GSConv modules to obtain an improved YOLOv9-SFF model. Among them, the LFRepBlock module is a module that realizes multiple feature fusions and integrates lightweight components to reduce the redundant computational amount between consecutive convolutional layers, ensure the effective transmission and fusion of feature information, and ensure the extraction ability of important features. Step S3, model training: Train the YOLOv9-SFF model with the sample dataset constructed in step S1 until the model converges based on a preset convergence condition to obtain a model for insect target detection and classification.
[0036] In this embodiment, a lightweight architecture is proposed in the design of the model, that is, using a lightweight Starnet to replace the original backbone of YOLOv9 to reduce the computational cost, using GSConv to replace the standard convolutional layer in the programmable gradient information of the original YOLOv9 model to further improve the computational efficiency of the model, and replacing some of the RepNCSPELAN4 modules for feature extraction and fusion with LFRepBlock modules. This module can realize multiple feature fusions, achieve high precision in complex backgrounds and small target tasks, reduce the occurrence of missed detections and misclassifications, and after replacing the RepNCSPELAN4 module with LFRepBlock, the number of parameters and GFLOPs of the model are greatly reduced, effectively solving the problems of long detection time and strong dependence on computing devices in the prior art, and having significant real-time detection advantages.
[0037] In this embodiment, the dataset construction is specifically implemented through the following steps.
[0038] The insect image dataset used in this embodiment was made through a laboratory electron microscope after random sampling in the natural wild, covering various insect groups such as Collembola and Isopoda, etc. For data collection, a suitable sampling environment was strictly selected to avoid background interference as much as possible. The insect samples were magnified using a high-resolution microscope to ensure that the details of the insect morphology were clearly visible. To ensure the diversity and accuracy of the data, each sample was photographed from various angles, and blurred or interfered images were screened. Eventually, a high-quality, multi-morphology dataset suitable for morphological feature detection was formed.
[0039] The collection of insect sample images can be achieved through the following steps:
[0040] A1. Select the animal specimen to be photographed, transfer it to a glass dish with forceps or a plastic dropper, and place the glass dish on the stage of a binocular microscope.
[0041] A2. Turn on the computer with an imaging system, adjust the magnification and focal length, and photograph the soil animals in the glass dish.
[0042] A3. After one photographing is completed, gently shake or stir the glass dish to rearrange the soil animals in the image, and then take another photograph. You can also adjust the position of the glass dish to obtain soil animal images from different angles.
[0043] A4. For individuals with different body colors, adjust the light in a timely manner to obtain clear images. For example, for individuals with white or light-colored bodies, the light should be dimmed; for individuals with dark or darker-colored bodies, the light should be brightened. When there are both light-colored and dark-colored individuals in the image, adjust gradually and select the light brightness with the clearest imaging for photographing.
[0044] A5. Name and number the photographed photos, save them in the folder of the corresponding order, and mark the classification information, relevant features, and the photographing date. The number can be a letter + number.
[0045] In this embodiment, in terms of dataset label production, the LabelImg tool can be used to manually label the insect individuals in each image, outline the insects and classify and label them according to their species. During the labeling process, the category label standard is strictly followed to ensure the accuracy and consistency of the labeling. The final generated labeling file adopts the YOLO format, which is convenient for subsequent model training and evaluation. After labeling and processing, the dataset is divided into a training set, a validation set, and a test set, with a ratio of 7:2:1, which is used for model training and testing. It is necessary to ensure that the proportion of the number of various insects in each subset remains consistent, and the numbers are 1175:334:184 respectively.
[0046] In this embodiment, during the process of building the model, compare and refer to Figure 1 and Figure 3 ,Figure 3 This is the improved YOLOv9-SFF model architecture diagram provided in this embodiment. The backbone network of the YOLOv9 model is replaced with a multi-layer StarNet backbone network. The improved YOLOv9-SFF main network mainly extracts features from the input image. Among them, StarNet is an efficient neural network architecture based on the star operation. The structure is as follows Figure 4 shown. The star operation maps the input features to an implicit high-dimensional non-linear feature space through element-wise multiplication without increasing the width of the network. The key innovation of this operation is that it can achieve low-latency and efficient feature expression in a compact network structure with low energy consumption, especially suitable for resource-constrained scenarios.
[0047] In this embodiment, the entire backbone network of the original YOLOv9 model is replaced with a StarNet backbone network. The backbone network of this embodiment is set as a hierarchical network of StarNet, directly using a convolutional layer to downsample the shape of the feature map, and repeating four StarNet_Blocks modules with different depths to extract features, StarNet networks with different widths and depths.
[0048] Refer to Figure 4 , this is the structure diagram of the StarNet_Blocks module provided in this embodiment. The image processing process of the StarNet_Blocks module is that after the feature map passes through DWConv, it passes through a convolutional layer (F1, F2) respectively, and then performs feature fusion through the star operation. Then, it passes through a convolutional layer (G) and DWConv in sequence, and the features extracted are connected with the original feature map in a residual connection to achieve feature fusion and then output to the lower-layer network. The residual connection helps to alleviate the problem of gradient disappearance in deep networks. Among them,
[0049] Design of the deep module (Block): The depthwise separable convolution (DW-conv) in each StarNet_Block module uses a relatively large convolutional kernel (7×7), which helps to capture a larger range of spatial context information. This design can better aggregate the surrounding background information when dealing with complex backgrounds, and improve the ability to identify small targets under background interference.
[0050] Multi-scale feature extraction: The number of channels (embed_dim) of the feature map is gradually increased through the downsampling modules inside multiple stages, and a downsampling operation is introduced in each stage. This pyramid structure allows the model to extract multi-scale features from different resolutions, which helps to capture information about targets of different sizes, thereby improving the detection effect of small targets.
[0051] Dual-branch feature fusion (f1, f2): In the StarNet_Block module, the f1 and f2 feature channels of the dual-branch structure are fused with different features through activation and multiplication operations. Such operations not only enhance the feature expression ability but also better retain detailed information, which is beneficial for extracting small targets in complex backgrounds.
[0052] Residual connection and DropPath: The residual connection helps alleviate the vanishing gradient problem in deep networks while retaining the original information of the input. The introduction of DropPath provides a random regularization effect, which helps the model enhance its generalization ability in complex scenarios and reduce false detections of the background.
[0053] The basic formula of the star operation in the StarNet_Block module is as follows, defined as the element-wise product of two linearly transformed input feature vectors:
[0054]
[0055] where W1 and W2 are weight matrices, X is the input feature, and * represents element-wise multiplication. This operation can map the input to an implicitly high-dimensional feature space of approximately (d + 2)(d + 1) / 2 dimensions. The expanded expression is:
[0056]
[0057] Through this element-wise multiplication, the star operation realizes the non-linear combination of features, greatly improving the network's expression ability. When multiple layers of star operations are stacked, the implicit dimension grows exponentially. For the l-th layer, the recursive formula of the star operation is:
[0058]
[0059] As the number of layers increases, the implicit dimension approaches infinity, thus achieving a powerful high-dimensional feature expression ability with low-dimensional inputs.
[0060] Preferably, in this embodiment, the 26th and 27th Conv modules in the programmable gradient information of the original model are replaced with GSConv modules to improve the accuracy and computational efficiency of the model's object detection, and the lightweight GSConv module makes the model applicable to scenarios with limited computing resources.
[0061] Programmable Gradient Information (PGI) is a new type of auxiliary supervision. The 26th and 27th layer Conv modules are replaced with GSConv modules. GSConv is a lightweight convolution technology. By combining depthwise convolution and ordinary convolution, GSConv reduces redundancy and breaks the local dependencies between channels caused by grouped convolution through channel shuffle, enhancing the diversity of feature maps and improving the module's ability to represent complex features. It improves the accuracy and computational efficiency of real-time object detection, is particularly suitable for edge computing and lightweight models, and at the same time enhances the ability to extract local details. The original module mainly focuses on repetitive standard convolution structures, which are less efficient than GSConv when dealing with complex backgrounds or small targets. The structure is as Figure 5 shown. The GSConv module is a module that combines standard convolution (SC), depthwise separable convolution (DWConv), and channel shuffle. It combines the advantages of standard convolution (SC) and depthwise separable convolution (DSC), and uses shuffle to break the local dependencies between channels caused by grouped convolution.
[0062] First, it uses grouped convolution to divide the input features into multiple groups, and each group performs convolution operations independently. This method significantly reduces the computational amount because each group only processes a part of the input channels. For standard convolution, the number of parameters of the convolution kernel is k×k×c in ×c out , and the number of parameters of grouped convolution is reduced to where g is the number of groups. Then, the feature map is further processed through a 5×5 convolution. Compared with conventional 3×3 or 1×1 convolutions, it can capture a larger receptive field and achieve a good balance between feature extraction ability and computational efficiency.
[0063] In this embodiment, a brand-new LFRepBlock module is designed to replace the 13th, 19th, 28th, and 34th layer RepNCSPELAN4 modules of the original YOLOv9 model to achieve model lightweighting and performance optimization, as follows.
[0064] Refer to Figure 6 shown, which is the structural diagram of the LFRepBlock module provided in this embodiment. The LFRepBlock module is an efficient lightweight module based on multiple feature fusions and designs. It decomposes the output operation of the GSConv module into two lightweight convolution operations, and combines a feature fusion method with lightweight design to efficiently fuse the feature maps output by each layer. Specifically:
[0065] Refer to Figure 6 and Figure 7 shown.Figure 7 It is the structural diagram of the RepNCSPELAN4 module in the existing YOLOv9 model. In this embodiment, the standard convolution module at the input end of the original RepNCSPELAN4 module is replaced with a GSConv module, and then two feature branches are output. The RepNCSP module and the standard convolution in one branch are replaced with a FasterBlock module. The other branch is processed through a convolutional layer and Pconv. Subsequently, these two branches are subjected to multiple feature fusions using concatenation and another convolutional layer. The output features are concatenated with the output of the GSConv module at the input end and then further feature fused through a convolutional layer. Among them, the standard convolution module at the input end is replaced with a GSConv module, which effectively reduces redundant calculations through standard convolution and depthwise separable convolution and retains the feature expression ability. Then, the RepNCSP module is replaced with a FasterBlock module, effectively solving the problem that the RepNCSP structure is relatively complex and contains multiple convolutional layers and residual connections. The FasterBlock significantly reduces the computational amount by reducing the module complexity and depth. At the same time, the ability to extract features by Pconv and the multi-layer perceptron MLP to enhance the non-linear feature expression ability can still maintain a good feature extraction effect. In this embodiment, the LFRepBlock module uses GSConv(SC+DSC) and multiple feature fusions to achieve more efficient features, which is suitable for processing small targets in complex backgrounds.
[0066] In a specific embodiment, the LFRepBlock module sequentially includes a GSConv module. Two branches are respectively set under the GSConv module. One branch is set with a FasterBlock module, and the other branch is set with a Conv2d module and a PConv module. The features output by the two branches are concatenated into an overall feature through a Concat module. After the concatenation operation, 1×1 convolution is used to further compress and streamline the feature information, so as to ensure that insects in complex backgrounds can be quickly identified through a lightweight network. The fused features are then concatenated with the features processed by the GSConv module through a Concat module. Finally, the feature map after a convolutional layer is output to the next network layer. This layer effectively reduces the computational amount while maintaining the efficiency of the model in complex backgrounds, so that the output result can not only contain sufficient feature information but also quickly respond to the detection task.
[0067] In this embodiment, the FasterNet Block module includes a main branch and a residual connection path. The main branch consists of two convolutional layers and a 1×1 convolution. The first convolutional layer uses a 3×3 PConv convolution to enhance feature extraction. The second convolutional layer uses a 1×1 convolutional kernel to expand the channel features for dimensionality increase. Finally, a convolutional layer with a 1×1 convolutional kernel is used to perform feature dimensionality reduction, refine, and integrate features. The output of the main branch is added to the input features retained by the residual connection after regularization to form the final output, which is then fed into the next layer of the network. In this way, the module can significantly reduce the number of network parameters and computational complexity while keeping the size of the output feature map unchanged.
[0068] In this embodiment, the processing principle of each layer of the LFRepBlock module:
[0069] 1) Input processing and channel dimensionality reduction (cv1): The image passes through the GSConv module. First, the input features are mapped from the channel number c1 to the c3 dimension through a 1×1 convolution, and then split into two paths through channel splitting, with each path having a channel number of c3 / / 2. This dimensionality reduction operation can reduce the computational overhead while retaining the main feature information, laying the foundation for subsequent processing. For insect images in complex backgrounds, the lightweight operation of this layer helps reduce unnecessary redundant information and concentrate on extracting the main target features.
[0070] 2) Local feature enhancement (cv2): Part of the feature channels are processed through the FasterBlock. The FasterBlock enhances the local features and the feature non-linear expression ability through partial convolution and the MLP structure. At the same time, while keeping the model lightweight, it can capture the subtle features of insects. Especially in the case of complex backgrounds, this step helps to highlight the differences between insects and the background and reduce misidentifications.
[0071] 3) Feature segmentation and lightweight convolution (cv3): In the second part of the channels, PConv takes advantage of the redundancy in the feature map and only applies conventional Conv to some of the input channels for spatial feature extraction while retaining the remaining channels. While retaining important information, it performs convolution on some features. This layer is very effective for detail processing in insect images and can further distinguish the texture, contour, etc. of insects from the background. Essentially, PConv can avoid frequent memory access, better utilize the lower FLOPs of PConv compared to conventional Conv, and make better use of the computing power on the device.
[0072] 4) Feature Fusion Part (cv4 + cv5): `cv4` simplifies the output fusion of `cv2` and `cv3` through lightweight 1x1 convolutions, avoiding excessive computational overhead and adjusting the feature weights to ensure a balance between performance and efficiency. Finally, the output layer of the `cv5` part is also optimized, reducing redundant calculations by simplifying the convolution operations and ensuring that the number of channels in the output matches that in the input under the premise of lightweight. These improvements, by reducing the number of channels, introducing lightweight modules, simplifying feature fusion, supporting cross-layer feature fusion, and output design, enable the model to significantly reduce the inference time and computational resource consumption while maintaining performance, making it particularly suitable for embedded devices or real-time detection tasks. Through branch processing and feature fusion, the model can still maintain a high detection accuracy when dealing with complex features, ensuring a good balance between lightweight and performance.
[0073] In this embodiment, the detailed processing flow of the LFRepBlock module:
[0074] The processing flow is divided into five parts. For ease of description, the following variables are defined: the input feature map is denoted as X, with a size of (C 1 , H, W), and the output feature map is denoted as Y, with a size of (C 2 , H″, W″). C 1 and C 2 represent the number of input and output channels respectively, C 3 and C 4 represent intermediate variables. H and W represent the height and width of the input feature map. k represents the convolution kernel size.
[0075] (1) Input Feature Figure X First, it passes through GSConv (k = 1) for channel transformation and feature recombination, converting the number of channels from C1 to C3, while reducing the number of parameters and enhancing feature diversity. Subsequently, the feature map is split into two branches, and the size of each branch is:
[0076]
[0077] (2) Lightweight processing of the branch feature maps. The first branch passes through the Faster Block for processing, and the size of the output feature map remains unchanged. The second branch passes through Conv2d (k = 1) and PConv for processing, and the size of the output feature map is (2). Both branches perform lightweight and efficient feature extraction simultaneously.
[0078]
[0079] Connect the feature maps of the two branches in step (2), and then further process them through Conv2d (k = 1) to enhance the module's ability to capture detailed and high-level semantic features, provide a more comprehensive feature representation for subsequent processing, and obtain an output feature map of size:
[0080] size(y Cat )=(C 4 ,H′,W′).
[0081] (3) Concatenate the output of GSConv in step (1) with the output of step (3), which not only retains low-level feature details but also retains high-level global semantics. This information complementarity effectively improves the object detection ability of the module. The concatenated feature map is further fused through Conv2d (k = 1) to enhance feature representation and interaction, and finally the output size is:
[0082] size(Y)=(C 2 ,H″,W″).
[0083] In this embodiment, the beneficial effects brought by each layer of the LFRepBlock module are as follows:
[0084] 1) Difference in feature extraction efficiency and convolutional structure: The LFRepBlock module uses GSConv (SC+DSC) and multiple feature fusions to achieve more efficient features, which is suitable for processing small targets in complex backgrounds. GSConv reduces redundant calculations by composing depth convolution and ordinary convolution, and at the same time improves the ability to extract local details. The original module mainly focuses on repetitive convolutional structures and is not as efficient as GSConv in processing complex backgrounds or small targets. In addition, GSConv will also replace the standard convolutional layer in PGI.
[0085] 2) Lightweight design and module fusion: LFRepBlock introduces Faster Block and PConv. Faster Block uses the combination of PConv and MLP to improve the retention of local information, computational efficiency, and feature non-linear expression ability, especially in complex backgrounds, it can better extract the details of small targets. PConv only performs more detailed convolutional processing on part of the input features while keeping some features unchanged, which can ensure that the model better retains and strengthens the detailed features of the target.
[0086] 3) Inter-layer fusion strategy: In LFRepBlock, cv4 is responsible for lightweight fusion of the outputs of cv2 and cv3, and cv5 is responsible for lightweight fusion of the outputs of cv4 and cv1. This process combines different channel features after convolutional operations, which can better integrate context information, retain both low-level feature semantics and high-level feature information, and retain key information.
[0087] 4) Multi-branch and parallel processing: The LFRepBlock utilizes the multi-branch strategy of cv1 to split the input features into two parts, which are processed through cv2 and cv3 respectively. PConv and Faster Block are further nested in cv3. After multiple parallel processings and then fusion, the model can capture information of different scales and regions more flexibly.
[0088] In this embodiment, the contribution of the LFRepBlock module to reducing false detections and missed detections is improved as follows:
[0089] 1) FasterBlock: This module processes features by introducing efficient spatial mixing operations (such as PConv) and multi-layer perceptrons (MLP), thereby improving the accuracy and discrimination ability of feature non-linear expression. Especially in the detection of targets of similar classes, it can capture more subtle feature differences and reduce false detections.
[0090] 2) Multiple parallel processings and fusion: The LFRepBlock adopts the multiple parallel processing strategy. After decomposing the input features, they are processed through different paths and then fused, which can capture information of multiple scales and regions. This strategy has a significant effect on distinguishing between complex backgrounds and similar classes, and helps to reduce missed detections and false detections.
[0091] In this embodiment, when using the YOLOv9-SFF of this embodiment to process insect images, the model will generate a rectangular bounding box for each detected target, indicating the position and size of the target in the image, and assign a class label to each target, specifying which class the target belongs to (pre-defined in the training stage). At the same time, the model will also attach a confidence score to each detection result, indicating the credibility of the model for this detection result. The higher the confidence, the more accurate the model's judgment of the target's class and position. Finally, the inference output of YOLOv9-SFF contains the corresponding target bounding box, class label and confidence, which are used for object detection, classification and other tasks.
[0092] Refer to Table 1 and Table 2 below. Table 1 shows a comparison case of the calculation amount and accuracy of the LFRepBlock module and the RepNCSPELAN4 module adopted in this embodiment during the insect order detection task. Table 2 shows the ablation comparison experiment results of the LFRepBlock module and the RepNCSPELAN4 module adopted in this embodiment.
[0093] The specific configuration for comparison implementation is as follows: Network training is conducted on an experimental platform running Windows 11 (64-bit), which is equipped with an Intel i7-14700KF CPU, an NVIDIA GeForce RTX 2080TI GPU (22GB), and 64GB of RAM. The experimental code is implemented using CUDA 12.2 and Python 3.9. The input size of the network is set to 640×640, and all hyperparameter settings are kept consistent. Training is initialized with a fixed random seed, and no pre-trained weights are applied. The batch size is set to 16, and the training process uses an auxiliary branch and a main branch. Other configurations remain at their default values. Evaluation metrics are used to evaluate the effectiveness of the proposed improvements, including parameters, precision, recall, average precision (AP), mAP, and GFLOP. Parameters refer to the total number of trainable parameters, and GFLOP represents the performance calculated in terms of floating-point operations per second.
[0094] The formulas for precision, recall, AP, and mAP are as follows.
[0095]
[0096] Among them, TP represents the number of positive samples correctly predicted as positive; FP represents the number of negative samples wrongly predicted as positive; TN represents the number of negative samples correctly predicted as negative; FN represents the number of positive samples wrongly predicted as negative. r represents recall.
[0097] Table 1. Comparative data of LFRepBlock compared with RepNCSPELAN4
[0098]
[0099] Table 2. Results of ablation experiments
[0100]
[0101] According to the comparative experiment results in Table 1, the LFRepBlock module has achieved a significant reduction in the number of parameters and computational volume compared with RepNCSPELAN4, with a reduction of 30.2% and 23.8% respectively, while the mAP val75It can still increase by 0.1%. This indicates that LFRepBlock not only has obvious advantages in model lightweighting but also can maintain a high detection accuracy. This can be attributed to the design of the LFRepBlock module, which optimizes the computational efficiency of the model by introducing lightweight convolution (GSConv) and partial convolution (PConv). In this design, GSConv effectively reduces the computational complexity of convolution operations by reducing grouped convolution; PConv further reduces redundant computations and avoids full-scale feature map convolution, thus accelerating the inference speed. These structural improvements enable the model to capture sufficient feature information while reducing the computational burden, ensuring the maintenance of detection performance. In addition, the introduced Faster Block improves the computational efficiency by using PConv, MLP, and Layer Scale, further reducing the number of parameters and computational resource consumption. Although the increase in mAP val75 is limited, more importantly, LFRepBlock significantly improves the inference speed (FPS is increased to 44) without sacrificing accuracy, which makes the improved YOLOv9-SFF more suitable for resource-constrained embedded devices and low-power environments in practical applications, such as real-time monitoring of agricultural pests and field species monitoring in ecological protection.
[0102] Table 2 shows the change process of the architecture in the ablation experiment. By gradually introducing new modules, specifically when GSConv is introduced alone, both the number of parameters and GFLOPs decrease slightly, and mAP val75 increases by 1.6%, which indicates that GSConv not only effectively reduces the computational amount but also improves the efficiency of the model in feature extraction. When part of RepNCSPELAN4 is replaced with LFRepBlock, the number of parameters and GFLOPs are significantly reduced to 45.06M and 216.3, showing the contribution of LFRepBlock in optimizing the inference speed. Although mAP val75 slightly decreases to 0.907, the overall detection performance still remains at a high level. When StarNet is introduced alone, the number of parameters and GFLOPs are reduced more significantly. After applying all improvement schemes, the number of parameters and GFLOPs are reduced by 30.01% and 29.8% respectively, and the FPS of the model is increased by 33.3%. Although the mAP val75 of the improved model slightly decreases, in some specific scenarios, its performance is better than the original model, especially in small object detection and complex background processing. The improved lightweight design makes the model more efficient and accurate in dealing with these scenarios, especially in real-time monitoring applications, where the improved model shows stronger advantages in terms of speed and accuracy. Therefore, although the accuracy is reduced, these improvements bring better comprehensive performance in practical applications.
[0103] In this embodiment, referring to Figure 8 as shown, it is a comparison chart of the detection results of insect images using YOLOv9 in the prior art and YOLOv9-SFF of this embodiment in most scenarios. Referring to Figure 9 , it is a comparison chart of the detection results of insect images with complex backgrounds, dense detection targets or small targets in some scenarios processed by YOLOv9 in the prior art and YOLOv9-SFF of this embodiment. Among them, Figure 8 and Figure 9 , from left to right vertically shown are the original image → YOLOv9 detection result image → YOLOv9-SFF detection result image.
[0104] Among them, through the recognition of the target area, the two models can classify the detected objects, so as to output the specific insect species. Although the improved YOLOv9-SFF model has a slight decrease in mAP val75 (the original model is 0.915 and the improved model is 0.901), their detection effects in conventional target detection tasks are basically equivalent. This is because the mAP values are all in the high-precision range, indicating that the models perform well in target detection tasks, such as Figure 8 . In addition, when processing images with complex backgrounds or dense detection targets containing small targets, Ours (YOLOv9-SFF model) performs better than YOLOv9, such as Figure 9 (horizontal comparison) group b and group d of the figure, which shows that although there are differences in accuracy in some complex scenarios, the overall detection effect is still satisfactory. In addition, by horizontally comparing the four groups of figures a-d, it can be seen that YOLOv9 shows misdetection ( Figure 9 -a and 9-c) and missed detection ( Figure 9 -b and 9-d) phenomena in the data set, while YOLOv9-SFF shows higher accuracy and reliability in small target detection. The introduction of the lightweight structure greatly improves the inference speed of the model, effectively making up for the impact of the accuracy drop. Generally speaking, YOLOv9-SFF not only maintains high detection ability in practical applications, but also can adapt to scenarios with high real-time requirements, showing excellent comprehensive performance.
[0105] The results of the comparative experiments show that the LFRepBlock module we proposed has achieved remarkable results in terms of lightweight. Compared with the original RepNCSPELAN4 module, the LFRepBlock module reduces the number of parameters by 30.2% and GFLOPs by 23.8%, while still achieving a slight improvement in the Precision and mAPval75 metrics. After being applied to all improvements, the number of parameters and GFLOPs of the model decreased by 30.1% and 29.8% respectively. This makes the LFRepBlock module and the improved model particularly suitable for applications in embedded devices or edge computing environments, such as real-time species monitoring in ecological protection and agricultural pest detection scenarios. Although the mAP val75 decreased by 1.4%, in actual image inference, the improved model can still accurately identify individual insects in crowded and complex scenes, especially for targets with small size and blurred features, where the detection accuracy has been significantly improved. In addition, there were misdetection and missed detection phenomena in the original YOLOv9 model during the detection process, while the improved model shows more stable and reliable performance in terms of accuracy.
[0106] Device Embodiment
[0107] According to an embodiment of the present invention, there is provided an insect target detection device based on an improved YOLOv9 model, as Figure 10 shown, which is a block diagram of the insect target detection device based on the improved YOLOv9 model provided in this embodiment. The insect target detection device based on the improved YOLOv9 model according to an embodiment of the present invention includes:
[0108] A dataset construction module 10, configured to collect magnified insect sample images through a high-resolution microscope and set category labels to obtain a sample dataset;
[0109] A detection model construction module 20, configured to use the YOLOv9 model as the basic model, replace the backbone network of the YOLOv9 model with a multi-layer StarNet backbone network to map the input features to an implicit high-dimensional non-linear feature space; replace the 13th, 19th, 28th, and 34th RepNCSPELAN4 modules of the original YOLOv9 model with LFRepBlock modules, and replace the 26th and 27th Conv modules with GSConv modules to obtain an improved YOLOv9-SFF model, where the LFRepBlock module is a module that realizes multiple feature fusions and incorporates lightweight components to reduce the redundant computational amount between consecutive convolutional layers, ensure the effective transmission and fusion of feature information, and ensure the extraction ability of important features.
[0110] The model training module 30 is used to train the YOLOv9-SFF model with the constructed sample data set until the model converges based on the preset convergence conditions, and obtain a model for insect target detection and classification.
[0111] In this embodiment, a lightweight architecture is proposed in the design of the model, that is, using the lightweight Starnet to replace the original backbone of YOLOv9 to reduce the computational cost, and using GSConv to replace some convolutional layers in the PGI of YOLOv9 to further improve the computational efficiency of the model. The RepNCSPELAN4 module for feature extraction and fusion is partially replaced by the LFRepBlock module, which uses multi-branch feature processing and efficiently fuses the feature maps output by each layer through a lightweight feature fusion method, achieving high precision in complex backgrounds and small target tasks, and reducing the occurrence of missed detections and misclassifications.
[0112] In this embodiment, for the YOLOv9-SFF model constructed by the detection model building module 20, the 26th and 27th layer Conv modules in the auxiliary supervision framework are also replaced by GSConv modules to improve the accuracy and computational efficiency of the model's target detection, and the lightweight GSConv module makes the model applicable to scenarios with limited computing resources.
[0113] In this embodiment, a new LFRepBlock module is designed to replace the 13th, 19th, 28th, and 34th layer RepNCSPELAN4 modules of the original YOLOv9 model to achieve model lightweighting and performance optimization, as follows.
[0114] Such as Figure 6As shown, the standard convolution module at the input end of the original RepNCSPELAN4 module is replaced with a GSConv module, and then two feature branches are output. The RepNCSP module and the standard convolution in one branch are replaced with a FasterBlock module. The other branch is processed through a convolutional layer and Pconv. Subsequently, the two branches are used for multiple feature fusions through concatenation and another convolutional layer. After the output features are concatenated with the output of the GSConv at the input end, they are further feature fused through a convolutional layer. Among them, the standard convolution module at the input end is replaced with a GSConv module, which effectively reduces redundant calculations through standard convolution and depthwise separable convolution and retains the feature expression ability. Then, the RepNCSP module is replaced with a FasterBlock module, effectively solving the problem that the RepNCSP structure is relatively complex and contains multiple convolutional layers and residual connections. The FasterBlock significantly reduces the computational amount by reducing the module complexity and depth. At the same time, the ability to extract features by PConv and the multi-layer perceptron MLP to enhance the non-linear feature expression ability can still maintain a good feature extraction effect. In this embodiment, the LFRepBlock module uses GSConv(SC+DSC) and multiple feature fusions to achieve more efficient features, which is suitable for processing small targets in complex backgrounds.
[0115] In this embodiment, in a specific embodiment, the LFRepBlock module sequentially includes a GSConv module. Two branches are respectively arranged under the GSConv module. A FasterBlock module is arranged in one branch, and a convolutional layer and a PConv module are arranged in the other branch. The features output by the two branches are fused into an overall feature through a Concat module. The fusion operation uses 1×1 convolution to further compress and streamline the feature information, so as to ensure that insects in complex backgrounds can be quickly recognized through a lightweight network, avoiding the loss or redundancy of features. The fused features are then passed through a convolutional layer and the features processed by the GSConv module are concatenated through a Concat module. Finally, they are feature fused through a convolutional layer, and the output feature map is output to the next network layer. This layer effectively reduces the computational amount while maintaining the efficiency of the model in complex backgrounds, so that the output result can contain sufficient feature information and can quickly respond to the detection task.
[0116] In this embodiment, when processing insect images using the YOLOv9-SFF of this embodiment, the model generates a rectangular bounding box for each detected target, representing the position and size of the target in the image, and assigns a class label to each target, indicating which class the target belongs to (predefined in the training phase). At the same time, the model also attaches a confidence score to each detection result, representing the credibility of the model for this detection result. The higher the confidence, the more accurate the model's judgment of the class and position of the target. Finally, the inference output of YOLOv9-SFF includes these bounding boxes, class labels, and confidences, which are used for object detection, classification, and other tasks.
[0117] The embodiment of the present invention is a device embodiment corresponding to the above method embodiment. The specific operations of each module processing step can be understood with reference to the description of the method embodiment and will not be elaborated here.
[0118] As Figure 11 shown, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the insect target detection method based on the improved YOLOv9 model in the above embodiment.
[0119] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the insect target detection method based on the improved YOLOv9 model in the above embodiment, or when the computer program is executed by the processor, it implements the insect target detection method based on the improved YOLOv9 model in the above embodiment. The computer program, when executed by the processor, implements the following method steps:
[0120] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0121] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present invention, and the content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.
Claims
1. An insect target detection method based on an improved YOLOv9 model, characterized in that The following steps are involved: Dataset construction: The enlarged insect sample images are collected through a high-resolution microscope, and the category labels are set to obtain the sample dataset; Detection model construction: Using the YOLOv9 model as the basic model, the backbone network of the YOLOv9 model is replaced with a multi-layer StarNet backbone network to map the input features to an implicit high-dimensional nonlinear feature space; the RepNCSPELAN4 modules of the 13th, 19th, 28th, and 34th layers of the original YOLOv9 model are replaced with LFRepBlock modules, and the Conv modules of the 26th and 27th layers are replaced with GSConv modules to obtain an improved detection model, where the LFRepBlock module is a module that implements multiple feature fusions and incorporates lightweight components to reduce the redundant calculation amount between consecutive convolutional layers, ensure the effective transmission and fusion of feature information, and ensure the ability to extract important features; as well as Model training: The YOLOv9-SFF model is trained using the constructed sample dataset until the model converges based on the preset convergence conditions to obtain a model for insect target detection and classification.
2. The insect target detection method based on the improved YOLOv9 model as claimed in claim 1, characterized in that: It also includes replacing the Conv modules of the 26th and 27th layers in the original YOLOv9 model's programmable gradient information with GSConv modules to improve the accuracy and computational efficiency of the model's target detection.
3. The insect target detection method based on the improved YOLOv9 model according to any one of claims 1 or 2, characterized in that: In the LFRepBlock module, the standard convolution module at the input end of the original RepNCSPELAN4 module is replaced with a GSConv module, and two branches are output, one of which is processed by Faster Block, and the other branch is processed by a convolution layer and PConv, and then the processed features are fused multiple times, including the fusion of the same hierarchical structure and the fusion of low-level and high-level features; wherein the standard convolution module at the input end is replaced with a GSConv module, which performs standard convolution, depth-separable convolution and channel shuffling.
4. The insect target detection method based on the improved YOLOv9 model according to any one of claims 1 or 2, characterized in that: The LFRepBlock module includes a GSConv module in sequence. Two branches are respectively set under the GSConv module, one branch is set with a FasterBlock module, and the other branch is set with a Conv2d module and a PConv module. The features output by the two branches are simultaneously fused with Conv2d through a Concat module, and then with the features processed by the GSConv module through a Concat module and a convolution layer for hierarchical feature fusion and then output to the next network layer.
5. The insect target detection method based on the improved YOLOv9 model as claimed in claim 1, characterized in that: The method of collecting the enlarged insect sample image by a high-resolution microscope comprises the following steps: A1. Select the animal specimen to be photographed, transfer it to a glass dish with tweezers or a plastic dropper, and place the glass dish on the stage of the binocular microscope; A2. Turn on the computer with the imaging system, adjust the magnification and focal length, and take photos of the soil animals in the glass dish; A3. After taking a picture, gently shake or stir the glass dish to rearrange the soil animals in the image, and then take another picture. You can also adjust the position of the glass dish to obtain images of soil animals from different angles. A4. For individuals with different body colors, adjust the lighting in time to obtain clear images for shooting; A5. The photos taken are numbered and named, saved in the corresponding target folder, and marked with classification information and related features.
6. The insect target detection method based on the improved YOLOv9 model as claimed in claim 1, characterized in that: The YOLOv9-SFF detection results include target bounding boxes, category labels, and confidence levels.
7. An insect target detection device based on an improved YOLOv9 model, characterized in that include: A dataset construction module is used to collect magnified insect sample images through a high-resolution microscope and set category labels to obtain a sample dataset; A detection model building module is used to use the YOLOv9 model as the basic model, replace the backbone network of the YOLOv9 model with a multi-layer StarNet backbone network, and realize mapping the input features to an implicit high-dimensional nonlinear feature space; replace the RepNCSPELAN4 modules of the 13th, 19th, 28th, and 34th layers of the original YOLOv9 model with LFRepBlock modules, and replace the Conv modules of the 26th and 27th layers with GSConv modules to obtain an improved detection model, wherein the LFRepBlock module is a module for realizing multiple feature fusion and integrating lightweight components, so as to reduce the redundant calculation amount between consecutive convolutional layers, ensure the effective transmission and fusion of feature information, and ensure the ability to extract important features; The model training module is used to train the YOLOv9-SFF model through the constructed sample data set until the model converges based on the preset convergence conditions to obtain a model for insect target detection and classification.
8. The insect target detection device based on the improved YOLOv9 model as claimed in claim 7, characterized in that: In the LFRepBlock module, the standard convolution module at the input end of the original RepNCSPELAN4 module is replaced with a GSConv module, and two branches are output, one of which is processed by Faster Block, and the other branch is processed by a convolution layer and PConv, and then the processed features are fused multiple times, including the fusion of the same hierarchical structure and the fusion of low-level and high-level features. Among them, the standard convolution module at the input end is replaced with a GSConv module, which performs standard convolution, depth-separable convolution and channel shuffling.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the insect target detection method based on the improved YOLOv9 model as described in any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the insect target detection method based on the improved YOLOv9 model as described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Improved StarNet-YOLOv13-based unmanned aerial vehicle field tobacco virus disease lightweight detection method
CN121767895A