Lightweight improved pepper flower small target detection method and system based on YOLOv8n

By introducing EMA lightweight attention mechanism and GSConv module into the YOLOv8n model, combined with the WIoU loss function, the YOLOv8n model is optimized, and the omissions and misjudgment problems of small object detection are solved, achieving the lightweight and efficient deployment of the model.

CN120298670APending Publication Date: 2025-07-11HUNAN AGRI UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510430497.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing YOLOv8n model is prone to omissions and misjudgments in small-size object detection, and has large parameters and high calculation costs, making it difficult to deploy efficiently in complex environments.

Method used

By introducing EMA lightweight attention mechanism and GSConv module to replace the traditional convolution module, and using the WIoU loss function to optimize the YOLOv8n model, a lightweight chili flower small object detection system is built.

Benefits of technology

It improves the accuracy and robustness of small object detection, while reducing the amount of model parameters and calculation complexity, making it suitable for resource-constrained terminal equipment deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298670A_ABST
    Figure CN120298670A_ABST
Patent Text Reader

Abstract

The invention discloses a small target detection method and system for pepper flowers based on YOLOv8n lightweight improvement, and the method comprises the steps: loading target detection image data through a corresponding database, converting the target detection image data into a VOC format data set trained by YOLO (You Only Look Once), and dividing a training set and a test set; the method comprises the following steps: by adding an EMA lightweight attention mechanism, introducing a GSConv (GSConv) lightweight convolution module and a WIoU (Light Intersection over Union) loss function to respectively replace a C2F convolution module in a trunk and an original CIoU (Complete Intersection over Union) loss function, and constructing a lightweight capsicum flower small target detection model based on improved YOLOv8n; training the small target detection model based on the training set to obtain an optimal small target detection model; inputting the test set into the optimal small target detection model, and outputting to obtain a small target detection result; by optimizing the YOLOv8n algorithm, the detection average precision mean value of the small-target pepper flowers is improved, the parameter quantity of the algorithm is reduced, the lightweight of the algorithm is realized, and the method is suitable for the requirements of mobile terminal deployment in the agricultural field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing in agricultural applications, and specifically to a small target detection technology for chili flowers based on the lightweight improvement of YOLOv8n. This technology is particularly applicable to the vision recognition system of chili flower pollination robots based on mobile terminals in facility agriculture. Background Technique

[0002] Traditional machine vision technology mainly relies on image processing and computer vision technology to implement steps such as image acquisition, preprocessing, feature extraction, and recognition through devices such as optical systems, image sensors, and computers. It has a wide range of applications in many detection fields and has the advantages of high precision and high speed. However, traditional machine vision also has some disadvantages, such as poor adaptability, low robustness, limited processing ability for complex and dynamic environments, and insufficient stability in the face of noise and interference. In recent years, the breakthrough in object detection technology based on deep learning has brought significant changes to the agricultural field. Especially the YOLO (You Only Look Once) algorithm, whose remarkable advantages include excellent real-time performance, simple and efficient architecture, multi-scale detection ability, effective utilization of global context information, and flexibility in supporting multi-task learning. It performs outstandingly in scenarios such as crop object detection, pose estimation, and real-time tracking, providing important support for precision agricultural management.

[0003] Currently, although single-stage object detection technologies represented by YOLOv8n have made significant optimization progress in network structure improvement, they are still insufficient in integrating context information, which leads to easy omission and misjudgment in the detection of small-sized objects. In addition, with the complication of application scenarios, YOLO series models face challenges of large parameter quantities and high computational costs. These models usually come with large parameter quantities and high computational complexity, which pose challenges in realizing end-to-end later deployment. Therefore, it has become an urgent problem to be solved to optimize the model structure through lightweight technology, reduce parameters and computational requirements, and develop a new method to lightweight improve YOLOv8n while enhancing the detection accuracy of small targets to achieve more efficient deployment of the model and further promote the application of object detection technology in the intelligentization and automation of agricultural production. Summary of the Invention

[0004] The technical problem solved by the present invention is to provide a method and system for detecting small targets of chili flowers based on the lightweight improvement of YOLOv8n. By optimizing the network structure of YOLOv8n, while improving the detection accuracy of small targets of chili flowers, it realizes the balance between the lightweight of the model and its performance, and is used to solve problems such as low detection accuracy, large and complex model volume, and difficult deployment in object detection tasks.

[0005] The method and system for detecting small pepper flower targets based on the lightweight improvement of YOLOv8n to solve the technical problems of the present invention are implemented by the following technical solutions.

[0006] In the present invention, the specific steps of the method for detecting small pepper flower targets based on the lightweight improvement of YOLOv8n are as follows: Step 1): Use the corresponding dataset to load the target detection image data, convert it into the YOLO training format, and then divide it into a training set, a validation set, and a test set; Step 2): Add the EMA lightweight attention mechanism, introduce the GSConv (GSConvolution, lightweight convolution) module to replace the original C2F convolution module in the backbone, introduce the WIoU (Weighted Intersection over Union) loss function to replace the original CIoU (Complete Intersection over Union) loss function in the backbone, and then construct a small pepper flower target detection model based on the lightweight improvement of YOLOv8n; Step 3): Train the small pepper flower target detection model based on the lightweight improvement of YOLOvn8 using the training set in Step 1), and then obtain the optimal small target detection model; Step 4): Input the test set in Step 1) into the optimal small target detection model obtained in Step 3), and then output the small target detection result.

[0007] In the present invention, the small pepper flower target detection model based on the lightweight improvement of YOLOv8n includes a backbone network, a neck network, and a detection head; The backbone network includes a first CBS convolution module, a second CBS convolution module, a first GSConv module, a third CBS convolution module, a second GSConv module, a fourth CBS convolution module, a third GSConv module, a fifth CBS convolution module, a fourth GSConv module, and an SPPF module connected in sequence. The neck network includes a first upsampling module (Upsample), a first aggregation module (Concat), a first C2F module, a second upsampling module (Upsample), a second aggregation module (Concat), a first EMA attention mechanism module, a second C2F module, a sixth CBS convolution module, a third aggregation module (Concat), a second EMA attention mechanism module, a third C2F module, a seventh CBS convolution module, a fourth aggregation module (Concat), a third EMA attention mechanism module, and a fourth C2F module connected in sequence; The second GSConv module is connected to the first aggregation module (Concat), the third GSConv module is connected to the second aggregation module (Concat), the SPPF module is respectively connected to the first upsampling module (Upsample) and the fourth aggregation module (Concat), and the first C2F module is connected to the third aggregation module (Concat).

[0008] In the present invention, the first GSConv module, the second GSConv module, the third GSConv module, and the fourth GSConv module all include depthwise separable convolution modules. The specific process of the depthwise separable convolution module is as follows: Process 1): Perform standard convolution processing on the input feature map, that is, pass the input feature map through a standard convolution operation to change the number of channels to meet the needs of subsequent operations; Process 2): Perform depthwise separable convolution processing on the input feature map, that is, process the input feature map through the depthwise separable convolution module. The depthwise separable convolution method can achieve feature extraction with less computational effort while keeping the number of channels unchanged; Process 3): Concatenate the input feature maps processed by Process 1) and Process 2), that is, concatenate the results of the standard convolution processing and the results of the depthwise separable convolution processing. Concatenating the feature maps extracted by the two different convolution methods can obtain a richer feature representation; Process 4): Perform a shuffle operation on the input feature map concatenated in Process 3), thereby breaking the independence of the input feature map in the channel dimension, promoting information exchange between different channels, and enhancing the model's ability to capture spatial information. This is one of the key innovations of the GSConv module.

[0009] In the present invention, the WIoU loss function is an improved loss function for object detection tasks. The WIoU loss function is innovated on the basis of the traditional CIoU loss function to improve the performance and robustness of the model in object detection.

[0010] In the present invention, the pepper flower small object detection system improved based on YOLOv8n lightweight includes a dataset construction module, a model construction module, a model training module, a model improvement module, and a model detection module; The dataset construction module loads object detection image data using the corresponding database, converts it into the YOLO training format, and divides it into a training set and a test set; The model construction module introduces the EMA lightweight attention mechanism and the GSConv module to optimize the YOLOv8n object detection algorithm, and constructs a pepper flower small object detection model improved based on YOLOv8n lightweight; The described model training module trains the small target detection model of chili flowers based on YOLOvn8 lightweight improvement using a training set, and then obtains the optimal small target detection model; The described model improvement module constructs a small target detection model of chili flowers based on YOLOvn8 lightweight improvement by introducing a GSConv lightweight convolution module, an EMA lightweight attention mechanism, and replacing and optimizing the loss function of the YOLOv8n object detection algorithm; The described model detection module inputs a test set into the optimal small target detection model, and then outputs small target detection results. Beneficial effects

[0011] Compared with the prior art, the present invention provides a method and system for small target detection of chili flowers based on YOLOv8n lightweight improvement. First, an EMA attention mechanism based on efficient multi-scale is introduced into the small target detection model, which has better detection performance compared with traditional attention mechanisms, and can have more superior performance in small target detection compared with some existing attention mechanisms, solving the problems of omission and misjudgment of traditional detection methods; Second, the GSConv is used to replace the traditional convolution module, which can effectively reduce the model parameters while ensuring that the model performance is not negatively affected, and complements the efficient multi-scale EMA attention mechanism to alleviate problems such as the increase in computational complexity caused by adding the attention mechanism; Third, there is a great improvement in the model parameters and detection accuracy, achieving a balance between lightweight and high performance compared with traditional models, laying a solid technical foundation for ensuring the accuracy of small target detection and transplanting it to terminal devices. Description of the drawings

[0012] Figure 1 It is a flowchart of the method for small target detection of chili flowers based on YOLOv8n lightweight improvement; Figure 2 It is a network structure diagram of YOLOv8n after lightweight improvement; Figure 3 It is a flowchart of the EMA attention mechanism; Figure 4 It is a structure diagram of the GSConv convolution module; Figure 5 It is a structure diagram of the WIoU loss function; Figure 6 It is a comparison diagram of the actual detection differences between the improved YOLOv8n detection algorithm and several other different object detection algorithms; Figure 7 It is an architecture diagram of the small target detection system of chili flowers based on YOLOv8n lightweight improvement. Detailed implementation manners

[0013] In order to make the technical means, creative features, achieved objectives and effects realized by the present invention easy to understand, the present invention will be further described below with reference to specific illustrations.

[0014] See Figures 1 to 6 The small target detection method for pepper flowers based on the lightweight improvement of YOLOv8n is as follows: Step 1): Load the target detection image data using the corresponding dataset, convert it into the YOLO training format, and then divide it into a training set, a validation set, and a test set; In this embodiment, the target detection image system selects 6419 pepper flower datasets taken by itself and names them VOC6419. Use self-made Python code to convert the target detection images into the YOLO training format, and divide the VOC6419 dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1. The train and value data of VOC6419 are used as the training set, with 6135 and 642 pictures respectively, a total of 5777 pictures for model training. The test data of VOC6419 is used as the test set, with a total of 642 pictures for model testing.

[0015] Step 2): Add the EMA lightweight attention mechanism, introduce the GSConv (GSConvolution, lightweight convolution) module to replace the original C2F convolution module in the backbone, and introduce the WIoU (Weighted Intersection over Union) loss function to replace the original CIoU (Complete Intersection over Union) loss function in the backbone, and then construct a small target detection model for pepper flowers based on the lightweight improvement of YOLOv8n; In this embodiment, constructing a small target detection model for pepper flowers based on the lightweight improvement of YOLOv8n is divided into three steps: The 2.1st step is to introduce three EMA lightweight attention mechanisms in the neck network of the YOLOv8n model. That is, the EMA lightweight attention mechanism includes the first EMA attention mechanism module, the second EMA attention mechanism module, and the third EMA attention mechanism module, and then construct an efficient multi-scale small target detection method to improve the model's ability in detecting small targets of pepper flowers.

[0016] In this embodiment, the EMA lightweight attention mechanism module reshapes part of the channel dimension into the batch dimension and groups the channel dimension into multiple sub-features, thereby distributing spatial semantic features within each feature group. This method combines feature grouping, parallel sub-network structures, and cross-space learning to capture multi-scale information and pixel-level pairwise relationships. The EMA lightweight attention mechanism includes the following three steps: Step 2.1.1: Given an input feature map X with a size of RC×H×W, the EMA lightweight attention mechanism first divides it into G smaller feature segments in the channel dimension, denoted as [X0, Xi, …, XG-1], and the size of each segment is reduced to RC / G×H×W, aiming to obtain rich semantic content from different perspectives. Next, the EMA lightweight attention mechanism uses three independent paths to extract the attention weights and feature description information of each segmented input feature map respectively.

[0017] Step 2.1.2: The first two paths in the EMA lightweight attention mechanism module use 1×1 convolutional layers, while the third path uses a 3×3 convolutional layer. In the 1×1 convolutional path, the channel information is encoded through one-dimensional global average pooling in two different directions, thereby achieving information exchange between channels. Relative to the 1×1 convolutional path, the 3×3 convolutional path omits one-dimensional global average pooling and GroupNorm, so more diverse feature representations can be obtained.

[0018] Step 2.1.3: The output results of the 1×1 and 3×3 paths are integrated and encoded for spatial information through two-dimensional global average pooling, and the two-dimensional global average pooling is calculated according to Equation (1). (1) In Equation (1), Zc represents the output result after pooling for the C-th channel, where H and W refer to the height and width of the input feature map respectively, and C represents the number of channels. Xc(i, j) refers to the pixel value at the i-th row and j-th column in the C-th channel.

[0019] The spatial attention weight values generated by the two parallel branches (such as 1x1 and 3x3 convolutions) are aggregated within the group. The output feature maps of each group are fused by combining the two independently obtained spatial attention weights, and the Sigmoid activation function is applied to capture the pairwise relationships at the pixel level to capture pixel-level pairwise relationships and obtain global context information, and finally an enhanced output feature map is generated.

[0020] The second step is to integrate the GSConv module into the backbone network of the YOLOv8 network architecture to replace the original C2F traditional convolutional unit. The GSConv module adopts depthwise separable convolution (DSC). The GSConv module can perform convolutional operations on each channel separately and aggregate the results. Compared with those traditional convolutional methods that require cross-channel calculations, depthwise separable convolution (DSC) can more efficiently reduce the number of model parameters and computational load. In addition, through its feature aggregation and mixing functions, the GSConv module ensures the complete retention of multi-channel information while expanding the number of channels, thus enhancing the semantic information captured by the model. Therefore, introducing the GSConv module not only maintains the performance of the model but also effectively reduces the increase in computational volume, making the model more lightweight and simplified.

[0021] Compared with the traditional standard convolutional module, the main shortcoming of depthwise separable convolution (DSC) is that it somewhat ignores the inter-channel relationships to some extent, resulting in the dispersion of information. The standard convolution convolves all channels of the input image simultaneously and finally integrates the information of all channels, while depthwise separable convolution (DSC) performs convolutional operations on each channel independently and then recombines the information of these channels. Although this does reduce the computational burden brought by multiple channels, it may also cause the loss of information between channels. To overcome this problem, this embodiment introduces the GSConv module, which uses a shuffle operation to incorporate the information generated by SC (obtained through dense convolutional operations) into the information generated by depthwise separable convolution (DSC). The input channel is C1, and the output channel is C2. The GSConv module increases the number of channels to C2 / 2 through a standard convolution and then applies depthwise separable convolution (DSC) for processing while keeping the number of channels unchanged. Then, the result of the standard convolution is concatenated and shuffled with the output of depthwise separable convolution (DSC). And in the shuffle step, although the channel information is randomly reorganized, the multi-channel information is retained, which enhances the extracted semantic information, strengthens the feature fusion, and improves the expressive ability of image features.

[0022] In this embodiment, the specific implementation steps of the GSConv module are as follows: Step 2.2.1: Divide the C1 channels of the input feature map into two parts to obtain two sets of sub-feature maps. Send these two sets of sub-feature maps to two different processing paths. One set of sub-feature maps will undergo depthwise separable convolution (DSC), while the other set will go through standard convolution operations. These two convolution operations are performed independently within their respective sub-channels. After independent completion, the convolution results of these two sub-channels are merged and then reorganized to form a feature map with dual channels. Finally, by performing a shuffle operation, feature information is exchanged between channels to enhance the information exchange between them, and ultimately a feature map with C2 channels is output. This process effectively improves the information flow and interaction between features. The operation time of the GSConv convolution module is calculated according to Equation (2). (2) In Equation (2), W and H are the width and height of the output feature map respectively; K1*K2 is the size of the convolution kernel; C1 and C2 are the number of channels of the input and output feature maps respectively.

[0023] Step 2.3 is to construct the WIoU loss function based on the CIoU loss function, as follows: Replace the original CIoU loss function with the newly constructed WIoU loss function. By introducing a dynamic non-monotonic focusing mechanism, this mechanism uses "outlier degree" to evaluate the quality of anchor boxes and designs a wise gradient gain allocation strategy accordingly. The gradient gain allocation strategy reduces the competitiveness of high-quality anchor boxes while also reducing the harmful gradients generated by low-quality examples, making the model training process pay more attention to anchor boxes of ordinary quality and effectively improving the performance and robustness of the model for detecting small targets of pepper flowers.

[0024] The WIoU (Weighted Intersection over Union) loss function is an improved loss function for object detection tasks, which is an innovation based on the traditional CIoU (Complete Intersection over Union) loss function to improve the performance and robustness of the model in object detection. The following are the specific steps for implementing the WIoU loss function module: Step 2.3.1: Calculate the outlier degree of the anchor box, which is achieved by comparing the IoU value of the anchor box with the ground truth bounding box; Step 2.3.2: Design a suitable gradient gain allocation strategy to ensure that the gradient can be effectively transmitted to the model parameters; Step 2.3.3: Appropriately adjust the hyperparameters during the training process to obtain the best performance.

[0025] In the application of chili flower small target detection, the WIoU loss function, through its dynamic non-monotonic focusing mechanism, can better handle low-quality examples in the training set, reducing the negative impact of these examples on the model performance. At the same time, by reasonably allocating gradient gains, the WIoU loss function can improve the learning efficiency of the model for normal-quality anchor boxes, thereby enhancing the detection performance and robustness for small targets such as chili flowers.

[0026] Through the above three steps, the YOLOv8n model is optimized to obtain a chili flower small target detection model (Chili Flower) based on the lightweight improvement of YOLOv8n. The network structure includes a backbone network (Backbone), a neck network (Neck), and a detection head (Head) part.

[0027] In this embodiment, the backbone network (Backbone) is composed of the first 10 layers of the network structure diagram. Its core function is to perform feature extraction. While maintaining the original backbone network architecture of YOLOv8, some of its traditional C2F convolution modules are replaced. The backbone network (Backbone) includes a first CBS convolution module, a second CBS convolution module, a first GSConv module, a third CBS convolution module, a second GSConv module, a fourth CBS convolution module, a third GSConv module, a fifth CBS convolution module, a fourth GSConv module, and an SPPF module connected in sequence; The second GSConv module is connected to the first aggregation module (Concat), the third GSConv module is connected to the second aggregation module (Concat), and the SPPF module is respectively connected to the first upsampling module (Upsample) and the fourth aggregation module (Concat); These consecutive modules work together to extract the key features of the image.

[0028] In this embodiment, the neck network (Neck) consists of modules 10 to 24, which are connected in sequence. While maintaining the original backbone network architecture of YOLOv8, an EMA lightweight attention mechanism module is introduced; The described Neck network includes a first Upsample module, a first Concat module, a first C2F module, a second Upsample module, a second Concat module, a first EMA attention mechanism module, a second C2F module, a sixth CBS convolution module, a third Concat module, a second EMA attention mechanism module, a third C2F module, a seventh CBS convolution module, a fourth Concat module, a third EMA attention mechanism module, and a fourth C2F module connected in sequence. The first C2F module is connected to the third Concat module; In the YOLO series of object detection models, the P3, P4, and P5 layers of the feature pyramid dynamically fuse and cooperate through a three-level multi-scale bidirectional path to achieve precise detection. Among them, the P3 layer (shallow features) captures the details of small objects (such as chili flower stamens) based on a high resolution (1 / 8 of the input size). Its C2F module fuses detail and semantic information through cross-stage residual connections, thereby enhancing the recognition ability of small chili flower targets; the P4 layer (mid-level features) balances detail and context at a medium resolution (1 / 16 of the input size), optimizes the computational efficiency through grouped convolution, accurately locates single chili flowers, and suppresses interference from branches and leaves; the P5 layer (deep features) analyzes the global morphology with a low resolution (1 / 32 of the input size) and a large receptive field, and uses dilated convolution to separate dense flowers, enhancing the interference resistance of chili flowers in the environment of branch and leaf occlusion; the three-level multi-scale features are dynamically fused through a bidirectional path, which can effectively handle challenges in complex agricultural scenarios such as strong light overexposure, branch and leaf occlusion, and flower overlapping occlusion in chili detection; The tasks of the second C2F module, the third C2F module, and the fourth C2F module are to extract features from the P3, P4, and P5 layers of the backbone network respectively, and then aggregate the features to integrate the information captured by the backbone network at each stage.

[0029] In this embodiment, the convolutional layers and bottleneck layers in the Neck network are carefully optimized, which not only enhances the feature extraction efficiency of the model but also improves the overall performance; in addition, by reducing the number of parameters and computational load of the model, the ability of the model to recognize small targets is successfully improved, and at the same time, the lightweight of the model is achieved.

[0030] In this embodiment, to maintain the initial settings of the YOLOv8 detection head, no changes are made to it. That is, the second C2F module (16), the third C2F module (20), and the fourth C2F module (24) are directly connected to three decoupled detection heads respectively. This architecture design enables the detection heads to directly utilize the feature information transmitted by the Neck network, thereby more accurately performing the object detection task.

[0031] Step 3): Train the small target detection model of pepper flowers with lightweight improvement based on YOLOvn8 using the training set in Step 1), and then obtain the optimal small target detection model.

[0032] In this embodiment, the parameter configuration used in the training process is shown in Table 1. Table 1 Test environment and parameter configuration Step 4): Input the test set in Step 1) into the optimal small target detection model obtained in Step 3), and then output the small target detection result.

[0033] In this embodiment, the GSConv (GSConvolution, lightweight convolution) module includes a first GSConv module, a second GSConv module, a third GSConv module, and a fourth GSConv module. The first GSConv module, the second GSConv module, the third GSConv module, and the fourth GSConv module all contain a depthwise separable convolution module. The specific process of the depthwise separable convolution module is as follows: Process 1): Perform standard convolution processing on the input feature map, that is, pass the input feature map through a standard convolution operation. This step is usually to change the number of channels to meet the needs of subsequent operations. Process 2): Perform depthwise separable convolution processing on the input feature map, that is, process the input feature map through the depthwise separable convolution module. The depthwise separable convolution method can achieve feature extraction with less computational effort while keeping the number of channels unchanged. Process 3): Concatenate the input feature maps processed by Process 1) and Process 2), that is, concatenate the result of the standard convolution processing with the result of the depthwise separable convolution processing. Concatenating the feature maps extracted by two different convolution methods can obtain a richer feature representation. Process 4): Shuffle the input feature map concatenated in Process 3). The shuffle operation helps to break the independence of the input feature map in the channel dimension, promote information exchange between different channels, and thus enhance the model's ability to capture spatial information. This is one of the key innovations of the GSConv module.

[0034] In this embodiment, the WIoU loss function is an improved loss function for object detection tasks. The WIoU loss function is an innovation based on the traditional CIoU loss function to improve the performance and robustness of the model in object detection.

[0035] See Figure 7The chili flower small target detection system based on the lightweight improvement of YOLOv8n includes a dataset construction module, a model construction module, a model training module, a model improvement module, and a model detection module; The described dataset construction module uses the corresponding database to load the target detection image data, converts it into the YOLO training format, and divides the training set and the test set; The described model construction module introduces the EMA lightweight attention mechanism and the GSConv module to optimize the YOLOv8n target detection algorithm, and constructs a chili flower small target detection model based on the lightweight improvement of YOLOv8n; The described model training module trains the chili flower small target detection model based on the lightweight improvement of YOLOvn8 on the training set, and then obtains the optimal small target detection model; The described model improvement module introduces the GSConv lightweight convolution module and the EMA lightweight attention mechanism, and replaces and optimizes the loss function of the YOLOv8n target detection algorithm, and constructs a chili flower small target detection model based on the lightweight improvement of YOLOvn8; The described model detection module inputs the test set into the optimal small target detection model, and then outputs the small target detection result.

[0036] In this embodiment, the model training module focuses on using the algorithm improved by the lightweight of YOLOv8n, and conducts in-depth optimization and training on a specific training set for the small target detection task, aiming to generate the optimal small target detection model. This process fully demonstrates the technical strategy of improving the model's ability to recognize small targets through a data-driven method; The chili flower small target detection system based on the lightweight improvement of YOLOv8n realizes the seamless connection from model learning to practical application, and effectively supports various application scenarios that require high-precision chili flower small target recognition capabilities.

[0037] In this embodiment, in order to verify the effectiveness of the lightweight GSConv convolution optimization strategy, the efficient multi-scale EMA lightweight attention mechanism, and the WIoU loss function adopted in this study, ablation experiments and comparative experiments are carried out in this embodiment. These experiments aim to evaluate the specific impact of the proposed improvement measures and other existing improvement methods on the performance of the YOLOv8n model. Through these experiments, the specific impact of various improvements on the model performance can be deeply understood.

[0038] In this embodiment, the evaluation criteria are first determined. The performance of the YOLO series models is usually measured by the following core indicators: Precision (P), Recall (R), and mean Average Precision (mAP). In this experiment, mAP@0.5 is selected as the key indicator to measure the performance in this embodiment. mAP@0.5 represents the average value of mAP when the IoU threshold is 0.5. The higher the mAP value, the higher the overall accuracy of the model. When exploring the balance between the lightweight and performance of the model, in addition to paying attention to the accuracy and speed of the model, the parameter scale and computational complexity of the model also need to be considered. Therefore, in this embodiment, two key indicators, FLOPs (Floating Point Operations) and Params (the number of model parameters), which measure the complexity of the model, are introduced. Through these two indicators, the trade-off between the resource consumption and performance of the model can be evaluated more comprehensively.

[0039] The FLOPs and Params indicators are specifically calculated according to equations (3)-(8): (3) (4) (5) (6) (7) (8) Among them, TP means that the model correctly predicts the actual chili flower as a chili flower, FP means that the model incorrectly predicts the actual non-chili flower as a chili flower, FN means that the model incorrectly predicts the actual chili flower as a non-chili flower, H×W is the size of the output feature map, C in is the input channel, K is the kernel size, C out is the output channel, r is the convolution kernel size, a is the input size, and v is the output size.

[0040] In this embodiment, the effectiveness of the proposed improved algorithm is gradually verified through ablation experiments and comparative experiments. Based on the YOLOv8n model, a series of improvement attempts and combinations are made on the EMA lightweight attention mechanism, GSConv lightweight convolution module, and WIoU loss function, and ablation tests are conducted on each improvement. The results of the ablation experiments are shown in Table 4. The specific ablation experiment settings refer to Table 2. Through this method, the specific impact of each different attention mechanism, convolution module, and loss function improvement on the model performance can be clarified, so as to evaluate its impact on the overall algorithm performance.

[0041] In this embodiment, through experimental training of basic models such as YOLOv5s, YOLOv7tiny, YOLOv8s, YOLOv8n, and YOLOv9 on a self - captured and established chili flower dataset, the experimental results are shown in Table 2. The experimental results show that YOLOv8n has excellent detection performance after training and relatively few parameters. Although YOLOv8n does not perform optimally in terms of mean average precision, overall evaluation shows that the YOLOv8n algorithm has the best effect and is more suitable for detecting small targets, achieving a balance between detection accuracy and lightweight. Therefore, it is considered to improve and optimize the YOLOv8n model. Table 2 Performance parameters of each model In this embodiment, according to the data in Table 2, it is decided whether to apply and replace different lightweight convolutions in the backbone network of YOLOv8n, introduce and apply the attention mechanism in the neck network of YOLOv8n, and finally replace and apply different loss function modules for ablation experiments to further test and improve the model performance. The ablation experiment results are shown in Table 3.

[0042] In this embodiment, by introducing three different attention mechanism modules, namely EMA, SE, and Biformer, a series of improvement experiments are carried out on the neck network of the YOLOv8n algorithm. The experimental results show that all these attention mechanisms improve the model's precision (P - value) and mean average precision (mAP - value) to a certain extent. Among them, the Biformer attention mechanism is the most significant in improving the mAP - value. However, this performance improvement is accompanied by an increase in model weights and floating - point operation counts. In the comparison between the EMA and SE modules, although the SE module has a slight advantage in improving mAP, it also leads to a slight increase in model parameters and weights. In contrast, the EMA module not only improves mAP but also helps reduce the model's parameters. Considering lightweight models and ease of deployment, the EMA attention mechanism will be selected and integrated into the YOLOv8n model as the improvement scheme for this embodiment.

[0043] In this embodiment, a series of improvement experiments were conducted on the backbone network of the YOLOv8n algorithm by introducing three different convolutional modules, namely DCNV3, GSConv, and SCConv. By introducing the DCNV3 convolutional module, although this module can significantly improve the mean average precision of the model, it is not beneficial for the lightweight of the model because the reduction of parameters and weights is not obvious. Then, the SCConv module was continued to be tried, which only brought a slight improvement in mAP and was not effective in reducing the number of parameters. Then, when the GSConv module was introduced, although the mAP level decreased slightly, it significantly reduced the number of parameters and weights of the model, making the model more suitable for lightweight chili flower pollination detection equipment. Therefore, the GSConv module became the preferred improvement measure in this embodiment due to its ability to streamline the model while maintaining performance.

[0044] In this embodiment, the loss function is crucial for the training efficiency and final performance of the model. To solve the problem of sample imbalance, enhance generalization ability, improve accuracy, and optimize the adaptability of the model to specific tasks, the effects of three new loss functions, namely EIoU, Shape-IoU, and WIoU, replacing the original CIoU loss function were compared and studied. The experimental results show that introducing the Shape-IoU loss function slightly reduces the computational amount and improves the frame rate, but does not significantly improve the P value and mAP value; the EIoU loss function slightly improves the mAP value, but the effect is limited; in contrast, the WIoU loss function not only significantly improves the mAP, but also maximally improves the detection frame rate, showing its applicability in the small target detection task of chili flowers and effectively enhancing the performance and robustness of the model. Table 3 Ablation Test Results In this embodiment, the EMA module ensures the streamlining and efficiency of the model while maintaining performance improvement. The GSConv module was selected because of its ability to lightweight the model while maintaining detection performance, and the WIoU loss function was adopted because of its significant effect in improving the model performance and detection frame rate. These improvements provide effective technical support for the small target detection of chili flowers; To evaluate the performance of the YOLOv8n-Chili Flower model in chili flower target detection, a series of performance comparison tests were conducted with other models such as YOLOv5s, YOLOv7tiny, YOLOv8s, YOLOv8n, and YOLOv9. The test results are shown in Table 4.

[0045] Table 4 Comparison of Training Results of Six Different Detection Algorithms In this embodiment, according to the test results in Table 4, although the average precision of the YOLOv8n-Chili Flower model is lower than that of YOLOv5s, significant improvements have been made in terms of model size, computational requirements, and detection speed. Compared with YOLOv7tiny, YOLOv8n-Chili Flower demonstrates superiority in all key performance indicators. Compared with YOLOv8s, while reducing the model parameters and computational volume, the recall rate and average precision of YOLOv8n-Chili Flower have increased by 0.3% and 0.2% respectively. Compared with the original YOLOv8n model, the improved YOLOv8n-Chili Flower has increased the recall rate and average precision by 0.9% and 0.6% respectively, while significantly reducing the number of model parameters and computational volume. Although it is slightly lower than YOLOv9 in terms of recall rate, precision, and detection accuracy, considering the large number of parameters and model weights of YOLOv9, which are not conducive to lightweight deployment and practical applications, the YOLOv8n-Chili Flower model is more suitable for practical applications. These improvements make YOLOv8n-Chili Flower an efficient object detection model that not only ensures detection performance but is also suitable for resource-constrained environments.

[0046] In this embodiment, after the YOLOv8n is improved by the method of the present invention, the detection effect of the YOLOv8n-Chili Flower model on small targets has been effectively improved, and the recall rate and average precision in small target detection have increased by 0.9 and 0.6 percentage points respectively. Moreover, the number of model parameters, computational volume, and weights of the model are 2.39M, 7.2 GFLOPs, and 5.0 Mb respectively, which are reduced by 20.6%, 12.2%, and 20.6% respectively compared with the original model.

[0047] In this embodiment, in the comparison of the detection effects of different algorithms, YOLOv5s had false detections in single-object detection scenarios. This may be due to the influence of light changes and climatic conditions in complex environments on image quality, which reduced the contrast between the flowers and the background and increased the risk of false detections. In multi-object, strong light, flower occlusion, and blurred scenarios, all algorithms showed high detection confidence, without missed detections or false detections. In low-light scenarios, YOLOv7tiny had missed detections due to insufficient light and occlusion. In the scenario of foliage occlusion, although there were missed detections, the improved algorithm demonstrated better robustness, effectively improving the mAP and recall rate, while reducing the model parameters and computational amount, proving its adaptability and lightweight advantages in different environments. Generally speaking, in the detection of chili flower images in the real experimental validation set, the improved algorithm reduced false detections and missed detections, improved the detection performance, and the model was more suitable for mobile deployment. The experimental results confirmed that compared with YOLOv8n, YOLOv8n-Chili Flower had a significant improvement in the accuracy of small-object detection and also enhanced the overall performance, achieving a good balance between lightweight and algorithm performance. In addition, the small-object detection system for chili flowers based on the lightweight improvement of YOLOv8n consisted of multiple key modules, and at the same time, the lightweight improvement was carried out on YOLOv8n as the basic model to meet the application requirements of end-to-end deployment.

[0048] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A small target detection method for pepper flowers based on the lightweight improvement of YOLOv8n, characterized in that The specific steps are as follows: Step 1): Load the object detection image data using the corresponding dataset, convert it into the YOLO training format, and then divide it into a training set, a validation set, and a test set; Step 2): Add the EMA lightweight attention mechanism, introduce the GSConv (GSConvolution, lightweight convolution) module to replace the original C2F convolution module in the backbone, and introduce the WIoU (Weighted Intersection over Union) loss function to replace the original CIoU (Complete Intersection over Union) loss function in the backbone, and then construct a pepper flower small object detection model based on the lightweight improvement of YOLOv8n; Step 3): Train the pepper flower small object detection model based on the lightweight improvement of YOLOvn8 using the training set in Step 1), and then obtain the optimal small object detection model; Step 4): Input the test set in Step 1) into the optimal small object detection model obtained in Step 3), and then output the small object detection result.

2. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 1, wherein The EMA lightweight attention mechanism includes a first EMA attention mechanism module, a second EMA attention mechanism module, and a third EMA attention mechanism module; the GSConv (GSConvolution, lightweight convolution) module includes a first GSConv module, a second GSConv module, a third GSConv module, and a fourth GSConv module; the pepper flower small object detection model based on the lightweight improvement of YOLOv8n includes a backbone network, a neck network, and a detection head.

3. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 1, characterized in that, The dataset loading is to convert the object detection image into the YOLO training format using self-made Python code. The dataset is divided into a training set, a validation set, and a test set according to a ratio. The train and value data of the dataset are used as the training set for model training, and the test data of the dataset is used as the test set for model testing; the WIoU loss function is an improved loss function for object detection tasks. The WIoU loss function is innovated on the basis of the traditional CIoU loss function to improve the performance and robustness of the model in object detection.

4. A small target detection method for pepper flowers based on the lightweight improvement of YOLOv8n, characterized in that, The backbone network includes a first CBS convolutional module, a second CBS convolutional module, a first GSConv module, a third CBS convolutional module, a second GSConv module, a fourth CBS convolutional module, a third GSConv module, a fifth CBS convolutional module, a fourth GSConv module, and an SPPF module connected in sequence; the neck network includes a first upsampling module (Upsample), a first aggregation module (Concat), a first C2F module, a second upsampling module (Upsample), a second aggregation module (Concat), a first EMA attention mechanism module, a second C2F module, a sixth CBS convolutional module, a third aggregation module (Concat), a second EMA attention mechanism module, a third C2F module, a seventh CBS convolutional module, a fourth aggregation module (Concat), a third EMA attention mechanism module, and a fourth C2F module connected in sequence.

5. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 4, characterized in that, The second GSConv module is connected to the first aggregation module (Concat), the third GSConv module is connected to the second aggregation module (Concat), the SPPF module is respectively connected to the first upsampling module (Upsample) and the fourth aggregation module (Concat), and the first C2F module is connected to the third aggregation module (Concat); the second C2F module (16), the third C2F module (20), and the fourth C2F module (24) are respectively directly connected to three decoupled detection heads, maintaining the initial settings of the YOLOv8 detection heads, and the detection heads directly utilize the feature information transmitted by the neck network to accurately perform the object detection task.

6. A small target detection method for pepper flowers based on the lightweight improvement of YOLOv8n, characterized in that, The first GSConv module, the second GSConv module, the third GSConv module, and the fourth GSConv module all contain depthwise separable convolution modules, and the specific process of the depthwise separable convolution module working is as follows: Process 1): Perform standard convolution processing on the input feature map, that is, pass the input feature map through a standard convolution operation to change the number of channels to meet the needs of subsequent operations; Process 2): Perform depthwise separable convolution processing on the input feature map, that is, process the input feature map through the depthwise separable convolution module. The depthwise separable convolution method can achieve feature extraction with less computational effort while keeping the number of channels unchanged; Process 3): Concatenate the input feature maps processed by Process 1) and Process 2), that is, concatenate the results of standard convolution processing with the results of depthwise separable convolution processing, and concatenate the feature maps extracted by the two different convolution methods to obtain a richer feature representation; Process 4): Perform a shuffle operation on the input feature map concatenated in Process 3), thereby breaking the independence of the input feature map in the channel dimension, promoting information exchange between different channels, and enhancing the model's ability to capture spatial information.

7. A small target detection method for pepper flowers based on the lightweight improvement of YOLOv8n, characterized in that, Building a pepper flower small object detection model based on the lightweight improvement of YOLOv8 is divided into three steps: Step 2.1 is to introduce three EMA lightweight attention mechanisms into the neck network of the YOLOv8n model, construct an efficient multi-scale small target detection method, and improve the model's ability in detecting small pepper flower targets; Step 2.2 is to integrate the GSConv module into the backbone network of the YOLOv8 network architecture to replace the original C2F traditional convolutional unit. The GSConv module adopts depthwise separable convolution (DSC). The GSConv module can perform convolution operations on each channel separately and aggregate the results, reducing the number of model parameters and computational load. In addition, through its feature aggregation and mixing functions, the GSConv module ensures the complete retention of multi-channel information while expanding the number of channels, thereby enhancing the semantic information captured by the model; Step 2.3 is to construct the WIoU loss function based on the CIoU loss function, use the newly constructed WIoU loss function to replace the original CIoU loss function, evaluate the quality of the anchor boxes by introducing a dynamic non-monotonic focusing mechanism and using its "degree of outlier", and then design a wise gradient gain allocation strategy. The gradient gain allocation strategy reduces the competitiveness of high-quality anchor boxes while reducing the harmful gradients generated by low-quality examples, making the model pay more attention to the anchor boxes of ordinary quality during the training process and improving the performance and robustness of the model in detecting small pepper flower targets.

8. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 7, characterized in that, The GSConv module uses a shuffle operation to integrate the information generated by SC (obtained through dense convolution operations) into the information generated by depthwise separable convolution (DSC). The input channel is C1, and the output channel is C2. The GSConv module increases the number of channels to C2 / 2 through a standard convolution, and then applies depthwise separable convolution (DSC) for processing, keeping the number of channels unchanged. Then, the result of the standard convolution is concatenated and shuffled with the output of depthwise separable convolution (DSC). And in the shuffle step, although the channel information is randomly reorganized, it strengthens the image feature fusion and improves the expression ability of the image features.

9. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 1 or 7, characterized in that The EMA lightweight attention mechanism distributes spatial semantic features within each feature group by reshaping part of the channel dimension into the batch dimension and grouping the channel dimension into multiple sub-features, and combines feature grouping with parallel sub-network structures and cross-space learning, thereby capturing multi-scale information and pixel-level pairwise relationships. The EMA lightweight attention mechanism includes the following three steps: In Step 2.1.1, upon receiving the input feature map X with a size of RC×H×W, the EMA lightweight attention mechanism first divides it into G smaller feature segments in the channel dimension, denoted as [X0, Xi, …, XG-1], and the size of each segment is reduced to RC / G×H×W, obtaining rich semantic content from different perspectives. Next, the EMA lightweight attention mechanism uses three independent paths to extract the attention weights and feature description information of each segmented input feature map respectively; In Step 2.1.2, the first two paths in the EMA lightweight attention mechanism adopt 1×1 convolutional layers, while the third path adopts a 3×3 convolutional layer. In the 1×1 convolutional path, the channel information is encoded through one-dimensional global average pooling in two different directions, thereby realizing information exchange between channels. Compared with the 1×1 convolutional path, the 3×3 convolutional path omits one-dimensional global average pooling and Group Normalization (GroupNorm), thereby obtaining more diverse feature representations; In Step 2.1.3, the spatial information of the output results of the 1×1 and 3×3 paths is integrated and encoded through two-dimensional global average pooling. The two-dimensional global average pooling is calculated according to Equation (1). In Equation (1), Zc represents the output result after pooling for the C-th channel, where H and W respectively refer to the height and width of the input feature map, and C represents the number of channels. Xc(i, j) refers to the pixel value at the i-th row and j-th column in the C-th channel; The spatial attention weight values generated by two parallel branches (such as 1x1 and 3x3 convolutions) are aggregated within the group. The output feature maps of each group are fused by combining two independently obtained spatial attention weights, and the Sigmoid activation function is applied to capture the mutual relationship at the pixel level to capture pixel-level pairwise relationships and obtain global context information, and finally an enhanced output feature map is generated.

10. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 1 or 7, characterized in that, The specific implementation steps of the GSConv module are as follows: In Step 2.2.1, the C1 channels of the input feature map are divided into two groups to obtain two groups of sub-feature maps. The two groups of sub-feature maps are sent to two different processing paths. One group of sub-feature maps will undergo depthwise separable convolution (DSC), while the other group undergoes standard convolution operations. These two convolution operations are performed independently within their respective sub-channels. After independent completion, the convolution results of the two sub-channels are merged and then reorganized to form a feature map with dual channels. Finally, by performing a shuffle operation, the feature information is exchanged between channels to enhance the information exchange between them, and finally a feature map with C2 channels is output. The operation time of the GSConv module is calculated according to Equation (2). In Equation (2), W and H are the width and height of the output feature map respectively; K1*K2 is the size of the convolution kernel; C1 and C2 are the number of channels of the input and output feature maps respectively.

11. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 1 or 7, characterized in that, The WIoU (Weighted Intersection over Union) loss function is an improved loss function for object detection tasks. It is an innovation based on the traditional CIoU (Complete Intersection over Union) loss function, which improves the performance and robustness of the model in object detection. The following are the specific steps for the implementation of the WIoU loss function module: In Step 2.3.1, the outlier degree of the anchor box is calculated by comparing the IoU value of the anchor box and the true bounding box; Step 2.3.2: Design a suitable gradient gain allocation strategy to ensure that the gradient can be effectively transmitted to the model parameters; Step 2.3.3: Appropriately adjust the hyperparameters during the training process to obtain the best performance.

12. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 7, characterized in that, The P3, P4, and P5 layers of the feature pyramid achieve precise detection through dynamic fusion and collaboration of a three-level multi-scale bidirectional path. The P3 layer (shallow features) captures the details of small objects (such as pepper flower stamens) based on high resolution (1 / 8 of the input size). Its C2F module fuses the detail and semantic information through cross-stage residual connections, thereby enhancing the recognition ability of small pepper flower targets; the P4 layer (middle features) balances detail and context at medium resolution (1 / 16 of the input size), optimizes the computational efficiency through grouped convolution, precisely locates single pepper flowers, and suppresses the interference of branches and leaves occlusion; the P5 layer (deep features) analyzes the global morphology with low resolution (1 / 32 of the input size) and large receptive field, uses dilated convolution to separate dense flowers, and improves the anti-interference ability of pepper flowers in the environment of branches and leaves occlusion; the tasks of the second C2F module, the third C2F module, and the fourth C2F module are to extract features from the P3, P4, and P5 layers of the backbone network respectively, and then aggregate the features to integrate the information captured by the backbone network at each stage.

13. A small target detection method for pepper flowers based on the lightweight improvement of YOLOv8n, characterized in that, Ablation experiments and comparative experiments are used to evaluate the effectiveness of the small pepper flower target detection method based on the lightweight improvement of YOLOv8n. Specifically, two key indicators, FLOPs (Floating Point Operations, i.e., the number of floating-point operations) and Params (the number of model parameters), which measure the model complexity, are introduced. The FLOPs and Params indicators are calculated according to equations (3)-(8) specifically: Among them, TP means that the model correctly predicts the actual chili flower as a chili flower, FP means that the model incorrectly predicts the actual non-chili flower as a chili flower, FN means that the model incorrectly predicts the actual chili flower as a non-chili flower, H×W is the size of the output feature map, C in is the input channel, K is the kernel size, C out is the output channel, r is the convolution kernel size, a is the input size, and v is the output size.

14. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 13, wherein, The ablation experiments and comparative experiments are based on the YOLOv8n model. A series of improvement attempts and combinations are made on the EMA lightweight attention mechanism, the GSConv lightweight convolution module, and the WIoU loss function, and ablation tests are conducted on each improvement. Through this method, the specific impact of each improvement of different attention mechanisms, convolution modules, and loss functions on the model performance can be clarified, so as to evaluate its impact on the overall algorithm performance.

15. The pepper flower small target detection method based on the lightweight improvement of YOLOv8n according to claim 13, characterized in that, Through the experimental training of basic models such as YOLOv5s, YOLOv7tiny, YOLOv8s, YOLOv8n, and YOLOv9, it is determined that YOLOv8n meets the comprehensive requirements of accuracy and lightweight balance for small target detection; through the introduction of three different attention mechanism modules, namely EMA, SE, and Biformer, a series of improvement experiments are carried out on the neck network of the YOLOv8n algorithm, and it is determined that EMA is the optimal choice for lightweight attention mechanism; through the introduction of three different convolution modules, namely DCNV3, GSConv, and SCConv, a series of improvement experiments are carried out on the backbone network of the YOLOv8n algorithm, and it is determined that the GSConv module is the preferred measure for improvement; through the introduction of three new loss functions, namely EIoU, Shape-IoU, and WIoU, to replace the original CIoU loss function for verification, it is determined that the WIoU loss function is the optimal choice.

16. The small target detection system for pepper flowers based on the lightweight improvement of YOLOv8n includes a dataset construction module, a model construction module, a model training module, a model improvement module, and a model detection module, and is characterized in that, The described dataset construction module uses the corresponding database to load the target detection image data, converts it into the YOLO training format, and divides it into a training set and a test set; the model construction module introduces the EMA lightweight attention mechanism and the GSConv module to optimize the YOLOv8n target detection algorithm, and constructs a small target detection model for pepper flowers based on the lightweight improvement of YOLOv8n; the model training module trains the small target detection model for pepper flowers based on the lightweight improvement of YOLOvn8 on the training set, and then obtains the optimal small target detection model; the model improvement module introduces the GSConv lightweight convolution module and the EMA lightweight attention mechanism, and replaces and optimizes the loss function of the YOLOv8n target detection algorithm to construct a small target detection model for pepper flowers based on the lightweight improvement of YOLOvn8; the model detection module inputs the test set into the optimal small target detection model, and then outputs the small target detection result; The described model training module focuses on using the algorithm with lightweight improvement of YOLOv8n, aiming at the small target detection task, and conducts in-depth optimization and training on a specific training set, and then generates the optimal small target detection model, which is a technical strategy to improve the model's ability to recognize tiny targets through a data-driven method.

Citation Information

Cited By

  • Motion injury image classification method and system based on deep learning

    CN120495793A

  • Self-supervised kiwi fruit flower cluster depth estimation method and device based on monocular vision

    CN120525941A

  • A Self-Supervised Method and Device for Estimating Kiwi Blossom Cluster Depth Based on Monocular Vision

    CN120525941B

  • Instrument reading method and system, computer readable storage medium and electronic equipment

    CN120544175A

  • Lightweight silkworm cocoon classification detection algorithm based on improved YOLOv8 network

    CN120563950A